Audio device and method of operation thereof

JP2024540011A5Pending Publication Date: 2025-10-27KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024525040
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-10-26
Filing Date
2022-10-19
Publication Date
2025-10-27

AI Technical Summary

Technical Problem

Existing methods for rendering reverberant audio in virtual reality applications are inflexible, computationally intensive, and often compromise audio quality when adapting to changes in position or orientation, lacking the necessary flexibility and efficiency for dynamic environments.

Method used

An audio device that modifies reverberation parameters such as delay and decay rate based on metadata, using a compensator to adjust parameter values and generate a diffuse reverberant signal efficiently, reducing computational load and improving adaptability.

Benefits of technology

The approach provides a more natural-sounding reverberant audio experience with reduced complexity and computational resources, allowing for flexible and efficient rendering of reverberation across varying positions and orientations in virtual reality environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The audio device comprises a receiver 501 for receiving audio data and metadata, the audio data including data of reverberation parameters of an environment. A modifier 503 generates a modified first parameter value of a first reverberation parameter, the first parameter being a reverberation delay parameter or a reverberation decay rate parameter. A compensator 505 generates a modified second parameter value of a second reverberation parameter in response to the modification of the first reverberation parameter. The second reverberation parameter is indicative of the energy of reverberation in the acoustic environment. A renderer 400 generates an audio output signal by rendering the audio data using the metadata, in particular a reverberation renderer 407 generates at least one reverberation signal component of the at least one audio output signal from at least one of the audio signals and in response to the first modified parameter value and the second modified parameter value. The compensation provides an improvement of the perceived reverberation while allowing flexible adaptation.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an apparatus and method for generating an audio output signal, and in particular, but not exclusively, to an apparatus and method for generating an audio output signal that includes a diffuse reverberant signal component that emulates the reverberant characteristics of an environment, for example as part of a virtual reality experience. [Background technology]

[0002] In recent years, the variety and range of experiences based on audiovisual content has increased significantly, and new services and techniques for using and consuming such content are continually being developed and introduced. In particular, many spatial and interactive services, applications and experiences have been developed to give users more engaging and immersive experiences.

[0003] Examples of such applications include Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR) applications, which are rapidly becoming mainstream with many solutions aimed at the consumer market and many standards being developed by various standards bodies. These standards activities are actively developing standards for various aspects of VR / AR / MR systems, including, for example, streaming, broadcasting, rendering, etc.

[0004] VR applications tend to provide a user experience that corresponds to the user being in a different world / environment / scene, whereas AR (including mixed reality MR) applications tend to provide a user experience that corresponds to the user being in a current environment but with additional information or virtual objects or information added. Thus, VR applications tend to provide fully immersive synthetically generated worlds / scenes, whereas AR applications tend to provide partially synthetic worlds / scenes overlaid on a real scene in which the user is physically present. However, these terms are often used interchangeably and overlap in many parts. In the following, the term virtual reality / VR will be used to refer to both virtual reality and augmented / mixed reality.

[0005] As an example, an increasingly popular service is the presentation of images and audio in a manner that allows the user to actively and dynamically interact with the system to change the parameters of the rendering, which can adapt to movements and changes in the user's position and orientation. A very attractive feature in many applications is the ability to change the viewer's effective viewing position and direction, e.g., to allow the viewer to move and "look around" within the scene being presented.

[0006] Such features in particular make it possible to provide a virtual reality experience to a user, where the user can move around (relatively) freely in a virtual environment and dynamically change his position and where he is looking. Typically, such virtual reality applications are based on a three-dimensional model of the scene, which is dynamically evaluated to provide a specific requested view. This approach is well known, for example, from gaming applications, such as the first-person shooter category, for computers and consoles.

[0007] Also, particularly in virtual reality applications, it is desirable that the images being presented are three-dimensional images, typically presented using a stereoscopic display. Indeed, to optimize the viewer's immersion, it is typically preferred that the user experience the presented scene as a three-dimensional scene. Indeed, it is desirable for a virtual reality experience to allow the user to choose his or her position, viewpoint, and moment in time relative to the virtual world.

[0008] In addition to the visual rendering, most VR / AR applications also provide a corresponding audio experience. In many applications, it is preferable to provide a spatial audio experience where the audio sources are perceived to arrive from positions corresponding to the positions of corresponding objects in the visual scene. Thus, it is preferable that the audio scene and the video scene are perceived as being consistent and both provide a complete spatial experience.

[0009] For example, many immersive experiences are provided by virtual audio scenes that are generated by headphone playback using binaural audio rendering techniques. In many scenarios, such headphone playback is based on head tracking, such that the rendering can be made responsive to the user's head movements, which greatly enhances the sense of immersion.

[0010] An important feature for many applications is how to generate and / or deliver audio that can provide a natural and realistic perception of the audio environment. For example, when generating audio for virtual reality applications, it is important not only to generate the desired audio sources, but also to modify these audio sources to provide a realistic perception of the audio environment, including attenuation, reflections, coloration, etc.

[0011] In room acoustics, or more generally environmental acoustics, sound waves reflect off the walls, floors, ceilings, objects, etc. of the environment, resulting in variations in delay and attenuation (usually frequency dependent) of the source signal, which reach the listener (i.e., the user of a VR / AR system) via different paths. The combined effect can be modeled by an impulse response, which may hereafter be referred to as Room Impulse Response (RIR) (although this term suggests a specific application of an acoustic environment in the form of a room, it tends to be used more generally with regard to acoustic environments, whether or not they correspond to a room).

[0012] A room impulse response typically consists of a direct sound, which depends on the distance from the sound source to the listener, followed by a reverberant part that characterizes the acoustic properties of the room, as shown in Figure 1. The size and shape of the room, the position of the sound source and the listener within the room, and the reflective properties of the room's surfaces all affect the characteristics of this reverberant part.

[0013] The reverberant part can be decomposed into two time domains that usually overlap. The first domain contains the so-called early reflections, which represent isolated reflections of the sound source that occur off walls or obstacles in the room before reaching the listener. As the time lag / (propagation) delay increases, the number of reflections present within a certain time interval increases, and the path contains second order or higher reflections (for example, when reflections are off multiple walls, or both a wall and the ceiling).

[0014] The second region in the reverberant section is where the density of these reflections increases to the point where they can no longer be separated by the human brain. This region is usually called the diffuse reverberation, late reverberation, or reverberation tail.

[0015] The reverberant portion contains cues that give the auditory system information about the distance of the sound source, the size and acoustic properties of the room. The relationship of the energy in the reverberant portion to the energy in the anechoic portion strongly determines the perceived distance of the sound source. The level and delay of the earliest reflections provide clues about how close the sound source is to a wall, and anthropometric filtering enhances the assessment of a particular wall, floor or ceiling.

[0016] The density of (early) reflections contributes to the perceived size of a room. The time it takes for the energy level of the reflections to drop by 60 dB is called the reverberation time, T 60 Reverberation time is often used as a measure of how quickly reflections in a room dissipate. Reverberation time provides information about the acoustic quality of a room, specifically whether the walls are highly reflective (e.g. a bathroom) or highly absorptive (e.g. a bedroom with furniture, carpets, and curtains).

[0017] Furthermore, the RIR depends on the user's anthropometric characteristics since it is filtered by the head, ears and shoulders, i.e., it is the head related impulse response (HRIR) and is therefore part of the binaural room impulse response (BRIR).

[0018] Since late reverberant reflections cannot be distinguished and separated by the listener, they are often simulated and represented parametrically, for example using parametric reverberators that use feedback delay networks, such as the well-known Jot reverberator.

[0019] For early reflections, the delay, which depends on the incidence direction and distance, is an important cue for humans to extract information about the relative position of the room and the sound source. Therefore, the simulation of early reflections must be clearer than late reverberation. Therefore, in efficient acoustic rendering algorithms, early reflections are simulated differently from late reverberation. A well-known method for early reflections is to mirror the sound source at each boundary of the room to generate virtual sound sources that represent the reflections.

[0020] While early reflections involve the position of the user and / or sound source relative to the room boundaries (walls, ceiling, floor), late reverberation tends to be homogenous throughout the room, as the acoustic response of the room is diffuse, which makes simulating late reverberation often more computationally efficient than early reflections.

[0021] The two main characteristics of a room-defined late reverberation are parameters describing the slope and amplitude of the impulse response versus time above a given level, both of which tend to be strongly frequency dependent in natural rooms.

[0022] Examples of parameters traditionally used to indicate the slope and amplitude of the impulse response corresponding to diffuse reverberation include the well-known T 60 These include the amplitude value and the reverberation level / energy. Recently, other indices of amplitude level have been proposed, specifically parameters that indicate the ratio of diffuse reverberation energy to total emitted source energy.

[0023] Such known approaches tend to provide an efficient description of the reverberation that allows the rendering side to accurately reproduce the reverberant characteristics of the environment. However, while these approaches tend to be advantageous when attempting to accurately render the reverberant characteristics of an environment, they tend to be suboptimal in some scenarios and in particular tend to be relatively inflexible. Typically, it tends to be difficult to adapt and modify the processing and / or the resulting reverberant components, in particular without degrading the (perceived) audio quality and / or requiring more computational resources than is recommended. Summary of the Invention [Problem to be solved by the invention]

[0024] Therefore, improved approaches for rendering reverberant audio of an environment would be advantageous, particularly approaches that allow for improved operation, more flexibility, less complexity, easier implementation, improved audio experience, improved audio quality, reduced computational load, improved suitability for various positions, improved performance for virtual / mixed / augmented reality applications, improved perceptual cues for diffuse reverberation, more and / or easier adaptability, more processing flexibility, improved rendering side customization, and / or improved performance and / or operation.

[0025] SUMMARY OF THE DISCLOSURE Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination. [Means for solving the problem]

[0026] According to an aspect of the present invention, there is provided an audio device comprising: a receiver configured to receive audio data and metadata of the audio data, the audio data comprising data of a plurality of audio signals representative of audio sources in an environment, the metadata comprising data of reverberation parameters of the environment; a modifier configured to generate a modified first parameter value by modifying an initial first parameter value of a first reverberation parameter, the first reverberation parameter being a parameter from the group consisting of a reverberation delay parameter and a reverberation decay rate parameter; and a modifier configured to generate a modified first parameter value of a second reverberation parameter in response to the modification of the first reverberation parameter. the compensator configured to generate modified second parameter values ​​by modifying initial second parameter values, the second reverberation parameters being included in the metadata and indicative of energy of reverberation in the acoustic environment; and a renderer configured to generate an audio output signal by rendering the audio data using the metadata, the renderer comprising a reverberation renderer configured to generate at least one reverberation signal component of the at least one audio output signal from at least one of the audio signals and in response to the first modified parameter value and the second modified parameter value.

[0027] The present invention provides improved and / or facilitated rendering of audio including reverberant components. In many embodiments and scenarios, the present invention produces a more natural sounding (diffuse) reverberant signal, providing an improved perception of an acoustic environment. Renderings of audio output signals and reverberant signal components are often produced with reduced complexity and reduced computational resource requirements.

[0028] This approach provides improved, increased, and / or facilitated flexibility and / or adaptation of the processed and / or rendered audio. Such adaptation is in many applications and embodiments substantially facilitated by adaptation performed by modifying parameter values. In particular, in many cases the algorithms, processes, and / or rendering operations are not changed and the required adaptation is achieved by merely modifying the parameter values. Adaptation or modification of the reverberation output and / or processing is further facilitated by modifying a second reverberation parameter (indicative of the energy of reverberation in the acoustic environment) based on how the reverberation delay parameter and / or the reverberation decay rate parameter are changed.

[0029] Modifying the reverberation delay parameter and / or the reverberation decay rate parameter provides a particularly efficient and advantageous operation and adaptation of the reverberation, and the second reverberation parameter is automatically compensated for this modification, thereby automatically reducing or eliminating unintended effects of the modification of the reverberation delay parameter and / or the reverberation decay rate parameter, e.g. reducing the perceptual effects of the adaptation and / or providing, e.g., a more consistent and / or harmonious audio signal output.

[0030] This approach allows for an effective representation of diffuse reverberant sound in an acoustic environment with a relatively small number of parameters.

[0031] This approach allows in many embodiments to generate diffuse reverberation signals that are independent of the position of the sound source and / or listener, allowing efficient generation of diffuse reverberation signals for dynamic applications where positions change, such as many virtual reality and augmented reality applications.

[0032] The audio device may be implemented in a single device or functional unit, or may be distributed across different devices or functions, for example the audio device may be implemented as part of a decoder functional unit, or may be distributed such that some functional elements are performed on the decoder side and others on the encoder side.

[0033] The compensator is configured to generate a modified second parameter value in response to a difference between the modified first parameter value and the initial first parameter value.

[0034] In many embodiments, the renderer comprises a further renderer for rendering a direct path component and / or an early reflection component of the audio signal, the renderer being configured to generate an output signal in response to a combination of the direct path component, the early reflection component and at least one reverberation signal.

[0035] The reverberation renderer is a diffuse reverberation renderer. The reverberation renderer is a parametric reverberation renderer such as a Feedback Delay Network (FDN) reverberator, specifically the Jot reverberator.

[0036] The metadata relates to the audio signal / audio source and / or the environment.

[0037] According to an optional feature of the invention, the compensator comprises a model of diffuse reverberation, the model being dependent on the first reverberation parameter and the second reverberation parameter, and the compensator is configured to determine a modified second parameter value in dependence on the model.

[0038] This approach provides a particularly efficient operation for generating a diffuse reverberant signal that reflects frequency dependence.

[0039] A model is a mathematical function / equation or set of functions / equations.

[0040] According to an optional feature of the invention, the first reverberation parameter is a reverberation decay rate.

[0041] The present invention provides performance and / or operational improvements that facilitate and / or improve adaptation and flexibility and allow enhanced control of the rendered reverberation. The reverberation decay rate parameter provides a particularly efficient adaptation, in particular allowing practical adaptation of the perceived reverberation characteristics in the environment.

[0042] The reverberation decay rate parameter is, for example, T 60 (or more generally T xx is any suitable integer) parameter.

[0043] According to an optional feature of the invention, the compensator is configured to modify the second parameter value to reduce a change in an amplitude measure of the reverberation decay rate resulting from modification of the first reverberation parameter.

[0044] This allows for particularly advantageous adaptation and compensation that is highly efficient yet typically of low complexity.

[0045] The amplitude criterion is a function of the reverberation decay rate and a second parameter.

[0046] According to an optional feature of the invention, the compensator is configured to modify the second parameter value such that an amplitude measure of the reverberation decay rate remains substantially unchanged upon modification of the first reverberation parameter.

[0047] This allows for particularly advantageous operation and / or performance.

[0048] According to an optional feature of the invention, the first reverberation parameter is a reverberation delay parameter indicative of a propagation time delay of reverberation in the environment.

[0049] The present invention provides improved performance and / or operation, which facilitates and / or improves adaptation and flexibility and allows enhanced control of the rendered reverberation. The reverberation delay parameter provides a particularly efficient adaptation, in particular allowing practical adaptation of the perceived reverberation characteristics in the environment.

[0050] The reverberation delay parameter is specifically a pre-delay parameter.

[0051] The propagation time delay indicates the time offset from a reference event in the propagation of waves in a room. Usually the reference event is the emission of sound energy at an audio source, but in some cases / embodiments it is the direct path response. More specifically, it indicates the lag of the room impulse response. In many embodiments it indicates the offset time at which a second reverberation parameter is calculated, which is indicative of the reverberant energy in the acoustic environment. This value is selected by analyzing the room impulse response represented by the reverberation parameter. For example, the propagation time delay indicates the delay between the emission at the sound source and the start of the diffuse late reverberant part of the signal (i.e. the sound after the early reflections), and is specified in seconds, or the lag of the room response to which the signal is diffuse, i.e. the same incidence level from all directions and similar levels across all positions in the room.

[0052] According to an optional feature of the invention, the second reverberation parameter is indicative of the energy of reverberation in the acoustic environment after a propagation time delay indicated by the first reverberation parameter.

[0053] This allows for particularly advantageous operation and / or performance.

[0054] According to an optional feature of the invention, the compensator is configured to determine a modified second parameter value to reduce a difference between the first reverberation energy measure and the second reverberation energy measure, the first reverberation energy measure being the energy of the reverberation after a modified delay, the modified delay being represented by the modified first parameter value and being determined from the reverberation model using the modified delay value and the modified second parameter value, and the second reverberation energy measure being the energy of the reverberation after a modified delay, the modified delay being determined from the reverberation model using the initial delay value and the initial second parameter value.

[0055] This allows for particularly advantageous operation and / or performance, as in many scenarios it makes it possible to reduce the perceptual effects of modifying the reverberation delay parameters on the rendered reverberation.

[0056] According to an optional feature of the invention, the compensator is configured to determine modified second reverberation parameter values ​​such that the first reverberation energy measure and the second reverberation energy measure are substantially the same.

[0057] This allows for particularly advantageous operation and / or performance, in that in many scenarios the perceptual effects of modifying the reverberation delay parameters on the rendered reverberation can be reduced, or even substantially eliminated.

[0058] According to an optional feature of the invention, the compensator is configured to modify the second parameter value so as to reduce the difference in reverberation amplitude as a function of time of delay beyond the delay indicated by the modified first parameter value.

[0059] This allows for particularly advantageous operation and / or performance, as in many scenarios it makes it possible to reduce the perceptual effects of modifying the reverberation delay parameters on the rendered reverberation.

[0060] In many embodiments, the reverberation renderer is configured to generate the at least one reverberation signal component to include only contributions corresponding to propagation delays beyond the propagation delay time indicated by the first modified reverberation parameter.

[0061] In some embodiments, the reverberation renderer is configured to generate the at least one reverberation signal component to include only contributions corresponding to portions of the room impulse response at times beyond a propagation delay time indicated by the first modified reverberation parameter.

[0062] According to an optional feature of the invention, the second parameter represents a level of diffuse reverberant sound relative to the total sound emission in the environment.

[0063] This provides particularly advantageous operation and / or performance.

[0064] In many embodiments, the second parameter represents the energy of the diffuse reverberant sound relative to the total emitted energy in the environment.

[0065] The relationship / ratio of the diffuse reverberation signal to the total signal is also called the diffuse reverberation signal level to total signal level ratio, or the diffuse reverberation level to total level ratio, or the emitted source energy to diffuse reverberation energy ratio (or variations / permutations thereof).

[0066] According to an optional feature of the invention, the second reverberation parameter represents the distance at which the energy of the direct response to sound propagation in the environment is equal to the energy of the reverberation in the environment.

[0067] This provides particularly advantageous operation and / or performance.

[0068] The second reverberation parameter is the critical distance parameter.

[0069] In some embodiments, the second parameter represents the amplitude at a given determined time / lag of the room impulse response to the environment.

[0070] According to an optional feature of the invention, the first reverberation parameter is one of the reverberation parameters of the metadata.

[0071] According to an optional feature of the invention, the renderer is configured to determine a level gain of the at least one reverberation signal component in dependence on a second parameter value.

[0072] This provides an efficient and advantageous generation of reverberant signal components in many scenarios. Level gain is, for example, a gain / scale factor that determines / sets / controls the level of a reverberant signal component.

[0073] This provides particularly advantageous operation and / or performance.

[0074] According to an aspect of the present invention, there is provided a method of operation for an audio device comprising: receiving audio data and metadata of the audio data, the audio data including data for a plurality of audio signals representing audio sources in an environment, the metadata including data for reverberation parameters of the environment; modifying an initial first parameter value of a first reverberation parameter by modifying the first parameter value, the first reverberation parameter being a parameter from the group consisting of a reverberation delay parameter and a reverberation decay rate parameter; generating a modified second parameter value by modifying an initial second parameter value of a second reverberation parameter in response to the modification of the first reverberation parameter, the second reverberation parameter being included in the metadata and indicative of energy of reverberation in the acoustic environment; and generating an audio output signal by rendering the audio data using the metadata, the rendering including generating at least one reverberation signal component of the at least one audio output signal from at least one of the audio signals and in response to the first modified parameter value and the second modified parameter value.

[0075] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter.

[0076] Embodiments of the invention will now be described, by way of example only, with reference to the drawings in which: [Brief description of the drawings]

[0077] [Figure 1] An example of a room impulse response is shown below. [Diagram 2] An example of a room impulse response is shown below. [Diagram 3] 1 illustrates examples of elements of a virtual reality system. [Figure 4] 4 illustrates an example of a renderer for generating audio output according to some embodiments of the present invention. [Diagram 5] 1 illustrates an example of an audio device for generating audio output according to some embodiments of the present invention. [Figure 6] An example of a room impulse response is shown below. [Figure 7] 4 shows an example of the amplitude and stored energy of a room impulse response. [Figure 8] An example of the reverberant part of a room impulse response is shown below. [Figure 9] An example of the reverberant part of a room impulse response is shown below. [Figure 10] An example of the reverberant part of a room impulse response is shown below. [Figure 11] An example of the reverberant part of a room impulse response is shown below. [Figure 12] An example of the reverberant part of a room impulse response is shown below. [Figure 13] 1 shows an example of a parametric reverberator. [Figure 14] An example of a reverberator is shown below. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0078] Although the following description focuses on audio processing and rendering for virtual reality applications, it will be understood that the principles and concepts described can be used in many other applications and embodiments.

[0079] Virtual experiences that allow users to move around in virtual worlds are becoming increasingly popular, and services are being developed to meet such demand.

[0080] In some systems, the VR application is provided locally to the viewer by a standalone device that does not, for example, use or even access any remote VR data or processing. For example, the device, such as a game console, comprises a store for storing scene data, an input for receiving / generating the viewer's pose, and a processor for generating corresponding images from the scene data.

[0081] In other systems, the VR application is implemented and executed remotely from the viewer. For example, a device local to the user detects / receives motion / pose data that is sent to a remote device that processes the data to generate a pose for the viewer. The remote device then generates a view image and corresponding audio signal appropriate for the user's pose based on scene data describing the scene. The view image and corresponding audio signal are sent to a device local to the viewer for presentation there. For example, the remote device directly generates a video stream (usually a stereo / 3D video stream) and a corresponding audio stream that is directly presented by the local device. Thus, in such an example, the local device does not perform any VR processing other than sending motion data and presenting the received video data.

[0082] In many systems, functionality is distributed between local and remote devices. For example, the local device processes the received input data and sensor data to generate a user pose that is continuously transmitted to the remote VR device. The remote VR device then generates a corresponding view image and a corresponding audio signal and transmits these to the local device for presentation. In other systems, the remote VR device does not directly generate the view image and the corresponding audio signal, but selects relevant scene data and transmits this to the local device, which then generates the presented view image and the corresponding audio signal. For example, the remote VR device identifies the nearest capture point, extracts the corresponding scene data (e.g., a set of object sources and their position metadata) and transmits this to the local device. The local device then processes the received scene data to generate an image and an audio signal for a particular current user pose. A user pose usually corresponds to a head pose, and a reference to a user pose is usually considered to correspond to a reference to a head pose as well.

[0083] In many applications, especially for broadcast services, audio sources transmit or stream scene data in the form of images (including video) and audio representations of the scene that are independent of the user pose. For example, signals and metadata corresponding to audio sources within a particular virtual room are transmitted or streamed to multiple clients. Each client then locally synthesizes an audio signal that corresponds to the current user pose. Similarly, audio sources transmit a general description of the audio environment, including a description of the audio sources in the environment and the acoustic properties of the environment. An audio representation is then generated locally and presented to the user, for example using binaural rendering and processing.

[0084] 3 shows such an example of a VR system in which a remote VR client device 301 cooperates with a VR server 303 over a network 305, such as the Internet. The server 303 is configured to support a potentially large number of client devices 301 simultaneously.

[0085] The VR Server 303 supports the broadcast experience by, for example, transmitting an image signal containing an image representation in the form of image data that is used by the client device to locally synthesize a view image corresponding to the appropriate user pose (a pose is referred to as a position and / or orientation). Similarly, the VR Server 303 can transmit an audio representation of the scene to locally synthesize audio for a user pose. In particular, as the user moves around in the virtual environment, the images and audio that are synthesized and presented to the user are updated to reflect the user's current (virtual) position and orientation in the (virtual) environment.

[0086] Thus, in many applications, such as that of FIG. 3, it is desirable to model a scene and generate efficient image and audio representations that can be efficiently included in a data signal that is transmitted or streamed to various devices that can locally synthesize views and audio for poses different from the capture pose.

[0087] In some embodiments, a model representing the scene is, for example, stored locally and used locally to synthesize appropriate images and audio. For example, an audio model of a room includes not only the acoustic properties of the room, but also indications of the properties of the audio sources that can be heard in the room. The model data is then used to synthesize audio appropriate for a particular location.

[0088] How an audio scene is represented and how this representation is used to generate audio are important issues. Audio rendering that aims to provide a natural and realistic effect to the listener typically includes a rendering of the acoustic environment. For many environments this includes a representation and rendering of the diffuse reverberation present in the environment, such as a room. It has been found that the rendering and representation of such diffuse reverberation has a significant effect on the perception of the environment, including whether the audio is perceived as representing a natural and realistic environment. In the following, advantageous approaches for representing audio scenes and rendering audio, in particular diffuse reverberant audio, are described.

[0089] The approach will be described with reference to an audio device comprising a renderer 400 as shown in Fig. 4. The audio device is configured to generate an audio output signal representative of audio in an acoustic environment. In particular, the audio device generates audio representative of the audio perceived by a user moving around in a virtual environment having several audio sources and given acoustic characteristics. Each audio source is represented by an audio signal representative of the sound from the audio source, and metadata describing characteristics of the audio source (such as providing a level indication of the audio signal). Additionally, metadata characterizing the acoustic environment is provided.

[0090] The renderer 400 comprises a pass renderer 401 for each audio source. Each pass renderer 401 is configured to generate a direct path signal component representing a direct path from the audio source to the listener. The direct path signal components are generated based on the positions of the listener and the audio source, in particular by scaling potentially frequency dependent audio signals for distance dependent audio sources and, for example, relative gains for audio sources in a particular direction relative to the user (e.g., non-omnidirectional sources).

[0091] In many embodiments, the renderer 401 also generates a direct path signal based on occlusion or diffraction (virtual) elements that lie between the sound source position and the user position.

[0092] In many embodiments, the path renderer 401 generates further signal components for each path that includes one or more reflections, for example by evaluating reflections off walls, ceilings, etc., as known to those skilled in the art. The direct path components and reflected path components are combined into a single output signal for each path renderer, thus generating a single signal representing the direct path reflections and early / discrete reflections for each audio source.

[0093] In some embodiments, the output audio signal of each audio source is a binaural signal, so that each output signal includes both a left-ear and a right-ear (sub) signal.

[0094] The output signals from the pass renderers 401 are provided to a combiner 403, which combines the signals from the different pass renderers 401 to generate a single combined signal. In many embodiments, a binaural output signal is generated, and the combiner performs a combination, such as a weighted combination, of the individual signals from the pass renderers 401, i.e. all right ear signals from the pass renderers 401 are added together to generate a combined right ear signal and all left ear signals from the pass renderers 401 are added together to generate a combined left ear signal.

[0095] The pass renderers and combiners are typically implemented in any suitable manner, including executable code for processing on a suitable computational resource, such as a microcontroller, microprocessor, digital signal processor, or central processing unit including supporting circuitry such as memory. It will be appreciated that multiple pass renderers may be implemented as parallel functional units, such as, for example, a bank of dedicated processing units, or as repeated operations for each audio source. Typically the same algorithm / code is executed for each audio source / signal.

[0096] In addition to the individual path audio components, the renderer 400 is further configured to generate a signal component representing the diffuse reverberation in the environment. The diffuse reverberation signal is generated in a specific example by combining the source signal with a downmix signal and then applying a reverberation algorithm to the downmix signal to generate the diffuse reverberation signal.

[0097] The audio device of Fig. 4 comprises a downmixer 405 which receives audio signals of multiple sound sources (usually all sound sources in the acoustic environment in which the reverberator is simulating diffuse reverberation) and combines them into a downmix. The downmix thus reflects all sounds generated in the environment. The coefficients / weights of the individual audio signals are for example set to reflect the level of the corresponding sound source.

[0098] The downmix is ​​provided to a reverberation renderer / reverberator 407 configured to generate a diffuse reverberation signal based on the downmix. The reverberator 407 is in particular a parametric reverberator such as the Jot reverberator. The reverberator 407 is coupled to a combiner 403 to which the diffuse reverberation signal is provided. The combiner 403 then combines the diffuse reverberation signal with the path signals representing the individual paths to generate a combined audio signal representing the combined sound in the environment as perceived by the listener.

[0099] The renderer is in this example part of an audio device configured to receive audio data and metadata of the environment and render audio representative of at least a part of the environment based on the received data. Figure 5 illustrates an example of such a device, and an approach for generating an audio output signal, in particular a reverberant signal component, based on the received audio data and metadata will be described with reference to Figures 4 and 5. The audio device of Figure 5 in particular corresponds to or is part of the client device 301 of Figure 3.

[0100] The audio device of Fig. 5 comprises a receiver 501 configured to receive data from one or more audio sources. The audio sources may be any suitable source for providing data, either internal or external. The receiver 501 comprises the functionality required to receive / acquire data, such as wireless functionality, network interface functionality, etc.

[0101] The receiver 501 receives data from any suitable source in any suitable form, including, for example, as part of an audio signal. The data is received from an internal or external source. The receiver 401 is configured to receive the room data, for example, via a network connection, a wireless connection, or any other suitable connection to the internal source. In many embodiments, the receiver receives data from a local source, such as a local memory. In many embodiments, the receiver 501 is configured to retrieve the room data from a local memory, such as, for example, a local RAM or ROM memory. In specific examples, the receiver 501 includes network capabilities for interfacing to a network 305 to receive data from the VR server 303.

[0102] The receiver 501 may be implemented in any suitable manner, including, for example, using discrete or dedicated electronic devices. The receiver 501 may be implemented as an integrated circuit, such as, for example, an application specific integrated circuit (ASIC). In some embodiments, the circuitry may be implemented as, for example, firmware or software running on a programmed processing unit, such as a suitable processor, for example, a central processing unit, digital signal processing unit, or microcontroller. In such embodiments, it will be appreciated that the processing unit may include on-board or external memory, clock driving circuits, interface circuitry, user interface circuitry, and the like. Such circuitry may further be implemented as part of the processing unit, as an integrated circuit, and / or as separate electronic circuitry.

[0103] The received data includes audio data of a plurality of audio signals representative of audio sources in the environment, the audio data specifically including a plurality of audio signals, each of the audio signals representing one audio source (and thus the audio signals describing sound from the audio sources).

[0104] In addition, the receiver 501 receives metadata of the audio source and / or the environment.

[0105] The metadata of each audio signal / source includes (relative) signal level indicators of the audio sources, which indicate the level / energy / amplitude of the audio source represented by the audio signal. The metadata of the audio sources also includes directivity data indicating the directionality of the sound radiation from the audio source. The directivity data of an audio signal describes, for example, the gain pattern, and in particular the relative gain / energy density of the audio source in different directions from the position of the audio source. The metadata also includes other data, such as, for example, an indication of the nominal starting position of the audio source, or of the current (or possibly static) position.

[0106] The receiver 501 further receives metadata indicative of the acoustic environment. In particular, the receiver 501 receives metadata including reverberation parameters describing the reverberant characteristics of the environment. In particular, the metadata includes an indication of a reverberation decay rate parameter and possibly also an indication of a reverberation delay parameter. The metadata further includes a reverberation energy parameter indicative of the energy / level of reverberation.

[0107] The diffuse reverberation characteristics, for example the Room Impulse Response (RIR), are represented by parameters that can be communicated to the renderer via parameter data.

[0108] A parameter that at least partially describes the reverberation of the environment is a reverberation delay parameter. The reverberation delay parameter indicates the delay of the reverberation from the audio source. In particular, the reverberation delay parameter indicates the start time (in the RIR) of the reverberant part of the RIR.

[0109] In many embodiments, the metadata includes an indication of when the diffuse reverberation signal should begin, i.e., it indicates a time delay associated with the diffuse reverberation signal. The time delay indication is specifically in the form of a pre-delay.

[0110] Predelay represents the delay / lag of the RIR and is defined to be the threshold between the early reflections and the diffuse and late reverberation. This threshold usually occurs as part of a smooth transition from (more or less) separate reflections to a fully coherent mixture of higher order reflections, so a suitable threshold is chosen using a suitable evaluation / decision process. This decision can be made automatically based on an analysis of the RIR or calculated based on the room dimensions and / or material properties.

[0111] Alternatively, a fixed threshold can be chosen, e.g. 80 ms into the RIR. The pre-delay can be given in seconds, milliseconds or samples. In the following description, it is assumed that the pre-delay is chosen at a point after the reverberation has actually diffused. However, even if this is not the case, the method described works well.

[0112] The pre-delay therefore indicates the onset of the diffuse reverberation response from the start of the sound source emission. For instance, in an example such as that shown in Figure 6, if the sound source starts emitting at t0 (e.g. t0=0), the direct sound reaches the user at t1>t0, the first reflection reaches the user at t2>t1, and the defined threshold between the early reflections and the diffuse reverberation reaches the user at t3>t2. The pre-delay is then t3-t0. The pre-delay is considered to reflect the propagation delay at the onset of the diffuse reverberation.

[0113] In many embodiments, the reverberation delay parameter is included in the metadata, for example in the form of a pre-delay. In other embodiments, however, it is a pre-defined or fixed parameter. For example, the bitstream complies with a suitable audio standard or specification that defines a standard pre-delay with reference to other reverberation parameters (e.g. decay rate or reverberation energy parameters) being given.

[0114] Another parameter that at least partially describes the reverberation of the environment is the reverberation decay rate parameter. The reverberation decay rate parameter indicates the rate of decrease in the level of the reverberation of the environment, and in particular the rate of decrease in the level of the reverberant portion of the RIR. In particular, the reverberation decay rate parameter indicates the slope of the reverberant portion of the RIR.

[0115] The reverberation decay rate parameter indicates the level variation of reverberation as a function of time / lag / delay, and in particular indicates the level of decay / reduction of reverberation (in particular the reverberant portion of the RIR) as a function of delay / time. In some embodiments, the reverberation decay rate parameter is a parameter indicating the average number of decibels (dB) of reverberation response reduction per unit time (e.g. per second) or may be a parameter in the linear amplitude or energy domain (e.g. 2 -γt ) is the exponential coefficient of the exponential equation describing the level decay in

[0116] The reverberation decay rate parameter varies among different embodiments. In many embodiments, it is, for example, T 60 , T 30 , or T 20 These parameters indicate the time it takes for the reverberant energy to decay by 60 dB (30 and 20 dB, respectively). For example, they are expressed as the time corresponding to a 60 dB drop in the energy decay curve (EDC) and are given by the integral equation:

number

[0117] Another parameter that at least partially describes the reverberation of an environment is a reverberation parameter that indicates the energy of the reverberation in the acoustic environment, in particular the energy of the reverberant part of the RIR. Such a parameter is also called a reverberation energy parameter. The reverberation energy parameter can be given, for example, as the reverberation energy relative to the total sound source energy, as the critical distance, as the reverberation amplitude relative to the total sound source energy.

[0118] In many embodiments, the reverberation of the environment, and in particular the (diffuse) reverberant part of the RIR, is characterized by a combination of reverberation delay, reverberation decay rate, and reverberation energy parameters. A set of such parameters describes the onset of reverberation, the progression of the level of reverberation over time, and the overall level of reverberation. One, several, or all of these parameters are received as part of the metadata.

[0119] The received audio data is rendered with the reverberant portion of the rendered audio controlled by the received reverberation parameters, resulting in an output audio signal having a reverberant component that corresponds to the reverberant component of the environment. However, the audio device of Figure 5 also includes the ability to locally adapt and customize the reverberation. In the audio device of Figure 5, this is achieved by including the ability to modify the reverberation delay and / or decay rate parameters before they are used to control the reverberation rendering by the renderer 400.

[0120] In the audio device of Figure 5, the receiver 501 is coupled to the renderer 400, to which the received audio data is directly fed. However, instead of being fed directly to the renderer 400, the metadata is first fed to a modifier 503 arranged to modify a first reverberation parameter, which may be a reverberation delay parameter or a reverberation decay rate parameter (in some cases both of these parameters are modified).

[0121] Thus, a first reverberation parameter initially has a given parameter value, which is modified by the modifier 503 to a (different) modified parameter value. For example, for a reverberation delay parameter, the initial delay value is modified to a modified delay value, which is typically either a smaller delay or a larger delay (although in some embodiments the modifier 503 is asymmetric and can only increase the delay or can only decrease the delay). Alternatively or additionally, for a reverberation decay rate parameter, the initial decay rate value is modified to a modified decay rate value, which is typically either a smaller decay rate / slope or a larger decay rate / slope (although in some embodiments the modifier 503 is asymmetric and can only increase the decay rate or can only decrease the decay rate).

[0122] The modification of the parameter values ​​is fully automatic and is determined by the device itself, for example depending on the current operating conditions. For example depending on the available computational resources, the amount of RIR processed by the path renderer 401 and the reverberation renderer 407, respectively, is dynamically modified by the modifier 503, which modifies the reverberation delay parameter. In other embodiments and applications, the modification is performed in response to user input, and in fact the user directly controls the modification of the reverberation parameters. For example, if a less reverberant experience is desired by the user, the reverberation decay rate parameter is modified by user input to a parameter value corresponding to a higher decay rate, so that the reverberation dies out faster. It will be understood that many other reasons, approaches and purposes for the modification are possible, and the described approach does not depend on a specific context or approach for modifying the reverberation parameters.

[0123] It has been found that such an approach of modifying the rendering, i.e. adapting and customizing the reverberant rendering, by modifying the second reverberant parameters describing the reverberant part of the RIR, while being very efficient and advantageous, is not optimal in all scenarios and in many scenarios may result in an audio rendering that is not perceived as ideal, e.g., in many scenarios artifacts, degradation in quality, perceptual distortions, and / or imbalance between different parts of the RIR occur.

[0124] Furthermore, it has been found that many drawbacks can be mitigated and even substantially eliminated by introducing a compensation that modifies reverberation parameters indicative of the energy of reverberation in the environment (reverberation energy parameters), in particular the reverberation parameters indicative of the energy / level of the reverberant part of the RIR. The compensation is based on a modification of the reverberation delay and / or decay rate parameters, in particular the difference between the modified parameter value of a first reverberation parameter and the original value of the first parameter. In particular, compensating the reverberation energy parameter from the received metadata results in an improved consistency with the modified reverberation parameters, allowing for example to perceive a more natural sounding reverberation and overall audio experience.

[0125] Thus, the apparatus of FIG. 5 comprises a compensator 505 configured to generate a modified second reverberation parameter value by modifying a reverberation value of the second reverberation parameter in response to modification of the first reverberation parameter, if the second reverberation parameter is provided as part of the metadata and the second reverberation parameter is a reverberation energy parameter indicative of the energy of reverberation in the acoustic environment.

[0126] The compensator 505 is configured to adapt the reverberation energy parameter to reflect, for example, that for a modified reverberation delay parameter, the energy may change if more or less RIR is rendered as diffuse reverberation rather than path reflection. As another example, for a change in the reverberation decay rate parameter, the reverberation energy parameter is modified to normalize the energy for the different decay rates.

[0127] Different parameters are used in different applications in the metadata and bitstream to indicate the energy of the diffuse reverberation. Usually, the energy of the diffuse part of the RIR tends to be indicated by a single parameter. However, in some cases, several parameters are used as alternatives or in combination. The energy measure is frequency dependent.

[0128] The particular reverberation energy parameters that are modified by the compensator will therefore also vary in different embodiments. In the following some particularly advantageous reverberation energy parameters are described.

[0129] Reverberation level / energy usually has a psychoacoustic relevance mainly in relation to the direct sound. The level difference between the two is an indication of the distance between the sound source and the user (or the RIR measurement point). As the distance increases, the direct sound is attenuated more, but the level of late reverberation remains the same (it is the same throughout the room). Similarly, for sound sources with directionality that depends on where the user is relative to the source, as the user moves around the source, the directivity affects the direct response but not the level of reverberation.

[0130] Therefore, reverberation levels are often not advantageously expressed in terms of direct sound, but rather a more general characteristic that is independent of the sound source and the user's position within the room.

[0131] In some embodiments, the reverberation energy parameter is a parameter indicative of the level of diffuse reverberant sound relative to the total emitted sound in the environment. The reverberation energy parameter indicates the ratio of the diffuse reverberant signal to the total signal, i.e. the Diffuse to Source Ratio (DSR) is used to express the amount or level of diffuse reverberant energy of a sound source received by a user as a ratio of the total emitted energy of that sound source. This is expressed in a way that the diffuse reverberant energy is appropriately conditioned to the level calibration of the signal being rendered and the corresponding metadata (e.g. pre-gain).

[0132] This representation ensures that there is a meaningful relationship to the signal levels used in the system, independent of the absolute positions and orientations of the listener and sound sources within the environment, independent of the relative position and orientation of the user to the sound sources and vice versa, and independent of any particular algorithm for rendering reverberation.

[0133] As described below, for such reverberation energy parameters, the exemplary rendering described calculates downmix coefficients that take into account both the directional pattern to impose the correct relative levels between the source signals and the DSR to achieve the correct level on the output of the reverberator 407.

[0134] DSR represents the ratio between the emitted sound source energy and the diffuse reverberation characteristics, specifically the energy or (initial) level of the diffuse reverberation signal.

[0135] The description mainly focuses on the DSR, which shows the diffuse reverberation energy relative to the total energy.

number

[0136] Hereafter, this is referred to as DSR (Diffusion-to-Source Ratio).

[0137] It is understood that the ratio and the inverse ratio provide the same information, i.e. any ratio can be expressed as an inverse ratio. The relationship of the diffuse reverberant signal to the total signal is thus expressed by the fraction of a value reflecting the level of diffuse reverberant sound divided by a value reflecting the total sound emission, or equivalently by the fraction of a value reflecting the total sound emission divided by a value reflecting the level of diffuse reverberant sound. It is also understood that various corrections of the estimates can be introduced, for example non-linear functions (e.g. logarithmic functions) can be applied.

[0138] Such an approach is consistent with current standard proposals. In preparation for the MPEG-I Audio Call for Proposals (CfP), the Encoder Input Format (EIF) was defined (MPEG output document N19211, section 3.9, "MPEG-I 6DoF Audio Encoder Input Format", MPEG 130). The EIF defines the reverberation level by the pre-delay and direct diffusion ratio (DDR), which, although named differently, is defined as the ratio between the emitted source energy and the pre-delay diffuse reverberation energy (DDR=DSR).

[0139] The diffuse reverberation energy is taken to be the energy produced by the room response from the beginning of the diffuse section, e.g. this is the energy of the RIR from the time indicated by the pre-delay to infinity. Note that subsequent room excitations add to the reverberation energy, so this is usually only measured directly by excitation with a Dirac pulse, or derived from the measured RIR.

[0140] Reverberation energy represents the energy at a single point in the diffuse field space, rather than being integrated over the entire space.

[0141] A particularly advantageous alternative to the above is to use the DSR, which indicates the initial amplitude of the diffuse sound relative to the total emitted sound energy in the environment. In particular, the DSR indicates the reverberation amplitude at the time indicated by the pre-delay.

[0142] The amplitude at the pre-delay is the maximum excitation of the room impulse response at the pre-delay or immediately after the pre-delay, for example within 5, 10, 20 or 50 ms after the pre-delay. The reason for choosing the maximum excitation within a certain range is that at the pre-delay time the room impulse response happens to be in the lower part of the response. The general trend is a decaying amplitude, and the maximum excitation at a short interval after the pre-delay is usually also the maximum excitation of the entire diffuse reverberation response.

[0143] Using the DSR to indicate the initial amplitude (e.g., within a 10 ms interval) makes it easier and more reliable to map the DSR to parameters of many reverberation algorithms. Thus, the DSR, in some embodiments,

number

[0144] In some embodiments, the reverberation energy parameter may represent the amplitude at a given time of the room impulse response to the environment. As in the above example, the amplitude may be given as a relative amplitude (e.g. relative to the total emitted energy) and / or the given time may be the start time of initialization of the diffuse reverberation part of the RIR.

[0145] Parameters in the DSR are expressed with respect to the same source signal level reference.

[0146] This is achieved, for example, by measuring (or simulating) the RIR of the room of interest with a microphone within certain known conditions (such as the distance between the sound source and the microphone, and the directivity pattern of the sound source, etc.) The sound source needs to emit a calibrated amount of energy, for example a Dirac impulse with known energy, into the room.

[0147] The calibration coefficients for the electrical and analog to digital conversion of the measuring equipment are measured or derived from specifications. It can also be calculated from the direct path response of the RIR, which can be predicted from the directional pattern of the sound source and the distance between the sound source and the microphone. The direct response has a specific energy in the digital domain and represents the emitted energy multiplied by a directional gain with respect to the direction of the microphone and a distance gain that depends on the microphone surface for a full sphere surface area with a radius equal to the distance between the sound source and the microphone.

[0148] Both elements must use the same digital level reference, for example a full-scale 1 kHz sine corresponds to 100 dBSPL.

[0149] Measuring the diffuse reverberation energy from the RIR and compensating it with a calibration factor gives the appropriate energy in the same region as the known emitted energy. Together with the emitted energy, the appropriate DSR can be calculated.

[0150] The reference distance indicates the distance at which the distance gain applied to the signal is 0 dB, i.e., no gain or attenuation is applied to compensate for the distance. The actual distance gain applied by the path renderer 401 can then be calculated by considering the actual distance against the reference distance.

[0151] The expression of the effect of distance on sound propagation is performed with reference to a given distance: doubling the distance reduces the energy density (energy per surface unit) by 6 dB; halving the distance induces an energy density (energy per surface unit) of 6 dB.

[0152] In order to determine the distance gain at a given distance, i.e., to determine how much the density has decreased or increased, it is necessary to know the distance that corresponds to a given level so that the relative change in the current distance can be determined.

[0153] Neglecting absorption in air and assuming no reflecting or occluding elements, the emitted energy of a sound source is constant on a sphere of any radius centered on the source position. The ratio of the surface corresponding to the actual distance to the reference distance indicates the attenuation of the energy. The linear signal amplitude gain at the rendering distance d can be expressed as b,

number

[0154] As an example, if the reference distance is 1 meter and the rendering distance is 2 meters, this formula results in approximately 6 dB of signal attenuation (or a gain of -6 dB).

[0155] The total radiated energy index represents the total energy emitted by an audio source. Audio sources usually radiate in all directions, but not equally in all directions. The integral of the energy density over a sphere around the source provides the total radiated energy. For a loudspeaker, the radiated energy can often be calculated knowing the voltage applied to the terminals and the loudspeaker coefficients that describe the impedance, energy losses, and transfer of electrical energy to sound pressure waves.

[0156] In some embodiments, the reverberation energy parameter represents the distance at which the energy of the direct response to sound propagation in the environment is equal to the energy of the reverberation in the environment. Such a parameter is, for example, a critical distance parameter.

[0157] Critical distance is considered / defined as the distance from a sound source to the (potential nominal / virtual / theoretical) point (or audio receiver (e.g. microphone)) where the energy of the direct response is equal to the energy of the reverberant response. This distance changes depending on the orientation of the receiver with respect to the sound source, if the directivity changes.

[0158] The energy of the reverberant sound is more or less independent of the position of the source and receiver in the room. Early reflections are still position dependent, but the further into the RIR the less position dependent their level becomes. Due to this property there exists a distance over which the direct sound of a sound source is as loud / has the same level as the reverberant sound of the same source.

[0159] Diffuse reverberation results in a uniform level throughout the room, regardless of the location of the audio source. The level of the direct path response depends strongly on the distance between the microphone / observer / listener location and the sound source. The attenuation of the direct response level of an audio source as a function of the distance to the microphone is very well defined. Therefore, the distance between the audio source and the microphone is often used to indicate the critical distance. This is the distance at which the direct response of the audio source decays to a level equal to the (constant) reverberation level. Critical distance is an acoustic property known to those skilled in the art.

[0160] Thus, in the approach of Fig. 5, the apparatus allows to modify certain reverberation metadata parameters (delay and decay rate) with a compensator and then adjust the associated reverberation energy metadata. The compensation is such that the relationship between the reverberation energy metadata and other metadata parameters remains similar to the original, e.g. according to suitable algorithms, criteria and measures. The modified / compensated reverberation parameters are fed to the renderer, with the rendering of the reverberation signal components being based on the modified reverberation parameter values ​​rather than the original values.

[0161] In many embodiments, the renderer 400 is specifically configured to determine a level gain of at least one reverberant signal component depending on the second parameter value. For example, the pass / signal processing performed by the renderer to generate the reverberant signal component includes a gain / scale factor that sets the energy level of the reverberant signal component. For example, the renderer 400 includes an energy normalization function followed (or preceded) by a variable gain that is applied to the reverberant signal component (or the input audio signal from which it is generated). The variable gain sets the overall level of the reverberant signal component. The renderer 400 is configured to determine the gain of the variable gain from the modified / compensated second parameter value.

[0162] In many embodiments, the compensator 505 comprises a model of the diffuse reverberation, which model is based on the reverberation parameters. The compensator 505 is configured to determine new values ​​based on this reverberation model, in particular modifying the parameters such that an evaluation of the model against the modified parameters provides a desired result, which is usually determined from the initial parameter values. For example, the compensated reverberation energy parameter value is determined such that a parameter or measure that can be determined from the model against the original parameter values ​​is not changed (or is changed in a desired way) for the combination of the modified reverberation decay rate parameter and / or reverberation delay parameter with the compensated reverberation energy parameter. Such a measure is for example the energy / level ratio between the energy of the direct path component of the RIR (or the energy of an initial time interval, such as the time / delay until reverberation begins) and the energy of the reverberant part. As another example, the measure is the initial reference amplitude.

[0163] Reverberation metadata is the decay rate (e.g., T 60 , T 30 , T 20 ) and in bitstreams that include reverberation energy measures (e.g. DSR), the energy measures must be explicitly or implicitly related to a particular choice of reverberation response / RIR. This usually involves starting at a certain lag / delay of the RIR and continuing to a sufficient distance of the RIR where the response amplitude is sufficiently close to the noise floor of the RIR (either due to the resolution of the digital representation or noise introduced by the measurement or the measurement device). Due to the exponential decay nature of reverberation, the main defining point of the reverberation energy is usually the starting lag of the energy measurement, which corresponds to the pre-delay parameter mentioned above.

[0164] The pre-delay value is provided along with other reverberation metadata, but is implied by the definition of the reverberation energy metric used in the application.

[0165] A common mathematical equation can usually be used as a simple model for the diffuse reverberation amplitude envelope. An exponential function usually fits well with the decaying amplitude envelope,

number

number

[0166] When the cumulative energy of such a function is calculated, it asymptotically approaches a final energy value, as shown in FIG.

[0167] Diffuse reverberation is usually very sparse as a function of time (many values ​​are lower than the amplitude index given by the exponential function), so to determine the reverberant energy from the above equation, a compensation is typically included, often simply as a scale factor.

[0168] Indeed, starting from a mathematical model, the energy calculated by the model is usually proportional to the reverberation energy. Therefore, the model is often not suitable for predicting the reverberation energy without (empirically derived) corrections. However, this proportionality can be improved by adjusting the pre-delay or T 60 can be used without correction to calculate the energy adjustment coefficients for the modification of . The reverberation energy can be calculated by the model using integration from the pre-delay to infinity (because the model does not include the noise floor) and can be solved analytically (

number

number

[0169] The model is, for example, used to determine the ratio of the energy predictions of the model before and after the modification, and the reverberation energy parameters are then adapted to reflect this change, for example simply compensated in the same ratio.

[0170] In some embodiments, the modifier 503 is specifically configured to modify a reverberation delay parameter, which indicates the propagation time delay of reverberation in the environment / RIR. Specifically, the modifier 503 is configured to modify the pre-delay. The pre-delay is typically used to indicate the start of the diffuse reverberant part of the RIR. Thus, the pre-delay indicates the time (delay) at which the RIR is dominated by the diffuse reverberant part, i.e., the part that is typically rendered by a diffuse reverberant renderer, such as the Jot reverberator. Thus, the pre-delay is typically used by the renderer to indicate which part of the RIR is rendered by the diffuse reverberant rendering function rather than the pass renderer. In the example of FIG. 4, the pre-delay is used to indicate the instantaneous time of the RIR that is rendered by the reverberator 407 and the pass renderer 401, respectively.

[0171] In some embodiments, the modifier 403 is configured to modify the pre-delay (either a default value or a value indicated by the received metadata) before rendering, thereby modifying the amount of RIR modeled by the diffuse reverberation renderer 407 and the amount rendered by the path renderer 401. As shown in Figures 8 and 9, which show the diffuse reverberation part of the RIR, the pre-delay before modification t pre is the original value tpre The new value t is either earlier (Figure 8) or later (Figure 9) than rend will be corrected to.

[0172] Such modifications may in some embodiments be performed manually, for example, to achieve a desired perceptual effect: for example, a pass renderer tends to provide a more accurate rendering, while the user adjusts the quality of the rendered audio, for example by modifying the pre-delay.

[0173] However, in some embodiments, the modification is automatic. For example, path rendering tends to require much more computational resources than diffuse reverberant rendering using a parametric reverberator. In some embodiments, the modifier is configured to determine the computational load of the device and / or to determine the amount of computational resources available for rendering (many approaches for determining such measures are known to those skilled in the art). The modifier is configured to modify the reverberation delay parameter / pre-delay depending on the available computational resources. In particular, the modifier increases the delay when the amount of available resources increases and decreases the delay when the amount of available resources decreases. For example, the delay (modification) is a monotonically decreasing function of the available computational resources.

[0174] The pre-delay parameters may be changed for reasons other than renderer configuration, such as transcoding metadata into a different format that requires matching implicit pre-delay values, or co-signal HRTFs with certain filter lengths.

[0175] Therefore, renderers including diffuse reverberation rendering will render the diffuse reverberation from a different lag than indicated by the metadata pre-delay (or default / nominal pre-delay). As a result, the required reverberation energy will be different than indicated by the received metadata, resulting in a different reverberation effect / experience than intended by the metadata. Often this difference is large.

[0176] In the described approach, the compensator 505 adjusts the reverberation energy parameters of the metadata such that the adjusted pre-delay represents a perceptually similar reverberation energy metadata corresponding to the rendering delay (or other target delay). The adjustment makes the reverberation energy with the updated pre-delay represent a similar reverberation effect / experience as the original reverberation energy metadata. For example, in Fig. 8 and Fig. 9, the grey areas indicate the reverberation energy to be provided by the diffusion reverberator. This is done by adjusting the pre-delay t pre to infinity. In Figure 8, the energy metadata value is too low to allow the reverberation rendering to start at an earlier lag (dashed triangles). In Figure 9, the energy metadata value is too high to allow the rendering to start at a later lag (dashed triangles).

[0177] In many embodiments, the modifier 505 is configured to modify the reverberation energy parameters such that the energy / amplitude / level of the reverberation during the portion of the RIR that is considered as the reverberant part and is specifically to be rendered by the reverberation renderer after modification of the reverberation delay parameters is similar or even the same when determined using the initial delay and energy indicated by the parameters and when determined using the modified delay and energy.

[0178] Specifically, in many embodiments, the compensator 505 is configured to determine modified reverberation energy parameter values ​​so as to reduce the difference between the first and second reverberation energy measures. Both energy measures are determined for the reverberation starting from the modified delay value, and both energy measures are determined using the same model, in particular the exponential decay reverberation model previously derived. However, the first measure is determined by evaluating the model using modified parameter values ​​of the reverberation delay parameter and the reverberation energy parameter, whereas the second measure is determined by evaluating the model using initial (unmodified / uncompensated) parameter values ​​of the reverberation delay parameter and the reverberation energy parameter. The compensator 505 sets in particular the modified reverberation energy parameter values ​​so that these energies are equal, and thus the energy of the reverberation after the modified delay matches the original value.

[0179] Thus, the first reverberation energy measure is determined as the energy of the reverberation after a modified delay represented by the modified reverberation delay parameter. It is determined from the reverberation model using the modified delay value and the modified reverberation energy parameter. The first reverberation energy measure is indicative of the energy of the reverberation after a modified delay calculated using the modified value.

[0180] The second reverberation energy measure is determined as the energy of the reverberation after a modified delay represented by the modified reverberation delay parameter and is determined using an initial delay value and an initial reverberation energy parameter from the same reverberation model. The second reverberation energy measure indicates the energy of the reverberation after a modified delay calculated using the initial value.

[0181] In many embodiments, the compensator 505 is configured to modify the reverberation energy parameter to reduce (or even eliminate) the difference in reverberation amplitude as a function of time after a modified delay (specifically, a rendering delay indicative of the portion of the RIR that is rendered by the reverberation renderer).

[0182] The reverberation renderer is configured to generate the reverberation signal components, as described above, such that they only include contributions that correspond to propagation delays that exceed the propagation delay time indicated by the correction delay. The reverberation renderer implements in particular the portion of the RIR that follows the correction delay time.

[0183] As a concrete example using the exponential model provided above, if the initial uncorrected pre-delay reverberation energy is proportional to the model energy (G corr ), the energy of the reverberation after the modified pre-delay is assumed to be similarly proportional (i.e. the compensation required to show sparseness is the same).

number

number

[0184] Energy conversion coefficients can be calculated with these formulas to scale the reverberation energy metadata from values ​​corresponding to the initial pre-delay to values ​​corresponding to a modified pre-delay (also called the rendering delay) and still describe the same reverberation characteristics.

number

[0185] From the formula, the conversion factor is n render >n pre When , it is smaller than 1, and n pre >n render We can see that when , it is greater than 1.

[0186] For example, we use DSR to calculate the reverberation rendering configuration. renderBefore using, the DSR parameters are compensated. DSR render =DSR metadata *G conv

[0187] In some embodiments, the corrector is T 60 The reverberation decay rate, such as a value, may be configured to modify the reverberation decay rate, such as a value that may be modified by a user, for example, to modify the perceptual experience of the environment by modifying the perceived amount of reverberation, for example, manually, such as to provide a modified perception, particularly for different artistic effects.

[0188] However, modifying the decay rate also affects the reverberation energy. 60 A shorter time will result in faster decay and therefore less reverberant energy.

[0189] Furthermore, the modified decay rate not only affects the decay rate of the reverberation response after the pre-delay, but also usually affects the decay before the pre-delay and therefore also the initial reverberation response amplitude at the pre-delay lag associated with the reverberation energy index. This is described by Fig. 10, Fig. 11 and Fig. 12, which illustrate the situation where the pre-modified / pre-compensated reverberation energy parameter shows an energy (indicated by grey triangles) that is inconsistent with the desired condition for rendering, i.e. the modified decay parameter. In Fig. 10, the unmodified reverberation energy parameter has a value that is too high to render reverberation with a shorter decay time (dashed triangle). In Fig. 11, the unmodified reverberation energy parameter has a value that is too low to render reverberation with a longer decay time (dashed triangle).

[0190] In the system of Figure 5, the compensator compensates the reverberation energy parameter to present a modified energy level corresponding to the modified reverberation decay rate parameter value. The compensator decreases the presented energy value for an increase in the decay rate and / or increases the presented energy value for a decrease in the decay rate.

[0191] In many embodiments, the compensator 505 adjusts the amplitude measure of the reverberation decay rate resulting from the modification of the first reverberation parameter (A in FIG. 12 ). 00 ), which in particular seeks to keep this reference amplitude substantially unchanged.

[0192] The amplitude criterion is a function of the reverberation decay rate and the reverberation energy parameters, and is taken as the value at t=0 of the RIR that results in the decay rate and energy level of the diffuse reverberant part of the RIR (i.e. the RIR after pre-delay) as indicated by the decay rate and reverberation energy indices.

[0193] This typically results in the reverberation energy parameters being modified to correspond to the modified decay rate similar to how the original reverberation energy metadata corresponds to the original decay rate.

[0194] As a specific example, the corrector 503 is T 60 Modify the values ​​to modify the room characteristics and accordingly modify the reverberation energy parameters in the form of DSR. For example, this is determined based on the model of reverberation presented earlier, which determines how the DSR should be adjusted. Usually, T 60 As changes in , the amplitude at the pre-delay time / onset of the diffuse reverberation also changes, and so does A0, as can be seen in figure 12. As a result, the DSR can be seen to have a double effect: the direct effect of the changed attenuation during reverberation, and the effect of the changed attenuation on the RIR up to the pre-delay and therefore on the amplitude A0 at the start of the reverberant part.

[0195] The change in A0 is determined by the effect of the changed decay rate before the pre-delay. Usually, the early part of the RIR depends heavily on the location of the source and receiver used to measure or model the RIR. This results in early decay, where for example, if the source and receiver are relatively close, a steep decay occurs in the early part of the RIR.

[0196] When it comes to tuning the reverberation parameters in diffuse reverberation modelling, it is often useful to ignore such aspects and assume a RIR with a consistent decay rate over its entire length, which is a good match when the source and receiver are relatively far apart.

[0197] To this end, one approach is to make the reference amplitude of the attenuation line t=t0, as shown in FIG.

number

[0198] Next, the modified A0 value of the modified reverberation delay parameter, A r is T60 r Modified T referring to 60 It can be calculated using:

number

number

[0199] The conversion factor for reverberation energy is

number

number

number

[0200] The transformation gain is applied by multiplication, similar to the modification of the reverberation delay parameter.

[0201] T 60 If is frequency dependent, then the conversion gain is frequency dependent.

[0202] In the above example, compensation of the reverberation energy parameters was achieved simply by determining a linear transformation or compensation factor and applying this to the reverberation energy parameters in the form of DSF parameters.

[0203] A similar approach is used when, for example, the reverberation energy parameter is a critical distance or an amplitude parameter.

[0204] For example, if the reverberation energy parameter is a critical distance parameter, this also implies a certain pre-delay at which the reverberation response energy is calculated. Therefore, the same transformation applies. For example, E pre =E cd E rend =E pre *G conv =E cd *G conv and E cd is the energy of the direct response at the critical distance, E pre is the reverberation energy measured from the predelay associated with the critical distance metadata, and E rend represents the reverberation energy from the rendering delay.

[0205] In instances where the reverberation energy parameter is expressed in amplitude, such as the amplitude of the ratio of the initial reverberation energy to the source energy (or total energy or source amplitude), the square root of the gain is taken, as is well known to those skilled in the art.

number

[0206] If both the reverberation delay and the reverberation decay rate parameters are changed, the compensations are combined, e.g. the transformation gains indicated for the various parameters above are combined, e.g. by simple multiplication.

[0207] Below, specific aspects of various embodiments of the approach of FIGS.

[0208] The renderer 407 specifically generates a downmix of the individual audio sources and generates reverberation by applying this signal to a parametric reverberator, such as the Jot reverberator of FIG. 13, which is set based on the reverberation parameters.

[0209] This approach is based on applying a reverberation process to the downmix signal, as described above and shown in Figure 14. Downmix coefficients are determined, which correspond to the weighting of that audio signal in the downmix. A downmix coefficient is the weight of an audio signal in the weighted combination that produces the downmix signal. A downmix coefficient is thus the relative weight for an audio signal when combining them to produce the downmix signal (which in many embodiments is a mono signal), e.g. the weight of a weighted sum.

[0210] The downmix coefficients are based on the ratio of the received diffuse reverberant signal to the total signal, ie the diffuse-to-source ratio, DSR.

[0211] The coefficients are further determined in response to a determined total emitted energy index, which is indicative of the total energy emitted from the audio source, the DSR being typically common for several, and typically all, audio signals, whereas the total emitted energy index is typically specific for each audio source.

[0212] The total emitted energy index usually indicates the normalized total emitted energy, which is independent of the signal content and is entirely defined by source characteristics such as directivity pattern, reference distance, etc. The same normalization is applied to all audio sources and direct and reflected path components. The total emitted energy index is therefore relative to the total emitted energy index of other audio sources / signals, or individual path components, or full-scale sample values ​​of the audio signal.

[0213] The total emission energy measure when combined with the DSR provides, for each audio source, a downmix coefficient that reflects the relative contribution from that audio source to diffuse reverberant sound. Thus, determining the downmix coefficient as a function of the DSR and the total emission energy measure provides a downmix coefficient that reflects the relative contribution to diffuse sound. Thus, when the downmix coefficients are used to generate a downmix signal, each of the audio sources is appropriately weighted, resulting in a downmix signal that reflects the overall sound generated in an environment where the acoustic environment is accurately modeled.

[0214] In many embodiments, the downmix coefficients as a function of the DSR and the total emitted energy measure combined with scaling according to the reverberator characteristics provide downmix coefficients that reflect the appropriate relative level of diffuse reverberation with respect to the corresponding path signal components.

[0215] The total emitted energy is determined from metadata received for the audio source.

[0216] The received metadata includes a signal reference level for each audio source providing an indication of the level of the audio. The signal reference levels are typically normalized or relative values ​​providing an indication of the signal reference level relative to the other audio sources or normalized reference levels. Thus, the signal reference levels typically do not indicate the absolute sound level of the audio source but rather its level relative to the other audio sources.

[0217] In a specific example, the signal reference level includes an indication in the form of a reference distance providing a distance at which the distance attenuation applied to the audio signal is 0 dB. Thus, when the distance between the audio source and the listener is equal to the reference distance, the received audio signal can be used without any distance-dependent scaling. At distances shorter than the reference distance, the attenuation is small and therefore a gain higher than 0 dB needs to be applied when determining the sound level at the listening position. At distances greater than the reference distance, the attenuation is large and therefore a gain higher than 0 dB needs to be applied when determining the sound level at the listening position. Similarly, for a constant distance between the audio source and the listening position, a higher gain is applied to audio signals associated with a larger reference distance than to audio signals associated with a smaller reference distance. The reference distance provides an indication of the signal reference level of a particular audio source, since audio signals are usually normalized to represent a meaningful reference distance or to exploit the full dynamic range (e.g., both a jet engine and a cricket are represented by audio signals that exploit the full dynamic range of the data words used).

[0218] In this example, the signal reference level is further indicated by a reference gain, called pre-gain, which is provided for each audio source and provides the gain that must be applied to the audio signal when determining the rendered audio level. Thus, pre-gain is used to further indicate level variations between different audio sources.

[0219] The metadata further includes directionality data indicating the directionality of sound radiation from the audio source represented by the audio signal. The directionality data for each audio source indicates the relative gain, relative to a signal reference level, in different directions from the audio source. The directionality data provides, for example, a full function or description of the radiation pattern from the audio source defining the gain in each direction. As another example, a simplified indicator is used, for example a single data value indicating a given pattern. As yet another example, the directionality data provides individual gain values ​​for a range of different directional intervals (e.g. segments of a sphere).

[0220] Thus, the audio levels can be generated by the metadata together with the audio signal. Specifically, the path renderer determines the signal content of the direct path by applying a gain to the audio signal, where the gain is a combination of a pre-gain, a distance gain determined as a function of the distance between the audio source and the listener and a reference distance, and a directional gain in the direction from the audio source to the listener.

[0221] For the generation of the diffuse reverberation signal, the metadata is used to determine a (normalized) total emitted energy measure of the audio source based on the signal reference level and directivity data of the audio source.

[0222] Specifically, the total emitted energy index is generated by integrating the directional gain over all directions (e.g., integrating over the surface of a sphere centered on the position of the audio source) and scaled by the signal reference level, specifically the distance gain and pre-gain.

[0223] The determined total emitted energy measure is then processed in the DSR to generate the downmix coefficients.

[0224] The downmix coefficients are then used to generate a downmix signal, in particular as a combination, in particular a sum, of the audio signals, each audio signal being weighted by the downmix coefficient of the corresponding audio signal.

[0225] The downmix is ​​typically produced as a mono signal and then fed into a reverberator to produce a diffuse reverberant signal.

[0226] It should be noted that the rendering and generation of the individual path signal components by the path renderer 401 is position dependent, e.g. with respect to determining distance gains and directivity gains, and the generation of the diffuse reverberant signal thereafter is position independent of both the sound source and the listener.

[0227] The total emitted energy index can be determined based on the signal reference level and the directivity data, without considering the positions of the sound source and the listener. In particular, the pre-gain and the reference distance of the sound source can be used to determine a directivity-independent signal reference level, for example normalized with respect to a full-scale sample of the audio signal, at a nominal distance from the sound source (the nominal distance is the same for all audio signals / sources). The integration of the directivity gain over all directions can be performed on a normalized sphere, for example, as in the case of a sphere at the reference distance. Thus, the total emitted energy index is independent of the positions of the sound source and the listener (reflecting that in an environment such as a room, the diffuse reverberant sound tends to be uniform). The total emitted energy index is then combined with the DSR to generate the downmix coefficients (in many embodiments, other parameters such as the parameters of the reverberator are also taken into account). Since the DSR is also position-independent, the diffuse reverberant signal is generated without any consideration of the specific positions of the sound source and the listener, as in the downmix and reverberation processes.

[0228] Such an approach provides high performance, natural-sounding audio perception without requiring excessive computational resources, and is particularly suitable for, for example, virtual reality applications where the user (and sound sources) move around in the environment, and thus the relative positions of the listener (and possibly some or all of the audio sources) change dynamically.

[0229] The reverberator determines the total emitted energy measure by considering the directional data of the audio source. It should be noted that when determining the diffuse reverberation signal of a sound source with varying source directivity, it is important to use the total emitted energy and not just the signal level or the signal reference level. Consider for example a source directivity corresponding to a very narrow beam with a directivity coefficient of 1 and coefficients of 0 in all other directions (i.e. the energy is only transmitted in a very narrow beam). In this case, the emitted source energy is very similar to the energy and signal reference level of the audio signal, since it represents the total energy. If another sound source with the same energy and signal reference level, but with an audio signal with omnidirectionality, is considered instead, the emitted energy of this sound source will be much higher than the audio signal energy and signal reference level. Thus, if both sound sources are active at the same time, the signal of the omnidirectional sound source should be much more strongly represented in the diffuse reverberation signal, i.e. in the downmix, than the very directional sound source.

[0230] The emitted energy is determined from integrating the energy density over the surface of a sphere surrounding the audio source. Ignoring the distance gain, i.e., integrating over a surface of a radius where the distance gain is 0 dB (i.e., the radius corresponding to the reference distance), the total emitted energy measure can be determined from the following formula:

number

[0231] p is independent of direction, so it moves out of the integral. Similarly, the signal x is independent of direction (the directional gain reflects that variation).

number

[0232] One particular approach for determining this integral is described in more detail below.

[0233] It is desirable to integrate the directional gain over a sphere.

number

[0234] Using a sphere of radius equal to the reference distance (r) results in 0 dB in distance gain, which means that distance gain / attenuation can be neglected.

[0235] In this example, a sphere is chosen because of its computational advantages, but the same energy can be determined from any closed surface of any shape that surrounds the source position. As long as appropriate distance and directivity gains are used in the integration, the effective surface is considered to be facing the source position (i.e., with the normal vector along the source position).

[0236] The surface integral needs to define a small surface dS. Therefore, defining a sphere with two parameters, azimuth (a) and elevation (e), provides the dimensions to do this. Using a coordinate system for the solution, f(a,e,r)=r*cos(e)*cos(a)*u x +r*cos(e)*cos(a)*u y +r*sin(e)*u z And then, Here, u x , u y , and u zare the unit basis vectors of the coordinate system.

[0237] The small surface dS is the magnitude of the cross product of the partial derivatives of the surface of the sphere with respect to the two parameters multiplied by the derivatives of each parameter. dS=|f a ×f e |It's da de.

[0238] This derivative determines the vector that is tangent to the sphere at the point of interest. f a =-r*cos(e)*sin(a)*u x +r*cos(e)*cos(a)*u y +0*u z f e =-r*sin(e)*cos(a)*u x -r*sin(e)*sin(a)*u y +r*cos(e)*u z

[0239] The cross product of the derivatives is a vector that is perpendicular to both. f a ×f e =(r 2 *cos(e)*cos(a)*cos(e)+0*sin(e)*sin(a))*u x +(-0*sin(e)*cos(a)+r 2 *cos(e)*sin(a)*cos(e))*u y +(r 2 *cos(e)*sin(a)*sin(e)*sin(a)+r 2 *cos(e)*cos(a)*sin(e)*cos(a))*u z =r 2 *cos 2 (e)*cos(a)*u x +r 2 *cos 2 (e)*sin(a)*u y +(r 2 *cos(e)*sin(e)*sin 2 (a)+r 2*cos(e)*sin(e)*cos 2 (a))*u z =r 2 *cos 2 (e)*cos(a)*u x +r 2 *cos 2 (e)*sin(a)*u y +(r 2 *cos(e)*sin(e)*(sin 2 (a)+cos 2 (a)))*u z =r 2 *cos 2 (e)*cos(a)*u x +r 2 *cos 2 (e)*sin(a)*u y +r 2 *cos(e)*sin(e)*u z

[0240] The magnitude of the cross product is the surface area of ​​the parallelogram spanned by the vectors f_a and f_e, i.e., the surface area of ​​a sphere, |f a ×f e |=sqrt((r 2 *cos 2 (e)*cos(a) 2 +(r 2 *cos 2 (e)*sin(a) 2 +(r 2 *cos(e)*sin(e) 2 ) =sqrt(r 4 *cos 4 (e)*cos 2 (a)+r 4 *cos 4 (e)*sin 2 (a)+r 4 *cos 2 (e)*sin 2 (e)) =sqrt(r 4 *cos 4 (e)*(cos 2 (a) + sin2 (a))+r 4 *cos 2 (e)*sin 2 (e)) =sqrt(r 4 *cos 4 (e)+r 4 *cos 2 (e)*sin 2 (e)) =sqrt(r 4 *cos 2 (e)*(cos 2 (e)+sin 2 (e))) =sqrt(r 4 *cos 2 (e)) =abs(r 2 *cos(e))=r 2 *cos(e) where e=[-0.5*pi,0.5*pi].

[0241] the result, dS=r 2 *cos(e)*da*de, where the first two terms define the normalized surface area, which when multiplied by da and de results in the actual surface, based on the size of the segments da and de. The double integral over the surface can be expressed in terms of azimuth and elevation angles. The surface dS is expressed in terms of a and e, as above. The two integrals can be performed over azimuth=0...2*pi (inner product), and elevation=-0.5*pi...0.5*pi (cross product).

number

[0242] In many practical embodiments, the directional pattern is not provided as an integrable function, but for example as a discrete set of sample points. For example, each sampled directional gain is associated with an azimuth angle and an elevation angle. Usually, these samples represent a grid on a sphere. One approach to handle this is to convert the integral into a summation, i.e., a discrete integration is performed. The integration is performed in this example as a summation over the points on the sphere where the directional gains are available. This gives the value of g(a,e), but da and de need to be chosen correctly so that there are no large errors due to overlaps or gaps.

[0243] In other embodiments, the directional pattern is provided as a limited number of non-uniformly spaced points in space, in which case the directional pattern is interpolated and resampled uniformly across the azimuth and elevation range of interest.

[0244] Another solution is to assume that g(a,e) is constant around its defined point and solve the integral analytically locally, for example for small azimuth and elevation ranges, e.g. halfway between adjacent defined points. This uses the integral above, but for different ranges of a and e, where g(a,e) is considered constant.

[0245] Experiments show that even when the directivity resolution is fairly coarse, simple summation produces small errors. Moreover, the errors are independent of radius. Linear spacing in azimuth between 10 points, and 10 linearly spaced points in elevation, results in a relative error of -20 dB.

[0246] The integral expressed above gives a result that scales with the radius of the sphere, and therefore with the reference distance. This dependence on radius is because it does not take into account the inverse effect of "distance gain" between two different radii. Doubling the radius reduces the distance of a given surface area (e.g. 1 cm 2) will be 6 dB lower. It can therefore be said that the integration needs to take into account the distance gain. However, the integration is done over a reference distance, which is defined as the distance over which the distance gain is reflected in the signal. In other words, the signal level indicated by the reference distance is not included as a scaling of the value to be integrated, but is reflected by the surface area over which the integration is performed, which varies with the reference distance (since the integration is performed over a sphere with a radius equal to the reference distance).

[0247] As a result, the integral above reflects the energy scaling factor of the audio signal (including any pre-gain or similar calibration adjustments) since the audio signal represents the correct signal reproduction energy at a fixed surface area of ​​a sphere with a radius equal to the reference distance (without directional gain).

[0248] This means that when the reference distance is larger, the total signal energy scaling factor is also larger without changing the signal, because the corresponding signal represents a sound source that is relatively larger than a sound source with the same signal energy, but at a small reference distance.

[0249] In other words, by performing the integration over the surface of a sphere with a radius equal to the reference distance, the signal level indicator provided by the reference distance is automatically taken into account. The larger the reference distance, the larger the surface area and the larger the total emitted energy indicator. The integration is specifically performed directly at the distance where the distance gain is 1.

[0250] The integral above is normalized to the surface units used and to the units used to express the reference distance r. If the reference distance r is expressed in meters, the result of the integral is m 2 It is provided in units of.

[0251] To relate the estimated emitted energy value to a signal, it must be expressed in surface units corresponding to the signal. The surface area of ​​the human ear may be more appropriate, since the level of the signal represents the level reproduced by the user at a reference distance. At the reference distance, this surface relative to the entire surface of a sphere relates to the portion of the energy of the sound source perceived by the person.

[0252] Therefore, the total emitted energy index, which represents the emitted source energy normalized to the full-scale samples in the audio signal, is

number

[0253] Using the DSR, which characterizes the diffuse acoustic properties of the space, and the calculated emitted source energy derived from the directivity, pre-gain, and reference distance metadata, the corresponding reverberation energy can be calculated.

[0254] The DSR is usually determined at the same reference level used by both its components, which may be the same as the total emission energy index or different. In any case, when such a DSR is combined with the total emission energy index, the resulting reverberation energy is also expressed as an energy normalized to a full-scale sample in the audio signal, when the total emission energy determined by the above integration is used. In other words, all the considered energies are essentially normalized to the same reference level so that they can be directly combined without the need for level adjustment. In particular, the determined total emission energy can be used directly together with the DSR to generate a level index of the diffuse reverberation generated from each sound source, which directly indicates the appropriate level with respect to the diffuse reverberation of the other audio sources and with respect to the individual path signal components.

[0255] As a specific example, the relative signal levels of the diffuse reverberant signal components of different sound sources are obtained directly by multiplying the DSR by the total emitted energy index.

[0256] In the described system, the adaptation of the contributions of the different audio sources to the diffuse reverberant signal is performed at least in part by adapting the downmix coefficients used to generate the downmix signal such that the relative contribution / energy level of the diffuse sound from each audio source reflects the diffuse reverberant energy determined for the sound sources.

[0257] As a specific example, if the DSR indicates an initial amplitude level, the downmix coefficients are determined to be proportional to (or equal to) the DSR multiplied by a total emitted energy index, and if the DSR indicates an energy level, the downmix coefficients are determined to be proportional to (or equal to) the square root of the DSR multiplied by a total emitted energy index.

[0258] As a specific example, for a signal having multiple input signal indices x, the downmix coefficients d xteeth,

number

number

[0259] Or, the downmix coefficient d x teeth, d x =E norm,x *DSR is calculated according to Where:

number

[0260] In many embodiments, the downmix coefficients are determined in part by combining the DSR with a total emission energy measure. Whether the DSR indicates the relationship of the total emission energy to the diffuse reverberation energy or the initial amplitude of the diffuse reverberation response, further adaptation of the downmix coefficients is often necessary to accommodate the particular reverberator algorithm used, which scales the signal so that the output of the reverberation processor reflects the desired energy or initial amplitude. For example, the density of reflections of the reverberation algorithm has a strong influence on the resulting reverberation energy, even if the input level remains the same. As another example, the initial amplitude of the reverberation algorithm is not equal to the amplitude of its excitation. Therefore, algorithm-specific, or algorithm- and configuration-specific, adjustments are required, which can be included in the downmix coefficients and are usually common to all sound sources. In some embodiments, these adjustments are applied to the downmix or included in the reverberator algorithm.

[0261] Once the downmix coefficients have been generated, the downmix signal is generated, for example by direct weighted combination or summation.

[0262] An advantage of the described approach is that it uses a conventional reverberator, for example reverberator 407, implemented by a feedback delay network, for example as implemented in the standard Jot reverberator.

[0263] As illustrated in Fig. 13, the principle of the feedback delay network uses one or more (usually several) feedback loops with different delays. An input signal, in this case the downmix signal, is fed into the loop, where the signal is fed back with an appropriate feedback gain. An output signal is extracted by combining the signals in the loop. Thus, the signal is repeated successively with different delays. By using delays that are relatively prime and having a feedback matrix that mixes the signals between the loops, it is possible to create patterns that resemble reverberation in real space.

[0264] To achieve a stable damped impulse response, the absolute values ​​of the elements in the feedback matrix must be less than 1. In many implementations, additional gains or filters are included in the loop. These filters can control the damping instead of the matrix. The use of filters has the advantage that the damping response differs for different frequencies.

[0265] In some embodiments where the reverberator output is rendered binaurally, the estimated reverberation is filtered by average HRTFs (Head Related Transfer Functions) for the left and right ears, respectively, to produce left and right channel reverberation signals. It can be seen that if HRTFs are available for multiple evenly spaced distances on a sphere around the user, the average HRTFs for the left and right ears are generated using the set of HRTFs with the largest distances. The use of average HRTFs is based on or reflects the consideration that reverberation is isotropic and comes from all directions. Thus, rather than including a pair of HRTFs for a given direction, an average over all HRTFs can be used. The averaging can be performed once for the left ear and once for the right ear, and the resulting filters are used to process the reverberator output for binaural rendering.

[0266] In some cases, the reverberator itself introduces coloration of the input signal, resulting in an output that does not have the desired output diffuse signal energy as described by the DSR. Therefore, the effect of this process is equalized as well. This equalization can be performed based on a filter that is analytically determined as the inverse of the frequency response of the reverberator operation. In some embodiments, the transfer function can be estimated using machine estimation learning techniques such as linear regression, line fitting, etc.

[0267] In some embodiments, the same approach is applied uniformly across the frequency bands. However, in other embodiments, frequency-dependent processing is performed. For example, one or more of the provided metadata parameters are frequency-dependent. In such examples, the device is configured to split the signal into different frequency bands corresponding to the frequency dependency, and the aforementioned processing is performed in each of the frequency bands individually.

[0268] Specifically, in some embodiments, the diffuse reverberant signal to total signal ratio DSR is frequency dependent. For example, different DSR values ​​are provided for ranges of individual frequency bands / bins, or the DSR is provided as a function of frequency. In such embodiments, the device is configured to generate frequency-dependent downmix coefficients that reflect the frequency dependency of the DSR. For example, downmix coefficients for individual frequency bands are generated. Similarly, a frequency-dependent downmix and a diffuse reverberant signal are generated as a result.

[0269] In the case of frequency-dependent DSR, the downmix coefficients are in other embodiments complemented by filters that filter the audio signals as part of the generation of the downmix. As another example, the DSR effect is separated into a frequency-independent (broadband) component that is used to generate frequency-independent downmix coefficients that are used to scale the individual audio signals when generating the downmix signal, and a frequency-dependent component that is applied to the downmix, for example by applying a frequency-dependent filter to the downmix. In some embodiments, such filters are combined with further coloration filters, for example as part of a reverberator algorithm. Figure 7 shows a correlation (u,v) filter and a coloration (h L ,h R ) filter, a feedback delay network dedicated to binaural output known as the Jot reverberator.

[0270] Thus, in some embodiments, the DSR comprises a frequency-dependent component part and a non-frequency-dependent component part, and the downmix coefficients are determined depending on the non-frequency-dependent component part (and independent of the frequency-dependent part). The downmix processing is then adapted based on the frequency-dependent component part, i.e. the reverberator is adapted depending on the frequency-dependent part.

[0271] In some embodiments, the directionality of sound radiation from one or more of the audio sources is frequency dependent; in such a scenario, a frequency-dependent total emitted energy is generated which, when combined with the DSR (which may or may not be frequency dependent), results in frequency-dependent downmix coefficients.

[0272] This is achieved, for example, by performing individual processing on separate frequency bands. In contrast to the processing of frequency-dependent DSR, the frequency dependence of the directivity usually needs to be performed before (or as part of) the generation of the downmix signal. This reflects that a frequency-dependent downmix is ​​usually required to include the frequency-dependent effect of directivity, since it usually varies from sound source to sound source. After integration, the net effect can vary significantly with frequency. That is, the total emission energy measure of a given sound source differs from sound source to sound source and has a substantial frequency dependence. Therefore, since different sound sources usually have different directivity patterns, the total emission energy measures of different sound sources also usually have different frequency dependences.

[0273] A concrete example of a possible approach is given below: By providing a DSR characterizing the diffuse acoustic properties of a space and determining the emitted source energy from the directivity, pre-gain and reference distance metadata, the corresponding desired reverberation energy can be calculated. E norm *DSR It can be determined as:

[0274] If the components for calculating the DSR use the same reference level (e.g., related to the full scale of the signal), the resulting reverberation energy will be calculated as E norm When using , it is also the energy normalized to the full-scale samples in the PCM signal and therefore corresponds to the energy of the diffuse reverberation impulse response (IR) that can be applied to the corresponding input signal in order to provide the correct level of reverberation in the signal representation used.

[0275] These energy values ​​can be used to determine the setting parameters of the reverberation algorithm, the downmix coefficients before the reverberation algorithm, or the downmix filter.

[0276] There are various techniques to generate reverberation. Feedback Delay Network (FDN) based algorithms such as the Jot reverberator are a suitable low-complexity approach. Alternatively, the noise sequence can be shaped to have the appropriate (frequency-dependent) decay and spectral shape. In both examples, a prototype IR (with at least an appropriate T60) can be adjusted so that its (frequency-dependent) level is corrected.

[0277] The reverberator algorithm is adjusted to produce an impulse response with unit energy (or unit initial amplitude of the DSR is related to the initial amplitude), or the reverberator algorithm includes a unique compensation, for example in the coloration filter of the Jot reverberator, or the downmix is ​​modified by a (possibly frequency dependent) adjustment or the downmix coefficients produced by the coefficient processor 507 are modified.

[0278] The compensation is determined by generating an impulse response without any such adjustments, but with all other configurations applied (such as appropriate reverberation time (T60) and reflection density (e.g. delay values ​​in the FDN)) and measuring the energy of that IR.

number

[0279] The compensation is the inverse of that energy. To include it in the downmix coefficients, e.g.

number

[0280] In many other embodiments, the compensation is derived from configuration parameters. For example, if the DSR is related to the initial reverberation amplitude, the first reflection can be derived from its configuration. Correlation filters are, by definition, energy-preserving, and coloration filters can also be designed to be so.

[0281] Assuming there is no net boost or attenuation due to the coloration filter, the reverberator will have an initial amplitude (A0) that depends on, for example, T60 and a minimum delay value, minDelay.

number

[0282] The prediction of the reverberation energy is also done heuristically.

[0283] As a general model of the diffuse reverberation energy, we can consider an exponential function A(t), where

number

[0284] Calculating the cumulative energy of such a function asymptotically approaches a final energy value, which has an almost perfectly linear relationship with T60.

[0285] The coefficients of the linear relationship are the sparseness of the function A (setting every third value to 0 will roughly halve the energy), the initial value A0 (the energy is

number

[0286] When generating multiple parametric reverberations with broadband T60 values ​​ranging from 0.1 to 2 seconds, the energy of the IR is approximately linear with the model. The scaling factor between the actual energy and the average of the exponential equation model is determined by the sparseness of the FDN response. This sparseness has the most impact at the beginning, although it decreases towards the end of the IR. Testing the above with multiple configurations of delay values, we found that there is an approximately linear relationship between the model reduction factor and the minimum difference between the delays configured in the FDN.

[0287] For example, in the particular implementation of the Jot reverberator, this is: SF=7.0208*MinDelayDiff+214.1928 The scaling factor SF is calculated by

[0288] The energy of the model is calculated by integrating from t=0 to infinity. This can be done analytically, and the result is

number

[0289] Combining the above, we get the following prediction for reverberation energy:

number

[0290] It will be appreciated that for clarity, the above description has described embodiments of the invention with reference to different functional circuits, units, and processors. However, it will be apparent that any suitable distribution of functionality between different functional circuits, units, or processors may be used without detracting from the invention. For example, functions shown to be performed by separate processors or controllers may be performed by the same processor or controller. Thus, references to specific functional units or circuits should not be considered as indicative of a strict logical or physical structure or organization, but merely as references to suitable means for providing the described functionality.

[0291] The invention can be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention is optionally implemented at least partly as computer software running on one or more data processors and / or digital signal processors. The elements and components of the embodiments of the invention may be physically, functionally and logically implemented in any suitable way. Indeed, functionality may be implemented in a single unit, in several units or as part of other functional units. Thus, the invention may be implemented in a single unit or may be physically and functionally distributed between different units, circuits and processors.

[0292] Although the present invention has been described in relation to some embodiments, it is not intended to be limited to the specific form described herein. Rather, the scope of the present invention is limited only by the appended claims. In addition, although features may appear to be described in relation to certain embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprises" does not exclude the presence of other elements or steps.

[0293] Moreover, although individually recited, a plurality of means, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. In addition, although individual features are included in different claims, they may be advantageously combined, and inclusion in different claims does not imply that the combination of features is not feasible and / or advantageous. Also, the inclusion of a feature in one category of claims does not imply a limitation to this category, but indicates that the feature is equally applicable to other claim categories, as appropriate. Furthermore, the order of features in the claims does not imply a particular order in which the features must function, and in particular the order of individual steps in a method claim does not imply that the steps must be performed in this order. Rather, the steps may be performed in any suitable order. In addition, a singular reference does not exclude a plurality. Thus, reference to "first", "second", etc. does not exclude a plurality. Reference signs in the claims are provided merely as a clarifying example and should not be construed as limiting the claims in any way.

Claims

1. a receiver for receiving audio data and metadata of said audio data, said audio data comprising data for a plurality of audio signals representing audio sources in an environment, and said metadata comprising data for reverberation parameters of said environment; a modifier for generating a modified first parameter value by modifying an initial first parameter value of a first reverberation parameter, the first reverberation parameter being a reverberation delay parameter; a compensator for generating a modified second parameter value by modifying an initial second parameter value of a second reverberation parameter in response to the modification of the first reverberation parameter, the second reverberation parameter being included in the metadata and indicative of reverberation energy in the acoustic environment; a renderer configured to generate audio output signals by rendering the audio data using the metadata, the renderer comprising a reverberation renderer configured to generate at least one reverberation signal component of at least one audio output signal from at least one of the audio signals and in response to the modified first parameter value and the modified second parameter value.

2. 2. The audio device of claim 1, wherein the compensator comprises a model of diffuse reverberation, the model being dependent on the first reverberation parameter and the second reverberation parameter, and the compensator determines a modified value of the second parameter depending on the model.

3. 3. The audio device of claim 1, wherein the first reverberation parameter is a reverberation delay parameter indicative of a propagation time delay of reverberation in the environment.

4. 4. An audio device according to claim 1, wherein the second reverberation parameter is indicative of the energy of reverberation in the acoustic environment after a propagation time delay indicated by the first reverberation parameter.

5. 5. The audio device of claim 1, wherein the compensator determines modified second parameter values ​​to reduce a difference between a first reverberation energy measure and a second reverberation energy measure, the first reverberation energy measure being the energy of the reverberation after a modified delay, the modified delay being represented by the modified first parameter value, the first reverberation energy measure being determined from a reverberation model using the modified delay value and the modified second parameter value, and the second reverberation energy measure being the energy of the reverberation after a modified delay and being determined from the reverberation model using an initial delay value and the initial second parameter value.

6. 6. The audio device of claim 5, wherein the compensator determines the modified second parameter value such that the first reverberation energy measure and the second reverberation energy measure are substantially the same.

7. 7. An audio device according to claim 1, wherein the compensator modifies the second parameter value to reduce differences in reverberation amplitude as a function of time for delays beyond a delay indicated by the modified first parameter value.

8. 8. An audio device according to claim 1, wherein the second reverberation parameter represents a level of diffuse reverberant sound relative to the total sound emission in the environment.

9. 8. An audio device according to claim 1, wherein the second reverberation parameter represents a distance at which the energy of a direct response to sound propagation in the environment is equal to the energy of reverberation in the environment.

10. 8. Audio device according to claim 1, wherein the first reverberation parameter is one of the reverberation parameters of the metadata.

11. 11. Audio device according to any one of claims 1 to 10, wherein the renderer determines a level gain of the at least one reverberant signal component depending on the second parameter value.

12. 1. A method of operation for an audio device, said method comprising: receiving audio data and metadata of said audio data, said audio data comprising data for a plurality of audio signals representing audio sources in an environment, and said metadata comprising data for reverberation parameters of said environment; modifying a first parameter value by modifying an initial first parameter value of a first reverberation parameter, wherein the first reverberation parameter is a reverberation delay parameter; generating modified second parameter values ​​by modifying initial second parameter values ​​of a second reverberation parameter in response to the modification of the first reverberation parameter, the second reverberation parameter being included in the metadata and indicative of reverberant energy in the acoustic environment; and generating audio output signals by rendering the audio data using the metadata, wherein the rendering includes generating at least one reverberation signal component of at least one audio output signal from at least one of the audio signals and in response to the first modified parameter value and the second modified parameter value.

13. A computer program comprising computer program code means for performing all the steps of the method according to claim 12 when the computer program is run on a computer.