Apparatus and method for generating diffusion reverberation signal

By receiving and processing audio signals and metadata to generate frequency-dependent diffused reverberation signals, this technology solves the problems of unnatural generation and high computational resources in existing technologies, achieving a more natural audio experience and reducing computational complexity, making it suitable for virtual reality applications.

CN121645128APending Publication Date: 2026-03-10KONINKLIJKE PHILIPS NV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2021-06-21
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies suffer from suboptimal or insufficient operation when generating diffuse reverberation signals, especially in virtual reality applications, where they struggle to effectively represent and render the audio environment, resulting in an unnatural audio experience and high computational resource consumption.

Method used

By receiving multiple audio signals and metadata representing sound sources in the environment, a diffuse reverberation signal is generated using directional data and signal level indication. The signal components are combined using a downmixer and a reverberator to generate a frequency-dependent diffuse reverberation signal, reducing computational complexity and resource requirements.

Benefits of technology

It achieves more natural diffusion reverberation signal generation, reduces computing resource requirements, supports virtual reality applications with dynamic positional changes, and improves audio experience quality and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121645128A_ABST
    Figure CN121645128A_ABST
Patent Text Reader

Abstract

An audio device that generates a diffusion reverberation signal includes a receiver that receives a plurality of audio signals representing sound sources and metadata including a relationship of the diffusion reverberation signal to a total signal, the relationship of the diffusion reverberation signal to the total source indicating a level of the diffusion reverberation sound relative to a total emitted sound in an environment. The metadata further includes, for each audio signal, a signal level indication and directivity data indicating directivity of sound radiation from a sound source represented by the audio signal. The circuitry determines a total emission energy indication based on the signal level indication and the directivity data, and determines a downmix coefficient based on the total emission energy and a relationship of the diffusion reverberation signal to the total signal. The downmixer generates a downmix signal by combining signal components of each audio signal, the signal components of each audio signal being generated by applying a downmix coefficient of each audio signal to the audio signal. The reverberator generates a diffused reverberation signal of the environment based on the downmix signal component.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application filed on June 21, 2021, with application number 202180044786.5 and entitled "Apparatus and Method for Generating Diffused Reverberation Signal". Technical Field

[0002] The present invention relates to apparatus and methods for processing audio data, and particularly, but not exclusively, to apparatus and methods for processing to generate diffuse reverberation signals for augmented / mixed / virtual reality applications. Background Technology

[0003] The types and scope of experiences based on audiovisual content have substantially increased in recent years as new services and methods of utilizing and consuming such content have been continuously developed and introduced. In particular, many spatial and interactive services, applications, and experiences are being developed to provide users with more engaging and immersive experiences.

[0004] Examples of such applications include the rapidly mainstreaming Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR) applications, with multiple solutions targeting the consumer market. Multiple standards are also under development by various standardization bodies. These standardization activities actively develop standards for all aspects of VR / AR / MR systems, including, for example, streaming, broadcasting, and rendering.

[0005] VR applications tend to provide a user experience corresponding to different worlds / environments / scenes, while AR (including Mixed Reality) applications tend to provide a user experience corresponding to the current environment, but with added extra information or virtual objects or data. Therefore, VR applications tend to provide a fully immersive, synthetically generated world / scene, while AR applications tend to provide a partially synthetic world / scene overlaid on the real-world scene where the user physically exists. However, the terms are often used interchangeably and have a high degree of overlap. In the following, the term Virtual Reality / VR will be used to refer to both Virtual Reality and Augmented / Mixed Reality.

[0006] As an example, an increasingly popular service delivers images and audio in such a way that users can actively and dynamically interact with the system to change rendering parameters, making this adaptable to changes in the user's position and orientation. A very appealing feature in many applications is the ability to change the viewer's effective viewing position and orientation, such as allowing the viewer to move and "look around" within the presented scene.

[0007] This feature allows virtual reality experiences to be offered to users. It allows users to move (relatively) freely within a virtual environment and dynamically change their position and what they are looking at. Typically, such virtual reality applications are based on a 3D model of the scene, where the model is dynamically evaluated to provide a specific requested view. This method is well-known from applications used, for example, in computer and console games (such as in the first-person shooter genre).

[0008] It is also desirable, particularly for virtual reality applications, that the presented images are three-dimensional, typically displayed using stereoscopic displays. In fact, to optimize the viewer's immersion, presenting the scene as a three-dimensional experience is generally preferred for the user. Ideally, the virtual reality experience should allow the user to choose their own position, viewpoint, and time relative to the virtual world.

[0009] In addition to visual rendering, most VR / AR applications also provide a corresponding audio experience. In many applications, the audio preferably provides a spatial audio experience where the sound source is perceived as arriving from a location corresponding to a corresponding object in the visual scene. Therefore, the audio and video scenes are preferably perceived as consistent and provide a full spatial experience.

[0010] For example, many immersive experiences are provided by virtual audio scenes generated through headphone playback using binaural audio rendering technology. In many scenarios, this headphone playback can be based on head tracking, allowing rendering to respond to the user's head movements, which greatly enhances the sense of immersion.

[0011] A key feature of many applications is how to generate and / or distribute audio that provides a natural and realistic perception of the audio environment. For example, when generating audio for virtual reality applications, it is important not only to generate the desired sound sources, but also to modify these sound sources to provide a realistic perception of the audio environment, including damping, reflections, shading, etc.

[0012] In room acoustics, or more generally, ambient acoustics, the reflection of sound waves from the walls, floors, ceilings, objects, etc., of the environment causes delayed and attenuated (usually frequency-dependent) versions of the sound source signal to reach the listener (i.e., the user of a VR / AR system) via different paths. The combined effect can be modeled by an impulse response, which may be referred to below as the room impulse response (RIR) (although the term is suggested for the specific use of acoustic environments in the form of a room, it tends to be used more generally relative to an acoustic environment, regardless of whether that acoustic environment corresponds to a room).

[0013] like Figure 1As illustrated, a room impulse response typically consists of the direct sound, which depends on the distance from the sound source to the listener, followed by a reverberation component that characterizes the acoustic properties of the room. The size and shape of the room, the positions of the sound source and the listener within the room, and the reflective properties of the room surfaces all contribute to the characteristics of this reverberation component.

[0014] The reverberation can be broken down into two typically overlapping time zones. The first zone contains so-called early reflections, which represent isolated reflections of the sound source on walls or obstacles inside the room before reaching the listener. As the time lag increases, the number of reflections present in a fixed time interval increases, and the path can include second-order or higher-order reflections (e.g., reflections may exit from several walls or both walls and ceiling, etc.).

[0015] The second region in the reverberation is where the density of these reflections increases to the point where they are no longer isolated by the human brain. This region is often referred to as diffuse reverberation, post-reverberation, or the reverberation tail.

[0016] The reverberation section contains cues to the auditory system about the distance of the source, as well as the size and acoustic properties of the room. The energy of the reverberation section relative to the energy of the anechoic section largely determines the perceived distance of the sound source. The level and delay of the earliest reflections can provide cues about how close the sound source is to the walls, and anthropometry filtering can enhance the assessment of a particular wall, floor, or ceiling.

[0017] The density of (early) reflections contributes to the perceived size of a room. The time it takes for reflections to decrease by 60 dB at the energy level (by the reverberation time T) 60 Reverberation time (RW) is a commonly used measure of how quickly reflections dissipate in a room. RW provides information about the acoustic properties of a room; such as, specifically, whether the walls are highly reflective (e.g., a bathroom) or whether there is significant absorption of sound (e.g., a bedroom with furniture, carpets, and curtains).

[0018] Furthermore, when the RIR is part of the binaural room impulse response (BRIR), the RIR may depend on the user's anthropometric properties, as the RIR is filtered by the head, ears, and shoulders; i.e., head-related impulse response (HRIR).

[0019] Since reflections in post-reverberation cannot be distinguished and isolated by the listener, they are usually simulated and represented parametrically using, for example, parametric reverberators employing feedback delay networks, as is well known in the Jot reverberator.

[0020] For early reflections, the incident direction and distance-dependent delay are crucial clues for humans to extract information about the relative positions of the room and sound sources. Therefore, the simulation of early reflections must be more explicit than that of later reverberation. Consequently, in efficient acoustic rendering algorithms, early reflections are simulated differently from later reverberation. A well-known method for early reflections is to mirror the sound sources in each of the room's boundaries to generate virtual sound sources representing the reflections.

[0021] For early reflections, the location of the user and / or sound source relative to the room's boundaries (walls, ceiling, floor) is relevant, while for later reverberation, the room's acoustic response is diffuse and therefore tends to be more uniform throughout the room. This allows for simulations of later reverberation to be generally more computationally efficient than those of early reflections.

[0022] The two main properties of room-defined post-reverberation are the T60 value and the reverberation level. In terms of diffuse reverberation impulse response, these values ​​represent the slope and amplitude of the impulse response. In a natural room, both are typically strongly frequency-dependent.

[0023] The T60 parameter is important for providing an impression of a room's reflectivity and size, while the reverberation level indicates the combined effect of multiple reflections at the room's boundaries. The reverberation level and its frequency behavior depend on the pre-delay, indicating where the difference lies between early reflections and later reverberation (see [link to relevant documentation]). Figure 2 ).

[0024] The reverberation level has a primary psychoacoustic correlation with the direct sound. The difference in level between the two is an indication of the distance between the sound source and the user (or the RIR measurement point). A greater distance will cause more attenuation of the direct sound, while the level of reverberation remains the same (it is the same throughout the room). Similarly, for a source with directionality that depends on where the user is relative to the source, directionality affects the direct response but not the level of reverberation as the user moves around the source.

[0025] A significant challenge and consideration for many systems, such as virtual reality applications, is how to effectively represent and distribute the audio environment. This is often achieved by providing signals representing the individual source signals, along with data that can parameterize the properties of the sound sources and the acoustic environment. This challenge is far from trivial and encompasses a range of issues.

[0026] Separating the descriptions of direct paths and diffuse reverberation has been proposed. However, the question of how to represent, distribute, and render / synthesize diffuse reverberation is currently receiving considerable attention.

[0027] Indicators for providing reverberation levels have been proposed, which are not related to the direct sound but rather to more general properties. A specific proposal has been proposed as part of the preparation for the MPEG-I Audio Proposal Call (CfP), in which the Encoder Input Format (EIF) has been defined (Section 3.9 of MPEG Output Document N19211, “MPEG-I 6DoF Audio EncoderInput Format”, MPEG 130). The EIF defines the reverberation level through pre-delay and the direct-to-diffusion ratio (DDR). The DDR is defined as the ratio between the diffused reverberation energy and the source energy after the pre-delay:

[0028] However, while such parameters may be useful, there are many substantial problems that need to be addressed. For example, there are currently no proposals on how specific parameters should be defined or determined. There is also no consideration of how DDR indicators can be used to render audio or, specifically, how they can be used to generate diffuse reverberation signals.

[0029] EP3402222 discloses a virtualization method for generating binaural signals in response to the channels of a multi-channel audio signal, which applies binaural room impulse response (BRIR) to each channel, including using a common post-mixing response for the downmixing of the channel by using at least one feedback delay network (FDN).

[0030] Therefore, current methods and proposals regarding how to represent and generate audio, and especially diffuse reverberation, are often suboptimal, insufficient, and / or incomplete. This is particularly true for applications such as virtual reality, where the location where audio should be generated can change significantly.

[0031] Therefore, methods for generating diffuse reverberation signals will be advantageous. In particular, methods that allow for improved operation, increased flexibility, reduced complexity, facilitated implementation, improved audio experience, improved audio quality, reduced computational burden, improved adaptability to changing locations, improved performance in virtual / mixed / augmented reality applications, improved perceptual cues of diffuse reverberation, and / or improved performance and / or operation will be advantageous. Summary of the Invention

[0032] Therefore, the present invention seeks to mitigate, alleviate or eliminate one or more of the above-mentioned disadvantages, either individually or in any combination.

[0033] According to one aspect of the invention, an audio apparatus is provided for generating a diffuse reverberation signal for an environment; the apparatus includes: a receiver arranged to receive a plurality of audio signals representing sound sources in the environment; a metadata receiver arranged to receive metadata for the plurality of audio signals, the metadata including: a relationship between the diffuse reverberation signal and a total signal, indicating the level of the diffuse reverberation sound relative to the total emitted sound in the environment, and for each audio signal: a signal level indication; directional data indicating the directionality of sound radiation from the sound source represented by the audio signal; circuitry arranged for, for each of the plurality of audio signals: determining a total emitted energy indication based on the signal level indication and the directional data, and determining a downmixing coefficient based on the total emitted energy and the relationship between the diffuse reverberation signal and the total signal; a downmixer arranged to generate a downmixed signal by combining signal components of each audio signal, the signal components of each audio signal being generated by applying the downmixing coefficient for each audio signal to the audio signal; and a reverberator for generating the diffuse reverberation signal for the environment based on the downmixed signal components.

[0034] In many embodiments, the present invention can provide improved and / or enhanced determination of diffuse reverberant signals. In many embodiments and scenarios, the present invention can generate more natural diffuse reverberant signals, thereby providing improved perception of the acoustic environment. The generation of diffuse reverberant signals can often be achieved with low complexity and low computational resource requirements. This method allows diffuse reverberant sounds in an acoustic environment to be efficiently represented by a relatively small number of parameters, which also provide an efficient representation of individual sources and individual path sound propagation from these sources, and specifically for direct path propagation.

[0035] In many embodiments, this method can allow the generation of diffuse reverberation signals independent of the source and / or listener location. This can allow for the efficient generation of diffuse reverberation signals for dynamic applications involving location changes, such as those used in many virtual reality and augmented reality applications.

[0036] The ratio of diffuse reverberant signal to total signal can also be called the ratio of diffuse reverberant signal level to total signal level, or the ratio of diffuse reverberant level to total level, or the ratio of source energy to diffuse reverberant energy (or its variations / arrangements).

[0037] Audio devices can be implemented in a single device or a single functional unit, or they can be distributed across different devices or functions. For example, an audio device can be implemented as part of a decoder functional unit, or it can be distributed with some functional elements that perform on the decoder side and other elements that perform on the encoder side.

[0038] According to an optional feature of the invention, the directionality of the sound radiation is frequency-dependent, and the circuit is arranged to determine the frequency-dependent total emitted energy and the frequency-dependent downmixing coefficient.

[0039] This method can provide a particularly efficient operation for generating diffuse reverberant signals that reflect frequency correlation.

[0040] According to an optional feature of the invention, the relationship between the diffused reverberation signal and the total signal is frequency-dependent, and the circuit is arranged to generate frequency-dependent downmixing coefficients.

[0041] This method can provide a particularly efficient operation for generating frequency-dependent diffused reverberation signals that reflect frequency correlations.

[0042] According to an optional feature of the invention, the relationship between the diffused reverberation signal and the total signal includes a frequency-dependent portion and a non-frequency-dependent portion, and wherein the circuit is arranged to generate the downmixing coefficients based on the non-frequency-dependent portion and to adjust the reverberator based on the frequency-dependent portion.

[0043] This method provides a particularly efficient operation for generating diffused reverberant signals that reflect frequency correlations, and can specifically reduce complexity and / or resource usage. For example, this method can allow frequency correlations to be reflected through a single filtering of the downmixed signal.

[0044] According to an optional feature of the invention, the circuit is arranged to determine the total emission energy indication for a first audio signal among the plurality of audio signals in response to scaling the signal level indication for the first audio signal by integrating a directional pattern of the sound source represented by the first audio signal.

[0045] In many embodiments, this can provide particularly advantageous operation. The scaling can be any function applied to the signal level indication in conjunction with a determined downmixing coefficient. This function typically increases monotonically based on the total transmit energy indication. The scaling can be linear or non-linear.

[0046] Scaling can be independent of the time variation of the signal, and therefore may not need to be updated with the instantaneous level of the audio signal, and may only need to be recalculated when the signal level indication or directional pattern changes.

[0047] According to an optional feature of the invention, the signal level indication for a first audio signal among the plurality of audio signals includes a reference distance, the reference distance indicating a distance reference gain relative to the first audio signal and the distance from the sound source represented by the first audio signal.

[0048] In many embodiments, this can provide particularly advantageous operation. The distance reference gain can be a predetermined value, and is typically common to at least some and generally all sound sources and signals. In many embodiments, the distance reference gain can be 0 dB.

[0049] According to an optional feature of the invention, the integration is performed with respect to the following distance from the sound source represented by the first audio signal, the distance being the reference distance.

[0050] This can provide a particularly effective method and can facilitate operations.

[0051] According to an optional feature of the invention, the relationship between the diffuse reverberation signal and the total signal indicates the energy of the diffuse reverberation sound relative to the energy of the total emitted sound in the environment.

[0052] In many embodiments, this can provide particularly advantageous operation.

[0053] According to an optional feature of the invention, the relationship between the diffuse reverberation signal and the total signal indicates the initial amplitude of the energy of the diffuse sound relative to the total emitted sound in the environment.

[0054] In many embodiments, this can provide particularly advantageous operation.

[0055] According to an optional feature of the invention, the downmixing coefficient determined for a first audio signal among the plurality of audio signals is independent of the location of a first sound source represented by the first audio signal.

[0056] This can provide particularly advantageous operation in many embodiments and can particularly facilitate operation for dynamic applications with sound sources that change position, such as for virtual reality applications.

[0057] According to an optional feature of the invention, the downmixing coefficient determined for the first audio signal among the plurality of audio signals is independent of the listener's position.

[0058] This can provide particularly advantageous operation in many embodiments and can particularly facilitate operation for dynamic applications with changing positions, such as for virtual reality applications.

[0059] In some embodiments, the processing of the audio device is independent of the location of the sound source. In some embodiments, the processing of the audio device is independent of the location of the listener.

[0060] In some embodiments, the processing of the audio device is independent of the listener's location within the area where the diffuse reverberation signal to total signal ratio applies.

[0061] In some embodiments, the update rate of the downmixing coefficients is lower than the update rate of the location of the first sound source represented by the first audio signal. In some embodiments, the update rate of the downmixing coefficients is lower than the update rate of the listener's location. The downmixing coefficients can be calculated at a time rate much lower than the update rate of the listener's location / sound source location.

[0062] According to an optional feature of the invention, the signal level indication for a first audio signal among the plurality of audio signals further includes a gain indication for the first audio signal, the gain indication indicating the gain to be applied to the first audio signal when rendering sound from a first sound source represented by the first audio signal, and wherein the circuitry is arranged to determine the downmixing coefficient for the first audio signal in response to the gain indication.

[0063] According to an optional feature of the invention, the audio device further includes a direct rendering circuit arranged to generate a direct path audio signal for the first audio signal in response to the signal level indication and the directionality data for the first audio signal among the plurality of audio signals.

[0064] In many embodiments, this can provide particularly advantageous operation.

[0065] According to an optional feature of the invention, the metadata also includes a delay indication, and the diffuse reverberation signal to total signal ratio (DSR) indicates the energy of diffuse reverberant sound in the environment having a longer delay than the delay indicated by the delay indication, relative to the energy of the total emitted sound.

[0066] The energy of diffuse reverberant sound in an environment with a delay longer than the delay indication can be reflected by the room impulse response contribution that occurs at least a certain delay after the corresponding sound is emitted at the sound source / determined by the room impulse response contribution that occurs at least a certain delay after the corresponding sound is emitted at the sound source / as the room impulse response contribution that occurs at least a certain delay after the corresponding sound is emitted at the sound source, wherein the certain delay is indicated by the delay indication.

[0067] In some embodiments, the diffuse reverberant signal-to-total signal ratio (DSR) indicates the energy of the diffuse reverberant sound relative to the energy of the total emitted sound in the environment, wherein the energy of the diffuse reverberant sound is determined by the contribution of the room response that occurs at least a certain delay after the corresponding sound is emitted at the sound source.

[0068] According to another aspect of the present invention, a method for generating a diffuse reverberation signal for an environment is provided, the method comprising: receiving a plurality of audio signals representing sound sources in the environment; receiving metadata for the plurality of audio signals, the metadata including: a relationship between the diffuse reverberation signal and a total signal, indicating the level of the diffuse reverberation sound relative to the total emitted sound in the environment, and for each audio signal: a signal level indication; directional data indicating the directionality of sound radiation from the sound source represented by the audio signal; for each of the plurality of audio signals: determining a total emitted energy indication based on the signal level indication and the directional data, and determining a downmixing coefficient based on the total emitted energy and the relationship between the diffuse reverberation signal and the total signal; generating a downmixing signal by combining signal components of each audio signal, the signal components of each audio signal being generated by applying the downmixing coefficient for each audio signal to the audio signal; and generating the diffuse reverberation signal for the environment based on the downmixing signal components.

[0069] These and other aspects, features and advantages of the invention will become apparent from reference to one or more embodiments described below, and will be set forth with reference to one or more embodiments described below. Attached Figure Description

[0070] Embodiments of the invention will now be described by way of example only with reference to the accompanying drawings, wherein...

[0071] Figure 1 An example of room impulse response is illustrated;

[0072] Figure 2 An example of room impulse response is illustrated;

[0073] Figure 3 The illustration shows an example of the components of a virtual reality system;

[0074] Figure 4 The illustration shows an example of an audio device for generating audio output according to some embodiments of the present invention;

[0075] Figure 5 An example of an audio reverberation device for generating a diffuse reverberation signal according to some embodiments of the present invention is illustrated;

[0076] Figure 6 An example of room impulse response is illustrated; and

[0077] Figure 7 An example of a reverb is shown. Detailed Implementation

[0078] The following description will focus on audio processing and generation for virtual reality applications; however, it should be understood that the principles and concepts described can be used in many other applications and embodiments.

[0079] Virtual experiences that allow users to move around in a virtual world are becoming increasingly popular, and services are being developed to meet this demand.

[0080] In some systems, VR applications can be provided locally to the viewer by a standalone device that, for example, does not use any remote VR data or processing, or even has any access to any remote VR data or processing. For example, the device (such as a game console) may include storage for storing scene data, input for receiving / generating viewer gestures, and a processor for generating corresponding images based on the scene data.

[0081] In other systems, VR applications can be implemented and executed remotely from the viewer. For example, a device local to the user can detect / receive motion / pose data transmitted to a remote device that processes the data to generate the viewer's posture. The remote device can then generate a suitable view image and corresponding audio signal for the user's posture based on scene data describing the scene. The view image and corresponding audio signal are then transmitted to the device local to the viewer where they are presented. For example, the remote device can directly generate a video stream (typically a stereoscopic / 3D video stream) and corresponding audio stream that are directly presented by the local device. Therefore, in such an example, the local device may not perform any VR processing other than transmitting motion data and presenting the received video data.

[0082] In many systems, functionality can be distributed across local and remote devices. For example, a local device can process received input and sensor data to generate user poses that are continuously transmitted to a remote VR device. The remote VR device can then generate corresponding view images and corresponding audio signals and transmit these to the local device for rendering. In other systems, instead of directly generating view images and corresponding audio signals, the remote VR device may select relevant scene data and transmit it to the local device, which can then generate the rendered view images and corresponding audio signals. For example, the remote VR device can identify the nearest capture point and extract the corresponding scene data (e.g., a set of object sources and their location metadata) and transmit this to the local device. The local device can then process the received scene data to generate images and audio signals specific to the current user pose. User poses typically correspond to head poses, and references to user poses can generally be considered equivalent to references to head poses.

[0083] In many applications, especially for broadcast services, a source can transmit or stream scene data as an image (including video) and audio representation of the scene, independent of the user's pose. For example, signals and metadata corresponding to sound sources within a virtual room can be transmitted or streamed to multiple clients. Individual clients can then locally synthesize audio signals corresponding to their current user pose. Similarly, a source can transmit a general description of the audio environment, including descriptions of sound sources and the acoustic characteristics of the environment. An audio representation can then be generated locally and presented to the user, for example, using binaural rendering and processing.

[0084] Figure 3 The illustration shows an example of a VR system in which a remote VR client device 301 communicates with a VR server 303, for example, via a network 305 (e.g., the Internet). The server 303 can be configured to simultaneously support a potentially large number of client devices 301.

[0085] VR server 103 can support broadcast experiences, for example, by transmitting image signals that include image representations in the form of image data. Client devices can use these image signals to locally synthesize view images corresponding to appropriate user poses (poses refer to position and / orientation). Similarly, VR server 303 can transmit audio representations of the scene, thereby allowing audio to be synthesized locally for user poses. Specifically, as the user moves around in the virtual environment, the synthesized images and audio presented to the user are updated to reflect the user's current (virtual) position and orientation in the (virtual) environment.

[0086] In many applications (such as) Figure 3 In applications such as scene modeling, it may be desirable to model the scene and generate effective image and audio representations that can be efficiently included in a data signal, which can then be transmitted or streamed to various devices that can synthesize views and audio locally for poses different from the captured pose.

[0087] In some embodiments, a model representing a scene can be stored locally, for example, and can be used locally to synthesize appropriate images and audio. For example, an audio model of a room can include the properties of sound sources that can be heard in the room and an indication of the room's acoustic properties. The model data can then be used to synthesize appropriate audio for a specific location.

[0088] How to represent an audio scene and how to use that representation to generate audio is a crucial question. Audio rendering aimed at providing listeners with natural and realistic effects often involves the rendering of the acoustic environment. For many environments, this includes the representation and rendering of diffuse reverberation present in the environment (such as in a room). It has been found that the rendering and representation of such diffuse reverberation has a significant impact on the perception of the environment, such as whether the audio is perceived as representing a natural and realistic environment. Below, advantageous methods for representing audio scenes and rendering audio based on that representation, particularly diffuse reverberant audio, will be described.

[0089] References Figure 4 The method is described using an audio apparatus. The audio apparatus is arranged to generate an audio output signal representing audio in an acoustic environment. Specifically, the audio apparatus can generate audio representing audio perceived by a user moving around in a virtual environment with multiple sound sources and given acoustic properties. Each sound source is represented by an audio signal representing the sound from the sound source and metadata that can describe the characteristics of the sound source, such as providing a level indication of the audio signal. Additionally, metadata is provided to characterize the acoustic environment.

[0090] The audio device includes a path renderer 401 for each sound source. Each path renderer 401 is arranged to generate a direct path signal component representing the direct path from the sound source to the listener. The direct path signal component is generated based on the positions of the listener and the sound source, and can be specifically generated by scaling the audio signal for the potential frequency dependence of the sound source according to distance and, for example, the relative gain of the sound source in a specific direction to the user (e.g., for a non-omnidirectional source).

[0091] In many embodiments, renderer 401 may also generate a direct path signal based on occlusion or diffraction (virtual) elements between the source location and the user location.

[0092] In many embodiments, path renderer 401 may also generate additional signal components for individual paths, wherein these signal components include one or more reflections. This can be done, for example, by evaluating reflections from walls, ceilings, etc., as will be known to those skilled in the art. The direct path and reflected path components can be combined into a single output signal for each path renderer, and thus a single signal representing the direct path and early / discrete reflections can be generated for each sound source.

[0093] In some embodiments, the output audio signal of each sound source can be a binaural signal, and therefore each output signal can include both left-ear and right-ear (sub) signals.

[0094] The output signal from path renderer 401 is provided to combiner 403, which combines the signals from different path renderers 401 to generate a single combined signal. In many embodiments, binaural output signals can be generated, and the combiner can perform the combination of individual signals from path renderers 401 (such as weighted combination), that is, all right-ear signals from path renderers 401 can be added together to generate a combined right-ear signal, and all left-ear signals from path renderers 401 can be added together to generate a combined left-ear signal.

[0095] Path renderers and combiners can be implemented in any suitable manner, typically comprising executable code for processing on suitable computing resources, such as microcontrollers, microprocessors, digital signal processors, or central processing units including supporting circuitry such as memory. It should be understood that multiple path renderers can be implemented as parallel functional units (such as, for example, a set of dedicated processing units), or as repetitive operations for each sound source. Typically, the same algorithm / code is executed for each sound source / signal.

[0096] In addition to the individual path audio components, the audio device is also configured to generate signal components representing diffuse reverberation in the environment. The diffuse reverberation signal is generated (effectively) by combining the source signals into a downmixed signal and then applying a reverberation algorithm to the downmixed signal to generate the diffuse reverberation signal.

[0097] Figure 4 The audio device includes a downmixer 405 that receives audio signals from multiple sound sources (typically all sources within an acoustic environment where a reverberator simulates diffuse reverberation) and combines them into a downmix. Therefore, the downmix reflects all the sounds generated in the environment. The downmix is ​​fed to a reverberator 407, which is arranged to generate a diffuse reverberation signal based on the downmix. The reverberator 407 may specifically be a parametric reverberator, such as a Jot reverberator. The reverberator 407 is coupled to a combiner 403, to which the diffuse reverberation signal is fed. The combiner 403 then continues to combine the diffuse reverberation signal with path signals representing individual paths to generate a combined audio signal representing the combined sounds in the environment as perceived by the listener.

[0098] References Figure 5 The illustrated audio reverberation device further illustrates the generation of a diffuse reverberation signal. The audio reverberation device may include... Figure 4 In the audio device, the downmixer 405 and the reverb 407 can be specifically implemented.

[0099] The audio reverberation device includes a receiver 501 arranged to receive audio scene data representing audio. The audio scene data specifically includes a plurality of audio signals, each of which represents a sound source (and thus the audio signal describes the sound from the sound source). Additionally, the receiver 501 receives metadata for each sound source. This metadata includes a (relative) signal level indication of the sound source, wherein the signal level indication may indicate the level / energy / amplitude of the sound source represented by the audio signal. The metadata of the source also includes directional data indicating the directionality of sound radiation from the sound source. The directional data of the audio signal may, for example, describe a gain pattern and may specifically describe the relative gain / energy density of the sound source in a direction different from the location of the sound source.

[0100] Receiver 501 also receives metadata indicating the acoustic environment. Specifically, receiver 501 receives the relationship between the diffuse reverberation signal and the total signal, and more specifically, receives the diffuse reverberation signal to total signal ratio (also referred to as the diffuse reverberation signal level to total signal level ratio, or in some cases, the diffuse reverberation signal level to total signal energy ratio, or the emitted energy to diffuse reverberation energy ratio), which indicates the level of diffuse reverberant sound relative to the total emitted sound in the acoustic environment. In the following text, for the sake of brevity, the diffuse reverberation signal to total signal ratio will also be referred to as the diffusion-to-source ratio (DSR) or the equivalent source-to-diffusion ratio (SDR) (the former will be used primarily in the following description).

[0101] It will be recognized that ratios and inverse ratios can provide the same information, i.e., any ratio can be expressed as an inverse ratio. Therefore, the relationship between the diffuse reverberant signal and the total signal can be expressed by dividing the value reflecting the level of the diffuse reverberant sound by the fraction reflecting the value of the total emitted sound, or equivalently by dividing the value reflecting the total emitted sound by the fraction reflecting the level of the diffuse reverberant sound. It should also be understood that various modifications to the estimate can be introduced; for example, nonlinear functions (such as logarithmic functions) can be applied.

[0102] This can be used in metadata to provide any indication of the relationship between the diffuse reverberant signal and the total signal, indicating the level of diffuse reverberant sound relative to the total emitted sound in an acoustic environment. The following description will focus on the relationship represented by the ratio between the level of the diffuse reverberant signal and the level of the total signal (e.g., energy or energy density). Therefore, the description will focus on an example of the diffuse reverberant signal to total signal ratio, which will also be referred to as DSR.

[0103] Receiver 501 can be implemented in any suitable manner, including, for example, using discrete or application-specific electronics. Receiver 501 can be implemented, for example, as an integrated circuit such as an application-specific integrated circuit (ASIC). In some embodiments, the circuitry can be implemented as a programmable processing unit, such as firmware or software running on a suitable processor, such as a central processing unit, digital signal processing unit, or microcontroller. Such circuitry can also be implemented as part of a processing unit, an integrated circuit, and / or discrete electronic circuitry.

[0104] Receiver 501 can receive audio scene data from any suitable source and in any suitable form, including, for example, as part of an audio signal. Data can be received from internal or external sources. Receiver 501 can be configured, for example, to receive room data via a network connection, a radio connection, or any other suitable connection to an internal source. In many embodiments, the receiver can receive data from a local source, such as local memory. In many embodiments, receiver 501 can, for example, be configured to retrieve room data from local memory, such as local RAM or ROM memory.

[0105] Receiver 501 can be coupled to path renderer 401 and can forward audio scene data to these path renderers 401 for use in generating path signal components (direct path and early reflections) as described above.

[0106] The audio reverb device also includes a downmixer 405, which is also fed audio scene data. The downmixer 405 includes an energy circuit / processor 505, a coefficient circuit / processor 507, and a downmixing circuit / processor 509.

[0107] The downmixer 405, and each of the actual power circuit / processor 505, coefficient circuit / processor 507, and downmixer circuit / processor 509, can be implemented in any suitable manner, including, for example, using discrete or dedicated electronics. The receiver 501 can be implemented, for example, as an integrated circuit such as an application-specific integrated circuit (ASIC). In some embodiments, the circuitry / processor can be implemented as a programmable processing unit, such as firmware or software running on a suitable processor, such as a central processing unit, digital signal processing unit, or microcontroller. It should be understood that in these embodiments, the processing unit may include on-board or external memory, clock drive circuitry, interface circuitry, user interface circuitry, etc. These circuits can also be implemented as part of the processing unit, integrated circuits, and / or discrete electronic circuitry.

[0108] Coefficient processor 507 is arranged to determine downmixing coefficients for at least some of the audio signals in the received audio signal. The downmixing coefficients of the audio signals may correspond to weights of those audio signals in the downmix. The downmixing coefficients may be weights of the audio signals in a weighted combination that generates the downmixed signal. Therefore, when combining audio signals to generate a downmixed signal (which is a mono signal in many embodiments), the downmixing coefficients may be relative weights of the audio signals; for example, they may be weights in a weighted summation.

[0109] The coefficient processor 507 is configured to generate downmixing coefficients based on the received diffuse reverberation signal to total signal ratio (i.e., diffuse to source ratio DSR).

[0110] The coefficients are further determined in response to the determined total emitted energy indication, which indicates the total energy emitted from the sound source. Although the DSR is generally common to some and usually all audio signals, the total emitted energy indication is usually specific to each sound source.

[0111] The total emitted energy indicator typically indicates the normalized total emitted energy. The same normalization can be applied to all sound sources as well as direct and reflected path components. Therefore, the total emitted energy indicator can be a total emitted energy indicator relative to other sound sources / signals, or a relative value relative to individual path components, or a relative value relative to the full-scale sample value of the audio signal.

[0112] When combined with DSR, the total emission energy indicator can provide downmixing coefficients for each sound source, reflecting the relative contribution to the diffuse reverberation sound from that source. Therefore, determining the downmixing coefficients as a function of DSR and the total emission energy indicator can provide downmixing coefficients reflecting the relative contribution to the diffuse sound. Thus, using downmixing coefficients to generate a downmixed signal results in a downmixed signal that reflects the total generated sound in the environment, where each of the sound sources is appropriately weighted, and where the acoustic environment is accurately modeled.

[0113] In many embodiments, the downmixing coefficient, which is a function of DSR and total emission energy indication, combined with scaling in response to the properties of the reverberator (407), can provide an appropriate relative level of downmixing coefficient that reflects the diffuse reverberant sound relative to the corresponding path signal component.

[0114] Energy processor 505 is coupled to coefficient processor 507 and is arranged to determine the total transmit energy indication from metadata received for the sound source.

[0115] The received metadata includes a signal reference level for each source, which provides an indication of the audio level. The signal reference level is typically a normalized or relative value, providing an indication of the signal reference level relative to other sound sources or relative to a normalized reference level. Therefore, the signal reference level may not typically indicate the absolute sound level of a source, but rather its relative level to other sound sources.

[0116] In a specific example, the signal reference level can include an indication in the form of a reference distance, providing the distance at which a 0 dB attenuation is applied to the audio signal. Therefore, for distances equal to the reference distance between the sound source and the listener, the received audio signal can be used without any distance-dependent scaling. For distances smaller than the reference distance, the attenuation is smaller, and therefore a gain greater than 0 dB should be applied when determining the sound level at the listening location. For distances greater than the reference distance, the attenuation is even greater, and therefore a gain greater than 0 dB should be applied when determining the sound level at the listening location. Equivalently, for a given distance between the sound source and the listening location, a higher gain will be applied to the audio signal associated with a higher reference distance compared to an audio signal associated with a shorter reference distance. Since audio signals are typically normalized to represent meaningful reference distances or utilize full dynamic range (e.g., jet engines and cricket balls would both be represented by audio signals utilizing the full dynamic range of the data words used), the reference distance provides an indication of the signal reference level for a particular sound source.

[0117] In this example, the signal reference level is also indicated by a reference gain called pregain. A reference gain is provided for each sound source, and the gain that should be applied to the audio signal when determining the rendered audio level is also provided. Therefore, pregain can be used to further indicate level variations between different sound sources.

[0118] The metadata also includes directional data indicating the directionality of sound radiation from a sound source represented by an audio signal. The directional data for each sound source can indicate the relative gain relative to a signal reference level in different directions from the sound source. The directional data can, for example, provide a complete function or description of the radiation pattern from the sound source, defining the gain in each direction. As another example, simplified indications can be used, such as a single data value indicating a predetermined pattern. As yet another example, the directional data can provide individual gain values ​​for a series of different directional intervals (e.g., segments of a sphere).

[0119] Therefore, metadata, together with the audio signal, can allow the generation of audio levels. Specifically, the path renderer can determine the signal components of the direct path by applying gain to the audio signal, where the gain is a combination of pregain, distance gain determined based on the distance between the sound source and the listener and a reference distance, and directional gain in the direction from the sound source to the listener.

[0120] Regarding the generation of diffuse reverberant signals, metadata is used to determine the (normalized) total emitted energy indication of the sound source based on the signal reference level and the directionality data of the sound source.

[0121] Specifically, the total emission energy indication can be generated by integrating the directional gain in all directions (e.g., on the surface of a sphere centered on the location of the sound source), and scaled by the signal reference level and specifically by the distance gain and pre-gain.

[0122] The determined total transmit energy indication is then fed to coefficient processor 507, where it is processed by DSR to generate undermixing coefficients.

[0123] Then, the downmixing processor 509 uses downmixing coefficients to generate a downmixed signal. Specifically, the downmixed signal can be generated as a combination of audio signals and summed, where each audio signal is weighted by the downmixing coefficients of the corresponding audio signal.

[0124] The downmix is ​​typically generated as a mono signal, which is then fed to reverb unit 407, which continues to generate a diffuse reverb signal.

[0125] It should be noted that although the rendering and generation of individual path signal components by the path renderer 401 is position-dependent, such as with regard to determining distance gain and directional gain, the generation of the diffuse reverberation signal can then be independent of the positions of both the source and the listener.

[0126] The total emitted energy indication can be determined based on signal reference level and directional data, regardless of the source and listener locations. Specifically, the source's pregain and reference distance can be used to determine the non-directionally correlated signal reference level at a nominal distance from the source (the nominal distance is the same for all audio signals / sources), and it is normalized relative to, for example, a full-scale sample of the audio signal. The directional gain can be integrated in all directions over the normalized sphere (e.g., over the sphere at the reference distance). Therefore, the total emitted energy indication will be independent of the source and listener locations (reflecting that diffuse reverberant sound tends to be uniform in an environment such as a room). The total emitted energy indication is then combined with the DSR to generate the downmixing coefficients (in many implementations, other parameters, such as reverberator parameters, may also be considered). Since the DSR is also location-independent, so is the downmixing and reverberation processing, thus a diffuse reverberant signal can be generated without considering the specific locations of the source and listener.

[0127] This approach can provide high-performance and natural-sounding audio perception without requiring excessive computational resources. It is particularly well-suited for applications such as virtual reality, where the user (and the source) can move around in the environment, and therefore the relative positions of the listener (and possibly some or all of the sound sources) can change dynamically.

[0128] The following text will describe this in more detail. Figure 4 and Figure 5 Specific aspects of various embodiments of the method.

[0129] In many embodiments, the metadata may also include an indication of when the diffuse reverberation signal should begin; that is, it may indicate the time delay associated with the diffuse reverberation signal. The time delay indication may specifically be in the form of a pre-delay.

[0130] Pre-delay can represent the delay / hysteresis in RIR and can be defined as a threshold between early reflections and diffusion, and later reverberation. Since this threshold typically occurs as part of a smooth transition from a mixture of (more or less) discrete reflections to higher-order reflections that constitute full disturbance, an appropriate evaluation / decision process can be used to select a suitable threshold. This determination can be performed automatically based on RIR analysis or calculated based on room dimensions and / or material properties.

[0131] Alternatively, a fixed threshold can be chosen, such as 80 ms into the RIR. The pre-delay can be indicated in seconds, milliseconds, or samples. In the description below, it is assumed that the pre-delay is chosen at the point after which the reverberation actually diffuses. However, if this is not the case, the described method can still work adequately.

[0132] Therefore, the pre-delay indicates the start of the diffuse reverberation response from the start of source emission. For example, for... Figure 6 In the example shown, if the source begins emitting at t0 (e.g., t0=0), the direct sound reaches the user at t1>t0, the first reflection reaches the user at t2>t1, and the defined threshold between the early reflections and the diffuse reverberation reaches the user at t3>t2. Therefore, the pre-delay is t3-t0.

[0133] In this system, the diffuse reverberation signal to total signal ratio (i.e., the diffusion to source ratio DSR) can be used to express the amount of diffuse reverberation energy or level of the source received by the user as a ratio to the total emitted energy of that source. It can be expressed as follows: the diffuse reverberation energy is appropriately adjusted for level calibration of the signal to be rendered and the corresponding metadata (e.g., pregain).

[0134] Expressing it in this way ensures that the value is independent of the absolute position and orientation of the listener and source in the environment, independent of the user's relative position and orientation relative to the source, and vice versa, independent of the specific algorithm used to render the reverberation, and has a meaningful relationship with the signal level used in the system.

[0135] The described method calculates a downmixing coefficient that takes into account the directional pattern to apply the correct relative level between the source signals and takes into account the DSR to achieve the correct level on the output of the reverberator 407.

[0136] DSR can represent the ratio between the energy of the emission source and the diffuse reverberation properties (such as the energy or (initial) level of the specific diffuse reverberant signal).

[0137] This description will focus primarily on the DSR, which indicates the diffuse reverberation energy relative to the total energy:

[0138] Diffusion reverberation energy can be considered as the energy generated by the room response from the beginning of the diffusion phase; for example, it can be the energy from the RIR indicated by the pre-delay up to infinity. Note that subsequent excitations of the room will sum to the reverberation energy, so this can typically only be measured directly using excitation with a Dirac pulse. Alternatively, it can be derived from the measured RIR.

[0139] Reverberation energy represents the energy at a single point in the diffuse field space, rather than the energy integrated over the entire space.

[0140] A particularly advantageous alternative to the above would be to use the DSR, which indicates the initial magnitude of the energy of the diffuse sound relative to the total emitted sound in the environment. Specifically, the DSR can indicate the reverberation magnitude at a time indicated by a pre-delay.

[0141] The amplitude at the pre-delay can be the maximum excitation of the room impulse response at or immediately following the pre-delay. For example, within 5, 10, 20, or 50 ms after the pre-delay. The reason for choosing a specific range of maximum excitation is that at the pre-delay time, the room impulse response may coincidentally be in the lower part of the response. When the overall trend is decay amplitude, the maximum excitation within a short interval after the pre-delay is usually also the maximum excitation of the entire diffuse reverberation response.

[0142] Using a DSR that indicates the initial amplitude (within, for example, 10-millisecond intervals) makes mapping the DSR to parameters easier and more robust in many reverberation algorithms. Therefore, in some embodiments, the DSR can be given as:

[0143] The parameters in a DSR are expressed relative to the same source signal horizontal reference.

[0144] This can be achieved, for example, by measuring (or simulating) the RIR of the room of interest with a microphone under certain known conditions, such as the distance between the source and the microphone and the directional pattern of the source. The source should emit a calibrated amount of energy into the room, such as a Dirac pulse with a known energy.

[0145] The calibration factors for electrical and analog-to-digital conversions in measuring equipment can be measured or derived from specifications. It can also be calculated based on the direct path response in the RIR, which can be predicted based on the source's directional pattern and the source-microphone distance. The direct path response has a certain energy in the digital domain and represents the emitted energy multiplied by the directional gain and distance gain in the microphone direction. This distance gain can depend on the microphone surface relative to the total spherical surface area with a radius equal to the source-microphone distance.

[0146] Both components should use the same digital level reference. For example, a full-scale 1kHz sine wave corresponds to 100 dBSPL.

[0147] Measuring the diffuse reverberation energy from the RIR and compensating it with a calibration factor yields the appropriate energy in the same domain as the known emission energy. Together with the emission energy, the appropriate DSR can be calculated.

[0148] The reference distance can indicate the distance at which a distance gain of 0dB is applied to the signal, meaning that no gain or attenuation should be applied to compensate for the distance. The actual distance gain applied by the path renderer 401 can then be calculated by considering the actual distance relative to the reference distance.

[0149] The effect of distance on sound propagation is performed with reference to a given distance. Doubling the distance reduces the energy density (energy per surface area) by 6 dB. Halving the distance causes the energy density (energy per surface area) to increase by 6 dB.

[0150] To determine the distance gain at a given distance, the distance corresponding to the given level must be known, so that the relative change in the current distance can be determined, i.e., how much the density has decreased or increased.

[0151] Ignoring absorption in the air and assuming no reflection or occlusion elements, the emitted energy of the source is constant on any sphere with any radius centered at the source location. The ratio of the actual distance to the reference distance on the surface indicates the energy attenuation. The linear signal amplitude gain at the rendering distance d can be expressed as b:

[0152] Where, r ref This is a reference distance.

[0153] As an example, if the reference distance is 1 meter and the rendering distance is 2 meters, this results in approximately 6 dB of signal attenuation (or -6 dB of gain).

[0154] Total emitted energy indicates the total energy emitted by a sound source. Typically, a sound source radiates in all directions, but not equally in all directions. The total emitted energy can be obtained by integrating the energy density over a sphere surrounding the source. In the case of a loudspeaker, the emitted energy can usually be calculated using the voltage applied to the terminals and the loudspeaker coefficient, which describes the impedance, energy loss, and the transfer of electrical energy to the sound pressure wave.

[0155] Energy processor 505 is configured to determine the total emitted energy indication by taking into account the directional data of the sound source. It should be noted that when determining the diffuse reverberant signal of a source that may have varying source directionality, it is important to use the total emitted energy, not just the signal level or signal reference level. For example, consider the source directionality corresponding to a very narrow beam with a directivity coefficient of 1 and coefficients of 0 for all other directions (i.e., energy is transmitted only in the very narrow beam). In this case, the emitted source energy can be very similar to the energy of the audio signal and the signal reference level, since this represents the total energy. If alternatively, another source with the same energy and signal reference level but omnidirectional directionality is considered, the emitted energy of that source will be much higher than the audio signal energy and the signal reference level. Therefore, in the case where both sources are activated simultaneously, the signal of the omnidirectional source should be represented much stronger in the diffuse reverberant signal and thus in the downmixing than that of the very directional source.

[0156] As described above, the energy processor 505 can determine the emitted energy by integrating the energy density over the surface of a sphere surrounding the sound source. Ignoring distance gain, i.e., integrating over a surface with a radius of 0 dB (i.e., where the radius corresponds to a reference distance), the total emitted energy indication can be determined according to the following formula:

[0157] in, It is a directional gain function. It is the pregain associated with the audio signal / source, and Indicates the level of the audio signal itself.

[0158] because It is direction-independent, therefore it can also extend beyond the integral. Similarly, the signal... Independent of direction (directional gain reflects this change). (It can be multiplied later because:)

[0159] And therefore the integral becomes independent of the signal.

[0160] A specific method for determining this integral is described in more detail below.

[0161] The goal is to integrate the directional gain over the sphere.

[0162] Using a sphere with a radius equal to the reference distance (r) means that the distance gain is 0 dB, and therefore the distance gain / attenuation can be ignored.

[0163] In this example, a sphere is chosen because it provides advantageous computation, but the same energy can be determined from any closed surface of any shape surrounding the source location. An effective surface facing the source location (i.e., having a normal vector consistent with the source location) is considered, provided that appropriate distance and directional gains are used in the integration.

[0164] The surface integral should define the small surface dS. Therefore, defining a sphere with two parameters (azimuth (a) and elevation (a)) provides the dimension for doing so. Using a coordinate system for our solution, we obtain:

[0165] f(a, e, r) = r * cos(e) * cos(a) * u x + r * cos(e) * cos(a) * u y + r *sin(e) * u z

[0166] Among them, u x u y and u z It is the unit basis vector of the coordinate system.

[0167] The small surface dS is the magnitude of the cross product of the partial derivatives of the sphere with respect to the two parameters multiplied by the differential of each parameter:

[0168] dS = |f a xf e | da de

[0169] The derivative determines the vector that is tangent to the sphere at the point of interest. f a = -r * cos(e) * sin(a) * u x + r * cos(e) * cos(a) * u y + 0 * u z f e = -r *sin(e) * cos(a) * u x - r * sin(e) * sin(a) * u y + r * cos(e) * u z

[0170] The cross product of derivatives is a vector perpendicular to both. f a xf e = (r 2 * cos(e) * cos(a) * cos(e) + 0 * sin(e) * sin(a)) * u x + (-0* sin(e) * cos(a) + r 2 * cos(e) * sin(a) * cos(e)) * u y + (r 2 * cos(e) * sin(a) *sin(e) * sin(a) + r 2 * cos(e) * cos(a) * sin(e) * cos(a)) * u z = r 2 * cos 2 (e) * cos(a) * u x + r 2 * cos 2 (e) * sin(a) * u y + (r 2* cos(e) *sin(e) * sin 2 (a) + r 2 * cos(e) * sin(e) * cos 2 (a)) * u z = r 2 * cos 2 (e) * cos(a) *u x + r 2 * cos 2 (e) * sin(a) * u y + (r 2 * cos(e) * sin(e) * (sin 2 (a) + cos 2 (a))) * u z =r 2 * cos 2 (e) * cos(a) * u x + r 2 * cos 2 (e) * sin(a) * u y + r 2 * cos(e) * sin(e) * u z

[0171] The magnitude of the cross product is the surface area of ​​the parallelogram spanned by vectors f_a and f_e, and therefore the surface area on a sphere: |f a xf e | = sqrt((r 2 * cos 2 (e) * cos(a)) 2 + (r 2 * cos 2 (e) * sin(a)) 2 + (r 2 *cos(e) * sin(e)) 2 ) = sqrt(r 4 * cos 4 (e) * cos 2 (a) + r 4 * cos 4 (e) * sin 2 (a) + r 4 * cos 2 (e) *sin 2(e))= sqrt(r 4 * cos 4 (e) * (cos 2 (a) + sin 2 (a)) + r 4 * cos 2 (e) * sin 2 (e))= sqrt(r 4 * cos 4 (e) + r 4 * cos 2 (e) * sin 2 (e))= sqrt(r 4 * cos 2 (e) * (cos 2 (e) + sin 2 (e)))=sqrt(r 4 * cos 2 (e))= abs(r 2 * cos(e)) = r 2 * cos(e) when e = [-0.5*pi, 0.5*pi]

[0172] lead to:

[0173] dS = r 2 * cos(e) * da * de

[0174] The first two terms define the normalized surface area, and by multiplying by da and de, it becomes the actual surface based on the size of segments da and de. The double integral on the surface can then be expressed according to the azimuth and elevation angles. As mentioned above, the surface dS is expressed according to a and e. Two integrals can be performed at azimuth = 0 ... 2*pi (inner integral) and elevation = -0.5*pi ... 0.5*pi (outer integral).

[0175] in, It is directionality as a function of azimuth and elevation. Therefore, if The result should be the surface of a sphere. (Analytically calculating the integral as proof leads to the expected result) ).

[0176] In many practical embodiments, the directional pattern may not be provided as an integrable function, but rather as, for example, a discrete set of sample points. For instance, the directional gain of each sample is associated with azimuth and elevation angles. Typically, these samples would represent a grid on a sphere. One way to handle this is to convert the integration into a summation, i.e., to perform a discrete integral. In this example, the integral could be implemented as a summation over the points on the sphere where the directional gain is available. This gives... The value, but requires correct selection. and This ensures that they do not cause large errors due to overlap or gaps.

[0177] In other embodiments, the directional pattern can be provided as a finite number of non-uniformly spaced points in space. In this case, the directional pattern can be interpolated and uniformly resampled within the range of azimuth and elevation angles of interest.

[0178] Alternative solutions could be hypothetical The integral is constant around its defined point and is solved locally analytically. This is useful for small azimuth and elevation ranges, such as the midpoint between adjacent defining points. This uses the aforementioned integral, but with a different range. and ,and Assume it to be a constant.

[0179] Experiments show that by direct summation, the error is small even when the directional resolution is quite coarse. Furthermore, the error is independent of the radius. For a linear interval of azimuth angles between 10 points, and for elevation points at 10 linear intervals, this results in a relative error of -20 dB.

[0180] The integral expressed above provides a result proportional to the radius of the sphere. Therefore, it is proportional to the reference distance. This dependence on the radius is because we haven't considered the reaction of the "distance gain" between two different radii. If the radius were doubled, the energy "flowing" through a fixed surface area (e.g., 1 cm²) would be 6 dB lower. Therefore, it could be said that the integral should take the distance gain into account. However, the integration is performed at a reference distance, defined as the distance that reflects the distance gain in the signal. In other words, the signal level indicated by the reference distance is not included as a scaling of the value being integrated, but rather reflected by the variation of the surface area on which the integration is performed with the reference distance (because the integration is performed on a sphere with a radius equal to the reference distance).

[0181] Therefore, the integral described above reflects the audio signal energy scaling factor (including any pre-gain or similar calibration adjustment), since the audio signal represents the correct signal playback energy at a fixed surface region on a sphere with a radius equal to the reference distance (without directional gain).

[0182] This means that if the reference distance is larger, the total signal energy scaling factor is also larger without changing the signal. This is because the corresponding signal represents a relatively louder sound source than a sound source with the same signal energy but at a smaller reference distance.

[0183] In other words, by performing integration on the surface of a sphere with a radius equal to the reference distance, the signal level indication provided by the reference distance is automatically taken into account. A higher reference distance will result in a larger surface area, and therefore a larger total transmit energy indication. Specifically, integration is performed directly at a distance where the range gain is 1.

[0184] The above integral produces a value normalized to the surface units used and the units used to indicate the reference distance r. If the reference distance r is expressed in meters, the result of the integral is in meters. 2 Provided to the organization.

[0185] To correlate the estimated emitted energy value with the signal, it should be expressed in surface units corresponding to the signal. Since the signal level represents the level that should be played for a user at a reference distance, the surface area of ​​the human ear may be more suitable. At the reference distance, this surface relative to the entire surface of a sphere will be associated with a portion of the source energy that will be perceived.

[0186] Therefore, the total emitted energy, representing the emitted source energy normalized to full-scale samples in an audio signal, can be indicated by the following formula:

[0187] in, The indicator is the energy determined by integrating the directional gain over the surface of a sphere with a radius equal to the reference distance. It is pre-gain, and It is a normalization scaling factor (to correlate the determined energy with the region of the human ear).

[0188] The corresponding reverberation energy can be calculated using the DSR, which characterizes the diffuse acoustic properties of the space, and the calculated source energy derived from metadata on directionality, pregain, and reference distance.

[0189] A DSR (Distributed Reverberation Scale) can typically be determined using the same reference level used by both of its components. This can be the same as or different from the total emitted energy indicator. In any case, when this DSR is combined with the total emitted energy indicator, the resulting reverberation energy, when using the total emitted energy determined by the integration described above, is also expressed as energy normalized to full-scale samples in the audio signal. In other words, all the energies considered are essentially normalized to the same reference level, allowing them to be directly combined without level adjustment. Specifically, the determined total emitted energy can be used directly with the DSR to generate a level indicator of the diffuse reverberation generated from each source, where the level indicator directly indicates the appropriate level of diffuse reverberation relative to other sound sources and relative to the individual path signal components.

[0190] As a specific example, the relative signal levels of diffuse reverberant signal components from different sources can be directly obtained by multiplying the DSR by the total emission energy indication.

[0191] In the described system, the adjustment of the contributions of different sound sources to the diffuse reverberation signal is performed at least in part by adjusting the downmixing coefficients used to generate the downmixing signal. Therefore, downmixing coefficients can be generated such that the relative contribution / energy level of the diffused sound from each sound source reflects the determined diffuse reverberation energy of the source.

[0192] As a specific example, if the DSR indicates an initial amplitude level, the downmixing factor can be determined to be proportional to (or equal to) the DSR multiplied by the total transmit energy indication. If the DSR indicates an energy level, the downmixing factor can be determined to be proportional to (or equal to) the square root of the DSR multiplied by the total transmit energy indication.

[0193] As a specific example, it is used to index one of multiple input signals. The signal provides an appropriately adjusted downmixing coefficient. It can be calculated using the following formula:

[0194] in, Refers to pregain, and Signal before pregain The normalized source energy. DSR represents the ratio of diffuse reverberation energy to source energy. Current mixing coefficient. Applied to input signals The resulting signal represents the signal as a function of filtering a reverberator with a reverberation response of unit energy. Direct path rendering and relative to other sources The direct path and diffuse reverberation energy are the signals. Provide the correct signal level for diffused reverberation energy.

[0195] Alternatively, the lower mixing coefficient It can be calculated using the following formula:

[0196] in, Referential signal The normalized emission source energy, and DSR represents the ratio of diffuse reverberation energy to the initial reverberation response amplitude. The current mixing coefficient... Applied to input signals At this point, the resulting signal represents the signal level corresponding to the initial level of the diffuse reverberation signal, and can be processed by a reverberator with a reverberation response starting at amplitude 1. Therefore, the output of the reverberator is relative to the signal... Direct path rendering and relative to other sources The direct path and diffuse reverberation energy are the signals. Provide the correct diffusion reverberation energy.

[0197] In many embodiments, the downmixing coefficients are partially determined by combining the DSR with a total emitted energy indication. Regardless of whether the DSR indicates the relationship between the total emitted energy and the diffuse reverberation energy or the initial amplitude of the diffuse reverberation response, further adjustments to the downmixing coefficients are often required to suit the specific reverberator algorithm used, which scales the signal so that the output of the reverberation processor reflects the desired energy or initial amplitude. For example, when the input level remains constant, the reflection density in the reverberation algorithm has a strong influence on the resulting reverberation energy. As another example, the initial amplitude of the reverberation algorithm may not be equal to the amplitude of its excitation. Therefore, algorithm-specific or algorithm and configuration-specific adjustments may be required. This can be included in the downmixing coefficients and is generally common to all sources. In some embodiments, these adjustments may be applied to downmixing or included in the reverberator algorithm.

[0198] Once the downmixing coefficients are generated, the downmixing processor 509 can generate the downmixed signal, for example, by direct weighted combination or summation.

[0199] The advantage of the described method is that it can use conventional reverb units. For example, reverb unit 407 can be implemented using a feedback delay network, such as in a standard Jot reverb unit.

[0200] like Figure 7As illustrated, the principle of a feedback delay network uses one or more (usually more than one) feedback loops with different delays. The input signal (in this case, the undermixed signal) is fed into the loops, where the signal is fed back with an appropriate feedback gain. The output signal is extracted by combining the signals in the loops. Thus, the signal is repeated continuously with different delays. Using coprime delays and having a feedback matrix that mixes the signal between loops, a pattern similar to reverberation in real space can be created.

[0201] The absolute values ​​of the elements in the feedback matrix must be less than 1 to achieve a stable decay impulse response. In many implementations, additional gain or filters are included in the loop. These filters control the decay instead of the matrix. Using filters can have different benefits for the decay response at different frequencies.

[0202] In some embodiments where the reverb output is binaurally rendered, the estimated reverb can be filtered separately by the average HRTF (Head-Related Transfer Function) of the left and right ears to produce left-channel and right-channel reverb signals. When the HRTF is available at uniformly spaced intervals on a sphere around the user for more than one distance, it is understood that the set of HRTFs with the maximum distance is used to generate the average HRTF for the left and right ears. Using the average HRTF can be based on / reflect the consideration that the reverb is isotropic and originates from all directions. Therefore, instead of including a pair of HRTFs for a given direction, the average of all HRTFs can be used. An averaging can be performed once for the left ear and once for the right ear, and the resulting filter can be used to process the reverb output for binaural rendering.

[0203] In some cases, the reverb itself can introduce coloration into the input signal, resulting in an output that does not have the desired output spread signal energy as described by DSR. Therefore, the effect of this process can also be equalized. This equalization can be performed based on a filter that is analytically determined to be the reciprocal of the frequency response of the reverb's operation. In some embodiments, machine estimation learning techniques such as linear regression, line fitting, etc., can be used to estimate the transfer function.

[0204] In some embodiments, the same method can be applied uniformly across the entire frequency band. However, in other embodiments, frequency-dependent processing can be performed. For example, one or more of the provided metadata parameters may be frequency-dependent. In such an example, the apparatus can be arranged to divide the signal into different frequency bands corresponding to frequency correlation, and the processing described above can be performed individually in each frequency band.

[0205] Specifically, in some embodiments, the diffuse reverberant signal to total signal ratio (DSR) is frequency-dependent. For example, different DSR values ​​can be provided for a series of discrete frequency bands / bands, or the DSR can be provided based on frequency. In such embodiments, the apparatus can be arranged to generate frequency-dependent downmixing coefficients that reflect the frequency dependence of the DSR. For example, downmixing coefficients for individual frequency bands can be generated. Similarly, frequency-dependent downmixed and diffuse reverberant signals can therefore be generated.

[0206] For frequency-dependent DSR, in other embodiments, the downmixing coefficients can be supplemented by filters that filter the audio signal as part of the downmixing generation. As another example, the DSR effect can be divided into frequency-independent (wideband) components and frequency-dependent components. The frequency-independent (wideband) components are used to generate frequency-independent downmixing coefficients, which are used to scale the individual audio signals when generating the downmixed signal. The frequency-dependent components can be applied to the downmixing, for example, by applying a frequency-dependent filter. In some embodiments, such filters can be combined with additional coloring filters, for example, as part of a reverb algorithm. Figure 7 The diagram illustrates the relationship between (u, v) and color (h). L , h R Example of a filter. This is a feedback delay network specifically designed for binaural output, called a Jot reverb.

[0207] Therefore, in some embodiments, the DSR may include a frequency-dependent component portion and a non-frequency-dependent component portion, and the coefficient processor 507 may be arranged to generate downmixing coefficients based on the non-frequency-dependent component portion (and independently of the frequency-dependent portion). The downmixing process can then be adjusted based on the frequency-dependent component portion; that is, the reverb unit can be adjusted according to the frequency-dependent portion.

[0208] In some embodiments, the directionality of sound radiation from one or more sound sources may be frequency-dependent, and in this case, energy processor 505 may be arranged to generate a frequency-dependent total emission energy that, when combined with a DSR (which may be frequency-dependent or independent), can produce a frequency-dependent downmixing coefficient.

[0209] This can be achieved, for example, by performing individual processing within a discrete frequency band. Directional frequency correlation, compared to frequency-dependent DSR processing, must typically be performed before (or as part of) generating the downmixed signal. This reflects the fact that frequency-dependent downmixing is usually required to include directional frequency-dependent effects, as these effects are typically different for different sources. After integration, it is possible that the net effect has a significant variation in frequency; that is, the total emission energy indication of a given source can have substantial frequency correlation, where this is different for different sources. Therefore, since different sources typically have different directional patterns, the total emission energy indication of different sources will typically also have different frequency correlations.

[0210] Specific examples of possible methods will be described below. Providing a DSR characterizing the diffuse acoustic properties of the space and determining the emission source energy based on directivity, pregain, and reference distance metadata allows for the calculation of the corresponding desired reverberation energy. For example, this can be determined as:

[0211] When the components used to calculate DSR use the same reference level (e.g., related to the full-scale of the signal), when using the same method as above for calculating the source energy... The resulting reverberation energy will also be the energy normalized for full-scale samples in the PCM signal, and thus corresponds to the energy of the impulse response (IR) of diffuse reverberation that can be applied to the corresponding input signal to provide the correct reverberation level in the signal representation used.

[0212] These energy values ​​can be used to determine the configuration parameters of the reverberation algorithm, downmixing coefficients, or downmixing filter before the reverberation algorithm is applied.

[0213] There are different ways to generate reverberation. Algorithms based on feedback delay networks (FDNs), such as the Jot reverberator, are suitable low-complexity methods. Alternatively, the noise sequence can be shaped to have appropriate (frequency-dependent) decay and spectral shape. In both examples, the prototype IR (with at least an appropriate T60) can be tuned so that its (frequency-dependent) level is corrected.

[0214] The reverb algorithms can be adjusted so that they produce impulse responses with unit energy (or the unit initial amplitude of the DSR can be correlated with the initial amplitude), or the reverb algorithms can include their own compensation, for example, in the coloring filter of the Jot reverb. Alternatively, the downmixing can be modified using (potentially frequency-dependent) tuning, or the downmixing coefficients generated by the coefficient processor 507 can be modified.

[0215] Compensation can be determined by generating an impulse response and measuring the energy of that IR without any such adjustment but with all other configurations applied, such as appropriate reverberation time (T60) and reflection density (e.g., the delay value in the FDN).

[0216] The compensation can be the reciprocal of the energy. To include it in the lower mixing coefficient, the square root is usually applied. For example:

[0217] In many other embodiments, compensation can be derived from configuration parameters. For example, the first reflection can be derived from its configuration when the DSR is relative to the initial reverberation amplitude. By definition, the correlation filter is energy-conserving, and the coloring filter can also be designed as such.

[0218] Assuming no net boost or attenuation from the coloring filter, the reverb can, for example, result in a value dependent on T60 and the minimum delay. initial amplitude .

[0219] Predicting reverberation energy can also be done heuristically.

[0220] As a general model for diffuse reverberation energy, an exponential function can be considered. :

[0221] for .in, It is a decay factor controlled by T60, and It is the amplitude at the pre-delay point.

[0222] Calculating the cumulative energy of a function like this will asymptotically approach some final energy value. The final energy value has an almost perfect linear relationship with T60.

[0223] The factor for a linear relationship depends on the sparsity of function A (setting every second value to 0 produces about half the energy) and the initial value. (Energy with) The diffusion tail can be reliably modeled using functions such as T60, reflection density (derived from the FDN delay), and sampling rate (linearly scaled with respect to fs). It can be calculated as equal to FDN as shown above. .

[0224] When generating multiple parametric reverberations with broadband T60 values ​​in the range of 0.1–2 s, the energy of the IR is nearly linear with the model. The scaling factor between the actual energy and the exponential equation model average is determined by the sparsity of the FDN response. This sparsity decreases towards the end of the IR but has the greatest effect at the beginning. Based on testing the above with various configurations of delay values, an almost linear relationship has been found between the model reduction factor and the minimum difference between the delays configured in the FDN.

[0225] For example, for a specific implementation of the Jot reverb, this can be equivalent to a scaling factor calculated by the following formula. :

[0226] The energy of the model is calculated by integrating from t=0 to infinity. This can be done analytically and results in:

[0227] Combining the above information, we obtain the following prediction for reverberation energy.

[0228] It will be appreciated that, for clarity, the above description has referenced various functional circuits, units, and processors to describe embodiments of the invention. However, it will be apparent that any suitable distribution of functionality among the different functional circuits, units, or processors can be used without diminishing the invention. For example, functionality illustrated as being performed by separate processors or controllers may be performed by the same processor or controller. Therefore, references to specific functional units or circuits should be considered only as references to suitable means for providing the described functionality and not as indications of a strict logical or physical structure or organization.

[0229] This invention can be embodied in any suitable form, including hardware, software, firmware, or any combination thereof. The invention can optionally be implemented, at least in part, as computer software running on one or more data processors and / or digital signal processors. Elements and components of embodiments of the invention can be implemented physically, functionally, and logically in any suitable manner. In practice, functionality can be implemented in a single unit, in multiple units, or as part of other functional units. Therefore, the invention can be implemented in a single unit or can be physically and functionally distributed among different units, circuits, and processors.

[0230] While the invention has been described in conjunction with some embodiments, it is not intended to be limited to the specific forms set forth herein. Rather, the scope of the invention is limited only by the accompanying drawings. Furthermore, although features may appear to be described in conjunction with specific embodiments, those skilled in the art will recognize that various features of the described embodiments can be combined according to the invention. In the claims, terms include those that do not exclude the presence of other elements or steps.

[0231] Furthermore, although listed individually, multiple means, elements, circuits, or method steps may be implemented, for example, by a single circuit, unit, or processor. Moreover, while individual features may be included in different claims, these may be advantageously combined, and inclusion in different claims does not imply that such a combination of features is not feasible and / or advantageous. Furthermore, inclusion of features in a claim of one class does not imply a limitation on that class, but rather indicates that the features are equally applicable, as appropriate, to other claim classes. Additionally, the order of features in a claim does not imply any particular order in which the features must operate, and in particular, the order of individual steps in a method claim does not imply that the steps must be performed in that order. Rather, the steps may be performed in any suitable order. Additionally, singular references do not exclude plural. Therefore, references to “a,” “an,” “first,” “second,” etc., do not exclude plural. Reference numerals in the claims are provided only as clarifying examples and should not be construed as limiting the scope of the claims in any way.

Claims

1. An audio device for generating a binaural signal comprising a diffuse reverberant signal for an environment; The audio apparatus comprises: a receiver (501) arranged to receive a plurality of audio signals representing sound sources in the environment; a metadata receiver (501) arranged to receive metadata for the plurality of audio signals, the metadata comprising: a diffuse reverb signal to total signal relationship measure indicative of a level of diffuse reverb sound relative to total emitted sound in the environment, and for each audio signal: a signal level indication; directionality data indicative of a directionality of sound radiation from the sound source represented by the audio signal; circuitry (505, 507) arranged to determine, for each audio signal of the plurality of audio signals, a total emitted energy indication based on the signal level indication and the directionality data; a downmixer (509) arranged to generate a downmix signal by combining signal components of each audio signal; a reverb (407) for generating, from downmix signal components, the diffuse reverb signal for the environment based on the diffuse reverb signal to total signal relationship measure; a direct path rendering circuit for generating, in response to the signal level indication and the directionality data for a first audio signal of the plurality of audio signals or in response to the total emitted energy indication, a direct path audio signal for the first audio signal; a combiner for generating, from the direct path audio signal, a binaural output signal to include left and right channel reverb signals produced from the diffuse reverb signal.

2. The audio apparatus of claim 1, wherein, The directionality of sound radiation is frequency dependent, and the circuitry is arranged to determine a frequency dependent total emitted energy.

3. The audio apparatus of any of the preceding claims, wherein, The diffuse reverb signal to total signal relationship is frequency dependent.

4. The audio apparatus of any of the preceding claims, wherein, The diffuse reverb signal to total signal relationship comprises a frequency dependent part and a non-frequency dependent part.

5. The audio apparatus of any of the preceding claims, wherein, The circuitry is arranged to determine the total emitted energy indication for a first audio signal of the plurality of audio signals in response to scaling the signal level indication for the first audio signal with a value determined by integrating a directivity pattern of the sound source represented by the first audio signal; the directivity pattern being determined based on directionality data.

6. The audio apparatus of any one of the preceding claims, wherein, The signal level indication for a first audio signal of the plurality of audio signals comprises a reference distance indicative of a distance from the audio source represented by the first audio signal at which a distance reference gain is applied.

7. The audio device of claim 6 when dependent on claim 5, wherein, The integrating is performed for a distance from the audio source represented by the first audio signal that is the reference distance.

8. The audio apparatus of any of the preceding claims, wherein, The diffuse reverb signal to total signal relationship is indicative of an energy of diffuse reverb sound relative to an energy of total emitted sound in the environment.

9. The audio apparatus of any one of claims 1-7, wherein, The diffuse reverb signal to total signal relationship is indicative of an initial amplitude of diffuse sound relative to an energy of total emitted sound in the environment.

10. The audio apparatus of any of the preceding claims, wherein, The signal level indication for a first audio signal of the plurality of audio signals further comprises a gain indication for the first audio signal, the gain indication indicating a gain to be applied to the first audio signal when rendering sound from a first audio source represented by the first audio signal.

11. The audio apparatus of any of the preceding claims, wherein, The metadata further comprises a delay indication, and the diffuse-reverberation signal-to-total-signal relation indicates an energy of diffuse-reverberant sound having a longer delay relative to an energy of total emitted sound in the environment than the delay indication.

12. A method of generating a binaural signal comprising a diffuse-reverberation signal for an environment, the method comprising: receiving a plurality of audio signals representing sound sources in the environment; receiving metadata for the plurality of audio signals, the metadata comprising: a diffuse-reverberation signal-to-total-signal relation measure indicating a level of diffuse-reverberant sound relative to total emitted sound in the environment, and for each audio signal: a signal level indication; directionality data indicating a directionality of sound radiation from the sound source represented by the audio signal; determining a total emitted energy indication for each audio signal of the plurality of audio signals based on the signal level indication and the directionality data, and generating a downmix signal by combining signal components of each audio signal; generating the diffuse-reverberation signal for the environment from downmix signal components based on the diffuse-reverberation signal-to-total-signal relation measure; generating a direct-path audio signal for a first audio signal of the plurality of audio signals in response to the signal level indication and the directionality data for the first audio signal or in response to the total emitted energy indication; generating the binaural output signal from the direct-path audio signal to comprise left and right channel reverberation signals generated from the diffuse-reverberation signal.

13. A computer program comprising program code means adapted to perform all the steps of claim 12 when the program is run on a computer.

Citation Information

Patent Citations

  • Generating binaural audio in response to multi-channel audio using at least one feedback delay network

    EP3402222A1