Device and method for generating diffuse reverberant signal
The audio apparatus and method efficiently generate diffuse reverberant signals by using downmix coefficients and reverberators to improve audio quality and adaptability in virtual reality applications, addressing the limitations of existing techniques.
Patent Information
- Application Number
- JP2025153751
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-06-22
- Filing Date
- 2025-09-17
- Publication Date
- 2025-12-03
AI Technical Summary
Current approaches for generating diffuse reverberant signals in virtual reality applications are suboptimal, insufficient, and incomplete, particularly in terms of flexibility, computational efficiency, and audio quality, with a lack of methods for defining specific parameters and rendering techniques.
An audio apparatus and method that generates diffuse reverberation signals by receiving audio signals and metadata, determining downmix coefficients based on diffuse reverberation-to-total signal ratios and directivity data, and combining these signals to create a downmix signal, which is then processed by a reverberator to generate a diffuse reverberation signal.
This approach allows for the efficient generation of natural-sounding diffuse reverberant signals with low complexity, improved audio quality, and adaptability to changing positions, suitable for dynamic virtual and augmented reality applications.
Smart Images

Figure 2025176160000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an apparatus and method for processing audio data, particularly but not exclusively to processing for generating diffuse reverberant signals for augmented / mixed / virtual reality applications. [Background technology]
[0002] In recent years, the variety and range of experiences based on audiovisual content has expanded significantly, and new services and methods for using and consuming such content are continually being developed and introduced. In particular, many spatial and interactive services, applications, and experiences have been developed to provide users with more immersive and engaging experiences.
[0003] Examples of such applications include virtual reality (VR), augmented reality (AR), and mixed reality (MR) applications, which are rapidly becoming mainstream, with many solutions aimed at the consumer market. Additionally, many standards are being developed by various standards organizations. Such standardization efforts are actively developing standards for various aspects of VR / AR / MR systems, including, for example, streaming, broadcasting, rendering, etc.
[0004] VR applications tend to provide a user experience that corresponds to the user being in a different world / environment / scene, while AR (including Mixed Reality MR) applications tend to provide a user experience that corresponds to the user in their current environment, but with the addition of additional information or virtual objects or information. Thus, VR applications tend to provide fully immersive synthetically generated worlds / scenes, while AR applications tend to provide partially synthetic worlds / scenes that are overlaid on the real scene in which the user is physically present. However, the terms are often used interchangeably and largely overlap. In the following, the term virtual reality / VR will be used to refer to both virtual reality and augmented / mixed reality.
[0005] As an example, an increasingly popular service is for a user to actively and dynamically interact with the system to change the parameters of the rendering, providing images and sound in a manner that can adapt to movement and changes in the user's position and orientation. A very attractive feature in many applications is the ability to change the viewer's effective viewing position and direction, for example, to allow the viewer to move and "look around" in the scene being presented.
[0006] Such functionality specifically allows providing a virtual reality experience to the user, whereby the user can move (relatively) freely around the virtual environment and dynamically change his or her position and where he or she is looking. Typically, such virtual reality applications are based on a three-dimensional model of the scene, which is dynamically evaluated to provide a specific requested view. This approach is well known from computer and console gaming applications, e.g., in the first-person shooter category.
[0007] Also, particularly in virtual reality applications, it is desirable that the images presented are three-dimensional images, typically presented using a stereoscopic display. Indeed, to optimize the viewer's sense of immersion, it is typically preferred for the user to experience the presented scene as a three-dimensional scene. Indeed, a virtual reality experience should preferably allow the user to choose their position, viewpoint, and moment relative to the virtual world.
[0008] In addition to the visual rendering, most VR / AR applications also provide a corresponding audio experience. In many applications, it is preferable to provide a spatial audio experience, where audio sources are perceived as arriving from positions corresponding to the positions of corresponding objects in the visual scene. Thus, the audio scene and the video scene are preferably perceived consistently, providing a complete spatial experience together.
[0009] For example, many immersive experiences are provided by virtual audio scenes generated by headphone playback using binaural audio rendering techniques. In many scenarios, such headphone playback is based on head tracking so that the rendering responds to the user's head movements, which greatly enhances the sense of immersion.
[0010] An important feature for many applications is a way to generate and / or distribute sound that can provide a natural and realistic perception of the sound environment. For example, when generating sound for virtual reality applications, it is important not only to generate the desired sound sources, but also to modify these sound sources to provide a realistic perception of the sound environment, including attenuation, reflections, coloration, etc.
[0011] In the case of room acoustics, or more generally environmental acoustics, reflections of sound waves from the walls, floor, ceiling, objects, etc. of the environment cause delayed and attenuated (usually frequency-dependent) versions of the source signal to reach the listener (i.e., the user of a VR / AR system) via different paths. The combined effect can be modeled by an impulse response, hereafter referred to as the Room Impulse Response (RIR) (although this term suggests the specific application of an acoustic environment in the form of a room, it tends to be used more generally with regard to acoustic environments, whether or not they correspond to a room).
[0012] As illustrated in Figure 1, a room impulse response typically consists of a direct sound component, which depends on the distance from the source to the listener, followed by a reverberant component that characterizes the acoustics of the room. The size and shape of the room, the position of the source and listener within the room, and the reflective properties of the room's surfaces all play a role in characterizing this reverberant component.
[0013] The reverberant part can be divided into two time domains that usually overlap. The first domain contains the so-called early reflections, which represent isolated reflections of the sound source off the walls and obstacles in the room before reaching the listener. As the time lag increases, the number of reflections that exist within a certain time interval increases, and the path may contain more than one order of reflections (for example, if the reflection is off multiple walls, or both a wall and the ceiling).
[0014] The second region in the reverberant section is where the density of these reflections increases to the point where they can no longer be separated by the human brain. This region is usually called the diffuse reverberation, late reverberation, or reverberation tail.
[0015] The reverberant part contains clues that give the auditory system information about the distance of a sound source and the size and acoustic properties of the room. The energy in the reverberant part relative to the energy in the anechoic part largely determines the perceived distance of a sound source. The level and delay of the earliest reflections provide clues about how close a sound source is to a wall, and anthropometric filtering enhances the assessment of a particular wall, floor, or ceiling.
[0016] The density of (early) reflections affects the perceived size of the room. Reverberation time T 60 The time it takes for the energy level of a reflection to drop by 60 dB, denoted by , is often used as a measure of how quickly reflections dissipate in a room. Reverberation time provides information about the acoustic properties of a room, specifically whether the walls are highly reflective (e.g., a bathroom) or highly sound absorbing (e.g., a bedroom with furniture, carpeting, and curtains).
[0017] Furthermore, the RIR is filtered by the head, ears, and shoulders, i.e., the RIR is a head-related impulse response (HRIR), and therefore depends on the anthropometric characteristics of the user when it is part of the binaural room impulse response (BRIR).
[0018] Late reverberant reflections cannot be distinguished or separated by the listener and are therefore often simulated and rendered parametrically using parametric reverberators that use feedback delay networks, such as the well-known Jot reverberator.
[0019] For early reflections, the delay, which depends on the incidence direction and distance, is an important clue for humans to extract information about the room and the relative position of the sound source. Therefore, the simulation of early reflections needs to be more specific than late reverberation. Therefore, in efficient acoustic rendering algorithms, early reflections are simulated differently from late reverberation. A well-known method for early reflections is to mirror the sound source at each room boundary and generate virtual sources to represent the reflections.
[0020] While early reflections involve the position of the user and / or sound source relative to the room boundaries (walls, ceiling, floor), late reverberation tends to be more uniform throughout the room because the acoustic response of the room is diffuse, which makes simulating late reverberation often more computationally efficient than early reflections.
[0021] Two main characteristics of the late reverberation defined by a room are the T60 value and the reverberation level. For a diffuse reverberation impulse response, these values describe the slope and amplitude of the impulse response, both of which are usually highly frequency dependent in natural rooms.
[0022] The T60 parameter is important in giving the impression of the reflectivity and size of the room, while the reverberation level indicates the combined effect of multiple reflections at the room boundaries. The reverberation level and its frequency behavior depend on the pre-delay and indicate where the distinction between early and late reflections is made (see Figure 2).
[0023] Reverberation level has a primarily psychoacoustic relevance in relation to the direct sound. The level difference between the two is an indication of the distance between the sound source and the user (or the RIR measurement point). As the distance increases, the direct sound attenuates more, but the level of late reverberation remains the same (it is the same throughout the room). Similarly, for sound sources with directionality that depends on where the user is relative to the source, as the user moves around the source, the directionality affects the direct response but not the level of reverberation.
[0024] An important challenge and consideration for many systems, such as virtual reality applications, is how to efficiently represent and distribute an audio environment. Often, the sounds of an environment are represented and distributed by providing signals representing individual audio source signals along with data that parametrically describe the characteristics of the audio sources and the acoustic environment. This challenge is not trivial, and a variety of problems are possible.
[0025] Separate descriptions of direct path and diffuse reverberation have been proposed, however the question of how to represent, distribute and render / synthesize diffuse reverberation is currently of great interest.
[0026] It has been proposed to provide an indication of reverberation level not related to the direct sound but by more general characteristics. A specific proposal (section 3.9 of MPEG output document N19211, "MPEG-I 6DoF Audio Encoder Input Format", MPEG 130) was made as part of the preparation of the MPEG-I Audio Call for Proposals (CfP), in which the Encoder Input Format (EIF) is defined. The EIF defines the reverberation level in terms of the pre-delay and direct-diffuse ratio (DDR). The DDR is defined as the ratio between the diffuse reverberation energy after the pre-delay and the radiated sound source energy.
number
[0027] However, while such parameters are useful, there are many substantive issues that need to be addressed. For example, there are currently no proposals on how to define or determine specific parameters. Also, there is no consideration of how DDR metrics are used to render audio, and specifically the methods used to generate diffuse reverberant signals.
[0028] EP3402222 discloses a virtualization method for generating binaural signals according to the channels of a multi-channel audio signal, which applies a binaural room impulse response (BRIR) to each channel, including applying a common late reverberation to a downmix of the channels by using at least one feedback delay network (FDN). Summary of the Invention [Problem to be solved by the invention]
[0029] Thus, current approaches and proposals on how to represent and generate sound, and specifically diffuse reverberation, tend to be suboptimal, insufficient, and / or incomplete, especially in virtual reality applications, for example, where the location from which sound is generated can vary significantly.
[0030] Therefore, approaches for generating diffuse reverberant signals would be advantageous, particularly approaches that allow for improved operation, increased flexibility, reduced complexity, easier implementation, improved audio experience, improved audio quality, reduced computational load, improved adaptability to varying locations, improved performance in virtual / mixed / augmented reality applications, improved perceptual cues of diffuse reverberation, and / or improved performance and / or operation.
[0031] SUMMARY OF THE INVENTION Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination. [Means for solving the problem]
[0032] According to an aspect of the present invention, there is provided an audio apparatus for generating a diffuse reverberation signal of an environment, the apparatus comprising: a receiver configured to receive a plurality of audio signals representing sound sources in the environment; a metadata receiver configured to receive metadata of the plurality of audio signals, the metadata including a diffuse reverberation signal to total signal relationship indicating the level of diffuse reverberation sound relative to total sound radiated in the environment, a signal level indicator and directivity data for each audio signal indicating the directionality of sound radiation from the sound source represented by the audio signal; circuitry configured to determine, for each of the plurality of audio signals, a total radiated energy indicator based on the signal level indicators and the directivity data, and a downmix coefficient based on the total radiated energy and the diffuse reverberation signal to total signal relationship; a downmixer configured to generate a downmix signal by combining signal components of each audio signal generated by applying the downmix coefficient of each audio signal to the audio signals; and a reverberator for generating the diffuse reverberation signal of the environment from the downmix signal components.
[0033] The present invention, in many embodiments, improves and / or facilitates the determination of diffuse reverberant signals. The present invention generates more natural-sounding diffuse reverberant signals that, in many embodiments and scenarios, provide an improved perception of the acoustic environment. The generation of diffuse reverberant signals often has low complexity and low computational resource requirements. This approach allows the diffuse reverberant sound in an acoustic environment to be effectively represented with a relatively small number of parameters, which also provides an efficient representation of individual sound sources and the sound propagation of individual paths from these, particularly direct path propagation.
[0034] This approach, in many embodiments, allows for the generation of diffuse reverberation signals that are independent of the position of the sound source and / or listener, thereby enabling efficient generation of diffuse reverberation signals for dynamic applications where positions change, such as many virtual and augmented reality applications.
[0035] The diffuse reverberation signal to total signal ratio is also referred to as the diffuse reverberation signal level to total signal level ratio, or the diffuse reverberation level to total level ratio, or the radiated sound source energy to diffuse reverberation energy ratio (or variations / permutations thereof).
[0036] The audio device may be implemented in a single device or single functional unit, or may be distributed across different devices or functions, for example the audio device may be implemented as part of a decoder functional unit, or may be distributed such that some functional elements are performed on the decoder side and other elements are performed on the encoder side.
[0037] According to an optional feature of the invention, the directionality of the sound radiation is frequency dependent and the circuitry is configured to generate a frequency dependent total radiation energy and a frequency dependent downmix coefficient.
[0038] This approach provides a particularly efficient operation for generating a diffuse reverberant signal that reflects frequency dependence.
[0039] According to an optional feature of the invention, the relationship between the diffuse reverberant signal and the total signal is frequency dependent, and the circuitry is configured to generate frequency dependent downmix coefficients.
[0040] This approach provides a particularly efficient operation for generating a frequency-dependent diffuse reverberant signal that reflects frequency dependence.
[0041] According to an optional feature of the invention, the relation of the diffuse reverberant signal to the total signal comprises a frequency dependent part and a non-frequency dependent part, and the circuit is configured to depend on the non-frequency dependent part to generate the downmix coefficients and to depend on the frequency dependent part to adapt the reverberator.
[0042] This approach provides a particularly efficient operation for generating a diffuse reverberant signal that reflects frequency dependence, in particular reducing complexity and / or resource usage, for example allowing the frequency dependence to be reflected by a single filtering of the downmix signal.
[0043] According to an optional feature of the invention, the circuitry is configured to determine a total radiant energy indicator for the first audio signal in response to scaling a signal level indicator for the first audio signal by a value determined by integrating a directivity pattern of a sound source represented by the first audio signal of the plurality of audio signals.
[0044] This provides a particularly advantageous operation in many embodiments. The scaling is any function applied to the signal level measure in connection with determining the downmix coefficients. This function typically increases monotonically as a function of the total radiant energy measure. The scaling may be linear or non-linear.
[0045] The scaling does not depend on the time variations of the signal, so it does not need to be updated at the instantaneous level of the audio signal, but only needs to be recalculated if the signal level indicator or directional pattern changes.
[0046] According to an optional feature of the invention, the signal level indicator of a first audio signal of the plurality of audio signals includes a reference distance, the reference distance indicating a distance from an audio source represented by the first audio signal with respect to a distance reference gain for the first audio signal.
[0047] This provides particularly advantageous operation in many embodiments. The distance reference gain is a predetermined value, typically common to at least some, and often all, audio sources and signals. In many embodiments, the distance reference gain is 0 dB.
[0048] According to an optional feature of the invention, the integration is performed over a distance that is a reference distance from the audio source represented by the first audio signal.
[0049] This provides a particularly efficient approach and is easy to operate.
[0050] According to an optional feature of the invention, the relationship of the diffuse reverberant signal to the total signal indicates the energy of the diffuse reverberant sound relative to the energy of the total radiated sound in the environment.
[0051] This provides particularly advantageous operation in many embodiments.
[0052] According to an optional feature of the invention, the diffuse signal versus total signal relationship indicates the initial amplitude of the diffuse sound relative to the energy of the total radiated sound in the environment.
[0053] This provides particularly advantageous operation in many embodiments.
[0054] According to an optional feature of the invention, the downmix coefficients determined for a first audio signal of the plurality of audio signals are independent of the position of the first audio source represented by the first audio signal.
[0055] This provides particularly advantageous operation in many embodiments, particularly facilitating operation in dynamic applications where the position of the sound source changes, such as virtual reality applications.
[0056] According to an optional feature of the invention, the downmix coefficients determined for a first audio signal of the plurality of audio signals are independent of the position of the listener.
[0057] This provides particularly advantageous operation in many embodiments, particularly facilitating operation for dynamic applications where positions change, such as virtual reality applications.
[0058] In some embodiments, the processing of the audio device is independent of the position of the sound source.In some embodiments, the processing of the audio device is independent of the position of the listener.
[0059] In some embodiments, the processing of the audio device does not depend solely on the position of the listener within the area to which the diffuse signal to total signal ratio is applied.
[0060] In some embodiments, the update rate of the downmix coefficients is slower than the update rate of the position of the first audio source represented by the first audio signal. In some embodiments, the update rate of the downmix coefficients is slower than the update rate of the listener position. The downmix coefficients are calculated at a time rate that is much slower than the update rate of the listener position / audio source position.
[0061] According to an optional feature of the invention, the signal level indicator for a first audio signal of the plurality of audio signals further includes a gain indicator for the first audio signal, the gain indicator indicating a gain to apply to the first audio signal when rendering sound from a first audio source represented by the first audio signal, and the circuitry is configured to determine a downmix coefficient for the first audio signal in response to the gain indicator.
[0062] According to an optional feature of the invention, the audio device further comprises a direct rendering circuit configured to generate a direct path audio signal for a first audio signal of the plurality of audio signals in response to the signal level indicator and the directivity data of the first audio signal.
[0063] This provides particularly advantageous operation in many embodiments.
[0064] According to an optional feature of the invention, the metadata further includes a delay index, the diffuse signal to total signal ratio (DSR) being indicative of the energy of the diffuse reverberant sound in the environment having a delay longer than that indicated by the delay index relative to the energy of the total radiated sound.
[0065] The energy of diffuse reverberant sound in an environment with a delay longer than the delay index is reflected by the contribution of a room impulse response, or determined as the contribution of a room impulse, that occurs at least a certain delay after the radiation of the corresponding sound at the sound source, the certain delay being indicated by the delay index.
[0066] In some embodiments, the diffuse signal to total signal ratio (DSR) indicates the energy of the diffuse reverberant sound relative to the energy of the total radiated sound in the environment, where the energy of the diffuse reverberant sound is determined by the room response contribution occurring at least a certain delay after the radiation of the corresponding sound at the sound source.
[0067] According to another aspect of the present invention, there is provided a method for generating a diffuse reverberation signal of an environment, the method comprising: receiving a plurality of audio signals representing sound sources in the environment; receiving metadata for the plurality of audio signals, the metadata including a diffuse reverberation signal to total signal relationship indicating the level of diffuse reverberation sound relative to total sound radiated in the environment, and, for each audio signal, a signal level indicator and directivity data indicating the directionality of sound radiation from the sound source represented by the audio signal; determining, for each of the plurality of audio signals, a total radiated energy indicator based on the signal level indicator and the directivity data and a downmix coefficient based on the total radiated energy and the diffuse reverberation signal to total signal relationship; generating a downmix signal by combining signal components of each audio signal generated by applying the downmix coefficient of each audio signal to the audio signals; and generating a diffuse reverberation signal of the environment from the downmix signal components.
[0068] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter.
[0069] Embodiments of the present invention will now be described, by way of example only, with reference to the drawings in which: [Brief explanation of the drawings]
[0070] [Figure 1] FIG. 10 is a diagram illustrating an example of a room impulse response. [Figure 2] FIG. 10 is a diagram illustrating an example of a room impulse response. [Figure 3] FIG. 1 illustrates an example of elements of a virtual reality system. [Figure 4]FIG. 1 illustrates an example of an audio device for generating audio output, according to some embodiments of the present invention. [Figure 5] FIG. 1 illustrates an example of an audio reverberator for generating a diffuse reverberant signal in accordance with some embodiments of the present invention. [Figure 6] FIG. 10 is a diagram illustrating an example of a room impulse response. [Figure 7] FIG. 1 illustrates an example of a reverberator. DETAILED DESCRIPTION OF THE INVENTION
[0071] Although the following description focuses on audio processing and generation for virtual reality applications, it should be understood that the principles and concepts described can be used in many other applications and embodiments.
[0072] Virtual experiences that allow users to move around in virtual worlds are becoming increasingly popular, and services are being developed to meet this demand.
[0073] In some systems, VR applications are provided locally to the viewer by a standalone device that does not use or even have access to remote VR data or processing, e.g., a device such as a game console that includes a store for storing scene data, an input for receiving / generating viewer poses, and a processor for generating corresponding images from the scene data.
[0074] In other systems, the VR application is implemented and executed remotely from the viewer. For example, a device local to the user detects / receives motion / pose data, which is transmitted to a remote device that processes the data to generate a pose for the viewer. The remote device then generates an appropriate view image and corresponding audio signal appropriate for the user pose based on scene data describing the scene. The view image and corresponding audio signal are then transmitted to a device local to the viewer to be presented. For example, the remote device directly generates a video stream (typically a stereoscopic / 3D video stream) and corresponding audio stream that are directly presented by the local device. Thus, in such an example, the local device does not perform VR processing except for transmitting motion data and presenting the received video data.
[0075] In many systems, functionality is distributed between a local device and a remote device. For example, the local device processes received input and sensor data to generate a user pose that is continuously transmitted to the remote VR device. The remote VR device then generates a corresponding view image and a corresponding audio signal and transmits them to the local device for presentation. In other systems, the remote VR device does not directly generate the view image and the corresponding audio signal, but selects relevant scene data and transmits it to the local device, which generates the presented view image and the corresponding audio signal. For example, the remote VR device identifies the nearest capture point, extracts the corresponding scene data (e.g., a set of object sources and their position metadata), and transmits it to the local device. The local device then processes the received scene data to generate an image and audio signal for a particular current user pose. A user pose typically corresponds to a head pose, and references to a user pose typically correspond to references to a head pose.
[0076] In many applications, particularly for broadcast services, an audio source transmits or streams scene data in the form of images (including video) and audio representations of a scene that are independent of user pose. For example, signals and metadata corresponding to audio sources within a particular virtual room are transmitted or streamed to multiple clients. Each client then locally synthesizes an audio signal corresponding to the current user pose. Similarly, the audio source transmits a general description of the audio environment, including a description of the audio sources within the environment and the acoustic properties of the environment. An audio representation is then generated locally and presented to the user, for example using binaural rendering and processing.
[0077] 3 shows such an example of a VR system in which remote VR client devices 301 interact with a VR server 303 over a network 305, such as the Internet. The server 303 is configured to support a potentially large number of client devices 301 simultaneously.
[0078] The VR Server 303 supports the broadcast experience by, for example, transmitting image signals containing image representations in the form of image data that are used by the client devices to locally synthesize view images corresponding to the appropriate user pose (pose refers to position and / or orientation). Similarly, the VR Server 303 can transmit an audio representation of the scene to locally synthesize audio for the user pose. Specifically, as the user moves around in the virtual environment, the images and audio synthesized and presented to the user are updated to reflect the user's current (virtual) position and orientation within the (virtual) environment.
[0079] Therefore, in many applications, such as that of Figure 3, it is desirable to model a scene and generate efficient image and audio representations that can be efficiently included in a data signal that is transmitted or streamed to various devices that can locally synthesize views and audio for poses different from the capture pose.
[0080] In some embodiments, models representing the scene are stored locally and used locally to synthesize appropriate images and sounds, for example. For example, an audio model of a room includes indications of the acoustic properties of the room as well as the properties of sound sources that can be heard in the room. The model data is then used to synthesize sounds appropriate for the particular location.
[0081] How a sound scene is represented and how this representation is used to generate sound are important issues. Sound rendering that aims to provide a natural and realistic effect to the listener typically involves rendering of the acoustic environment. For many environments, this includes representing and rendering the diffuse reverberation present in the environment, such as a room. The rendering and representation of such diffuse reverberation is known to have a significant effect on the perception of the environment, including whether the sound is perceived as representing a natural and realistic environment. In the following, an advantageous approach is described for representing a sound scene and rendering sound, particularly diffuse reverberant sound, based on this representation.
[0082] This approach will be described with reference to an audio device as illustrated in Figure 4. The audio device is configured to generate an audio output signal representing sound in an acoustic environment. Specifically, the audio device generates sound representing the sound perceived by a user moving around in a virtual environment having several audio sources and given acoustic characteristics. Each audio source is represented by an audio signal representing the sound from the audio source and metadata describing the characteristics of the audio source (such as providing a level indicator of the audio signal). Additionally, metadata characterizing the acoustic environment is provided.
[0083] The audio device includes a path renderer 401 for each audio source. Each path renderer 401 is configured to generate direct path signal components representing the direct path from the audio source to the listener. The direct path signal components are generated based on the positions of the listener and audio source, specifically by scaling potentially frequency-dependent audio signals for distance-dependent audio sources and relative gains, for example, for audio sources in specific directions relative to the user (e.g., non-omnidirectional audio sources).
[0084] In many embodiments, the renderer 401 also generates a direct path signal based on occlusion or diffraction (virtual) elements between the source position and the user position.
[0085] In many embodiments, the path renderer 401 generates additional signal components for each path that includes one or more reflections, for example, by evaluating reflections off walls, ceilings, etc., as known to those skilled in the art. The direct path components and reflected path components are combined into a single output signal for each path renderer, thus generating a single signal representing the direct path reflections and early / discrete reflections for each audio source.
[0086] In some embodiments, the output audio signal of each audio source is a binaural signal, so that each output signal includes both a left-ear and a right-ear (sub) signal.
[0087] The output signals from the path renderers 401 are provided to a combiner 403, which combines the signals from the different path renderers 401 to generate a single combined signal. In many embodiments, a binaural output signal is generated and the combiner performs a combination, such as a weighted combination, of the individual signals from the path renderers 401, i.e., all right ear signals from the path renderers 401 are added together to generate a combined right ear signal, and all left ear signals from the path renderers 401 are added together to generate a combined left ear signal.
[0088] The path renderers and combiners are typically implemented in any suitable manner, including executable code for processing on a suitable computational resource, such as a microcontroller, microprocessor, digital signal processor, or central processing unit including supporting circuitry such as memory. It will be appreciated that multiple path renderers may be implemented as parallel functional units, such as, for example, a bank of dedicated processing units, or as repeated operations for each audio source. Typically, the same algorithm / code is executed for each audio source / signal.
[0089] In addition to the individual path audio components, the audio device is further configured to generate signal components representing the diffuse reverberation in the environment. The diffuse reverberation signal is (effectively) generated by combining the audio source signal with a downmix signal and then applying a reverberation algorithm to the downmix signal to generate the diffuse reverberation signal.
[0090] The audio device of FIG. 4 includes a downmixer 405 that receives audio signals from multiple sound sources (typically all sound sources in the acoustic environment for which the reverberator is simulating diffuse reverberation) and combines them into a downmix. The downmix thus reflects all sounds generated in the environment. The downmix is provided to a reverberator 407 that is configured to generate a diffuse reverberation signal based on the downmix. The reverberator 407 is specifically a parametric reverberator, such as a Jot reverberator. The reverberator 407 is coupled to a combiner 403 that is provided with the diffuse reverberation signal. The combiner 403 then combines the diffuse reverberation signal with the path signals representing the individual paths to generate a combined audio signal that represents the combined sound in the environment as perceived by the listener.
[0091] The generation of a diffuse reverberant signal will be further described with reference to an audio reverberator as illustrated in Figure 5. The audio reverberator is included in the audio device of Figure 4 and specifically implements the downmixer 405 and the reverberator 407.
[0092] The audio reverberator comprises a receiver 501 configured to receive audio scene data representing audio. The audio scene data specifically comprises a plurality of audio signals, each of which represents one audio source (hence the audio signals describe the sound from the audio source). In addition, the receiver 501 receives metadata for each of the audio sources. This metadata includes (relative) signal level indicators of the audio sources, indicating the level / energy / amplitude of the audio source represented by the audio signal. The audio source metadata further includes directivity data, which indicates the directionality of sound radiation from the audio source. The directivity data of an audio signal describes, for example, a gain pattern, specifically the relative gain / energy density of the audio source in different directions from the position of the audio source.
[0093] The receiver 501 further receives metadata indicative of the acoustic environment. In particular, the receiver 501 receives a diffuse reverberation signal to total signal ratio (also referred to as the diffuse reverberation signal level to total signal level ratio, or in some cases the diffuse reverberation signal level to total signal energy ratio, or the radiated energy to diffuse reverberation energy ratio), which indicates the relationship between the diffuse reverberation signal and the total signal, in particular the level of the diffuse reverberation sound relative to the total radiated sound in the acoustic environment. For simplicity, the diffuse reverberation signal to total signal ratio is also referred to below as the diffuse-to-source ratio DSR, or equivalently, the source-to-diffusion ratio SDR (the former will be mainly used in the following description).
[0094] It should be understood that ratios and inverse ratios provide the same information, i.e., any ratio can be expressed as an inverse ratio. Thus, the relationship of the diffuse reverberant signal to the total signal is expressed by the fraction of a value reflecting the level of diffuse reverberant sound divided by a value reflecting total sound radiation, or equivalently, by the fraction of a value reflecting total sound radiation divided by a value reflecting the level of diffuse reverberant sound. It should also be understood that various modifications of the estimates can be introduced, for example, non-linear functions (e.g., logarithmic functions) can be applied.
[0095] Any measure of the diffuse reverberant signal-to-total signal relationship that indicates the level of diffuse reverberant sound relative to the total radiated sound in the acoustic environment can be used and provided in the metadata. The following description focuses on the relationship expressed by the ratio between the level of the diffuse reverberant signal and the level of the total signal ratio (e.g., energy or energy density). Therefore, this description focuses on the example of the diffuse reverberant signal-to-total signal ratio, also referred to as DSR.
[0096] The receiver 501 may be implemented in any suitable manner, including, for example, using discrete or dedicated electronics. The receiver 501 may be implemented as an integrated circuit, such as, for example, an application specific integrated circuit (ASIC). In some embodiments, the circuitry may be implemented as a programmed processing unit, such as firmware or software running on a suitable processor, such as, for example, a central processing unit, a digital signal processing unit, or a microcontroller. In such embodiments, it will be understood that the processing unit may include on-board or external memory, clock driver circuitry, interface circuitry, user interface circuitry, etc. Such circuitry may further be implemented as part of the processing unit, as an integrated circuit, and / or as discrete electronic circuitry.
[0097] The receiver 501 receives audio scene data from any suitable audio source in any suitable form, including, for example, as part of an audio signal. The data may be received from an internal source or an external source. The receiver 401 may be configured to receive the room data, for example, via a network connection, a wireless connection, or any other suitable connection to an internal source. In many embodiments, the receiver receives the data from a local source, such as local memory. In many embodiments, the receiver 501 is configured to retrieve the room data from local memory, for example, local RAM or ROM memory.
[0098] The receiver 501 is coupled to the path renderer 401 and forwards the audio scene data thereto for generating the path signal components (direct path and early reflections) as described above.
[0099] The audio reverberator further comprises a downmixer 405 to which the audio scene data is also supplied. The downmixer 405 comprises an energy circuit / processor 505, a coefficient circuit / processor 507 and a downmix circuit / processor 509.
[0100] The downmixer 405, and indeed each of the energy circuit / processor 505, coefficient circuit / processor 507, and downmix circuit / processor 509, may be implemented in any suitable manner, including, for example, using discrete or dedicated electronics. The receiver 501 may be implemented as an integrated circuit, such as, for example, an application specific integrated circuit (ASIC). In some embodiments, the circuit / processor is implemented as a programmed processing unit, such as firmware or software running on a suitable processor, such as, for example, a central processing unit, a digital signal processing unit, or a microcontroller. It will be appreciated that in such embodiments, the processing unit includes on-board or external memory, clock driving circuits, interface circuitry, user interface circuitry, etc. Such circuitry may further be implemented as part of the processing unit, as an integrated circuit, and / or as a separate electronic circuit.
[0101] The coefficient processor 507 is configured to determine downmix coefficients for at least some of the received audio signals. The downmix coefficients of an audio signal correspond to the weighting of that audio signal in the downmix. The downmix coefficients are the weights of the audio signals in the weighted combination that generates the downmix signal. Thus, the downmix coefficients are the relative weights of the audio signals when combined to generate the downmix signal (which in many embodiments is a mono signal), e.g., weights in a weighted sum.
[0102] The coefficient processor 507 is configured to generate downmix coefficients based on the received diffuse reverberant signal to total signal ratio, ie the diffuse to source ratio DSR.
[0103] This coefficient is further determined as a function of a determined total radiant energy index, which indicates the total energy radiated from the audio source. The DSR is typically common to some, typically all, audio signals, whereas the total radiant energy index is typically specific to each audio source.
[0104] The total radiant energy index usually indicates the normalized total radiant energy. The same normalization is applied to all audio sources and direct and reflected path components. The total radiant energy index is therefore relative to the total radiant energy index of other audio sources / signals, or individual path components, or the full-scale sample value of the audio signal.
[0105] The total radiant energy index, when combined with the DSR, provides a downmix coefficient for each sound source that reflects the relative contribution of that sound source to diffuse reverberant sound. Therefore, determining the downmix coefficient as a function of the DSR and the total radiant energy index provides a downmix coefficient that reflects the relative contribution to diffuse sound. Therefore, when the downmix coefficients are used to generate a downmix signal, each of the sound sources is appropriately weighted, resulting in a downmix signal that reflects the overall sound generated in an environment whose acoustic environment is accurately modeled.
[0106] In many embodiments, the downmix coefficients as a function of the DSR and the total radiant energy index combined with scaling according to the characteristics of the reverberator (407) provide downmix coefficients that reflect the appropriate relative level of diffuse reverberation with respect to the corresponding path signal components.
[0107] The energy processor 505 is coupled to the coefficient processor 507 and is configured to determine a total radiant energy measure from the metadata received about the audio source.
[0108] The received metadata includes a signal reference level for each audio source that provides an indication of the level of the audio. The signal reference level is typically a normalized or relative value that provides an indication of the signal reference level relative to other audio sources or normalized reference levels. Thus, the signal reference level typically does not indicate the absolute sound level of an audio source, but rather its level relative to other audio sources.
[0109] In a specific example, the signal reference level includes an indication in the form of a reference distance, providing a distance at which 0 dB of distance attenuation is applied to the audio signal. Thus, when the distance between the audio source and the listener is equal to the reference distance, the received audio signal can be used without distance-dependent scaling. At distances shorter than the reference distance, attenuation is small, and therefore a gain greater than 0 dB needs to be applied when determining the sound level at the listening position. At distances greater than the reference distance, attenuation is large, and therefore a gain greater than 0 dB needs to be applied when determining the sound level at the listening position. Similarly, for a constant distance between the audio source and the listening position, a higher gain is applied to audio signals associated with a longer reference distance than to audio signals associated with a shorter reference distance. Because audio signals are typically normalized to represent a meaningful reference distance or to utilize their full dynamic range (e.g., a jet engine and a cricket are both represented by audio signals that utilize the full dynamic range of the data words used), the reference distance provides an indication of the signal reference level for a particular audio source.
[0110] In this example, the signal reference level is further indicated by a reference gain, referred to as pre-gain, which is provided for each audio source and provides the gain that needs to be applied to the audio signal when determining the rendered audio level. Thus, pre-gain is used to further indicate level variations between different audio sources.
[0111] The metadata further includes directionality data indicating the directionality of sound radiation from the sound source represented by the audio signal. The directionality data for each sound source indicates the relative gain relative to a signal reference level in different directions from the sound source. The directionality data provides, for example, a complete function or description of the radiation pattern from the sound source defining the gain in each direction. As another example, a simplified indicator is used, such as a single data value indicating a given pattern. As yet another example, the directionality data provides individual gain values for a range of different directional intervals (e.g., segments of a sphere).
[0112] Thus, the audio level can be generated by the metadata along with the audio signal. Specifically, the path renderer determines the signal component of the direct path by applying a gain to the audio signal, where the gain is a combination of a pre-gain, a distance gain determined as a function of the distance between the audio source and the listener and a reference distance, and a directional gain in the direction from the audio source to the listener.
[0113] For the generation of the diffuse reverberant signal, the metadata is used to determine a (normalized) total radiant energy index of the audio source based on the signal reference level and directivity data of the audio source.
[0114] Specifically, the total radiant energy index is generated by integrating the directional gain over all directions (e.g., integrating over the surface of a sphere centered at the location of the sound source) and scaled by the signal reference level, specifically the distance gain and pre-gain.
[0115] The determined total radiant energy indicator is then fed to the coefficient processor 507 and processed in a DSR to generate the downmix coefficients.
[0116] The downmix coefficients are then used by the downmix processor 509 to generate a downmix signal. Specifically, the downmix signal is generated as a combination, specifically a sum, of audio signals where each audio signal is weighted by the downmix coefficient of the corresponding audio signal.
[0117] The downmix is typically produced as a mono signal and then fed to a reverberator 407 to produce a diffuse reverberant signal.
[0118] It should be noted that the rendering and generation of the individual path signal components by the path renderer 401 is position dependent, e.g., with respect to determining distance gain and directional gain, and the subsequent generation of the diffuse reverberant signal is position independent of both the sound source and the listener.
[0119] The total radiated energy index can be determined based on the signal reference level and directivity data, without taking into account the positions of the audio source and the listener. Specifically, the pre-gain and the reference distance of the audio source can be used to determine a directivity-independent signal reference level, for example, normalized with respect to a full-scale sample of the audio signal, at a nominal distance from the audio source (the nominal distance is the same for all audio signals / audio sources). The integration of the directivity gain over all directions can be performed over a normalized sphere, for example, as in the case of a sphere at the reference distance. Thus, the total radiated energy index is independent of the positions of the audio source and the listener (reflecting that in an environment such as a room, diffuse reverberant sound tends to be uniform). The total radiated energy index is then combined with the DSR to generate downmix coefficients (in many embodiments, other parameters, such as reverberator parameters, can also be taken into account). Because the DSR is also position-independent, a diffuse reverberant signal is generated without taking into account the specific positions of the audio source and the listener, just like the downmix and reverberation processes.
[0120] Such an approach provides high-performance, natural-sounding audio perception without requiring excessive computational resources, and is particularly suitable for, for example, virtual reality applications where the user (and audio sources) move through the environment, thus dynamically changing the relative positions of the listener (and possibly some or all of the audio sources).
[0121] Following certain aspects of various embodiments of the approach of Figures 4 and 5 are described in more detail.
[0122] In many embodiments, the metadata further includes an indicator that indicates when the diffuse reverberation signal should start, i.e., it indicates a time delay associated with the diffuse reverberation signal. The time delay indicator is particularly in the form of a pre-delay.
[0123] Predelay describes the delay / lag in the RIR and is defined to be the threshold between early reflections and the diffuse, late reverberant sound. This threshold usually occurs as part of a smooth transition from (more or less) individual reflections to a fully coherent mixture of higher order reflections, so an appropriate threshold is chosen using a suitable evaluation / decision process. This decision can be made automatically based on an analysis of the RIR, or calculated based on the room dimensions and / or material properties.
[0124] Alternatively, a fixed threshold can be chosen, for example 80 ms into the RIR. The predelay can be expressed in seconds, milliseconds, or samples. In the following description, it is assumed that the predelay is chosen at a point after the reverberation has actually diffused. However, even if this is not the case, the described method works well.
[0125] Therefore, the pre-delay indicates the onset of the diffuse reverberation response from the start of the sound source radiation. For example, if the sound source starts radiating at t0 (e.g., t0=0), as shown in Figure 6, the direct sound reaches the user at t1 (>t0), the first reflection reaches the user at t2 (>t1), and the defined threshold between the early reflections and the diffuse reverberation reaches the user at t3 (>t2). In that case, the pre-delay is t3-t0.
[0126] The system uses the diffuse reverberant signal to total signal ratio, or DSR, to express the amount of diffuse reverberant energy received by the user, or the level of a sound source, as a ratio of the total radiant energy of that sound source, so that the diffuse reverberant energy is appropriately adjusted for level calibration of the rendered signal and corresponding metadata (e.g., pre-gain).
[0127] Expressing it in this way ensures that the values are independent of the absolute position and orientation of the listener and sound source within the environment, independent of the relative position and orientation of the user to the sound source and vice versa, independent of the particular algorithm for rendering the reverberation, and have a meaningful link to the signal levels used in the system.
[0128] The described approach takes into account both directional patterns and calculates downmix coefficients that impose the correct relative levels between the source signals and the DSR to achieve the correct level at the output of the reverberator 407.
[0129] The DSR represents the ratio between the radiated sound source energy and the diffuse reverberation characteristics, specifically the energy or (initial) level of the diffuse reverberation signal.
[0130] This description focuses primarily on the DSR, which indicates the diffuse reverberation energy relative to the total energy.
number
[0131] The diffuse reverberation energy is considered to be the energy generated by the room response from the beginning of the diffuse section; for example, it is the energy of the RIR from the time indicated by the predelay to infinity. Note that subsequent room excitation adds to the reverberation energy, and therefore it can usually only be measured directly by excitation with a Dirac pulse. Alternatively, it can be derived from the measured RIR.
[0132] Reverberation energy represents the energy at a single point in diffuse field space, rather than being integrated over the entire space.
[0133] A particularly advantageous alternative to the above is to use the DSR, which indicates the initial amplitude of the diffuse sound relative to the total radiated sound energy in the environment. Specifically, the DSR indicates the reverberation amplitude at the time indicated by the pre-delay.
[0134] The amplitude at the pre-delay is the maximum excitation of the room impulse response at the pre-delay or immediately after the pre-delay, for example within 5, 10, 20 or 50 milliseconds after the pre-delay. The reason for choosing the maximum excitation within a particular range is that at the pre-delay time, the room impulse response happens to be in the low part of the response. The general trend is for decaying amplitudes, and the maximum excitation in the short interval after the pre-delay is usually also the maximum excitation of the entire diffuse reverberation response.
[0135] Using the DSR to indicate the initial amplitude (e.g., within a 10 ms interval) makes it easier and more reliable to map the DSR to parameters of many reverberation algorithms. Thus, in some embodiments, the DSR is
number
[0136] Parameters in DSR are expressed relative to the same source signal level reference.
[0137] This can be achieved, for example, by measuring (or simulating) the RIR of the room of interest using a microphone within certain known conditions (such as the distance between the sound source and the microphone, the directivity pattern of the sound source, etc.) The sound source should radiate a calibrated amount of energy, e.g., a Dirac impulse with known energy, into the room.
[0138] The calibration coefficients for the electrical and analog-to-digital conversion of the measurement equipment are measured or derived from specifications. They can also be calculated from the directivity pattern of the sound source and the direct path response of the RIR, which can be predicted from the distance between the source and the microphone. The direct response has a specific energy in the digital domain and represents the radiated energy multiplied by a directional gain relative to the direction of the microphone and a distance gain that depends on the microphone surface for a full spherical surface area with a radius equal to the distance between the sound source and the microphone.
[0139] Both elements must use the same digital level reference, e.g., a full-scale 1 kHz sine corresponds to 100 dBSPL.
[0140] Measuring the diffuse reverberation energy from the RIR and compensating it with a calibration factor gives the appropriate energy in the same area as the known radiant energy. With the radiant energy, the appropriate DSR can be calculated.
[0141] The reference distance indicates the distance at which the distance gain applied to the signal is 0 dB, i.e., no gain or attenuation is applied to compensate for the distance. The actual distance gain applied by the path renderer 401 can then be calculated by considering the actual distance against the reference distance.
[0142] The expression of the effect of distance on sound propagation is performed with reference to a given distance: doubling the distance reduces the energy density (energy per surface unit) by 6 dB; halving the distance induces an increase in the energy density (energy per surface unit) by 6 dB.
[0143] To determine the distance gain at a particular distance, i.e., to determine how much the density has decreased or increased, it is necessary to know the distance corresponding to a particular level so that the relative change in current distance can be determined.
[0144] Neglecting absorption in air and assuming no reflections or occlusions, the radiant energy of a sound source is constant on a sphere of any radius centered at the source position. The ratio of the surface corresponding to the actual distance to the reference distance indicates the energy attenuation. The linear signal amplitude gain at a rendering distance d can be expressed as b,
number
[0145] As an example, if the reference distance is 1 meter and the rendering distance is 2 meters, this formula results in a signal attenuation of approximately 6 dB (or a gain of -6 dB).
[0146] The total radiated energy index describes the total energy radiated by a sound source. Sound sources usually radiate in all directions, but not equally in all directions. The integral of the energy density over a sphere around the sound source gives the total radiated energy. For a loudspeaker, the radiated energy can often be calculated knowing the voltage applied to its terminals and the loudspeaker coefficients that describe the impedance, energy losses, and transfer of electrical energy to sound pressure waves.
[0147] The energy processor 505 is configured to determine the total radiated energy measure by taking into account the directivity data of the sound source. It should be noted that when determining the diffuse reverberation signal of a sound source with varying source directivity, it is important to use the total radiated energy, not just the signal level or signal reference level. For example, consider a sound source directivity corresponding to a very narrow beam with a directivity coefficient of 1 and coefficients of 0 in all other directions (i.e., energy is transmitted only in a very narrow beam). In this case, the radiated source energy represents the total energy and is very similar to the energy and signal reference level of the sound signal. If another sound source with the same energy and signal reference level but an omnidirectional sound signal is instead considered, the radiated energy of this sound source will be much higher than the sound signal energy and signal reference level. Therefore, if both sound sources are active simultaneously, the signal of the omnidirectional sound source should be much more strongly represented in the diffuse reverberation signal, i.e., in the downmix, than the highly directional sound source.
[0148] As previously mentioned, the energy processor 505 determines the radiated energy by integrating the energy density over the surface of a sphere surrounding the sound source. Ignoring distance gain, i.e., integrating over a surface at a radius where distance gain is 0 dB (i.e., the radius corresponding to the reference distance), the total radiated energy index can be determined from the following equation:
number
[0149] p is independent of direction, so it moves out of the integral. Similarly, the signal x is independent of direction (the directional gain reflects that variation).
number
[0150] One particular approach for determining this integral is described in more detail below.
[0151] It is desirable to integrate the directional gain over a sphere.
number
[0152] Using a sphere with a radius equal to the reference distance (r) means that there will be 0 dB in distance gain and distance gain / attenuation can be neglected.
[0153] In this example, a sphere is chosen for its computational convenience, but the same energy can be determined from any closed surface of any shape that surrounds the source position. As long as appropriate distance and directivity gains are used in the integration, the effective surface is considered to be facing the source position (i.e., with a normal vector along the source position).
[0154] The surface integral requires defining a small surface dS. Therefore, defining a sphere using two parameters, azimuth (a) and elevation (e), gives us the dimensions to do this. Using a coordinate system for the solution, f(a,e,r)=r*cos(e)*cos(a)*u x +r*cos(e)*cos(a)*u y +r*sin(e)*u z And where u x ,u y , and u z are the unit basis vectors of the coordinate system.
[0155] The small surface dS is the magnitude of the cross product of the partial derivatives of the spherical surface with respect to the two parameters multiplied by the derivative of each parameter. dS=|f a ×f e |da de.
[0156] This derivative is the vector tangent to the sphere at the point of interest f a =-r*cos(e)*sin(a)*u x +r*cos(e)*cos(a)*u y +0*u z and, f e =-r*sin(e)*cos(a)*u x -r*sin(e)*sin(a)*u y +r*cos(e)*u z Determine.
[0157] The cross product of derivatives is a vector that is perpendicular to both.
[0158] f a ×f e =(r 2 *cos(e)*cos(a)*cos(e)+0*sin(e)*sin(a))*u x +(-0*sin(e)*cos(a)+r 2 *cos(e)*sin(a)*cos(e))*u y +(r 2 *cos(e)*sin(a)*sin(e)*sin(a)+r 2 *cos(e)*cos(a)*sin(e)*cos(a))*u z =r 2 *cos 2 (e)*cos(a)*u x +r 2 *cos 2 (e)*sin(a)*u y +(r 2 *cos(e)*sin(e)*sin 2 (a)+r 2 *cos(e)*sin(e)*cos 2 (a))*u z =r 2 *cos 2 (e)*cos(a)*u x +r 2 *cos 2 (e)*sin(a)*u y +(r2 *cos(e)*sin(e)*(sin 2 (a)+cos 2 (a)))*u z =r 2 *cos 2 (e)*cos(a)*u x +r 2 *cos 2 (e)*sin(a)*u y +r 2 *cos(e)*sin(e)*u z
[0159] The magnitude of the cross product is the surface area of the parallelogram spanned by the vectors f_a and f_e, i.e., the surface area of the sphere, |f a ×f e |=sqrt((r 2 *cos 2 (e)*cos(a)) 2 +(r 2 *cos 2 (e)*sin(a)) 2 +(r 2 *cos(e)*sin(e)) 2 ) =sqrt(r 4 *cos 4 (e)*cos 2 (a)+r 4 *cos 4 (e)*sin 2 (a)+r 4 *cos 2 (e)*sin 2 (e)) =sqrt(r 4 *cos 4 (e)*(cos 2 (a)+sin 2 (a))+r 4 *cos 2 (e)*sin 2 (e)) =sqrt(r 4 *cos 4 (e)+r 4 *cos 2 (e)*sin 2 (e)) =sqrt(r 4 *cos 2 (e)*(cos 2 (e)+sin 2 (e))) =sqrt(r 4 *cos 2 (e)) =abs(r 2 *cos(e) =r 2 *cos(e) where e=[-0.5*pi,0.5*pi].
[0160] As a result, dS=r 2 *cos(e)*da*de, where the first two terms define the normalized surface area, which when multiplied by da and de, results in the actual surface, based on the size of the segments da and de. The double integral over the surface can be expressed in terms of azimuth and elevation angles. The surface dS is expressed in terms of a and e, as above. The two integrals can be performed over azimuth = 0...2*pi (inner product), and elevation = -0.5*pi...0.5*pi (cross product).
number
[0161] In many practical embodiments, the directional pattern is not provided as an integrable function, but rather as a discrete set of sample points. For example, each sampled directional gain is associated with an azimuth angle and an elevation angle. Typically, these samples represent a grid on a sphere. One approach to dealing with this is to convert the integral to a summation, i.e., a discrete integration is performed. In this example, the integration is performed as a summation over the points on the sphere where the directional gain is available. This gives the value of g(a,e), but da and de need to be chosen correctly so that there are no large errors due to overlaps or gaps.
[0162] In other embodiments, the directivity pattern is provided as a limited number of non-uniformly spaced points in space, in which case the directivity pattern is interpolated and uniformly resampled across the range of azimuth and elevation angles of interest.
[0163] Another solution is to assume that g(a,e) is constant around that defined point and then analytically solve the integral locally, for example, for small azimuth and elevation ranges, e.g., halfway between adjacent defined points. This uses the integral above, but for different ranges of a and e, where g(a,e) is considered constant.
[0164] Experiments show that simple summation produces small errors even when the directivity resolution is fairly coarse. Furthermore, the error is independent of radius. Linear spacing in azimuth between 10 points and 10 linearly spaced points in elevation results in a relative error of -20 dB.
[0165] The above integral provides a result that scales with the radius of the sphere. It therefore scales with the reference distance. This dependency on radius does not take into account the inverse effect of "distance gain" between two different radii. Doubling the radius results in 6 dB less energy "flowing" through a given surface area (e.g., 1 cm2). Therefore, one could say that the integration must take distance gain into account. However, the integration is performed over a reference distance, defined as the distance at which the distance gain is reflected in the signal. In other words, the signal level indicated by the reference distance is not included in the scaling of the value being integrated, but is instead reflected by the surface area over which the integration is performed, which varies with the reference distance (since the integration is performed over a sphere with a radius equal to the reference distance).
[0166] As a result, the integral above reflects the energy scaling factor (including pre-gain or similar calibration adjustment) of the audio signal, since the audio signal represents the correct signal reproduction energy at a fixed surface area of a sphere with a radius equal to the reference distance (without directional gain).
[0167] This means that if the reference distance is large, the total signal energy scaling factor will also be large without changing the signal, since the corresponding signal will represent a source that is relatively larger than a source with the same signal energy, but at a small reference distance.
[0168] In other words, the signal level indicator provided by the reference distance is automatically taken into account by performing the integration over the surface of a sphere with a radius equal to the reference distance. The larger the reference distance, the larger the surface area and the larger the total radiant energy indicator. The integration is specifically performed directly at a distance where the distance gain is 1.
[0169] The integral above is normalized to the surface units used and to the units used to express the reference distance r. If the reference distance r is expressed in meters, the result of the integral is m 2 It is provided in units of
[0170] To relate an estimated radiant energy value to a signal, it must be expressed in surface units corresponding to the signal. The surface area of the human ear may be more appropriate, since the signal level represents the level reproduced by a user at a reference distance. At the reference distance, this surface relative to the entire surface of a sphere relates to the portion of the energy of the sound source perceived by a person.
[0171] Therefore, the total radiated energy index, which represents the radiated source energy normalized to a full-scale sample in the audio signal, is:
number
[0172] Using the DSR, which characterizes the diffuse acoustic properties of the space, and the calculated radiated source energy derived from the directivity, pre-gain, and reference distance metadata, the corresponding reverberation energy can be calculated.
[0173] The DSR is typically determined at the same reference level used by both of its components, which may be the same as or different from the total radiant energy index. In any case, when such a DSR is combined with the total radiant energy index, the resulting reverberation energy is also expressed as an energy normalized to a full-scale sample in the audio signal, if the total radiant energy determined by the above integration is used. In other words, all considered energies are essentially normalized to the same reference level so that they can be directly combined without the need for level adjustment. Specifically, the determined total radiant energy can be used directly with the DSR to generate a level index of the diffuse reverberation generated from each sound source, which directly indicates the appropriate level relative to the diffuse reverberation of other sound sources and for each individual path signal component.
[0174] As a specific example, the relative signal levels of the diffuse reverberant signal components of different sound sources are obtained directly by multiplying the DSR by the total radiated energy index.
[0175] In the described system, adaptation of the contributions of different audio sources to the diffuse reverberant signal is performed at least in part by adapting the downmix coefficients used to generate the downmix signal, such that the relative contribution / energy level of the diffuse sound from each audio source reflects the diffuse reverberant energy determined for the audio sources.
[0176] As a specific example, if the DSR indicates an initial amplitude level, the downmix coefficient is determined to be proportional to (or equal to) the DSR multiplied by the total radiant energy index, and if the DSR indicates an energy level, the downmix coefficient is determined to be proportional to (or equal to) the square root of the DSR multiplied by the total radiant energy index.
[0177] As a specific example, for a signal with multiple input signal indices x, the downmix coefficients d x teeth,
number
number
[0178] Or, the downmix coefficient d x is d x =E norm,x *Calculated according to DSR, where:
number
[0179] In many embodiments, the downmix coefficients are determined in part by combining the DSR with a total radiant energy measure. Regardless of whether the DSR indicates the relationship of total radiant energy to diffuse reverberant energy or the initial amplitude of the diffuse reverberant response, further adaptation of the downmix coefficients is often necessary to accommodate the specific reverberator algorithm used, scaling the signal so that the output of the reverberation processor reflects the desired energy or initial amplitude. For example, the density of reflections in a reverberation algorithm has a strong influence on the reverberant energy generated, even if the input level remains the same. As another example, the initial amplitude of a reverberation algorithm is not equal to the amplitude of its excitation. Therefore, algorithm-specific, or algorithm- and configuration-specific, adjustments are required, which can be included in the downmix coefficients and are typically common to all sound sources. In some embodiments, these adjustments are applied to the downmix or included in the reverberator algorithm.
[0180] Once the downmix coefficients are generated, the downmix processor 509 generates the downmix signal, for example by direct weighted combination or summation.
[0181] An advantage of the described approach is that it uses a conventional reverberator, for example, reverberator 407, implemented by a feedback delay network, such as is implemented in a standard Jot reverberator.
[0182] As illustrated in Figure 7, the principle of a feedback delay network uses one or more (usually several) feedback loops with different delays. An input signal, in this case a downmix signal, is fed into the loop, where the signal is fed back with an appropriate feedback gain. The output signal is extracted by combining the signals in the loop. Thus, the signal is repeated successively with different delays. By using relatively prime delays and having a feedback matrix that mixes the signals between the loops, it is possible to create patterns that resemble reverberation in a real space.
[0183] To achieve a stable damped impulse response, the absolute values of the elements in the feedback matrix must be less than 1. In many implementations, additional gains or filters are included in the loop. These filters can control the damping instead of the matrix. Using filters has the advantage that the damping response varies with frequency.
[0184] In some embodiments in which the reverberator output is rendered binaurally, the estimated reverberation is filtered by average HRTFs (head-related transfer functions) for each of the left and right ears to generate left and right channel reverberation signals. It can be appreciated that if HRTFs are available for multiple uniformly spaced distances on a sphere around the user, the average HRTFs for the left and right ears are generated using the set of HRTFs with the largest distances. The use of average HRTFs is based on or reflects the consideration that reverberation is isotropic and arrives from all directions. Thus, rather than including a pair of HRTFs for a given direction, an average over all HRTFs can be used. The averaging can be performed once for the left ear and once for the right ear, and the resulting filters are used to process the reverberator output for binaural rendering.
[0185] In some cases, the reverberator itself introduces coloration of the input signal, resulting in an output that does not have the desired output diffuse signal energy as described by DSR. Therefore, the effect of this process is also equalized. This equalization can be performed based on a filter that is analytically determined as the inverse of the frequency response of the reverberator operation. In some embodiments, the transfer function can be estimated using machine learning techniques such as linear regression, line fitting, etc.
[0186] In some embodiments, the same approach is applied uniformly across frequency bands. However, in other embodiments, frequency-dependent processing is performed. For example, one or more of the provided metadata parameters are frequency-dependent. In such examples, the device is configured to split the signal into different frequency bands corresponding to the frequency dependence, and the aforementioned processing is performed in each of the frequency bands individually.
[0187] Specifically, in some embodiments, the diffuse reverberant signal to total signal ratio DSR is frequency-dependent. For example, different DSR values are provided for individual frequency band / bin ranges, or the DSR is provided as a function of frequency. In such embodiments, the device is configured to generate frequency-dependent downmix coefficients that reflect the frequency dependency of the DSR. For example, downmix coefficients for individual frequency bands are generated. Similarly, a frequency-dependent downmix and a diffuse reverberant signal are generated as a result.
[0188] In the case of frequency-dependent DSR, the downmix coefficients are in other embodiments complemented by filters that filter the audio signals as part of generating the downmix. As another example, the DSR effect is separated into a frequency-independent (broadband) component that is used to generate frequency-independent downmix coefficients that are used to scale the individual audio signals when generating the downmix signal, and a frequency-dependent component that is applied to the downmix, for example, by applying a frequency-dependent filter to the downmix. In some embodiments, such filters are combined with further coloration filters, for example, as part of a reverb algorithm. Figure 7 shows a correlation (u,v) filter and a coloration (h L ,h R ) filter, known as the Jot reverberator, a feedback delay network dedicated to binaural output.
[0189] Therefore, in some embodiments, the DSR comprises a frequency-dependent component part and a non-frequency-dependent component part, and the coefficient processor 507 is configured to generate downmix coefficients depending on the non-frequency-dependent component part (and independent of the frequency-dependent part). The processing of the downmix is then adapted based on the frequency-dependent component part, i.e. the reverberator is adapted depending on the frequency-dependent part.
[0190] In some embodiments, the directionality of sound radiation from one or more of the audio sources is frequency dependent, and in such a scenario the energy processor 505 is configured to generate a frequency dependent total radiation energy which, when combined with the DSR (which may or may not be frequency dependent), results in a frequency dependent downmix coefficient.
[0191] This is achieved, for example, by performing individual processing on separate frequency bands. In contrast to frequency-dependent DSR processing, frequency-dependence on directivity typically needs to be performed before (or as part of) generating the downmix signal. This reflects the need for frequency-dependent downmixing to include the frequency-dependent effects of directionality, which typically vary depending on the sound source. After integration, the net effect can vary significantly with frequency. That is, the total radiant energy index of a given sound source varies from source to source and has substantial frequency dependence. Therefore, because different sound sources typically have different directional patterns, the total radiant energy index of different sound sources also typically has different frequency dependence.
[0192] A specific example of a possible approach is described below: By providing a DSR that characterizes the diffuse acoustic properties of a space and determining the radiated source energy from the directivity, pre-gain, and reference distance metadata, the corresponding desired reverberation energy can be calculated. For example, this can be expressed as E norm *Can be determined as DSR.
[0193] If the components for calculating the DSR use the same reference level (e.g., related to the full scale of the signal), the resulting reverberation energy will be calculated as E as calculated above for the radiated source energy. norm When using , it also results in an energy normalized to a full-scale sample in the PCM signal and therefore corresponds to the energy of the diffuse reverberation impulse response (IR) that can be applied to the corresponding input signal to provide the correct level of reverberation in the signal representation used.
[0194] These energy values can be used to determine the setting parameters of the reverberation algorithm, the downmix coefficients before the reverberation algorithm, or the downmix filter.
[0195] There are various techniques for generating reverberation. Feedback delay network (FDN)-based algorithms such as the Jot reverberator are a suitable low-complexity approach. Alternatively, noise sequences can be shaped to have the appropriate (frequency-dependent) decay and spectral shape. In both examples, a prototype IR (with at least the appropriate T60) can be adjusted so that its (frequency-dependent) level is correct.
[0196] The reverberator algorithm is adjusted to generate an impulse response with unit energy (or unit initial amplitude in DSR is related to the initial amplitude), or the reverberator algorithm includes a unique compensation, for example in the coloration filter of the Jot reverberator, or the downmix is modified by a (possibly frequency dependent) adjustment, or the downmix coefficients generated by the coefficient processor 507 are modified.
[0197] The compensation is determined by generating an impulse response without such adjustments, but with all other configurations applied (such as the appropriate reverberation time (T60) and reflection density (e.g., delay value in the FDN)), and measuring the energy of that IR.
number
[0198] The compensation is the inverse of that energy. To include it in the downmix coefficients, e.g.
number
[0199] In many other embodiments, the compensation is derived from configuration parameters. For example, if the DSR is related to the initial reverberation amplitude, the first reflection can be derived from its configuration. Correlation filters are, by definition, energy-preserving, and coloration filters can also be designed to be so.
[0200] Assuming there is no net boost or attenuation due to the coloration filter, the reverberator will have an initial amplitude (A0) that depends on, for example, T60 and a minimum delay value, minDelay.
number
[0201] The prediction of the reverberation energy is also done heuristically.
[0202] As a general model of the diffuse reverberation energy, we can consider the exponential function A(t), where
number
[0203] Calculating the cumulative energy of such a function asymptotically approaches a final energy value, which has an almost perfectly linear relationship with T60.
[0204] The coefficients of the linear relationship depend on the sparseness of the function A (setting every third value to 0 will roughly halve the energy), the initial value A0 (the energy is A0 2 ), and the sample rate (f s The diffusion tail can be reliably modeled with such a function using T60, reflection density (derived from the FDN delay), and sample rate. The A0 of the model can be calculated as above and is equal to the A0 of the FDN.
[0205] When generating multiple parametric reverberations with broadband T60 values ranging from 0.1 to 2 seconds, the energy of the IR is approximately linear with the model. The scaling factor between the actual energy and the average of the exponential equation model is determined by the sparseness of the FDN response. This sparseness decreases towards the end of the IR but has the most impact at the beginning. Testing the above with multiple configurations of delay values, we found that there is an approximately linear relationship between the model reduction factor and the minimum difference between the delays configured in the FDN. For example, in the particular implementation of the Jot reverberator, this results in a scaling factor SF calculated by SF=7.0208*MinDelayDiff+214.1928.
[0206] The energy of the model is calculated by integrating from t=0 to infinity. This can be done analytically and the result is
number
[0207] Combining the above, we get the following prediction for reverberation energy:
number
[0208] It should be understood that, for clarity, the above description has described embodiments of the invention with reference to different functional circuits, units, and processors. However, it will be apparent that any suitable distribution of functionality between different functional circuits, units, or processors may be used without detracting from the invention. For example, functionality shown to be performed by separate processors or controllers may be performed by the same processor or controller. References to specific functional units or circuits should therefore be seen merely as references to suitable means for providing the described functionality, rather than to any strict logical or physical structure or organization.
[0209] The invention can be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention is optionally implemented at least partly as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the invention may be physically, functionally and logically implemented in any suitable way. Indeed, functionality may be implemented in a single unit, in multiple units or as part of other functional units. Thus, the invention may be implemented in a single unit or may be physically and functionally distributed between different units, circuits and processors.
[0210] Although the present invention has been described in connection with several embodiments, it is not intended to be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the appended claims. In addition, although features may appear to be described in connection with particular embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.
[0211] Furthermore, although individually listed, a plurality of means, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. Moreover, although individual features are included in different claims, they may be advantageously combined, and inclusion in different claims does not imply that a combination of features is not feasible and / or advantageous. Furthermore, the inclusion of a feature in one category of claims does not imply limitation to this category, but indicates that the feature is equally applicable to other claim categories, as appropriate. Furthermore, the order of features in the claims does not imply a particular order in which the features must function, and in particular the order of individual steps in method claims does not imply that the steps must be performed in this order. Rather, steps may be performed in any suitable order. Moreover, singular references do not exclude pluralities; thus, references to "first," "second," etc. do not exclude pluralities. Reference signs in the claims are provided merely as a clarifying example and should not be construed as limiting the claims in any way.
Claims
1. 1. An audio device for generating a diffuse reverberant signal of an environment, the audio device comprising: a receiver for receiving a plurality of audio signals representative of sound sources within the environment; a metadata receiver for receiving metadata of the plurality of audio signals, the metadata comprising: a measure of the relationship between the diffuse reverberant signal and the total signal, which indicates the level of diffuse reverberant sound relative to the total radiated sound in the environment; For each audio signal, a signal level indicator; a metadata receiver including directivity data indicative of the directionality of sound radiation from the sound source represented by the audio signal; for each of the plurality of audio signals, a total radiant energy index based on the signal level index and the directional data; a circuit for determining a downmix coefficient based on the total radiant energy and the relationship of the diffuse reverberant signal to the total signal; a downmixer for generating a downmix signal by combining signal components of each audio signal generated by applying the downmix coefficients of each audio signal to said audio signal; a reverberator for generating the diffuse reverberant signal of the environment from the downmix signal components.
2. 2. An audio device according to claim 1, wherein the directionality of sound radiation is frequency dependent, and the circuitry determines a frequency dependent total radiated energy and a frequency dependent downmix coefficient.
3. 3. An audio device according to claim 1 or 2, wherein the relation between the diffuse reverberant signal and the total signal is frequency dependent, and the circuit determines frequency dependent downmix coefficients.
4. 4. The audio device of claim 1, wherein the relationship of the diffuse reverberant signal to the total signal comprises a frequency-dependent part and a non-frequency-dependent part, and the circuit determines the downmix coefficients in dependence on the non-frequency-dependent part and adapts the reverberator in dependence on the frequency-dependent part.
5. 5. The audio device of claim 1, wherein the circuit determines the total radiated energy measure of the first audio signal in response to scaling the signal level measure of the first audio signal by a value determined by integrating a directivity pattern of the sound source represented by a first audio signal of the plurality of audio signals, the directivity pattern being determined based on directivity data.
6. 6. The audio device of claim 1, wherein the signal level indicator for a first audio signal of the plurality of audio signals includes a reference distance, the reference distance indicating a distance from an audio source represented by the first audio signal for a distance reference gain for the first audio signal.
7. 7. An audio device according to claim 6 when dependent on claim 5, wherein the integration is performed over a distance which is the reference distance from the audio source represented by the first audio signal.
8. 8. An audio device according to any one of claims 1 to 7, wherein the diffuse reverberant signal to total signal relationship indicates the energy of diffuse reverberant sound relative to the energy of total radiated sound in the environment.
9. 9. An audio device according to any one of claims 1 to 8, wherein the diffuse reverberation signal versus total signal relationship is indicative of the initial amplitude of diffuse sound relative to the energy of total radiated sound in the environment.
10. 10. The audio device of claim 1, wherein the downmix coefficients determined for a first audio signal of the plurality of audio signals are independent of a position of a first audio source represented by the first audio signal.
11. 11. An audio device according to any one of the preceding claims, wherein the downmix coefficients determined for a first audio signal of the plurality of audio signals are independent of a listener's position.
12. 12. The audio device of claim 1, wherein the signal level indicator for a first audio signal of the plurality of audio signals further includes a gain indicator for the first audio signal, the gain indicator indicating a gain to apply to the first audio signal when rendering sound from a first audio source represented by the first audio signal, and the circuit determines the downmix coefficients for the first audio signal in response to the gain indicator.
13. 13. The audio device of claim 1, further comprising a direct rendering circuit configured to generate a direct path audio signal of a first audio signal of the plurality of audio signals in response to the signal level indicator and the directivity data of the first audio signal.
14. 14. The audio device of claim 1, wherein the metadata further comprises a delay index, and wherein the relationship of the diffuse reverberant signal to the total signal is indicative of the energy of diffuse reverberant sound having a longer delay than the delay index relative to the energy of total radiated sound in the environment.
15. 1. A method for generating a diffuse reverberant signal of an environment, the method comprising: receiving a plurality of audio signals representing sound sources within the environment; receiving metadata for the plurality of audio signals, the metadata comprising: a measure of the relationship between the diffuse reverberant signal and the total signal, which indicates the level of diffuse reverberant sound relative to the total radiated sound in the environment; For each audio signal, a signal level indicator; receiving metadata including directional data indicative of a directionality of sound emission from the sound source represented by the audio signal; for each of the plurality of audio signals, a total radiant energy index based on the signal level index and the directional data; determining a downmix coefficient based on the total radiant energy and the relationship of the diffuse reverberant signal to the total signal; generating a downmix signal by combining signal components of each audio signal generated by applying the downmix coefficients of each audio signal to said audio signals; and generating the diffuse reverberant signal of the environment from the downmix signal components.
16. A computer program comprising computer program code means for performing all the steps of the method according to claim 15 when the computer program is executed on a computer.