Generating binaural audio in response to multi-channel audio using at least one feedback delay network

By applying BRIR to each channel of a multi-channel audio signal using an FDN in a filter bank domain, the method improves out-of-head localization and natural sound simulation in headphone virtualization systems.

JP2025123226AActive Publication Date: 2025-08-22DOLBY LABORATORIES LICENSING CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025080881
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2014-05-05
Filing Date
2025-05-14
Publication Date
2025-08-22
Estimated Expiration
2034-12-18

AI Technical Summary

Technical Problem

Existing headphone virtualizers using feedback delay networks (FDNs) lack flexibility in simulating the microstructure of early reflections and often introduce excessive tonal distortion and reverberation, failing to achieve effective out-of-head localization.

Method used

Implement a binaural room impulse response (BRIR) to each channel of a multi-channel audio input signal, using a feedback delay network (FDN) to add common late reverberation, and control frequency-dependent acoustic attributes like reverberation decay time and interaural coherence through a filter bank domain implementation.

Benefits of technology

Enhances out-of-head localization and produces a more natural-sounding binaural virtualization by systematically controlling spatial acoustic attributes, reducing computational complexity and memory requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025123226000001_ABST
    Figure 2025123226000001_ABST
Patent Text Reader

Abstract

To provide a method and system for generating binaural audio in response to multi-channel audio using at least one feedback delay network.SOLUTION: A virtualization method for generating a binaural signal in response to channels of a multi-channel audio input signal applies a binaural room impulse response (BRIR) to each channel by using at least one feedback delay network (FDN) to apply a common late reverberation to a downmix of the channels. To each channel, a direct response and early reflection portion of a single-channel BRIR for the channel are applied. The downmix of the channels is processed in a second processing path including at least one FDN which applies the common late reverberation. The common late reverberation emulates collective macro attributes of late reverberation portions of at least some of single-channel BRIRs.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to Chinese Patent Application No. 201410178258.0, filed April 29, 2014, U.S. Provisional Patent Application No. 61 / 923,579, filed January 3, 2014, and U.S. Provisional Patent Application No. 61 / 988,617, filed May 5, 2014, the contents of each of which are incorporated herein by reference in their entirety. 1. Field of the Invention The present invention relates to methods (sometimes referred to as headphone virtualization methods) and systems for generating binaural signals in response to a multi-channel audio input signal by applying a Binaural Room Impulse Response (BRIR) to each channel (e.g., to all channels) of a set of channels of the input signal. In some embodiments, at least one feedback delay network (FDN) applies a late reverberant portion of the downmix BRIR to the downmix of said channels. [Background technology]

[0002] 2. Background of the invention Headphone virtualization (or binaural rendering) is a technique that aims to deliver a surround sound experience or immersive sound field using standard stereo headphones.

[0003] Early headphone virtualizers applied head-related transfer functions (HRTFs) to convey spatial information in binaural rendering. HRTFs are a set of directional and distance-dependent filter pairs that characterize how sound travels from a specific point in space (the source location) to a listener's ears in an anechoic environment. Essential spatial cues, such as interaural time difference (ITD), interaural level difference (ILD), head shadowing effect, and spectral peaks and notches due to shoulder and pinna reflections, can be perceived in the rendered HRTF-filtered binaural content. Due to the size constraints of the human head, HRTFs do not provide sufficient or robust cues for source distances beyond approximately one meter. As a result, virtualizers based solely on HRTFs typically do not achieve good out-of-head localization or perceived distance.

[0004] Many acoustic events in everyday life occur in reverberant environments, where audio signals reach the listener's ears through various reflected paths in addition to the direct path (from source to ear) modeled by the HRTF. Reflections introduce profound effects on the auditory experience, such as distance, room size, and other attributes of the space. To convey this information in binaural rendering, the virtualizer must apply room reverberation in addition to the cues in the direct path HRTF. The binaural room impulse response (BRIR) characterizes the transformation of an audio signal from a specific point in space to the listener's ears in a specific acoustic environment. In theory, the BRIR contains all acoustic cues relevant to spatial perception.

[0005] Figure 1 shows the full frequency range of each channel (X1,…,X N, X 1 , ..., X 2 , ... , X 3 , ... , X 4 , ... , X 5 , ... , X 6 , ... , X 7 , ... , X 8 , ... , X 9 , ... , X 10 , ... , X 11 , ... , X 12 , ... , X 13 , ... , X 14 , ... , X 15 , ... , X 16 , ... , X 17 , ... , X 18 , ... , X 19 , ... , X 20 , ... , X 21 , ... , X 22 , ... , X 23 , ... , X 24 , ... , X 25 , ... , X 26 , ... , X 27 , ... , X 28 , ... , X 29 , ... , X N Each of the channels X1, X2, X3, X4, X5, X6, X7, X8, X9, X10, X11, X12, X13, X14, X15, X16, X17, X18, X19, X20, X21, X22, X23, X24, X25, X36, X37, X40, X41, X42, X43, X54, X55, X66, X77, X88, X99, X10, X11, X12, X13, X14, X15, X26, X27, X38, X44, X55, X68, X79, X89, X99, X10, X11, X12, X13, X25, X39, X40, X41, X56, X57, X69, X69, X70, X81, X99, X10, X11, X12, X13, X25, X39, X42, X57, X69, X70, X81, X99, X10, X11, X12, X13, X25, X38, X41, X58, X69, X70, X82, X99, X10, X11, X12, X13, X25, X39, X42, X58, X69, X70, X81, X99, X10, X11, X12, X13, X25, X39, X43, X59, X69, X70, X82, X99, X10, X11, X12, X13, X25, X39, X41, X59, X69, X70, X81, X99, X10, X11, N BRIR N (BRIR for the corresponding source direction), and so on. The output of each BRIR subsystem (each of subsystems 2, ..., 4) is a time-domain signal comprising a left channel and a right channel. The left channel outputs of the BRIR subsystems are mixed in summation element 6, and the right channels of the BRIR subsystems are mixed in summation element 8. The output of element 6 is the left channel L of the binaural audio signal output from the virtualizer, and the output of element 8 is the right channel R of the binaural audio signal output from the virtualizer.

[0006] The multi-channel audio input signal may also include a low-frequency effects (LFE) or subwoofer channel, identified in FIG. 1 as the “LFE” channel. Conventionally, the LFE channel is not convolved with the BRIR; instead, it is attenuated (e.g., by more than −3 dB) in gain stage 5 of FIG. 1 , and the output of gain stage 5 is mixed equally (by summing elements 6 and 8) with each channel of the virtualizer's binaural output signal. An additional delay stage may be required in the LFE path to time-align the output of stage 5 with the outputs of the BRIR subsystems (2, …, 4). Alternatively, the LFE channel may simply be ignored (i.e., not presented to or processed by the virtualizer). For example, the FIG. 2 embodiment of the present invention (described below) simply ignores any LFE channel in the multi-channel audio input signal it processes. Many consumer headphones are unable to accurately reproduce the LFE channel.

[0007] In some typical virtualizers, the input signal is transformed from the time domain to the frequency domain into the QMF (quadrature mirror filter) domain to generate channels of QMF-domain frequency components. These frequency components are filtered in the QMF domain (e.g., in the QMF-domain implementation of subsystems 2, ..., 4 in Figure 1), and the resulting frequency components are then transformed back to the time domain (e.g., in the final stage of each of subsystems 2, ..., 4 in Figure 1). The audio output of the virtualizer is thus a time-domain signal (e.g., a time-domain binaural signal).

[0008] In general, each full-frequency-range channel of a multi-channel audio signal input to a headphone virtualizer is assumed to represent audio content emanating from a sound source at a known position relative to the listener's ears. The headphone virtualizer is configured to apply a binaural room impulse response (BRIR) to each such channel of the input signal. Each BRIR can be decomposed into two parts: a direct response and a reflected response. The direct response is the HRTF corresponding to the direction of arrival (DOA) of the sound source, adjusted with appropriate gain and delay due to the distance (between the sound source and the listener), and optionally augmented with parallax effects for small distances.

[0009] The remaining part of the BRIR models reflections. Early reflections are typically first or second order reflections and have a relatively sparse temporal distribution. The microstructure of each first or second order reflection (e.g., ITD and ILD) is important. For late reflections (sound reflected from three or more surfaces before reaching the listener), the echo density increases with the number of reflections, making the micro-attributes of individual reflections less observable. For increasingly later reflections, the macrostructure (e.g., reverberation decay rate, interaural coherence, and overall reverberation spectral distribution) becomes more important. Therefore, reflections can be further segmented into two parts: early reflections and late reverberations.

[0010] The delay of the direct response is the source distance from the listener divided by the speed of sound, and its level is inversely proportional to the source distance (in the absence of walls or large surfaces near the source position). On the other hand, the delay and level of late reverberation are generally insensitive to source position. For practical reasons, the virtualizer may choose to time-align direct responses from sources with different distances and / or compress their dynamic range. However, the time and level relationships between the direct response, early reflections, and late reverberation within the BRIR should be maintained.

[0011] The effective length of a typical BRIR reaches hundreds of milliseconds or more in many acoustic environments. Direct application of the BRIR requires convolution with a filter with thousands of taps, which is computationally expensive. Additionally, without parameterization, achieving sufficient spatial resolution requires large memory space to store BRIRs for different source positions. Last but not least, sound source positions can change over time and / or the listener's position and orientation can change over time. Accurate simulation of such motion requires time-varying BRIR impulse responses. Proper interpolation and application of such time-varying filters can be difficult when the impulse responses of these filters have many taps.

[0012] To implement a spatial reverberator configured to apply simulated reverberation to one or more channels of a multi-channel audio input signal, filters with a well-known filter structure known as a feedback delay network (FDN) can be used. The structure of an FDN is simple: several reverberation tanks (e.g., the FDN in Figure 4) are connected by a gain element g1 and a delay line z -n1 The FDN has a number of reverberation tanks (each with a delay and gain), each with its own delay and gain. In a typical implementation of the FDN, the outputs from all reverberation tanks are mixed by a unitary feedback matrix, and the output of the matrix is ​​fed back and summed with the reverberation tank's input. Gain adjustments may be made to the reverberation tank outputs. The reverberation tank outputs (or their gain-adjusted versions) can be suitably remixed for multichannel or binaural playback. Natural-sounding reverberation can be generated and applied by the FDN with a compact computational and memory footprint. Therefore, the FDN has been used in virtualizers to supplement the direct response generated by HRTFs.

[0013] For example, the commercially available "Dolby Mobile" headphone virtualizer includes a reverberator with an FDN-based architecture that is operable to add reverberation to each channel of a five-channel audio signal (having left front, right front, center, left surround, and right surround channels) and filter each reverberated channel with a different filter pair from a set of five head-related transfer function ("HRTF") filter pairs. The "Dolby Mobile" headphone virtualizer is also operable to generate a two-channel "reverberated" binaural audio output (a reverberated two-channel virtual surround sound output) in response to a two-channel audio input signal. When the reverberated binaural output is rendered and played through a pair of headphones, it is perceived at the listener's eardrums as HRTF-filtered reverberated sound from five loudspeakers located at left front, right front, center, left rear (surround), and right rear (surround) positions. The virtualizer upmixes the downmixed two-channel audio input (without using any spatial cue parameters received with the audio input) to generate five upmixed audio channels, adds reverberation to the upmixed channels, and downmixes the five reverberated channel signals to generate the virtualizer's two-channel reverberated output. The reverberation for each upmixed channel is filtered with a different pair of HRTF filters. Summary of the Invention [Problem to be solved by the invention]

[0014] In virtualizers, the FDN is configured to achieve a certain reverberation decay time and echo density. However, the FDN lacks the flexibility to simulate the microstructure of early reflections. Furthermore, in typical virtualizers, tuning and configuring the FDN is mostly trial and error.

[0015] Headphone virtualizers that do not simulate all reflection paths (early and late) cannot achieve effective out-of-head localization. The inventors have come to recognize that virtualizers using FDNs that attempt to simulate all reflection paths (early and late) typically have, at best, limited success in simulating both early reflections and late reverberation and adding both to the audio signal. The inventors have also come to recognize that virtualizers that use FDNs but do not have the ability to properly control spatial acoustic attributes such as reverberation decay time, interaural coherence, and direct-to-late ratio may achieve some degree of out-of-head localization, but at the cost of introducing excessive tonal distortion and reverberation. [Means for solving the problem]

[0016] In a first class of embodiments, the present invention is a method for generating a binaural signal in response to a set of channels of a multi-channel audio input signal (e.g., each of those channels or each of the full frequency range channels). The method includes: (a) applying a binaural room impulse response (BRIR) to each channel of the set (e.g., by convolving each channel of the set with the BRIR corresponding to that channel) to thereby generate a filtered signal, including by using at least one feedback delay network (FDN) to add common late reverberation to a downmix (e.g., a monophonic downmix) of the channels of the set; and (b) combining the filtered signals to generate the binaural signal. Typically, a bank of FDNs is used to add the common late reverberation to the downmix (e.g., each FDN adds common late reverberation to a different frequency band). Typically, step (a) involves applying to each channel of the set the "direct response and early reflection" portion of a single-channel BRIR for that channel, the common late reverberation being generated to emulate the collective macro-attributes of the late reverberation portions of at least some (e.g. all) of the single-channel BRIRs.

[0017] A method for generating a binaural signal in response to a multi-channel audio input signal (or in response to some set of channels of such a signal) is sometimes referred to herein as a "headphone virtualization" method, and a system configured to perform such a method is sometimes referred to herein as a "headphone virtualizer" (or "headphone virtualization system" or "binaural virtualizer").

[0018] In a first class of exemplary implementations, each FDN is implemented in the filter bank domain (e.g., the hybrid complex quadrature mirror filter (HCQMF) domain, the quadrature mirror filter (QMF) domain, or other transform or subband domain that may include decimation). In some such embodiments, frequency-dependent spatial acoustic attributes of the binaural signal are controlled by controlling the configuration of each FDN used to add late reverberation. Typically, for efficient binaural rendering of the audio content of a multi-channel signal, a monophonic downmix of the channels is used as input to the FDN. Exemplary embodiments of the first class include adjusting FDN coefficients corresponding to frequency-dependent attributes (e.g., reverberation decay time, interaural coherence, modal density, and direct-to-late ratio) by providing a control value to a feedback delay network to set at least one of the input gain, reverberation tank gain, reverberation tank delay, or output matrix parameters of each FDN. This enables better matching of the acoustic environment and a more natural-sounding output.

[0019] In a second class of embodiments, the present invention is a method for generating a binaural signal in response to a multi-channel audio input signal having channels by applying a binaural room impulse response (BRIR) to each channel of a set of channels of the input signal (e.g., each of the channels of the input signal or each full frequency range channel of the input signal). This includes processing each channel of the set in a first processing path configured to model and apply to each channel the direct response and early reflections of a single-channel BRIR for that channel, and processing a downmix (e.g., a monophonic downmix) of the channels of the set in a second processing path (parallel to the first processing path) configured to model and apply a common late reverberation to the downmix. Typically, the common late reverberation is generated to emulate collective macro-attributes of the late reverberant portions of at least some (e.g., all) of the single-channel BRIRs. Typically, the second processing path includes at least one FDN (e.g., one FDN for each of a plurality of frequency bands). Typically, the mono downmix is ​​used as the input to all reverberation tanks of each FDN implemented by the second processing path. Typically, a mechanism is provided for systematic control of the macro-attributes of each FDN to better simulate the acoustic environment and produce a more natural-sounding binaural virtualization. Because most such macro-attributes are frequency-dependent, each FDN is typically implemented in the hybrid complex quadrature mirror filter (HCQMF) domain, the frequency domain, or another filter bank domain, with a different or independent FDN used for each frequency band. The primary benefit of implementing an FDN in the filter bank domain is that it allows the application of reverberation with frequency-dependent reverberation attributes. In various embodiments, the FDN is implemented using any of a variety of filter banks in any of a wide variety of filter bank domains.These include, but are not limited to, real or complex-valued quadrature mirror filters (QMF), finite impulse response filters (FIR filters), infinite impulse response filters (IIR filters), discrete Fourier transforms (DFT), (modified) cosine or sine transforms, wavelet transforms, or crossover filters. In a preferred implementation, the filter banks or transforms used include decimation (e.g., reducing the sampling rate of the frequency-domain signal representation) to reduce the computational complexity of the FDN process.

[0020] Some embodiments of the first class (and the second class) implement one or more of the following features.

[0021] 1. A filterbank domain (e.g., hybrid complex quadrature mirror filter domain) FDN implementation or a hybrid filterbank domain FDN implementation and a time domain late reverberation filter implementation, which typically allows independent adjustment of FDN parameters and / or settings for each frequency band (which allows simple and flexible control of frequency-dependent acoustic attributes), for example by providing the ability to vary the reverberation tank delay in different bands to vary the modal density as a function of frequency.

[0022] 2. The particular downmix process used to generate the downmixed (e.g., monophonic downmixed) signal (from the multichannel input audio signal) that is processed in the second processing path depends on the source distance of each channel and the treatment of the direct response to maintain the proper level and timing relationship between the direct and late responses.

[0023] 3. An all-pass filter (APF) is applied in a second processing path (e.g., at the input or output of the bank of FDNs) to introduce phase diversity and increased echo density without changing the spectrum and / or timbre of the resulting reverberation.

[0024] 4. To overcome the problems associated with quantized delays in the downsample-factor grid, fractional delays are implemented in the feedback path of each FDN in a complex-valued multi-rate structure.

[0025] 5. In FDN, the reverberation tank outputs are linearly mixed directly into the binaural channels using output mix coefficients that are set based on the desired interaural coherence in each frequency band. Optionally, the mapping of reverberation tanks to binaural output channels alternates across frequency bands to achieve balanced delays between the binaural channels. Optionally, a normalization factor is applied to the reverberation tank outputs to equalize their levels while preserving fractional delays and overall power.

[0026] 6. Frequency-dependent reverberation decay time and / or modal density are controlled by setting the proper combination of reverberation tank delay and gain in each frequency band to simulate a real room.

[0027] 7. For each frequency band (e.g., at either the input or output of the associated processing path), one scaling factor is applied, which is: Controlling the frequency-dependent direct-to-late ratio (DLR) to match the DLR of the actual room (a simple model may be used to calculate the required scaling factor based on the target DLR and reverberation decay time, e.g., T60); providing low-frequency attenuation to mitigate excessive combing artifacts and / or low-frequency rumble; and / or This is to apply diffuse field spectral shaping to the FDN response.

[0028] 8. Simple parametric models are implemented to control essential frequency-dependent attributes of late reverberation, such as reverberation decay time, interaural coherence and / or direct-to-late ratio.

[0029] Aspects of the present invention include methods and systems that perform (or are configured to perform or support) binaural virtualization of audio signals (e.g., audio signals whose audio content consists of speaker channels and / or object-based audio signals).

[0030] In another class of embodiments, the present invention is a method and system for generating a binaural signal in response to a set of channels of a multi-channel audio input signal, by applying a binaural room impulse response (BRIR) to each channel of the set to thereby generate a filtered signal, including by using a single feedback delay network (FDN) to add common late reverberation to a downmix of the channels of the set; and combining the filtered signals to generate the binaural signal. The FDN is implemented in the time domain. In some such embodiments, the time-domain FDN: an input filter having an input coupled to receive the downmix, the input filter configured to generate a first filtered downmix in response to the downmix; an all-pass filter coupled and configured to provide a second filtered downmix in response to the first filtered downmix; a reverberation application subsystem having a first output and a second output, the reverberation application subsystem including a collection of reverberation tanks, each reverberation tank having a different delay, the reverberation application subsystem coupled and configured to generate a first unmixed binaural channel and a second unmixed binaural channel in response to the second filtered downmix, and to present the first unmixed binaural channel at the first output and the second unmixed binaural channel at the second output; an interaural cross-correlation coefficient (IACC) filtering and mixing stage coupled to the reverberation application subsystem and configured to generate first and second mixed binaural channels in response to the first and second unmixed binaural channels.

[0031] The input filter may be implemented to generate (preferably as a cascade of two filters configured to generate) the first filtered downmix so that each BRIR has a direct-to-late ratio (DLR) that at least substantially matches a target DLR.

[0032] Each reverberation tank may be configured to generate a delayed signal and may include a reverberation filter (e.g., implemented as a shelf filter or a cascade of shelf filters) coupled and configured to apply gain to the signal propagating in each reverberation tank so that the delayed signal has a gain that at least substantially matches a target delayed gain. 60 This is to achieve the following characteristics:

[0033] In some embodiments, the first unmixed binaural channel leads the second unmixed binaural channel, and the reverberation tank includes a first reverberation tank configured to generate a first delayed signal having the shortest delay and a second reverberation tank configured to generate a second delayed signal having a second shortest delay. The first reverberation tank is configured to apply a first gain to the first delayed signal, and the second reverberation tank is configured to apply a second gain to the second delayed signal, the second gain being different from the first gain, and the second gain being different from the first gain, and application of the first gain and the second gain results in an attenuation of the first unmixed binaural channel relative to the second unmixed binaural channel. Typically, the first mixed binaural channel and the second mixed binaural channel exhibit a recentered stereo image. In some embodiments, the IACC filtering and mixing stage is configured to generate the first mixed binaural channel and the second mixed binaural channel such that the first mixed binaural channel and the second mixed binaural channel have IACC characteristics that at least substantially match a target IACC characteristic.

[0034] Exemplary embodiments of the present invention provide a simple, unified framework for supporting both input audio consisting of speaker channels and object-based input audio. In embodiments in which BRIR is applied to input signal channels that are object channels, the "direct response and early reflections" processing performed for each object channel assumes a source direction indicated by the metadata provided with the audio content of that object channel. In embodiments in which BRIR is applied to input signal channels that are speaker channels, the "direct response and early reflections" processing performed for each speaker channel assumes a source direction corresponding to that speaker channel (i.e., the direction of the direct path from the assumed location of the corresponding speaker to the assumed listener location). Regardless of whether the input channel is an object channel or a speaker channel, the "late reverberation" processing is performed on a downmix of the input channel (e.g., a monophonic downmix) and does not assume any particular source direction for the audio content of the downmix.

[0035] Other aspects of the invention are a headphone virtualizer configured (e.g., programmed) to perform any embodiment of the inventive method, a system (e.g., a stereo, multi-channel or other decoder) including such a virtualizer, and a computer-readable medium (e.g., a disk) storing code for implementing any embodiment of the inventive method. [Brief explanation of the drawings]

[0036] [Figure 1] FIG. 1 is a block diagram of a typical headphone virtualization system. [Figure 2] FIG. 1 is a block diagram of a system including an embodiment of the headphone virtualization system of the present invention. [Figure 3] FIG. 2 is a block diagram of another embodiment of the headphone virtualization system of the present invention. [Figure 4]FIG. 4 is a block diagram of an FDN of a type that may be included in a typical implementation of the system of FIG. 3. [Figure 5] 1 is a graph of reverberation decay time (T60) in milliseconds as a function of frequency in Hz that may be achieved by an embodiment of a virtualizer of the present invention in which the T60 values ​​at each of two specific frequencies (fA and fB) are set as follows: T60,A=320 ms at fA=10 Hz and T60,B=150 ms at fB=2.4 kHz. [Figure 6] 1 is a graph of interaural coherence (Coh) as a function of frequency in Hz that may be achieved by an embodiment of the virtualizer of the present invention in which the control parameters Cohmax, Cohmin and fC are set to have values ​​of Cohmax=0.95, Cohmin=0.05 and fC=700 Hz. [Figure 7] 1 is a graph of the direct-to-late ratio (DLR) in dB at a source distance of 1 meter as a function of frequency in Hz that can be achieved by an embodiment of the virtualizer of the present invention in which the control parameters DLR1K, DLRslope, DLRmin, HPFslope and fT are set to have values ​​of DLR1K=18 dB, DLRslope=6 dB per frequency decade, DLRmin=18 dB, HPFslope=6 dB per frequency decade and fT=200 Hz. [Figure 8] FIG. 10 is a block diagram of another embodiment of the late reverberation processing subsystem of the headphone virtualization system of the present invention. [Figure 9] FIG. 1 is a block diagram of a time-domain implementation of an FDN of the type included in some embodiments of the system of the present invention. [Figure 9A] FIG. 10 is a block diagram of an example implementation of the filter 400 of FIG. 9. [Figure 9B] FIG. 10 is a block diagram of an example implementation of the filter 406 of FIG. 9. [Figure 10] FIG. 2 is a block diagram of an embodiment of the headphone virtualization system of the present invention in which the late reverberation processing subsystem 221 is implemented in the time domain. [Figure 11]10 is a block diagram of an embodiment of elements 422, 423, and 424 of the FDN of FIG. 9. A is a graph of the frequency response of a typical implementation of filter 500 (R1), the frequency response of a typical implementation of filter 501 (R2), and the frequency response of filters 500 and 501 connected in parallel. [Figure 12] 10 is a graph of an example of an IACC characteristic (curve "I") and a target IACC characteristic (curve "IT") that may be achieved by an implementation of the FDN of FIG. [Figure 13] 10 is a graph of the T60 characteristics that may be achieved by one implementation of the FDN of FIG. 9 by appropriately implementing each of filters 406, 407, 408, and 409 as shelf filters. [Figure 14] 10 is a graph of the T60 characteristics that may be achieved by one implementation of the FDN of FIG. 9 by appropriately implementing each of filters 406, 407, 408, and 409 as a cascade of two IIR shelf filters. DETAILED DESCRIPTION OF THE INVENTION

[0037] Notation and Nomenclature Throughout this disclosure, including the claims, the expression performing an operation "on" a signal or data (e.g., filtering, scaling, transforming, or applying a gain to the signal or data) is used broadly to refer to performing the operation either directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or pre-processing prior to performing the operation).

[0038] Throughout this disclosure, including the claims, the term "system" is used broadly to refer to an apparatus, system, or subsystem. For example, a subsystem that implements a virtualizer may be referred to as a virtualizer system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M of the inputs and the other XM inputs are received from external sources) may also be referred to as a virtualizer system (or virtualizer).

[0039] Throughout this disclosure, including the claims, the term "processor" is used broadly to refer to a system or device that is programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio or video or other image data). Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipelined processing on audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.

[0040] Throughout this disclosure, including the claims, the expression "analysis filterbank" is used broadly to refer to a system (e.g., a subsystem) configured to apply a transform (e.g., a time-domain to frequency-domain transform) to a time-domain signal to generate values ​​(e.g., frequency components) indicative of the content of the time-domain signal in each of a set of frequency bands. Throughout this disclosure, including the claims, the expression "filterbank domain" is used broadly to refer to the domain of frequency components generated by a transform or analysis filterbank (e.g., the domain in which such frequency components are processed). Examples of filterbank domains include (but are not limited to) the frequency domain, the quadrature mirror filter (QMF) domain, and the hybrid complex quadrature mirror filter (HCQMF) domain. Examples of transforms that may be applied by a analysis filterbank include (but are not limited to) the discrete cosine transform (DCT), the modified discrete cosine transform (MDCT), the discrete Fourier transform (DFT), and the wavelet transform. Examples of analysis filter banks include (but are not limited to) quadrature mirror filters (QMF), finite impulse response filters (FIR filters), infinite impulse response filters (IIR filters), crossover filters, and other suitable filters with multirate structures.

[0041] Throughout this disclosure, including the claims, the term "metadata" refers to data that is separate and distinct from the corresponding audio data (the audio content of a bitstream that also includes the metadata). The metadata is associated with audio data and indicates at least one feature or characteristic of the audio data (e.g., what type(s) of processing have been or should be performed on the audio data, or the trajectory of an object represented by the audio data). The association of the metadata with the audio data is time-synchronous. In this manner, current (most recently received or updated) metadata may indicate that the corresponding audio data contemporaneously has the indicated characteristics and / or includes the results of the indicated type of audio data processing.

[0042] Throughout this disclosure, including the claims, the terms "couple" or "coupled" are used to mean a direct or indirect connection. Thus, when a first device couples to a second device, the connection may be through a direct connection or through an indirect connection via other devices and connections.

[0043] Throughout this disclosure, including the claims, the following expressions have the following definitions.

[0044] Speaker and loudspeaker are used interchangeably to refer to any sound-producing transducer. This definition includes loudspeakers implemented as multiple transducers (e.g., a woofer and a tweeter).

[0045] Speaker Feed: An audio signal applied directly to a loudspeaker or to an amplifier and loudspeaker in series.

[0046] Channel (or "audio channel"): A monophonic audio signal. Such a signal can typically be rendered to be equivalent to applying the signal directly to a loudspeaker in a desired or nominal position. The desired position may be static, as is typically the case with a physical loudspeaker, or it may be dynamic.

[0047] Audio Program: A collection of one or more audio channels (at least one speaker channel and / or at least one object channel) and optionally associated metadata (e.g., metadata describing a desired spatial audio presentation).

[0048] Speaker Channel (or "Speaker Feed Channel"): An audio channel associated with a specified loudspeaker (in a desired or nominal position) or associated with a specified speaker zone within a defined speaker configuration. A speaker channel is rendered to be equivalent to applying the audio signal directly to the specified loudspeaker (in a desired or nominal position) or to a speaker within a specified speaker zone.

[0049] Object Channel: An audio channel that describes the sound emitted by an audio source (sometimes referred to as an audio "object"). Typically, an object channel determines a parametric audio source description (e.g., metadata describing the parametric audio source description is included within or provided along with the object channel). The source description may determine the sound emitted by the source (as a function of time), the apparent position of the source as a function of time (e.g., 3D spatial coordinates), and optionally at least one additional parameter characterizing the source (e.g., apparent source size or width).

[0050] Object-Based Audio Program: An audio program that includes a collection of one or more object channels (and optionally at least one speaker channel) and optionally associated metadata (e.g., metadata indicating the trajectory of the audio object(s) emitting the sound(s) represented by the object channels, or otherwise indicating the desired spatial audio presentation of the sound(s) represented by the object channels, or metadata indicating the identity of at least one audio object that is the source of the sound(s) represented by the object channels).

[0051] Rendering: The process of converting an audio program into one or more speaker feeds, or converting an audio program into one or more speaker feeds and then converting the speaker feeds into sound using one or more loudspeakers. (In the latter case, rendering is sometimes referred to herein as rendering "with" loudspeakers.) An audio channel can be trivially rendered ("at" the desired location) by applying the signal directly to a physical loudspeaker at the desired location. Alternatively, one or more audio channels can be rendered using one of a variety of virtualization techniques designed to be substantially equivalent (to a listener) to such a trivial rendering. In this latter case, each audio channel may be converted into one or more speaker feeds to be applied to a loudspeaker or loudspeakers at a known location, typically different from the desired location, so that the sound produced by the loudspeakers in response to the feed will be perceived to emanate from the desired location. Examples of such virtualization techniques include binaural rendering via headphones (e.g., using the "Dolby Headphone" process to simulate up to 7.1 channel surround sound for headphone wearers) and wave field synthesis.

[0052] The notation used herein to refer to a multi-channel audio signal as an "xy" or "xyz" channel signal indicates that the signal has "x" full-frequency speaker channels (corresponding to speakers nominally located in the horizontal plane of the ears of the intended listener), "y" LFE (or subwoofer) channels, and optionally also "z" full-frequency overhead speaker channels (corresponding to speakers located above the head of the intended listener, e.g., on or near the ceiling of the room).

[0053] The expression "IACC" is used herein to refer to the interaural cross-correlation coefficient in its usual sense, which is a measure of the difference between the arrival times of audio signals at the listener's ears, typically represented by a number ranging from a first value indicating that the arriving signals are equal in magnitude and exactly opposite in phase, through intermediate values ​​indicating that the arriving signals are dissimilar, to a maximum value indicating identical arriving signals with the same amplitude and phase.

[0054] Detailed Description of the Preferred Embodiments Many embodiments of the present invention are technically possible, and it will be clear to one skilled in the art how to implement them from this disclosure. Embodiments of the systems and methods of the present invention will be described with reference to Figures 2-14.

[0055] 2 is a block diagram of a system (20) including one embodiment of the headphone virtualization system of the present invention. The headphone virtualization system (sometimes referred to as a virtualizer) virtualizes N full frequency range channels (X1, ..., X N ) is configured to apply a binaural room impulse response (BRIR) to channels X1,…,X N Each of the channels (which may be speaker or object channels) corresponds to a particular source direction and distance relative to the intended listener, and the system of Figure 2 is configured to convolve each such channel with the BRIR for the corresponding source direction and distance.

[0056] The system 20 is coupled to receive an encoded audio program and then transmits it over N full frequency range channels (X1, ..., X N, 14, 15 of a virtualization system (having elements 12, ..., 14, 15, 16, 18 coupled as shown). The decoder may include subsystems (not shown in FIG. 2 ) coupled and configured to decode the programs, including by restoring the encoded programs from the virtualized program, and provide them to elements 12, ..., 14, 15 of a virtualization system (having elements 12, ..., 14, 15, 16, 18 coupled as shown). The decoder may include additional subsystems, some of which may perform functions unrelated to the virtualization functions performed by the virtualization system and some of which may perform functions related to the virtualization functions. For example, the latter functions may include extracting metadata from the encoded programs and providing the metadata to a virtualization control subsystem that uses the metadata to control elements of the virtualizer system.

[0057] Subsystem 12 (together with subsystem 15) is configured to convolve channel X1 with BRIR1 (the BRIR for the corresponding source direction and distance), and subsystem 14 (together with subsystem 15) is configured to convolve channel X N BRIR N 2. The output of each of the N-2 other BRIR subsystems 12, ..., 14, 15 is a time-domain signal including a left channel and a right channel. Summing elements 16 and 18 are coupled to the outputs of elements 12, ..., 14, 15. Summing element 16 is configured to combine (mix) the left channel outputs of the BRIR subsystems, and summing element 18 is configured to combine (mix) the right channel outputs of the BRIR subsystems. The output of element 16 is the left channel L of the binaural audio signal output from the virtualizer of FIG. 2, and the output of element 18 is the right channel R of the binaural audio signal output from the virtualizer of FIG. 2.

[0058] An important feature of exemplary embodiments of the present invention becomes apparent from comparing the FIG. 2 embodiment of the headphone virtualizer of the present invention with the conventional headphone virtualizer of FIG. 1. By way of comparison, the systems of FIG. 1 and FIG. 2 exhibit the same direct response and early reflection portions (i.e., the associated EBRIRs of FIG. 2) when each is presented with the same multi-channel audio input signal. i ) with BRIR i The full frequency range of each channel X of the input signal i Each BRIR applied by the system of Figure 1 or Figure 2 is configured to apply (although not necessarily with the same degree of success) i is the direct response and early reflection part (e.g., EBRIR1, ..., EBRIR applied by subsystems 12-14 in Figure 2). N The embodiment of FIG. 2 (and other exemplary embodiments of the present invention) uses multiple single-channel BRIRs, i.e., BRIRs. i We assume that the late reverberant part of the input signal can be shared across source directions and thus across all channels, and that the same late reverberation (i.e., common late reverberation) can be applied to the downmix of all full frequency range channels of the input signal. This downmix can be a monophonic (mono) downmix of all input channels, but alternatively it can be a stereo or multi-channel downmix obtained from the input channels (e.g., from a subset of the input channels).

[0059] More specifically, subsystem 12 of FIG. 2 is configured to convolve input signal channel X1 with EBRIR1 (the direct response and early reflection BRIR portion for the corresponding source direction), and subsystem 14 convolves input signal channel X N EBRIR NThe late reverberation subsystem 15 of FIG. 2 is configured to generate a mono downmix of all full-frequency range channels of the input signal and convolve the downmix with LBRIR (a common late reverberation for all of the downmixed channels). The output of each BRIR subsystem (subsystems 12, ..., 14, 15, respectively) of the virtualizer of FIG. 2 includes a left channel and a right channel (of the corresponding speaker channel or binaural signal generated from the downmix). The left channel outputs of the BRIR subsystems are combined (mixed) in summing element 16, and the right channel outputs of the BRIR subsystems are combined (mixed) in summing element 18.

[0060] Assuming that appropriate level adjustments and time alignments are implemented in subsystems 12, ..., 14, 15, summing element 16 can be implemented to simply sum corresponding left binaural channel samples (left channel outputs of subsystems 12, ..., 14, 15) to generate the left channel of the binaural output signal. Similarly, assuming that appropriate level adjustments and time alignments are implemented in subsystems 12, ..., 14, 15, summing element 18 can be implemented to simply sum corresponding right binaural channel samples (right channel outputs of subsystems 12, ..., 14, 15) to generate the right channel of the binaural output signal.

[0061] Subsystem 15 of FIG. 2 can be implemented in any of a variety of ways, but typically includes at least one feedback delay network configured to add common late reverberation to the monophonic downmix of the input signal channels presented to it. Typically, each of subsystems 12, ..., 14 includes a common late reverberation network configured to add a common late reverberation network to the monophonic downmix of the input signal channels (X i ) for the direct response and early reflection part (EBRIR) of the single channel BRIR. i), a common late reverberation is generated to emulate the collective macro-attributes of at least some (e.g., all) of the late reverberant portions of those single-channel BRIRs (whose "direct response and early reflection portions" are applied by subsystems 12, ..., 14). For example, one implementation of subsystem 15 has the same structure as subsystem 200 of Figure 3, including a bank of feedback delay networks (203, 204, ..., 205) configured to apply a common late reverberation to a monophonic downmix of the input signal channels presented to it.

[0062] Similarly, subsystems 12, ..., 14 of Figure 2 can be implemented in any of a variety of ways (in the time domain or filter bank domain), and the preferred implementation for any particular application will depend on various considerations such as (for example) performance, computation, and memory. In one exemplary implementation, each of subsystems 12, ..., 14 is configured to convolve the channel presented to it with an FIR filter corresponding to the direct and early responses associated with that channel. Gains and delays are appropriately set so that the outputs of subsystems 12, ..., 14 may be simply and efficiently combined with the output of subsystem 15.

[0063] Figure 3 is a block diagram of another embodiment of a headphone virtualization system of the present invention. The embodiment of Figure 3 is similar to the embodiment of Figure 2, in that two time-domain signals (left and right channels) are output from a direct response and early reflections processing subsystem 100, and two time-domain signals (left and right channels) are output from a late reverberation processing subsystem 200. A summing element 210 is coupled to the outputs of subsystems 100 and 200. Element 210 is configured to combine (mix) the left channel outputs of subsystems 100 and 200 to generate a left channel L of the binaural audio signal output from the virtualizer of Figure 3, and to combine (mix) the right channel outputs of subsystems 100 and 200 to generate a right channel R of the binaural audio signal output from the virtualizer of Figure 3. Assuming appropriate level adjustment and time alignment are implemented in subsystems 100 and 200, element 210 can be implemented to simply sum corresponding left channel samples output from subsystems 100 and 200 to generate the left channel of the binaural output signal, and to simply sum corresponding right channel samples output from subsystems 100 and 200 to generate the right channel of the binaural output signal.

[0064] In the system shown in Figure 3, channel X of the multi-channel audio input signal i are directed to and processed in two parallel processing paths: one through the direct response and early reflection processing subsystem 100, and the other through the late reverberation processing subsystem 200. The system of FIG. i BRIR i Each BRIR is configured to apply ican be decomposed into two parts: a direct response and early reflections portion (applied by subsystem 100) and a late reverberation portion (applied by subsystem 200). In operation, direct response and early reflections processing subsystem 100 thus generates the direct response and early reflections portion of the binaural audio signal that is output from the virtualizer, and late reverberation processing subsystem ("late reverberation generator") 200 thus generates the late reverberation portion of the binaural audio signal that is output from the virtualizer. The outputs of subsystems 100 and 200 are mixed (by summing subsystem 210) to generate a binaural audio signal that is typically presented from subsystem 210 to a rendering system (not shown) where it undergoes binaural rendering for playback over headphones.

[0065] Typically, when rendered and played through a pair of headphones, a typical binaural audio signal output from element 210 is perceived at a listener's eardrums as sound from "N" loudspeakers (where N > 2, and N is typically 2, 5, or 7) located in any of a wide variety of positions, including positions in front of, behind, and above the listener. Reproduction of the output signals produced in operation of the system of Figure 3 can give the listener the experience of sound coming from more than two (e.g., five or seven) "surround" sources, at least some of which may be virtual.

[0066] Direct response and early reflections processing subsystem 100 can be implemented in any of a variety of ways (in the time domain or filter bank domain), and the preferred implementation for any particular application will depend on various considerations such as (for example) performance, computation, and memory. In one exemplary implementation, subsystem 100 is configured to convolve each channel presented to it with an FIR filter corresponding to the direct and early responses associated with that channel. Gains and delays are appropriately set so that the output of subsystem 100 may be simply and efficiently combined with the output of subsystem 200 (in element 210).

[0067] As shown in FIG. 3, the late reverberation generator 200 includes a downmix subsystem 201, an analysis filter bank 202, a bank of FDNs (FDNs 203, 204, ..., 205), and a synthesis filter bank 207, coupled as shown. Subsystem 201 is configured to downmix the channels of a multi-channel input signal to a mono downmix, and analysis filter bank 202 is configured to apply a transform to the mono downmix and split the mono downmix into "K" frequency bands, where K is an integer. The filter bank domain values ​​(output from filter bank 202) in each different frequency band are submitted to a different one of FDNs 203, 204, ..., 205 (there are "K" FDNs, each coupled and configured to apply the late reverberation portion of the BRIR to the filter bank domain values ​​submitted thereto). The filter bank domain values ​​are preferably decimated in time to reduce the computational complexity of the FDNs.

[0068] In principle, each input channel (to subsystem 100 and subsystem 201 in FIG. 3 ) can be processed by its own FDN (or bank of FDNs) to simulate the late reverberation portion of its BRIR. Despite the fact that the late reverberation portions of BRIRs associated with different sound source positions are typically very different in terms of the root-mean-square of their impulse responses, their statistical attributes, such as their average power spectrum, their energy decay structure, modal density, peak density, etc., are often very similar. Therefore, since the late reverberation portions of a set of BRIRs are typically perceptually very similar across channels, it is possible to use a common FDN or bank of FDNs (e.g., FDNs 203, 204, ..., 205) to simulate the late reverberation portions of two or more BRIRs. In a typical embodiment, such a common FDN (or bank of FDNs) is used, and its input consists of one or more downmixes constructed from the input channels. In the exemplary implementation of FIG. 2, the downmix is ​​a monophonic downmix of all input channels (presented at the output of subsystem 201).

[0069] Referring to the embodiment of FIG. 2, each of FDNs 203, 204, ..., 205 is implemented in the filterbank domain and is coupled and configured to process different frequency bands of values ​​output from analysis filterbank 202 to generate left and right reverberated signals for each band. For each band, the left reverberated signal is a sequence of filterbank domain values, and the right reverberated signal is another sequence of filterbank domain values. Synthesis filterbank 207 applies a frequency-domain to time-domain transform to 2K sequences of filterbank domain values ​​(e.g., frequency components in the QMF domain) and assembles the transformed values ​​into a left-channel time-domain signal (representing the audio content of the mono downmix with late reverberation) and a right-channel time-domain signal (also representing the audio content of the mono downmix with late reverberation). These left-channel and right-channel signals are output to element 210.

[0070] In a typical implementation, each of FDNs 203, 204, ..., 205 is implemented in the QMF domain, and filter bank 202 transforms the mono downmix from subsystem 201 into the QMF domain (e.g., the hybrid complex quadrature mirror filter (HCQMF) domain), such that the signal presented from filter bank 202 to the input of each of FDNs 203, 204, ..., 205 is a sequence of QMF-domain frequency components. In such an implementation, the signal presented from filter bank 202 to FDN 203 is a sequence of QMF-domain frequency components in a first frequency band, the signal presented from filter bank 202 to FDN 204 is a sequence of QMF-domain frequency components in a second frequency band, and the signal presented from filter bank 202 to FDN 205 is a sequence of QMF-domain frequency components in the "K"th frequency band. When analysis filterbank 202 is so implemented, synthesis filterbank 207 applies a QMF-domain to time-domain transform to the 2K sequences of output QMF-domain frequency components from the FDN, producing left-channel and right-channel post-reverberated time-domain signals that are output to element 210.

[0071] 3, if K=3, there are six inputs to synthesis filter bank 207 (left and right channels, each comprising frequency-domain or QMF-domain samples output from FDNs 203, 204, and 205) and two outputs from 207 (left and right channels, each consisting of time-domain samples). In this example, filter bank 207 is typically implemented as two synthesis filter banks: one (representing the three left channels from FDNs 203, 204, and 205) configured to generate the time-domain left channel signal output from filter bank 207, and a second (representing the three right channels from FDNs 203, 204, and 205) configured to generate the time-domain right channel signal output from filter bank 207.

[0072] Optionally, control subsystem 209 is coupled to each of FDNs 203, 204, ..., 205 and configured to provide control parameters to each of those FDNs to determine the late reverberant parts (LBRIR) applied by subsystem 200. Examples of such control parameters are described below. In some implementations, control subsystem 209 may be operable in real time (i.e., in response to user commands provided thereto by an input device) to implement real-time variation of the late reverberant parts (LBRIR) applied by subsystem 200 to the monophonic downmix of the input channels.

[0073] For example, if the input signal to the system of Figure 2 is a 5.1 channel signal (whose full frequency range channels are in the following channel order: L, R, C, Ls, Rs), all full frequency range channels have the same source distance, and the downmix subsystem 201 can be implemented as the following downmix matrix, which simply sums the full frequency range channels to form a mono downmix:

[0074]

number

[0075]

number

[0076]

number

[0077]

number

[0078] Next, a discussion of the specific implementations of downmix subsystem 201 and subsystems 100 and 200 of the virtualizer of FIG. 3 is provided.

[0079] The downmix process implemented by subsystem 201 depends on the source distance (between the sound source and the assumed listener position) for each channel to be downmixed and the treatment of the direct response. d teeth: t d =d / v s where d is the distance between the sound source and the listener, and v s is the speed of sound. Furthermore, the gain of the direct response is proportional to 1 / d. If these rules are preserved in treating the direct responses of channels with different source distances, subsystem 201 can implement a straightforward downmix of all channels, since the delay and level of late reverberation are generally insensitive to the source position.

[0080] For practical reasons, the virtualizer (e.g., virtualizer subsystem 100 of FIG. 3) may be implemented to time-align the direct responses for input channels with different source distances. To preserve the relative delay between the direct response and the late reverberation for each channel, the channel with source distance d is time-aligned by (dmax-d) / v before being downmixed with the other channels. s where dmax represents the maximum possible source distance.

[0081] The virtualizer (e.g., virtualizer subsystem 100 of FIG. 3) may also be implemented to compress the dynamic range of the direct response. For example, the direct response for a channel with source distance d is d -1 Instead of factor d -α where 0≦α≦1. To preserve the level difference between the direct response and the late reverberation, the downmix subsystem 201 scales the channel with source distance d by a factor d before downmixing it with the other scaled channels. 1-α may need to be implemented to scale by

[0082] The feedback delay network of FIG. 4 is an exemplary implementation of FDN 203 (or 204 or 205) of FIG. 3. The system of FIG. 4 includes four reverberation tanks (each with gain stage g i and delay line z -ni ), although variations of this system (and other FDNs used in virtualizer embodiments of the present invention) implement more or less than four reverberation tanks.

[0083] The FDN of FIG. 4 includes an input gain element 300, an all-pass filter (APF) 301 coupled to the output of element 300, summing elements 302, 303, 304, and 305 coupled to the output of APF 301, and four reverberation tanks (each reverberation tank is coupled to the output of a different one of elements 302, 303, 304, and 305) respectively. k (one of elements 306) and a delay line z coupled to it. -Mk (one of elements 307) and its associated gain element 1 / g k (one of elements 309, 0≦k−1≦3). Unitary matrix 308 is coupled to the output of delay line 307 and configured to provide a feedback output to each second input of elements 302, 303, 304, and 305. The outputs of two of gain elements 309 (the first and second reverberation tanks) are presented to inputs of summing element 310, the output of which is presented to one input of output mixing matrix 312. The outputs of the other two of gain elements 309 (the third and fourth reverberation tanks) are presented to inputs of summing element 311, the output of which is presented to the other input of output mixing matrix 312.

[0084] Element 302 is a delay line z -n1 The output of the matrix 308 corresponding to -n1 Element 303 is configured to apply feedback from the output of delay line z -n2The output of the matrix 308 corresponding to -n2 Element 304 is configured to apply feedback from the output of delay line z -n3 The output of the matrix 308 corresponding to -n3 Element 305 is configured to apply feedback from the output of delay line z -n4 The output of the matrix 308 corresponding to -n4 (applying feedback from the output of

[0085] The input gain element 300 of the FDN of Figure 4 is coupled to receive one frequency band of the converted monophonic downmix signal (filterbank domain signal) output from the analysis filterbank 202 of Figure 3. The input gain element 300 applies a gain (scaling) factor G to the filterbank domain signal presented to it. in Collectively, the scaling factor G for all frequency bands (implemented by all FDNs 203, 204, ..., 205 in Figure 3) is applied. in controls the spectral shaping and level of the late reverberation. The input gain G in Setting a standard often takes into account the following goals: BRIR direct-to-late ratio (DLR) applied to each channel to match the real room; Necessary low frequency attenuation to mitigate excessive combing artifacts and / or low frequency rumbling; Matching diffuse-field spectral envelopes.

[0086] If the direct response (applied by subsystem 100 of FIG. 3) provides unitary gain in all frequency bands, then the specific DLR (power ratio) is: G in=sqrt(ln(10 6 ) / (T60*DLR)) So that G in where T60 is the reverberation decay time, defined as the time it takes for the reverberation to decay by 60 dB (which is determined by the reverberation delay and reverberation gain discussed below), and "ln" represents the natural logarithm function.

[0087] Input gain factor G in may depend on the content being processed. One application of such content dependence is to ensure that the energy of the downmix in each time / frequency segment is equal to the sum of the energies of the individual channel signals being downmixed, regardless of any correlation that exists between the input channel signals. In that case, the input gain factor is

number

[0088] In a typical QMF-domain implementation of the FDN of FIG. 4, the signal presented to the reverberation tank input from the output of all-pass filter (APF) 301 is a sequence of QMF-domain frequency components. To produce a more natural-sounding FDN output, APF 301 is applied to the output of gain element 300 to introduce phase diversity and increased echo density. Alternatively or additionally, one or more all-pass filters may be applied to individual inputs to downmix subsystem 201 (of FIG. 3) before the inputs are downmixed in subsystem 201 and processed by the FDN, or in the reverberation tank feedforward or feedback paths depicted in FIG. 4 (e.g., delay line z in each reverberation tank). -Mk ) or may be applied to the output of the FDN (i.e., to the output of output matrix 312).

[0089] Reverberation Tank Delay z -ni When implementing i should be relatively prime. The sum of the delays should be large enough to provide sufficient modal density to avoid an artificial sounding output, yet the shortest delay should be short enough to avoid excessive time gaps between the late reverberation and other components of the BRIR.

[0090] Typically, the reverberation tank output is initially panned to either the left or right binaural channel. Usually, the sets of reverberation tank outputs panned to the two binaural channels are equal in number and mutually exclusive. It is also desirable to balance the timing of the two binaural channels. Thus, if the reverberation tank output with the shortest delay goes to one binaural channel, the reverberation tank output with the second shortest delay goes to the other channel.

[0091] The reverberation tank delay can be varied across the frequency band to vary the modal density as a function of frequency. Generally, lower frequency bands require higher modal densities and therefore longer reverberation tank delays.

[0092] Reverberation Tank Gain g i The amplitude of and the reverberation tank delay jointly determine the reverberation delay time of the FDN in Figure 4: T 60 =-3n i / log 10 (|g i |) / F FRM where F FRM is the frame rate of the filter bank 202 (in FIG. 3). The phase of the reverberation tank gain introduces a fractional delay to overcome problems associated with the reverberation tank delay being quantized into the downsample factor lattice of the filter bank.

[0093] The unitary feedback matrix 308 provides an even mix between the reverberation tanks in the feedback path.

[0094] To equalize the level of the reverberation tank output, gain element 309 provides a normalized gain 1 / |g i | is applied to the output of each reverberation tank to remove the level effect of the reverberation tank gain while preserving the fractional delay introduced by its phase.

[0095] The output mixing matrix 312 (matrix M out The output mixing matrix 312 (also identified as |Coh|) is a 2x2 matrix configured to mix the unmixed binaural channels from the initial panning (the outputs of elements 310 and 311, respectively) to achieve output left and right binaural channels (L and R signals presented at the output of matrix 312) with the desired interaural coherence. The unmixed binaural channels are largely uncorrelated after the initial panning, since they do not contain any common reverberation tank outputs. If the desired interaural coherence is Coh, where |Coh|≦1, then the output mixing matrix 312 is

number

number

[0096] For an FDN for each individual frequency band in the virtualizer of the present invention, if the target acoustic attributes T60, Coh, and DLR defined above are known, each FDN (each FDN may have the structure shown in FIG. 4) can be configured to achieve the target attributes. In particular, in some embodiments, the input gain (G in ) and the gain and delay of the reverberation tank (g i and n i ) and the output matrix M out can be set (e.g., by control values ​​presented to it by the control subsystem 209 of FIG. 3). In practice, it is often sufficient to set the frequency-dependent attributes by a model with simple control parameters to generate natural-sounding late reverberation that matches a particular acoustic environment.

[0097] Next, the target reverberation decay time (T) for the FDN for each particular frequency band of an embodiment of the virtualizer of the present invention is calculated. 60 ) for each of a small number of frequency bands is the target reverberation decay time (T 60 ) the level of the FDN response decays exponentially with time. T 60 is inversely proportional to the decay factor df (defined as the dB attenuation per unit time), i.e.: T 60 =60 / df is.

[0098] The decay factor df is frequency dependent and generally increases linearly on a logarithmic frequency scale. Therefore, the reverberation decay time is also a function of frequency and generally decreases as frequency increases. Therefore, the decay time T for two frequency points is 60 Once the value of is determined (for example, set), T 60 The curve is determined, for example, at frequency point f A and f B The reverberation decay time for each is T 60,A and T 60,B If so, T 60 The curve is defined as follows:

[0099]

number

[0100] Next, we will describe an example of how the target interaural coherence (Coh) for the FDN for each specific frequency band of an embodiment of the virtualizer of the present invention can be achieved by setting a few control parameters. The interaural coherence (Coh) of late reverberation roughly follows the pattern of a diffuse sound field. It is found that the crossover frequency f C It can be modeled by a sinc function up to the crossover frequency and a constant above the crossover frequency. A simple model for the Coh curve is:

[0101]

number

[0102] Next, we provide an example of how a target direct-to-late ratio (DLR) for the FDN for each particular frequency band of an embodiment of the virtualizer of the present invention can be achieved by setting a few control parameters. The direct-to-late ratio (DLR) in dB generally increases linearly with logarithmic frequency, and the DLR 1K It is controlled by setting the DLR (in dB at 1 kHz) and DLRslope (in dB per decade of frequency). However, a low DLR in the low frequency range often leads to excessive combing artifacts. To mitigate these artifacts, two correction mechanisms are added to control the DLR: Minimum DLR floor, DLRmin (in dB); and transition frequency f T and the slope of the attenuation curve below that HPF slope A high-pass filter defined by the frequency (in dB per decade of frequency).

[0103] The resulting DLR curve, in dB, is defined as follows:

[0104]

number

[0105] Variations of the embodiments disclosed herein have one or more of the following features: The virtualizer of the present invention is implemented in the time domain or has a hybrid implementation with FDN-based impulse response acquisition and FIR-based signal filtering; The inventive virtualizer is implemented to allow the application of energy compensation as a function of frequency during the downmix stage that generates the downmixed input signal for the late reverberation processing subsystem; The virtualizer of the present invention is implemented to allow manual or automatic control of the late reverberation attributes that are applied in response to external factors (ie, in response to the setting of control parameters).

[0106] For applications where system latency is critical and the delays introduced by the decomposition and synthesis filter banks are prohibitive, the filter bank domain FDN structures of exemplary embodiments of the virtualizer of the present invention can be converted to the time domain, and each FDN structure can be implemented in the time domain in certain embodiments of the present virtualizer. In the time domain implementation, the input gain factor (Gin ), reverberation tank gain (g i ) and normalized gain (1 / |g i The subsystem applying |) is replaced by a filter with a similar magnitude response to allow frequency-dependent control. The output mixing matrix (M out ) is also replaced by a matrix of filters. Unlike other filters, the phase response of this matrix of filters is crucial, as it can affect power conservation and interaural coherence. The reverberation tank delay in the time-domain implementation may need to be slightly changed (from its value in the filter-bank domain implementation) to avoid sharing the filter-bank stride as a common factor. Due to various constraints, the performance of the time-domain implementation of the FDN of our virtualizer may not exactly match that of the filter-bank implementation.

[0107] With reference to Figure 8, we now describe a hybrid (filterbank domain and time domain) implementation of our inventive late reverberation processing subsystem of our virtualizer. This hybrid implementation of our inventive late reverberation processing subsystem is a variation on the late reverberation processing subsystem 200 of Figure 4, implementing FDN-based impulse response capture and FIR-based signal filtering.

[0108] Figure 8 includes elements 201, 202, 203, 204, 205, and 207 that are identical to the identically numbered elements of subsystem 200 of Figure 3. The above description of these elements will not be repeated with reference to Figure 8. In the embodiment of Figure 8, unit impulse generator 211 is coupled to provide the input signal (pulse) to analysis filter bank 202. LBRIR filter 208 (mono input, stereo output), implemented as an FIR filter, applies the appropriate late reverberation portion of the BRIR (LBRIR) to the monophonic downmix output from subsystem 201. Elements 211, 202, 203, 204, 205, and 207 are thus processing sidechains to LBRIR filter 208.

[0109] Whenever the setting of the late reverberation part LBRIR is modified, impulse generator 211 is operated to present a unit impulse to element 202, and the resulting output from filter bank 207 is captured and presented to filter 208 (to set filter 208 to apply the new LBRIR determined by the output of filter bank 207). To accelerate the time lapse between the LBRIR setting change and the time the new LBRIR takes effect, samples of the new LBRIR can begin replacing the old LBRIR as they become available. To reduce the inherent latency of the FDN, leading zeros in the LBRIR can be discarded. These options provide flexibility and allow the hybrid implementation to potentially offer performance improvements (over that offered by a filter bank domain implementation) at the cost of additional computations from FIR filtering.

[0110] For applications where system latency is critical but computational power is not a major issue, a sidechain filterbank domain late reverberator (e.g., implemented by elements 211, 202, 203, 204, ..., 205 in Figure 8) can be used to capture the effective FIR impulse response applied by filter 208. FIR filter 208 implements this captured FIR response and can apply it directly to the mono downmix of the input channels (during input channel virtualization).

[0111] The various FDN parameters, and thus the resulting late reverberation attributes, can be manually tuned and then incorporated into embodiments of the late reverberation processing subsystem of the present invention as fixed configurations, for example, by one or more presets that can be adjusted by a user of the system (e.g., by manipulating the control subsystem 209 of FIG. 3). However, given the high-level description of late reverberation, its relationship to the FDN parameters, and the ability to modify its behavior, a wide variety of methods are envisioned for controlling various embodiments of FDN-based late reverberation processors, including (but not limited to):

[0112] 1. An end user may manually control the FDN parameters, for example, through a user interface on a display (e.g., implemented by an embodiment of the control subsystem 209 of FIG. 3), or may switch between presets using physical controls (e.g., implemented by an embodiment of the control subsystem 209 of FIG. 3). In this way, the end user can adapt the room simulation according to preferences, environment, or content.

[0113] 2. The author of the audio content to be virtualized may provide settings or desired parameters that are conveyed along with the content itself, for example, by metadata provided along with the input audio signal. Such metadata may be parsed and used (e.g., by an embodiment of control subsystem 209 of FIG. 3) to control the associated FDN parameters. Thus, the metadata may indicate attributes such as reverberation time, reverberation level, direct-to-reverberant ratio, etc., and these attributes may be time-varying and indicated by the time-varying metadata.

[0114] 3. The playback device may be aware of its location or environment through one or more sensors. For example, a mobile device may use a GSM network, a global positioning system (GPS), known WiFi access points, or any other location service to determine where the device is located. Data indicative of the location and / or environment may then be used (e.g., by an embodiment of control subsystem 209 in FIG. 3 ) to control associated FDN parameters. In this manner, the FDN parameters may be modified in response to the device's location, e.g., to mimic the physical environment.

[0115] 4. Cloud services or social media may be used to derive the most common settings used by consumers in a certain environment in relation to the location of the playback device. Furthermore, users may upload their current settings, associated with their (known) location, to cloud or social media services and make them available to other users or themselves.

[0116] 5. The playback device may include other sensors such as cameras, light sensors, microphones, accelerometers, gyroscopes, etc. to determine the user's activity and the environment in which the user is located, in order to optimize the FDN parameters for that particular activity and / or environment.

[0117] 6. FDN parameters may be controlled by audio content. Audio classification algorithms or manually annotated content may indicate whether segments of audio contain speech, music, sound effects, silence, etc. FDN parameters may be adjusted according to such labels. For example, the direct-to-reverberant ratio may be reduced for dialogue to improve dialogue intelligibility. Furthermore, video analysis may be used to determine the location of the current video segment, and FDN parameters may be adjusted accordingly to better simulate the environment depicted in the video. And / or 7. A solid-state playback system may use a different FDN setting than a mobile device. For example, the setting may be device dependent. A solid-state system in a living room may simulate a typical (highly reverberant) living room scenario with distant sources, while a mobile device may render the content closer to the listener.

[0118] Some implementations of the virtualizer of the present invention include an FDN (e.g., the FDN implementation of FIG. 4) configured to apply fractional delays in addition to integer sample delays. For example, in one such implementation, a fractional delay element is connected in each reverberation tank in series with a delay line that adds an integer delay equal to an integer number of sample periods (e.g., each fractional delay element is positioned after or otherwise in series with one of the delay lines). The fractional delay can be approximated in each frequency band by a phase shift (unit complex multiplication) corresponding to a fraction of the sample period f=τ / T, where f is the delay fraction, τ is the desired delay for that band, and T is the sample period for that band. How to add fractional delays in the context of applying reverberation in the QMF domain is well known.

[0119] In a first class of embodiments, the present invention is a headphone virtualization method that generates binaural signals in response to a set of channels (e.g., each of those channels or each of the full frequency range channels) of a multi-channel audio input signal. The method includes: (a) applying a binaural room impulse response (BRIR) to each channel of the set (e.g., by convolving each channel of the set with the BRIR corresponding to the channel in subsystems 100 and 200 of FIG. 3 or in subsystems 12, ..., 14, and 15 of FIG. 2), thereby generating a filtered signal (e.g., the output of subsystems 100 and 200 of FIG. 3 or the output of subsystems 12, ..., 14, and 15 of FIG. 2), by using at least one feedback delay network (e.g., FDNs 203, 204, ..., 205 of FIG. 3) to add a common late reverberation to a downmix (e.g., a monophonic downmix) of the channels of the set; and (b) combining the filtered signals (e.g., in subsystem 210 of FIG. 3 or a subsystem including elements 16 and 18 of FIG. 2) to generate a binaural signal. Typically, a bank of FDNs is used to add the common late reverberation to the downmix (e.g., each FDN adds late reverberation to a different frequency band). Typically, step (a) involves applying to each channel of the set (e.g., in subsystem 100 of FIG. 3 or subsystems 12, ..., 14 of FIG. 2) the "direct response and early reflection" portion of a single-channel BRIR for that channel, wherein the common late reverberation is generated to emulate the collective macro-attributes of the late reverberation portions of at least some (e.g., all) of the single-channel BRIRs.

[0120] In typical implementations of the first class, each FDN is implemented in the hybrid complex quadrature mirror filter (HCQMF) domain or the quadrature mirror filter (QMF) domain. In some such embodiments, the frequency-dependent spatial acoustic attributes of the binaural signal are controlled by controlling the configuration of each FDN used to add late reverberation (e.g., using control subsystem 209 in FIG. 3 ). Typically, for efficient binaural rendering of the audio content of a multichannel signal, a monophonic downmix of the channels (e.g., the downmix generated by subsystem 201 in FIG. 3 ) is used as input to the FDN. Typically, the downmix process is controlled based on the source distance for each channel (i.e., the distance between the assumed source of the channel's audio content and the assumed user position) and relies on the treatment of the direct response corresponding to the source distance to preserve the temporal and level structure of each BRIR (i.e., each BRIR determined by the direct response and early reflection portions of the single-channel BRIR for a given channel and the common late reverberation for the downmix including that channel). The channels to be downmixed can be time-aligned and scaled in various ways during the downmix, but the proper level and time relationship between the direct response, early reflections, and the common late reverberation part of the BRIR for each channel should be maintained. In embodiments using a single FDN bank to generate the common late reverberation part for all downmixed channels (to generate the downmix), the proper gain and delay need to be applied (for each downmixed channel) during the generation of the downmix.

[0121] Typical embodiments of this class include adjusting FDN coefficients corresponding to frequency-dependent attributes (e.g., reverberation decay time, interaural coherence, modal density, and direct-to-late ratio), which allows for better matching of the acoustic environment and a more natural-sounding output.

[0122] In a second class of embodiments, the present invention is a method for generating a binaural signal in response to a multi-channel audio input signal by applying a binaural room impulse response (BRIR) to each channel of a set of channels of the input signal (e.g., each of the channels of the input signal or each full frequency range channel of the input signal) (e.g., by convolving each channel with a corresponding BRIR), processing each channel of the set in a first processing path (e.g., implemented by subsystem 100 of FIG. 3 or subsystems 12, ..., 14 of FIG. 2) configured to model and apply to each channel the direct response and early reflections (e.g., EBRIR applied by subsystems 12, 14, or 15 of FIG. 2) of a single-channel BRIR for that channel, and processing a downmix (e.g., a monophonic downmix) of the channels of the set in a second processing path (e.g., implemented by subsystem 200 of FIG. 3 or subsystem 15 of FIG. 2) parallel to the first processing path. The second processing path is configured to model and apply a common late reverberation (e.g., LBRIR applied by subsystem 15 in FIG. 2 ) to the downmix. Typically, the common late reverberation emulates the collective macro-attributes of at least some (e.g., all) of the late reverberant portions of the single-channel BRIR. Typically, the second processing path includes at least one FDN (e.g., one FDN for each of multiple frequency bands). Typically, the mono downmix is ​​used as the input to all reverberation tanks of each FDN implemented by the second processing path. Typically, a mechanism for systematic control of the macro-attributes of each FDN is provided (e.g., control subsystem 209 in FIG. 3 ) to better simulate the acoustic environment and produce a more natural-sounding binaural virtualization. Because most such macro-attributes are frequency-dependent, each FDN is typically implemented in the hybrid complex quadrature mirror filter (HCQMF) domain, the frequency domain, or another filter bank domain, with a different FDN used for each frequency band.The primary benefit of implementing the FDN in the filter bank domain is that it allows for the application of reverberation with frequency-dependent reverberation characteristics. In various embodiments, the FDN is implemented in any of a wide variety of filter bank domains using any of a variety of filter banks, including, but not limited to, quadrature mirror filters (QMF), finite impulse response filters (FIR filters), infinite impulse response filters (IIR filters), or crossover filters.

[0123] Some embodiments of the first class (and the second class) implement one or more of the following features.

[0124] 1. An FDN implementation in the filter bank domain (e.g., hybrid complex quadrature mirror filter domain) (e.g., the FDN implementation in Figure 4) or an FDN implementation in the hybrid filter bank domain and a late reverberation filter implementation in the time domain (e.g., the structure described with reference to Figure 8). This typically allows independent adjustment of the FDN parameters and / or settings for each frequency band (which allows for simple and flexible control of frequency-dependent acoustic attributes), for example by providing the ability to vary the reverberation tank delay in different bands to change the modal density as a function of frequency.

[0125] 2. The particular downmix process used to generate the processed downmixed (e.g., monophonic downmixed) signal in the second processing path (from the multichannel input audio signal) depends on the source distance of each channel and the treatment of the direct response to maintain the proper level and timing relationship between the direct and late responses.

[0126] 3. An all-pass filter (e.g., APF 301 in Figure 4) is applied in a second processing path (e.g., at the input or output of the bank of FDNs) to introduce phase diversity and increased echo density without changing the spectrum and / or timbre of the resulting reverberation.

[0127] 4. To overcome the problems associated with quantized delays in the downsample-factor grid, fractional delays are implemented in the feedback path of each FDN in a complex-valued multi-rate structure.

[0128] 5. In FDN, the reverberation tank outputs are linearly mixed directly into the binaural channels (e.g., by matrix 312 in FIG. 4) using output mix coefficients that are set based on the desired interaural coherence in each frequency band. Optionally, the mapping of reverberation tanks to binaural output channels alternates across frequency bands to achieve balanced delays between the binaural channels. Optionally, a normalization factor is applied to the reverberation tank outputs to equalize their levels while preserving fractional delays and overall power.

[0129] 6. Frequency-dependent reverberation decay time is controlled by setting the proper combination of reverberation tank delay and gain in each frequency band to simulate a real room.

[0130] 7. For each frequency band (e.g., at either the input or output of the associated processing path), one scaling factor is applied (e.g., by elements 306 and 309 in Figure 4), which: Controlling the frequency-dependent direct-to-late ratio (DLR) to match the DLR of the actual room (a simple model may be used to calculate the required scaling factor based on the target DLR and reverberation decay time, e.g., T60); providing low frequency attenuation to mitigate excessive combing artifacts; and / or Apply diffuse field spectral shaping to the FDN response.

[0131] 8. Simple parametric models are implemented (e.g., by control subsystem 209 in Figure 3) to control essential frequency-dependent attributes of late reverberation, such as reverberation decay time, interaural coherence, and / or direct-to-late ratio.

[0132] In some embodiments (e.g., for applications where system latency is critical and the delays introduced by the decomposition and synthesis filter banks are prohibitive), the filter bank domain FDN structure of the exemplary embodiment of the inventive system (e.g., the FDN of FIG. 4 for each frequency band) is replaced by an FDN structure implemented in the time domain (e.g., the FDN 220 of FIG. 10, which may be implemented as shown in FIG. 9). In the time domain embodiment of the inventive system, the input gain factor (G in ), reverberation tank gain (g i ) and normalized gain (1 / |g i The subsystem in a filter bank domain implementation that applies |) is replaced by a time-domain filter (and / or gain element) to allow frequency-dependent control. The output mixing matrix of a typical filter bank domain implementation (e.g., output mixing matrix 312 of FIG. 4) is replaced (in a typical time-domain embodiment) by a set of time-domain filters' outputs (e.g., elements 500-503 in the FIG. 11 implementation of element 424 of FIG. 9). Unlike other filters in a typical time-domain embodiment, the phase response of this set of filters' outputs is typically critical (because power conservation and interaural coherence can be affected by the phase response). In some time-domain embodiments, the reverberation tank delays are varied (e.g., slightly varied) from their values ​​in the corresponding filter bank domain implementation (e.g., to avoid sharing the filter bank stride as a common factor).

[0133] FIG. 10 is a block diagram of an embodiment of a headphone virtualization system of the present invention similar to FIG. 3 , except that elements 202-207 of FIG. 3 are replaced in the system of FIG. 10 by a single FDN 220 implemented in the time domain (e.g., FDN 220 of FIG. 10 may be implemented similarly to the FDN of FIG. 9 ). In FIG. 10 , two time-domain signals (left and right channels) are output from direct response and early reflection processing subsystem 100, and two time-domain signals (left and right channels) are output from late reverberation processing subsystem 221. A summing element 210 is coupled to the outputs of subsystems 100 and 200. Element 210 is configured to combine (mix) the left channel outputs of subsystems 100 and 221 to generate the left channel L of the binaural audio signal output from the virtualizer of FIG. 10 , and to combine (mix) the right channel outputs of subsystems 100 and 221 to generate the right channel R of the binaural audio signal output from the virtualizer of FIG. 10 . Assuming appropriate level adjustment and time alignment are implemented in subsystems 100 and 221, element 210 can be implemented to simply sum corresponding left channel samples output from subsystems 100 and 221 to generate the left channel of the binaural output signal, and to simply sum corresponding right channel samples output from subsystems 100 and 221 to generate the right channel of the binaural output signal.

[0134] In the system of Figure 10, (channel X i The multi-channel audio input signal (having a direct response and early reflections processing subsystem 100) is directed to and processed in two parallel processing paths: one through the direct response and early reflections processing subsystem 100, and the other through the late reverberation processing subsystem 221. The system of FIG. 10 provides a i BRIR i Each BRIR is configured to apply ican be decomposed into two parts: a direct response and early reflections portion (applied by subsystem 100) and a late reverberation portion (applied by subsystem 221). In operation, direct response and early reflections processing subsystem 100 thus generates the direct response and early reflections portion of the binaural audio signal that is output from the virtualizer, and late reverberation processing subsystem ("late reverberation generator") 221 thus generates the late reverberation portion of the binaural audio signal that is output from the virtualizer. The outputs of subsystems 100 and 221 are mixed (by subsystem 210) to generate a binaural audio signal that is typically presented from subsystem 210 to a rendering system (not shown) where it undergoes binaural rendering for playback over headphones.

[0135] The downmix subsystem 201 (of the late reverberation processing subsystem 221) is configured to downmix the channels of the multi-channel input signal to a mono downmix (which is a time-domain signal), and the FDN 220 is configured to apply the late reverberation part to the mono downmix.

[0136] With reference to Figure 9, we now describe an example of a time-domain FDN that can be used as FDN 220 of the virtualizer of Figure 10. The FDN of Figure 9 includes an input filter 400 coupled to receive a mono downmix of all channels of a multi-channel audio input signal (e.g., generated by subsystem 201 of the system of Figure 10). The FDN of Figure 9 includes an all-pass filter (APF) 401 (corresponding to APF 301 of Figure 4) coupled to the output of filter 400, an input gain element 401A coupled to the output of filter 401, summing elements 402, 403, 404, and 405 (which correspond to summing elements 302, 303, 304, and 305 of Figure 4) coupled to the output of element 401A, and four reverberation tanks. Each reverberation tank is coupled to the output of a different one of elements 402, 403, 404 and 405 and has one of reverberation filters 406 and 406A, 407 and 407A, 408 and 408A and 409 and 409A coupled to it, one of delay lines 410, 411, 412 and 413 (corresponding to delay line 307 in Figure 4), and one of gain elements 417, 418, 419 and 420 coupled to the output of one of these delay lines.

[0137] A unitary matrix 415 (corresponding to unitary matrix 308 of FIG. 4 and typically implemented identically to matrix 308) is coupled to the outputs of delay lines 410, 411, 412, and 413. Matrix 415 is configured to provide a feedback output to the second input of each of elements 402, 403, 404, and 405.

[0138] When the delay (n1) added by line 410 is shorter than the delay (n2) added by line 411, which is shorter than the delay (n3) added by line 412, which is shorter than the delay (n4) added by line 413, the outputs of gain elements 417 and 419 (of the first and third reverberation tanks) are presented to the inputs of summing element 422 and the outputs of gain elements 418 and 420 (of the second and fourth reverberation tanks) are presented to the inputs of summing element 423. The output of element 422 is presented to one input of IACC and mixing filter 424 and the output of element 423 is presented to the other input of IACC filtering and mixing stage 424.

[0139] An example implementation of gain elements 417-420 and elements 422, 423, and 424 of Figure 9 will be described with reference to a typical implementation of elements 310 and 311 and output mixing matrix 312 of Figure 4. The output mixing matrix 312 of Figure 4 (matrix M out 9 embodiment) is a 2×2 matrix configured to mix the unmixed binaural channels from the initial panning (the outputs of elements 310 and 311, respectively) to generate left and right binaural output channels (left ear “L” and right ear “R” signals presented at the output of matrix 312) with the desired interaural coherence. This initial panning is implemented by elements 310 and 311, each of which combines two reverberation tank outputs to generate one of the unmixed binaural channels, with the reverberation tank output with the shortest delay presented to the input of element 310 and the reverberation tank output with the second shortest delay presented to the input of element 311. Elements 422 and 423 of the FIG. 9 embodiment perform (on the time-domain signals presented to their inputs) the same type of initial panning that elements 310 and 311 (in each frequency band) of the FIG. 4 embodiment perform on the streams of filterbank-domain components (in the associated frequency bands) presented to their inputs.

[0140] The unmixed binaural channels (outputs from elements 310 and 311 in FIG. 4 or elements 422 and 423 in FIG. 9), which are largely uncorrelated because they do not contain any common reverberation tank outputs, may be mixed (by matrix 312 in FIG. 4 or stage 424 in FIG. 9) to implement a panning pattern that achieves the desired interaural coherence for the left and right binaural output channels. However, because the reverberation tank delay is different in each FDN (i.e., the FDN in FIG. 9 or the FDN implemented for each different frequency band in FIG. 4), one unmixed binaural channel (the output of one of elements 310 and 311 or 422 and 423) always leads the other unmixed binaural channel (the output of the other of elements 310 and 311 or 422 and 423).

[0141] Thus, in the embodiment of Figure 4, if the combination of reverberation tank delay and panning pattern were the same across all frequency bands, a sound image bias would result. This bias can be mitigated if the panning pattern is alternated across frequency bands, such that the mixed binaural output channels lead and lag each other in alternating frequency bands. For example, if the desired interaural coherence is Coh, and |Coh| < 1, then the output mixing matrix 312 in odd-numbered frequency bands will map the two inputs presented to it into the following form:

number

number

[0142] Alternatively, the above sound image bias in the binaural output channels can be mitigated by implementing matrix 312 to be identical in the FDN for all frequency bands, provided that the channel order of its inputs is switched for alternating frequency bands (e.g., for odd frequency bands, the output of element 310 may be presented to a first input of matrix 312 and the output of element 311 may be presented to a second input of matrix 312, and for even frequency bands, the output of element 311 may be presented to a first input of matrix 312 and the output of element 310 may be presented to a second input of matrix 312).

[0143] In the embodiment of FIG. 9 (and other time-domain embodiments of the FDN of the system of the present invention), it is not trivial to alternate panning based on frequency to address the sound image bias that would otherwise result when the unmixed binaural channel output from element 422 always leads (lags) the unmixed binaural channel output from element 423. This sound image bias is addressed in typical time-domain embodiments of the FDN of the system of the present invention differently than it is typically addressed in filterbank-domain embodiments of the FDN of the system of the present invention. In particular, in the embodiment of FIG. 9 (and other time-domain embodiments of the FDN of the system of the present invention), the relative gains of the unmixed binaural channels (e.g., outputs from elements 422 and 423 of FIG. 9) are determined by gain elements (e.g., elements 417, 418, 419, and 420 of FIG. 9) to compensate for the sound image bias that would otherwise result due to the unbalanced timing described above. The stereo image is re-centered by implementing a gain element (e.g., element 417) to attenuate the earliest arriving signal (which is panned to one side, e.g., by element 422) and a gain element (e.g., element 418) to boost the next earliest arriving signal (which is panned to the other side, e.g., by element 423). Thus, the reverberation tank including gain element 417 applies a first gain to the output of element 417, and the reverberation tank including gain element 418 applies a second gain (different from the first gain) to the output of element 418. The first gain and the second gain thereby attenuate the first unmixed binaural channel (output from element 422) relative to the second unmixed binaural channel (output from element 423).

[0144] More specifically, in a typical implementation of the FDN of FIG. 9, the four delay lines 410, 411, 412, and 413 have successively increasing lengths and successively increasing delay values n1, n2, n3, and n4, respectively. In this implementation, the filter 417 applies a gain of g1. Thus, the output of the filter 417 is a delayed version of the input to the delay line 410 to which the gain of g1 is applied. Similarly, the filter 418 applies a gain of g2, the filter 419 applies a gain of g3, and the filter 420 applies a gain of g4. Thus, the output of the filter 418 is a delayed version of the input to the delay line 411 to which the gain of g2 is applied, the output of the filter 419 is a delayed version of the input to the delay line 412 to which the gain of g3 is applied, and the output of the filter 420 is a delayed version of the input to the delay line 413 to which the gain of g4 is applied.

[0145] In this implementation, the following selection of gain values: g1 = 0.5, g2 = 0.5, g3 = 0.5, g4 = 0.5 may lead to an undesirable bias to one side of the output sound image (i.e., to the left or right channel) as indicated by the binaural channel output from element 424. According to an embodiment of the present invention, the values g1, g2, g3, g4 (applied by elements 417, 418, 419, and 420, respectively) are selected as follows to center the sound image: g1 = 0.38, g2 = 0.6, g3 = 0.5, g4 = 0.5. Thus, according to an embodiment of the present invention, the output stereo image attenuates the earliest arriving signal (which is panned to one side by element 422 in the current example) with respect to the second latest arriving signal (i.e., select g1 < g3), and boosts the second earliest signal (which is panned to the other side by element 423 in the current example) with respect to the latest arriving signal (i.e., select g4 < g2), thereby re-centering it.

[0146] The typical implementation of the time-domain FDN of FIG. 9 has the following differences and similarities with respect to the filter bank region (CQMF region) FDN of FIG. 4.

[0147] The same unitary feedback matrix A (matrix 308 in Figure 4 and matrix 415 in Figure 9).

[0148] Similar reverberation tank delay n i (i.e., the delay in the CQMF implementation in Figure 4 is 1 / T s where is the sampling rate (1 / T s is typically equal to 48KHz), n1=17*64T s =1088*T s , n2=21*64T s =1344*T s , n3=26*64T s =1664*T s , n4=29*64T s =1856*T s while the delay in the time domain implementation may be n1=1089*T s , n2=1345*T s , n3=1663*T s , n4= 185*T s (Note that in a typical CQMF implementation there is a practical constraint that each delay be some integer multiple of the duration of a block of 64 samples, but in the time domain there is more flexibility in the choice of each delay, and therefore more flexibility in the choice of delay for each reverberation tank.)

[0149] Similar all-pass filter implementations (i.e., similar implementations of filter 301 in FIG. 4 and filter 401 in FIG. 9). For example, an all-pass filter can be implemented by a cascade of several (e.g., three) all-pass filters. For example, each cascade of all-pass filters may have g=0.6, such that

number

[0150] In some implementations of the time-domain FDN of FIG. 9 , input filter 400 is implemented such that the direct-to-late ratio (DLR) of the BRIR applied by the system of FIG. 9 matches (at least substantially) a target DLR, and such that the DLR of the BRIR applied by a virtualizer including the system of FIG. 9 (e.g., the virtualizer of FIG. 10 ) can be changed by replacing filter 400 (or controlling the configuration settings of filter 400). For example, in some embodiments, filter 400 is implemented as a cascade of filters (e.g., first filter 400A and second filter 400B coupled as shown in FIG. 9A ) that implement the target DLR and, optionally, desired DLR control. For example, the cascade filters are IIR filters (e.g., filter 400A is a first-order Butterworth high-pass filter (IIR filter) configured to match a target low-frequency characteristic, and filter 400B is a second-order low-shelf IIR filter configured to match a target high-frequency characteristic). As another example, the cascade of filters are IIR and FIR filters (e.g., filter 400A is a second-order Butterworth high-pass filter (IIR filter) configured to match the target low-frequency characteristics, and filter 400B is a 14th-order FIR filter configured to match the target high-frequency characteristics). Typically, the direct signal is fixed, and filter 400 modifies the late signal to achieve the target DLR. All-pass filter (APF) 401 is preferably implemented to perform the same function as APF 301 in FIG. 4, i.e., to introduce phase diversity and increased echo density to produce a more natural-sounding FDN output. While input filter 400 controls the amplitude response, APF 401 typically controls the phase response.

[0151] In Figure 9, filter 406 and gain element 406A together implement a reverberation filter, filter 407 and gain element 407A together implement another reverberation filter, filter 408 and gain element 408A together implement another reverberation filter, and filter 409 and gain element 409A together implement another reverberation filter. Each of filters 406, 407, 408 and 409 in Figure 9 is preferably implemented as a filter with a maximum gain value close to 1 (unity gain), and each of gain elements 406A, 407A, 408A and 409A has a maximum gain value (with associated reverberation tank delay n i 409. Specifically, gain element 406A is configured to apply an attenuation gain to the output of a corresponding one of filters 406, 407, 408, and 409 that matches the desired attenuation (after the reverberation tank delay n i gain element 407A is configured to apply a decay gain (decaygain1) to the output of filter 406 such that the output of delay line 410 (after reverberation tank delay n2) has a first target decayed gain; gain element 407A is configured to apply a decay gain (decaygain2) to the output of filter 407 such that the output of delay line 411 (after reverberation tank delay n2) has a second target decayed gain; and gain element 408A is configured to apply a decay gain (decaygain3) to the output of filter 407 such that the output of delay line 411 (after reverberation tank delay n2) has a second target decayed gain. gain element 409A is configured to apply a decay gain (decaygain3) to the output of filter 408 such that the output of delay line 412 (after reverberation tank delay n3) has a third target decayed gain, and gain element 409A is configured to apply a decay gain (decaygain4) to the output of filter 409 such that the output of delay line 413 (after reverberation tank delay n4) has a fourth target decayed gain.

[0152] Each of filters 406, 407, 408, and 409 and each of elements 406A, 407A, 408A, and 409A of the system of Figure 9 are preferably implemented to achieve a target T60 characteristic of the BRIR applied by a virtualizer (e.g., the virtualizer of Figure 10) comprising the system of Figure 9 (each of filters 406, 407, 408, and 409 is preferably implemented as an IIR filter, e.g., a shelf filter or a cascade of shelf filters), where T60 is the reverberation decay time (T 60 ) represents a frequency characteristic of a signal. For example, in some embodiments, filters 406, 407, 408, and 409 are each implemented as a shelf filter (e.g., a shelf filter with Q=0.3 and a shelf frequency of 500 Hz to achieve the T60 characteristic shown in FIG. 13; T60 in FIG. 13 has units of seconds) or as a cascade of two IIR shelf filters (e.g., one with shelf frequencies of 100 Hz and 1000 Hz to achieve the T60 characteristic shown in FIG. 14; T60 in FIG. 14 has units of seconds). The shape of each shelf filter is determined to match a desired transition curve from low to high frequencies. When filter 406 is implemented as a shelf filter (or a cascade of multiple shelf filters), the reverberation filter comprising filter 406 and gain element 406A is also a shelf filter (or a cascade of shelf filters). Similarly, when each of filters 407, 408 and 409 is implemented as a shelf filter (or a cascade of shelf filters), filter 407 (or 408 or 409) and each reverberation filter with its corresponding gain element (407A, 408A or 409A) is also a shelf filter (or a cascade of shelf filters).

[0153] Figure 9B is an example of filter 406 implemented as a cascade of first shelf filter 406B and second shelf filter 406C coupled as shown in Figure 9B. Each of filters 407, 408, and 409 may be implemented similarly to the Figure 9B implementation of filter 406.

[0154] In some embodiments, the decay gain applied by elements 406A, 407A, 408A, and 409A i ) is determined as follows:

[0155]

number

[0156] 11 is an embodiment of the following elements of FIG. 9: elements 422 and 423 and IACC (interaural cross-correlation coefficient) filtering and mixing stage 424. Element 422 is coupled and configured to sum the outputs of filters 417 and 419 (of FIG. 9) and present the summed signal to the input of low-shelf filter 500, and element 422 is coupled and configured to sum the outputs of filters 418 and 420 (of FIG. 9) and present the summed signal to the input of high-pass filter 501. The outputs of filters 500 and 501 are summed (mixed) in element 502 to generate a binaural left-ear output signal, and the outputs of filters 500 and 501 are mixed in element 502 (the output of filter 500 is subtracted from the output of filter 501 in element 502) to generate a binaural right-ear output signal. Elements 502 and 503 mix (add and subtract) the filtered outputs of filters 500 and 501 to produce a binaural output signal that achieves the target IACC characteristic (within acceptable accuracy). In the embodiment of FIG. 11, low-shelf filter 500 and high-pass filter 501 are each typically implemented as first-order IIR filters. In one example where filters 500 and 501 have such an implementation, the embodiment of FIG. 11 may achieve the exemplary IACC characteristic plotted as curve "I" in FIG. 12. This is shown as "I" in FIG. 12. T It is a good match to the target IACC characteristic plotted as ".

[0157] Figure 11A is a graph of the frequency response (R1) of a typical implementation of filter 500 of Figure 11, the frequency response (R2) of a typical implementation of filter 501 of Figure 11, and the response of filters 500 and 501 connected in parallel. From Figure 11A, it is clear that the combined response is desirably flat across the range of 100 Hz to 10,000 Hz.

[0158] Thus, in one class of embodiments, the present invention is a system (e.g., the system of FIG. 10) and method for generating a binaural signal (e.g., the output of element 210 of FIG. 10) in response to a set of channels of a multi-channel audio input signal by applying a binaural room impulse response (BRIR) to each channel of the set to thereby generate a filtered signal, including by using a single feedback delay network (FDN) to add common late reverberation to a downmix of the channels of the set; and combining the filtered signals to generate the binaural signal. The FDN is implemented in the time domain. In some such embodiments, a time-domain FDN (e.g., FDN 220 of FIG. 10, configured as in FIG. 9) may: an input filter (e.g., filter 400 of FIG. 9) having an input coupled to receive the downmix, the input filter configured to generate a first filtered downmix in response to the downmix; an all-pass filter (e.g., all-pass filter 401 of FIG. 9) coupled and configured to provide a second filtered down-mix in response to the first filtered down-mix; a reverberation application subsystem (e.g., all elements other than elements 400, 401, and 424 of FIG. 9) having a first output (e.g., the output of element 422) and a second output (e.g., the output of element 423), the reverberation application subsystem including a collection of reverberation tanks, each reverberation tank having a different delay, the reverberation application subsystem coupled and configured to generate a first unmixed binaural channel and a second unmixed binaural channel in response to the second filtered downmix, and to present the first unmixed binaural channel at the first output and the second unmixed binaural channel at the second output; an interaural cross-correlation coefficient (IACC) filtering and mixing stage (e.g., stage 424 of FIG. 9, which may be implemented as elements 500, 501, 502, 503 of FIG. 11) coupled to the reverberation application subsystem and configured to generate first and second mixed binaural channels in response to the first and second unmixed binaural channels.

[0159] The input filter may be implemented to generate (preferably as a cascade of two filters configured to generate) the first filtered downmix so that each BRIR has a direct-to-late ratio (DLR) that at least substantially matches a target DLR.

[0160] Each reverberation tank may be configured to generate a delayed signal and may include a reverberation filter (e.g., implemented as a shelf filter or a cascade of shelf filters) coupled and configured to apply gain to a signal propagating in each reverberation tank such that the delayed signal has a gain that at least substantially matches a target delayed gain for the delayed signal. 60 This is to achieve the following characteristics:

[0161] In some embodiments, the first unmixed binaural channel leads the second unmixed binaural channel, and the reverberation tank includes a first reverberation tank configured to generate a first delayed signal having the shortest delay (e.g., the reverberation tank of FIG. 9 including delay line 410) and a second reverberation tank configured to generate a second delayed signal having a second shortest delay (e.g., the reverberation tank of FIG. 9 including delay line 411). The first reverberation tank is configured to apply a first gain to the first delayed signal, and the second reverberation tank is configured to apply a second gain to the second delayed signal, the second gain being different from the first gain, the second gain being different from the first gain, and application of the first gain and the second gain resulting in an attenuation of the first unmixed binaural channel relative to the second unmixed binaural channel. Typically, the first mixed binaural channel and the second mixed binaural channel exhibit a recentered stereo image. In some embodiments, the IACC filtering and mixing stage is configured to generate the first mixed binaural channel and the second mixed binaural channel such that the first mixed binaural channel and the second mixed binaural channel have IACC characteristics that at least substantially match a target IACC characteristic.

[0162] Aspects of the present invention include methods and systems (e.g., system 20 of FIG. 2 or the systems of FIG. 3 or 10) that perform (or are configured to perform or support) binaural virtualization of audio signals (e.g., audio signals whose audio content consists of speaker channels and / or object-based audio signals).

[0163] In some embodiments, a virtualizer of the present invention is or includes a general-purpose processor coupled to receive or generate input data representing a multi-channel audio input signal and programmed with software (or firmware) or otherwise configured (e.g., in response to control data) to perform any of a variety of operations on the input data, including method embodiments of the present invention. Such a general-purpose processor is typically coupled to an input device (e.g., a mouse and / or keyboard), memory, and a display device. For example, the system of FIG. 3 (or system 20 of FIG. 2 or a virtualizer system having elements 12, ..., 14, and 15 of system 20) can be implemented in a general-purpose processor, with the input being audio data representing N channels of the audio input signal and the output being audio data representing two channels of a binaural audio signal. A conventional digital-to-analog converter (DAC) can operate on the output data to generate analog versions of the binaural signal channels for playback through speakers (e.g., a pair of headphones).

[0164] While particular embodiments of the present invention and applications of the present invention have been described herein, it will be apparent to those skilled in the art that many variations to these embodiments and applications described herein are possible without departing from the scope of the invention as described and claimed herein. While certain forms of the present invention have been illustrated and described, it is to be understood that the invention is not limited to the specific embodiments described and illustrated, or to the specific methods described.

Claims

1. 1. A method for generating a binaural signal in response to a set of channels of a multi-channel audio input signal, the method comprising: applying a binaural room impulse response (BRIR) to each channel of the set, thereby generating a filtered signal; combining the filtered signals to generate the binaural signal; applying BRIR to each channel of the set includes introducing, using a late reverberation generator, a common late reverberation into a downmix of the channels of the set in response to a control value presented to the late reverberation generator, the common late reverberation emulating collective macro-attributes of late reverberant portions of single-channel BRIRs shared across at least some channels of the set; a content-dependent energy equalization factor is applied to the downmix, and a center channel of the multi-channel audio input signal is panned to both the left channel of the downmix and the right channel of the downmix. method.

2. The method of claim 1 , wherein applying a BRIR to each channel of the set comprises applying to each channel of the set a direct response and early reflection portion of a single-channel BRIR for that channel.

3. 2. The method of claim 1, wherein the late reverberation generator comprises a bank of feedback delay networks for adding the common late reverberation to the downmix, each feedback delay network of the bank adding late reverberation to a different frequency band of the downmix.

4. The method of claim 3 , wherein each of said feedback delay networks is implemented in a complex quadrature mirror filter domain.

5. The method of claim 1 , wherein the late reverberation generator includes a single feedback delay network for adding the common late reverberation to the downmix of the channels of the set, the feedback delay network being implemented in the time domain.

6. 1. A system for generating a binaural signal in response to a set of channels of a multi-channel audio input signal, the system comprising: applying a binaural room impulse response (BRIR) to each channel of the set, thereby generating a filtered signal; combining the filtered signals to generate the binaural signal; having one or more processors, applying BRIR to each channel of the set includes introducing, using a late reverberation generator, a common late reverberation into a downmix of the channels of the set in response to a control value presented to the late reverberation generator, the common late reverberation emulating collective macro-attributes of late reverberant portions of single-channel BRIRs shared across at least some channels of the set; a content-dependent energy equalization factor is applied to the downmix, and a center channel of the multi-channel audio input signal is panned to both the left channel of the downmix and the right channel of the downmix. system.

7. The system of claim 6 , wherein applying a BRIR to each channel of the set comprises applying to each channel of the set a direct response and early reflection portion of a single-channel BRIR for that channel.

8. 7. The system of claim 6, wherein the late reverberation generator comprises a bank of feedback delay networks configured to add the common late reverberation to the downmix, each feedback delay network of the bank adding late reverberation to a different frequency band of the downmix.

9. The system of claim 8 , wherein each of said feedback delay networks is implemented in a complex quadrature mirror filter domain.

10. 7. The system of claim 6, wherein the late reverberator comprises a feedback delay network implemented in the time domain, and the late reverberator is configured to process the downmix in the time domain in the feedback delay network to add the common late reverberation to the downmix.

11. 10. A non-transitory computer-readable storage medium having a sequence of instructions that, when executed by an audio signal processing apparatus, causes the audio signal processing apparatus to perform the method of claim 1.

12. A computer program product comprising instructions for carrying out the method of claim 1 when executed on a computer.

Citation Information

Patent Citations

  • Sound compensation device

    JP2007336080A

  • Compact Side Information for Parametric Coding of Spatial Audio

    JP2008527431A

  • Signal generation for binaural signals

    JP2011529650A

  • Reverberation device and method for reverberating audio signals

    JP2013508760A