Generation of binaural audio in response to multichannel audio using at least one feedback delay network
Patent Information
- Application Number
- ES2023195452T
- Authority / Receiving Office
- ES · ES
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2014-12-18
- Filing Date
- 2014-12-18
- Publication Date
- 2026-08-31
- Estimated Expiration
- 2034-12-18
Smart Images

Figure 00000028_0000 
Figure 00000028_0001 
Figure 00000029_0000
Abstract
Description
Generation of binaural audio in response to multichannel audio using at least one feedback delay network Cross-reference to related application This application is a European Divisional Application of European Divisional Application EP 20205638.8 (reference: D13171EP03), for which EPO Form 1001 was filed on November 4, 2020. This application claims priority from Chinese Patent Application No. 201410178258,0 filed on April 29, 2014, U.S. Provisional Patent Application No. 61 / 923,579 filed on January 3, 2014, and U.S. Provisional Patent Application No. 61 / 988,617 filed on May 5, 2014. Background of the invention 1. Field of the invention The invention relates to methods (sometimes referred to as headphone virtualization methods) and systems for generating a binaural signal in response to a multi-channel audio input signal by applying a binaural impulse response (BRIR) to each channel of a set of channels (e.g., to all channels) of the input signal. In some embodiments, at least one feedback delay network (FDN) applies a late reverb portion of a downmix BRIR to a downmix of the channels. 2. Background of the invention Headphone virtualization (or binaural playback) is a technology that aims to provide a surround sound experience or immersive sound field using standard stereo headphones. Early headphone virtualizers employed a head-related transfer function (HRTF) to convey spatial information in binaural playback. An HRTF is a set of direction- and distance-dependent filter pairs that characterize how sound is transmitted from a specific point in space (the sound source location) to both ears of a listener in an anechoic environment. Essential spatial cues such as interaural time difference (ITD), interaural level difference (ILD), head shadowing, and spectral peaks and notches due to shoulder and ear reflections can be perceived in the HRTF-filtered binaural content. Due to the limitations of human head size, HRTFs do not provide sufficient or robust cues with respect to distances from the source beyond approximately one meter.As a result, virtualizers based solely on an HRTF typically fail to achieve good outsourcing or perceived distance. Most acoustic events in our daily lives occur in reverberant environments where, in addition to the direct path (from source to ear) modeled by the HRTF, audio signals also reach the listener's ears through various reflection paths. Reflections profoundly impact an audience's perception of an auditorium, such as distance, room size, and other spatial attributes. To convey this information in binaural playback, a virtualizer must be applied to the room reverberation in addition to the signals in the direct path HTRF. A room's binaural impulse response (BRIR) characterizes the transformation of audio signals from a specific point in space to the listener's ears in a specific acoustic environment. In theory, BRIRs encompass all acoustic signals related to spatial perception. Figure 1 is a block diagram of a conventional headphone virtualizer configured to apply a one-room binaural impulse response (BRIR) to each full-range frequency channel (X1, ..., XN) of a multi-channel audio input signal. Each of the channels X1, ..., XN is a speaker channel corresponding to a different source direction relative to a hypothetical listener (i.e., the direction of the direct path from a hypothetical speaker position to the hypothetical listener position), and each such channel is convolved by the BRIR for the corresponding source direction. The acoustic path from each channel needs to be simulated for each ear. Therefore, throughout the rest of this document, the term BRIR will refer to one impulse response or a pair of impulse responses associated with the left and right ears.Therefore, subsystem 2 is configured to convolve channel X1 with BRIR1 (the BRIR for the corresponding source direction), subsystem 4 is configured to convolve channel XN with BRIRRN (the BRIR for the corresponding source direction), and so on. The output of each BRIR subsystem (each of subsystems 2, ..., 4) is a time-domain signal that includes a left and a right channel. The left channel outputs of the BRIR subsystems are mixed in addition element 6, and the right channel outputs of the BRIR subsystems are mixed in addition element 8. The output of element 6 is the left channel, L, of the virtualizer's binaural audio signal output, and the output of element 8 is the right channel, R, of the virtualizer's binaural audio signal output. The input signal of the multichannel audio may also include low-frequency effects (LFE) or a bass channel, identified in Figure 1 as the "LFE" channel. Conventionally, the LFE channel is not convolved with a BRIR, but instead is attenuated in gain stage 5 of Figure 1 (for example, by -3dB or more), and the output of stage 5 is mixed equally (via elements 6 and 8) into each of the channels of the virtualizer's binaural output signal. An additional delay stage in the LFE path may be required to time-align the output of stage 5 with the outputs of the BRIR subsystems (2, ..., 4). Alternatively, the LFE channel can simply be ignored (i.e., not imposed or processed by the virtualizer). For example, the embodiment of Figure 2 of the invention (described below) simply ignores any LFE channel of the multichannel audio input signal processed in this way.Many consumer headphones are not capable of accurately reproducing an LFE channel. In some conventional virtualizers, the input signal undergoes time-domain to frequency-domain transformation within the QMF domain (quadrature mirror filter) to generate the frequency component channels in the QMF domain. These frequency components are then filtered (e.g., in the QMF domain implementations of subsystems 2, ..., 4 in Figure 1) in the QMF domain, and the resulting frequency components are typically transformed back to the time domain (e.g., in a final stage of each of subsystems 2, ..., 4 in Figure 1) so that the audio output of the virtualizer is a time-domain signal (e.g., a binaural time-domain signal). In general, each full-range frequency channel of a multichannel audio signal input to a headphone virtualizer is assumed to be indicative of the audio content emitted from a sound source at a known location relative to the listener's ears. The headphone virtualizer is configured to apply a binaural impulse response (BRIR) to each of these input signal channels. Each BRIR can be broken down into two parts: the direct response and the reflections. The direct response is the HRTF corresponding to the direction of arrival (DOA) of the sound source, adjusted with appropriate gain and delay due to distance (between the sound source and the listener), and optionally augmented with parallax effects for short distances. The remaining part of the BRIR model shapes reflections. Early reflections are typically primary or secondary reflections and have a relatively dispersed temporal distribution. The microstructure (e.g., ITD and ILD) of each primary or secondary reflection is important. For later reflections (sound reflected from more than two surfaces before reaching the listener), echo density increases with the number of reflections, and the micro-attributes of individual reflections become difficult to observe. For increasingly late reflections, macrostructure (e.g., reverberation decay rate, interaural coherence, and overall reverberation spectral distribution) becomes more important. Because of this, reflections can be further segmented into two parts: early reflections and late reverberations. The delay of the direct response is the distance from the listener to the source divided by the speed of sound, and its level (in the absence of walls or large surfaces near the source) is inversely proportional to the distance to the source. On the other hand, the delay and level of late reverberations are generally insensitive to the source's location. For practical reasons, virtualizers may choose to time-align the direct responses of sources with different distances and / or compress their dynamic range. However, the temporal and level relationship between the direct response, early reflections, and late reverberation within a BRIR (Broadband Reverberation Interface) should be maintained. The effective length of a typical BRIR extends to hundreds of milliseconds or more in more acoustic environments. Direct application of BRIRs requires convolution with a filter of thousands of pulses, which is computationally expensive. Furthermore, without parameterization, it would require a large amount of memory to store the BRIRs for different source positions to achieve sufficient spatial resolution. Last but not least, the locations of the sound source can change over time, and / or the position and orientation of the listener can vary over time. Approximate simulation of such movements requires time-varying BRIR impulse responses. Interpolating and appropriately applying such time-varying filters can be challenging if the impulse responses of these filters have many pulses. A filter with the well-known filter structure known as a feedback delay network (FDN) can be used to implement a spatial reverberator configured to apply simulated reverb to one or more channels of a multi-channel audio input signal. The structure of an FDN is simple. It comprises several reverb tanks (for example, the reverb tank comprising the gain element g1 and the delay line z-n1, in the FDN of Figure 4), each reverb tank having a delay and a gain. In a typical FDN implementation, the outputs from all the reverb tanks are mixed by a unity feedback matrix, and the matrix outputs are fed back and summed with the inputs to the reverb tanks.Gain adjustments can be made to the outputs of the reverb tanks, and the outputs (or gain-adjusted versions thereof) can be remixed appropriately for multichannel or binaural playback. Natural-sounding reverb can be generated and applied by a FDN with a compact computational and memory footprint. FDNs have therefore been used in virtualizers to augment the direct response produced by the HRTF. For example, the commercially available Dolby Mobile Headphone Virtualizer includes a reverberator with an FDN-based structure that can apply reverb to each channel of a five-channel audio signal (having front left, front right, center, left surround, and right surround channels) and filter each reverberated channel using a different pair of filters from a set of five head-related transfer function ("HRTF") filter pairs. The Dolby Mobile Headphone Virtualizer can also operate in response to a two-channel audio signal to generate a two-channel "reverberated" binaural audio output (a two-channel virtual surround sound output to which reverb has been applied).When the reverberated binaural output is processed and played back through a pair of headphones, it is perceived in the listener's eardrum as a sound filtered by the HRTF, reverberating from the five speakers in the front left, front right, center, rear left (surround), and rear right (surround) positions. The virtualizer upmixes a two-channel audio input (without using any spatial signal parameters received with the audio input) to generate five upmixed audio channels, applies reverb to the upmixed channels, and downmixes the five reverberated channel signals to generate the virtualizer's two-channel reverberated output. The reverb for each upmixed channel is filtered through a different pair of HRTF filters. In a virtualizer, a sound diffusion density (FDN) can be configured to achieve a specific reverberation decay time and echo density. However, the FDN lacks the flexibility to simulate the microstructure of early reflections. Furthermore, in conventional virtualizers, the tuning and configuration of FDNs has been primarily heuristic. Headphone virtualizers that do not simulate all reflection paths (early and late) cannot achieve effective externalization. The inventors have acknowledged that virtualizers employing FDNs that attempt to simulate all reflection paths (early and late) generally have only limited success in simulating both early reflections and late reverberation and in applying both to an audio signal. The inventors have also acknowledged that virtualizers employing FDNs but lacking the ability to appropriately control spatial acoustic attributes such as reverberation decay time, interaural coherence, and direct-to-late ratio can achieve a degree of externalization at the cost of introducing excessive timbre distortion and reverberation. Patent document WO 2012 / 093352 A1 discloses an audio system and a method for operating a virtual audio signal space. It further discloses the application of a binaural impulse response (BRIR) to each channel of an array and the combination of filtered signals to generate the binaural signal. By applying the BRIRs, a common late reverb is applied to a downmix of the channels, thus saving computational resources. Brief description of the invention In a first class of embodiments, the invention is a method as defined in claim 1. A method for generating a binaural signal in response to a multichannel audio input signal (or in response to a set of channels of said signal) is sometimes referred to herein as a "headphone virtualization" method, and the system configured to perform said method is sometimes referred to herein as a "headphone virtualizer" (or "headphone virtualization system" or "binaural virtualizer"). In typical first-class implementations, each of the FDNs is implemented in the filter bank domain (e.g., in the hybrid complex quadrature mirror filter (HCQMF) domain or the quadrature mirror filter (QMF) domain, or another transform or sub-band domain that may include decimation). In some of these implementations, the frequency-dependent spatial acoustic attributes of the binaural signal are controlled by adjusting the configuration of each FDN used to apply late reverb. Typically, a monophonic downmix of the channels is used as the input to the FDNs for efficient binaural processing of the multichannel signal's audio content.Typical implementations in the first class include a step of adjusting FDN coefficients corresponding to frequency-dependent attributes (e.g., reverb decay time, interaural coherence, modal density, and direct-to-late ratio), for example, by imposing control values on the feedback delay network to a set of at least one of the input gain, reverb tank gains, reverb tank delays, or output matrix parameters for each FDN. This allows for better matching of acoustic environments and more natural sound outputs. In a second class of embodiments not covered by the appended claims, the invention is a method for generating a binaural signal in response to a multi-channel audio input signal having channels, by applying a one-room binaural impulse response (BRIR) to each channel of a set of channels of the input signal (e.g., each of the input signal channels or each full-frequency range channel of the input signal), including: processing each channel of the set in a first processing path configured to model, and apply to said channel, a forward response and an early reflection portion of a single-channel BRIR for the channel; and processing a downmix (e.g., a monophonic (mono) downmix) of the channels of the set in a second processing path (in parallel with the first processing path) configured to model,and apply a common late reverb to the downmix. Typically, the common late reverb is generated to emulate the collective macro attributes of the late reverb portions of at least some (e.g., all) of the single-channel BRIRs. Typically, the second processing path includes at least one FDN (e.g., one FDN for each of the multiple frequency bands). Typically, a mono downmix is used as the input to all the reverb tanks of each FDN implemented by the second processing path. Typically, mechanisms are provided for the systematic control of the macro attributes of each FDN to better simulate acoustic environments and produce a more natural-sounding binaural virtualization. Since most of these macro attributes are frequency-dependent, each FDN is typically implemented in the hybrid complex quadrature mirror filter (HCQMF) domain.the frequency domain, the domain, or another filter bank domain, and a different or independent FDN is used for each frequency band. The main benefit of implementing FDNs in the filter bank domain is to allow the application of reverb with frequency-dependent reverb properties. In various embodiments, FDNs are implemented in any of a wide variety of filter bank domains, using any of a variety of filter banks, including, but not limited to, complex-valued quadrature mirror (QMF) filters, finite impulse response (FIR) filters, infinite impulse response (IIR) filters, discrete Fourier transforms (DFTs), and modified cosine or sine transforms, wavelet transforms, or crossover filters. In a preferred embodiment, the transform or filter bank employed includes decimation (e.g.,a decrease in the sampling rate of the signal representation in the frequency domain) to reduce the computational complexity of the FDN process. Some implementations in the first class (and in the second class) implement one or more of the following features: 1. An FDN implementation in the filter bank domain (e.g., in the hybrid complex quadrature mirror filter domain), or an FDN implementation in the hybrid filter bank domain and a late reverb filter implementation in the time domain, which typically allows independent tuning of the FDN parameters and / or settings for each frequency band (allowing simple and flexible control of frequency-dependent acoustic attributes), e.g., by providing the ability to vary the reverb tank delays in different bands to change the modal density as a function of frequency: 2.The specific downmixing process, employed to generate (from the multi-channel input audio signal) the downmixed signal (e.g., mono downmix) processed in the second processing path, depends on the distance to the source of each channel and the handling of the forward response to maintain an appropriate level and timing relationship between the forward and late responses; 3. An all-pass filter (APF) is applied in the second processing path (e.g., at the input or output of an FDN bank) to introduce phase diversity and increase echo density without changing the spectrum and / or timbre of the resulting reverberation; 4. Fractional delays are implemented in the feedback path for each FDN with a multi-rate, complex-value structure to overcome the problems related to quantized delays for the sample decay factor grid; 5. In FDNs, the reverb tank outputs are linearly downmixed directly into the binaural channels, using output downmix coefficients that are adjusted based on the desired interaural coherence in each frequency band. Optionally, the mapping of the reverb tanks to the binaural output channels alternates between frequency bands to achieve a balanced delay between the binaural channels. Also optionally, normalization factors are applied to the reverb tank outputs to equalize their levels while preserving fractional delay and overall energy. 6. The frequency- and / or modal density-dependent reverb decay time is controlled by appropriate combinations of reverb tank delay settings and gains in each frequency band to simulate real rooms. 7. A frequency band scaling factor is applied (for example, at either the input or output of the relevant processing path), to: control the frequency-dependent direct-to-late relationship (DLR) that matches that of an actual room (a simple model can be used to calculate the required scaling factor based on the target DLR and decay time, e.g., T60); provide low-frequency attenuation to mitigate excessive comb artifacts and / or low-frequency noise; and / or apply diffuse field spectral shaping to the FDN responses; 8. Simple parametric models are implemented to control essential frequency-dependent attributes of late reverberation, such as reverberation decay time, interaural coherence, and / or direct relationship to late. Aspects of the invention include procedures and systems that perform (or are configured to perform, or support the performance of) binaural virtualization of audio signals (e.g., audio signals whose audio content consists of speaker channels, and / or object-based audio signals). In another class of embodiments not covered by the appended claims, the invention is a system for generating a binaural signal in response to an array of channels of a multi-channel audio input signal, including applying a binaural room impulse response (BRIR) to each channel of the array, thereby generating filtered signals; including using a unique feedback delay network (FDN) to apply a common late reverb to a downmix of the array's channels; and combining the filtered signals to generate the binaural signal. The FDN is implemented in the time domain. In some of these embodiments, the time-domain FDN includes: an inlet filter having a coupled inlet for receiving the downmix, wherein the inlet filter is configured to generate a first filtered downmix in response to the downmix; an all-pass filter, coupled and configured to a second filtered downmix in response to the first filtered downmix; a reverb application subsystem, having a first output and a second output, wherein the reverb application subsystem comprises a set of reverb tanks, each of the reverb tanks having a different delay, and wherein the reverb application subsystem is coupled and configured to generate a first unmixed binaural channel and a second unmixed binaural channel in response to the second filtered downmix, to impose the first unmixed binaural channel on the first output, and to impose the second unmixed binaural channel on the second output; and an interaural cross-correlation coefficient (IACC) filtering and mixing stage coupled to the reverb application subsystem and configured to generate a first mixed binaural channel and a second mixed binaural channel in response to the first unmixed binaural channel and the second unmixed binaural channel. The input filter can be implemented to generate (preferably as a cascade of two filters configured to generate) the first filtered downmix such that each BRIR has a direct-to-late ratio (DLR) that matches, at least substantially, a target DLR. Each reverb tank can be configured to generate a delayed signal, and can include a reverb filter (e.g., implemented as an attenuator filter or a cascade of attenuator filters) coupled and configured to apply gain to a signal propagating into each of the reverb tanks, to cause the delayed signal to have a gain that matches, at least substantially, a target decay gain for that delayed signal, in an effort to achieve a target reverb decay time characteristic (e.g., a T60 characteristic) of each BRIR. In some embodiments, the first unmixed binaural channel guides the second unmixed binaural channel. The reverb tanks include a first reverb tank configured to generate a first delayed signal having the shortest delay and a second reverb tank configured to generate a second delayed signal having the second shortest delay. The first reverb tank is configured to apply a first gain to the first delayed signal, and the second reverb tank is configured to apply a second gain to the second delayed signal. The second gain is different from the first gain, and the application of the first gain and the second gain results in an attenuation of the first unmixed binaural channel relative to the second unmixed binaural channel.Typically, the first and second unmixed binaural channels indicate a refocused stereo image. In some implementations, the IACC filtering and mixing stage is configured to generate the first and second mixed binaural channels in such a way that their IACC characteristics match, at least substantially, the target IACC characteristics. Typical embodiments of the invention provide a simple, unified structure to support both consistent speaker channel input audio and object-based input audio. In embodiments where BRIR is applied to input signal channels that are object channels, the "direct response and early reflection" processing performed on each object channel results in a source address indicated by metadata provided with the object channel's audio content. In embodiments where BRIR is applied to input signal channels that are speaker channels, the "direct response and early reflection" processing performed on each speaker channel results in a source address corresponding to the speaker channel (i.e., the direction of a direct path from the assumed position of the corresponding speaker to the assumed position of the listener).Regardless of whether the input channels are object or speaker channels, the "late reverb" processing is performed on the downmix (e.g., a mono downmix) of the input channels and does not imply any specific source direction for the audio content of the downmix. Other aspects not covered by the appended claims are a headphone virtualizer configured (e.g., programmed) to perform any embodiment of the inventive process, and a system (e.g., a stereo, multichannel, or other decoder) comprising such a virtualizer. According to claim 9, one aspect is a computer-readable medium (e.g., a disk) that stores code for implementing any embodiment of the inventive process. Brief description of the drawings Figure 1 is a block diagram of a conventional headphone virtualization system. Figure 2 is a block diagram of a system that includes an implementation of the inventive headphone virtualization system. Figure 3 is a block diagram of another embodiment of the inventive headphone virtualization system. Figure 4 is a block diagram of an FDN of one type included in a typical implementation of the system in Figure 3. Figure 5 is a graph of the reverberation decay time (T60) in milliseconds as a function of the frequency in Hz, which can be achieved by an inventive virtualizer realization for which the value of T60 at each of the two specified frequencies (fA and fB) is set as follows: T60, A = 320 ms with fA = 10 Hz, and T60, B = 150 ms with fB = 2.4 kHz. Figure 6 is a graph of Interaural Coherence (Coh) as a function of frequency in Hz, which can be achieved by an inventive virtualizer realization for which the control parameters Cohmax, Cohmin, and fC are set to have the following values: Cohmax = 0.95, Cohmin = 0.05, and fC = 700 Hz. Figure 7 is a graph of the direct-to-late ratio (DLR) with a distance to the source of one meter, in dB, as a function of the frequency in Hz, which can be achieved by an inventive virtualizer implementation for which the control parameters DLR1K, DLRslope, DLRmin, HPFslope, and fT are set to have the following values: DLR1K = 18 dB, DLRslope = 6 dB / 10x frequency, DLRmin = 18 dB, HPFslope = 6 dB / 10x frequency, and fT = 200 Hz. Figure 8 is a block diagram of another embodiment of a late reverb processing subsystem of the inventive headphone virtualization system. Figure 9 is a block diagram of a time-domain implementation of an FDN, of a type included in some embodiments of the inventive system. Figure 9A is a block diagram of an example implementation of filter 400 in Figure 9. Figure 9B is a block diagram of an example implementation of filter 406 in Figure 9. Figure 10 is a block diagram of an implementation of the inventive headphone virtualization system, in which the time-domain late reverb processing subsystem 221 is implemented. Figure 11 is a block diagram of an implementation of elements 422, 423, and 424 of the FDN in Figure 9. Figure 11A is a graph of the frequency response (R1) of a typical implementation of the 500 filter in Figure 11, the frequency response (R2) of a typical implementation of the 501 filter in Figure 11, and the response of the 500 and 501 filters connected in parallel. Figure 12 is a graph of an example of an IACC feature (curve "I") that can be achieved by implementing the FDN of Figure 9, and a target IACC feature (curve "IT"). Figure 13 is a graph of a T60 feature that can be achieved by implementing the FDN of Figure 9, by appropriately implementing each of filters 406, 407, 408, and 409 as an attenuator filter. Figure 14 is a graph of a T60 feature that can be achieved by implementing the FDN of Figure 9, by appropriately implementing each of filters 406, 407, 408, and 409 as a cascade of two IIR attenuating filters. Notation and nomenclature Throughout this description, including in the claims, the expression "performing an operation 'on' a signal or data (e.g., filtering, scaling, transforming, or applying gain to the signal or data) is used in a broad sense to denote performing the operation directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that undergoes preliminary filtering or pre-processing before performing the operation itself). Throughout this description, including in the claims, the term "system" is used in a broad sense to denote a service, system, or subsystem. For example, a subsystem that implements a virtualizer may be referred to as a virtualizing system, and a system that includes such a subsystem (for example, a system that generates X output signals in response to multiple inputs, wherein the subsystem generates M of the inputs and the other X - M inputs are received from an external source) may also be referred to as a virtualizing system (or a virtualizer). Throughout this description, including in the claims, the term "processor" is used in a broad sense to denote a system or device that is programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio, video, or other image data). Examples of processors include a field-programmable gate array (or other configurable integrated circuit or chipset), a digital signal processor programmed and / or otherwise configured to perform serial processing on audio or other sound data, a programmable general-purpose processor or computer, and a programmable microprocessor chip or chipset. Throughout this description, including the claims, the term "analysis filter bank" is used in a broad sense to denote a system (e.g., a subsystem) configured to apply a transformation (e.g., a time-domain to a frequency-domain transformation) to a time-domain signal to generate values (e.g., frequency components) indicative of the signal's time-domain content in each of a set of frequency bands. Throughout this description, including the claims, the term "filter bank domain" is used in a broad sense to denote the domain of the frequency components generated by a transformation or analysis filter bank (e.g., the domain in which such frequency components are processed).Examples of filter bank domains include (but are not limited to) the frequency domain, the quadrature mirror filter (QMF) domain, and the hybrid complex quadrature mirror filter (HCQMF) domain. Examples of transforms that can be applied through filter bank analysis include (but are not limited to) the discrete cosine transform (DCT), the modified discrete cosine transform (MDCT), the discrete Fourier transform (DFT), and the wavelet transform. Examples of filter bank analysis include (but are not limited to) quadrature mirror filters (QMF), finite impulse response filters (FIR filters), infinite impulse response filters (IIR filters), crossover filters, and filters with other suitable multi-rate structures. Throughout this description, including in the claims, the term "metadata" refers to data separate and distinct from the corresponding audio data (the audio content of a bitstream that also includes the metadata). The metadata is associated with the audio data and indicates at least one characteristic of the audio data (for example, what type or types of processing have already been, or should be, performed on the audio data, or the path of an object indicated by the audio data). The association of the metadata with the audio data is synchronous in time. Therefore, current metadata (the most recently received or updated metadata) can indicate that the corresponding audio data both has a specified characteristic and / or comprises the results of a specified type of audio data processing. Throughout this description, including in the claims, the term "couples" or "coupled" is used to indicate either a direct or indirect connection. Therefore, if a first device couples with a second device, that connection may be through a direct connection, or through an indirect connection via other devices and connections. Throughout this description, including in the claims, the following expressions have the following definitions: Loudspeaker or public address system are used synonymously to denote any sound-emitting transducer. This definition includes loudspeakers implemented as multiple transducers (e.g., woofer and tweeter). speaker feed: an audio signal to be applied directly to a speaker, or an audio signal to be applied to an amplifier and a speaker in series; channel (or audio channel): a monophonic audio signal. Such a signal can normally be reproduced in a manner equivalent to applying the signal directly to a loudspeaker at a desired or nominal position. The desired position can be static, as is normally the case with physical loudspeakers, or dynamic; audio program: a set of one or more audio channels (at least one speaker channel and / or at least one object channel) and optionally also the associated metadata (for example, metadata describing a desired spatial audio presentation); Speaker channel (or "speaker feed channel"): An audio channel that is associated with a designated loudspeaker (in a desired or nominal position), or with a designated loudspeaker zone within a defined loudspeaker configuration. A loudspeaker channel is reproduced in such a manner as to be equivalent to applying the audio signal directly to the designated loudspeaker (in the desired or nominal position) or to a loudspeaker in the designated loudspeaker zone. object channel: An audio channel indicative of the sound emitted by an audio source (sometimes referred to as an audio "object"). Typically, an object channel determines a parametric description of the audio source (for example, metadata indicative of the parametric audio source description is included in or provided with the object channel). The source description may determine the sound emitted by the source (as a function of time), the evident position (for example, 3D spatial coordinates) of the source as a function of time, and optionally at least one additional parameter (for example, the size or width of the evident source) that characterizes the source.audio program: object-based: an audio program comprising a set of one or more object channels (and optionally also comprising at least one speaker channel) and optionally also associated metadata (for example, metadata indicating a path of an audio object that emits the sound indicated by an object channel, or otherwise metadata indicating a desired spatial audio presentation of the sound indicated by an object channel, or metadata indicating an identification of at least one audio object that is a sound source indicated by an object channel); and; Playback: the process of converting an audio program into one or more speaker feeds, or the process of converting an audio program into one or more speaker feeds and converting the speaker feed(s) into sound using one or more loudspeakers (in the latter case, the processing is sometimes referred to herein as processing "by" the loudspeaker(s)). An audio channel may be trivially processed ("at" a desired position) by applying the signal directly to a physical loudspeaker at the desired position, or one or more audio channels may be processed using one of a variety of virtualization techniques designed to be substantially equivalent (to the listener) to such trivial processing.In the latter case, each audio channel can be converted into one or more speaker feeds to be applied to the speaker or speakers at known locations, which are generally different from the desired location, so that the sound emitted by the speaker or speakers in response to the feed or feeds will be perceived as emanating from the desired location. Examples of such virtualization techniques include binaural processing through headphones (for example, using Dolby Headphone processing that simulates up to 7.1 surround sound channels for the headphone carrier) and wavefield synthesis. The notation that a multi-channel audio signal is an "xy" or "xyz" channel signal denotes in the present memory that the signal has "x" full-frequency speaker channels (corresponding to the speakers nominally positioned in the horizontal plane of the presumed listener's ears), "y" LFE (or subwoofer) channels, and optionally also "z" full-frequency overhead speaker channels (corresponding to the speakers positioned above the presumed listener's head, e.g., on or near the ceiling of the room). The expression "IACC" in this memory denotes the interaural cross-correlation coefficient in its usual sense, which is a measure of the difference between the arrival times of the audio signal to the ears of a listener, normally indicated by a number in a range from a first value that indicates that the arriving signals are equal in magnitude and exactly out of phase, to an intermediate value that indicates that the arriving signals have no similarity, to a maximum value that indicates that the arriving signals are identical having the same amplitude and phase. Detailed description of preferred realizations Many embodiments of the present invention are technologically feasible. It will be evident to those skilled in the art from the present description how to implement them. The embodiments of the system and the inventive process will be described with reference to Figure 2-14. Figure 2 is a block diagram of a system (20) that includes an implementation of the inventive headphone virtualization system. The headphone virtualization system (sometimes referred to as a virtualizer) is configured to apply a one-room binaural impulse response (BRIR) to N full-range frequency channels (X1, ..., XN) of a multi-channel audio input signal. Each of the channels X1, ..., XN (which may be speaker channels or object channels) corresponds to a specific direction and distance from the source relative to a hypothetical listener, and the system in Figure 2 is configured to convolve each of these channels with a BRIR for the corresponding direction and distance from the source. System 20 may be a decoder coupled to receive an encoded audio program, and including a coupled subsystem (not shown in Figure 2) configured to decode the program by recovering the N full-range frequency channels (X1, ..., XN) thereof and providing them with elements 12, ..., 14, and 15 of the virtualization system (comprising elements 12, ..., 14, 15, 16, and 18, coupled as shown). The decoder may include additional subsystems, some of which perform functions unrelated to the virtualization function performed by the virtualization system, and some of which may perform functions related to the virtualization function.For example, the latter functions may include extracting metadata from the coded program, and providing the metadata to a virtualization control subsystem that uses the metadata to control elements of the virtualizing system. Subsystem 12 (with subsystem 15) is configured to convolve channel X1 with BRIR1 (the BRIR for direction and distance to the corresponding source), subsystem 14 (with subsystem 15) is configured to convolve channel XN with BRIRRN (the BRIR for direction to the corresponding source), and so on for each of the other N-2 BRIR subsystems. The output of each of subsystems 12, ..., 14, and 15 is a time-domain signal that includes a left channel and a right channel. In addition, elements 16 and 18 are coupled to the outputs of elements 12, ..., 14, and 15. Furthermore, element 16 is configured to combine (mix) the left channel outputs of the BRIR subsystems, and furthermore, element 18 is configured to combine (mix) the right channel outputs of the BRIR subsystems.The output of element 16 is the left channel, L, of the binaural audio signal emitted from the virtualizer in Figure 2, and the output of element 18 is the right channel, R, of the binaural audio signal emitted from the virtualizer in Figure 2. The important features of typical embodiments of the invention become apparent from a comparison of the embodiment of the inventive headphone virtualizer in Figure 2 with the conventional headphone virtualizer in Figure 1. For comparison purposes, we assume that the systems in Figure 1 and Figure 2 are configured such that, when the same multi-channel audio input is imposed on each, the systems apply a BRIR, having the same forward and early reflection portion (i.e., the relevant EBRIRI in Figure 2), to each full-range frequency channel, Xi, of the input signal (although not necessarily with the same degree of success). Each BRIRi applied by the system in Figure 1 or Figure 2 can be decomposed into two parts: a forward and early reflection portion (e.g., one of the EBIR1 portions, ..., EBRIRN applied by subsystems 12-14 of Figure 2), and a late reverb portion. The embodiment of Figure 2 (and other typical embodiments of the invention) assumes that the late reverb portions of the single-channel BRIRs, BRIRi, can be shared across the source directions and thus all channels, thereby applying the same late reverb (i.e., a common late reverb) to a downmix of all channels across the full frequency range of the input signal. This downmix can be a monophonic (mono) downmix of all channels, but alternatively it can be a stereo or multi-channel downmix obtained from the input channels (e.g., from a subset of the input channels). More specifically, subsystem 12 in Figure 2 is configured to convolve channel X1 of the input signal with EBRIR1 (the forward response and early reflection portion of the BRIR for the corresponding source direction), subsystem 14 is configured to convolve channel XN with EBRIRN (the forward response and early reflection portion of the BRIR for the corresponding source direction), and so on. The late reverb subsystem 15 in Figure 2 is configured to generate a mono downmix of all full-range frequency channels of the input signal and to convolve the downmix with LBRIR (a common late reverb for all downmixed channels). The output of each BRIR subsystem of the virtualizer in Figure 2 (each of subsystems 12, ..., 14, and 15) includes a left channel and a right channel (of a binaural signal generated from the corresponding speaker channel or downmix). The left channel outputs of the BRIR subsystems are combined (mixed) in addition element 16, and the right channel outputs of the BRIR subsystems are combined (mixed) in addition element 18. Addition element 16 can be implemented to simplify the summation of the samples from the Left binaural channel (the Left channel outputs of subsystems 12, ..., 14, and 15) to generate the Left channel of the binaural output signal, assuming that appropriate level adjustments and time alignments are implemented in subsystems 12, ..., 14, and 15. Similarly, addition element 18 can also be implemented to simplify the summation of the samples from the Right binaural channel (the Right channel outputs of subsystems 12, ..., 14, and 15) to generate the Right channel of the binaural output signal, assuming that appropriate level adjustments and time alignments are implemented in subsystems 12, ..., 14, and 15. Subsystem 15 in Figure 2 can be implemented in a variety of ways, but it typically includes at least one feedback delay network configured to apply the common late reverb to a monophonic downmix of the input signal channels imposed upon it. Typically, where each of subsystems 12, ..., 14 applies a forward response and early reflection (EBRIRi) portion of a single-channel BRIR to the channel (Xi) it processes, the common late reverb is generated to emulate the collective macro attributes of the late reverb portions of at least some (e.g., all) of the single-channel BRIRs (whose "forward response and early reflection portions" are applied by subsystems 12, ..., 14). For example, one implementation of subsystem 15 has the same structure as subsystem 200 in Figure 3, which includes a bank of networks (203, 204, ..., 205) of feedback delay configured to apply a common late reverb to a monophonic downmix of the input signal channels imposed on it. The subsystems 12, ..., 14 in Figure 2 can be implemented in any of a variety of ways (in either the time domain or the filter bank domain), with the preferred implementation for any given application depending on various considerations, such as (for example) performance, computation, and memory. In an exemplary implementation, each of the subsystems 12, ..., 14 is configured to convolve the channel imposed upon it with an FIR filter corresponding to the forward and early responses associated with the channel, with the gain and delay set appropriately so that the outputs of subsystems 12, ..., 14 can be efficiently simplified and combined with those of subsystem 15. Figure 3 is a block diagram of another embodiment of the inventive headphone virtualization system. The embodiment in Figure 3 is similar to that in Figure 2, with two time-domain signals (left and right channels) outputting from a forward response and early reflection processing subsystem 100, and two time-domain signals (left and right channels) outputting from a late reverb processing subsystem 200. The addition element 210 is coupled to the outputs of subsystems 100 and 200.Element 210 is configured to combine (mix) the left channel outputs of subsystems 100 and 200 to generate the left channel, L, of the binaural audio signal output of the virtualizer in Figure 3, and to combine (mix) the right channel outputs of subsystems 100 and 200 to generate the right channel, R, of the binaural audio signal output of the virtualizer in Figure 3. Element 210 can be implemented to simplify the summation of the left channel samples emitted from subsystems 100 and 200 to generate the left channel of the binaural output signal, and to simplify the summation of the right channel samples emitted from subsystems 100 and 200 to generate the right channel of the binaural output signal, assuming that appropriate level adjustments and time alignments are implemented in subsystems 100 and 200. In the system of Figure 3, the channels, Xi, of the multichannel audio signal are directed to, and processed in, two parallel processing paths: one through subsystem 100 for processing the direct response and early reflection; the other through subsystem 200 for processing the late reverberation. The system in Figure 3 is configured to apply a BRIRi to each channel, Xi. Each BRIRi can be decomposed into two parts: a direct response and early reflection portion (applied by subsystem 100), and a late reverberation portion (applied by subsystem 200).In operation, the direct response and early reflection processing subsystem 100 thus generates the direct response and early reflection portions of the binaural audio signal output from the virtualizer, and the late reverb processing subsystem 200 ("late reverb generator") thus generates the late reverb portion of the binaural audio signal output from the virtualizer. The outputs of subsystems 100 and 200 are mixed (via the addition subsystem 210) to generate the binaural audio signal, which is typically imposed from subsystem 210 to a processing system (not shown) where the binaural processing is experienced for playback through headphones. Normally, when processed and played back through a pair of headphones, a standard binaural audio signal emitted from element 210 is perceived at the listener's eardrum as sound from "N" loudspeakers (where N ≥ 2 and N is typically equal to 2, 5, or 7) in any of a wide variety of positions, including positions in front of, behind, and above the listener. Playback of the output signals generated in the operation of the system in Figure 3 can give the listener the experience that the sound is coming from more than two (e.g., five or seven) "surround" sources. At least some of these sources are virtual. The forward and early reflection subsystem 100 can be implemented in a variety of ways (in either the time domain or the filter bank domain), with the preferred implementation for any given application depending on various considerations, such as (for example) performance, computation, and memory. In one exemplary implementation, subsystem 100 is configured to convolve each channel imposed upon it with an FIR filter corresponding to the forward and early responses associated with the channel, with the gain and delay set appropriately so that the output of subsystem 100 can be combined simply and efficiently (in element 210) with these subsystem 200. As shown in Figure 3, the late reverb generator 200 includes a downmix subsystem 201, a filter bank 202, an FDN bank (FDN 203, 204, ..., 205), and a synthesis filter bank 207, coupled as shown. The subsystem 201 is configured to downmix the channels of the multi-channel input signal into a mono downmix, and the analysis filter bank 202 is configured to apply a transformation to the mono downmix to divide the mono downmix into "K" frequency bands, where K is an integer. The values in the filter bank domain (filter bank output 202) in each of the different frequency bands are imposed on a different one of the FDNs 203, 204, ..., 205 (there are "K" of these FDNs, each coupled and configured to apply a late reverb portion of a BRIR to the filter bank domain values imposed on it).The values in the filter bank domain are preferably decimated over time to reduce the computational complexity of the FDN. In principle, each input channel (to subsystem 100 and subsystem 201 in Figure 3) can be processed in its own FDN (or bank of FDNs) to simulate the late reverberation portion of its BRIR. Despite the fact that the late reverberation portions of the BRIRs associated with different sound source locations are typically very different in terms of root mean square differences in impulse responses, their statistical attributes, such as their mean energy spectrum, energy decay structure, modal density, peak density, and the like, are often very similar. Therefore, the late reverberation portion of a set of BRIRs is usually perceptually quite similar across channels, and it is thus possible to use a common FDN or bank of FDNs (e.g., FDNs 203, 204, ..., 205) to simulate the late reverberation portion of two or more BRIRs.In typical implementations, a common FDN (or FDN bank) is used, and its input consists of one or more downmixes constructed from the input channels. In the exemplary implementation of Figure 2, the downmix is a monophonic downmix (imposed on the output of subsystem 201) of all input channels. With reference to the implementation in Figure 2, each of the FDNs 203, 204, ..., 205 is implemented in the filter bank domain and coupled and configured to process a different frequency band from the values emitted by the analysis filter bank 202, in order to generate the left and right reverberated signals for each band. For each band, the left reverberated signal is a sequence of values in the filter bank domain, and the right reverberated signal is another sequence of values in the filter bank domain.The synthesis filter bank 207 is coupled and configured to apply a frequency-to-time domain transform to the 2K sequences of filter bank domain values (e.g., QMF frequency domain components) output from the FDNs, and to validate the transformed values as a left-channel time-domain signal (indicative of the audio content of the mono downmix to which the late reverb is applied) and a right-channel time-domain signal (also indicative of the audio content of the mono downmix to which the late reverb is applied). These left- and right-channel signals are output to element 210. In a typical implementation each of the FDNs 203, 204, ..., 205, is implemented in the QMF domain, and the filter bank 202 transforms the mono downmix of subsystem 201 into the QMF domain (e.g., in the hybrid complex quadrature mirror filter (HCQMF) domain), so that the signal imposed from the filter bank 202 to an output of each FDN 203, 204, ..., 205 is a sequence of frequency components in the QMF domain. In this implementation, the signal imposed from filter bank 202 to FDN 203 is a sequence of frequency components in the QMF domain in a first frequency band, the signal imposed from filter bank 202 to FDN 204 is a sequence of frequency components in the QMF domain in a second frequency band, and the signal imposed from filter bank 202 to FDN 205 is a sequence of frequency components in the QMF domain in a "K-th" frequency band.When the analysis filter bank 202 is implemented in this way, the synthesis filter bank 207 is configured to apply a transform from the QMF domain to the time domain to the 2K output sequences of the frequency components in the QMF domain of the FDN, to generate the time domain signals with late reverberation of the left and right channels that are output to element 210. For example, if K=3 in the system of Figure 3, then there are six inputs to the synthesis filter bank 207 (left and right channels, comprising samples in the frequency domain or in the QMF domain, output from each FDN 203, 204, and 205) and two outputs from 207 (left and right channels, each consisting of samples in the time domain). In this example, filter bank 207 would normally be implemented as two synthesis filter banks: one (to which the three left channels of FDN 203, 204, and 205 would be imposed) configured to generate the output of the left channel signal in the time domain from filter bank 207; and a second one (to which the three right channels of FDN 203, 204, and 205 would be imposed) configured to generate the output of the right channel signal in the time domain from filter bank 207. Optionally, the control subsystem 209 is coupled to each of the FDNs 203, 204, ..., 205 and configured to impose control parameters on each FDN to determine the late reverb portion (LBRIR) applied by subsystem 200. Examples of these control parameters are described later. It is envisaged that in some implementations, the control subsystem 209 will be operable in real time (for example, in response to user commands imposed on it by an input device) to implement real-time variation of the late reverb portion (LBRIR) applied by subsystem 200 to the mono downmix of the input channels. For example, if the input signal to the system in Figure 2 is a 5.1-channel signal (whose full-range frequency channels are in the following channel order: L, R, C, Ls, Rs), all full-range frequency channels are the same distance from the source, and the downmix subsystem 201 can be implemented as the following downmix matrix, which simply sums the full-range frequency channels to form a mono downmix: After filtering, I pass everything (in element 301 in each of the FDN 203, 204, ..., and 205), the mono downmix mixes the four reverb tanks up in such a way as to conserve energy: Alternatively (as an example), we can choose to assign the left-side channels to the first two reverb tanks, the right-side channels to the last two reverb tanks, and the center channel to all reverb tanks. In this case, the downmix subsystem 201 would be implemented to form two downmix signals: In this example, the upmix of the reverb tanks (in each of the FDN 203, 204, ..., and 205) is: Since there are two downmix signals, the all-pass filtering (at element 301 in each of the FDNs 203, 204, ..., and 205) needs to be applied twice. Diversity would be introduced for the late responses of (L, LS), (R, Rs), and C, even though they all have the same macro attributes. When the input signal channels are at different distances from the source, it would still be necessary to apply the appropriate delays and gains in the downmix process. Next, we will describe considerations for specific implementations of the downmixing subsystem 201, and the virtualizer subsystems 100 and 200 in Figure 3. The downmixing process implemented by subsystem 201 depends on the distance to the source (between the sound source and the assumed listener position) for each channel to be mixed, and the handling of the forward response. The forward response delay td is: where d is the distance between the sound source and the listener, and vs is the speed of sound. Furthermore, the gain of the forward response is proportional to 1 / d. If these rules are maintained when handling the forward responses of channels with different distances from the source, subsystem 201 can implement a direct downmix of all channels, since the delay and late reverb level are generally insensitive to the source location. For practical reasons, virtualizers (e.g., subsystem 100 of the virtualizer in Figure 3) can be implemented to time-align the forward responses of input channels with varying distances from the source. To maintain the relative delay between the forward response and the late reverb for each channel, a channel with distance d from the source should be delayed by (dmax - d) / vs before being downmixed with other channels. Here, dmax denotes the maximum possible distance from the source. Virtualizers (for example, subsystem 100 of the virtualizer in Figure 3) can be implemented to compress the dynamic range of direct responses. For example, the direct response for a channel with a distance d from the source can be scaled by a factor of d-1, where d ≥ 0, instead of d-1. To maintain the level of difference between the direct response and the late reverb, it may be necessary to implement the downmixing subsystem 201 to scale a channel with a distance d from the source by a factor of d-1 before downmixing it with the other scaled channels. The feedback delay network in Figure 4 is an exemplary implementation of FDN 203 (or 204, or 205) in Figure 3. Although the system in Figure 4 has four reverberation tanks (each including a gain stage, gi, and a delay line, z-ni, coupled to the output of the gain stage), variations on the system (and other FDNs employed in the inventive virtualizer's implementations) implement more or fewer than four reverberation tanks. The FDN of Figure 4 includes an input gain element 300, an all-pass filter 301 coupled to the output of element 300, addition elements 302, 303, 304, and 305 coupled to the output of APF 301, and four reverb tanks (each comprising a gain element, gk (one of elements 306), a delay line, z-Mk (one of elements 307) coupled thereto, and a gain element, 1 / gk (one of elements 309) coupled thereto, where 0 ≤ k-1 ≤ 3) each coupled to the output of one different element 302, 303, 304, and 305. The unit matrix 308 is coupled to the outputs of the delay lines 307 and is configured to impose a feedback output on a second entry of each of the elements 302, 303, 304, and 305.The outputs of the two gain elements 309 (from the first and second reverb tanks) are imposed on the inputs of the add element 310, and the output of element 310 is imposed on one input of the output mix matrix 312. The outputs of the other two gain elements 309 (from the third and fourth reverb tanks) are imposed on the inputs of the add element 311, and the output of element 311 is imposed on the other input of the output mix matrix 312. Element 302 is configured to add the output of array 308 corresponding to delay line z-n1 (i.e., to apply feedback from the output of delay line z-n1 through array 308) to the input of the first reverb tank. Element 303 is configured to add the output of array 308 corresponding to delay line z-n2 (i.e., to apply feedback from the output of delay line z-n2 through array 308) to the input of the second reverb tank. Element 304 is configured to add the output of array 308 corresponding to delay line z-n3 (i.e., to apply feedback from the output of delay line z-n3 through array 308) to the input of the third reverb tank.Element 305 is configured to add the output of array 308 corresponding to z-n4 delay line (i.e., to apply feedback from the output of z-n4 delay line through array 308) to the input of the fourth reverb tank. The input gain element 300 of the FDN in Figure 4 is coupled to receive a frequency band of the transformed monophonic downmix signal (a signal in the filter bank domain) that is an output of the analysis filter bank 202 in Figure 3. The input gain element 300 applies a gain (scaling) factor, Gentrada, to the signal in the filter bank domain imposed on it. Collectively, the Gentrada scaling factors (implemented by all FDNs 203, 204, ..., 205 in Figure 3) for all frequency bands control the spectral shaping and late reverb level. Setting the input gains, Gentrada, in all FDNs of the virtualizer in Figure 3 often takes into account the following objectives: a direct-to-latency ratio (DLR) of the BRIR applied to each channel, matching actual rooms; a low-frequency attenuation necessary to mitigate excessive blending artifacts and / or low-frequency noise; and to match the envelope of the diffuse field spectra. If we assume that the direct response (applied by subsystem 100 in Figure 3) provides unity gain in all frequency bands, a specific DLR (power ratio) can be achieved by adjusting Gentrada to be: where T60 is the reverb decay time defined as the time it takes for the reverb to decay by 60 dB (it is determined by the reverb delays and reverb gains discussed above), and "ln" denotes the natural logarithmic function. The input gain factor, Gentrada, can be dependent on the content being processed. One application of this content dependence is to ensure that the downmix energy in each time / frequency segment is equal to the sum of the energies of the individual channel signals being downmixed, regardless of any correlation that may exist between the input channel signals. In this case, the input gain factor can be (or can be multiplied by) a term similar to or equal to: wherein i is an index over all downmix samples of a time / frequency strip or sub-strip, and (i) are the downmix samples for the strip, and xi (j) is the input signal (for channel Xi) imposed on the input of the downmix subsystem 201. In a typical QMF-domain implementation of the FDN in Figure 4, the signal imposed from the output of the all-pass filter 301 (APF) to the reverb tank inputs is a sequence of frequency components in the QMF domain. To generate a more natural-sounding FDN output, the APF 301 is applied to the output of the gain element 300 to introduce phase diversity and increase echo density. Alternatively, or additionally, one or more all-pass delay filters may be applied to: the individual inputs to the downmix subsystem 201 (of Figure 3) before they are mixed in subsystem 201 and processed by the FDN; or in the forward feed or feedback paths of the reverb tank depicted in Figure 4 (e.g., in addition to or instead of the z-Mk delay lines in each reverb tank; or the FDN outputs (i.e., to the outputs of the 312 output array). In implementing the reverb tank delays, z-ni, the reverb delays (ni) should be mutually prime numbers to prevent the reverb modes from aligning at the same frequency. The sum of the delays should be large enough to provide sufficient modal density to avoid an artificial sound. However, the shorter delays must be short enough to prevent excessive time hopping between the late reverb and the other components of the BRIR. Typically, the reverb tank outputs are distributed to either the left or right binaural channel. The sets of reverb tank outputs distributed to the two binaural channels are usually equal in number and mutually exclusive. It's also desirable to balance the timing of the two binaural channels. Therefore, if the reverb tank output with the shortest delay goes to one binaural channel, the output with the second shortest delay goes to the other channel. The reverberation tank delays can vary across frequency bands to change the modal density as a function of frequency. Generally, lower frequency bands require higher modal density, hence longer reverberation tank delays. The reverberation tank gain amplitudes, gi, and the reverberation tank delays together determine the reverberation decay time of the FDN in Figure 4. where FFRM is the frame rate of filter bank 202 (from Figure 3). The reverb tank gain phases introduce fractional delays to overcome the reverb tank delay-related problems that are quantized for the filter bank sample decay factor grid. The 308 unit feedback array provides a uniform downmix between the reverb tanks in the feedback path. To equalize the levels of the reverb tank outputs, the 309 gain elements apply a normalization gain, 1 / |gi|, to the output of each reverb tank, to eliminate the impact of the reverb tank gain levels while maintaining the fractional delays introduced by their phases. The output downmix matrix 312 (also defined as Mout) is a 2x2 matrix configured to downmix the unmixed binaural channels (the outputs of elements 310 and 311, respectively) from the initial distribution to achieve the left and right output binaural channels (the L and R signals imposed on the output of matrix 312) that have the desired interaural coherence. The unmixed binaural channels are close to being uncorrelated after the initial distribution since they do not consist of any common output from the reverb tank. If the desired interaural coherence is Coh, where |Coh| ≥ 1, the output downmix matrix 312 can be defined as: Since the reverb tank delays are different, one of the unmixed binaural channels would constantly lead the other. If the combination of reverb tank delays and the distribution pattern were identical across the frequency bands, a sound image bias would occur. This bias can be mitigated if the distribution pattern alternates across the frequency bands so that the mixed binaural channels lead and lead each other in alternating frequency bands.This can be achieved by implementing the output mixing matrix 312 to have the shape set out in the previous paragraph in the odd frequency bands (i.e., in the first frequency band (processed by FDN 203 in Figure 3), the third frequency band, and so on), and to have the following shape in the even frequency bands (i.e., in the second frequency band (processed by FDN 204 in Figure 3), the fourth frequency band, and so on): where the definition of β remains the same. It should be noted that the matrix 312 can be implemented to be identical in the FDNs for all frequency bands, but the channel order of its inputs can be switched to alternate one of the frequency bands (for example, the output of element 310 can be imposed on the first input of matrix 312 and the output of element 311 can be imposed on the second input of matrix 312 in odd frequency bands, and the output of element 311 can be imposed on the first input of matrix 312 and the output of element 310 can be imposed on the second input of matrix 312 in even frequency bands). In the case where the frequency bands overlap (partially), the width of the frequency range over which the 312 matrix shape alternates can be increased (for example, it could alternate once for every two or three consecutive bands), or the value of β in the above expressions (for the 312 matrix shape) can be adjusted to ensure that the mean coherence equals the desired value to compensate for the spectral overlap of consecutive frequency bands. If the previously defined target acoustic attributes T60, Coh, and DLR are known for the FDN for each specific frequency band in the inventive virtualizer, each of the FDNs (each of which may have the structure shown in Figure 4) can be configured to achieve the target attributes. Specifically, in some embodiments, the input gain (Gentrada), the reverb tank gains and delays (gi and ni), and the output matrix parameters Msalida for each FDN can be set (for example, by means of control values imposed thereon by the control subsystem 209 in Figure 3) to achieve the target attributes according to the relationships described herein.In practice, establishing frequency-dependent attributes using models with simple control parameters is often sufficient to generate natural-sounding late reverberation that matches specific acoustic environments. The following describes an example of how a target reverberation decay time (T60) for the FDN can be determined for each specific frequency band of an implementation of the inventive virtualizer, by determining the target reverberation decay time (T60) for each of a small number of frequency bands. The response level of the FDN decays exponentially over time. T60 is inversely proportional to the decay factor, df (defined as dB of decay per unit of time): The decay factor, df, depends on the frequency and generally increases linearly compared to the logarithmic frequency scale. Therefore, the reverberation decay time is also a function of frequency, generally decreasing as the frequency increases. Thus, if the T60 values are determined (e.g., established) for two frequency points, the T60 curve for all frequencies is determined. For example, if the reverberation decay times for frequency points fA and fB are T60, A and T60, B, respectively, the T60 curve is defined as: Figure 5 shows an example of the T60 curve that can be achieved by an implementation of the inventive virtualizer for which the T60 value at each of the two specific frequencies (fA and fB) is set: T60, A = 320 ms with fA = 10 Hz, and T60, B = 150 ms with fB = 2.4 kHz. Next, we describe an example of how the target Interaural Coherence (Coh) for the FDN can be achieved for each specific frequency band of an implementation of the inventive virtualizer by setting a small number of control parameters. The target Interaural Coherence (Coh) of the late reverb largely follows the pattern of a diffuse sound field. It can be modeled by a sine function up to a cutoff frequency fC, and a constant above the cutoff frequency. A simple model for the Coh curve is: where the parameters Cohmin and Cohmax satisfy -1 COhmin < Cohmax 1, and control the range of the Coh. The optimal cutoff frequency fC depends on the size of the listener's head. An fC that is too large leads to an internalized sound source image, while a value that is too small leads to a dispersed or split sound source image. Figure 6 is an example of a Coh curve that can be achieved by an implementation of the inventive virtualizer for which the control parameters Cohmax, Cohmin, and fC are set to have the following values: Cohmax = 0.95, Cohmin = 0.05, and fC = 700 Hz. Next, we will describe an example of how a target direct-to-late ratio (DLR) for the FDN can be achieved for each specific frequency band of an inventive virtualizer implementation by setting a small number of control parameters. The direct-to-late ratio (DLR), in dB, generally increases linearly compared to the logarithmic frequency scale. It can be controlled by setting the DLR1K (DLR in dB @ 1 kHz) and the DLRpend (in dB per 10x frequency). However, a low DLR in the lower frequency range often results in excessive combing artifacts. To mitigate these artifacts, two modification mechanisms are added to the DLR control: a minimum floor DLR, DLRmin (in dB); and a high-pass filter defined by a transition frequency, fT, and the slope of the attenuation curve below it, HPFpend (in dB per 10x frequency). The resulting DLR curve is defined as: It should be noted that the DLR changes with the distance to the source, even in the same acoustic environment. Therefore, both the DLR1K and DLRmin values in this document are for a nominal distance to the source, such as 1 meter. Figure 7 is an example of a DLR curve for a distance to the source of 1 meter, achieved using an implementation of the inventive virtualizer with the control parameters DLR1K, DLRpend, DLRmin, HPFpend, and fT set to the following values: DLR1K = 18 dB, DLRpend = 6 dB / 10x frequency, DLRmin = 18 dB, HPFpend = 6 dB / 10x frequency, and fT = 200 Hz. The variations on the realizations described in this document have one or more of the following characteristics: The inventive virtualizer's FDNs are implemented in the time domain, or have a hybrid implementation with FDN-based impulse response capture and FIR-based signal filtering; The inventive virtualizer is implemented to allow the application of power compensation as a function of frequency during the execution of the downmix step that generates the downmixed input signal for the late reverb processing subsystem; and The inventive virtualizer is implemented to allow manual or automatic control of the attributes of the applied late reverberation in response to external factors (i.e., in response to the setting of the control parameters). For applications where system latency is critical and the delay caused by analysis and synthesis filter banks is prohibitive, the filter bank domain FDN structure of typical inventive virtualizer implementations can be transposed to the time domain, and each FDN structure can be implemented in the time domain in a class of virtualizer implementations. In the time-domain implementations, the subsystems that apply the input gain factor (Gin), the reverb tank gains (gi), and the gains (1 / |gi|) are replaced by filters with similar amplitude responses to allow frequency-dependent controls. The output mix matrix (Mout) is also replaced by a filter matrix.Unlike the other filters, the phase response of this filter array is critical, as energy conservation and interaural coherence can be affected by it. The reverb tank delays in a time-domain implementation may need to be slightly varied (from their values in a filter bank-domain implementation) to avoid sharing the filter bank step as a common factor. Due to various limitations, the performance of frequency-domain implementations of the inventive virtualizer's FDNs does not exactly match that of its filter bank-domain implementations. With reference to Figure 8, we will now describe a hybrid implementation (filter bank domain and time domain) of the inventive late reverb processing subsystem of the inventive virtualizer. This hybrid implementation of the inventive late reverb processing subsystem is a variation of the late reverb processing subsystem 200 in Figure 4, which implements impulse response capture based on an FDN and signal filtering based on an FIR. The implementation of Figure 8 includes elements 201, 202, 203, 204, 205, and 207, which are identical to the identically numbered elements of subsystem 200 in Figure 3. The preceding description of these elements will not be repeated with reference to Figure 8. In the implementation of Figure 8, the unity impulse generator 211 is coupled to impose an input signal (a pulse) on the analysis filter bank 202. A LBRIR filter 208 (mono input, stereo output) implemented as an FIR filter applies the appropriate late reverb portion of the BRIR (the LBRIR) to the mono downmixed output of subsystem 201. Therefore, elements 211, 202, 203, 204, 205, and 207 are a processing sidechain to the LBRIR filter 208. Whenever the LBRIR setting of the late reverb portion needs to be modified, the pulse generator 211 is operated to impose a unit pulse on element 202, and the resulting output from filter bank 207 is captured and imposed on filter 208 (to set filter 208 to apply the new LBRIR determined by the output of filter bank 207). To accelerate the time interval between the change in LBRIR setting and the moment the new LBRIR takes effect, samples of the new LBRIR can begin replacing the old LBRIR as they become available. To shorten the inherent latency of the FDRs, the leading zeros of the LBRIR can be discarded. These options provide flexibility and allow the hybrid implementation to provide a potential performance improvement (relative to that provided by the filter bank-domain implementation), at the cost of one additional computation to the FIR filtering. For applications where system latency is critical, but computational load is less important, the late reverb processor in the sidechain filter bank domain (e.g., the one implemented by elements 211, 202, 203, 204, ..., 205, and 207 in Figure 8) can be used to capture the effective FIR impulse response to be applied by filter 208. FIR filter 208 can then implement this captured FIR response and apply it directly to the mono downmix of the input channels (during input channel virtualization). The various parameters of the FDN, and therefore the resulting attributes of late reverberation, can be manually tuned and subsequently hardwired into an implementation of the inventive late reverberation processing subsystem, for example, by means of one or more presets that can be adjusted (for example, by operating control subsystem 209 in Figure 3) by the system user. However, given the high-level description of late reverberation, its relationship to the FDN parameters, and the ability to modify its behavior, a wide variety of procedures are conceived for controlling the various implementations of the FDN-based late reverberation processor, including (but not limited to) the following: 1. The end user can manually control the FDN parameters, for example, through a user interface in a presentation element (e.g., implemented by an implementation of the control subsystem 209 in Figure 3) or by toggling presets using physical controls (e.g., implemented by an implementation of the control subsystem 209 in Figure 3). In this way, the end user can adapt the room simulation according to their taste, the environment, or the content; 2. The author of the audio content to be virtualized can provide the desired settings or parameters that are carried along with the content itself, for example, through metadata provided with the input audio signal. Such metadata can be analyzed and used (for example, by means of an implementation of the control subsystem 209 in Figure 3) to control the relevant parameters of the FDN. The metadata can therefore be indicative of properties such as reverberation time, reverberation level, direct ratio to reverberation, and so on, and these properties can be time-varying, signaled by time-varying metadata. 3. A playback device can be aware of its location or environment by means of one or more sensors. For example, a mobile device can use GSM networks, the Global Positioning System (GPS), known Wi-Fi access points, or any other location service to determine where the device is. Subsequently, location and / or environment data can be used (for example, by means of an implementation of the control subsystem 209 in Figure 3) to control the relevant FDN parameters. Thus, the FDN parameters can be modified in response to the device's location, for example, to mimic the physical environment: 4. Regarding the location of the playback device, a cloud service or social media platform can be used to derive the most common settings used by consumers in a given environment. Additionally, users can upload their current settings to a cloud service or social media platform, associated with their (known) location, to make them available to other users, or to themselves. 5. A playback device may contain other sensors such as a camera, a light sensor, a microphone, an accelerometer, a gyroscope, to determine the user's activity and the environment in which the user is, in order to optimize the FDN parameters for that specific activity and / or environment; 6. The FDN parameters can be controlled by the audio content. Audio classification algorithms, or manually annotated content, can indicate whether audio segments comprise words, music, sound effects, silence, and the like. The FDN parameters can then be adjusted according to these labels. For example, the direct ratio to reverb can be reduced for dialogue to improve its intelligibility. Additionally, video analytics can be used to determine the location of a current video segment, and the FDN parameters can be adjusted accordingly to more closely simulate the environment depicted in the video; and / or 7. A solid-state playback system may use different FDN settings than a mobile device; for example, the settings may be device-dependent. A solid-state system in a living room might simulate a typical (and quite reverberant) living room environment with distant sources, while a mobile device might play content closer to the listener. Some implementations of the inventive virtualizer include FDNs (for example, an implementation of the FDN in Figure 4) that are configured to apply a fractional delay as well as an integer sample delay. For example, in such an implementation, a fractional delay element is connected in each reverb tank in series with a delay line that applies an integer delay equal to an integer number of sample periods (for example, each fractional delay element is positioned after or otherwise in series with one of the delay lines). The fractional delay can be approximated by a phase shift (complex unity multiplication) in each frequency band corresponding to a fraction of the sample period: f = Δf / T, where f is the fractional delay, T is the desired delay for the band, and T is the sample period for the band.It is well known how to apply fractional delay in the context of applying reverberation in the QMF domain. In a further class of embodiments, not covered by the claims, the invention is a headphone virtualization method for generating a binaural signal in response to a set of channels (e.g., each of the channels, or each of the full-frequency range channels) of a multi-channel audio input signal, comprising the steps of: (a) applying a binaural one-room impulse response (BRIR) to each channel of the set (e.g., by convolving each channel of the set with a BRIR corresponding to that channel, in subsystems 100 and 200 of Figure 3, or in subsystems 12, ..., 14, and 15 of Figure 2), thereby generating filtered signals (e.g., the outputs of subsystems 100 and 200 of Figure 3, or the outputs of subsystems 12, ..., 14, and 15 of Figure 2), including by using at least one network of feedback delay (for example, FDN 203, 204, ...(a) applying a common late reverb to a downmix (e.g., a monophonic downmix) of the ensemble channels; and (b) combining the filtered signals (e.g., in subsystem 210 of Figure 3, or the subsystem comprising elements 16 and 18 of Figure 2) to generate the binaural signal. Typically, a bank of FDNs is used to apply the common late reverb to the downmix (e.g., with each FDN applying late reverb to a different frequency band). Typically, step (a) includes applying to each ensemble channel a "direct response and early reflection" portion of a single-channel BRIR to that channel (e.g., in subsystem 100 of Figure 3 or subsystems 12, ..., 14 of Figure 2), and the common late reverb has been generated to emulate the collective macro attributes of the late reverb portions of at least some (e.g., all) of the single-channel BRIRs. Each of the FDNs is implemented in either the hybrid complex quadrature mirror filter (HCQMF) or quadrature mirror filter (QMF) domain. In some of these implementations, the frequency-dependent spatial acoustic attributes of the binaural signal are controlled (for example, using the control subsystem 209 in Figure 3) by controlling the configuration of each FDN used to apply the late reverb. Typically, a monophonic downmix of the channels (for example, the downmix generated by subsystem 201 in Figure 3) is used as the input to the FDNs for efficient binaural processing of the multichannel signal's audio content.Typically, the downmixing process is controlled based on the distance to the source for each channel (i.e., the distance between the assumed source of the audio content and the assumed user position) and relies on managing the direct responses corresponding to those distances to the source to preserve the temporal and level structure of each BRIR (i.e., each BRIR determined by the direct response and early reflection portions of a single-channel BRIR for a channel, along with the late reverb for a downmix that includes the channel). Although the channels being downmixed can be time-aligned and scaled in different ways during the downmix, the temporal and level relationship between the direct response, early reflection, and common late reverb portions of the BRIR for each channel should be maintained.In implementations where a single FDN bank is used to generate the common late reverb portion for all channels being downmixed (to generate a downmix), appropriate gain and delay must be applied (to each downmixed channel) during the generation of the downmix. Typical implementations in this class include an adjustment step (e.g., using the control subsystem 209 in Figure 3) of the FDN coefficients corresponding to frequency-dependent attributes (e.g., reverberation decay time, interaural coherence, modal density, and direct-to-late ratio). This allows for better matching of acoustic environments and more natural sound outputs. In a further class of embodiments not covered by the appended claims, the invention is a method for generating a binaural signal in response to a multi-channel audio input signal, by applying a binaural one-room impulse response (BRIR) to each channel (e.g., by convolving each channel with the corresponding BRIR) of a set of input signal channels (e.g., each of the input signal channels or each of the full-frequency range channels of the input signal), including: processing each channel of the set in a first processing path (e.g., implemented by subsystem 100 of Figure 3 or subsystems 12, ...14 in Figure 2) is configured to model and apply to each channel a portion of the forward response and early reflection (e.g., the EBRIR applied by subsystem 12, 14, or 15 in Figure 2) of a single-channel BRIR for that channel; and to process a downmix (e.g., a monophonic downmix) of the channels in the ensemble on a second processing path (e.g., implemented by subsystem 200 in Figure 3 or subsystem 15 in Figure 2), in parallel with the first processing path. The second processing path is configured to model and apply to the downmix a common late reverb (e.g., the LBRIR applied by subsystem 15 in Figure 2). Typically, the common late reverb emulates the collective macro attributes of the late reverb portions of at least some (e.g., all) of the single-channel BRIRs.Typically, the second processing path includes at least one FDN (for example, one FDN for each of the multiple frequency bands). A mono downmix is typically used as the input to all reverb tanks for each FDN implemented by the second processing path. Mechanisms (for example, control subsystem 209 in Figure 3) are typically provided for the systematic control of each FDN's macro attributes to better simulate acoustic environments and produce a more natural binaural virtualization of the sound. Since most of these attributes are frequency-dependent, each FDN is typically implemented in the hybrid complex quadrature mirror filter (HCQMF) domain, the frequency domain, or another domain of the filter bank, and a different FDN is used for each frequency band.A primary benefit of implementing FDNs in the filter bank domain is that they allow the application of reverb with frequency-dependent reverb properties. In various implementations, FDNs are implemented in any of a wide variety of filter bank domains, using any of a variety of filter banks, including, but not limited to, quadrature mirror filters (QMFs), finite impulse response filters (FIR filters), infinite impulse response filters (IIR filters), or crossover filters. Some implementations of these additional classes implement one or more of the following features: 1. An implementation of an FDN (e.g., the FDN implementation of Figure 4) in the filter bank domain (e.g., in the domain of the hybrid complex quadrature mirror filter), or an implementation of an FDN in the hybrid filter bank domain and an implementation of the late reverb filter in the time domain (e.g., the structure described with reference to Figure 8), which typically allows independent adjustment of the parameters and / or settings of the FDN for each frequency band (allowing simple and flexible control of frequency-dependent acoustic attributes), e.g., by providing the ability to vary the reverb tank delays in different bands to change the modal density as a function of frequency; 2. The specific downmixing process, employed to generate (from the multi-channel input audio signal) the downmixed signal (e.g., mono downmix) processed in the second processing path, depends on the distance to the source of each channel and the handling of the forward response to maintain the appropriate level and timing relationship between the forward and late responses; 3. An all-pass filter (e.g., the APF 301 in Figure 4) is applied in the second processing path (e.g., at the input or output of an FDN bank) to introduce phase diversity and increase echo density without changing the spectrum and / or timbre of the resulting reverberation; 4. Fractional delays are implemented in the feedback path of each FDN in a multi-rate structure, with a complex value, to overcome problems related to delays quantized to the sample decay factor grid; 5. In FDNs, the reverb tank outputs are linearly downmixed directly into the binaural channels (e.g., using the 312 matrix in Figure 4) using output mix coefficients set based on the desired interaural coherence in each frequency band. Optionally, the reverb tanks are mapped to the binaural output channels, alternating between frequency bands to achieve a balanced delay between the binaural channels. Also optionally, normalization factors are applied to the reverb tank outputs to equalize their levels while preserving fractional delay and overall energy. 6. The frequency-dependent reverberation decay time is controlled (e.g., using the control subsystem 209 in Figure 3) by setting an appropriate combination of reverb tank gains and delays in each frequency band to simulate real rooms; 7. A scaling factor (e.g., by means of elements 306 and 309 in Figure 4) is applied per frequency band (e.g., at either the input or output of the relevant processing path), to: control the direct-to-late frequency-dependent ratio (DLR) that matches that of a real room (a simple model can be used to calculate the required scaling factor based on the target DLR and reverberation decay time, e.g., T60); provide low-frequency attenuation to mitigate excessive combing artifacts; and / or apply diffuse field spectral shaping to the FDN responses; 8. Simple parametric models are implemented (e.g., by means of the control subsystem 209 in Figure 3) to control essential frequency-dependent attributes of late reverberation, such as reverberation decay time, interaural coherence, and / or direct relationship to late. In some implementations (e.g., for applications where system latency is critical and the delay caused by the analysis and synthesis filter banks is prohibitive), the filter bank domain FDN structures of typical implementations of the inventive system (e.g., the FDN in Figure 4 in each frequency band) are replaced by time-domain FDN structures (e.g., the FDN 220 in Figure 10, which can be implemented as shown in Figure 9). In the time-domain implementations of the inventive system, the subsystems of the filter bank domain implementations that apply an input gain factor (Gentrada), the reverb tank gains (gi), and the normalization gains (1 / |gi|) are replaced by time-domain filters (and / or gain elements) to allow frequency-dependent controls.The downmix matrix of a filter bank-domain implementation (e.g., the downmix matrix 312 in Figure 4) is replaced (in typical time-domain implementations) by an output set of time-domain filters (e.g., elements 500–503 of the implementation in Figure 11 from element 424 in Figure 9). Unlike the other filters in typical time-domain implementations, the phase response of this output set of filters is usually critical (since energy conservation and interaural coherence could be affected by the phase response).In some time-domain implementations, the reverb tank delays vary (e.g., vary slightly) from their values in a corresponding filter bank-domain implementation (e.g., to avoid sharing the filter bank step as a common factor). Figure 10 is a block diagram of an implementation of the inventive headphone virtualization system similar to that in Figure 3, except that elements 202–207 of the system in Figure 3 have been replaced in the system of Figure 10 by a single FDN 220 implemented in the time domain (e.g., the FDN 220 of Figure 10 can be implemented as the FDN of Figure 9). In Figure 10, two signals (left and right channels) are output in the time domain from the direct response and early reflection subsystem 100, and two signals (left and right channels) are output in the time domain from the late reverb processing subsystem 221. The addition element 210 is coupled to the outputs of subsystems 100 and 200.Element 210 is configured to combine (mix) the left channel outputs of subsystems 100 and 221 to generate the left channel, L, of the binaural audio signal output of the virtualizer in Figure 10, and to combine (mix) the right channel outputs of subsystems 100 and 221 to generate the right channel, R, of the binaural audio signal output of the virtualizer in Figure 10. Element 210 can be implemented to simply sum the corresponding left channel samples emitted from subsystems 100 and 221 to generate the left channel of the binaural output signal, and to simply sum the corresponding right channel samples emitted from subsystems 100 and 221 to generate the right channel of the binaural output signal, assuming that appropriate level adjustments and time alignments are implemented in subsystems 100 and 221. In the system of Figure 10, the multi-channel audio input signal (which has the channels, Xi) is directed to, and undergoes processing in, two parallel processing paths: one through subsystem 100 for processing the direct response and early reflection; the other through subsystem 221 for processing the late reverb. The system in Figure 10 is configured to apply a BRIR to each channel Xi. Each BRIRi can be decomposed into two parts: a direct response and early reflection portion (applied by subsystem 100), and a late reverb portion (applied by subsystem 221).In operation, the direct response and early reflection subsystem 100 generates the direct response and early reflection portions of the binaural audio signal output from the virtualizer, and the late reverb processing subsystem 221 ("late reverb generator") generates the late reverb portion of the binaural audio signal output from the virtualizer. The outputs of subsystems 100 and 221 are mixed (via subsystem 210) to generate the binaural audio signal, which is typically imposed from subsystem 210 to a processing system (not shown) where it undergoes binaural processing for playback through the headphones. The downmix subsystem 201 (of the reverb processing subsystem 221) is configured to downmix the channels of the multi-channel input signal into a mono downmix (which is a time-domain signal), and the FDN 220 is configured to apply the late reverb portion to the mono downmix. With reference to Figure 9, we will now describe an example of a time-domain FDN that can be used as the FDN 220 of the virtualizer in Figure 10. The FDN in Figure 9 includes the input filter 400, which is coupled to receive a mono downmix (for example, generated by subsystem 201 of the system in Figure 10) from all channels of a multi-channel audio input signal. The FDN in Figure 9 also includes an all-pass filter (APF) 401 (corresponding to APF 301 in Figure 4) coupled to the output of filter 400, input gain element 401A coupled to the output of filter 401, addition elements 402, 403, 404, and 405 (corresponding to elements 302, 303, 304, and 305 in Figure 4) coupled to the output of element 401A, and four reverberation tanks.Each reverb tank is coupled to the output of one different element 402, 403, 404, and 405, and comprises one of the reverb filters 406, and 406A, 407 and 407A, 408 and 408A, and 409 and 409A, one of the delay lines 410, 411, 412, and 413 (corresponding to the delay lines 307 in Figure 4) coupled to it, and one of the gain elements 417, 418, 419, and 420 coupled to the output of one of the delay lines. The unit matrix 415 (corresponding to the unit matrix 308 in Figure 4, and normally implemented to be identical to matrix 308) is coupled to the outputs of delay lines 410, 411, 412, and 413. Matrix 415 is configured to impose a feedback output on a second input of each of elements 402, 403, 404, and 405. When the delay (n1) applied by line 410 is less than that applied (n2) by line 411, the delay applied by line 411 is less than that applied (n3) by line 412, and the delay applied by line 412 is less than that applied (n4) by line 413, the outputs 417 and 419 of the gain elements (from the first and third reverb banks) are imposed on the inputs of the addition element 422, and the outputs 418 and 420 of the gain elements (from the second and fourth reverb banks) are imposed on the inputs of the addition element 423. The output of element 422 is imposed on one input of the IACC filter and mix stage 424, and the output of element 423 is imposed on the other input of the IACC filter and mix stage 424. Examples of the implementations of gain elements 41-420 and elements 422, 423, and 424 in Figure 9 will be described with reference to the typical implementation of elements 310 and 311 and the output mix matrix 312 in Figure 4. The output mix matrix 312 in Figure 4 (also identified as Mout) is a 2 x 2 matrix configured to mix the unmixed binaural channels (the outputs of elements 310 and 311, respectively) from the initial distribution to generate left and right binaural output channels (the left ear, "L", and right ear, "R" signals imposed on the output of matrix 312) that have the desired interaural coherence.This initial distribution is implemented by elements 310 and 311, each of which combines two outputs of the reverb tank to generate one of the unmixed binaural channels, with the reverb tank output having the smallest delay imposed on the input of element 310 and the reverb tank output having the second smallest delay imposed on the input of element 311. Elements 422 and 423 of the implementation in Figure 9 perform the same type of initial distribution (on the time-domain signals imposed on their inputs) that elements 310 and 311 (in each frequency band) of the implementation in Figure 4 perform on the component flows in the filter bank domain (in the relevant frequency band) imposed on their inputs. The unmixed binaural channels (output from elements 310 and 311 in Figure 4, or from elements 422 and 423 in Figure 9), which are close to being uncorrelated since they are not composed of any output from the common reverb tank, can be mixed (by array 312 in Figure 4 or stage 424 in Figure 9) to implement a distribution pattern that achieves a desired interaural coherence for the left and right binaural output channels. However, since the reverb tank delays are different in each FDN (that is, the FDN in Figure 9, or the FDN implemented for each different frequency band in Figure 4), one unmixed binaural channel (the output of one of elements 310 and 311, or 422 and 423) constantly guides the other unmixed binaural channel (the output of the other of elements 310 and 311, or 422 and 423). Therefore, in the implementation of Figure 4, if the combination of the reverb tank delays and the distribution pattern is identical across all frequency bands, it would result in a skewed sound image. This skew can be mitigated if the distribution pattern is alternated across the frequency bands in such a way that the mixed binaural output channels lead and drag each other in the alternating frequency bands. For example, if the desired interaural coherence is Coh, where |Coh| ≥ 1, the 312 output downmix matrix can be implemented in the odd bands to multiply the two inputs imposed on it by a matrix of the following form: arcse sinn (Coh) / 2, and the 312 output downmixing matrix can be implemented in the even frequency bands to multiply the two inputs imposed on it by a matrix having the following form: where Ο = arcsin (Coh) / 2. Alternatively, the sound image bias indicated above in binaural output channels can be mitigated by implementing the 312 matrix to be identical in the FDN for all frequency bands, if the order of the channel inputs is switched to alternate some of the frequency bands (e.g., the output element 310 can be imposed on the first input of the 312 matrix and the output of element 311 can be imposed on the second input of the 312 matrix in the odd frequency bands, and the output of element 311 can be imposed on the first input of the 312 matrix and the output of element 310 can be imposed on the second input of the 312 matrix in the even frequency bands). In the embodiment of Figure 9 (and other time-domain embodiments of an FDN of the inventive system), it is not trivial to alternate the frequency-based distribution to encompass the sound image skew that would otherwise result when the unmixed binaural channel output of element 422 constantly leads (or delays) the unmixed binaural channel output of element 423. This sound image skew is encompassed in a typical time-domain embodiment of an FDN of the inventive system in a different way than it is normally encompassed in the filter bank-domain embodiment of an FDN of the inventive system.Specifically, in the implementation of Figure 9 (and some other time-domain implementations of an FDN of the inventive system), the relative gains of the unmixed binaural channels (e.g., the output of elements 422 and 423 in Figure 9) are determined by the gain elements (e.g., elements 417, 418, 419, and 420 in Figure 9) to compensate for the sound image skew that would otherwise result from the observed unbalanced timing. By implementing a gain element (e.g., element 417) to attenuate the earliest arriving signal (which has been distributed to one side, e.g., by element 422) and implementing a gain element (e.g., element 418) to boost the next earliest signal (which has been distributed to the other side, e.g., by element 423), the stereo image is re-centered.Therefore, the reverb tank that includes gain element 417 applies a first gain to the output of element 417, and the reverb tank that includes gain element 418 applies a second gain (different from the first gain) to the output of element 418, so that the first gain and the second gain attenuate the first unmixed binaural channel (output of element 422) relative to the second unmixed binaural channel (output of element 423). More specifically, in a typical implementation of the FDN in Figure 9, the four delay lines 410, 411, 412, and 413 have increased lengths, with values n1, n2, n3, and n4, respectively. In this implementation, filter 417 applies a gain of g1. Therefore, the output of filter 417 is a delayed version of the input to delay line 410 with the gain g1 applied. Similarly, filter 418 applies a gain of g2, filter 419 applies a gain of g3, and filter 420 applies a gain of g4. Therefore, the output of filter 418 is a delayed version of the input to delay line 411 with a gain of g2 applied, and the output of filter 419 is a delayed version of the input to delay line 412 with a gain of g3 applied, and the output of filter 420 is a delayed version of the input to delay line 413 with a gain of g4 applied. In this implementation, the choice of the following gain values may result in an undesirable bias of the output sound image (indicated by the output of the binaural channels of element 424) to one side (i.e., to the left or right channel): g1 = 0.5, g2 = 0.5, g3 = 0.5, and g4 = 0.5. According to one embodiment of the invention, the gain values g1, g2, g3, and g4 (applied by elements 417, 418, 419, and 420, respectively) are chosen as follows to center the sound image: g1 = 0.38, g2 = 0.6, g3 = 0.5, and g4 = 0.5.Therefore, the output stereo image is re-centered according to an embodiment of the invention by attenuating the earliest arriving signal (which has been distributed to one side, by means of element 422 in the example) in relation to the second earliest arriving signal (i.e., by choosing g1 < g3), and increasing the second earliest signal (which has been distributed to the other side, by means of element 423 in the example), in relation to the latest arriving signal (i.e., by choosing g4 < g2). Typical time-domain FDN implementations in Figure 9 have the following differences and similarities to the filter bank domain (CQMF domain) FDN in Figure 4: the same unitary feedback matrix, A (matrix 308 in Figure 4 and matrix 415 in Figure 9); similar reverberation tank delays, nor (that is, the delays in the CQMF implementation of Figure 4 can be n1 = 17*64Ts = 1088*Ts, n2 = 21*64Ts = 1344*Ts, n3 = 26*64Ts = 1664*Ts, and n4 = 29*64Ts = 1856*Ts, where 1 / Ts is the sampling rate (1 / Ts is normally equal to 48 KHz), whereas the delays in the time-domain implementation can be: n1 = 1089*Ts, n2 = 1345*Ts, n3 = 1663*Ts, and n4 = 185*Ts. Note that in typical CQMF implementations there is the practical limitation that each delay is an integer multiple of the duration of a block of 64 samples (the sampling rate is normally 48 KHz), but in the time domain there is no more flexibility in the choice of each delay and therefore no more flexibility in the choice of delay for each reverb tank); Similar implementations of the all-pass filter (that is, similar implementations of filter 301 in Figure 4 and filter 401 in Figure 9). For example, the all-pass filter can be implemented by cascading several (for example, 3) all-pass filters. For example, each cascaded all-pass filter might be of the form where g = 0, 6. The all-pass filter 301 of Figure 4 can be implemented by three all-pass filters cascaded with suitable sampling block delays (e.g., n1 = 64*Ts, n2 = 128*Ts, and n3 = 196*Ts), where all the all-pass filters 401 of Figure 9 (the all-pass filters in the time domain) can be implemented by three all-pass filters cascaded with similar delays (e.g., n1 = 61*Ts, n2 = 127*Ts, and n3 = 191*Ts). In some time-domain FDN implementations of Figure 9, the input filter 400 is implemented to cause the BRIR's direct-to-late relationship (DLR) to be applied by the system in Figure 9 to match (at least substantially) a target DLR, and so that the BRIR DLR to be applied by a virtualizer that includes the system in Figure 9 (for example, the virtualizer in Figure 10) can be changed by the replacement filter 400 (or a control filter 400 configuration). For example, in some embodiments, filter 400 is implemented as a cascade of filters (for example, a first filter 400A and a second filter 400B, coupled as shown in Figure 9A) to implement the target DLR and optionally also to implement control of the desired DLR.For example, the filters in the cascade are IIR filters (e.g., filter 400A is a first-order Butterworth high-pass filter (an IIR filter) configured to match the target low-frequency characteristics, and filter 400B is an IIR low-pass filter configured to match the target high-frequency characteristics). As another example, the filters in the cascade are IIR and FIR filters (e.g., filter 400A is a second-order Butterworth high-pass filter (an IIR filter) configured to match the low-frequency characteristics, and filter 400B is a 14th-order FIR filter configured to match the target high-frequency characteristics). Typically, the forward signal is fixed, and filter 400 modifies the backward signal to achieve the target DLR.The all-pass filter (APF) 401 is preferably implemented to perform the same function as the APF 301 in Figure 4, primarily to introduce phase diversity and increase echo density to generate a more natural-sounding FDN output. The APF 401 typically controls the phase response, while the input filter 400 controls the amplitude response. In Figure 9, filter 406 and gain element 406A together implement a reverb filter, filter 407 and gain element 407A together implement another reverb filter, filter 408 and gain element 408A together implement another reverb filter, and filter 409 and gain element 409A together implement yet another reverb filter. Each of the filters 406, 407, 408, and 409 in Figure 9 is preferably implemented as a filter with a maximum gain value close to one (unity gain), and each of the gain elements 406A, 407A, 408A, and 409A is configured to apply a decay gain to the output of the corresponding filter 406, 407, 408, and 409 that matches the desired decay (after the relevant reverb tank delay, ni). Specifically,Gain element 406A is configured to apply a decay gain (decaygain1) to the output of filter 406 to cause the output of element 406A to have a gain such that the output of delay line 410 (after the reverb tank delay, n1) has a first target decay gain. Gain element 407A is configured to apply a decay gain (decaygain2) to the output of filter 407 to cause the output of element 407A to have a gain such that the output of delay line 411 (after the reverb tank delay, n2) has a second target decay gain. Gain element 408A is configured to apply a decay gain (decaygain3) to the output of filter 408 to cause the output of element 408A to have a gain such that the output of line 412 delay (after the reverb tank delay,n3) has a third target decay gain and the gain element 409A is configured to apply a decay gain (decaygain4) to the output of filter 409 to cause the output of element 409A to have a gain such that the output of delay line 413 (after the reverb tank delay, n4) has a fourth target decay gain. Each of the filters 406, 407, 408, and 409, and each of the elements 406A, 407A, 408A, and 409A of the system in Figure 9 are preferably implemented (with each of the filters 406, 407, 408, and 409 preferably implemented as an IIR filter, e.g., a limiting filter or a cascade of limiting filters) to achieve a target T60 characteristic of the BRIR to be applied by a virtualizer that includes the system in Figure 9 (e.g., the virtualizer in Figure 10), where "T60" denotes the reverberation decay time (T60).For example, in some embodiments each of filters 406, 407, 408, and 409 is implemented as either a limiter filter (e.g., a limiter filter having Q = 0.3 and a cutoff frequency of 500 Hz, to achieve the T60 characteristic shown in Figure 13, where T60 has units of seconds) or as a cascade of two IIR attenuator filters (e.g., having cutoff frequencies of 100 Hz and 1000 Hz, to achieve the T60 characteristic shown in Figure 14, where T60 has units of seconds). The shape of each attenuator filter is determined to match the desired switching curve from low frequency to high frequency. When the 406 filter is implemented as an attenuator filter (or a cascade of attenuator filters), the reverb filter comprising the 406 filter and the 406A gain element is also an attenuator filter (or a cascade of attenuator filters).Similarly, when each of filters 407, 408, and 409 is implemented as an attenuator filter (or a cascade of attenuator filters), each reverb filter comprising filter 407 (or 408 or 409) and the corresponding gain element (407A, 408A, or 409A) is also an attenuator filter (or a cascade of attenuator filters). Figure 9B is an example of filter 406 implemented as a cascade of a first attenuator filter 406B and a second attenuator filter 406C, coupled as shown in Figure 9B. Each of filters 407, 408, and 409 can be implemented as the Figure 9B implementation of filter 406. In some embodiments, the decay gains (decaygaini) applied by elements 406A, 407A, 408A, and 409A are determined as follows: where i is the reverb tank index (i.e., element 406A applies decay gain 1, element 407A applies decay gain 2, and so on), ni is the delay of the i-th reverb tank (e.g., n1 is the delay applied by delay line 410). Fs is the sampling rate, T is the desired reverb decay time (T60) to a predetermined low frequency. Figure 11 is a block diagram of an implementation of the following elements of Figure 9: elements 422 and 423, and IACC (interaural cross-correlation coefficient) filtering and mixing stage 424. Element 422 is coupled and configured to sum the outputs of filters 417 and 419 (from Figure 9) and to impose the summed signal on the input of the low-frequency attenuator filter 500. Element 422 is also configured to sum the outputs of filters 418 and 420 (from Figure 9) and to impose the summed signal on the input of the high-pass filter 501. The outputs of filters 500 and 501 are summed (mixed) in element 502 to generate the binaural left-ear output signal, and the outputs of filters 500 and 501 are mixed in element 502 (the output of filter 500 is subtracted from the output of filter 501) to generate the binaural right-ear output signal. Elements 502 and 503 mix (add and subtract) the filtered outputs of filters 500 and 501 to generate binaural output signals that achieve (within acceptable accuracy) the target IACC characteristic.In the implementation of Figure 11, each of the low-frequency attenuator filter 500 and the high-pass filter 501 is typically implemented as a first-order IIR filter. In an example where filters 500 and 501 have this implementation, the implementation of Figure 11 can achieve the exemplary IACC characteristic labeled "I" in Figure 12, which is a good match to the target IACC characteristic labeled "IT" in Figure 12. Figure 11A is a graph of the frequency response (R1) of a typical implementation of the 500 filter from Figure 11, the frequency response (R2) of a typical implementation of the 501 filter from Figure 11, and the response of the 500 and 501 filters connected in parallel. It is evident from Figure 11A that the combined response is desirably flat over the 100 Hz–10,000 Hz range. Therefore, in a class of embodiments not covered by the appended claims, the invention is a system (e.g., that of Figure 10) and a method for generating a binaural signal (e.g., the output of element 210 in Figure 10) in response to an array of channels of a multi-channel audio input signal, including by applying a binaural room impulse response (BRIR) to each channel of the array, thereby generating filtered signals, including the use of a unique feedback delay network (FDN) to apply a common late reverb to a downmix of the channels of the array; and combining the filtered signals to generate the binaural signal. The FDN is implemented in the time domain. In some such embodiments, the time-domain FDN (e.g., FDN 220 of Figure 10, configured as in Figure 9) includes: an inlet filter (for example, the 400 filter in Figure 9) having a coupled inlet to receive the downmix, wherein the inlet filter is configured to generate a first filtered downmix in response to the downmix; an all-pass filter (e.g., filter 401 in Figure 9) having a coupled input to receive the downmix, wherein the input filter is configured to generate a first filtered downmix in response to the downmix; a reverb application subsystem (e.g., all elements in Figure 9 other than elements 400, 401, and 424), having a first output (e.g., the output of element 422) and a second output (e.g., the output of element 423), wherein the reverb application subsystem comprises a set of reverb tanks, each of the reverb tanks having a different delay, and wherein the reverb application subsystem is coupled and configured to generate a first unmixed binaural channel and a second unmixed binaural channel in response to the second filtered downmix, to impose the first unmixed binaural channel on the first output, and to impose the second unmixed binaural channel on the second output;and an interaural cross-correlation coefficient (IACC) filtering and mixing stage (e.g., stage 424 in Figure 9, which can be implemented as elements 500, 501, 502, and 503 in Figure 11) coupled to the reverb application subsystem and configured to generate a first unmixed binaural channel and a second unmixed binaural channel in response to the first unmixed binaural channel and the second unmixed binaural channel. The input filter can be implemented to generate (preferably as a cascade of two filters configured to generate) the first filtered downmix in such a way that each BRIR has a direct-to-late ratio (DLR) that matches, at least substantially, a target DLR. Each reverb tank can be configured to generate the delayed signal, and can include a reverb filter (e.g., implemented as an attenuator filter or a cascade of attenuator filters) coupled and configured to apply gain to a signal propagating into each of the reverb tanks, to cause the delayed signal to have a gain that matches, at least substantially, the target decay gain for that delayed signal, in an effort to achieve a target reverb decay time characteristic (e.g., the T60 characteristic) of each BRIR. In some embodiments, the first unmixed binaural channel guides the second unmixed binaural channel; the reverb tanks include a first reverb tank (e.g., the reverb tank in Figure 9 including delay line 410) configured to generate a first delayed signal having the smallest delay and a second reverb tank (e.g., the reverb tank in Figure 9 including delay line 411) configured to generate a second delayed signal having the second smallest delay, wherein the first reverb tank is configured to apply a first gain to the first delayed signal, the second reverb tank is configured to apply a second gain to the second delayed signal, the second gain being different from the first gain.The application of the first and second gains results in the attenuation of the first unmixed binaural channel relative to the second unmixed binaural channel. Normally, the first and second unmixed binaural channels are indicative of a re-centered stereo image. In some implementations, the IACC filtering and mixing stage is configured to generate the first and second unmixed binaural channels such that they have an IACC characteristic that at least substantially matches the target IACC characteristic. Aspects of the invention include procedures and systems (for example, the system 20 of Figure 2, or the system of Figure 3, or of Figure 10) that perform (or are configured to perform, or support the performance of) binaural virtualization of audio signals (for example, audio signals whose audio content consists of speaker channels, and / or object-based audio signals). In some embodiments, the inventive virtualizer is or includes a coupled general-purpose processor for receiving or generating input data indicative of the multi-channel audio signal, and programmed with software (or firmware) and / or otherwise configured (e.g., in response to control data) to perform any of a variety of operations on the input data, including an embodiment of the inventive process. Such a general-purpose processor would normally be coupled to an input device (e.g., a mouse and / or keyboard), memory, and a display device. For example, the system in Figure 3 (or system 20 in Figure 2, or the virtualizer system comprising elements 12, ...Steps 14, 15, 16, and 18 of system 20) could be implemented on a general-purpose processor, with the inputs being audio data indicative of N channels of the input audio signal, and the outputs being audio data indicative of the two channels of a binaural audio signal. A conventional digital-to-analog converter (DAC) could operate on the output data to generate the analog versions of the binaural signal channels for playback by loudspeakers (e.g., a pair of headphones). Although the specific embodiments and applications of the present invention have been described herein, it will be evident to those of ordinary skill in the art that many variations on the embodiments and applications described herein are possible without leaving the scope of the invention described and claimed herein.
Claims
1. A method for generating a binaural signal in response to an array of channels of a multi-channel audio input signal, the method comprising: applying a binaural impulse response of a room, BRIR (12, , 15), to each channel of the array, thereby generating filtered signals; and combining (16, 18) the filtered signals to generate the binaural signal, wherein applying the BRIR to each channel of the array comprises using a late reverb generator (200) to apply, in response to a reverb time imposed on the late reverb generator (200), a common late reverb portion to a downmix of the channels of the array, wherein the common late reverb portion emulates collective macro attributes of the late reverb portions of at least some of the channel BRIRs, and wherein a content-dependent power equalization factor is applied to the downmix. 2.The method according to claim 1, wherein the late reverb generator (200) comprises a bank of feedback delay networks (203, 204, 205) for applying the common late reverb portion to the downstream mix, with each feedback delay network (203, 204, 205) in the bank applying late reverb to a different frequency band of the downstream mix.
3. The method according to claim 2, wherein each of the feedback delay networks (203, 204, 205) is implemented in the complex quadrature mirror filter domain.
4. The method according to claim 1, wherein the late reverb generator (200) comprises a feedback delay network (220) for applying the common late reverb portion to the downmix of the ensemble channels, wherein the feedback delay network (220) is implemented in the time domain. 5.A system (20) for generating a binaural signal in response to an array of channels of a multi-channel audio input signal, the system (20) comprising one or more processors that: apply a binaural room impulse response, BRIR, to each channel of the array, thereby generating filtered signals; and combine the filtered signals to generate the binaural signal, wherein applying the BRIR to each channel of the array comprises using a late reverb generator (200) to apply, in response to a reverb time imposed on the late reverb generator (200), a common late reverb portion to a downmix of the array channels, wherein the common late reverb portion emulates collective macro attributes of the late reverb portions of at least some channel BRIRs, and wherein a content-dependent power equalization factor is applied to the downmix. 6.The system according to claim 5, wherein the late reverb generator (200) includes a bank of feedback delay networks (203, 204, 205) configured to apply the late reverb portion to the downstream mix, with each feedback delay network (203, 204, 205) in the bank applying late reverb to a different frequency band of the downstream mix.
7. The system according to claim 6, wherein each of the feedback delay networks (203, 204, 205) is implemented in the complex quadrature mirror filter domain. 8.The system according to claim 5, wherein the late reverb generator (200) includes a feedback delay network (220) implemented in the time domain, and the late reverb generator (200) is configured to process the downstream mix in the time domain in said feedback delay network (220) to apply the common late reverb portion to said downstream mix.
9. A computer-readable means comprising instructions that, when executed by a computer, cause the computer to carry out the procedure according to any one of claims 1 to 4.