Generation of binaural audio responsive to multi-channel audio using at least one feedback delay network

The method addresses the limitations of conventional headphone virtualizers by using a bank of FDNs to apply a common late reverberation to binaural signals, improving external head localization and sound quality.

JP7683101B2Active Publication Date: 2025-05-26DOLBY LABORATORIES LICENSING CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024130673
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2014-05-05
Filing Date
2024-08-07
Publication Date
2025-05-26
Estimated Expiration
2034-12-18

AI Technical Summary

Technical Problem

Conventional headphone virtualizers using Feedback Delay Networks (FDN) struggle to accurately simulate the microstructure of early reflections and lack flexibility in tuning, leading to limited success in external head localization and introducing excessive color distortion and reverberation.

Method used

A method for generating a binaural signal by applying a binaural room impulse response (BRIR) to each channel of a multi-channel audio input signal, using a bank of FDNs to add a common late reverberation to a downmix, and combining the filtered signals to generate a binaural signal, allowing for control of frequency-dependent spatial acoustic attributes.

Benefits of technology

This approach enhances external head localization and produces a more natural-sounding output by better matching the acoustic environment, while reducing excessive color distortion and reverberation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007683101000015
    Figure 0007683101000015
  • Figure 0007683101000016
    Figure 0007683101000016
  • Figure 0007683101000017
    Figure 0007683101000017
Patent Text Reader

Abstract

To provide a method and system for generating a binaural signal in response to channels of a multi-channel audio signal.SOLUTION: A method includes applying a binaural room impulse response (BRIR) to each channel, which includes by using at least one feedback delay network (FDN) to apply a common late reverberation to a downmix of the channels. Input signal channels are processed in a first processing path to apply to each channel a direct response and early reflection portion of a single-channel BRIR for the channel, and the downmix of the channels is processed in a second processing path including at least one FDN which applies the common late reverberation. The common late reverberation emulates collective macro attributes of late reverberation portions of at least some of the single-channel BRIRs.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority of Chinese Patent Application No. 201410178258.0, filed on April 29, 2014, US Provisional Patent Application No. 61 / 923,579, filed on January 3, 2014, and US Provisional Patent Application No. 61 / 988,617, filed on May 5, 2014. The content of each application is hereby incorporated by reference in its entirety. 1. Field of the Invention The present invention relates to a method (sometimes referred to as a headphone virtualization method) and a system for generating a binaural signal in response to a multi - channel audio input signal by applying a binaural room impulse response (BRIR) to each channel of a set of channels of an input signal (e.g., to all channels). In some embodiments, at least one feedback delay network (FDN) applies a late reverberation portion of a down - mixed BRIR to the down - mix of the channels.

Background Art

[0002] 2. Background of the Invention Headphone virtualization (or binaural rendering) is a technology aimed at delivering a surround - sound experience or an immersive sound field using standard stereo headphones.

[0003] Early headphone virtualizers applied the head-related transfer function (HRTF) to convey spatial information in binaural rendering. The HRTF is a set of direction- and distance-dependent filter pairs that characterize how sound travels from a specific point (sound source location) in space to both ears of a listener in an anechoic environment. Essential spatial cues such as the interaural time difference (ITD), the interaural level difference (ILD), the head shadowing effect, and spectral peaks and notches due to shoulder and pinna reflections can be perceived in the rendered HRTF-filtered binaural content. Due to the size constraints of the human head, the HRTF does not provide sufficient or robust cues for source distances beyond approximately one meter. As a result, virtualizers based solely on the HRTF typically do not achieve good external head localization or perceived distance.

[0004] Many of the acoustic events in daily life occur in reverberant environments. In a reverberant environment, in addition to the direct path (from source to ear) modeled by the HRTF, the audio signal reaches the listener's ears through various reflection paths. Reflections introduce profound effects on the auditory experience, such as distance, room size, and other attributes of the space. To convey this information in binaural rendering, the virtualizer needs to apply room reverberation in addition to the cues in the direct-path HRTF. The binaural room impulse response (BRIR) characterizes the transformation of an audio signal from a specific point in space to the listener's ears in a particular acoustic environment. In theory, the BRIR contains all the acoustic cues related to spatial perception.

[0005] FIG. 1 shows each full-frequency range channel (X 1 , …, X N) is a block diagram of one type of conventional headphone virtualizer configured to apply a binaural room impulse response (BRIR). Channel X 1 , …, X N each corresponds to a speaker channel corresponding to a different source direction for an assumed listener (i.e., the direction of the direct path from the assumed position of the corresponding speaker to the assumed listener position), and each such channel is convolved by the BRIR for the corresponding source direction. The acoustic path from each channel needs to be simulated for each ear. Thus, for the remainder of this paper, the term BRIR refers to either a single impulse response or a pair of impulse responses associated with the left and right ears. Thus, subsystem 2 is configured to convolve channel X 1 with BRIR 1 (the BRIR for the corresponding source direction), and subsystem 4 is configured to convolve channel X N with BRIR N (the BRIR for the corresponding source direction), and so on. The output of each BRIR subsystem (each of subsystems 2, …, 4) is a time-domain signal including a left channel and a right channel. The left-channel outputs of the BRIR subsystems are mixed in adder element 6, and the right-channel outputs of the BRIR subsystems are mixed in adder element 8. The output of element 6 is the left channel L of the binaural audio signal output from the virtualizer, and the output of element 8 is the right channel R of the binaural audio signal output from the virtualizer.

[0006] The multi-channel audio input signal may also include a low frequency effects (LFE) or subwoofer channel, which is identified as the "LFE" channel in FIG. 1. In the normal way, the LFE channel is not convolved with the BRIR. Instead, it is attenuated (e.g., by more than -3 dB) at the gain stage 5 in FIG. 1, and the output of the gain stage 5 is equally (by the adding elements 6 and 8) mixed into each channel of the binaural output signal of the virtualizer. An additional delay stage may be required in the LFE path to time-align the output of stage 5 with the output of the BRIR subsystem (2, …, 4). Alternatively, the LFE channel may simply be ignored (i.e., not presented to or processed by the virtualizer). For example, the embodiment of FIG. 2 of the present invention (described later) simply ignores any LFE channel of the multi-channel audio input signal it processes. Many consumer headphones are not able to accurately reproduce the LFE channel.

[0007] In some conventional virtualizers, the input signal undergoes a conversion from the time domain to the frequency domain to be in the QMF (quadrature mirror filter) domain, generating channels of QMF domain frequency components. These frequency components are filtered in the QMF domain (e.g., in the QMF domain implementation of the subsystems 2, …, 4 in FIG. 1), and the resulting frequency components are then converted back to the time domain (e.g., at the respective final stages of the subsystems 2, …, 4 in FIG. 1). Thereby, the audio output of the virtualizer is a time domain signal (e.g., a time domain binaural signal).

[0008] Generally, each full - frequency - range channel of the multi - channel audio signal input into the headphone virtualizer is assumed to represent audio content emitted from a sound source located at a known position relative to the listener's ears. The headphone virtualizer is configured to apply a binaural room impulse response (BRIR) to each such channel of the input signal. Each BRIR can be decomposed into two parts: the direct response and the reflections. The direct response is the HRTF corresponding to the direction of arrival (DOA) of the sound source, adjusted with appropriate gain and delay due to the distance (between the sound source and the listener), and optionally enhanced with a parallax effect for small distances.

[0009] The remaining part of the BRIR models the reflections. Early reflections are typically first - or second - order reflections and have a relatively sparse temporal distribution. The microstructure (e.g., ITD and ILD) of each first - or second - order reflection is important. For late reflections (sounds reflected from more than three surfaces before reaching the listener), the echo density increases with the number of reflections, and the microscopic attributes of individual reflections become difficult to observe. For increasingly later reflections, the macrostructure (e.g., reverberation decay rate, inter - aural coherence, and overall spectral distribution of reverberation) becomes more important. Therefore, the reflections can be further segmented into two parts: early reflections and late reverberation.

[0010] The delay of the direct response is the source distance divided by the speed of sound, and its level is inversely proportional to the source distance (when there are no walls or large surfaces near the source position). On the other hand, the delay and level of late reverberation are generally not sensitive to the source position. For practical reasons, the virtualizer can choose to time - align the direct responses from sources with different distances and / or compress their dynamic range. However, the temporal and level relationships between the direct response, early reflections, and late reverberation within the BRIR should be maintained.

[0011] The effective length of a typical BRIR reaches several hundred milliseconds or more in many acoustic environments. The direct application of BRIR requires convolution with a filter of thousands of taps, which is computationally expensive. In addition, without parameterization, a large memory space is required to store the BRIRs for different source positions to achieve sufficient spatial resolution. Finally, but not least, the source position can change over time and / or the position and orientation of the listener can change over time. Accurate simulation of such movements requires a time-varying BRIR impulse response. Proper interpolation and application of such time-varying filters can be difficult when the impulse responses of these filters have many taps.

[0012] To implement a spatial reverberator configured to apply simulated reverberation to one or more channels of a multi-channel audio input signal, a filter having a well-known filter structure known as a feedback delay network (FDN) can be used. The structure of the FDN is simple. It has several reverberation tanks (for example, the reverberation tanks having gain element g 1 and delay line z -n1 in the FDN of FIG. 4), and each reverberation tank has a delay and a gain. In a typical implementation of the FDN, the outputs from all the reverberation tanks are mixed by a unitary feedback matrix, and the output of the matrix is fed back and summed with the input of the reverberation tank. Gain adjustment may be made to the reverberation tank outputs. The reverberation tank outputs (or their gain-adjusted versions) can be suitably remixed for multi-channel or binaural playback. An FDN with a compact computation and memory footprint can generate and apply a natural-sounding reverberation. Therefore, the FDN has been used in virtualizers to supplement the direct response generated by the HRTF.

[0013] For example, a commercially available "Dolby Mobile" headphone virtualizer includes a reverberator having an FDN-based structure that is operable to add reverberation to each channel of a five-channel audio signal (having left front, right front, center, left surround and right surround channels), and filter each reverberation-added channel using a different filter pair of a set of five head-related transfer function ("HRTF") filter pairs. The "Dolby Mobile" headphone virtualizer is also operable to generate a two-channel "reverberation-added" binaural audio output (a two-channel virtual surround sound output with added reverberation) in response to a two-channel audio input signal. When the reverberation-added binaural output is rendered and played back by a headphone pair, it is perceived at the listener's eardrums as HRTF-filtered reverberation-added sound from five loudspeakers located at left front, right front, center, left rear (surround) and right rear (surround) positions. The virtualizer upmixes a downmixed two-channel audio input (without using any spatial cue parameters received with the audio input) to generate five upmixed audio channels, adds reverberation to the upmixed channels, and downmixes the five reverberation-added channel signals to generate the two-channel reverberation-added output of the virtualizer. The reverberation for each upmixed channel is filtered in a different pair of HRTF filters. Summary of the Invention Problems to be Solved by the Invention

[0014] In a virtualizer, the FDN is configured to achieve a certain reverberation decay time and echo density. However, the FDN lacks the flexibility to simulate the microstructure of early reflections. Furthermore, in a conventional virtualizer, most of the tuning and configuration settings of the FDN are trial and error.

[0015] Headphone virtualizers that do not simulate all reflection paths (early and late) cannot achieve effective external head localization. The inventors have come to recognize that virtualizers using an FDN that attempts to simulate all reflection paths (early and late) typically achieve only limited success in simulating both early reflections and late reverberations and adding both to the audio signal. The inventors have also recognized that virtualizers that use an FDN but do not have the ability to properly control spatial acoustic attributes such as reverberation decay time, interaural coherence, and direct-to-late ratio may achieve some degree of external head localization but at the cost of introducing excessive color distortion and reverberation.

Means for Solving the Problems

[0016] In a first class of embodiments, the present invention is a method for generating a binaural signal in response to a set of channels of a multi-channel audio input signal (e.g., each of those channels or each of the full frequency range channels). The method includes: (a) applying a binaural room impulse response (BRIR) to each channel of the set (e.g., by convolving each channel of the set with the BRIR corresponding to the channel), thereby generating a filtered signal, including by using at least one feedback delay network (FDN) to add a common late reverberation to a downmix (e.g., a monophonic downmix) of the channels of the set; and (b) combining the filtered signals to generate a binaural signal. Typically, a bank of FDNs is used to add the common late reverberation to the downmix (e.g., each FDN adds a common late reverberation to a different frequency band). Typically, step (a) includes applying to each channel of the set the "direct response and early reflections" portion of a single-channel BRIR for that channel, and the common late reverberation is generated to emulate the collective macro attributes of the late reverberation portion of at least a portion (e.g., all) of the single-channel BRIRs.

[0017] A method of generating a binaural signal in response to a multi-channel audio input signal (or in response to a set of channels of such a signal) is sometimes referred to herein as a "headphone virtualization" method, and a system configured to perform such a method is sometimes referred to herein as a "headphone virtualizer" (or "headphone virtualization system" or "binaural virtualizer").

[0018] In a first class of typical implementations, each FDN is implemented in a filter bank region (e.g., a hybrid complex quadrature mirror filter (HCQMF) region or a quadrature mirror filter (QMF) region or other transform or sub-band region that may include decimation). In some such embodiments, the frequency-dependent spatial acoustic attributes of the binaural signal are controlled by controlling the configuration of each FDN used to add late reverberation. Typically, for efficient binaural rendering of the audio content of a multi-channel signal, a channel monophonic downmix is used as the input to the FDN. A first class of typical embodiments includes adjusting the FDN coefficients corresponding to frequency-dependent attributes (e.g., reverberation decay time, interaural coherence, mode density, and direct-to-late ratio) by presenting control values to a feedback delay network to set at least one of, for example, the input gain, reverberation tank gain, reverberation tank delay, or output matrix parameter of each FDN. This enables better matching of the acoustic environment and a more natural-sounding output.

[0019] In a second class of embodiments, the invention is a method of generating a binaural signal in response to a multi-channel audio input signal having a plurality of channels. This is by applying a binaural room impulse response (BRIR) to each channel of a selected set of channels of the input signal (e.g., each of the channels of the input signal or each of the input signal's full frequency range channels). This includes processing each channel of the set in a first processing path configured to model the direct response and early reflections of a single channel BRIR for that channel and apply it to that channel, and processing a downmix of the channels of the set (e.g., a monophonic (mono) downmix) in a second processing path (parallel to the first processing path) configured to model and apply a common late reverberation to the downmix. Typically, the common late reverberation is generated to emulate the collective macro attributes of the late reverberation portions of at least some (e.g., all) of the single channel BRIRs. Typically, the second processing path includes at least one FDN (e.g., one FDN for each of a plurality of frequency bands). Typically, the mono downmix is used as the input to all of the reverberation tanks of each FDN implemented by the second processing path. Typically, a mechanism for systematic control of the macro attributes of each FDN is provided to better simulate the acoustic environment and result in a more natural sounding binaural virtualization. Since most such macro attributes are frequency dependent, each FDN is typically implemented in a hybrid complex quadrature mirror filter (HCQMF) domain, frequency domain, domain or another filter bank domain, and different or independent FDNs are used for each frequency band. The main benefit of implementing the FDN in a filter bank domain is that it allows the application of reverberation with frequency dependent reverberation attributes. In various embodiments, the FDN is implemented in any of a wide variety of filter bank domains using any of a variety of filter banks.It includes, but is not limited to, real or complex-valued quadrature mirror filters (QMFs), finite impulse response filters (FIR filters), infinite impulse response filters (IIR filters), discrete Fourier transforms (DFTs), (modified) cosine or sine transforms, wavelet transforms, or crossover filters. In certain preferred implementations, the filter bank or transform used includes decimation (e.g., a reduction in the sampling rate of the frequency-domain signal representation) to reduce the computational complexity of the FDN process.

[0020] Some embodiments of the first class (and the second class) implement one or more of the following features.

[0021] 1. An FDN implementation in the filter bank region (e.g., the hybrid complex quadrature mirror filter region) or a hybrid filter bank region FDN implementation and a time-domain late reverberation filter implementation. This typically allows for independent adjustment of the FDN parameters and / or settings for each frequency band (which enables simple and flexible control of frequency-dependent acoustic attributes). This provides, for example, the ability to vary the reverberation tank delay in different bands so as to vary the mode density as a function of frequency.

[0022] 2. The particular downmixing process used to generate a downmixed (e.g., monophonic downmixed) signal that is processed in a second processing path (from the multi-channel input audio signal) depends on the handling of the direct response to maintain the appropriate level and timing relationship between the source distance of each channel and the direct and late responses.

[0023] 3. An all-pass filter (APF) is applied in the second processing path (e.g., at the input or output of a bank of FDNs) to introduce phase diversity and increased echo density without changing the resulting reverberation spectrum and / or timbre.

[0024] 4. To overcome problems related to quantization of the delay in the downsample - factor grid, fractional delay is implemented in the feedback path of each FDN in the complex - valued multirate structure.

[0025] 5. In the FDN, the reverberation tank output is linearly mixed directly into the binaural channels using output mixing coefficients set based on the desired inter - aural coherence in each frequency band. Optionally, the mapping of the reverberation tank to the binaural output channels alternates across frequency bands to achieve an equalized delay between the binaural channels. Also optionally, a normalization factor is applied to the reverberation tank output to equalize its level while preserving fractional delay and overall power.

[0026] 6. The frequency - dependent reverberation decay time and / or mode density is controlled by setting an appropriate combination of reverberation tank delay and gain in each frequency band to simulate an actual room.

[0027] 7. One scaling factor is applied for each frequency band (e.g., either at the input or output of the associated processing path). This is to: Control the frequency - dependent direct - to - late ratio (DLR) that matches the DLR of the actual room (a simple model may be used to calculate the required scaling factor based on the target DLR and reverberation decay time, e.g., T60); Provide low - frequency attenuation to mitigate excessive combing artifacts and / or low - frequency rumble; and / or Apply diffuse field spectral shaping to the FDN response.

[0028] 8. A simple parametric model is implemented to control the essential frequency-dependent attributes of late reverberation, such as reverberation decay time, interaural coherence, and / or direct-to-late ratio.

[0029] Aspects of the present invention include methods and systems for performing (or configured to perform or supporting the performance of) binaural virtualization of an audio signal (e.g., an audio signal consisting of audio content from speaker channels and / or an object-based audio signal).

[0030] In another class of embodiments, the present invention is a method and system for generating a binaural signal in response to a selected set of channels of a multi-channel audio input signal. This includes applying a binaural room impulse response (BRIR) to each channel of the selected set, thereby generating a filtered signal, including by using a single feedback delay network (FDN) to add a common late reverberation to the downmix of the channels of the selected set; and combining the filtered signals to generate a binaural signal. The FDN is implemented in the time domain. In some such embodiments, the time domain FDN includes: An input filter having an input coupled to receive the downmix, the input filter being configured to generate a first filtered downmix in response to the downmix; An all-pass filter coupled and configured to perform a second filtering of the downmix in response to the first filtered downmix; A reverberation application subsystem having a first output and a second output, the reverberation application subsystem including a set of reverberation tanks, each reverberation tank having a different delay, the reverberation application subsystem generating a first unmixed binaural channel and a second unmixed binaural channel in response to the second filtered downmix, presenting the first unmixed binaural channel at the first output, and presenting the second unmixed binaural channel at the second output, coupled and configured; Including an interaural cross-correlation coefficient (IACC) filtering and mixing stage coupled to the reverberation application subsystem and configured to generate a first mixed binaural channel and a second mixed binaural channel in response to the first unmixed binaural channel and the second unmixed binaural channel.

[0031] The input filter may be implemented (preferably as a cascade of two filters configured to generate it) to generate the first filtered downmix such that each BRIR has a direct-to-reverberant ratio (DLR) that at least substantially matches the target DLR.

[0032] Each reverberation tank may be configured to generate a delayed signal, and a gain is added to the signal propagating in each reverberation tank such that the delayed signal has a gain that at least substantially matches the target delayed gain. It may include a reverberation filter (e.g., implemented as a shelf filter or a cascade of shelf filters). This is to achieve the target reverberation decay time characteristic (e.g., T 60 Characteristic) of each BRIR.

[0033] In some embodiments, the first unmixed binaural channel is ahead of the second unmixed binaural channel, and the reverberation tank includes a first reverberation tank configured to generate a first delayed signal having the shortest delay and a second reverberation tank configured to generate a second delayed signal having the second shortest delay. The first reverberation tank is configured to apply a first gain to the first delayed signal, the second reverberation tank is configured to apply a second gain to the second delayed signal, the second gain is different from the first gain, and the application of the first gain and the second gain results in attenuation of the first unmixed binaural channel relative to the second unmixed binaural channel. Typically, the first mixed binaural channel and the second mixed binaural channel exhibit a re-centered stereo image. In some embodiments, the IACC filtering and mixing stage is configured to generate the first mixed binaural channel and the second mixed binaural channel such that the first mixed binaural channel and the second mixed binaural channel have IACC characteristics that at least substantially match the target IACC characteristics.

[0034] Exemplary embodiments of the present invention provide a simple and unified framework for supporting both input audio consisting of speaker channels and object-based input audio. In embodiments where BRIR is applied to an input signal channel that is an object channel, the "direct response and early reflection" processing performed on each object channel assumes a source direction indicated by metadata provided with the audio content of that object channel. In embodiments where BRIR is applied to an input signal channel that is a speaker channel, the "direct response and early reflection" processing performed on each speaker channel assumes a source direction corresponding to that speaker channel (i.e., the direction of the direct path from the assumed position of the corresponding speaker to the assumed listener position). Regardless of whether the input channel is an object channel or a speaker channel, the "late reverberation" processing is performed on a downmix of the input channel (e.g., a monophonic downmix) and does not assume any specific source direction for the audio content of the downmix.

[0035] Another aspect of the present invention is a headphone virtualizer configured (e.g., programmed) to execute any embodiment of the method of the present invention, a system (e.g., a stereo, multi-channel, or other decoder) including such a virtualizer, and a computer-readable medium (e.g., a disk) storing code for implementing any embodiment of the method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0036]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 9A

Figure 9B

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

DETAILED DESCRIPTION OF THE INVENTION

[0037] 〈Notation and Nomenclature〉 Throughout this disclosure, including the claims, the expression of performing an operation "on" a signal or data (e.g., filtering, scaling, transforming, or applying a gain to the signal or data) is used in a broad sense to represent performing the operation directly on the signal or data or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or preprocessing prior to the execution of the operation).

[0038] Throughout the present disclosure, including the claims, the term "system" is used in a broad sense to represent an apparatus, a system, or a subsystem. For example, a subsystem that implements a virtualizer may be referred to as a virtualizer system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to a plurality of inputs, where the subsystem generates M of the inputs and the other X - M inputs are received from an external source) may also be referred to as a virtualizer system (or a virtualizer).

[0039] Throughout the present disclosure, including the claims, the term "processor" is used in a broad sense to represent a system or an apparatus that is programmable or otherwise configurable to perform operations on data (e.g., audio or video or other image data), for example, using software or firmware. Examples of processors include field - programmable gate arrays (or other configurable integrated circuits or chip sets), digital signal processors programmed and / or otherwise configured to perform pipelined processing on audio or other sound data, programmable general - purpose processors or computers, and programmable microprocessor chips or chip sets.

[0040] Throughout the present disclosure, including the claims, the expression "analysis filter bank" is used in a broad sense to represent a system (e.g., a subsystem) configured to apply a transformation (e.g., a transformation from the time domain to the frequency domain) to a time domain signal to generate values (e.g., frequency components) indicative of the content of the time domain signal in each of a set of frequency bands. Throughout the present disclosure, including the claims, the expression "filter bank domain" is used in a broad sense to represent the domain of frequency components generated by a transformation or analysis filter bank (e.g., the domain in which such frequency components are processed). Examples of filter bank domains include (but are not limited to) the frequency domain, the quadrature mirror filter (QMF) domain, and the hybrid complex quadrature mirror filter (HCQMF) domain. Examples of transformations that may be applied by an analysis filter bank include (but are not limited to) the discrete cosine transform (DCT), the modified discrete cosine transform (MDCT), the discrete Fourier transform (DFT), and the wavelet transform. Examples of analysis filter banks include (but are not limited to) quadrature mirror filters (QMFs), finite impulse response filters (FIR filters), infinite impulse response filters (IIR filters), crossover filters, and filters having other suitable multirate structures.

[0041] Throughout the present disclosure, including the claims, the term "metadata" refers to data that is distinct from the corresponding audio data (the audio content of a bitstream that also includes the metadata). Metadata is associated with the audio data and indicates at least one characteristic or property of the audio data (e.g., what type(s) of processing, if any, has already been performed on the audio data, or should be performed, or the trajectory of an object represented by the audio data). The association of the metadata with the audio data is time synchronous. Thus, the current (most recently received or updated) metadata may indicate that the corresponding audio data includes the results of audio data processing having the characteristics and / or of the type indicated.

[0042] Throughout the present disclosure including the claims, the terms "coupled" or "coupling" are used to mean a direct or indirect connection. Thus, when a first device is coupled to a second device, the connection may be through a direct connection or through an indirect connection via other devices and connections.

[0043] Throughout the present disclosure including the claims, the following expressions have the following definitions.

[0044] Speaker and loudspeaker are used synonymously to represent any transducer that emits sound. This definition includes loudspeakers implemented as multiple transducers (e.g., woofers and tweeters).

[0045] Speaker feed: an audio signal applied directly to a loudspeaker or an audio signal applied to a series of amplifiers and a loudspeaker.

[0046] Channel (or "audio channel"): a monophonic audio signal. Such a signal can typically be rendered to be equivalent to applying the signal directly to a loudspeaker at a desired or nominal location. The desired location may be static, as is typically the case with physical loudspeakers, or it may be dynamic.

[0047] Audio program: a collection of one or more audio channels (at least one speaker channel and / or at least one object channel) and optionally associated metadata (e.g., metadata describing a desired spatial audio presentation).

[0048] Speaker Channel (or "Speaker Feed Channel"): An audio channel associated with a specified loudspeaker (at a desired or nominal location) or associated with a specified speaker zone within a defined speaker layout. The speaker channel is rendered such that it is equivalent to directly applying the audio signal to a specified loudspeaker (at a desired or nominal location) or to speakers within a specified speaker zone.

[0049] Object Channel: An audio channel indicating the sound emitted by an audio source (sometimes referred to as an audio "object"). Typically, the object channel determines a parametric audio source description (e.g., metadata indicating the parametric audio source description is included within or provided with the object channel). The source description may determine the sound emitted by the source (as a function of time), the apparent position of the source (e.g., 3D spatial coordinates) as a function of time, and optionally at least one additional parameter characterizing the source (e.g., apparent source size or width).

[0050] Object - based Audio Program: An audio program that includes a collection of one or more object channels (and optionally also at least one speaker channel) and optionally also associated metadata (e.g., metadata indicating the trajectory of the audio object emitting the sound indicated by the object channel, or metadata indicating the desired spatial audio presentation of the sound indicated by the object channel in some other way or metadata indicating the identification information of at least one audio object that is the source of the sound indicated by the object channel).

[0051] Rendering: The process of converting an audio program into one or more speaker feeds, or the process of converting an audio program into one or more speaker feeds and converting those speaker feeds into sound using one or more loudspeakers. (In the latter case, rendering is sometimes referred to in this document as rendering "by" the loudspeakers.) An audio channel can be trivially rendered (at the desired location) by applying the signal directly to a physical loudspeaker at the desired location. Alternatively, one or more audio channels can be rendered using one of a variety of virtualization techniques designed to be substantially equivalent to such trivial rendering (for the listener). In this latter case, each audio channel may generally be converted into one or more speaker feeds to be applied to a loudspeaker (single or plural) at a known location different from the desired location, such that the sound emitted by the loudspeaker in response to the feed is perceived to be emanating from the desired location. Examples of such virtualization techniques include binaural rendering via headphones (e.g., using "Dolby Headphone" processing to simulate up to 7.1 channels of surround sound for a headphone wearer) and wave field synthesis.

[0052] The notation in this document that a multi-channel audio signal is an "x.y" or "x.y.z" channel signal means that the signal has "x" full-frequency speaker channels (corresponding to speakers nominally located in the horizontal plane of the assumed listener's ears), "y" LFE (or subwoofer) channels, and optionally also "z" full-frequency overhead speaker channels (corresponding to speakers located above the head of the assumed listener, e.g., on or near the ceiling of a room).

[0053] The expression "IACC" represents, in this document, the interaural cross-correlation coefficient in its normal sense. This is an indicator of the difference in the arrival times of the audio signals at the listener's ears and is typically represented by a number within a range from a first value indicating that the arriving signals are exactly out of phase with equal magnitude, through intermediate values indicating that the arriving signals have no similarity, to a maximum value indicating the same arriving signals with the same amplitude and phase.

[0054] 〈Detailed Description of the Preferred Embodiment〉 Many embodiments of the present invention are technically possible. How to implement them from this disclosure will be clear to those skilled in the art. Embodiments of the system and method of the present invention will be described with reference to FIGS. 2 to 14.

[0055] FIG. 2 is a block diagram of a system (20) including an embodiment of the headphone virtualization system of the present invention. This headphone virtualization system (sometimes referred to as a virtualizer) is configured to apply a binaural room impulse response (BRIR) to N full-frequency range channels (X 1 , …, X N ) of a multi-channel audio input signal. Each of the channels X 1 , …, X N (which can be speaker channels or object channels) corresponds to a specific source direction and distance with respect to the assumed listener, and the system of FIG. 2 is configured to convolve each such channel with the BRIR for the corresponding source direction and distance.

[0056] System 20 is coupled to receive an encoded audio program, and then the N full-frequency range channels (X 1 , …, X NThe decoder may include a subsystem (not shown in FIG. 2) that is coupled and configured to decode the program, including by restoring ), and provide them to elements 12, …, 14, 15 of a virtualization system (having elements 12, … 14, 15, 16, 18 coupled as shown in the figure). The decoder may include additional subsystems, some of which may perform functions not related to the virtualization functions executed by the virtualization system, and some of which may perform functions related to the virtualization functions. For example, the latter functions may include extracting metadata from the encoded program and providing the metadata to a virtualization control subsystem that uses the metadata to control elements of the virtualizer system.

[0057] Subsystem 12 (along with subsystem 15) is for channel X 1 to be convolved with a BRIR 1 (BRIR for the corresponding source direction and distance), and subsystem 14 (along with subsystem 15) is for channel X N to be convolved with a BRIR N (BRIR for the corresponding source direction), and the same applies to each of the N - 2 other BRIR subsystems. The output of each of subsystems 12, …, 14, 15 is a time - domain signal including a left channel and a right channel. Addition elements 16 and 18 are coupled to the outputs of elements 12, …, 14, 15. Addition element 16 is configured to combine (mix) the left - channel outputs of the BRIR subsystems, and addition element 18 is configured to combine (mix) the right - channel outputs of the BRIR subsystems. The output of element 16 is the left channel L of the binaural audio signal output from the virtualizer of FIG. 2, and the output of element 18 is the right channel R of the binaural audio signal output from the virtualizer of FIG. 2.

[0058] An important feature of a typical embodiment of the present invention becomes apparent from comparing the embodiment of FIG. 2 of the headphone virtualizer of the present invention with the normal headphone virtualizer of FIG. 1. For purposes of comparison, the systems of FIGS. 1 and 2 are assumed to be configured such that when the same multi-channel audio input signal is presented to each of them, those systems have the same direct response and early reflection portions (i.e., the associated EBRIR of FIG. 2 i ) of the BRIR i are applied to each full frequency range channel X i of the input signal (although not necessarily with the same degree of success). Each BRIR i applied by the system of FIG. 1 or FIG. 2 can be decomposed into two parts: a direct response and early reflection portion (e.g., the EBRIR 1 applied by subsystems 12-14 of FIG. 2 N , …, one of the EBRIR i and a late reverberation portion. The embodiment of FIG. 2 (and other typical embodiments of the present invention) assumes that the late reverberation portions of a plurality of single channel BRIRs, i.e., BRIR

[0059] can cross the source direction and thus be shared across all channels, and the same late reverberation (i.e., common late reverberation) can be applied to the downmix of all full frequency range channels of the input signal. This downmix can be a monophonic (mono) downmix of all input channels, but alternatively can be a stereo or multi-channel downmix obtained from the input channels (e.g., from a subset of the input channels).

[0059] More specifically, subsystem 12 of FIG. 2 is configured to convolve input signal channel X 1 with EBRIR 1 (the direct response and early reflection BRIR portion for the corresponding source direction), and subsystem 14 is configured to convolve input signal channel X N with EBRIR N(Convolved with the direct response and early reflection BRIR portion for the corresponding source direction), and so on. The late reverberation subsystem 15 of FIG. 2 generates a mono-downmix of all full-frequency range channels of the input signal and is configured to convolve the downmix with an LBRIR (common late reverberation for all of the channels to be downmixed). The output of each BRIR subsystem (each of subsystems 12, …, 14, 15) of the virtualizer of FIG. 2 includes a left channel and a right channel (of the corresponding speaker channel or the binaural signal generated from the downmix). The left channel outputs of those BRIR subsystems are combined (mixed) in an adding element 16, and the right channel outputs of those BRIR subsystems are combined (mixed) in an adding element 18.

[0060] Assuming that appropriate level adjustment and time alignment are implemented in subsystems 12, …, 14, 15, the adding element 16 can be implemented to simply sum the corresponding left binaural-channel samples (left channel outputs of subsystems 12, …, 14, 15) to generate the left channel of the binaural output signal. Similarly, assuming that appropriate level adjustment and time alignment are also implemented in subsystems 12, …, 14, 15, the adding element 18 can also be implemented to simply sum the corresponding right binaural-channel samples (right channel outputs of subsystems 12, …, 14, 15) to generate the right channel of the binaural output signal.

[0061] The subsystem 15 of FIG. 2 can be implemented in any of a variety of ways, but typically includes at least one feedback delay network configured to add a common late reverberation to a monophonic downmix of the input signal channels presented thereto. Typically, each of subsystems 12, …, 14 has a direct response and early reflection portion (EBRIR i ) of a single-channel BRIR for the channel (X i) When applying, the common reverberation is generated to emulate the collective macro attributes of the late reverberation parts of at least some (e.g., all) of those single-channel BRIRs (wherein their "direct response and early reflection parts" are applied by subsystems 12, …, 14). For example, a certain implementation of subsystem 15 has the same structure as subsystem 200 in FIG. 3, including a bank of feedback delay networks (203, 204, …, 205) configured to apply a common reverberation to the monophonic downmix of the input signal channels presented thereto.

[0062] Similarly, subsystems 12, …, 14 in FIG. 2 can be implemented in any of various ways (in the time domain or filter bank domain), and the preferred implementation for any particular application depends on various circumstances such as (for example) performance, computation, and memory. In one exemplary implementation, each of subsystems 12, …, 14 is configured to convolve the channels presented thereto with FIR filters corresponding to the direct and early responses associated with those channels. The gains and delays are set appropriately so that the outputs of subsystems 12, …, 14 can be simply and efficiently combined with the output of subsystem 15.

[0063] Figure 3 is a block diagram of another embodiment of the headphone virtualization system of the present invention. The embodiment of FIG. 3 is similar to the embodiment of FIG. 2, and two (left and right channel) time domain signals are output from the direct response and early reflection processing subsystem 100, and two (left and right channel) time domain signals are output from the late reverberation processing subsystem 200. An adding element 210 is coupled to the outputs of subsystems 100 and 200. Element 210 combines (mixes) the left channel outputs of subsystems 100 and 200 to generate the left channel L of the binaural audio signal output from the virtualizer of FIG. 3, and combines (mixes) the right channel outputs of subsystems 100 and 200 to generate the right channel R of the binaural audio signal output from the virtualizer of FIG. 3. Assuming that appropriate level adjustment and time alignment are implemented in subsystems 100 and 200, element 210 can be implemented to simply sum the corresponding left channel samples output from subsystems 100 and 200 to generate the left channel of the binaural output signal, and simply sum the corresponding right channel samples output from subsystems 100 and 200 to generate the right channel of the binaural output signal.

[0064] In the system of FIG. 3, channel X of the multi-channel audio input signal i is directed to two parallel processing paths where it is processed. One passes through the direct response and early reflection processing subsystem 100, and the other passes through the late reverberation processing subsystem 200. The system of FIG. 3 i is configured to apply a BRIR i to each channel X. Each BRIR iIt can be decomposed into two parts: a direct response and early reflection part (applied by subsystem 100) and a late reverberation part (applied by subsystem 200). In operation, the direct response and early reflection processing subsystem 100 thus generates the direct response and early reflection part of the binaural audio signal output from the virtualizer, and the late reverberation processing subsystem (the "late reverberation generator") 200 thus generates the late reverberation part of the binaural audio signal output from the virtualizer. The outputs of subsystems 100 and 200 are mixed (by adder subsystem 210) to generate a binaural audio signal, which is typically presented from subsystem 210 to a rendering system (not shown) where it undergoes binaural rendering for playback through headphones.

[0065] Typically, when rendered and played back by a pair of headphones, a typical binaural audio signal output from element 210 is perceived at the listener's eardrums as sound from "N" loudspeakers located anywhere in a wide variety of positions, including positions in front of, behind, and above the listener (where N ≥ 2 and N is typically 2, 5, or 7). The playback of the output signal generated in the operation of the system of FIG. 3 can give the listener the experience of sound coming from more than two (e.g., five or seven) "surround" sources. At least some of these sources are virtual.

[0066] The direct response and early reflection processing subsystem 100 can be implemented in any of a variety of ways (in the time domain or filter bank domain), and the preferred implementation for any particular application depends on various circumstances such as (for example) performance, computation, and memory. In one exemplary implementation, the subsystem 100 is configured to convolve each channel presented to it with an FIR filter corresponding to the direct and early responses associated with that channel. The gains and delays are set appropriately so that the output of the subsystem 100 may be simply and efficiently combined with the output of the subsystem 200 (at element 210).

[0067] As shown in FIG. 3, the late reverberation generator 200 includes a downmixing subsystem 201, an analysis filter bank 202, a bank of FDNs (FDNs 203, 204, …, 205), and a synthesis filter bank 207 connected as shown in the figure. The subsystem 201 is configured to downmix the channels of a multi-channel input signal to a mono-downmix, and the analysis filter bank 202 is configured to apply a transform to the mono-downmix to divide the mono-downmix into “K” frequency bands. Here, K is an integer. The filter bank domain values (output from the filter bank 202) in each different frequency band are presented to different ones of the FDNs 203, 204, …, 205 (there are “K” of these FDNs, each connected and configured to apply the late reverberation portion of the BRIR to the filter bank domain value presented to it). The filter bank domain values are preferably decimated in time so as to reduce the computational complexity of the FDNs.

[0068] In principle, each input channel (to subsystems 100 and 201 of FIG. 3) can be processed by a unique FDN (or bank of FDNs) to simulate the late reverberation portion of its BRIR. Despite the fact that the late reverberation portions of BRIRs associated with different sound source positions are typically very different in terms of root mean square in the impulse response, their statistical attributes such as their average power spectrum, their energy decay structure, mode density, peak density, etc. are often very similar. Thus, the late reverberation portions of a set of BRIRs are typically perceptually very similar across channels, so it is possible to use one common FDN or bank of FDNs (e.g., FDNs 203, 204, …, 205) to simulate the late reverberation portions of two or more BRIRs. In a typical embodiment, such one common FDN (or bank of FDNs) is used, and the input thereto is composed of one or more downmixes constructed from the input channels. In the exemplary implementation of FIG. 2, the downmix is the monophonic downmix of all input channels (presented at the output of subsystem 201).

[0069] Referring to the embodiment of FIG. 2, each of FDNs 203, 204, …, 205 is implemented in a filter bank region, and is combined and configured to process different frequency bands among the values output from the decomposition filter bank 202 to generate left and right reverberated signals for each band. For each band, the left reverberated signal is a sequence of filter bank region values, and the right reverberated signal is another sequence of filter bank region values. The synthesis filter bank 207 applies the conversion from the frequency domain to the time domain to 2K sequences of filter bank region values (e.g., frequency components of the QMF region), and collects the converted values to form a left channel time domain signal (indicating the mono-downmixed audio content with late reverberation applied) and a right channel time domain signal (also indicating the mono-downmixed audio content with late reverberation applied). These left and right channel signals are output to element 210.

[0070] In a typical implementation, each of FDNs 203, 204, …, 205 is implemented in the QMF region, and the filter bank 202 converts the mono-downmix from subsystem 201 to the QMF region (e.g., the hybrid complex quadrature mirror filter (HCQMF) region), whereby the signals presented to each of the inputs of FDNs 203, 204, …, 205 from the filter bank 202 are sequences of QMF region frequency components. In such an implementation, the signal presented to FDN 203 from the filter bank 202 is a sequence of QMF region frequency components in the first frequency band, the signal presented to FDN 204 from the filter bank 202 is a sequence of QMF region frequency components in the second frequency band, and the signal presented to FDN 205 from the filter bank 202 is a sequence of QMF region frequency components in the "K"th frequency band. When the decomposition filter bank 202 is implemented in this way, the synthesis filter bank 207 applies the conversion from the QMF region to the time domain to 2K sequences of output QMF region frequency components from the FDNs, and generates left and right channel late reverberated time domain signals output to element 210.

[0071] For example, in the system of FIG. 3, if K = 3, there are six inputs to the synthesis filter bank 207 (left and right channels including the frequency domain or QMF domain samples output from each of FDNs 203, 204, and 205) and two outputs from 207 (left and right channels each consisting of time domain samples). In this example, the filter bank 207 is typically implemented as two synthesis filter banks. One (presented with the three left channels from FDNs 203, 204, and 205) is configured to generate the time domain left channel signal output from the filter bank 207, and the second one (presented with the three right channels from FDNs 203, 204, and 205) is configured to generate the time domain right channel signal output from the filter bank 207.

[0072] Optionally, the control subsystem 209 is coupled to each of FDNs 203, 204, …, 205 and is configured to present control parameters to each of those FDNs to determine the late reverberation portion (LBRIR) applied by the subsystem 200. Examples of such control parameters are described below. In some implementations, it is contemplated that the control subsystem 209 is operable in real time (i.e., in response to user commands presented thereto by the input device) to implement real-time variations of the late reverberation portion (LBRIR) applied to the monophonic downmix of the input channels by the subsystem 200.

[0073] For example, if the input signal to the system of FIG. 2 is a 5.1 channel signal (the full frequency range channels of which are in the following channel order: L, R, C, Ls, Rs), all full frequency range channels have the same source distance, and the downmix subsystem 201 can be implemented as the following downmix matrix. This simply sums the full frequency range channels to form a mono-downmix.

[0074]

Number

[0075]

Number

[0076]

Number

[0077]

Number

[0078] Next, the downmix subsystem 201 of the virtualizer in FIG. 3 and the individual implementations of subsystems 100 and 200 will be considered.

[0079] The downmix process implemented by subsystem 201 depends on the source distance (between the sound source and the assumed listener position) for each channel to be downmixed and on the handling of the direct response. The delay t d of the direct response is: t d = d / v s where d is the distance between the sound source and the listener, and v s is the speed of sound. Furthermore, the gain of the direct response is proportional to 1 / d. If these rules are preserved in the handling of the direct responses of channels with different source distances, subsystem 201 can implement a straight downmix of all channels. This is because the delay and level of the late reverberation are generally not sensitive to the source position.

[0080] For practical reasons, the virtualizer (e.g., subsystem 100 of the virtualizer in FIG. 3) may be implemented to time-align the direct responses for input channels with different source distances. To preserve the relative delay between the direct response and the late reverberation for each channel, the channel with source distance d should be delayed by (dmax - d) / v s before being downmixed with the other channels. Here, dmax represents the maximum possible source distance.

[0081] The virtualizer (e.g., subsystem 100 of the virtualizer in FIG. 3) may also be implemented to compress the dynamic range of the direct response. For example, the direct response for a channel with source distance d may be scaled by a factor d -1 instead of d -α where 0 ≦ α ≦ 1. To preserve the level difference between the direct response and the late reverberation, the downmix subsystem 201 may need to be implemented to scale the channel with source distance d by a factor d 1-α before downmixing it with the other scaled channels.

[0082] The feedback delay network of FIG. 4 is an exemplary implementation of the FDN 203 (or 204 or 205) of FIG. 3. The system of FIG. 4 has four reverberation tanks (each with a gain stage g i and a delay line z -ni ), but variations of this system (and other FDNs used in embodiments of the virtualizer of the present invention) implement more or fewer than four reverberation tanks.

[0083] The FDN of FIG. 4 includes an input gain element 300, an all-pass filter (APF) 301 coupled to the output of element 300, addition elements 302, 303, 304, and 305 coupled to the output of APF 301, and four reverberation tanks respectively coupled to the outputs of different ones of elements 302, 303, 304, and 305 (each reverberation tank has a gain element g k (one of elements 306), a delay line z -Mk coupled thereto (one of elements 307), and a gain element 1 / g k coupled thereto (one of elements 309), where 0 ≦ k - 1 ≦ 3). A unitary matrix 308 is coupled to the output of delay line 307 and is configured to present a feedback output to the respective second inputs of elements 302, 303, 304, and 305. The outputs of two of the gain elements 309 (the first and second reverberation tanks) are presented to the input of an addition element 310, and the output of element 310 is presented to one input of an output mixing matrix 312. The outputs of the other two of the gain elements 309 (the third and fourth reverberation tanks) are presented to the input of an addition element 311, and the output of element 311 is presented to the other input of output mixing matrix 312.

[0084] Element 302 is configured to add the output of matrix 308 corresponding to delay line z -n1 to the input of the first reverberation tank (i.e., apply feedback from the output of delay line z -n1 via matrix 308). Element 303 is configured to add the output of delay line z -n2The output of the matrix 308 corresponding to is added to the input of the second reverberation tank (i.e., feedback from the output of the delay line z -n2 is applied). Element 304 adds the output of the matrix 308 corresponding to the delay line z -n3 to the input of the third reverberation tank (i.e., feedback from the output of the delay line z -n3 is applied). Element 305 adds the output of the matrix 308 corresponding to the delay line z -n4 to the input of the fourth reverberation tank (i.e., feedback from the output of the delay line z -n4 is applied).

[0085] The input gain element 300 of the FDN in FIG. 4 is coupled to receive one frequency band of the converted monophonic downmix signal (filter bank region signal) output from the decomposed filter bank 202 in FIG. 3. The input gain element 300 applies a gain (scaling) factor G in to the filter bank region signal presented thereto. Collectively, the scaling factors G in for all frequency bands (implemented by all of the FDNs 203, 204,..., 205 in FIG. 3) control the spectral shaping and level of the late reverberation. Setting the input gain G in in all of the FDNs of the virtualizer in FIG. 3 often takes into account the following goals: The direct-to-late ratio (DLR) of the BRIR applied to each channel that matches the actual room; The necessary low-frequency attenuation to mitigate excessive coming artifacts and / or low-frequency rumble; Matching of the diffuse field spectral envelope.

[0086] Assuming that the direct response (applied by the subsystem 100 in FIG. 3) provides unitary gain in all frequency bands, a particular DLR (power ratio) is: G in=sqrt(ln(10 6 ) / (T60*DLR)) is achieved by setting G in in this way. Here, T60 is the reverberation decay time defined as the time it takes for the reverberation to decay by 60 dB (which is determined by the reverberation delay and reverberation gain discussed below), and "ln" represents the natural logarithm function.

[0087] The input gain factor G in may depend on the content being processed. One application of such content dependence is to ensure that, regardless of any correlation present between the input channel signals, the energy of the downmix in each time / frequency segment is equal to the sum of the energies of the individual channel signals being downmixed. In that case, the input gain factor is

Number

[0088] In a typical QMF domain implementation of the FDN of FIG. 4, the signal presented from the output of the all-pass filter (APF) 301 to the input of the reverberation tank is a sequence of QMF domain frequency components. To generate a more natural-sounding FDN output, the APF 301 is applied to the output of the gain element 300 to introduce phase diversity and increased echo density. Alternatively or additionally, one or more all-pass filters may be applied to the individual inputs to the downmixing subsystem 201 (of FIG. 3) before the input is downmixed in the subsystem 201 and processed by the FDN, or in the reverberation tank feedforward or feedback paths depicted in FIG. 4 (e.g., to the delay line z in each reverberation tank -Mk in addition to or instead of), or may be applied to the output of the FDN (i.e., to the output of the output matrix 312).

[0089] When implementing the reverberation tank delay z -ni to avoid the reverberation modes aligning at the same frequency, the reverberation delays n i should be relatively prime to each other. The total delay should be large enough to provide sufficient mode density to avoid an artificially-sounding output. However, the shortest delay should be short enough to avoid an excessive time gap between the late reverberation and the other components of the BRIR.

[0090] Typically, the reverberation tank outputs are initially panned to either the left or right binaural channel. Usually, the set of reverberation tank outputs panned to the two binaural channels is the same number and mutually exclusive. It is also desirable to balance the timing of the two binaural channels. Thus, if the reverberation tank output with the shortest delay goes to one binaural channel, the reverberation tank output with the second shortest delay will go to the other channel.

[0091] The reverberation tank delay can vary across the frequency band so as to change the mode density as a function of frequency. Generally, lower frequency bands require higher mode densities and thus longer reverberation tank delays.

[0092] Reverberation tank gain g i The amplitude and the reverberation tank delay together determine the reverberation delay time of the FDN of FIG. 4: T 60 =-3n i / log 10 (|g i |) / F FRM where F FRM is the frame rate of the filter bank 202 (of FIG. 3). The phase of the reverberation tank gain introduces a fractional delay so as to overcome problems associated with the fact that the reverberation tank delay is quantized to the downsample factor lattice of the filter bank.

[0093] The unitary feedback matrix 308 provides an equal mixing among the respective reverberation tanks in the feedback path.

[0094] To equalize the levels of the reverberation tank outputs, the gain element 309 applies the normalized gain 1 / |g i | to the output of each reverberation tank, removing the level effect of the reverberation tank gain while preserving the fractional delay introduced by its phase.

[0095] The output mixing matrix 312 (also specified as matrix M out ) is a 2×2 matrix configured to mix the unmixed binaural channels (outputs of elements 310 and 311 respectively) from the initial panning to achieve the left and right binaural channels of the output (the L and R signals presented at the output of matrix 312) with the desired interaural coherence. The unmixed binaural channels are almost uncorrelated since after the initial panning they contain no common reverberation tank output. If the desired interaural coherence is Coh and |Coh|≦1, the output mixing matrix 312 is

Number

Number

[0096] For the FDN for each individual frequency band in the virtualizer of the present invention, if the target acoustic attributes T60, Coh, and DLR defined above are known, each FDN (each FDN may have the structure shown in FIG. 4) can be configured to achieve the target attributes. In particular, in some embodiments, the input gain (G in ) and the gain and delay of the reverberation tank (g i and n i ) and the parameters of the output matrix M out can be set (for example, by the control values presented thereto by the control subsystem 209 in FIG. 3). In practice, it is often sufficient to set the frequency-dependent attributes by a model with simple control parameters in order to generate a natural-sounding late reverberation that matches a specific acoustic environment.

[0097] Next, an example of how the target reverberation decay time (T 60 ) for the FDN for each specific frequency band of an embodiment of the virtualizer of the present invention can be determined by determining the target reverberation decay time (T 60 ) for each of a small number of frequency bands. The level of the FDN response decays exponentially with time. T 60 is inversely proportional to the decay factor df (defined as dB decay per unit time), that is: T 60 = 60 / df That is.

[0098] The attenuation factor df depends on frequency and generally increases linearly with respect to the logarithmic frequency scale. Therefore, the reverberation decay time is also a function of frequency and generally decreases as the frequency increases. Thus, if the values of T 60 are determined (e.g., set) for two frequency points, the T 60 curve for all frequencies is determined. For example, if the reverberation decay times for frequency points f A and f B are T 60,A and T 60,B respectively, the T 60 curve is defined as follows.

[0099]

Equation

[0100] Next, an example of how the target interaural coherence (Coh) for the FDN for each specific frequency band of an embodiment of the virtualizer of the present invention can be achieved by setting a small number of control parameters is described. The interaural coherence (Coh) of the late reverberation generally follows the pattern of a diffuse sound field. It can be modeled by a sinc function up to the crossover frequency f C and a constant above the crossover frequency. A simple model for the Coh curve is as follows.

[0101]

Equation

[0102] Next, an example of how the target direct-to-reverberant ratio (DLR) for the FDN for each specific frequency band of a certain embodiment of the virtualizer of the present invention can be achieved by setting a small number of control parameters is described. The direct-to-reverberant ratio (DLR) in dB generally increases linearly with respect to the logarithmic frequency, and is controlled by setting DLR 1K (DLR in dB at 1 kHz) and DLRslope (in dB per decade of frequency). However, a low DLR in the low frequency range often leads to excessive commingling artifacts. To mitigate the artifacts, two correction mechanisms for controlling DLR are added: a minimum DLR floor, DLRmin (in dB); and a transition frequency f T and a high-pass filter defined by the slope HPF slope (in dB per decade of frequency) of the attenuation curve below it.

[0103] The resulting DLR curve in dB is defined as follows.

[0104]

Equation

[0105] Modifications of the embodiments disclosed in this document have one or more of the following features: The virtualizer of the present invention is implemented in the time domain or has a hybrid implementation with FDN-based impulse response capture and FIR-based signal filtering; The virtualizer of the present invention is implemented to allow the application of energy compensation as a function of frequency during the execution of the downmix stage that generates a downmixed input signal for the late reverberation processing subsystem; The virtualizer of the present invention is implemented to allow manual or automatic control of the late reverberation attributes applied in response to external factors (i.e., in response to the setting of control parameters).

[0106] For applications where system latency is critical and the delay caused by the analysis and synthesis filter banks is prohibitive, the filter bank region FDN structure of a typical embodiment of the virtualizer of the present invention can be converted to the time domain, and each FDN structure can be implemented in the time domain in certain classes of embodiments of this virtualizer. In the time domain implementation, the input gain factor (Gin ) The subsystem that applies the reverberation tank gain (g i ) and the normalized gain (1 / |g i |) is replaced by a filter with a similar amplitude response to allow frequency-dependent control. The output mixing matrix (M out ) is also replaced by a matrix of filters. Unlike other filters, the phase response of this matrix of filters is crucial. This is because the power preservation and the interaural coherence can be affected by the phase response. The reverberation tank delay in the time-domain implementation may need to be slightly changed to avoid sharing the filter bank stride as a common factor (compared to the value in the filter bank domain implementation). Due to various constraints, the execution of the time-domain implementation of the FDN of the virtualizer of the present invention may not exactly match that of its filter bank implementation.

[0107] Referring to FIG. 8, next, a hybrid (filter bank domain and time domain) implementation of the late reverberation processing subsystem of the virtualizer of the present invention will be described. This hybrid implementation of the late reverberation processing subsystem of the present invention is a modification to the late reverberation processing subsystem 200 of FIG. 4 and implements impulse response capture based on FDN and signal filtering based on FIR.

[0108] FIG. 8 includes elements 201, 202, 203, 204, 205, and 207 that are identical to the elements with the same reference numerals in the subsystem 200 of FIG. 3. The above description of these elements will not be repeated with reference to FIG. 8. In the embodiment of FIG. 8, the unit impulse generator 211 is coupled to present an input signal (pulse) to the decomposition filter bank 202. The LBRIR filter 208 (mono input, stereo output) implemented as an FIR filter applies the appropriate late reverberation portion of the BRIR (LBRIR) to the monophonic downmix output from the subsystem 201. Thus, the elements 211, 202, 203, 204, 205, and 207 are the processing side chain for the LBRIR filter 208.

[0109] Whenever the settings of the late reverberation part LBRIR are modified, the impulse generator 211 is always operated to present a unit impulse to the element 202, the resulting output from the filter bank 207 is captured, and presented to the filter 208 (to set the filter 208 to apply the new LBRIR determined by the output of the filter bank 207). To accelerate the passage of time from the LBRIR setting change until the new LBRIR becomes effective, the samples of the new LBRIR can begin to replace the old LBRIR as they become available. To shorten the inherent latency of the FDN, the initial zeros of the LBRIR can be discarded. These options provide flexibility and allow for potential performance improvements (compared to the performance provided by the filter bank area implementation) at the expense of the calculations added by the FIR filtering in the hybrid implementation.

[0110] For applications where system latency is critical but computational power is not as much of an issue, a side-chain filter bank area late reverberation processor (such as that implemented by elements 211, 202, 203, 204, …, 205 in FIG. 8) can be used to supplement the effective FIR impulse response applied by the filter 208. The FIR filter 208 implements this captured FIR response and can apply it directly to the mono-downmix of the input channels (during the virtualization of the input channels).

[0111] Various FDN parameters, and thus the resulting late reverberation attributes, can be manually tuned and then incorporated as a fixed configuration into an embodiment of the late reverberation processing subsystem of the present invention. For example, by one or more presets that can be adjusted by a user of the system (e.g., by operating the control subsystem 209 of FIG. 3). However, given a high-level description of late reverberation, its relationship to FDN parameters, and the ability to modify its behavior, a wide variety of ways to control various embodiments of an FDN-based late reverberation processor are envisioned. These include, but are not limited to, the following.

[0112] 1. The end user can manually control the FDN parameters via a user interface on a display (e.g., implemented by an embodiment of the control subsystem 209 of FIG. 3), or switch presets using physical controls (e.g., implemented by an embodiment of the control subsystem 209 of FIG. 3). In this way, the end user can adapt the room simulation according to their preferences, the environment, or the content.

[0113] 2. The author of the audio content to be virtualized may provide settings or desired parameters to be transmitted along with the content itself, for example, by metadata provided along with the input audio signal. Such metadata can be parsed and used (e.g., by an embodiment of the control subsystem 209 of FIG. 3) to control the relevant FDN parameters. Thus, the metadata may indicate attributes such as reverberation time, reverberation level, direct-to-reverberation ratio, etc., and these attributes may vary over time and be indicated by time-varying metadata.

[0114] 3. The playback device may recognize its location or environment by one or more sensors. For example, a mobile device may use a GSM network, a Global Positioning System (GPS), known WiFi access points, or any other location service to determine where the device is. Subsequently, data indicating the location and / or environment may be used to control relevant FDN parameters (e.g., by the embodiment of the control subsystem 209 in FIG. 3). Thus, the FDN parameters may be modified in response to the location of the device, for example, to mimic the physical environment.

[0115] 4. Cloud services or social media may be used to derive the most common settings that consumers use in certain environments in relation to the location of the playback device. Additionally, the user may associate their current settings with their (known) location and upload them to a cloud or social media service to make them available for other users or themselves.

[0116] 5. The playback device may include other sensors such as a camera, a light sensor, a microphone, an accelerometer, a gyroscope, etc. to determine the user's activities and the environment the user is in. This is to optimize the FDN parameters for that specific activity and / or environment.

[0117] 6. The FDN parameters may be controlled by audio content. An audio classification algorithm or manually annotated content may indicate whether audio segments include speech, music, sound effects, silence, etc. The FDN parameters may be adjusted according to such labels. For example, the direct-to-reverberation ratio may be reduced for dialog to improve dialog intelligibility. Additionally, video analysis may be used to determine the position of the current video segment, and the FDN parameters may be appropriately adjusted to better simulate the environment depicted in the video. And / or 7. The semiconductor playback system may use an FDN setting different from that of the mobile device. For example, the setting may be device-dependent. The in-room semiconductor system may simulate an in-room scenario with a typical (fairly reverberant) distant source, while the mobile device may render content closer to the listener.

[0118] Some implementations of the virtualizer of the present invention include an FDN (e.g., the FDN implementation of FIG. 4) configured to apply a fractional delay in addition to an integer sample delay. For example, in one such implementation, a fractional delay element is connected in series with a delay line that adds an integer delay equal to an integer number of sample periods within each reverberation tank (e.g., each fractional delay element is positioned after or otherwise in series with one of the delay lines). The fractional delay can be approximated by a phase shift (unit complex multiplication) corresponding to a fraction f = τ / T of the sample period, where f is the delay fraction, τ is the desired delay for that band, and T is the sample period for that band, in each frequency band. In the context of applying reverberation in the QMF domain, how to add a fractional delay is well known.

[0119] In a first class of embodiments, the present invention is a headphone virtualization method that generates a binaural signal in response to a set of channels of a multi-channel audio input signal (e.g., each of those channels or each of all frequency range channels). The method comprises: (a) applying a binaural room impulse response (BRIR) to each channel of the set (e.g., by convolving each channel of the set with the BRIR corresponding to the channel in subsystems 100 and 200 of FIG. 3 or in subsystems 12, …, 14, 15 of FIG. 2), thereby generating a filtered signal (e.g., the output of subsystems 100 and 200 of FIG. 3 or the output of subsystems 12, …, 14, 15 of FIG. 2), the step including using at least one feedback delay network (e.g., FDN 203, 204, …, 205 of FIG. 3) to add a common late reverberation to the downmixing (e.g., monophonic downmixing) of the channels of the set; and (b) combining the filtered signals (e.g., in subsystem 210 of FIG. 3 or in a subsystem including elements 16 and 18 of FIG. 2) to generate a binaural signal. Typically, a bank of FDNs is used to add the common late reverberation to the downmixing (e.g., each FDN adds late reverberation to a different frequency band). Typically, step (a) includes (e.g., in subsystem 100 of FIG. 3 or in subsystems 12, …, 14 of FIG. 2) applying the “direct response and early reflections” portion of a single-channel BRIR for that channel to each channel of the set, and the common late reverberation is generated to emulate a collective macro attribute of the late reverberation portion of at least a part (e.g., all) of the single-channel BRIR.

[0120] In a typical implementation of the first class, each FDN is implemented in a hybrid complex quadrature mirror filter (HCQMF) region or a quadrature mirror filter (QMF) region. In some such embodiments, the frequency-dependent spatial acoustic attributes of the binaural signal are controlled by controlling the configuration of each FDN used to add reverberation (e.g., using the control subsystem 209 of FIG. 3). Typically, for efficient binaural rendering of the audio content of a multi-channel signal, a channel's monophonic downmix (e.g., the downmix generated by subsystem 201 of FIG. 3) is used as the input to the FDN. Typically, the downmix process is controlled based on the source distance for each channel (i.e., the distance between the assumed source of the channel's audio content and the assumed user position), and depends on the handling of the direct response corresponding to the source distance to preserve the temporal and level structure of each BRIR (i.e., the direct response and early reflection portions of a single-channel BRIR for a channel and each BRIR determined by the common late reverberation for the downmix including that channel). The channels to be downmixed can be time-aligned in various ways and scaled during the downmix, but the proper level and temporal relationships between the direct response, early reflections, and common late reverberation portions of the BRIR for each channel should be maintained. In embodiments that use a single FDN bank to generate a common late reverberation portion for all channels to be downmixed (to generate the downmix), appropriate gains and delays need to be applied during the generation of the downmix (for each channel to be downmixed).

[0121] Typical embodiments of this class include the step of adjusting the FDN coefficients corresponding to frequency-dependent attributes (e.g., reverberation decay time, interaural coherence, mode density, and direct-to-late ratio). This enables better matching of the acoustic environment and a more natural-sounding output.

[0122] In a second class of embodiments, the invention is a method of generating a binaural signal in response to a multi-channel audio input signal. This is by applying a binaural room impulse response (BRIR) to each channel of a selected set of channels of the input signal (e.g., each of the channels of the input signal or each of the input signal's full frequency range channels), e.g., by convolving each channel with a corresponding BRIR. This includes processing each channel of the set in a first processing path (e.g., implemented by subsystem 100 of FIG. 3 or subsystems 12, …, 14 of FIG. 2) configured to model and apply to each channel the direct response and early reflections of a single-channel BRIR for that channel (e.g., the EBRIR applied by subsystems 12, 14, or 15 of FIG. 2), and processing a downmix of the channels of the set (e.g., a monophonic downmix) in a second processing path parallel to the first processing path (e.g., implemented by subsystem 200 of FIG. 3 or subsystem 15 of FIG. 2). The second processing path is configured to model and apply a common late reverberation (e.g., the LBRIR applied by subsystem 15 of FIG. 2) to the downmix. Typically, the common late reverberation emulates the collective macro attributes of the late reverberation portions of at least some (e.g., all) of the single-channel BRIRs. Typically, the second processing path includes at least one FDN (e.g., one FDN for each of a plurality of frequency bands). Typically, a mono-downmix is used as the input to all of the reverberation tanks of each FDN implemented by the second processing path. Typically, a mechanism for systematic control of the macro attributes of each FDN is provided to produce a binaural virtualization that better simulates the acoustic environment and sounds more natural (e.g., control subsystem 209 of FIG. 3). Since most such macro attributes are frequency-dependent, each FDN is typically implemented in a hybrid complex quadrature mirror filter (HCQMF) domain, frequency domain, domain, or another filter bank domain, and different FDNs are used for each frequency band.The main benefit of implementing the FDN in the filter bank region is that it allows the application of reverberation with frequency-dependent reverberation attributes. In various embodiments, the FDN is implemented in any of a wide variety of filter bank regions using any of a variety of filter banks. It includes, but is not limited to, quadrature mirror filters (QMFs), finite impulse response filters (FIR filters), infinite impulse response filters (IIR filters), or crossover filters.

[0123] Some embodiments of the first class (and the second class) implement one or more of the following features.

[0124] 1. FDN implementation in a filter bank region (e.g., a hybrid complex orthogonal mirror filter region) (e.g., the FDN implementation of FIG. 4) or FDN implementation in a hybrid filter bank region and late reverberation filter implementation in the time domain (e.g., the structure described with reference to FIG. 8). This typically allows independent adjustment of the FDN parameters and / or settings for each frequency band (which enables simple and flexible control of frequency-dependent acoustic attributes). This provides, for example, the ability to vary the reverberation tank delay in various bands so as to vary the mode density as a function of frequency.

[0125] 2. The particular downmix process used to generate a downmixed (e.g., monophonic downmixed) signal processed in a second processing path (from a multi-channel input audio signal) depends on the handling of the direct response to maintain the proper level and timing relationship between the source distance of each channel and the direct and late responses.

[0126] 3. In order to introduce phase diversity and increased echo density without changing the resulting reverberation spectrum and / or timbre, an all-pass filter (e.g., APF 301 in FIG. 4) is applied in a second processing path (e.g., at the input or output of a bank of FDNs).

[0127] 4. To overcome problems related to quantization of delays in the downsample-factor grid, fractional delay is implemented in the feedback path of each FDN in a complex-valued multirate structure.

[0128] 5. In the FDN, the reverberation tank output is linearly mixed directly into the binaural channels (e.g., by matrix 312 in FIG. 4) using output mixing coefficients set based on the desired interaural coherence in each frequency band. Optionally, the mapping of the reverberation tank to the binaural output channels alternates across frequency bands to achieve equalized delays between the binaural channels. Optionally, a normalization factor is applied to the reverberation tank output to equalize its level while preserving fractional delay and overall power.

[0129] 6. The frequency-dependent reverberation decay time is controlled by setting an appropriate combination of reverberation tank delays and gains in each frequency band to simulate an actual room.

[0130] 7. One scaling factor is applied for each frequency band (e.g., by elements 306 and 309 in FIG. 4), either at the input or output of the relevant processing path. This allows: Controlling the frequency-dependent direct-to-late ratio (DLR) that matches the DLR of an actual room (a simple model may be used to calculate the required scaling factor based on the target DLR and reverberation decay time, e.g., T60); Providing low-frequency attenuation to mitigate excessive combing artifacts; and / or Applying diffuse field spectral shaping to the FDN response.

[0131] 8. A simple parametric model is implemented (e.g., by the control subsystem 209 of FIG. 3) to control the essential frequency-dependent attributes of late reverberation such as reverberation decay time, interaural coherence, and / or direct-to-late ratio.

[0132] In some embodiments (e.g., for applications where system latency is critical and the delays introduced by the analysis and synthesis filter banks are prohibitive), the filter bank region FDN structure (e.g., the FDN of FIG. 4 in each frequency band) of a typical embodiment of the system of the present invention is replaced by an FDN structure implemented in the time domain (e.g., the FDN 220 of FIG. 10 that can be implemented as shown in FIG. 9). In the time-domain embodiment of the system of the present invention, the subsystems of the filter bank region embodiment that apply the input gain factor (G in ), the reverberation tank gain (g i ), and the normalization gain (1 / |g i |) are replaced by time-domain filters (and / or gain elements) to allow frequency-dependent control. The output mixing matrix of a typical filter bank region implementation (e.g., the output mixing matrix 312 of FIG. 4) is replaced (in a typical time-domain embodiment) by a set of outputs of time-domain filters (e.g., elements 500 - 503 of the implementation of FIG. 11 of element 424 of FIG. 9). Unlike other filters in a typical time-domain embodiment, the phase response of this set of outputs of the filter is typically crucial (since power conservation and interaural coherence can be affected by the phase response). In some time-domain embodiments, the reverberation tank delay is varied (e.g., slightly varied) from the value in the corresponding filter bank region implementation (e.g., to avoid sharing a filter bank stride as a common factor).

[0133] FIG. 10 is a block diagram of an embodiment of a headphone virtualization system of the present invention similar to FIG. 3, but in the system of FIG. 10, elements 202 to 207 of FIG. 3 are replaced by a single FDN 220 implemented in the time domain (for example, FDN 220 of FIG. 10 may be implemented in the same manner as the FDN of FIG. 9). In FIG. 10, two (left and right channel) time domain signals are output from the direct response and early reflection processing subsystem 100, and two (left and right channel) time domain signals are output from the late reverberation processing subsystem 221. An addition element 210 is coupled to the outputs of subsystems 100 and 200. Element 210 combines (mixes) the left channel outputs of subsystems 100 and 221 to generate the left channel L of the binaural audio signal output from the virtualizer of FIG. 10, and combines (mixes) the right channel outputs of subsystems 100 and 221 to generate the right channel R of the binaural audio signal output from the virtualizer of FIG. 10. Assuming that appropriate level adjustment and time alignment are implemented in subsystems 100 and 221, element 210 can be implemented to simply sum the corresponding left channel samples output from subsystems 100 and 221 to generate the left channel of the binaural output signal, and simply sum the corresponding right channel samples output from subsystems 100 and 221 to generate the right channel of the binaural output signal.

[0134] In the system of FIG. 10, a multi-channel audio input signal (having channel X i ) is directed to two parallel processing paths where it is processed. One passes through the direct response and early reflection processing subsystem 100, and the other passes through the late reverberation processing subsystem 221. The system of FIG. 10 is configured to apply a BRIR i to each channel X i . Each BRIR iIt can be decomposed into two parts: a direct response and early reflection part (applied by subsystem 100) and a late reverberation part (applied by subsystem 221). In operation, the direct response and early reflection processing subsystem 100 thus generates the direct response and early reflection part of the binaural audio signal output from the virtualizer, and the late reverberation processing subsystem (the "late reverberation generator") 221 thus generates the late reverberation part of the binaural audio signal output from the virtualizer. The outputs of subsystems 100 and 221 are mixed (by subsystem 210) to generate a binaural audio signal, which is typically presented from subsystem 210 to a rendering system (not shown) and undergoes binaural rendering for playback by headphones in the rendering system.

[0135] The downmix subsystem 201 (of the late reverberation processing subsystem 221) is configured to downmix the channels of the multi-channel input signal to a mono-downmix (which is a time-domain signal), and the FDN 220 is configured to apply the late reverberation part to the mono-downmix.

[0136] Referring to FIG. 9, next, an example of a time-domain FDN that can be used as the FDN 220 of the virtualizer in FIG. 10 will be described. The FDN in FIG. 9 includes an input filter 400 coupled to receive a mono-downmix of all channels of a multi-channel audio input signal (e.g., generated by the subsystem 201 of the system in FIG. 10). The FDN in FIG. 9 includes an all-pass filter (APF) 401 (corresponding to the APF 301 in FIG. 4) coupled to the output of the filter 400, an input gain element 401A coupled to the output of the filter 401, and addition elements 402, 403, 404, and 405 (corresponding to the addition elements 302, 303, 304, and 305 in FIG. 4) coupled to the output of the element 401A, and four reverberation tanks. Each reverberation tank is coupled to the output of a different one of the elements 402, 403, 404, and 405 and has one of the reverberation filters 406 and 406A, 407 and 407A, 408 and 408A, and 409 and 409A coupled thereto, one of the delay lines 410, 411, 412, and 413 (corresponding to the delay line 307 in FIG. 4) coupled thereto, and one of the gain elements 417, 418, 419, and 420 coupled to the output of one of these delay lines.

[0137] A unitary matrix 415 (corresponding to the unitary matrix 308 in FIG. 4 and typically implemented to be identical to the matrix 308) is coupled to the outputs of the delay lines 410, 411, 412, and 413. The matrix 415 is configured to present a feedback output to the respective second inputs of the elements 402, 403, 404, and 405.

[0138] When the delay (n1) added by line 410 is shorter than the delay (n2) added by line 411, the delay added by line 411 is shorter than the delay (n3) added by line 412, and the delay added by line 412 is shorter than the delay (n4) added by line 413, the outputs of gain elements 417 and 419 (of the first and third reverberation tanks) are presented at the inputs of adder element 422, and the outputs of gain elements 418 and 420 (of the second and fourth reverberation tanks) are presented at the inputs of adder element 423. The output of element 422 is presented at one input of IACC and mixing filter 424, and the output of element 423 is presented at the other input of IACC filtering and mixing stage 424.

[0139] An example implementation of the gain elements 417 - 420 and elements 422, 423, and 424 of FIG. 9 will be described with reference to a typical implementation of the elements 310 and 311 and output mixing matrix 312 of FIG. 4. The output mixing matrix 312 of FIG. 4 (also specified as matrix M out is a 2×2 matrix configured to mix the unmixed binaural channels from the initial panning (the outputs of elements 310 and 311 respectively) to generate left and right binaural output channels (the left ear "L" and right ear "R" signals presented at the output of matrix 312) with the desired interaural coherence. This initial panning is implemented by elements 310 and 311. Each of them combines two reverberation tank outputs to generate one of the unmixed binaural channels, and the reverberation tank output with the shortest delay is presented at the input of element 310, and the reverberation tank output with the second shortest delay is presented at the input of element 311. Elements 422 and 423 of the embodiment of FIG. 9 perform the same type of initial panning as that performed by elements 310 and 311 (in each frequency band) of the embodiment of FIG. 4 on the stream of filter bank region components (in the relevant frequency bands) of the time domain signals presented at their inputs.

[0140] Since there is no common reverberation tank output at all, the unmixed binaural channels (output from elements 310 and 311 in FIG. 4 or from elements 422 and 423 in FIG. 9), which are almost uncorrelated, may be mixed (by matrix 312 in FIG. 4 or stage 424 in FIG. 9) to implement a panning pattern that achieves the desired interaural coherence for the left and right binaural output channels. However, since the reverberation tank delay is different in each FDN (i.e., the FDN in FIG. 9 or the FDN implemented for each different frequency band in FIG. 4), one of the unmixed binaural channels (the output of one of elements 310 and 311 or 422 and 423) is always ahead of the other unmixed binaural channel (the output of the other of elements 310 and 311 or 422 and 423).

[0141] Thus, in the embodiment of FIG. 4, if the combination of the reverberation tank delay and the panning pattern is the same across all frequency bands, an acoustic bias will result. This bias can be alleviated if the panning pattern is alternated across frequency bands such that the mixed binaural output channels advance and lag relative to each other in alternating frequency bands. For example, if the desired interaural coherence is Coh and |Coh|≦1, the output mixing matrix 312 in the odd-numbered frequency bands may be implemented to multiply the two inputs presented to it by a matrix having the form

Number

Number

[0142] Alternatively, if the above-described audio and video bias in the binaural output channels has its input channel order switched for alternating frequency bands (e.g., for odd frequency bands, the output of element 310 may be presented to the first input of matrix 312, and the output of element 311 may be presented to the second input of matrix 312; for even frequency bands, the output of element 311 may be presented to the first input of matrix 312, and the output of element 310 may be presented to the second input of matrix 312), it can be mitigated by implementing matrix 312 such that it is the same in the FDN for all frequency bands.

[0143] In the embodiment of FIG. 9 (and other time domain embodiments of the FDN of the system of the present invention), it is not trivial to alternate panning based on frequency to address the audio bias that would ordinarily result when the unmixed binaural channel output from element 422 is always ahead (delayed) of the unmixed binaural channel output from element 423. This audio bias is addressed in a different way than it is typically addressed in the filter bank region embodiments of the FDN of the system of the present invention in typical time domain embodiments of the FDN of the system of the present invention. In particular, in the embodiment of FIG. 9 (and other time domain embodiments of the FDN of the system of the present invention), the relative gains of the unmixed binaural channels (e.g., the outputs from elements 422 and 423 of FIG. 9) are determined by gain elements (e.g., elements 417, 418, 419, and 420 of FIG. 9) to compensate for the audio bias that would ordinarily result for the above unbalanced timing. By implementing a gain element (e.g., element 417) to attenuate the signal that arrives earliest (which is panned to one side, e.g., by element 422), and implementing a gain element (e.g., element 418) to boost the next earliest signal (which is panned to the other side, e.g., by element 423), the stereo image is recentered. Thus, the reverberation tank including gain element 417 applies a first gain to the output of element 417, and the reverberation tank including gain element 418 applies a second gain (different from the first gain) to the output of element 418. Thereby, the first gain and the second gain attenuate the first unmixed binaural channel (output from element 422) relative to the second unmixed binaural channel (output from element 423).

[0144] More specifically, in a typical implementation of the FDN of FIG. 9, the four delay lines 410, 411, 412, and 413 have sequentially increasing lengths and sequentially increasing delay values n1, n2, n3, and n4, respectively. In this implementation, filter 417 applies a gain of g 1 Thus, the output of filter 417 is g 1is the delayed version of the input to delay line 410 to which the gain of 2 is applied. Similarly, filter 418 applies a gain of g 3 filter 419 applies a gain of g 4 and filter 420 applies a gain of g 2 Thus, the output of filter 418 is the delayed version of the input to delay line 411 to which the gain of 3 is applied, the output of filter 419 is the delayed version of the input to delay line 412 to which the gain of 4 is applied, and the output of filter 420 is the delayed version of the input to delay line 413 to which the gain of

[0145] In this implementation, the selection of the following gain values: g 1 = 0.5, g 2 = 0.5, g 3 = 0.5, g 4 = 0.5 may lead to an undesirable bias to one side of the output audio image (i.e., to the left or right channel) as indicated by the binaural channel output from element 424. According to an embodiment of the present invention, the values g 1 g 2 g 3 g 4 are selected as follows to center the audio image: g 1 = 0.38, g 2 = 0.6, g 3 = 0.5, g 4 = 0.5. Thus, the output stereo image, according to an embodiment of the present invention, attenuates the earliest arriving signal (which is panned to one side by element 422 in the current example) with respect to the second earliest arriving signal (i.e., select as g 1 < g 3 ), and boosts the second earliest signal (which is panned to the other side by element 423 in the current example) with respect to the latest arriving signal (i.e., g 4 < g 2By selecting as such), it is recentered.

[0146] The typical implementation of the time-domain FDN in FIG. 9 has the following differences and similarities with respect to the filter bank region (CQMF region) FDN in FIG. 4.

[0147] The same unitary feedback matrix A (matrix 308 in FIG. 4 and matrix 415 in FIG. 9).

[0148] Similar reverberation tank delays n i (That is, the delay in the CQMF implementation in FIG. 4 is 1 / T s assuming that is the sampling rate (1 / T s is typically equal to 48 KHz), n 1 = 17 * 64T s = 1088 * T s n 2 = 21 * 64T s = 1344 * T s n 3 = 26 * 64T s = 1664 * T s n 4 = 29 * 64T s = 1856 * T s may also be, while the delay in the time-domain implementation is n 1 = 1089 * T s n 2 = 1345 * T s n 3 = 1663 * T s n 4 = 185 * T s may also be. Note that in a typical CQMF implementation, there is an actual constraint that each delay is some integer multiple of the duration of a 64-sample block, while in the time domain, there is more flexibility in the selection of each delay, and thus more flexibility in the selection of the delay of each reverberation tank).

[0149] Similar all-pass filter implementations (i.e., similar implementations of filter 301 in FIG. 4 and filter 401 in FIG. 9). For example, the all-pass filter can be implemented by a cascade of several (e.g., three) all-pass filters. For example, each cascaded all-pass filter can be in the form of

Number

[0150] In some implementations of the time-domain FDN of FIG. 9, the input filter 400 is implemented to match (at least substantially) the direct-to-late ratio (DLR) of the BRIR applied by the system of FIG. 9 to the target DLR, and so that the DLR of the BRIR applied by a virtualizer (such as the virtualizer of FIG. 10) including the system of FIG. 9 can be changed by replacing the filter 400 (or controlling the configuration settings of the filter 400). For example, in some embodiments, the filter 400 is implemented as a cascade of filters that implement the target DLR and optionally implement the desired DLR control (e.g., a first filter 400A and a second filter 400B coupled as shown in FIG. 9A). For example, the cascade of filters is an IIR filter (e.g., filter 400A is a first-order Butterworth high-pass filter (IIR filter) configured to match the target low-frequency characteristics, and filter 400B is a second-order low-shelf IIR filter configured to match the target high-frequency characteristics). As another example, the cascade of filters is an IIR and an FIR filter (e.g., filter 400A is a second-order Butterworth high-pass filter (IIR filter) configured to match the target low-frequency characteristics, and filter 400B is a 14th-order FIR filter configured to match the target high-frequency characteristics). Typically, the direct signal is fixed and the filter 400 modifies the late signal to achieve the target DLR. The all-pass filter (APF) 401 is preferably implemented to perform the same function as the APF 301 of FIG. 4, i.e., to introduce phase diversity and increased echo density to generate a more natural-sounding FDN output. The input filter 400 controls the amplitude response while the APF 401 typically controls the phase response.

[0151] In FIG. 9, filter 406 and gain element 406A together implement a reverberation filter, filter 407 and gain element 407A together implement another reverberation filter, filter 408 and gain element 408A together implement another reverberation filter, and filter 409 and gain element 409A together implement another reverberation filter. Each of filters 406, 407, 408, and 409 in FIG. 9 is preferably implemented as a filter having a maximum gain value (unity gain) close to 1, and each of gain elements 406A, 407A, 408A, and 409A is configured to apply an attenuation gain to the output of the corresponding one of filters 406, 407, 408, and 409 that matches a desired attenuation (after reverberation tank delay n i ). Specifically, gain element 406A is configured to apply an attenuation gain (decaygain 1 ) to the output of filter 406 such that the output of delay line 410 (after reverberation tank delay n i ) has a first target attenuated gain at the output of element 406A, gain element 407A is configured to apply an attenuation gain (decaygain 2 ) to the output of filter 407 such that the output of delay line 411 (after reverberation tank delay n 2 ) has a second target attenuated gain at the output of element 407A, gain element 408A is configured to apply an attenuation gain (decaygain 3 ) to the output of filter 408 such that the output of delay line 412 (after reverberation tank delay n 3 ) has a third target attenuated gain at the output of element 408A, and gain element 409A is configured to apply an attenuation gain (decaygain 4 ) to the output of filter 409 such that the output of delay line 413 (after reverberation tank delay n 4 ) has a fourth target attenuated gain at the output of element 409A.

[0152] Each of the filters 406, 407, 408, and 409 and each of the elements 406A, 407A, 408A, and 409A of the system of FIG. 9 are preferably implemented to achieve the target T60 characteristic of the BRIR applied by a virtualizer (e.g., the virtualizer of FIG. 10) that includes the system of FIG. 9 (each of the filters 406, 407, 408, and 409 is preferably an IIR filter, e.g., implemented as a shelf filter or in cascade with shelf filters). Here, T60 represents the reverberation decay time (T 60 ). For example, in some embodiments, each of the filters 406, 407, 408, and 409 is implemented as a shelf filter (e.g., a shelf filter with Q = 0.3 and a shelf frequency of 500 Hz to achieve the T60 characteristic shown in FIG. 13; in FIG. 13, T60 is in seconds) or in cascade with two IIR shelf filters (e.g., those with shelf frequencies of 100 Hz and 1000 Hz to achieve the T60 characteristic shown in FIG. 14; in FIG. 14, T60 is in seconds). The shape of each shelf filter is determined to match the desired curve of change from low frequency to high frequency. When filter 406 is implemented as a shelf filter (or in cascade with multiple shelf filters), the reverberation filter having filter 406 and gain element 406A is also a shelf filter (or in cascade with shelf filters). Similarly, when each of filters 407, 408, and 409 is implemented as a shelf filter (or in cascade with shelf filters), each reverberation filter having filter 407 (or 408 or 409) and the corresponding gain element (407A, 408A, or 409A) is also a shelf filter (or in cascade with shelf filters).

[0153] FIG. 9B is an example of filter 406 implemented as a cascade of a first shelf filter 406B and a second shelf filter 406C coupled as shown in FIG. 9B. Each of filters 407, 408, 409 may be implemented in the same manner as the implementation of filter 406 in FIG. 9B.

[0154] In some embodiments, the decay gain applied by elements 406A, 407A, 408A, 409A (decaygain i ) is determined as follows.

[0155]

Equation

[0156] FIG. 11 is an embodiment of the following elements of FIG. 9: elements 422 and 423 and the IACC (Interaural Cross-Correlation coefficient) filtering and mixing stage 424. Element 422 is configured to sum the outputs of filters 417 and 419 (of FIG. 9) and couple the summed signal to the input of a low shelf filter 500, and element 422 is configured to sum the outputs of filters 418 and 420 (of FIG. 9) and couple the summed signal to the input of a high-pass filter 501. The outputs of filters 500 and 501 are added (mixed) at element 502 to generate a binaural left ear output signal, and the outputs of filters 500 and 501 are mixed at element 502 (the output of filter 500 is subtracted from the output of filter 501 at element 502) to generate a binaural right ear output signal. Elements 502 and 503 mix (add and subtract) the filtered outputs of filters 500 and 501 to generate a binaural output signal that achieves the target IACC characteristic (within an acceptable accuracy range). In the embodiment of FIG. 11, each of the low shelf filter 500 and the high-pass filter 501 is typically implemented as a first-order IIR filter. In an example where filters 500 and 501 have such an implementation, the embodiment of FIG. 11 can achieve the exemplary IACC characteristic plotted as curve "I" in FIG. 12. This is a good match to the target IACC characteristic plotted as "I" in FIG. 12. T This is a good match to the target IACC characteristic plotted as "I" in FIG. 12.

[0157] FIG. 11A is a graph of the frequency response (R1) of a typical implementation of the filter 500 of FIG. 11, the frequency response (R2) of a typical implementation of the filter 501 of FIG. 11, and the response of the filters 500 and 501 connected in parallel. From FIG. 11A, it is clear that the combined response is desirably flat across the range of 100 Hz to 10,000 Hz.

[0158] Thus, in certain embodiments, the present invention is a system (e.g., the system of FIG. 10) and method for generating a binaural signal (e.g., the output of element 210 in FIG. 10) in response to a set of channels of a multi-channel audio input signal. This includes the step of applying a binaural room impulse response (BRIR) to each channel of the set, thereby generating a filtered signal, including by using a single feedback delay network (FDN) to add a late reverberation common to the downmix of the channels of the set; and the step of combining the filtered signals to generate the binaural signal. The FDN is implemented in the time domain. In some such embodiments, the time domain FDN (e.g., the FDN 220 of FIG. 10 configured as in FIG. 9) is: An input filter (e.g., filter 400 of FIG. 9) having an input coupled to receive the downmix, the input filter being configured to generate a first filtered downmix in response to the downmix; An all-pass filter (e.g., all-pass filter 401 of FIG. 9) coupled and configured to perform a second filtering of the downmix in response to the first filtered downmix; A reverberation application subsystem (e.g., all elements of FIG. 9 other than elements 400, 401, and 424) having a first output (e.g., the output of element 422) and a second output (e.g., the output of element 423), the reverberation application subsystem including a set of reverberation tanks, each reverberation tank having a different delay, the reverberation application subsystem being coupled and configured to generate a first unmixed binaural channel and a second unmixed binaural channel in response to the second filtered downmix, presenting the first unmixed binaural channel at the first output and presenting the second unmixed binaural channel at the second output; Coupled to the reverberation application subsystem and configured to generate a first mixed binaural channel and a second mixed binaural channel in response to the first unmixed binaural channel and the second unmixed binaural channel, including an interaural cross-correlation coefficient (IACC) filtering and mixing stage (e.g., stage 424 of FIG. 9 that may be implemented as elements 500, 501, 502, 503 of FIG. 11).

[0159] The input filter may be implemented (preferably as a cascade of two filters configured to generate it) to generate the first filtered downmix such that each BRIR has a direct-to-reverberant ratio (DLR) that at least substantially matches the target DLR.

[0160] Each reverberation tank may be configured to generate a delayed signal, and a gain is applied to the signal propagating in each reverberation tank such that the delayed signal has a gain that at least substantially matches the target delayed gain for the delayed signal, and may include a configured reverberation filter (e.g., implemented as a shelf filter or a cascade of shelf filters). This is to achieve the target reverberation decay time characteristic (e.g., T 60 Characteristic) of each BRIR.

[0161] In some embodiments, the first unmixed binaural channel is ahead of the second unmixed binaural channel, and the reverberation tank includes a first reverberation tank (e.g., the reverberation tank of FIG. 9 including delay line 410) configured to generate a first delayed signal having the shortest delay, and a second reverberation tank (e.g., the reverberation tank of FIG. 9 including delay line 411) configured to generate a second delayed signal having the second shortest delay. The first reverberation tank is configured to apply a first gain to the first delayed signal, the second reverberation tank is configured to apply a second gain to the second delayed signal, the second gain is different from the first gain, and the application of the first gain and the second gain results in attenuation of the first unmixed binaural channel relative to the second unmixed binaural channel. Typically, the first mixed binaural channel and the second mixed binaural channel exhibit a re-centered stereo image. In some embodiments, the IACC filtering and mixing stage is configured to generate the first mixed binaural channel and the second mixed binaural channel such that the first mixed binaural channel and the second mixed binaural channel have IACC characteristics that at least substantially match the target IACC characteristics.

[0162] Aspects of the present invention include methods and systems (e.g., system 20 of FIG. 2 or system of FIG. 3 or FIG. 10) that perform (or are configured to perform or support the performance of) binaural virtualization of audio signals (e.g., audio signals where the audio content consists of speaker channels and / or object-based audio signals).

[0163] In some embodiments, the virtualizer of the present invention is coupled to receive or generate input data indicative of a multi-channel audio input signal, and is programmed with software (or firmware) to perform any of a variety of processes, including embodiments of the method of the present invention, on the input data or otherwise configured (e.g., in response to control data), and is a general-purpose processor or includes such a general-purpose processor. Such a general-purpose processor is typically coupled to an input device (e.g., a mouse and / or keyboard), memory, and a display device. For example, the system of FIG. 3 (or a virtualizer system having the elements 12, …, 14, 15 of the system 20 of FIG. 2 or the system 20) can be implemented in a general-purpose processor, where the input is audio data indicative of N channels of the audio input signal and the output is audio data indicative of two channels of a binaural audio signal. A conventional digital-to-analog converter (DAC) can act on the output data to generate an analog version of the binaural signal channels for reproduction by a speaker (e.g., a headphone pair).

[0164] Individual embodiments of the present invention and applications of the present invention are described herein, but it will be apparent to those skilled in the art that many variations to these embodiments and applications described herein are possible without departing from the scope of the invention described and claimed in this application. Although certain forms of the present invention have been shown and described, it should be understood that the present invention is not limited to the particular embodiments described and shown or to the particular methods described.

Claims

1. 1. A method for generating a binaural signal in response to a set of channels of a multi-channel audio input signal, the method comprising: applying a binaural room impulse response (BRIR) to each channel of the set, thereby generating a filtered signal; and combining the filtered signals to generate the binaural signal; applying the BRIR to each channel of the set includes introducing, using a late reverberation generator, a common late reverberation into a downmix of the channels of the set in response to a control value presented to the late reverberation generator, the common late reverberation emulating collective macro-properties of late reverberant portions of single channel BRIRs shared across at least some channels of the set; a content dependent energy equalization factor is applied to the downmix, such that a centre channel of the multi-channel audio input signal is mixed into a left channel of the downmix by a factor of 1 / √2 and into a right channel of the downmix by a factor of 1 / √2; method.

2. The method of claim 1 , wherein applying a BRIR to each channel of the set comprises applying to each channel of the set a direct response and early reflection portion of a single-channel BRIR for that channel.

3. 2. The method of claim 1, wherein the late reverberation generator comprises a bank of feedback delay networks for adding the common late reverberation to the downmix, each feedback delay network of the bank adding late reverberation to a different frequency band of the downmix.

4. The method of claim 3 , wherein each of said feedback delay networks is implemented in a complex quadrature mirror filter domain.

5. The method of claim 1 , wherein the late reverberation generator comprises a single feedback delay network for adding the common late reverberation to the downmix of the channels of the set, the feedback delay network being implemented in the time domain.

6. 1. A system for generating a binaural signal in response to a set of channels of a multi-channel audio input signal, the system comprising: applying a binaural room impulse response (BRIR) to each channel of the set, thereby generating a filtered signal; combining the filtered signals to generate the binaural signal; having one or more processors, applying the BRIR to each channel of the set includes introducing, using a late reverberation generator, a common late reverberation into a downmix of the channels of the set in response to a control value presented to the late reverberation generator, the common late reverberation emulating collective macro-properties of late reverberant portions of single channel BRIRs shared across at least some channels of the set; a content dependent energy equalization factor is applied to the downmix, such that a centre channel of the multi-channel audio input signal is mixed into a left channel of the downmix by a factor of 1 / √2 and into a right channel of the downmix by a factor of 1 / √2; system.

7. The system of claim 6 , wherein applying a BRIR to each channel of the set comprises applying to each channel of the set a direct response and early reflection portion of a single-channel BRIR for that channel.

8. 7. The system of claim 6, wherein the late reverberation generator comprises a bank of feedback delay networks configured to add the common late reverberation to the downmix, each feedback delay network of the bank adding late reverberation to a different frequency band of the downmix.

9. The system of claim 8 , wherein each of said feedback delay networks is implemented in a complex quadrature mirror filter domain.

10. 7. The system of claim 6, wherein the late reverberator comprises a feedback delay network implemented in the time domain, and the late reverberator is configured to process the downmix in the time domain in the feedback delay network to add the common late reverberation to the downmix.

11. 13. A non-transitory computer readable storage medium having a sequence of instructions which, when executed by an audio signal processing apparatus, causes the audio signal processing apparatus to perform the method of claim 1.

Citation Information

Patent Citations

  • Sound compensation device

    JP2007336080A

  • Signal generation for binaural signals

    JP2011529650A

  • Reverberation device and method for reverberating audio signals

    JP2013508760A