Generation of binaural audio in response to multichannel audio using at least one feedback delay network.
Patent Information
- Application Number
- JP2025080881
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2014-05-05
- Filing Date
- 2025-05-14
- Publication Date
- 2026-09-14
- Estimated Expiration
- 2034-12-18
Smart Images

Figure 0007920365000015 
Figure 0007920365000016 
Figure 0007920365000017
Abstract
Description
[Technical Field]
[0001] Cross-references to related applications This application claims priority to China Patent Application No. 201410178258.0 filed on 29 April 2014, U.S. Provisional Patent Application No. 61 / 923,579 filed on 3 January 2014, and U.S. Provisional Patent Application No. 61 / 988,617 filed on 5 May 2014. The contents of each application are incorporated herein by reference in their entirety. 1. Field of Invention The present invention relates to a method (sometimes referred to as a headphone virtualization method) and system for generating a binaural signal in response to a multi-channel audio input signal by applying a binaural room impulse response (BRIR) to each channel (for example, all channels) of a set of channels in an input signal. In some embodiments, at least one feedback delay network (FDN) applies the late reverberation portion of the downmixed BRIR to the downmix of the channels. [Background technology]
[0002] 2. Background of the Invention Headphone virtualization (or binaural rendering) is a technique that aims to deliver a surround sound experience or an immersive sound field using standard stereo headphones.
[0003] Early headphone virtualizers applied head-related transfer functions (HRTFs) to convey spatial information in binaural rendering. An HRTF is a set of direction- and distance-dependent filter pairs that characterize how sound travels from a specific point in space (sound source location) to the listener's ears in an anechoic environment. Essential spatial cues such as interaural time difference (ITD), interaural level difference (ILD), head shadowing effect, and spectral peaks and notches due to shoulder and auricular reflexes can be perceived in the rendered HRTF-filtered binaural content. Due to the constraints of human head size, HRTFs do not provide sufficient or robust cues for source distances beyond approximately one meter. As a result, virtualizers based solely on HRTFs typically do not achieve good out-of-head localization or perceived distance.
[0004] Many acoustic events in daily life occur in reverberant environments. In reverberant environments, audio signals reach the listener's ears not only through the direct path (from source to ear) modeled by the HRTF, but also through various reflective paths. Reflections introduce profound influences on the auditory experience, such as distance, room size, and other spatial attributes. To convey this information in binaural rendering, the virtualizer must apply room reverberation in addition to the cues in the direct path HRTF. The binaural room impulse response (BRIR) characterizes the transformation of audio signals from a specific point in space to the listener's ear in a particular acoustic environment. Theoretically, the BRIR contains all acoustic cues related to spatial perception.
[0005] Figure 1 shows the full frequency range of each channel (X1, ..., X) of a multi-channel audio input signal. NThis is a block diagram of one type of standard headphone virtualizer configured to apply a binaural room impulse response (BRIR) to channels X1, ..., X N Each of these is a speaker channel corresponding to a different source direction relative to the assumed listener (i.e., the direction of the direct path from the assumed position of the corresponding speaker to the assumed listener position), and each such channel is convolved with BRIR for the corresponding source direction. The acoustic path from each channel needs to be simulated for each ear. Therefore, for the remainder of this paper, the term BRIR refers to either a single impulse response or a pair of impulse responses associated with the left and right ears. Thus, subsystem 2 is configured to convolve channel X1 with BRIR1 (BRIR for the corresponding source direction), and subsystem 4 is configured to convolve channel X N BRIR N It is configured to convolve with (BRIR for the corresponding source direction), and so on. The output of each BRIR subsystem (subsystems 2, ..., 4) is a time-domain signal containing the left and right channels. The left channel outputs of the BRIR subsystems are mixed in summing element 6, and the right channels of the BRIR subsystems are mixed in summing element 8. The output of element 6 is the left channel L of the binaural audio signal output from the virtualizer, and the output of element 8 is the right channel R of the binaural audio signal output from the virtualizer.
[0006] A multi-channel audio input signal may also include low-frequency effects (LFE) or subwoofer channels, which are identified as the "LFE" channel in Figure 1. Typically, the LFE channel is not convolved with the BRIR, but instead is attenuated (e.g., by -3dB or more) in gain stage 5 in Figure 1, and the output of gain stage 5 is mixed equally (by summing elements 6 and 8) with each channel of the virtualizer's binaural output signal. An additional delay stage may be required in the LFE path to time-align the output of stage 5 with the outputs of the BRIR subsystems (2, ..., 4). Alternatively, the LFE channel may simply be ignored (i.e., not presented to or processed by the virtualizer). For example, the embodiment of the present invention in Figure 2 (described later) simply ignores any LFE channels in the multi-channel audio input signal it processes. Many consumer headphones cannot accurately reproduce LFE channels.
[0007] In some conventional virtualizers, the input signal is converted from the time domain to the frequency domain into a quadrature mirror filter (QMF) domain, generating channels of QMF domain frequency components. These frequency components are then filtered in the QMF domain (for example, in the QMF domain implementations of subsystems 2, ..., 4 in Figure 1), and the resulting frequency components are then converted back to the time domain (for example, in the final stages of subsystems 2, ..., 4 in Figure 1). Thus, the audio output of the virtualizer is a time domain signal (for example, a time domain binaural signal).
[0008] Generally, each channel across its entire frequency range in a multi-channel audio signal input to a headphone virtualizer is assumed to represent audio content emitted from a sound source located at a known position relative to the listener's ear. The headphone virtualizer is configured to apply a binaural room impulse response (BRIR) to each such channel of the input signal. Each BRIR can be decomposed into two parts: the direct response and the reflection. The direct response is an HRTF corresponding to the direction of arrival (DOA) of the sound source, adjusted with appropriate gain and delay due to the distance (between the sound source and the listener), and optionally amplified with a parallax effect for small distances.
[0009] The remaining part of BRIR models reflections. Early reflections are typically first or second-order reflections and have a relatively sparse temporal distribution. The microstructure of each first or second-order reflection (e.g., ITD and ILD) is important. For late reflections (sound reflected from three or more surfaces before reaching the listener), the echo density increases with increasing number of reflections, making it difficult to observe the micro-attributes of individual reflections. For increasingly later reflections, the macrostructure (e.g., reverberation decay rate, interaural coherence, and the overall spectral distribution of reverberation) becomes more important. For this reason, reflections can be further segmented into two parts: early reflections and late reverberation.
[0010] The delay of the direct response is the distance from the listener to the source divided by the speed of sound, and its level is inversely proportional to the source distance (unless there is a wall or large surface near the source). On the other hand, the delay and level of late reverberation are generally not sensitive to the source location. For practical reasons, virtualizers may choose to time-align direct responses from sources at different distances and / or compress their dynamic range. However, the temporal and level relationships between direct response, early reflection, and late reverberation within the BRIR should be preserved.
[0011] The effective length of a typical BRIR can reach several hundred milliseconds or more in many acoustic environments. Direct application of BRIRs requires convolution with filters of thousands of taps, which is computationally expensive. In addition, without parameterization, achieving sufficient spatial resolution requires a large memory space to store BRIRs for different source positions. Lastly, but not to be underestimated, the sound source position can change over time and / or the listener's position and orientation can change over time. Accurate simulation of such motion requires time-varying BRIR impulse responses. Proper interpolation and application of such time-varying filters can be difficult when the impulse responses of these filters have many taps.
[0012] To implement a spatial reverberator configured to apply simulated reverberation to one or more channels of a multi-channel audio input signal, a filter with a well-known filter structure known as a feedback delay network (FDN) can be used. The structure of an FDN is simple. Several reverberation tanks (for example, in the FDN of Figure 4, the gain element g1 and the delay line z) -n1 The FDN has reverberation tanks, each with delay and gain. In a typical implementation of FDN, the outputs from all reverberation tanks are mixed by a unitary feedback matrix, and the output of the matrix is fed back and summed with the inputs of the reverberation tanks. Gain adjustment may be applied to the reverberation tank outputs. The reverberation tank outputs (or their gain-adjusted versions) can be suitably remixed for multi-channel or binaural playback. With a compact computation and memory footprint, FDNs can generate and apply natural-sounding reverberations. For this reason, FDNs have been used in virtualizers to supplement direct responses generated by HRTFs.
[0013] For example, a commercially available "Dolby Mobile" headphone virtualizer includes a reverberator with an FDN-based structure capable of adding reverberation to each channel of a five-channel audio signal (with left front, right front, center, left surround, and right surround channels) and filtering each reverberated channel using different filter pairs of a set of five head-related transfer function ("HRTF") filter pairs. The "Dolby Mobile" headphone virtualizer can also operate to generate a two-channel "reverberated" binaural audio output (a two-channel virtual surround sound output with added reverberation) in response to a two-channel audio input signal. When the reverberated binaural output is rendered and played back by the headphone pair, it is perceived at the listener's eardrum as HRTF-filtered, reverberated sound from five loudspeakers located at left front, right front, center, left rear (surround), and right rear (surround) positions. The virtualizer upmixes the downmixed two-channel audio input (without using any spatial cue parameters received with the audio input) to generate five upmixed audio channels, adds reverb to the upmixed channels, and downmixes the five reverb-enhanced channel signals to generate the virtualizer's two-channel reverb-enhanced output. The reverb for each upmixed channel is filtered through different pairs of HRTF filters. [Overview of the project] [Problems that the invention aims to solve]
[0014] In virtualizers, FDNs are configured to achieve a certain reverberation decay time and echo density. However, FDNs lack the flexibility to simulate the microstructure of early reflections. Furthermore, in typical virtualizers, tuning and configuring FDNs is largely trial and error.
[0015] Headphone virtualizers that do not simulate all reflection paths (early and late) cannot achieve effective out-of-head localization. The inventors have come to realize that virtualizers using FDNs that attempt to simulate all reflection paths (early and late) typically achieve only limited success in simulating both early and late reverberations and adding both to the audio signal. The inventors have also come to realize that virtualizers using FDNs but lacking the ability to properly control spatial acoustic attributes such as reverberation decay time, interaural coherence, and direct-to-late ratio may achieve some degree of out-of-head localization, but at the cost of introducing excessive tonal distortion and reverberation. [Means for solving the problem]
[0016] In a first-class embodiment, the present invention is a method for generating a binaural signal in response to a set of channels of a multi-channel audio input signal (for example, each of those channels or each of the channels across the entire frequency range). The method includes: (a) applying a binaural chamber impulse response (BRIR) to each channel of the set (for example, by convolving each channel of the set with the BRIR corresponding to the channel) to generate a filtered signal, comprising using at least one feedback delay network (FDN) to add a common late reverberation to a downmix of the channels of the set (for example, a monophonic downmix); and (b) combining the filtered signals to generate a binaural signal. Typically, a bank of FDNs is used to add the common late reverberation to the downmix (for example, each FDN adds a common late reverberation to a different frequency band). Typically, step (a) includes applying the “direct response and early reflection” portion of a single-channel BRIR for each channel in the set, and the common late reverberation is generated to emulate the collective macro-attributes of at least some (e.g., all) of the late reverberation portions of the single-channel BRIR.
[0017] Methods for generating binaural signals in response to multi-channel audio input signals (or in response to a set of channels of such signals) are sometimes referred to in this paper as “headphone virtualization” methods, and systems configured to perform such methods are sometimes referred to in this paper as “headphone virtualizers” (or “headphone virtualization systems” or “binaural virtualizers”).
[0018] In a typical implementation of the first class, each FDN is implemented in a filter bank region (e.g., a hybrid complex quadrature mirror filter (HCQMF) region, a quadrature mirror filter (QMF) region, or other transformation or subband region that may include decimation). In some such embodiments, frequency-dependent spatial acoustic attributes of the binaural signal are controlled by controlling the configuration of each FDN used to add late reverberation. Typically, a monophonic downmix of channels is used as input to the FDN for efficient binaural rendering of audio content of a multi-channel signal. A typical embodiment of the first class includes a step of adjusting FDN coefficients corresponding to frequency-dependent attributes (e.g., reverberation decay time, interaural coherence, mode density, and direct-to-late ratio) by presenting control values to a feedback delay network to set at least one of the following parameters for each FDN: input gain, reverberation tank gain, reverberation tank delay, or output matrix parameter. This allows for better matching of the acoustic environment and a more natural-sounding output.
[0019] In a second class embodiment, the present invention is a method for generating a binaural signal in response to a multi-channel audio input signal having multiple channels. This is achieved by applying a binaural chamber impulse response (BRIR) to each channel of a set of channels in the input signal (e.g., each of the channels in the input signal or the entire frequency range channels of each of the input signals). This includes processing each channel in the set in a first processing path configured to model and apply the direct response and early reflection of a single-channel BRIR for that channel to each channel, and processing a downmix of the channels in the set (e.g., a monophonic (mono) downmix) in a second processing path (parallel to the first processing path) configured to model and apply a common late reverberation to the downmix. Typically, the common late reverberation is generated to emulate the collective macro-attributes of at least some (e.g., all) of the late reverberation portions of the single-channel BRIR. Typically, the second processing path includes at least one FDN (e.g., one FDN for each of several frequency bands). Typically, a mono downmix is used as the input to all reverberation tanks of each FDN implemented by a second processing path. Typically, a mechanism is provided for systematic control of the macro attributes of each FDN to better simulate the acoustic environment and produce binaural virtualization that sounds more natural. Since most such macro attributes are frequency-dependent, each FDN is typically implemented in the hybrid complex quadrature mirror filter (HCQMF) domain, frequency domain, domain, or another filter bank domain, with different or independent FDNs used for each frequency band. The main benefit of implementing FDNs in the filter bank domain is that it allows for the application of reverberation with frequency-dependent reverberation attributes. In various embodiments, FDNs are implemented in any of a wide variety of filter bank domains using any of the diverse filter banks.This includes, but is not limited to, real or complex-valued quadrature mirror filters (QMFs), finite impulse response filters (FIR filters), infinite impulse response filters (IIR filters), discrete Fourier transforms (DFTs), (modified) cosine or sine transforms, wavelet transforms, or crossover filters. In some preferred implementations, the filter bank or transform used may involve decimation (e.g., reducing the sampling rate of the frequency-domain signal representation) to reduce the computational complexity of the FDN process.
[0020] Some embodiments of the first class (and the second class) implement one or more of the following features:
[0021] 1. FDN implementations in the filter bank region (e.g., hybrid complex quadrature mirror filter region) or hybrid filter bank region FDN implementations and time-domain late reverberation filter implementations. This typically allows independent tuning of FDN parameters and / or settings for each frequency band (which enables simple and flexible control of frequency-dependent acoustic attributes). This is, for example, by providing the ability to vary the reverberation tank delay in different bands so that the mode density changes as a function of frequency.
[0022] 2. The specific downmixing process used to generate a downmixed (e.g., monophonically downmixed) signal processed in a second processing path (from a multi-channel input audio signal) depends on the source distance of each channel and the handling of the direct response to maintain the appropriate level and timing relationship between the direct and late responses.
[0023] 3. To introduce phase diversity and increased echo density without altering the resulting reverberation spectrum and / or timbre, an all-pass filter (APF) is applied in a second processing path (for example, at the input or output of the FDN bank).
[0024] 4. To overcome the problems related to quantized delays in the downsample-factor grid, fractional delays are implemented in the feedback paths of each FDN in a complex-valued multirate structure.
[0025] 5. In FDN, the reverberation tank output is directly and linearly mixed into the binaural channels using an output mixing coefficient set based on the desired interaural coherence in each frequency band. Optionally, the mapping of the reverberation tank to the binaural output channels alternates across frequency bands to achieve balanced delays between the binaural channels. Optionally, a normalization factor is applied to the reverberation tank output to equalize its level while preserving fractional delays and overall power.
[0026] 6. Frequency-dependent reverberation decay time and / or mode density are controlled by setting the appropriate combination of reverberation tank delay and gain in each frequency band to simulate a real room.
[0027] 7. One scaling factor is applied to each frequency band (for example, at either the input or output of the relevant processing path). This is: Control the frequency-dependent direct-to-late ratio (DLR) to match the actual room's DLR (a simple model may be used to calculate the required scaling factor based on the target DLR and reverberation decay time, e.g., T60); Provides low-frequency attenuation to mitigate excessive combing artifacts and / or low-frequency rumble; and / or This is to apply diffuse field spectral shaping to the FDN response.
[0028] 8. Simple parametric models are implemented to control the intrinsic frequency-dependent attributes of late reverberation, such as reverberation decay time, biaural coherence, and / or direct-to-late ratio.
[0029] Aspects of the present invention include methods and systems for performing (or being configured to perform or supporting) binaural virtualization of audio signals (for example, audio signals and / or object-based audio signals where audio content consists of speaker channels).
[0030] In another class of embodiments, the present invention relates to a method and system for generating a binaural signal in response to a set of channels of a multi-channel audio input signal. This includes a step of applying a binaural chamber impulse response (BRIR) to each channel of the set, thereby generating a filtered signal, by using a single feedback delay network (FDN) to add a common late reverberation to the downmix of the channels of the set; and a step of combining the filtered signals to generate a binaural signal. The FDN is implemented in the time domain. In some such embodiments, the time-domain FDN is: An input filter having an input coupled to receive the downmix, the input filter being configured to generate a first filtered downmix in response to the downmix; A pass filter configured to perform a second filtered downmix in response to the first filtered downmix is coupled with; A reverberation application subsystem having a first output and a second output, wherein the reverberation application subsystem comprises a collection of reverberation tanks, each having a different delay, and the reverberation application subsystem is configured to generate a first unmixed binaural channel and a second unmixed binaural channel in response to the second filtered downmix, and to exhibit the first unmixed binaural channel at the first output and the second unmixed binaural channel at the second output; It includes an interaural cross-correlation coefficient (IACC) filtering and mixing stage coupled to the reverberation application subsystem and configured to generate a first mixed binaural channel and a second mixed binaural channel in response to the first unmixed binaural channel and the second unmixed binaural channel.
[0031] The input filter may be implemented to generate the first filtered downmix such that each BRIR has a direct-to-late ratio (DLR) that at least substantially matches the target DLR (preferably as a cascade of two filters configured to generate it).
[0032] Each reverberation tank may be configured to generate a delayed signal and may include a reverberation filter (implemented, for example, as a shelf filter or a cascade of shelf filters) configured to add gain to the signal propagating in each reverberation tank so that the delayed signal has a gain that matches at least substantially the target delayed gain. The target reverberation decay time characteristics of each BRIR (e.g., T 60 This is to achieve the characteristic.
[0033] In some embodiments, the first unmixed binaural channel is more advanced than the second unmixed binaural channel, and the reverberation tank includes a first reverberation tank configured to produce a first delayed signal with the shortest delay, and a second reverberation tank configured to produce a second delayed signal with the second shortest delay. The first reverberation tank is configured to apply a first gain to the first delayed signal, and the second reverberation tank is configured to apply a second gain to the second delayed signal, wherein the second gain is different from the first gain, and the application of the first and second gains results in attenuation of the first unmixed binaural channel relative to the second unmixed binaural channel. Typically, the first mixed binaural channel and the second mixed binaural channel exhibit a re-centered stereo image. In some embodiments, the IACC filtering and mixing stage is configured to generate the first mixed binaural channel and the second mixed binaural channel such that the first mixed binaural channel and the second mixed binaural channel have IACC characteristics that at least substantially match the target IACC characteristics.
[0034] A typical embodiment of the present invention provides a simple and unified framework for supporting both speaker channel-based and object-based input audio. In embodiments where BRIR is applied to input signal channels that are object channels, the “direct response and early reflection” processing performed on each object channel assumes a source direction indicated by the metadata provided with the audio content of that object channel. In embodiments where BRIR is applied to input signal channels that are speaker channels, the “direct response and early reflection” processing performed on each speaker channel assumes a source direction corresponding to that speaker channel (i.e., the direction of the direct path from the assumed location of the corresponding speaker to the assumed listener location). Regardless of whether the input channel is an object channel or a speaker channel, the “late reverberation” processing is performed on the downmix of the input channel (e.g., a monophonic downmix) and does not assume any specific source direction for the audio content of the downmix.
[0035] Other aspects of the present invention include a headphone virtualizer configured (e.g., programmed) to perform any embodiment of the method of the present invention, a system (e.g., stereo, multichannel, or other decoder) including such a virtualizer, and a computer-readable medium (e.g., disk) for storing code to implement any embodiment of the method of the present invention. [Brief explanation of the drawing]
[0036] [Figure 1] This is a block diagram of a typical headphone virtualization system. [Figure 2] This is a block diagram of a system including one embodiment of the headphone virtualization system of the present invention. [Figure 3] This is a block diagram of another embodiment of the headphone virtualization system of the present invention. [Figure 4]Figure 3 is a block diagram of FDN types that can be included in a typical implementation of the system. [Figure 5] This is a graph of the reverberation decay time (T60) in milliseconds as a function of frequency in Hz, which can be achieved by one embodiment of the virtualizer of the present invention, where the value of T60 at each of two specific frequencies (fA and fB) is set as follows: T60, A = 320 ms at fA = 10 Hz and T60, B = 150 ms at fB = 2.4 kHz. [Figure 6] This is a graph of interauricular coherence (Coh) as a function of frequency in Hz, which can be achieved by one embodiment of the virtualizer of the present invention, where the control parameters Cohmax, Cohmin, and fC are set to Cohmax=0.95, Cohmin=0.05, and fC=700Hz. [Figure 7] This is a graph of the direct-to-late ratio (DLR) in dB units at a source distance of 1 meter, as a function of frequency in Hz, which can be achieved by one embodiment of the virtualizer of the present invention, where the control parameters DLR1K, DLRslope, DLRmin, HPFslope, and fT are set to have values of DLR1K = 18 dB, DLRslope = 6 dB for every 10 times the frequency, DLRmin = 18 dB, HPFslope = 6 dB for every 10 times the frequency, and fT = 200 Hz. [Figure 8] This is a block diagram of another embodiment of the late reverberation processing subsystem of the headphone virtualization system of the present invention. [Figure 9] This is a block diagram of a time-domain implementation of an FDN of a type included in some embodiments of the system of the present invention. [Figure 9A] Figure 9 is a block diagram of an example implementation of filter 400. [Figure 9B] This is a block diagram of an example implementation of filter 406 in Figure 9. [Figure 10] This is a block diagram of one embodiment of the headphone virtualization system of the present invention in which the late reverberation processing subsystem 221 is implemented in the time domain. [Figure 11]Figure 9 is a block diagram of embodiments of elements 422, 423, and 424 of the FDN. A is a graph of the frequency response (R1) of a typical implementation of filter 500, the frequency response (R2) of a typical implementation of filter 501, and the frequency response of filters 500 and 501 connected in parallel. [Figure 12] Figure 9 is a graph showing examples of IACC characteristics (curve "I") and target IACC characteristics (curve "IT") that can be achieved by a certain implementation of FDN. [Figure 13] Figure 9 shows a graph of the T60 characteristics that can be achieved by a certain implementation of the FDN by appropriately implementing filters 406, 407, 408, and 409 as shelf filters. [Figure 14] Figure 9 shows a graph of the T60 characteristics that can be achieved by a certain implementation of the FDN, by appropriately implementing filters 406, 407, 408, and 409 as cascades of two IIR shelf filters. [Modes for carrying out the invention]
[0037] <Notation and Nomenclature> Throughout this disclosure, including the claims, the expression "performing an operation on" a signal or data (e.g., filtering, scaling, transforming, or applying gain to a signal or data) is used broadly to mean performing the operation directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or preprocessing prior to the performance of the operation).
[0038] Throughout this disclosure, including the claims, the term “system” is used in a broad sense to mean a device, system, or subsystem. For example, a subsystem implementing a virtualizer may be called a virtualizer system, and a system containing such a subsystem (for example, a system that generates X output signals in response to a plurality of inputs, wherein the subsystem generates M of the inputs and the other XM inputs are received from an external source) may also be called a virtualizer system (or virtualizer).
[0039] Throughout this disclosure, including the claims, the term “processor” is used in a broad sense to mean a system or device that is programmable or otherwise configurable (e.g., using software or firmware) to perform operations on data (e.g., audio or video or other image data). Examples of processors include field-programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipelining operations on audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.
[0040] Throughout this disclosure, including the claims, the expression “decomposed filter bank” is used broadly to represent a system (e.g., subsystem) configured to apply a transformation (e.g., a time-domain to frequency-domain transformation) to a time-domain signal to generate values (e.g., frequency components) that represent the content of the time-domain signal in each of a set of frequency bands. Throughout this disclosure, including the claims, the expression “filter bank region” is used broadly to represent a region of frequency components generated by a transformation or decomposed filter bank (e.g., a region in which such frequency components are processed). Examples of filter bank regions include (but are not limited to) the frequency domain, the quadrature mirror filter (QMF) region, and the hybrid complex quadrature mirror filter (HCQMF) region. Examples of transformations that may be applied by a decomposed filter bank include (but are not limited to) the discrete cosine transform (DCT), the modified discrete cosine transform (MDCT), the discrete Fourier transform (DFT), and the wavelet transform. Examples of decomposed filter banks include (but are not limited to) quadrature mirror filters (QMFs), finite impulse response filters (FIR filters), infinite impulse response filters (IIR filters), crossover filters, and other filters with suitable multirate structures.
[0041] Throughout this disclosure, including the claims, the term “metadata” refers to data separate from the corresponding audio data (the audio content, including the metadata, in a bitstream). The metadata is associated with the audio data and indicates at least one feature or characteristic of the audio data (for example, what type(s) of processing has already been performed on the audio data, or should be performed on the audio data, or the trajectory of objects indicated by the audio data). The association of metadata with audio data is time-synchronous. Thus, the current (most recently received or updated) metadata may indicate that the corresponding audio data simultaneously includes the results of audio data processing having and / or of the indicated type of the indicated feature.
[0042] Throughout this disclosure, including the claims, the terms “to combine” or “to be combined” are used to mean direct or indirect connection. Thus, when a first device combines with a second device, the connection may be through a direct connection or through an indirect connection via other devices and connections.
[0043] Throughout this disclosure, including the claims, the following expressions have the following definitions:
[0044] The terms speaker and loudspeaker are used synonymously to refer to transducers that emit any sound. This definition includes loudspeakers that are implemented as multiple transducers (e.g., woofers and tweeters).
[0045] Speaker feed: An audio signal applied directly to a loudspeaker, or an audio signal applied to a loudspeaker and an amplifier in series.
[0046] Channel (or "audio channel"): A monophonic audio signal. Such a signal can typically be rendered to be equivalent to directly applying it to a loudspeaker at a desired or nominal location. The desired location may be static, as is typical with physical loudspeakers, or it may be dynamic.
[0047] Audio program: A set of one or more audio channels (at least one speaker channel and / or at least one object channel) and optionally associated metadata (e.g., metadata describing a desired spatial audio presentation).
[0048] Speaker Channel (or "Speaker Feed Channel"): An audio channel associated with a designated loudspeaker (in a desired or nominal position) or a designated speaker zone within a defined speaker configuration. The speaker channel is rendered to be equivalent to directly applying its audio signal to the designated loudspeaker (in a desired or nominal position) or to the speakers within the designated speaker zone.
[0049] Object Channel: An audio channel that represents the sound emitted by an audio source (sometimes referred to as an audio "object"). Typically, an object channel determines a parametric audio source description (for example, metadata indicating a parametric audio source description is included in or provided with the object channel). The source description may determine the sound emitted by the source (as a function of time), the apparent location of the source as a function of time (e.g., 3D spatial coordinates), and optionally at least one additional parameter that characterizes the source (e.g., apparent source size or width).
[0050] Object-based audio program: An audio program that includes one or more object channels (and optionally at least one speaker channel) and optionally associated metadata (for example, metadata indicating the trajectory of an audio object that emits the sound indicated by the object channels, or metadata indicating a desired spatial audio presentation of the sound indicated by the object channels, or metadata indicating identification information of at least one audio object that is the source of the sound indicated by the object channels).
[0051] Rendering: The process of converting an audio program into one or more speaker feeds, or the process of converting an audio program into one or more speaker feeds and then converting those speaker feeds into sound using one or more loudspeakers. (In the latter case, rendering is sometimes referred to in this paper as rendering "by" loudspeakers.) Audio channels can be trivially rendered ("at" the desired location) by directly applying the signal to physical loudspeakers at the desired location. Alternatively, one or more audio channels can be rendered using one of a variety of virtualization techniques designed to be substantially equivalent (to the listener) to such trivial rendering. In the latter case, each audio channel may be converted into one or more speaker feeds to be applied to one or more loudspeakers at known locations, generally different from the desired location, so that the sound emitted by the loudspeakers in response to the feeds is perceived as originating from the desired location. Examples of such virtualization techniques include binaural rendering via headphones (for example, using "Dolby Headphone" processing to simulate surround sound up to 7.1 channels for headphone wearers) and wave field synthesis.
[0052] In this paper, the notation "xy" or "xyz" channel signal for a multichannel audio signal means that the signal has "x" full-frequency speaker channels (corresponding to speakers nominally located in the horizontal plane of the assumed listener's ears), "y" LFE (or subwoofer) channels, and optionally "z" full-frequency overhead speaker channels (corresponding to speakers located above the assumed listener's head, for example, on or near the ceiling of the room).
[0053] In the present paper, the expression "IACC" represents the interaural cross-correlation coefficient in its ordinary meaning. This is an indicator of the difference between the arrival times of audio signals at the listener's ears, and is typically represented by a number ranging from a first value, which indicates that the arriving signals are equal in magnitude and exactly out of phase, through an intermediate value, which indicates that the arriving signals have no similarity, up to a maximum value, which indicates that the arriving signals are identical with the same amplitude and phase.
[0054] <Detailed description of preferred embodiments> Many embodiments of the present invention are technically possible. It will be clear to those skilled in the art how to implement them based on the present disclosure. Embodiments of the system and method of the present invention will be described with reference to FIGS. 2 to 14.
[0055] FIG. 2 is a block diagram of a system (20) including an embodiment of the headphone virtualization system of the present invention. The headphone virtualization system (sometimes referred to as a virtualizer) is configured to process N full-frequency-range channels (X1, ..., X N ) of a multi-channel audio input signal to apply a binaural room impulse response (BRIR). The channels X1, ..., X N (which may be speaker channels or object channels) each correspond to a specific source direction and distance relative to an intended listener, and the system of FIG. 2 is configured to convolve each such channel with the BRIR for the corresponding source direction and distance.
[0056] The system 20 is coupled to receive an encoded audio program, and therefrom extracts N full-frequency-range channels (X1, ..., X NThe decoder may include subsystems (not shown in Figure 2) that are coupled and configured to decode the program, including by restoring the encoded program, and provide them to elements 12, ..., 14, 15 of the virtualization system (having coupled elements 12, ..., 14, 15, 16, 18 as shown in the figure). The decoder may include additional subsystems, some of which perform functions unrelated to the virtualization functions performed by the virtualization system, and some of which perform functions related to the virtualization functions. For example, the latter functions may include extracting metadata from the encoded program and providing said metadata to a virtualization control subsystem that uses said metadata to control elements of the virtualizer system.
[0057] Subsystem 12 (together with subsystem 15) is configured to convolve channel X1 with BRIR1 (BRIR for the corresponding source direction and distance), and subsystem 14 (together with subsystem 15) channel X N BRIR N It is configured to convolve with (BRIR for the corresponding source direction), and similarly for each of the N-2 other BRIR subsystems. The output of subsystems 12, ..., 14, and 15 is a time-domain signal including the left and right channels. Adding elements 16 and 18 are coupled to the outputs of elements 12, ..., 14, and 15. Adding element 16 is configured to combine (mix) the left channel outputs of the BRIR subsystems, and adding element 18 is configured to combine (mix) the right channel outputs of the BRIR subsystems. The output of element 16 is the left channel L of the binaural audio signal output from the virtualizer in Figure 2, and the output of element 18 is the right channel R of the binaural audio signal output from the virtualizer in Figure 2.
[0058] An important feature of a typical embodiment of the present invention becomes clear when comparing the embodiment of the headphone virtualizer of the present invention shown in Figure 2 with the conventional headphone virtualizer shown in Figure 1. For comparison, when the same multi-channel audio input signal is presented to the systems in Figures 1 and 2, they produce the same direct response and early reflection portion (i.e., the relevant EBRIR in Figure 2). i ) has BRIR i The entire frequency range of each input signal channel X i Assume that the system is configured to apply to (though not necessarily to the same degree of success) each BRIR applied by the system in Figure 1 or Figure 2. i This includes the direct response and early reflection portions (for example, EBRIR1, ..., applied by subsystems 12-14 in Figure 2). N It can be broken down into two parts: one of the first parts and the later reverberation part. The embodiment in Figure 2 (and other typical embodiments of the present invention) is a multiple single-channel BRIR, i.e., BRIR i It is assumed that the late reverberation portion of the signal can traverse the source direction and thus be shared across all channels, and that the same late reverberation (i.e., a common late reverberation) can be applied to the downmix of all channels across the entire frequency range of the input signal. This downmix can be a monophonic (mono) downmix of all input channels, but alternatively, it may be a stereo or multichannel downmix obtained from the input channels (e.g., from a subset of the input channels).
[0059] More specifically, subsystem 12 in Figure 2 is configured to convolve the input signal channel X1 with EBRIR1 (the direct response and early reflection BRIR portion for the corresponding source direction), and subsystem 14 is configured to convolve the input signal channel X N to EBRIR NThe late reverberation subsystem 15 in Figure 2 is configured to convolve with (the direct response and early reflection BRIR portion for the corresponding source direction), and so on. The late reverberation subsystem 15 in Figure 2 is configured to generate a mono downmix of all full-frequency range channels of the input signal and to convolve this downmix with LBRIR (a common late reverberation for all channels being downmixed). The output of each BRIR subsystem of the virtualizer in Figure 2 (each of subsystems 12, ..., 14, 15) includes the left and right channels (of the binaural signal generated from the corresponding speaker channels or downmix). The left channel outputs of those BRIR subsystems are combined (mixed) in summing element 16, and the right channel outputs of those BRIR subsystems are combined (mixed) in summing element 18.
[0060] Assuming that appropriate level adjustment and time alignment are implemented in subsystems 12, ..., 14, and 15, the summing element 16 can be implemented to simply sum the corresponding left binaural channel samples (left channel outputs of subsystems 12, ..., 14, and 15) to generate the left channel of the binaural output signal. Similarly, also assuming that appropriate level adjustment and time alignment are implemented in subsystems 12, ..., 14, and 15, the summing element 18 can also be implemented to simply sum the corresponding right binaural channel samples (right channel outputs of subsystems 12, ..., 14, and 15) to generate the right channel of the binaural output signal.
[0061] Subsystem 15 in Figure 2 can be implemented in various ways, but typically includes at least one feedback delay network configured to add a common late reverberation to the monophonic downmix of the input signal channels presented thereto. Typically, each of subsystems 12, ..., 14 processes the channel (X) being processed. i ) Direct response and early reflection portion (EBRIR) of single-channel BRIR iWhen applying the common late reverberation, the common late reverberation is generated to emulate the collective macro-attributes of at least some (e.g., all) of the late reverberation portions of those single-channel BRIRs (whose "direct response and early reflection portions" are applied by subsystems 12, ..., 14). For example, one implementation of subsystem 15 has the same structure as subsystem 200 in Figure 3, including a bank of feedback delay networks (203, 204, ..., 205) configured to apply the common late reverberation to the monophonic downmix of the input signal channels presented to it.
[0062] Similarly, subsystems 12, ..., 14 in Figure 2 can be implemented in any way (in the time domain or filter bank domain), and the preferred implementation for any particular application depends on various factors such as performance, computation, and memory. In one exemplary implementation, each of subsystems 12, ..., 14 is configured to convolve the channel presented to it with FIR filters corresponding to the direct and early responses associated with that channel. The gains and delays are set appropriately so that the outputs of subsystems 12, ..., 14 may be combined simply and efficiently with the output of subsystem 15.
[0063] Figure 3 is a block diagram of another embodiment of the headphone virtualization system of the present invention. The embodiment in Figure 3 is similar to the embodiment in Figure 2, with two time-domain signals (left and right channels) output from the direct response and early reflection processing subsystem 100, and two time-domain signals (left and right channels) output from the late reverberation processing subsystem 200. An additive element 210 is coupled to the outputs of subsystems 100 and 200. Element 210 is configured to combine (mix) the left channel outputs of subsystems 100 and 200 to generate the left channel L of the binaural audio signal output from the virtualizer in Figure 3, and to combine (mix) the right channel outputs of subsystems 100 and 200 to generate the right channel R of the binaural audio signal output from the virtualizer in Figure 3. Assuming that appropriate level adjustment and time alignment are implemented in subsystems 100 and 200, element 210 can be implemented to generate the left channel of the binaural output signal by simply summing the corresponding left channel samples output from subsystems 100 and 200, and to generate the right channel of the binaural output signal by simply summing the corresponding right channel samples output from subsystems 100 and 200.
[0064] In the system shown in Figure 3, channel X of the multi-channel audio input signal i The signal is directed to two parallel processing paths, where it undergoes processing. One passes through the direct response and early reflection processing subsystem 100, and the other through the late reverberation processing subsystem 200. The system in Figure 3 is shown for each channel X i BRIR i It is configured to apply each BRIR. iThe signal can be broken down into two parts: a direct response and early reflection portion (applied by subsystem 100) and a late reverberation portion (applied by subsystem 200). In operation, the direct response and early reflection processing subsystem 100 generates the direct response and early reflection portion of the binaural audio signal output from the virtualizer, and the late reverberation processing subsystem ("late reverberation generator") 200 generates the late reverberation portion of the binaural audio signal output from the virtualizer. The outputs of subsystems 100 and 200 are mixed (by the summing subsystem 210) to generate a binaural audio signal, which is typically presented from subsystem 210 to a rendering system (not shown), where it undergoes binaural rendering for playback by headphones.
[0065] Typically, when rendered and played back through a pair of headphones, a typical binaural audio signal output from element 210 is perceived by the listener's eardrum as sound from "N" loudspeakers located at any of a wide variety of positions, including in front of, behind, and above the listener (where N ≥ 2, and N is typically 2, 5, or 7). Playback of the output signal generated in the operation of the system in Figure 3 can give the listener the experience of sound coming from two or more (e.g., five or seven) "surround" sources. At least some of these sources are virtual.
[0066] The direct response and early reflection processing subsystem 100 can be implemented in any way (in the time domain or filter bank domain), and the preferred implementation for any particular application depends on various factors such as performance, computation, and memory. In one exemplary implementation, subsystem 100 is configured to convolve each channel presented to it with an FIR filter corresponding to the direct and early response associated with that channel. The gain and delay are set appropriately so that the output of subsystem 100 may be combined simply and efficiently with the output of subsystem 200 (in element 210).
[0067] As shown in Figure 3, the late reverberation generator 200 includes a downmix subsystem 201, a decomposition filter bank 202, banks of FDNs (FDNs 203, 204, ..., 205), and a composite filter bank 207, combined as shown in the figure. The subsystem 201 is configured to downmix the channels of a multichannel input signal to a mono downmix, and the decomposition filter bank 202 is configured to apply a transformation to the mono downmix to divide the mono downmix into "K" frequency bands, where K is an integer. The filter bank domain values (output from filter bank 202) in each different frequency band are presented to different FDNs 203, 204, ..., 205 (there are "K" of these FDNs, each combined and configured to apply the late reverberation portion of BRIR to the filter bank domain values presented to it). The filter bank domain values are preferably decimated over time to reduce the computational complexity of the FDNs.
[0068] In principle, each input channel (to subsystems 100 and 201 in Figure 3) can be processed by its own FDN (or bank of FDNs) to simulate the late reverberation portion of its BRIR. Despite the fact that the late reverberation portions of BRIRs associated with different source locations are typically very different in terms of the root mean square in the impulse response, their statistical attributes such as their average power spectrum, energy decay structure, mode density, and peak density are often very similar. Therefore, since a set of late reverberation portions of BRIRs are typically very similar perceptually across channels, it is possible to use one common FDN or bank of FDNs (e.g., FDNs 203, 204, ..., 205) to simulate the late reverberation portions of two or more BRIRs. In a typical embodiment, such one common FDN (or bank of FDNs) is used, and its inputs consist of one or more downmixes constructed from the input channels. In the exemplary implementation shown in Figure 2, the downmix is a monophonic downmix of all input channels (presented at the output of subsystem 201).
[0069] Referring to the embodiment in Figure 2, each of the FDNs 203, 204, ..., 205 is implemented in the filter bank domain and is configured to process different frequency bands of values output from the decomposed filter bank 202 to generate left and right reverberated signals for each band. For each band, the left reverberated signal is a sequence of filter bank domain values, and the right reverberated signal is another sequence of filter bank domain values. The combined filter bank 207 applies a frequency-to-time domain conversion to 2K sequences of filter bank domain values (e.g., frequency components of the QMF domain), and collects the converted values to a left channel time-domain signal (representing the audio content of a mono downmix with late reverberation applied) and a right channel time-domain signal (also representing the audio content of a mono downmix with late reverberation applied). These left and right channel signals are output to element 210.
[0070] In a typical implementation, each of the FDNs 203, 204, ..., 205 is implemented in the QMF domain, and filter bank 202 converts the mono downmix from subsystem 201 to the QMF domain (e.g., the Hybrid Complex Quadrature Mirror Filter (HCQMF) domain), so that the signals presented from filter bank 202 to the inputs of each of the FDNs 203, 204, ..., 205 are sequences of QMF domain frequency components. In such an implementation, the signal presented from filter bank 202 to FDN 203 is a sequence of QMF domain frequency components in the first frequency band, the signal presented from filter bank 202 to FDN 204 is a sequence of QMF domain frequency components in the second frequency band, and the signal presented from filter bank 202 to FDN 205 is a sequence of QMF domain frequency components in the "K" frequency band. When the decomposed filter bank 202 is implemented in this manner, the composite filter bank 207 applies a QMF-to-time domain conversion to a sequence of 2K output QMF-domain frequency components from the FDN to generate the late-reverberation-added time-domain signals for the left and right channels output to element 210.
[0071] For example, in the system shown in Figure 3, if K=3, there are six inputs to the composite filter bank 207 (left and right channels, each containing frequency-domain or QMF-domain samples output from FDNs 203, 204, and 205, respectively) and two outputs from 207 (left and right channels, each consisting of time-domain samples). In this example, filter bank 207 is typically implemented as two composite fill banks. One (exhibiting three left channels from FDNs 203, 204, and 205) is configured to generate the time-domain left channel signal output from filter bank 207, and the second (exhibiting three right channels from FDNs 203, 204, and 205) is configured to generate the time-domain right channel signal output from filter bank 207.
[0072] Optionally, the control subsystem 209 is coupled to each of the FDNs 203, 204, ..., 205 and configured to present control parameters to each of those FDNs to determine the late-blooming reverberation (LBRIR) applied by subsystem 200. Examples of such control parameters are described below. In some implementations, the control subsystem 209 may be capable of operating in real time (i.e., in response to user commands presented to it by the input device) to implement real-time variation of the late-blooming reverberation (LBRIR) applied by subsystem 200 to the monophonic downmix of the input channel.
[0073] For example, if the input signal to the system in Figure 2 is a 5.1-channel signal (with all frequency-range channels in the following channel order: L, R, C, Ls, Rs), then all frequency-range channels have the same source distance, and the downmix subsystem 201 can be implemented as the following downmix matrix. This simply involves summing up all frequency-range channels to form a mono downmix.
[0074]
number
[0075]
number
[0076]
number
[0077]
number
[0078] Next, we will discuss the individual implementations of the virtualizer downmix subsystem 201 and subsystems 100 and 200 shown in Figure 3.
[0079] The downmixing process implemented by subsystem 201 depends on the source distance (between the sound source and the assumed listener position) for each channel to be downmixed, and on how the direct response is handled. The delay t of the direct response d teeth: t d =d / v s Here, d is the distance between the sound source and the listener, and v s is the speed of sound. Furthermore, the gain of the direct response is proportional to 1 / d. If these rules are conserved in handling the direct responses of channels with different source distances, subsystem 201 can implement a straight downmix of all channels, because the delay and level of late reverberation are generally not sensitive to the source location.
[0080] For practical reasons, the virtualizer (e.g., subsystem 100 of the virtualizer in Figure 3) may be implemented to time-align the direct responses for input channels with different source distances. To preserve the relative delay between the direct response and the late reverberation for each channel, channels with source distance d are downmixed with other channels by (dmax-d) / v s It should be delayed by only that much. Here, dmax represents the maximum possible source distance.
[0081] The virtualizer (for example, the virtualizer subsystem 100 in Figure 3) may also be implemented to compress the dynamic range of the direct response. For example, the direct response for a channel with source distance d is d -1 Instead of factor d -α It may be scaled by , where 0 ≤ α ≤ 1. To preserve the level difference between the direct response and the late reverberation, the downmix subsystem 201 downmixes a channel with source distance d with other scaled channels by factor d 1-α It may be necessary to implement scaling by [a specific method / system].
[0082] The feedback delay network in Figure 4 is an exemplary implementation of the FDN 203 (or 204 or 205) in Figure 3. The system in Figure 4 consists of four reverberation tanks (each with a gain stage g i and delay line z -ni While it includes, variations of this system (and other FDNs used in embodiments of the virtualizer of the present invention) implement more than four or fewer than four reverberation tanks.
[0083] The FDN in Figure 4 includes an input gain element 300, an all-pass filter (APF) 301 coupled to the output of element 300, summing elements 302, 303, 304, and 305 coupled to the output of APF 301, and four reverberation tanks coupled to the outputs of different elements 302, 303, 304, and 305 respectively (each reverberation tank has a gain element g k (One of the 306 elements) and the delay line z connected to it -Mk (One of the 307 elements) and the gain element 1 / g coupled to it. k (One of the elements 309) has (0 ≤ k-1 ≤ 3). A unitary matrix 308 is coupled to the output of the delay line 307 and is configured to provide a feedback output to the second inputs of elements 302, 303, 304, and 305, respectively. The outputs of two of the gain elements 309 (the first and second reverberation tanks) are presented to the input of the summing element 310, and the output of element 310 is presented to one input of the output mixing matrix 312. The outputs of the other two of the gain elements 309 (the third and fourth reverberation tanks) are presented to the input of the summing element 311, and the output of element 311 is presented to the other input of the output mixing matrix 312.
[0084] Element 302 is the delay line z -n1 The output of matrix 308 corresponding to this is applied to the input of the first reverberation tank (i.e., the delay line z via matrix 308). -n1 It is configured to apply feedback from the output of Element 303, which is a delay line z -n2The output of matrix 308 corresponding to this is applied to the input of the second reverberation tank (i.e., the delay line z via matrix 308). -n2 It is configured to apply feedback from the output of Element 304, which is a delay line z -n3 The output of matrix 308 corresponding to is applied to the input of the third reverberation tank (i.e., the delay line z via matrix 308). -n3 It is configured to apply feedback from the output of Element 305, which is a delay line z -n4 The output of matrix 308 corresponding to this is applied to the input of the fourth reverberation tank (i.e., the delay line z via matrix 308). -n4 It is configured to apply feedback from the output of [the system].
[0085] The input gain element 300 of the FDN in Figure 4 is coupled to receive one frequency band of the converted monophonic downmix signal (filter bank region signal) output from the decomposed filter bank 202 in Figure 3. The input gain element 300 applies a gain (scaling) factor G to the filter bank region signal presented to it. in The scaling factor G is applied collectively to all frequency bands (implemented by all FDNs 203, 204, ..., 205 in Figure 3). in This controls the spectral shaping and level of the late reverberation. The input gain G in all FDNs of the virtualizer in Figure 3. in Setting these often involves considering the following goals: The direct-to-late ratio (DLR) of BRIR applied to each channel to match the actual room; Necessary low-frequency attenuation to mitigate excessive combing artifacts and / or low-frequency rumbling; Matching of diffusion field spectral envelopes.
[0086] Assuming that the direct response (as applied by subsystem 100 in Figure 3) provides unitary gain across all frequency bands, the specific DLR (power ratio) is: G in=sqrt(ln(10 6 ) / (T60*DLR)) G in This can be achieved by setting the following: Here, T60 is the reverberation decay time, defined as the time it takes for the reverberation to decay by 60 dB (which is determined by the reverberation delay and reverberation gain discussed below), and "ln" represents the natural logarithm function.
[0087] Input gain factor G in This may depend on the content being processed. One application of such content dependency is to ensure that the downmix energy in each time / frequency segment is equal to the sum of the energies of the individual channel signals being downmixed, regardless of any correlations present between the input channel signals. In this case, the input gain factor is
number
[0088] In the typical QMF domain implementation of the FDN shown in Figure 4, the signal presented from the output of the whole-pass filter (APF) 301 to the input of the reverberation tank is a sequence of QMF domain frequency components. To produce a more natural-sounding FDN output, the APF 301 is applied to the output of the gain element 300 to introduce phase diversity and increased echo density. Alternatively or additionally, one or more whole-pass filters may be applied to individual inputs to the downmix subsystem 201 (in Figure 3) before the inputs are downmixed in subsystem 201 and processed by the FDN, or in the reverberation tank feedforward or feedback path depicted in Figure 4 (for example, the delay line z in each reverberation tank). -Mk It may be applied in addition to or instead of to the output of FDN (i.e., to the output of output matrix 312).
[0089] Reverberation tank delay z -ni When implementing this, to avoid the reverberation modes aligning at the same frequency, a reverberation delay of n i The elements should be relatively prime. The sum of the delays should be large enough to provide sufficient mode density to avoid an artificially sounding output. However, the shortest delay should be short enough to avoid an excessive time gap between the late reverberation and the other components of the BRIR.
[0090] Typically, the reverberation tank output is initially panned to either the left or right binaural channel. Usually, there are equal numbers of reverberation tank outputs panned to the two binaural channels, and these are mutually exclusive. It is also desirable to balance the timing of the two binaural channels. Therefore, if the reverberation tank output with the shortest delay goes to one binaural channel, the reverberation tank output with the second shortest delay will go to the other channel.
[0091] The reverberation tank delay can vary across frequency bands to change the mode density as a function of frequency. Generally, lower frequency bands require a higher mode density and therefore a longer reverberation tank delay.
[0092] Reverberation tank gain g i The amplitude and reverberation tank delay, combined, determine the reverberation delay time of the FDN in Figure 4: T 60 = -3n i / log 10 (|g i |) / F FRM Here, F FRM This is the frame rate of filter bank 202 (in Figure 3). The phase of the reverberation tank gain is modified by introducing a fractional delay to overcome problems related to the reverberation tank delay being quantized on the downsample factor grid of the filter bank.
[0093] The unitary feedback matrix 308 provides an even mixture among the reverberation tanks in the feedback path.
[0094] To equalize the level of the reverberation tank output, the gain element 309 has a normalized gain of 1 / |g i Apply | to the output of each reverberation tank to remove the level effect of the reverberation tank gain while preserving the fractional delay introduced by its phase.
[0095] Output mixed matrix 312 (matrix M) out The output mixing matrix 312 is a 2x2 matrix configured to mix the unmixed binaural channels (outputs of elements 310 and 311, respectively) from the initial panning to achieve left and right binaural channels (L and R signals presented in the output of matrix 312) with the desired interaural coherence. The unmixed binaural channels are almost uncorrelated after the initial panning, as they contain no common reverberation tank output. If the desired interaural coherence is Coh and |Coh| ≤ 1, then the output mixing matrix 312 is
number
number
[0096] For each individual frequency band in the virtualizer of the present invention, if the target acoustic attributes T60, Coh, and DLR defined above are known, each FDN (each FDN may have the structure shown in Figure 4) can be configured to achieve the target attributes. In particular, in some embodiments, the input gain (G) for each FDN can be configured to achieve the target attributes according to the relationships described herein. in ) and the gain and delay of the reverberation tank (g i and n i ) and output matrix M out The parameters can be set (for example, by the control values applied to them by the control subsystem 209 in Figure 3). In practice, it is often sufficient to set frequency-dependent attributes by a model with simple control parameters in order to generate natural-sounding late reverberations that match a particular acoustic environment.
[0097] Next, the target reverberation decay time (T) for FDN for each specific frequency band in one embodiment of the virtualizer of the present invention 60 ) for each of a few frequency bands, the target reverberation decay time (T 60 An example of how this can be determined is given by determining the FDN response level. The FDN response level decays exponentially over time. 60 It is inversely proportional to the decay factor df (defined as dB decay per unit time), i.e.: T 60 =60 / df That is the case.
[0098] The attenuation factor df is frequency-dependent and generally increases linearly with respect to the logarithmic frequency scale. Therefore, the reverberation decay time is also a function of frequency and generally decreases as frequency increases. Accordingly, for two frequency points, T 60 if the value of is determined (for example, set), then T for all frequencies 60 the curve is determined. For example, at frequency points f A and f B the reverberation decay times for are respectively T 60,A and T 60,B if so, then T 60 the curve is defined as follows.
[0099] [Mathematical formula] Figure 5 shows T for two specific frequencies (f A and f B respectively), T at 60 value is f A = 10 Hz, T 60,A = 320 ms and f B = 2.4 kHz, T 60,B = 150 ms, an example of a T 60 curve that can be achieved by an embodiment of the virtualizer of the present invention is shown.
[0100] Next, an example will be described of how the target interaural coherence (Coh) for the FDN for each specific frequency band in an embodiment of the virtualizer of the present invention can be achieved by setting a small number of control parameters. The interaural coherence (Coh) of late reverberation generally follows the pattern of a diffuse sound field. It can be modeled by a sinc function up to the crossover frequency f C and a constant above the crossover frequency. A simple model for the Coh curve is as follows.
[0101] [Mathematical formula] Here, the parameter Cohmin and Coh max satisfy -1≦Coh min <Coh max ≦1, and controls the range of Coh. The optimal crossover frequency f C depends on the head size of the listener. An excessively high f C leads to a sound source image localized inside the head, while an excessively low f C leads to a diffused or split sound source image. FIG. 6 shows the control parameter Coh max , Coh min and f C having the following values: Coh max =0.95, Coh min =0.05 and f C is an example of a Coh curve that can be achieved by an embodiment of the present invention set to have =700 Hz.
[0102] Next, an example is described of how a target direct-to-late ratio (DLR) for an FDN for each specific frequency band in an embodiment of the virtualizer of the present invention can be achieved by setting a small number of control parameters. The direct-to-late ratio (DLR) in dB generally increases linearly with respect to logarithmic frequency, and is controlled by setting DLR 1K (DLR in dB at 1 kHz) and DLRslope (in dB per decade of frequency). However, low DLR in the low frequency range often leads to excessive combing artifacts. To mitigate the artifacts, two correction mechanisms for controlling DLR are added: a minimum DLR floor, DLRmin (in dB); and a transition frequency f T and the slope of the attenuation curve below it HPF slope a high-pass filter defined by (in dB per decade of frequency).
[0103] The resulting DLR curve, in dB, is defined as follows.
[0104]
Mathematical Expression
[0105] Modifications of the embodiments disclosed in this paper have one or more of the following characteristics: The virtualizer of the present invention is implemented in the time domain, or has a hybrid implementation with FDN-based impulse response capture and FIR-based signal filtering; The virtualizer of the present invention is implemented to allow the application of energy compensation as a function of frequency during the execution of a downmixing stage that generates a downmixed input signal for a late reverberation processing subsystem; The virtualizer of the present invention is implemented to allow manual or automatic control of late reverberation attributes applied in response to external factors (i.e., in response to the setting of control parameters).
[0106] For applications where system latency is critical and delays caused by decomposition and synthesis filter banks are unacceptable, the filter bank region FDN structure of a typical embodiment of the virtualizer of the present invention can be converted to the time domain, and each FDN structure can be implemented in the time domain in some classes of embodiments of the virtualizer. In the time domain implementation, the input gain factor (Gin ), reverberation tank gain (g i ) and normalized gain (1 / |g i The subsystem to which |) is applied is replaced by a filter with a similar amplitude response to allow frequency-dependent control. Output mixing matrix (M out ) is also replaced by the filter matrix. Unlike other filters, the phase response of this filter matrix is crucial, as it can affect power conservation and biaural coherence. The reverberation tank delay in the time-domain implementation may need to be slightly altered (from its value in the filter bank domain implementation) to avoid sharing the filter bank stride as a common factor. Due to various constraints, the execution of the time-domain implementation of the FDN of the virtualizer of the present invention may not exactly match that of its filter bank implementation.
[0107] Referring to Figure 8, a hybrid (filter bank domain and time domain) implementation of the late reverberation processing subsystem of the present invention for the virtualizer of the present invention is described next. This hybrid implementation of the late reverberation processing subsystem of the present invention is a variation of the late reverberation processing subsystem 200 of Figure 4 and implements impulse response capture based on FDN and signal filtering based on FIR.
[0108] Figure 8 includes elements 201, 202, 203, 204, 205, and 207, which are identical to the same referenced elements in subsystem 200 of Figure 3. The above descriptions of these elements are not repeated in the reference to Figure 8. In the embodiment of Figure 8, a unit impulse generator 211 is coupled to produce an input signal (pulse) to the decomposed filter bank 202. An LBRIR filter 208 (mono input, stereo output), implemented as an FIR filter, applies the appropriate late reverberation portion of the BRIR (LBRIR) to the monophonic downmix output from subsystem 201. Thus, elements 211, 202, 203, 204, 205, and 207 are processing sidechains for the LBRIR filter 208.
[0109] Whenever the setting of the late reverberation portion LBRIR is modified, the impulse generator 211 is made to produce a unit impulse to element 202, the resulting output from filter bank 207 is captured and presented to filter 208 (to set filter 208 to apply the new LBRIR determined by the output of filter bank 207). To accelerate the time elapsed between the LBRIR setting change and the time when the new LBRIR becomes effective, samples of the new LBRIR can begin replacing the old LBRIR as they become available. To reduce the intrinsic latency of the FDN, the initial zeros of the LBRIR can be discarded. These options provide flexibility and allow the hybrid implementation to provide potential performance improvements (compared to the performance provided by the filter bank region implementation) at the cost of the additional computations from the FIR filtering.
[0110] For applications where system latency is critical but computational power is not a major concern, a sidechain filter bank region late reverberation processor (for example, implemented by elements 211, 202, 203, 204, ..., 205 in Figure 8) can be used to capture the effective FIR impulse response applied by filter 208. FIR filter 208 implements this captured FIR response and can be directly applied to the mono downmix of the input channel (during the virtualization of the input channel).
[0111] Various FDN parameters, and thus the resulting late reverberation attributes, can be manually tuned and then incorporated as a fixed configuration into embodiments of the late reverberation processing subsystem of the present invention. For example, by one or more presets that can be adjusted by the user of the system (for example, by operating the control subsystem 209 in Figure 3). However, given a high level of description of late reverberation, its relationship with FDN parameters, and the ability to modify its behavior, a wide variety of methods can be conceived for controlling various embodiments of FDN-based late reverberation processors. These include, but are not limited to, the following:
[0112] 1. The end user may manually control the FDN parameters by a user interface on a display (for example, implemented by the embodiment of the control subsystem 209 in Figure 3), or by switching presets using physical controls (for example, implemented by the embodiment of the control subsystem 209 in Figure 3). In this way, the end user can adapt the room simulation according to their preferences, environment, or content.
[0113] 2. The author of the audio content to be virtualized may provide settings or desired parameters that are transmitted along with the content itself, for example, by metadata provided with the input audio signal. Such metadata may be parsed and used to control the relevant FDN parameters (for example, by the embodiment of the control subsystem 209 in Figure 3). Thus, the metadata may indicate attributes such as reverberation time, reverberation level, direct-to-reverberation ratio, etc., and these attributes may change over time and be indicated by time-varying metadata.
[0114] 3. The playback device may recognize its location or environment by one or more sensors. For example, a mobile device may use a GSM network, Global Positioning System (GPS), a known WiFi access point, or any other location service to determine where it is located. The location and / or environment data may then be used to control the relevant FDN parameters (for example, by the embodiment of the control subsystem 209 in Figure 3). Thus, the FDN parameters may be modified in response to the device's location to mimic, for example, the physical environment.
[0115] 4. Cloud services or social media may be used to derive the most common settings used by consumers in certain environments, in relation to the location of the playback device. Furthermore, users may upload their current settings, associated with their (known) location, to cloud or social media services to make them available to other users or to themselves.
[0116] 5. The playback device may include other sensors such as cameras, light sensors, microphones, accelerometers, and gyroscopes to determine the user's activity and the environment in which the user is located, in order to optimize the FDN parameters for that particular activity and / or environment.
[0117] 6. FDN parameters may be controlled by audio content. Audio classification algorithms or manually annotated content may indicate whether audio segments include speech, music, sound effects, silence, etc. FDN parameters may be adjusted according to such labels. For example, the direct-to-reverberation ratio may be reduced for dialogue to improve dialogue intelligibility. Furthermore, video analysis may be used to determine the position of the current video segment, and FDN parameters may be adjusted accordingly to better simulate the environment depicted in the video. and / or 7. The semiconductor playback system may use different FDN settings than the mobile device. For example, the settings may be device-dependent. A semiconductor system in a living room may simulate a typical living room scenario with a distant source (with considerable reverberation), while the mobile device may render the content closer to the listener.
[0118] Some implementations of the virtualizer of the present invention include FDNs (e.g., the FDN implementation in Figure 4) configured to apply fractional delays in addition to integer sample delays. For example, in one such implementation, fractional delay elements are connected in series with delay lines that add integer delays equal to an integer number of sample periods within each reverberation tank (e.g., each fractional delay element is located after one of the delay lines or in other ways in series with it). The fractional delay can be approximated in each frequency band by a phase shift (unit complex multiplication) corresponding to a certain percentage of the sample period f = τ / T, where f is the delay fraction, τ is the desired delay for that band, and T is the sample period for that band. How fractional delays are added in the context of applying reverberation in the QMF domain is well known.
[0119] In a first-class embodiment, the present invention is a headphone virtualization method that generates a binaural signal in response to a set of channels of a multi-channel audio input signal (for example, each of those channels or each of the channels across the entire frequency range). The method includes: (a) applying a binaural chamber impulse response (BRIR) to each channel of the set (for example, in subsystems 100 and 200 in Figure 3 or by convolving each channel of the set with the BRIR corresponding to the channel in subsystems 12, ..., 14, 15 in Figure 2) to generate a filtered signal (for example, the outputs of subsystems 100 and 200 in Figure 3 or the outputs of subsystems 12, ..., 14, 15 in Figure 2), which includes using at least one feedback delay network (for example, FDNs 203, 204, ..., 205 in Figure 3) to add a common late reverberation to a downmix of the channels of the set (for example, a monophonic downmix); and (b) combining the filtered signals (for example, in subsystem 210 in Figure 3 or a subsystem including elements 16 and 18 in Figure 2) to generate a binaural signal. Typically, a bank of FDNs is used to add the common late reverberation to the downmix (for example, each FDN adds late reverberation to a different frequency band). Typically, step (a) includes applying the “direct response and early reflection” portion of a single-channel BRIR for each channel of the set (for example, in subsystem 100 in Figure 3 or subsystems 12, ..., 14 in Figure 2), and the common late reverberation is generated to emulate the collective macro-attributes of at least some (e.g., all) of the late reverberation portions of the single-channel BRIR.
[0120] In a typical implementation of the first class, each FDN is implemented in the hybrid complex quadrature mirror filter (HCQMF) domain or the quadrature mirror filter (QMF) domain. In some such embodiments, the frequency-dependent spatial acoustic attributes of the binaural signal are controlled by controlling the configuration of each FDN used to add late reverberation (e.g., using the control subsystem 209 in Figure 3). Typically, for efficient binaural rendering of audio content of a multi-channel signal, a monophonic downmix of the channels (e.g., the downmix generated by subsystem 201 in Figure 3) is used as input to the FDN. Typically, the downmixing process is controlled based on the source distance for each channel (i.e., the distance between the assumed source of the audio content of the channel and the assumed user position) and relies on the handling of the direct response corresponding to the source distance to preserve the temporal and level structure of each BRIR (i.e., each BRIR determined by the direct response and early reflection portion of a single-channel BRIR for a given channel, as well as a common late reverberation for the downmix containing that channel). The channels to be downmixed can be time-aligned and scaled in various ways during the downmix, but the appropriate level and temporal relationships between the direct BRIR response, early reflections, and common late reverberation portion for each channel should be maintained. In embodiments that use a single FDN bank to generate a common late reverberation portion for all channels being downmixed (to generate the downmix), appropriate gain and delay must be applied (for each channel being downmixed) during the downmix generation.
[0121] A typical embodiment of this class includes a step of adjusting FDN coefficients corresponding to frequency-dependent attributes (e.g., reverberation decay time, interaural coherence, mode density, and direct-to-late ratio). This allows for better matching of the acoustic environment and a more natural-sounding output.
[0122] In a second class embodiment, the present invention is a method for generating a binaural signal in response to a multi-channel audio input signal. This is achieved by applying a binaural chamber impulse response (BRIR) to each channel of a set of channels in the input signal (e.g., each of the channels in the input signal or the entire frequency range channels of each of the input signals) (e.g., by convolving each channel with the corresponding BRIR). This includes processing each channel of the set in a first processing path (e.g., implemented by subsystem 100 in Figure 3 or subsystems 12, ..., 14 in Figure 2) configured to model and apply to each channel a direct response and early reflection of a single-channel BRIR for that channel (e.g., EBRIR applied by subsystems 12, 14 or 15 in Figure 2), and processing a downmix of the channels of the set (e.g., a monophonic downmix) in a second processing path (e.g., implemented by subsystem 200 in Figure 3 or subsystem 15 in Figure 2) in parallel with the first processing path. The second processing path is configured to model a common late reverberation (e.g., the LBRIR applied by subsystem 15 in Figure 2) and apply it to the downmix. Typically, the common late reverberation emulates the collective macro attributes of at least some (e.g., all) of the late reverberation portions of the single-channel BRIR. Typically, the second processing path includes at least one FDN (e.g., one FDN for each of several frequency bands). Typically, a mono downmix is used as the input to all reverberation tanks for each FDN implemented by the second processing path. Typically, a mechanism is provided for systematic control of the macro attributes of each FDN (e.g., control subsystem 209 in Figure 3) to better simulate the acoustic environment and produce a more natural-sounding binaural virtualization. Since most such macro attributes are frequency-dependent, each FDN is typically implemented in a hybrid complex quadrature mirror filter (HCQMF) domain, frequency domain, domain, or another filter bank domain, with a different FDN used for each frequency band.A major benefit of implementing FDN in the filter bank domain is that it allows for the application of reverberation with frequency-dependent reverberation attributes. In various embodiments, FDN is implemented in any of the broad and diverse filter bank domains using any of the diverse filter banks. This includes, but is not limited to, quadrature mirror filters (QMFs), finite impulse response filters (FIR filters), infinite impulse response filters (IIR filters), or crossover filters.
[0123] Some embodiments of the first class (and the second class) implement one or more of the following features:
[0124] 1. An FDN implementation in the filter bank region (e.g., the hybrid complex quadrature mirror filter region) (e.g., the FDN implementation in Figure 4) or an FDN implementation in the hybrid filter bank region and a time-domain late reverberation filter implementation (e.g., the structure described with reference to Figure 8). This typically allows independent adjustment of the FDN parameters and / or settings for each frequency band (which enables simple and flexible control of frequency-dependent acoustic attributes). This is, for example, by providing the ability to vary the reverberation tank delay in various bands so that the mode density changes as a function of frequency.
[0125] 2. The specific downmixing process used to generate the downmixed (e.g., monophonically downmixed) signal processed in the second processing path (from a multi-channel input audio signal) depends on the source distance of each channel and the handling of the direct response to maintain the appropriate level and timing relationship between the direct and late responses.
[0126] 3. To introduce phase diversity and increased echo density without altering the resulting reverberation spectrum and / or timbre, a full-pass filter (e.g., APF 301 in Figure 4) is applied in a second processing path (e.g., at the input or output of the FDN bank).
[0127] 4. To overcome the problems related to quantized delays in the downsample-factor grid, fractional delays are implemented in the feedback paths of each FDN in a complex-valued multirate structure.
[0128] 5. In FDN, the reverberation tank output is directly and linearly mixed into the binaural channel (for example, by matrix 312 in Figure 4) using an output mixing coefficient set based on the desired interaural coherence in each frequency band. Optionally, the mapping of the reverberation tank to the binaural output channel alternates across frequency bands to achieve balanced delays between the binaural channels. Optionally, a normalization factor is also applied to the reverberation tank output to equalize its level while preserving fractional delays and overall power.
[0129] 6. Frequency-dependent reverberation decay time is controlled by setting the appropriate combination of reverberation tank delay and gain in each frequency band to simulate a real room.
[0130] 7. One scaling factor is applied for each frequency band (for example, at either the input or output of the relevant processing path) (for example, by elements 306 and 309 in Figure 4). This results in: Control the frequency-dependent direct-to-late ratio (DLR) to match the actual room's DLR (a simple model may be used to calculate the required scaling factor based on the target DLR and reverberation decay time, e.g., T60); Provides low-frequency attenuation to mitigate excessive combing artifacts; and / or Diffuse field spectral shaping is applied to the FDN response.
[0131] 8. A simple parametric model is implemented (for example, by the control subsystem 209 in Figure 3) to control the intrinsic frequency-dependent attributes of late reverberation, such as reverberation decay time, biaural coherence, and / or direct-to-late ratio.
[0132] In some embodiments (for example, for applications where system latency is deterministic and delays caused by decomposition and synthesis filter banks are prohibited), the filter bank domain FDN structure of a typical embodiment of the system of the present invention (e.g., the FDN in Figure 4 for each frequency band) is replaced by a time-domain implemented FDN structure (e.g., the FDN 220 in Figure 10, which may be implemented as shown in Figure 9). In the time-domain embodiment of the system of the present invention, the input gain factor (G in ), reverberation tank gain (g i ) and normalized gain (1 / |g i In a filter bank domain embodiment to which |) is applied, the subsystem is replaced by a time domain filter (and / or gain element) to allow frequency-dependent control. The output mixing matrix of a typical filter bank domain implementation (e.g., output mixing matrix 312 in Figure 4) is replaced (in a typical time domain embodiment) by the output set of a time domain filter (e.g., elements 500-503 of the implementation in Figure 11, element 424 in Figure 9). Unlike other filters in a typical time domain embodiment, the phase response of this output set of the filter is typically critical (because the phase response can affect power conservation and biaural coherence). In some time domain embodiments, the reverberation tank delay is varied (e.g., slightly) from its value in the corresponding filter bank domain implementation (e.g., to avoid sharing the filter bank stride as a common factor).
[0133] Figure 10 is a block diagram of an embodiment of the headphone virtualization system of the present invention similar to Figure 3, except that elements 202-207 in Figure 3 are replaced in the system of Figure 10 by a single FDN 220 implemented in the time domain (for example, the FDN 220 in Figure 10 may be implemented similarly to the FDN in Figure 9). In Figure 10, two time-domain signals (left and right channels) are output from the direct response and early reflection processing subsystem 100, and two time-domain signals (left and right channels) are output from the late reverberation processing subsystem 221. Adding elements 210 are coupled to the outputs of subsystems 100 and 200. Element 210 is configured to combine (mix) the left channel outputs of subsystems 100 and 221 to generate the left channel L of the binaural audio signal output from the virtualizer in Figure 10, and to combine (mix) the right channel outputs of subsystems 100 and 221 to generate the right channel R of the binaural audio signal output from the virtualizer in Figure 10. Assuming that appropriate level adjustment and time alignment are implemented in subsystems 100 and 221, element 210 can be implemented to generate the left channel of the binaural output signal by simply summing the corresponding left channel samples output from subsystems 100 and 221, and to generate the right channel of the binaural output signal by simply summing the corresponding right channel samples output from subsystems 100 and 221.
[0134] In the system shown in Figure 10, (Channel X i The multi-channel audio input signal (having X) is directed to two parallel processing paths where it is processed. One passes through the direct response and early reflection processing subsystem 100, and the other passes through the late reverberation processing subsystem 221. The system in Figure 10 has each channel X i BRIR i It is configured to apply each BRIR. iThe signal can be broken down into two parts: a direct response and early reflection portion (applied by subsystem 100) and a late reverberation portion (applied by subsystem 221). In operation, the direct response and early reflection processing subsystem 100 thus generates the direct response and early reflection portion of the binaural audio signal output from the virtualizer, and the late reverberation processing subsystem ("late reverberation generator") 221 thus generates the late reverberation portion of the binaural audio signal output from the virtualizer. The outputs of subsystems 100 and 221 are mixed (by subsystem 210) to generate a binaural audio signal, which is typically presented from subsystem 210 to a rendering system (not shown), where it undergoes binaural rendering for playback by headphones.
[0135] The downmix subsystem 201 (of the late reverberation processing subsystem 221) is configured to downmix the channels of the multi-channel input signal to a mono downmix (which is a time-domain signal), and the FDN 220 is configured to apply the late reverberation portion to the mono downmix.
[0136] Referring to Figure 9, we now describe an example of a time-domain FDN that can be used as FDN 220 of the virtualizer in Figure 10. The FDN in Figure 9 includes an input filter 400 coupled to receive a mono downmix of all channels of a multi-channel audio input signal (e.g., generated by subsystem 201 of the system in Figure 10). The FDN in Figure 9 includes a full-pass filter (APF) 401 coupled to the output of filter 400 (corresponding to APF 301 in Figure 4), an input gain element 401A coupled to the output of filter 401, summing elements 402, 403, 404 and 405 coupled to the output of element 401A (these correspond to summing elements 302, 303, 304 and 305 in Figure 4), and four reverberation tanks. Each reverberation tank has a reverberation filter 406 and 406A, 407 and 407A, 408 and 408A and 409 and 409A coupled to the output of different elements 402, 403, 404 and 405, one of the reverberation filters 406 and 406A, 407 and 407A, one of the delay lines 410, 411, 412 and 413 coupled thereto (corresponding to delay line 307 in Figure 4), and one of the gain elements 417, 418, 419 and 420 coupled to the output of one of these delay lines.
[0137] A unitary matrix 415 (corresponding to the unitary matrix 308 in Figure 4, and typically implemented to be identical to matrix 308) is coupled to the outputs of delay lines 410, 411, 412, and 413. Matrix 415 is configured to provide a feedback output to the second inputs of elements 402, 403, 404, and 405, respectively.
[0138] When the delay (n1) added by line 410 is shorter than the delay (n2) added by line 411, the delay added by line 411 is shorter than the delay (n3) added by line 412, and the delay added by line 412 is shorter than the delay (n4) added by line 413, the outputs of gain elements 417 and 419 (of the first and third reverberation tanks) are presented to the input of summing element 422, and the outputs of gain elements 418 and 420 (of the second and fourth reverberation tanks) are presented to the input of summing element 423. The output of element 422 is presented to one input of the IACC and mixing filter 424, and the output of element 423 is presented to the other input of the IACC filtering and mixing stage 424.
[0139] An example of the implementation of gain elements 417-420 and elements 422, 423, and 424 in Figure 9 will be described with reference to a typical implementation of elements 310 and 311 and the output mixing matrix 312 in Figure 4. Figure 4 Output Mixing Matrix 312 (Matrix M out The matrix (also identified as the matrix 312) is a 2x2 matrix configured to mix the unmixed binaural channels from the initial panning (the outputs of elements 310 and 311, respectively) to generate left and right binaural output channels with the desired interaural coherence (left ear "L" and right ear "R" signals presented at the output of matrix 312). This initial panning is implemented by elements 310 and 311, each combining two reverberation tank outputs to generate one of the unmixed binaural channels, with the reverberation tank output with the shortest delay being presented to the input of element 310 and the reverberation tank output with the second shortest delay being presented to the input of element 311. Elements 422 and 423 in the embodiment of Figure 9 perform the same type of initial panning as elements 310 and 311 in the embodiment of Figure 4 (in each frequency band) perform on the stream of filter bank domain components (in the relevant frequency bands) presented to their inputs (for the time-domain signals presented to their inputs).
[0140] Since they do not contain any common reverberation tank output, the unmixed binaural channels (output from elements 310 and 311 in Figure 4 or elements 422 and 423 in Figure 9), which are almost uncorrelated, may be mixed (by matrix 312 in Figure 4 or row 424 in Figure 9) to implement a panning pattern that achieves the desired interaural coherence for the left and right binaural output channels. However, since the reverberation tank delays differ in each FDN (i.e., the FDN in Figure 9 or the FDNs implemented for each different frequency band in Figure 4), one unmixed binaural channel (one output from elements 310 and 311 or 422 and 423) is always ahead of the other unmixed binaural channel (the other output from elements 310 and 311 or 422 and 423).
[0141] Thus, in the embodiment of Figure 4, if the combination of reverberation tank delay and panning pattern is identical across all frequency bands, an image bias will result. This bias can be mitigated if the panning pattern is alternating across frequency bands so that the mixed binaural output channels lead and lag each other in alternating frequency bands. For example, if the desired interaural coherence is Coh and |Coh| ≤ 1, then the output mixing matrix 312 in the odd-numbered frequency bands takes the two inputs presented to it in the following form
number
number
[0142] Alternatively, the above-mentioned sound image bias in the binaural output channel can be mitigated by implementing matrix 312 to be identical in FDN for all frequency bands, provided that the channel order of its inputs is switched for alternating frequency bands (for example, in odd frequency bands, the output of element 310 may be presented to the first input of matrix 312 and the output of element 311 to the second input of matrix 312; in even frequency bands, the output of element 311 may be presented to the first input of matrix 312 and the output of element 310 to the second input of matrix 312).
[0143] In the embodiment of Figure 9 (and other time-domain embodiments of the FDN of the present invention), it is not trivial to alternate the panning based on frequency to address the image bias that would normally result when the unmixed binaural channel output from element 422 is always ahead (lagging) the unmixed binaural channel output from element 423. This image bias is addressed in a different way in a typical time-domain embodiment of the FDN of the present invention than is typically addressed in the filter bank domain embodiment of the FDN of the present invention. In particular, in the embodiment of Figure 9 (and other time-domain embodiments of the FDN of the present invention), the relative gain of the unmixed binaural channels (e.g., outputs from elements 422 and 423 in Figure 9) is determined by the gain elements (e.g., elements 417, 418, 419 and 420 in Figure 9) to compensate for the image bias that would normally result from the aforementioned unbalanced timing. The stereo image is recentered by implementing a gain element (e.g., element 417) to attenuate the earliest arriving signal (which is, for example, panned to one side by element 422), and another gain element (e.g., element 418) to boost the next earliest signal (which is, for example, panned to the other side by element 423). Thus, a reverberation tank containing gain element 417 applies a first gain to the output of element 417, and a reverberation tank containing gain element 418 applies a second gain (different from the first gain) to the output of element 418. The first and second gains then attenuate the first unmixed binaural channel (output from element 422) with respect to the second unmixed binaural channel (output from element 423).
[0144] More specifically, in the exemplary implementation of the FDN of FIG. 9, the four delay lines 410, 411, 412, and 413 have sequentially increasing lengths, and have sequentially increasing delay values n1, n2, n3, and n4, respectively. In this implementation, filter 417 applies a gain of g1. Thus, the output of filter 417 is a delayed version of the input to delay line 410, to which the gain of g1 has been applied. Similarly, filter 418 applies a gain of g2, filter 419 applies a gain of g3, and filter 420 applies a gain of g4. Thus, the output of filter 418 is a delayed version of the input to delay line 411, to which the gain of g2 has been applied, the output of filter 419 is a delayed version of the input to delay line 412, to which the gain of g3 has been applied, and the output of filter 420 is a delayed version of the input to delay line 413, to which the gain of g4 has been applied.
[0145] In this implementation, selecting the following gain values: g1=0.5, g2=0.5, g3=0.5, g4=0.5 may lead to undesirable bias on one side (i.e., to the left or right channel) of the output sound image represented by the binaural channel output from element 424. According to an embodiment of the present invention, the values g1, g2, g3, g4 (applied by elements 417, 418, 419, and 420, respectively) are selected as follows to center the sound image: g1=0.38, g2=0.6, g3=0.5, g4=0.5. Thus, according to an embodiment of the present invention, the output stereo image attenuates the earliest arriving signal, which is panned to one side by element 422 in the present example, relative to the second latest arriving signal (i.e., selecting g1 such that g1<g3), and boosts the second earliest signal, which is panned to the other side by element 423 in the present example, relative to the latest arriving signal (i.e., selecting g4 such that g4<g2), thereby recentering the output stereo image.
[0146] The exemplary implementation of the time-domain FDN in FIG. 9 has the following differences and similarities relative to the filter bank domain (CQMF domain) FDN of FIG. 4.
[0147] The same unitary feedback matrix A (matrix 308 in Figure 4 and matrix 415 in Figure 9).
[0148] Similar reverberation tank delay n i (That is, the delay in the CQMF implementation in Figure 4 is 1 / T) s Assuming that is the sampling rate (1 / T s (typically equal to 48KHz), n1 = 17 * 64T s =1088*T s n² = 21 * 64T s =1344*T s n3 = 26 * 64T s =1664*T s n4 = 29 * 64T s =1856*T s This may also be the case, while the delay in the time-domain implementation is n1 = 1089 * T s n2 = 1345 * T s n3 = 1663 * T s n4 = 185*T s This is also acceptable. Note that while a typical CQMF implementation imposes the practical constraint that each delay is some integer multiple of the duration of a 64-sample block, there is greater flexibility in the time domain regarding the choice of each delay, and therefore greater flexibility in the choice of delays for each reverberation tank.
[0149] Similar pass-through filter implementations (i.e., similar implementations of filter 301 in Figure 4 and filter 401 in Figure 9). For example, a pass-through filter can be implemented by a cascade of several (e.g., three) pass-through filters. For example, each cascaded pass-through filter, assuming g=0.6,
number
[0150] In some implementations of the time-domain FDN in Figure 9, the input filter 400 is implemented to match (at least substantially) the direct-to-late ratio (DLR) of the BRIR applied by the system in Figure 9 to a target DLR, and to allow the DLR of the BRIR applied by a virtualizer including the system in Figure 9 (e.g., the virtualizer in Figure 10) to be modified by replacing (or controlling the configuration settings of) the filter 400. For example, in some embodiments, the filter 400 is implemented as a cascade of filters (e.g., a first filter 400A and a second filter 400B coupled as shown in Figure 9A) that implement the target DLR and optionally implement desired DLR control. For example, the filters in the cascade are IIR filters (e.g., filter 400A is a first-order Butterworth high-pass filter (IIR filter) configured to match the target low-frequency characteristics, and filter 400B is a second-order low-shelf IIR filter configured to match the target high-frequency characteristics). Another example of this cascaded filter is an IIR and FIR filter (for example, filter 400A is a second-order Butterworth high-pass filter (IIR filter) configured to match the target low-frequency characteristics, and filter 400B is a 14th-order FIR filter configured to match the target high-frequency characteristics). Typically, the direct signal is fixed, and filter 400 modifies the later signal to achieve the target DLR. The full-pass filter (APF) 401 is preferably implemented to perform the same function as APF 301 in Figure 4, i.e., to introduce phase diversity and increased echo density to produce a more natural-sounding FDN output. The input filter 400 controls the amplitude response, while APF 401 typically controls the phase response.
[0151] In Figure 9, filter 406 and gain element 406A together implement a reverberation filter, filter 407 and gain element 407A together implement another reverberation filter, filter 408 and gain element 408A together implement another reverberation filter, and filter 409 and gain element 409A together implement another reverberation filter. Each of filters 406, 407, 408, and 409 in Figure 9 is preferably implemented as a filter with a maximum gain value (unit gain) close to 1, and each of the gain elements 406A, 407A, 408A, and 409A has (related reverberation tank delay n i Later, the filters 406, 407, 408, and 409 are configured to apply attenuation gain to the outputs of their corresponding filters that match the desired attenuation. Specifically, the gain element 406A applies (reverberation tank delay n) to the output of element 406A. i The filter 406 is configured to apply a decay gain (decaygain1) to the output of the delay line 410 (after the reverberation tank delay n2) so that the output has the decayed gain of the first target, and the gain element 407A is configured to apply a decay gain (decaygain2) to the output of the filter 407 so that the output of the delay line 411 (after the reverberation tank delay n2) has the decayed gain of the second target, and the gain element 408A is configured to apply a decay gain (decaygain2) to the output of the filter 407, and the gain element 408A The output of the filter 408 is configured to have a decay gain (decaygain3) applied to it so that the output of delay line 412 (after reverberation tank delay n3) has a decayed gain of the third target, and the gain element 409A is configured to have a decay gain (decaygain4) applied to the output of the filter 409 so that the output of delay line 413 (after reverberation tank delay n4) has a decayed gain of the fourth target.
[0152] Each of the filters 406, 407, 408, and 409 and each of the elements 406A, 407A, 408A, and 409A in the system of Figure 9 is preferably implemented to achieve a target T60 characteristic of BRIR applied by a virtualizer including the system of Figure 9 (e.g., the virtualizer of Figure 10) (each of the filters 406, 407, 408, and 409 is preferably implemented as an IIR filter, e.g., a shelf filter or a series of shelf filters). Here, T60 is the reverberation decay time (T 60 ) represents. For example, in some embodiments, filters 406, 407, 408, and 409 are each implemented as a shelf filter (for example, a shelf filter with Q=0.3 and a shelf frequency of 500Hz to achieve the T60 characteristics shown in Figure 13; T60 in Figure 13 is in seconds) or as a cascade of two IIR shelf filters (for example, one with shelf frequencies of 100Hz and 1000Hz to achieve the T60 characteristics shown in Figure 14; T60 in Figure 14 is in seconds). The shape of each shelf filter is determined to match the desired change curve from low frequency to high frequency. When filter 406 is implemented as a shelf filter (or a cascade of multiple shelf filters), the reverberation filter having filter 406 and gain element 406A is also a shelf filter (or a cascade of shelf filters). Similarly, when filters 407, 408, and 409 are each implemented as shelf filters (or a cascade of shelf filters), each reverberation filter having filter 407 (or 408 or 409) and its corresponding gain element (407A, 408A, or 409A) is also a shelf filter (or a cascade of shelf filters).
[0153] Figure 9B shows an example of filter 406 implemented as a cascaded configuration of a first shelf filter 406B and a second shelf filter 406C, as shown in Figure 9B. Filters 407, 408, and 409 may each be implemented similarly to the implementation of filter 406 in Figure 9B.
[0154] In some embodiments, the decay gain applied by elements 406A, 407A, 408A, and 409A is i ) is determined as follows:
[0155]
number
[0156] Figure 11 shows an embodiment of the following elements of Figure 9: elements 422 and 423 and IACC (interaural cross-correlation coefficient) filtering and mixing stage 424. Element 422 is configured to sum the outputs of filters 417 and 419 (in Figure 9) and present the summed signal as input to a low-shelf filter 500. Element 422 is configured to sum the outputs of filters 418 and 420 (in Figure 9) and present the summed signal as input to a high-pass filter 501. The outputs of filters 500 and 501 are added (mixed) in element 502 to generate a binaural left ear output signal, and the outputs of filters 500 and 501 are mixed in element 502 (the output of filter 500 is subtracted from the output of filter 501 in element 502) to generate a binaural right ear output signal. Elements 502 and 503 mix (add and subtract) the filtered outputs of filters 500 and 501 to generate a binaural output signal that achieves the target IACC characteristics (within an acceptable range of precision). In the embodiment of Figure 11, each of the low-shelf filter 500 and the high-pass filter 501 is typically implemented as a first-order IIR filter. In one example where filters 500 and 501 have such implementations, the embodiment of Figure 11 can achieve the exemplary IACC characteristics plotted as curve "I" in Figure 12. T This is a good match for the target IACC characteristics plotted as "[...]."
[0157] Figure 11A shows the frequency response (R1) of a typical implementation of filter 500 in Figure 11, the frequency response (R2) of a typical implementation of filter 501 in Figure 11, and a graph of the response of filters 500 and 501 connected in parallel. From Figure 11A, it is clear that the combined response is desirablely flat across the range of 100 Hz to 10,000 Hz.
[0158] Thus, in one class of embodiments, the present invention is a system (e.g., the system in Figure 10) and method for generating a binaural signal (e.g., the output of element 210 in Figure 10) in response to a set of channels of a multi-channel audio input signal. This includes a step of applying a binaural chamber impulse response (BRIR) to each channel of the set, thereby generating a filtered signal, by using a single feedback delay network (FDN) to add a common late reverberation to the downmix of the channels of the set; and a step of combining the filtered signals to generate the binaural signal. The FDN is implemented in the time domain. In some such embodiments, the time-domain FDN (e.g., FDN 220 in Figure 10, configured as in Figure 9) is: An input filter (for example, filter 400 in Figure 9) having an input coupled to receive the downmix, wherein the input filter is configured to generate a first filtered downmix in response to the downmix; A pass filter (e.g., pass filter 401 in Figure 9) is coupled and configured to perform a second filtered downmix in response to the first filtered downmix; A reverberation application subsystem (for example, all elements except elements 400, 401, and 424 in Figure 9) having a first output (for example, the output of element 422) and a second output (for example, the output of element 423), wherein the reverberation application subsystem includes a collection of reverberation tanks, each having a different delay, and the reverberation application subsystem is configured to generate a first unmixed binaural channel and a second unmixed binaural channel in response to the second filtered downmix, and to exhibit the first unmixed binaural channel at the first output and the second unmixed binaural channel at the second output; It includes an interaural cross-correlation coefficient (IACC) filtering and mixing stage (for example, stage 424 in Figure 9, which may be implemented as elements 500, 501, 502, and 503 in Figure 11) coupled to the reverberation application subsystem and configured to generate a first mixed binaural channel and a second mixed binaural channel in response to the first unmixed binaural channel and the second unmixed binaural channel.
[0159] The input filter may be implemented to generate the first filtered downmix such that each BRIR has a direct-to-late ratio (DLR) that at least substantially matches the target DLR (preferably as a cascade of two filters configured to generate it).
[0160] Each reverberation tank may be configured to generate a delayed signal and may include a reverberation filter (implemented, for example, as a shelf filter or a cascade of shelf filters) configured to add gain to the signal propagating in each reverberation tank so that the delayed signal has a gain that matches at least substantially the target delayed gain for the delayed signal. The target reverberation decay time characteristics of each BRIR (e.g., T 60 This is to achieve the characteristic.
[0161] In some embodiments, the first unmixed binaural channel is more advanced than the second unmixed binaural channel, and the reverberation tank includes a first reverberation tank (e.g., the reverberation tank in Figure 9 including delay line 410) configured to produce a first delayed signal having the shortest delay, and a second reverberation tank (e.g., the reverberation tank in Figure 9 including delay line 411) configured to produce a second delayed signal having the second shortest delay. The first reverberation tank is configured to apply a first gain to the first delayed signal, and the second reverberation tank is configured to apply a second gain to the second delayed signal, wherein the second gain is different from the first gain, and the application of the first and second gains results in attenuation of the first unmixed binaural channel relative to the second unmixed binaural channel. Typically, the first mixed binaural channel and the second mixed binaural channel exhibit a re-centered stereo image. In some embodiments, the IACC filtering and mixing stages are configured to generate the first mixed binaural channel and the second mixed binaural channel such that they have IACC characteristics that at least substantially match a target IACC characteristic.
[0162] Aspects of the present invention include methods and systems (for example, System 20 in Figure 2 or the System in Figure 3 or Figure 10) that perform (or are configured to perform or support the performance of) binaural virtualization of audio signals (for example, audio signals and / or object-based audio signals in which audio content consists of speaker channels).
[0163] In some embodiments, the virtualizer of the present invention is or includes a general-purpose processor coupled to receive or generate input data representing a multi-channel audio input signal, and programmed with software (or firmware) to perform any of the various processing, including embodiments of the method of the present invention, on the input data, or otherwise configured (e.g., in response to control data). Such a general-purpose processor is typically coupled to an input device (e.g., mouse and / or keyboard), memory, and a display device. For example, the system in Figure 3 (or system 20 in Figure 2 or a virtualizer system having elements 12, ..., 14, 15 of system 20) can be implemented in a general-purpose processor, where the input is audio data representing N channels of the audio input signal, and the output is audio data representing two channels of a binaural audio signal. A typical digital-to-analog converter (DAC) can act on the output data to generate an analog version of the binaural signal channels for playback by speakers (e.g., a pair of headphones).
[0164] While specific embodiments and applications of the present invention are described herein, it will be apparent to those skilled in the art that many modifications are possible to these embodiments and applications without departing from the scope of the invention described and claimed herein. Although certain forms of the present invention are shown and described, it should be understood that the present invention is not limited to the specific embodiments or methods described and shown herein.
Claims
1. A method for generating a binaural signal in response to a set of channels of a multi-channel audio input signal, the method being: The steps include: applying a binaural intra-room impulse response (BRIR) to each channel of the aforementioned set to generate a filtered signal; The step includes the step of generating the binaural signal by combining the filtered signals, Applying BRIR to each channel of the set includes using a late reverberation generator to introduce a common late reverberation into the downmix of the channels of the set in response to a control value presented to the late reverberation generator, the common late reverberation emulates the collective macro-attributes of the late reverberation portion of a single-channel BRIR shared across at least some channels of the set. A content-dependent energy equalizer is applied to the downmix, and the center channel of the multi-channel audio input signal is panned to both the left channel and the right channel of the downmix. method.
2. The method according to claim 1, wherein applying BRIR to each channel of the set comprises applying the direct response and early reflection portion of a single-channel BRIR for each channel of the set.
3. The method according to claim 1, wherein the late reverberation generator includes a bank of feedback delay networks for adding the common late reverberation to the downmix, and each feedback delay network in the bank adds late reverberation to a different frequency band of the downmix.
4. The method according to claim 3, wherein each of the feedback delay networks is implemented in the complex orthogonal mirror filter domain.
5. The method according to claim 1, wherein the late reverberation generator includes a single feedback delay network for adding the common late reverberation to the downmix of the channels of the set, the feedback delay network being implemented in the time domain.
6. A system that generates a binaural signal in response to a set of channels of a multi-channel audio input signal, wherein the system: A binaural intra-intra-iron response (BRIR) is applied to each channel of the aforementioned set to generate a filtered signal; The filtered signals are combined to generate the binaural signal. It has one or more processors, Applying BRIR to each channel of the set includes using a late reverberation generator to introduce a common late reverberation into the downmix of the channels of the set in response to a control value presented to the late reverberation generator, the common late reverberation emulates the collective macro-attributes of the late reverberation portion of a single-channel BRIR shared across at least some channels of the set. A content-dependent energy equalizer is applied to the downmix, and the center channel of the multi-channel audio input signal is panned to both the left channel and the right channel of the downmix. system.
7. The system according to claim 6, wherein applying BRIR to each channel of the set includes applying the direct response and early reflection portions of a single-channel BRIR for each channel of the set.
8. The system according to claim 6, wherein the late reverberation generator includes a bank of feedback delay networks configured to add the common late reverberation to the downmix, each feedback delay network in the bank adding late reverberation to a different frequency band of the downmix.
9. The system according to claim 8, wherein each of the feedback delay networks is implemented in the complex orthogonal mirror filter domain.
10. The system according to claim 6, wherein the late reverberation generator includes a time-domain implemented feedback delay network, and the late reverberation generator is configured to process the downmix in the time domain within the feedback delay network in order to add the common late reverberation to the downmix.
11. A non-temporary computer-readable storage medium having a sequence of instructions, wherein when an audio signal processing device executes the sequence of instructions, the audio signal processing device performs the method according to claim 1.
12. A computer program product comprising instructions for performing the method described in claim 1 when executed on a computer.
Citation Information
Patent Citations
Sound compensation device
JP2007336080A
Compact Side Information for Parametric Coding of Spatial Audio
JP2008527431A
Signal generation for binaural signals
JP2011529650A
Reverberation device and method for reverberating audio signals
JP2013508760A