Crosstalk Processing b-chain
The system addresses speaker configuration asymmetries through spatial enhancement and b-chain processing, achieving a balanced and compelling stereo listening experience by adjusting frequency response, time alignment, and signal level.
Patent Information
- Application Number
- JP2024207610
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-09-21
- Filing Date
- 2024-11-28
- Publication Date
- 2026-02-05
- Estimated Expiration
- 2038-11-26
AI Technical Summary
Existing audio systems often fail to provide an optimal listening experience due to asymmetries in speaker configuration, such as unequal distances, rotational offsets, and unequal frequency and amplitude responses, leading to a flawed stereo listening experience.
A system that includes spatial enhancement and b-chain processing to adjust for these asymmetries by applying N-band equalization, delay, and gain adjustments to balance frequency response, time alignment, and signal level, using components like subband spatial processors, crosstalk compensation, and b-chain processors to enhance audio signals for left and right speakers.
Restores a near-ideal spatial soundstage by correcting for system-specific asymmetries, providing a more compelling and balanced stereo listening experience even in non-ideal configurations.
Smart Images

Figure 0007811628000017 
Figure 0007811628000018 
Figure 0007811628000019
Abstract
Description
[Technical Field]
[0001] The subject matter described herein relates to audio signal processing, and more particularly to dealing with asymmetries (geometric and physical) when applying audio crosstalk cancellation to loudspeakers. [Background technology]
[0002] An audio signal may be output to a rendering system and / or room acoustics that are less than optimally configured. Figure 1A shows an example of an ideal loudspeaker and listener configuration for an ideal transaural configuration, i.e., a two-channel stereo speaker system with one listener in an empty soundproof room. As shown in Figure 1A, listener 140 is in the ideal position (i.e., the "sweet spot") to experience the rendered audio from left loudspeaker 110L and right loudspeaker 110R most accurately, spatially and timbrally, with respect to the content creator's original intent.
[0003] However, there are various situations in which the ideal "sweet spot" condition is not met or cannot be achieved by an audio output device. As shown in Figure 1B, what has been described so far includes situations in which the head position of listener 140 is laterally offset from the ideal "sweet spot" listening position between loudspeaker 110L and loudspeaker 110R. Or, as shown in Figure 1C, listener 140 is in the ideal position, but the distances between each of loudspeakers 110L and 110R and the head position of listener 140 are not equal. Furthermore, as shown in Figure 1D, listener 140 is in the ideal position, but the frequency and amplitude responses of loudspeaker 110L and loudspeaker 110R are not equal (i.e., the rendering system is "unmatched"). In another example, the physical location of listener 140 and loudspeakers 110L and 110R may be ideal, but one or more of loudspeakers 110L and 110R may be rotationally offset from the ideal angle relative to right loudspeaker 110R, as shown in FIG. 1E. Summary of the Invention
[0004] Exemplary embodiments relate to b-chain processing of spatially enhanced audio signals that adjust for asymmetries of various speakers or environments. Examples of asymmetries may include time delay between one speaker and a listener different from another speaker, signal level (perceived and intended) between one speaker and a listener different from another speaker, or frequency response between one speaker and a listener different from another speaker.
[0005] In an exemplary embodiment, a system for enhancing input audio signals for left and right speakers includes a spatial enhancement processor and a b-chain processor. The spatial enhancement processor generates a spatially enhanced signal by gain-adjusting spatial and non-spatial components of the input audio signal. The b-chain processor determines asymmetries between the left and right speakers in frequency response, time alignment, and signal level at the listening position. The b-chain processor generates a left output channel for the left speaker and a right output channel for the right speaker by applying N-band equalization to the spatial enhancement signal to adjust for the asymmetry in frequency response, applying a delay to the spatial enhancement signal to adjust for the asymmetry in time alignment, and applying a gain to the spatial enhancement signal to adjust for the asymmetry in signal level.
[0006] In an embodiment, the b-chain processor applies N-band equalization by applying one or more filters to at least one of the left spatially enhanced channel and the right spatially enhanced channel, where the one or more filters balance the frequency responses of the left and right speakers and may include at least one of a low-shelf filter and a high-shelf filter, a band-pass filter, a band-stop filter, a peak-notch filter, a low-pass filter, and a high-pass filter.
[0007] In an embodiment, the b-chain processor adjusts at least one of the delay and the gain in response to changes in listening position.
[0008] An embodiment may include a non-transitory computer-readable medium storing instructions that, when executed by a processor, configure the processor to generate a spatial enhancement signal by gain adjusting spatial and non-spatial components of an input audio signal comprising a left input channel for a left speaker and a right input channel for a right speaker, determine asymmetry between the left and right speakers, apply N-band equalization to the spatial enhancement signal to adjust for the frequency response asymmetry, apply a delay to the spatial enhancement signal to adjust for the time alignment asymmetry, and apply a gain to the spatial enhancement signal to adjust for the signal level asymmetry, thereby generating a left output channel for the left speaker and a right output channel for the right speaker.
[0009] An embodiment may include a method for processing input audio signals for left and right speakers, the method may include generating a spatial enhancement signal by gain adjusting spatial and non-spatial components of an input audio signal including a left input channel for a left speaker and a right input channel for a right speaker, determining asymmetry between the left and right speakers in frequency response, time alignment, and signal level at a listening position, and generating a left output channel for the left speaker and a right output channel for the right speaker by applying N-band equalization to the spatial enhancement signal to adjust for the asymmetry in frequency response, applying a delay to the spatial enhancement signal to adjust for the asymmetry in time alignment, and applying a gain to the spatial enhancement signal to adjust for the asymmetry in signal level. [Brief explanation of the drawings]
[0010] [Figure 1A] 1 illustrates the position of loudspeakers relative to a listener according to some embodiments. [Figure 1B] 1 illustrates the position of loudspeakers relative to a listener according to some embodiments. [Figure 1C] 1 illustrates the position of loudspeakers relative to a listener according to some embodiments. [Figure 1D] 1 illustrates the position of loudspeakers relative to a listener according to some embodiments. [Figure 1E] 1 illustrates the position of loudspeakers relative to a listener according to some embodiments. [Figure 2] FIG. 1 is a schematic block diagram of an audio processing system according to some embodiments. [Figure 3] FIG. 2 is a schematic block diagram of a spatial enhancement processor according to some embodiments. [Figure 4] FIG. 2 is a schematic block diagram of a subband spatial processor according to some embodiments. [Figure 5] FIG. 2 is a schematic block diagram of a crosstalk compensation processor according to some embodiments. [Figure 6] FIG. 1 is a schematic block diagram of a crosstalk cancellation processor according to some embodiments. [Figure 7] FIG. 2 is a schematic block diagram of a b-chain processor according to some embodiments. [Figure 8] 1 is a flowchart of a method for b-chain processing of an input audio signal according to some embodiments. [Figure 9] 1 illustrates a non-ideal head position and mismatched loudspeakers according to some embodiments. [Figure 10A] 10 illustrates the frequency response of the non-ideal head position and unmatched loudspeakers shown in FIG. 9 according to some embodiments. [Figure 10B] 10 illustrates the frequency response of the non-ideal head position and unmatched loudspeakers shown in FIG. 9 according to some embodiments. [Figure 11] FIG. 1 is a schematic block diagram of a computer system according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0011] The drawings and detailed description depict various non-limiting embodiments for purposes of illustration only.
[0012] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. The following description sets forth certain specific details to provide a thorough understanding of various embodiments. However, the described embodiments may be practiced without these specific details. In other instances, specific methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.
[0013] Embodiments of the present disclosure relate to an audio processing system that provides spatial enhancement and b-chain processing. Spatial enhancement can include applying sub-band spatial processing and crosstalk cancellation to an input audio signal. The b-chain processing restores the perceived spatial soundstage of transaurally rendered audio on a non-ideally configured stereo loudspeaker rendering system.
[0014] A digital audio system, such as might be used in a movie theater or for personal headphones, can be thought of as two parts: an a-chain and a b-chain. In a movie theater environment, for example, the a-chain typically contains the audio recording on the film print, which is available in Dolby Analog and a choice of digital formats such as Dolby Digital, DTS, and SDDS. Additionally, devices that capture and process the audio from the film print and prepare it for amplification are part of the a-chain.
[0015] The b-chain includes hardware and software systems for applying multi-channel volume control, equalization, time alignment, and amplification to loudspeakers to correct for and / or minimize the effects of suboptimally configured rendering system installation, room acoustics, or listener position. The b-chain processing can be configured analytically or parametrically to optimize the perceived quality of the listening experience, with the general goal of bringing the listener closer to an "ideal" experience.
[0016] Exemplary Audio System 2 is a schematic block diagram of an audio processing system 200 according to some embodiments. The audio processing system 200 applies subband spatial processing, crosstalk cancellation processing, and b-chain processing to an input audio signal X comprising a left input channel XL and a right input channel XR to generate an output audio signal O comprising a left output channel OL and a right output channel OR. The output audio signal O restores the perceived spatial sound stage for a transaurally rendered input audio signal X on a non-ideally configured stereo loudspeaker rendering system.
[0017] The audio processing system 200 includes a spatial enhancement processor 205 connected to a b-chain processor 240. The spatial enhancement processor 205 includes a subband spatial processor 210, a crosstalk compensation processor 220, and a crosstalk cancellation processor 230 connected to the subband spatial processor 210 and the crosstalk compensation processor 220.
[0018] The subband spatial processor 210 generates a spatially enhanced audio signal by gain adjusting the mid and side subband components of the left input channel XL and the right input channel XR. The crosstalk compensation processor 220 performs crosstalk compensation to compensate for spectral deficiencies or artifacts of the crosstalk cancellation applied by the crosstalk cancellation processor 230. The crosstalk cancellation processor 230 performs crosstalk cancellation on the combined output of the subband spatial processor 210 and the crosstalk compensation processor 220 to generate a left enhancement channel AL and a right enhancement channel AR. Additional details regarding the spatial enhancement processor 210 are described below with respect to Figures 3-6.
[0019] The b-chain processor 240 includes a speaker matching processor 250 connected to a delay and gain processor 260. In particular, the b-chain processor 240 can adjust for the overall delay time differences between loudspeakers 110L and 110R and the listener's head, the (perceived and intended) signal level differences between loudspeakers 110L and 110R and the listener's head, and the frequency response differences between loudspeakers 110L and 110R and the listener's head.
[0020] The speaker matching processor 250 receives the left enhancement channel AL and the right enhancement channel AR and performs speaker balancing for devices that do not provide matched speaker pairs, such as those of a mobile device or other types of left / right speaker pairs. In an embodiment, the speaker matching processor 250 applies equalization and gain or attenuation to each of the left enhancement channel AL and the right enhancement channel AR to provide a spectrally and perceptually balanced stereo image from the perspective of an ideal listening sweet spot. The delay and gain processor 260 receives the output of the speaker matching processor 250 and applies equalization and gain or attenuation to each of the channels AL and AR to time align and perceptually balance the spatial image from a particular listener's head position given actual physical asymmetries in the rendering / listening system (e.g., an off-center head position and / or unequal loudspeaker-to-head distances). The processing applied by the speaker matching processor 250 and the delay and gain processor 260 may occur in different orders. Additional details regarding the b-chain processor 240 are discussed below with respect to FIG.
[0021] Spatial Enhancement Processor Example FIG. 3 is a schematic block diagram of a spatial enhancement processor 205 according to some embodiments. The spatial enhancement processor 205 spatially enhances an input audio signal and performs crosstalk cancellation on the spatially enhanced audio signal. To that end, the spatial enhancement processor 205 receives an input audio signal X including a left input channel XL and a right input channel XR. In an embodiment, the input audio signal X is provided from a source component of a digital bitstream (e.g., PCM data, etc.). The source component can be a computer, a digital audio player, an optical disc player (e.g., DVD, CD, Blu-ray, etc.), a digital audio streamer, or other source of a digital audio signal. The spatial enhancement processor 205 processes the input channel XL and the input channel XR to generate an output audio signal A including two output channels AL and AR. The output audio signal A is a spatially enhanced audio signal of the input audio signal X through crosstalk compensation and crosstalk cancellation. Although not shown in FIG. 3, the spatial enhancement processor 205 may further include an amplifier that amplifies the output audio signal A from the crosstalk cancellation processor 230 and provides the signal A to an output device that converts the output channels AL and AR into sound, such as, for example, loudspeaker 110R and loudspeaker 110R.
[0022] The spatial enhancement processor 205 includes a subband spatial processor 210, a crosstalk compensation processor 220, a combiner 222, and a crosstalk cancellation processor 230. The spatial enhancement processor 205 performs crosstalk compensation and subband spatial processing on the input audio input channels XL and XR, combines the results of the subband spatial processing with the results of the crosstalk compensation, and then performs crosstalk cancellation on the combined signal.
[0023] The subband spatial processor 210 includes a spatial frequency band divider 310, a spatial frequency band processor 320, and a spatial frequency band combiner 330. The spatial frequency band divider 310 is connected to the input channel XL and the input channel XR and the spatial frequency band processor 320. The spatial frequency band divider 310 receives the left input channel XL and the right input channel XR and processes the input channels into spatial (or "side") components Ys and non-spatial (or "mid") components Ym. For example, the spatial component Ys can be generated based on the difference between the left input channel XL and the right input channel XR. The non-spatial component Ym can be generated based on the sum of the left input channel XL and the right input channel XR. The spatial frequency band divider 310 provides the spatial component Ys and the non-spatial component Ym to the spatial frequency band processor 320.
[0024] The spatial frequency band processor 320 is connected to the spatial frequency band divider 310 and the spatial frequency band combiner 330. The spatial frequency band processor 320 receives the spatial components Ys and the non-spatial components Ym from the spatial frequency band divider 310 and enhances the received signal. In particular, the spatial frequency band processor 320 generates an enhanced spatial component Es from the spatial components Ys and an enhanced non-spatial component Em from the non-spatial components Ym.
[0025] For example, the spatial frequency band processor 320 applies subband gains to the spatial components Ys to generate enhanced spatial components Es and to the non-spatial components Ym to generate enhanced non-spatial components Em. In some embodiments, additionally or alternatively, the spatial frequency band processor 320 provides subband delays to the spatial components Ys to generate the enhanced spatial components Es and to the non-spatial components Ym to generate the enhanced non-spatial components Em. The subband gains and / or delays may be different for different (e.g., n) subbands of the spatial components Ys and the non-spatial components Ym or may be the same (e.g., for two or more subbands). The spatial frequency band processor 320 adjusts the gains and / or delays of different subbands of the spatial components Ys and the non-spatial components Ym relative to each other to generate the enhanced spatial components Es and the enhanced non-spatial components Em. The spatial frequency band processor 320 then provides the enhanced spatial components Es and the enhanced non-spatial components Em to the spatial frequency band combiner 330 .
[0026] The spatial frequency band combiner 330 is connected to the spatial frequency band processor 320 and further connected to the combiner 222. The spatial frequency band combiner 330 receives the enhanced spatial components Es and the enhanced non-spatial components Em from the spatial frequency band processor 320 and combines the enhanced spatial components Es and the enhanced non-spatial components Em into a left spatial enhancement channel EL and a right spatial enhancement channel ER. For example, the left spatial enhancement channel EL may be generated based on the sum of the enhanced spatial components Es and the enhanced non-spatial components Em, and the right spatial enhancement channel ER may be generated based on the difference between the enhanced non-spatial components Em and the enhanced spatial components Es. The spatial frequency band combiner 330 provides the left spatial enhancement channel EL and the right spatial enhancement channel ER to the combiner 222.
[0027] The crosstalk compensation processor 220 performs crosstalk compensation to compensate for spectral imperfections and artifacts of crosstalk cancellation. The crosstalk compensation processor 220 receives the input channels XL and XR and performs processing to compensate for artifacts in the subsequent crosstalk cancellation of the enhanced non-spatial component Em and the enhanced spatial component Es performed by the crosstalk cancellation processor 230. In an embodiment, the crosstalk compensation processor 220 may perform enhancement on the non-spatial component Xm and the spatial component Xs by applying a filter to generate a crosstalk-compensated signal Z comprising a left crosstalk compensation channel ZL and a right crosstalk compensation channel ZR. In another embodiment, the crosstalk compensation processor 220 may perform enhancement only on the non-spatial component Xm.
[0028] The combiner 222 combines the left spatial enhancement channel EL with the left crosstalk compensation channel ZL to generate a left enhancement compensation channel TL, and combines the right spatial enhancement channel ER with the right crosstalk compensation channel ZR to generate a right enhancement compensation channel TR. The combiner 222 is connected to the crosstalk cancellation processor 230 and provides the left enhancement compensation channel TL and the right enhancement compensation channel TR to the crosstalk cancellation processor 230.
[0029] The crosstalk cancellation processor 230 receives the left enhancement-compensated channel TL and the right enhancement-compensated channel TR and performs crosstalk cancellation on the channels TL, TR to generate an output audio signal A comprising a left output channel OL and a right output channel OR.
[0030] Additional details regarding the subband spatial processor 210 are described below with respect to FIG. 4, additional details regarding the crosstalk compensation processor 220 are described below with respect to FIG. 5, and additional details regarding the crosstalk cancellation processor 230 are described below with respect to FIG. 6.
[0031] 4 is a schematic block diagram of a sub-band spatial processor 210 according to some embodiments. The sub-band spatial processor 210 includes a spatial frequency band divider 310, a spatial frequency band processor 320, and a spatial frequency band combiner 330. The spatial frequency band divider 310 is connected to the spatial frequency band processor 320, and the spatial frequency band processor 320 is connected to the spatial frequency band combiner 330.
[0032] The spatial frequency band divider 310 includes an L / RM / S converter 402 that receives a left input channel XL and a right input channel XR and converts these inputs into spatial components Xs and non-spatial components Xm. The spatial components Xs can be generated by subtracting the left input channel XL and the right input channel XR. The non-spatial components Xm can be generated by adding the left input channel XL and the right input channel XR.
[0033] The spatial frequency band processor 320 receives the non-spatial components Xm and applies a set of subband filters to generate non-spatial enhancement subband components Em. Additionally, the spatial frequency band processor 320 receives the spatial subband components Xs and applies a set of subband filters to generate non-spatial enhancement subband components Em. The subband filters may include various combinations of peak filters, notch filters, low-pass filters, high-pass filters, low-shelf filters, high-shelf filters, band-pass filters, band-stop filters, and / or all-pass filters.
[0034] In an embodiment, the spatial frequency band processor 320 includes a subband filter for each of the n frequency subbands of the non-spatial component Xm and a subband filter for each of the n frequency subbands of the spatial component Xs. For example, for n=4, the spatial frequency band processor 320 includes a series of subband filters for the non-spatial component Xm, including a mid equalization (EQ) filter 404(1) for subband(1), a mid EQ filter 404(2) for subband(2), a mid EQ filter 404(3) for subband(3), and a mid EQ filter 404(4) for subband(4). Each mid EQ filter 404 applies a filter to a frequency subband portion of the non-spatial component Xm to generate an enhanced non-spatial component Em.
[0035] Additionally, the spatial frequency band processor 320 includes a series of subband filters for the frequency subbands of the spatial component Xs, including a side equalization (EQ) filter 406(1) for subband(1), a side EQ filter 406(2) for subband(2), a side EQ filter 406(3) for subband(2), and a side EQ filter 406(4) for subband(3). Each side EQ filter 406 applies a filter to a frequency subband portion of the spatial component Xs to produce an enhanced spatial component Es.
[0036] Each of the n frequency subbands for the non-spatial component Xm and the spatial component Xs may correspond to a frequency range. For example, frequency subband (1) corresponds to 0 to 300 Hz, frequency subband (2) corresponds to 300 to 510 Hz, frequency subband (3) corresponds to 510 to 2700 Hz, and frequency subband (4) corresponds to 2700 Hz to the Nyquist frequency. In some embodiments, the n frequency subbands are a unified set of important bands. The important bands may be defined using a corpus of audio samples from various musical genres. The long-term average energy ratio of the mid-to-side components in the 24 Bark scale critical bands is determined from the samples. Contiguous frequency bands with similar long-term average ratios are then grouped to form a set of important bands. The range of the frequency subbands and the number of frequency subbands are adjustable. In some embodiments, each of the n frequency subbands may include a set of important bands.
[0037] In an embodiment, the mid EQ filter 404 or the side EQ filter 406 may include a biquad filter having a transfer function defined by Equation 1.
[0038]
number
[0039] where z is a complex variable and a0, a1, a2, b0, b1, and b2 are digital filter coefficients. The filter may be implemented using a direct form I topology defined by Equation 2.
[0040]
number
[0041] where X is the input vector and Y is the output. Other topologies could be advantageous for some processors depending on maximum word length and saturation behavior.
[0042] It is then possible to implement any second-order filter with real input and output values using the fourth order. To design a discrete-time filter, a continuous-time filter is designed and transformed to discrete time via a bilinear transform. Furthermore, compensation for any resulting shifts in center frequency and bandwidth can be achieved using frequency distortion.
[0043] For example, the peak filter may include an S-plane transfer function defined by Equation 3:
[0044]
number
[0045] where S is a complex variable, A is the amplitude of the peak, and Q is the filter "quality" (canonically derived as
[0046]
number
[0047] ) The digital filter coefficients are as follows:
[0048]
number
[0049] where ω0 is the center frequency of the filter in radians and
[0050]
number
[0051] This is expressed as:
[0052] The spatial frequency band combiner 330 receives the mid and side components, applies a gain to each component, and converts the mid and side components into left and right channels. For example, the spatial frequency band combiner 330 receives the enhanced non-spatial components Em and the enhanced spatial components Es and performs a global mid-and-side gain before converting the enhanced non-spatial components Em and the enhanced spatial components Es into the left spatial enhancement channel EL and the right spatial enhancement channel ER.
[0053] Specifically, spatial frequency band combiner 330 includes a global mid gain 408, a global side gain 410, and an M / SL / R converter 412 coupled to global mid gain 408 and global side gain 410. Global mid gain 408 receives and applies a gain to the enhanced non-spatial components Em, and global side gain 410 receives and applies a gain to the enhanced non-spatial components Es. M / SL / R converter 412 receives the enhanced non-spatial components Em from global mid gain 408 and the enhanced spatial components Es from global side gain 410 and converts these inputs into a left spatial enhancement channel EL and a right spatial enhancement channel ER.
[0054] 5 is a schematic block diagram of a crosstalk compensation processor 220 according to some embodiments. The crosstalk compensation processor 220 receives left and right input channels and generates left and right output channels by applying crosstalk compensation on the input channels. The crosstalk compensation processor 220 includes an L / RM / S converter 502, a mid component processor 520, a side component processor 530, and an M / SL / R converter 514.
[0055] When the crosstalk compensation processor 220 is part of the audio system 202, 400, 500, or 504, the crosstalk compensation processor 220 receives input channels XL and XR and performs pre-processing to generate left and right crosstalk compensation channels ZL and ZR. The channels ZL and ZR may be used to compensate for crosstalk processing artifacts, such as crosstalk cancellation or simulation. The L / RM / S converter 502 receives the left input audio channel XL and the right input audio channel XR and generates non-spatial and spatial components Xm and Xs of the input channels XL and XR. In general, the left and right channels may be summed to generate the non-spatial components of the left and right channels and subtracted to generate the spatial components of the left and right channels.
[0056] The mid component processor 520 includes multiple filters 540, such as m mid filters 540(a), 540(b), and 540(m). Each of the m mid filters 540 processes one of m frequency bands of the non-spatial component Xm and the spatial component Xs. The mid component processor 520 processes the non-spatial component Xm to generate the mid crosstalk compensation channel Zm. In an embodiment, the mid filter 540 is configured using a frequency response plot of the non-spatial component Xm due to crosstalk processing through simulation. Furthermore, by analyzing the frequency response plot, spectral impairments, such as peaks and troughs in the frequency response plot that occur as crosstalk processing artifacts, can be estimated above a preset threshold (e.g., 10 dB). These artifacts are primarily due to the summation of delayed and inverted contralateral signals with the corresponding ipsilateral signals in the crosstalk processing, thereby effectively introducing a comb-filter-like frequency response into the final rendering result. The mid crosstalk compensation channel Zm can be generated by the mid component processor 520 to compensate for the estimated peaks or troughs, where each of the m frequency bands corresponds to a peak or trough. Specifically, based on the specific delays, filtering frequencies, and gains applied in the crosstalk processing, the peaks and troughs move up or down in the frequency response, causing amplification or attenuation of energy in specific regions of the spectrum. Each mid filter 540 can be configured to tune to one or more peaks and troughs.
[0057] The side component processor 530 includes multiple filters 550, such as m side filters 550(a), 550(b), 550(m), and so on. The side component processor 530 processes the spatial components Xs to generate the side crosstalk compensation channels Zs. In an embodiment, a frequency response plot of the spatial components Xs due to crosstalk processing can be obtained by simulation. By analyzing the frequency response plot, spectral impairments, such as peaks and troughs in the frequency response plot that occur as artifacts of the crosstalk processing, can be estimated above a predetermined threshold (e.g., 10 dB). The side crosstalk compensation channels Zs can be generated by the side component processor 530 to compensate for the estimated peaks or troughs. Specifically, the peaks and troughs move up or down in the frequency response based on the specific delays, filtering frequencies, and gains applied in the crosstalk processing, causing amplification or attenuation of energy in specific regions of the spectrum. Each side filter 550 can be configured to tune to one or more peaks and troughs. In some embodiments, the mid component processor 520 and the side component processor 530 may include different numbers of filters.
[0058] In an embodiment, the mid-filter 540 and the side-filter 550 may include fourth-order filters having transfer functions defined by Equation 4.
[0059]
number
[0060] where z is a complex variable and a0, a1, a2, b0, b1, and b2 are the digital filter coefficients. One way to implement such a filter is the direct form I topology defined in Equation 5.
[0061]
number
[0062] where X is the input vector and Y is the output. Other topologies may be used depending on the maximum word length and saturation behavior.
[0063] Then, using biquads, second-order filters with real-valued inputs and outputs can be implemented. To design a discrete-time filter, a continuous-time filter is designed and converted to discrete time by a bilinear transformation. Furthermore, shifts in center frequency and bandwidth can be compensated using frequency distortion.
[0064] For example, the peak filter has a complex plane transfer function defined by Equation 6.
[0065]
number
[0066] where s is a complex variable, A is the amplitude of the peak, Q is the filter "quality", and the digital filter coefficients are defined as follows:
[0067]
number
[0068] where ω0 is the center frequency of the filter in radians and
[0069]
number
[0070] This is expressed as:
[0071] Furthermore, the filter quality Q can be defined by Equation 7:
[0072]
number
[0073] however,
[0074]
number
[0075] is the bandwidth, f c is the center frequency.
[0076] The M / SL / R converter 514 receives the mid crosstalk compensation channel Zm and the side crosstalk compensation channel Zs and generates the left crosstalk compensation channel ZL and the right crosstalk compensation channel ZR. In general, the mid channel and the side channel may be added to generate the left channel of the mid and side components, and the mid channel and the side channel may be subtracted to generate the right channel of the mid and side components.
[0077] 6 is a schematic block diagram of a crosstalk cancellation processor 230 according to some embodiments, which receives the left enhancement-compensated channel TL and the right enhancement-compensated channel TR from the combiner 222 and performs crosstalk cancellation on the channels TL, TR to produce left output channel AL and right output channel AR.
[0078] The crosstalk cancellation processor 230 includes an in-out-band divider 610, inverters 620 and 622, contralateral estimators 630 and 640, combiners 650 and 652, and an in-out-band combiner 660. These components split the input channels TL, TR into in-band and out-of-band components, perform crosstalk cancellation on the in-band components, and generate the output channels AL, AR.
[0079] By dividing the input audio signal T into different frequency band components and performing crosstalk cancellation on selective components (e.g., in-band components), crosstalk cancellation can be performed in specific frequency bands while eliminating degradation in other frequency bands. If crosstalk cancellation is performed without dividing the input audio signal T into different frequency bands, the audio signal after crosstalk cancellation may exhibit significant attenuation or amplification of non-spatial and spatial components at low frequencies (e.g., below 350 Hz), high frequencies (e.g., above 12,000 Hz), or both. Selectively performing crosstalk cancellation in-band (e.g., 250 Hz to 14,000 Hz), where the majority of impactful spatial cues reside, can maintain balanced overall energy across the spectrum in the mix, especially for non-spatial components.
[0080] The in-out band divider 610 separates the input channels TL and TR into in-band channels TL,In and TR,In and out-of-band channels TL,Out and TR,Out, respectively. In particular, the in-out band divider 610 divides the left enhancement-compensated channel TL into a left in-band channel TL,In and a left out-of-band channel TL,Out. Similarly, the in-out band divider 610 divides the right enhancement-compensated channel TR into a right in-band channel TR,In and a right out-of-band channel TR,Out. Each in-band channel encompasses a portion of each input channel corresponding to a frequency range, for example, 250 Hz to 14 kHz. The range of the frequency band can be adjusted according to speaker parameters, etc.
[0081] The inverter 620 and the contralateral estimator 630 operate together to generate a left contralateral cancellation component SL to compensate for the contralateral sound component due to the left in-band channel TL,In. Similarly, the inverter 622 and the contralateral estimator 640 operate together to generate a right contralateral cancellation component SR to compensate for the contralateral sound component due to the right in-band channel TR,In.
[0082] In one approach, the inverter 620 receives the in-band channel TL,In and inverts the polarity of the received in-band channel TL,In to generate an inverted in-band channel TL,In′. The contralateral estimator 630 receives the inverted in-band channel TL,In′ and extracts, through filtering, the portion of the inverted in-band channel TL,In′ that corresponds to the contralateral sound component. Because filtering is performed on the inverted in-band channel TL,In′, the portion extracted by the contralateral estimator 630 is the inverse of the portion of the in-band channel TL,In that is attributable to the contralateral sound component. Thus, the portion extracted by the contralateral estimator 630 becomes the left contralateral cancellation component SL, which can be added to the contralateral in-band channel TR,In to reduce the contralateral sound component attributable to the in-band channel TL,In. In some embodiments, the inverter 620 and the contralateral estimator 630 are implemented in a different order.
[0083] The inverter 622 and the contralateral estimator 640 perform a similar operation on the in-band channel TR,In to generate the right contralateral cancellation component SR, and therefore a detailed description thereof will be omitted herein for the sake of brevity.
[0084] In one exemplary implementation, the contralateral estimator 630 includes a filter 632, an amplifier 634, and a delay unit 636. The filter 632 receives the inverted input channel TL,In′ and extracts a portion of the inverted in-band channel TL,In′ corresponding to the contralateral sound component through a filtering function. An example of a filter implementation is a notch or highshelf filter with a center frequency selected between 5000 and 10000 Hz and a Q selected between 0.5 and 1.0. The gain in decibels (GdB) can be derived from Equation 8: G dB =-3.0-log 1.333 (D) Equation 8 where D is the amount of delay by the delay unit 636 in samples, eg, a sampling rate of 48 KHz.
[0085] Another implementation is a low pass filter with a corner frequency selected in the range 5000-10000 Hz and a Q selected in the range 0.5-1.0. Additionally, amplifier 634 has a corresponding gain factor G L,In The amplifier 634 amplifies the extracted portion by a delay unit 636, which delays the amplified output from the amplifier 634 according to a delay function D to generate a left contralateral cancellation component SL. The contralateral estimator 640 includes a filter 642, an amplifier 644, and a delay unit 646. This unit outputs the inverted in-band channel T R,In In one example, the contralateral estimators 630, 640 generate the left contralateral cancellation component SL and the right contralateral cancellation component SR according to the following equations: S L =D[G L,In *F[T L,In ']] expression 9 S R =D[G R,In *F[T R,In ']] Expression 10 where F[] is the filter function and D[] is the delay function.
[0086] The crosstalk cancellation settings can be determined by speaker parameters, such as the filter center frequency, delay, amplifier gain, and filter gain, depending on the angle relative to the listener between the two speakers 110. In some embodiments, values between speaker angles are used to interpolate other values.
[0087] Combiner 650 combines the right contralateral cancellation component SR with the left in-band channel TL,In to generate a left in-band compensation channel UL, and combiner 652 combines the left contralateral cancellation component SL with the right in-band channel TR,In to generate a right in-band compensation channel UR. In-outband combiner 660 combines the left in-band compensation channel UL with the out-of-band channel TL,Out to generate a left output channel AL, and combines the right in-band compensation channel UR with the out-of-band channel TR,Out to generate a right output channel AR.
[0088] Thus, the left output channel AL includes a right contralateral cancellation component SR corresponding to the inverse of the portion of the in-band channels TR,In attributable to contralateral sound, and the right output channel AR includes a left contralateral cancellation component SL corresponding to the inverse of the portion of the in-band channels TL,In attributable to contralateral sound. In this configuration, the wavefront of the ipsilateral sound component output by loudspeaker 110R responsive to right output channel AR at the right ear can cancel the wavefront of the contralateral sound component output by loudspeaker 110L responsive to left output channel AL. Similarly, the wavefront of the ipsilateral sound component output by loudspeaker 110L responsive to left output channel AL at the left ear can cancel the wavefront of the contralateral sound component output by loudspeaker 110R responsive to right output channel AR. Thus, the contralateral sound component can be reduced to enhance spatial detectability.
[0089] Exemplary b-chain processor 7 is a schematic block diagram of a b-chain processor 240 according to some embodiments. The b-chain processor 240 includes a speaker matching processor 250 and a delay and gain processor 260. The speaker matching processor 250 includes an N-band equalizer (EQ) 702 connected to a left amplifier 704 and a right amplifier 706. The delay and gain processor 260 includes a left delay 708 connected to a left amplifier 712 and a right delay 710 connected to a right amplifier 714.
[0090] Assuming that the orientation of the listener 140 remains fixed toward the center of an ideal spatial image (e.g., a virtual lateral center of the sound field, predetermined symmetry, matching, and equidistant loudspeakers, etc.), as shown in Figures 1A-1E, the transformation relationship between the ideal spatial image and the actually rendered spatial image can be explained based on the fact that (a) the overall time delay between one speaker and the listener 140 is different from that of another speaker, (b) the (perceived and intended) signal level between one speaker and the listener 140 is different from that of another speaker, and (c) the frequency response between one speaker and the listener 140 is different from that of another speaker.
[0091] The b-chain processor 240 corrects for the above relative differences in delay, signal level, and frequency response, resulting in a restoration of a near-ideal spatial image, as if the listener 140 (e.g., head position) and / or rendering system were ideally configured.
[0092] The b-chain processor 240 receives as input an audio signal A, including a left enhancement channel AL and a right enhancement channel AR, from the spatial enhancement processor 205. The input to the b-chain processor 240 can include any stereo audio stream that has been transaurally processed for a given listener / speaker configuration under ideal conditions (as illustrated in FIG. 1A). If audio signal A has no spatial asymmetry, and if no other anomalies are present in the system, the spatial enhancement processor 205 provides a dramatically enhanced sound field to the listener 140. However, if asymmetries are present in the system, as described above and illustrated in FIGS. 1B-1E, the b-chain processor 240 can be applied to maintain the enhanced sound field under non-ideal conditions.
[0093] While an ideal listener / speaker configuration involves a pair of loudspeakers with matched left and right speaker-to-head distances, many real-world setups fail to meet these criteria, resulting in a flawed stereo listening experience. For example, a mobile device may include front-facing earpiece loudspeakers with limited bandwidth (e.g., 1000–8000 Hz frequency response) and orthogonally oriented (downward or sideways) micro-loudspeakers (e.g., 200–20,000 Hz frequency response). Here, the speaker systems are mismatched in two ways: due to different audio driver performance characteristics (e.g., signal level, frequency response, etc.) and due to mismatched time alignment with respect to the “ideal” listener position due to non-parallel speaker orientation. Another example is when a listener using a stereo desktop speaker system does not position either the loudspeakers or the speakers themselves in the ideal configuration (e.g., as shown in Figures 1B, 1C, or 1E). The b-chain processor 240 thus helps to adjust the characteristics of each channel, addressing any associated system-specific asymmetries, resulting in a more perceptually compelling transaural sound field.
[0094] After spatial enhancement processing or other processing has been applied to the stereo input signal X, which has been tuned under the assumption of an ideally configured system (i.e., sweet-spot listener, matching, symmetrically placed loudspeakers, etc.), the speaker matching processor 250 provides practical loudspeaker balancing for devices that do not offer matched speaker pairs, as is the case in most mobile devices. The N-band EQ 702 of the speaker matching processor 250 receives the left enhancement channel AL and the right enhancement channel AR and applies equalization to each of the channels AL and AR.
[0095] In an embodiment, the N-band EQ 702 provides a variety of EQ filter types, such as low-shelf filters, high-shelf filters, band-pass filters, band-stop filters, peak-notch filters, low-pass filters, and high-pass filters. For example, if one loudspeaker in a stereo pair is angled away from the ideal listener sweet spot, that loudspeaker will exhibit significant high-frequency attenuation from the listener sweet spot. One or more bands of the N-band EQ 702 can be applied to the loudspeaker channel to restore high-frequency energy when viewed from the sweet spot (e.g., via a high-shelf filter), achieving a close match to the characteristics of the other forward-facing loudspeaker. In another scenario, if both loudspeakers are front-facing but one loudspeaker has a significantly different frequency response, EQ tuning can be applied to both the left and right channels to achieve spectral balance between the two. Applying such adjustments can be equivalent to "rotating" the desired speaker to match the orientation of the other forward-facing speaker. In an embodiment, the N-band EQ 702 includes a filter for each of n bands that are processed independently. The number of bands can vary. In an embodiment, the number of bands corresponds to the subbands of the subband spatial processing.
[0096] In an embodiment, the speaker asymmetry may be predefined for a particular set of speakers, with the known asymmetry used as a basis for selecting parameters for the N-band EQ 702. In another example, the speaker asymmetry may be determined based on testing the speakers, for example, by using a test audio signal, recording the sound produced by the speakers from the signal, analyzing the recorded sound, etc.
[0097] Left amplifier 704 is connected to N-band EQ 702 to receive the left channel, and right amplifier 706 is connected to N-band EQ 702 to receive the right channel. Amplifiers 704 and 706 address asymmetries in the loudspeaker loudspeaker and dynamic range capabilities by adjusting the output gain on one or both channels. This is particularly useful for balancing loudspeaker loudspeaker distances from the listening position and for balancing unmatched loudspeaker pairs with widely different sound pressure level (SPL) output characteristics.
[0098] The delay and gain processor 260 receives the left and right output channels of the speaker matching processor 250 and applies a time delay and gain or attenuation to one or more of the channels. To that end, the delay and gain processor 260 includes a left delay 708 that receives the left channel output from the speaker matching processor 250 and applies a time delay, and a left amplifier 712 that applies gain or attenuation to the left channel to generate a left output channel OL. Additionally, the delay and gain processor 260 includes a right delay 710 that receives the right channel output from the speaker matching processor 250 and applies a time delay, and a right amplifier 714 that applies gain or attenuation to the right channel to generate a right output channel OR. As previously mentioned, the speaker matching processor 250 ignores time-based asymmetries present in the actual configuration, focusing on perceptually balancing the left / right spatial images from the perspective of an ideal listener "sweet spot" and providing balanced SPL and frequency responses for each driver from that position. After this speaker matching is achieved, the delay and gain processor 260 time-aligns and further perceptually balances the spatial image from a particular listener's head position given the actual physical asymmetries of the rendering / listening system (e.g., off-center head position and / or unequal speaker-to-head distances).
[0099] The delay and gain values applied by the delay and gain processor 260 may be set to accommodate static system configurations, such as, for example, a cell phone using orthogonally oriented loudspeakers, or a listener that is laterally offset from the ideal sweet spot in front of the speakers, such as, for example, a home theater soundbar.
[0100] Additionally, the delay and gain values applied by the delay-and-gain processor 260 may be dynamically adjusted based on the changing spatial relationship between the listener's head and the loudspeakers, as may occur in gaming scenarios that use physical movement as an element of gameplay (e.g., positional tracking using a depth camera in games, artificial reality systems, etc.). In an embodiment, the audio processing system includes a camera, light sensor, proximity sensor, or other suitable device used to determine the position of the listener's head relative to the speakers. The determined user's head position may be used to determine the delay and gain values of the delay-and-gain processor 260.
[0101] An audio analysis routine can provide the appropriate inter-speaker delays and gains used to configure the b-chain processor 240, resulting in a time-aligned and perceptually balanced left / right stereo image. In embodiments, where no measurable data is available from such analysis methods, intuitive manual user control, or automatic control via computer vision or other sensor input, can be achieved using a mapping such as defined by Equations 11 and 12 below.
[0102]
number
[0103]
number
[0104] where delayDelta and delay are in milliseconds and gain is in decibels. The column vectors for delay and gain assume that the first component relates to the left channel and the second component relates to the right channel. Thus,
[0105]
number
[0106] indicates that the left speaker delay is greater than or equal to the right speaker delay, and delayDelta<0 indicates that the left speaker delay is less than the right speaker delay.
[0107] In embodiments, instead of applying attenuation to a channel, the same amount of gain may be applied to the opposite channel, or a combination of gain applied to one channel and attenuation applied to the other channel. For example, gain may be applied to the left channel rather than attenuating it. For close-range listening, such as occurs in mobile, desktop PC and console gaming, and home theater scenarios, the difference in distance between the listener's position and each speaker is small enough, and therefore the SPL delta between the listener's position and each speaker is small enough, that any of the above mappings will help to successfully restore a transaural spatial image while maintaining an overall sound field of acceptable size compared to an ideal listener / speaker configuration.
[0108] Exemplary Audio System Processing 8 is a flowchart of a method 800 for processing an input audio signal according to some embodiments. The method 800 may have fewer or additional steps, and the steps may be performed in a different order.
[0109] The audio processing system 200 (e.g., the spatial enhancement processor 205) enhances an input audio signal to generate an enhanced signal 802. The enhancement may include spatial enhancement. For example, the spatial enhancement processor 205 applies subband spatial processing, crosstalk compensation processing, and crosstalk cancellation processing to an input audio signal X including a left input channel XL and a right input channel XR to generate an enhanced signal A including a left enhancement channel AL and a right enhancement channel AR. Here, the audio processing system 200 applies spatial enhancement by gain adjusting the mid (non-spatial) and side (spatial) subband components of the input audio signal X, and the enhanced signal A is referred to as a "spatially enhanced signal." The audio processing system 200 may perform other types of enhancement to generate the enhanced signal A.
[0110] Audio processing system 200 (e.g., N-band EQ 702 of speaker matching processor 250 of b-chain processor 240) applies N-band equalization to enhancement signal A to adjust for frequency response asymmetry between the left and right speakers 804. N-band EQ 702 may apply one or more filters to left enhancement channel AL, right enhancement channel AR, or both left channel AL and right channel AR. The one or more filters applied to left enhancement channel AL and / or right enhancement channel AR balance the frequency responses for the left and right speakers. In an embodiment, the frequency response balancing may be used to adjust for rotational offsets from ideal angles of the left and right speakers. In an embodiment, N-band EQ 702 adjusts for left and right speaker asymmetry and determines filter parameters for applying the N-band EQ based on the determined asymmetry.
[0111] The audio processing system 200 (e.g., left amplifier 704 and / or right amplifier 706) applies 806 a gain to at least one of the left enhancement channel AL and the right enhancement channel AR to adjust for asymmetries in signal levels between the left and right speakers. The applied gain can be a positive gain or a negative gain (also called attenuation) to address asymmetries in the loudness and dynamic range capabilities of the speakers or in unmatched speaker pairs with different sound pressure level (SPL) output characteristics.
[0112] The audio processing system 200 (e.g., the delay and gain processor 260 of the b-chain processor 240) applies delay and gain to the enhancement signal A to adjust for the listening position 808. The listening position may include the user's position relative to the left and right speakers. The user refers to the listener of the speakers. The delay and gain time-align and perceptually balance the spatial image output from the speaker matching processor 250 to the listener's position given the actual physical asymmetry of the rendering / listening system (e.g., off-center head position and / or unequal loudspeaker-to-head distances). For example, to the left enhancement channel AL, the left delay 708 may apply a delay, and the left amplifier 712 may apply a gain. To the right enhancement channel AR, the right delay 710 may apply a delay, and the right amplifier 714 may apply a gain. In an embodiment, the delay may be applied to one of the left enhancement channel AL or the right enhancement channel AR, and the gain may be applied to one of the left enhancement channel AL or the right enhancement channel AR.
[0113] The audio processing system 200 (e.g., the delay and gain processor 260 of the b-chain processor 240) adjusts 810 at least one of the delay and gain in response to changes in the listening position. For example, the user's spatial position relative to the left and right speakers may change. The audio processing system 200 monitors the listener's position over time, determines the gain and delay to be applied to the enhancement signal O based on the listener's position, and adjusts the delay and gain to be applied to the enhancement signal O in response to changes in the listener's position over time to generate the left output channel OL and the right output channel OR.
[0114] The various asymmetry adjustments may be performed in different orders. For example, adjustments for asymmetries in speaker characteristics (e.g., frequency response) may be performed before, after, or in conjunction with adjustments for asymmetries at the listening position relative to the speaker's position or orientation. The audio processing system determines asymmetries between the left and right speakers in frequency response, time alignment, and signal level at the listening position, applies N-band equalization to a spatial enhancement signal to adjust for the asymmetry between the left and right speakers in frequency response, applies delay to the spatial enhancement signal to adjust for the asymmetry in time alignment, and applies gain to the spatial enhancement signal to adjust for the asymmetry in signal level to generate a left output channel for the left speaker and a right output channel for the right speaker.
[0115] In embodiments, rather than applying multiple gains or delays to adjust for different causes of asymmetry (e.g., speaker characteristics or listening position), a single gain and a single delay are used to adjust for multiple types of asymmetry due to gain or time delay differences between speakers and resulting in advantageous listening positions. However, it may be beneficial to separate the processing for speaker asymmetry and listening position asymmetry to reduce processing needs. For example, once the frequency responses of the speakers are known, the same filter values may be used to adjust the speakers, while separate time delay and signal level adjustments are made for changes in listening position (e.g., user movement, etc.).
[0116] 9 illustrates a non-ideal head position and unmatched loudspeakers according to some embodiments. The listener 140 is at different distances from the left speaker 910L and the right speaker 910R. Furthermore, the frequency and / or amplitude characteristics of the speakers 910L and 910R are not equivalent. FIG. 10A illustrates the frequency response of the left speaker 910L, and FIG. 10B illustrates the frequency response of the right speaker 910R.
[0117] 9, 10A, and 10B, to correct for speaker asymmetry between speakers 910L and 910R and the position of listener 140 relative to each of speakers 910L and 910R, the components of b-chain processor 240 may use the following configuration: N-band EQ 702 may apply a high-shelf filter with a cutoff frequency of 4,500 Hz, a Q of 0.7, and a slope of -6 dB to the left enhancement channel AL, and a high-shelf filter with a cutoff frequency of 6,000 Hz, a Q of 0.5, and a slope of +3 dB to the right enhancement channel AR. Left delay 708 may apply a delay of 0 milliseconds, right delay 710 may apply a delay of 0.27 milliseconds, left amplifier 712 may apply a gain of 0 dB, and right amplifier 714 may apply a gain of -0.40625 dB.
[0118] Exemplary Computing System It is noted that the systems and processes described herein may be embodied in embedded electronic circuits or systems. Furthermore, the systems and processes may be embodied in a computing system that includes one or more processing systems (such as, for example, digital signal processors), memories (such as, for example, programmed read-only memories or programmable solid-state memories), or other circuits, such as, for example, application-specific integrated circuits (ASICs) or field-programmable gate array (FPGA) circuits.
[0119] FIG. 11 illustrates an example computer system 1100 according to an embodiment. Audio system 200 may be implemented on system 1100. At least one processor 1102 is illustrated connected to a chipset 1104. Chipset 1104 includes a memory controller hub 1120 and an I / O (input / output) controller hub 1122. Memory 1106 and a graphics adapter 1112 are connected to memory controller hub 1120, and a display device 1118 is connected to the graphics adapter 1112. A storage device 1108, a keyboard 1110, a pointing device 1114, and a network adapter 1116 are connected to I / O controller hub 1122. Other embodiments of computer 1100 have different architectures. For example, in some embodiments, memory 1106 is directly connected to processor 1102.
[0120] The storage device 1108 includes one or more temporary computer-readable storage media, such as a hard drive, a compact disc read-only memory (CD-ROM), a DVD, or a solid-state memory device. The memory 1106 holds instructions and data used by the processor 1102. For example, the memory 1106 may store instructions that, when executed by the processor 1102, cause or configure the processor 1102 to perform functions described herein, such as, for example, method 800. The pointing device 1114 is used in combination with the keyboard 1110 to input data into the computer system 1100. The graphics adapter 1112 displays images and other information on the display device 1118. In embodiments, the display device 1118 includes touchscreen capabilities for receiving user inputs and selections. The network adapter 1116 connects the computer system 1100 to a network. Some embodiments of the computer 1100 have different and / or other components than those shown in FIG. 11 . For example, the computer system 1100 may be a server lacking a display device, keyboard, and other components, or may use other types of input devices.
[0121] Additional Considerations The disclosed configurations may include a number of benefits and / or advantages. For example, an input signal can be output to unmatched loudspeakers while maintaining or enhancing the sense of spatiality of the sound field. A high-quality listening experience can be achieved even when the speakers are unmatched, even when the listener is not in an ideal listening position relative to the speakers.
[0122] Upon reading this disclosure, those skilled in the art will still recognize additional and alternative embodiments of the principles disclosed herein. Thus, while particular embodiments and applications have been illustrated and described, it is to be understood that the disclosed embodiments are not limited to the precise structure and components disclosed herein. Various modifications, changes, and variations that will be apparent to those skilled in the art may be made in the arrangement, operation, and details of the methods and apparatus disclosed herein without departing from the scope described herein.
[0123] The steps, operations, or processes described herein may be performed or implemented by one or more hardware or software modules, alone or in combination with other devices. In one embodiment, a software module is implemented by a computer-readable medium (e.g., a non-transitory computer-readable medium) containing computer program code and can be executed by a computer processor to perform some or all of the steps, operations, or processes described. [Explanation of symbols]
[0124] 110L Left loudspeaker 110R Right Loudspeaker 200 Audio Processing System 910L left speaker 910R right speaker 1100 Computer System
Claims
1. 1. A system for enhancing an input audio signal for a left speaker and a right speaker, comprising: determining, using one or more sensors, a spatial relationship between a listener's head position and the left and right speakers; determining a listener-perceived disparity in at least one of amplitude characteristics, time characteristics, or frequency characteristics of spatial images produced by the left speaker and the right speaker based on the determined spatial relationship; and applying at least one of a delay and a gain to the input audio signal to generate a left output channel for the left speaker and a right output channel for the right speaker to adjust for the listener's perceived imbalance, wherein at least one of the delay and the gain is dynamically adjusted based on changes in the spatial relationship between the listener's head and the left and right speakers as determined using the one or more sensors. A processing circuit configured as follows: A system comprising:
2. 10. The system of claim 1, wherein the processing circuitry is further configured to determine, from a viewpoint of a particular position, asymmetry between the left speaker and the right speaker in at least one of frequency response, time alignment, or signal level.
3. 3. The system of claim 2, wherein the processing circuitry is configured to apply additional gain to the input audio signal to adjust for the asymmetry due to the mismatch between the left speaker and the right speaker.
4. 10. The system of claim 1, wherein the processing circuitry is further configured to apply N-band equalization to the input audio signal to adjust for asymmetries in the frequency responses of the left and right speakers by applying one or more filters to at least one of the left or right channels of the input audio signal.
5. The system of claim 4 , wherein the one or more filters balance the frequency response of the left and right speakers.
6. The one or more filters low-shelf and high-shelf filters, A bandpass filter and A bandstop filter, A peak notch filter, Low-pass and high-pass filters 5. The system of claim 4, comprising at least one of:
7. 2. The system of claim 1, wherein the processing circuitry is configured to apply the delay or gain to the input audio signal by applying the delay or gain to one of a left channel or a right channel of the input audio signal.
8. 2. The system of claim 1, wherein the at least one of the delay and the gain adjusts for the determined spatial relationship being unequal distances from the listener's head to the left speaker and the right speaker.
9. 10. The system of claim 1, wherein the processing circuitry is further configured to apply at least one of crosstalk compensation or crosstalk cancellation to the input audio signal.
10. 10. The system of claim 9, wherein the processing circuitry is configured to apply the at least one of the crosstalk compensation or the crosstalk cancellation to the input audio signal before applying the at least one of the delay and the gain to the input audio signal.
11. 10. The system of claim 1, wherein the processing circuitry is further configured to gain adjust spatial and non-spatial components of the input audio signal.
12. 1. A method for enhancing an input audio signal for a left speaker and a right speaker by a processing circuit, comprising: determining, using one or more sensors, a spatial relationship between a listener's head position and the left and right speakers; determining a listener-perceived disparity in at least one of amplitude characteristics, time characteristics, or frequency characteristics of spatial images produced by the left speaker and the right speaker based on the determined spatial relationship; applying at least one of a delay and a gain to the input audio signal to generate a left output channel for the left speaker and a right output channel for the right speaker to adjust for the listener's perceived imbalance, wherein at least one of the delay and the gain is dynamically adjusted based on changes in the spatial relationship between the listener's head and the left and right speakers as determined using the one or more sensors; A method comprising:
13. 13. The method of claim 12, further comprising determining, from a viewpoint of a particular position, asymmetry between the left speaker and the right speaker in at least one of frequency response, time alignment, or signal level.
14. 14. The method of claim 13, further comprising applying additional gain to the input audio signal to adjust for the asymmetry due to the mismatch between the left and right speakers.
15. 13. The method of claim 12, further comprising applying N-band equalization to the input audio signal to adjust for asymmetries in the frequency responses of the left and right speakers by applying one or more filters to at least one of the left or right channels of the input audio signal.
16. 16. The method of claim 15, wherein the one or more filters balance the frequency response of the left and right speakers.
17. The one or more filters low-shelf and high-shelf filters, A bandpass filter and A bandstop filter, A peak notch filter, Low-pass and high-pass filters 16. The method of claim 15, comprising at least one of:
18. 13. The method of claim 12, wherein applying the delay or gain to the input audio signal comprises applying the delay or gain to one of a left channel or a right channel of the input audio signal.
19. 13. The method of claim 12, wherein the at least one of the delay and the gain adjusts for the determined spatial relationship being unequal distances from the listener's head to the left speaker and the right speaker.
20. 13. The method of claim 12, further comprising applying at least one of crosstalk compensation or crosstalk cancellation to the input audio signal.
21. 21. The method of claim 20, wherein the at least one of the crosstalk compensation or the crosstalk cancellation is applied to the input audio signal before applying the at least one of the delay and the gain to the input audio signal.
22. 13. The method of claim 12, further comprising gain adjusting spatial and non-spatial components of the input audio signal.
23. A non-transitory computer-readable medium storing a program for causing a processing circuit to perform the method of any one of claims 12 to 22.
Citation Information
Patent Citations
Sound image controller
JP1997046800A
Acoustic processor
JP2003092799A
Acoustic apparatus
JP2007028198A
Speaker array unit and microphone array unit
JP2007081642A
Sound reproduction system and method
JP2008113118A