Crosstalk Processing B-Chain
The system addresses geometric and physical asymmetries in audio systems by using spatial enhancement and b-chain processing to adjust frequency response, time alignment, and signal levels, improving the spatial soundstage in non-ideal configurations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- BOOMCLOUD 360 INC
- Filing Date
- 2026-01-26
- Publication Date
- 2026-04-21
AI Technical Summary
Existing audio systems often fail to provide an optimal listening experience due to geometric and physical asymmetries between speakers and listeners, such as mismatched distances, angles, and unequal frequency responses, leading to suboptimal spatial sound reproduction.
A system employing spatial enhancement and b-chain processing, including subband spatial processing, crosstalk cancellation, and b-chain processing, to adjust frequency response, time alignment, and signal levels, using filters and gain adjustments to balance audio signals across speakers.
Enhances the perceived spatial soundstage by correcting asymmetries, providing a more balanced and compelling listening experience even in non-ideal configurations.
Smart Images

Figure 2026067957000001_ABST
Abstract
Description
[Technical Field]
[0001] The subject matter described herein relates to audio signal processing, and more specifically, to addressing (geometric and physical) asymmetries when applying speech crosstalk cancellation to speakers. [Background technology]
[0002] Audio signals may be output to rendering systems and / or room acoustics that are not optimally configured. Figure 1A shows an example of an ideal transaural configuration, i.e., an ideal loudspeaker and listener configuration for a 2-channel stereo speaker system with one listener in an empty soundproof room. As shown in Figure 1A, listener 140 is in the ideal position (i.e., the "sweet spot") to experience the rendered audio from the left loudspeaker 110L and the right loudspeaker 110R, which is most accurately reproduced spatially and timbrally with respect to the content creator's original intent.
[0003] However, there are various situations in which the ideal "sweet spot" conditions are not met or cannot be achieved by the audio output device. As shown in Figure 1B, what has been described so far includes situations in which the listener 140's head position is laterally offset from the ideal "sweet spot" listening position between loudspeaker 110L and loudspeaker 110R. Or, as shown in Figure 1C, the listener 140 is in the ideal position, but the distances between each loudspeaker 110R and the listener 140's head position are not equal. Furthermore, as shown in Figure 1D, the listener 140 is in the ideal position, but the frequency response and amplitude response of loudspeaker 110L and loudspeaker 110R are not equal (i.e., the rendering system is "unmatched"). In another example, while the physical positions of the listener 140 and the loudspeakers 110L and 110R may be ideal, as shown in Figure 1E, one or more loudspeakers 110L and 110R may be offset as a rotation from the ideal angle with respect to the right loudspeaker 110R. [Overview of the project]
[0004] An exemplary embodiment relates to the b-chaining of a spatially enhanced audio signal to compensate for asymmetries between different speakers or environments. Examples of asymmetries may include time delays between one speaker and different listeners of another speaker, (perceived and intended) signal levels between one speaker and different listeners of another speaker, or frequency responses between one speaker and different listeners of another speaker.
[0005] In an exemplary embodiment, a system for enhancing input audio signals for left and right speakers includes a spatial enhancement processor and a b-chain processor. The spatial enhancement processor generates a spatially enhanced signal by gain-adjusting the spatial and non-spatial components of the input audio signal. The b-chain processor determines the asymmetry between the left and right speakers in terms of frequency response, time alignment, and signal levels at the listening position. The b-chain processor generates a left output channel for the left speaker and a right output channel for the right speaker by: applying N-band equalization to the spatial enhancement signal to adjust the asymmetry in the frequency response; applying delay to the spatial enhancement signal to adjust the asymmetry in the time alignment; and applying gain to the spatial enhancement signal to adjust the asymmetry in the signal levels.
[0006] In an embodiment, the b-chain processor applies N-band equalization by applying one or more filters to at least one of the enhanced channels in the left space and the enhanced channels in the right space. The one or more filters balance the frequency responses of the left and right speakers and may include at least one of the following filters: low-shelf filters and high-shelf filters, band-pass filters, band-stop filters, peak-notch filters, low-pass filters and high-pass filters.
[0007] In one embodiment, the b-chain processor adjusts at least one of the delay and gain in response to a change in the listening position.
[0008] The embodiment may include a non-temporary, computer-readable medium that stores instructions that, when executed by the processor, generate a spatial enhancement signal by gain-adjusting the spatial and non-spatial components of an input audio signal including a left input channel for the left speaker and a right input channel for the right speaker; determine the asymmetry between the left and right speakers; apply N-band equalization to the spatial enhancement signal to adjust the asymmetry of the frequency response; apply delay to the spatial enhancement signal to adjust the asymmetry of the time alignment; and apply gain to the spatial enhancement signal to adjust the asymmetry of the signal level, thereby generating a left output channel for the left speaker and a right output channel for the right speaker.
[0009] Embodiments may include a method for processing input audio signals for a left speaker and a right speaker. The method may include generating a spatial enhancement signal by gain-adjusting the spatial and non-spatial components of the input audio signal, which includes a left input channel for the left speaker and a right input channel for the right speaker; determining the asymmetry between the left and right speakers in frequency response, time alignment, and signal level at the listening position; generating a left output channel for the left speaker and a right output channel for the right speaker by applying N-band equalization to the spatial enhancement signal to adjust the asymmetry in frequency response; applying delay to the spatial enhancement signal to adjust the asymmetry in time alignment; and applying gain to the spatial enhancement signal to adjust the asymmetry in signal level. [Brief explanation of the drawing]
[0010] [Figure 1A] The positions of loudspeakers relative to listeners according to several embodiments are illustrated below. [Figure 1B] The positions of loudspeakers relative to listeners according to several embodiments are illustrated below. [Figure 1C] The positions of loudspeakers relative to listeners according to several embodiments are illustrated below. [Figure 1D] The positions of loudspeakers relative to listeners according to several embodiments are illustrated below. [Figure 1E] The positions of loudspeakers relative to listeners according to several embodiments are illustrated below. [Figure 2] This is a schematic block diagram of an audio processing system according to several embodiments. [Figure 3] This is a schematic block diagram of a spatial enhancement processor according to several embodiments. [Figure 4] This is a schematic block diagram of a subband space processor according to several embodiments. [Figure 5] This is a schematic block diagram of a crosstalk compensation processor according to several embodiments. [Figure 6] This is a schematic block diagram of a crosstalk cancellation processor according to several embodiments. [Figure 7] This is a schematic block diagram of a b-chain processor according to several embodiments. [Figure 8] This is a flowchart of a method for b-chaining input audio signals according to several embodiments. [Figure 9] Examples of less-than-ideal head positions and mismatched loudspeakers according to several embodiments are illustrated. [Figure 10A] Figure 9 illustrates the frequency response of a loudspeaker with an unideal head position and mismatch, as shown in several embodiments. [Figure 10B] Figure 9 illustrates the frequency response of a loudspeaker with an unideal head position and mismatch, as shown in several embodiments. [Figure 11] This is a schematic block diagram of a computer system according to several embodiments. [Modes for carrying out the invention]
[0011] The drawings and the detailed description depict various non-limiting embodiments for illustrative purposes only.
[0012] Here, embodiments are referred to in detail, and examples thereof are shown in the accompanying drawings. The following description shows certain specific details in order to provide a thorough understanding of the various embodiments. However, the described embodiments can be implemented without these specific details. In other instances, clear methods, procedures, components, circuits, and networks are not described in detail so as not to unnecessarily obscure aspects of the embodiments.
[0013] Embodiments of the present disclosure relate to an audio processing system that provides spatial enhancement and b-chain processing. Spatial enhancement may include applying sub-band spatial processing and crosstalk cancellation to an input audio signal. B-chain processing restores the perceived spatial sound stage of audio rendered transaurally on a non-ideally configured stereo loudspeaker rendering system.
[0014] For example, a digital audio system that can be used in a movie theater or personal headphones can be considered as two parts: an a-chain and a b-chain. For example, in an environment such as a movie theater, the a-chain typically includes the audio recording on the film print available for Dolby analog and further selection from digital formats such as Dolby Digital, DTS, and SDDS. Additionally, a device that acquires and processes audio from the film print to be ready for amplification is part of the a-chain.
[0015] A b-chain involves hardware and software systems for applying multi-channel volume control, equalization, time alignment, and amplification to loudspeakers to correct and / or minimize the effects of poorly configured rendering system installation, room acoustics, or listener positioning. B-chain processing can be configured analytically or parametrically to optimize the perceived quality of the listening experience, with the general objective of bringing the listener closer to an "ideal" experience.
[0016] Exemplary audio system Figure 2 is a schematic block diagram of an audio processing system 200 according to several embodiments. The audio processing system 200 applies subband spatial processing, crosstalk cancellation processing, and b-chain processing to an input audio signal X including a left input channel XL and a right input channel XR to generate an output audio signal O including a left output channel OL and a right output channel OR. The output audio signal O reconstructs the perceived spatial soundstage for the input audio signal X, which is transaurally rendered on a non-idealized stereo loudspeaker rendering system.
[0017] The audio processing system 200 includes a spatial enhancement processor 205 connected to a b-chain processor 240. The spatial enhancement processor 205 includes a subband spatial processor 210, a crosstalk compensation processor 220, and a crosstalk cancellation processor 230 connected to the subband spatial processor 210 and the crosstalk compensation processor 220.
[0018] The subband spatial processor 210 generates a spatially enhanced audio signal by gain-adjusting the mid and side subband components of the left input channel XL and the right input channel XR. The crosstalk compensation processor 220 performs crosstalk compensation to compensate for spectral defects or artifacts of the crosstalk cancellation applied by the crosstalk cancellation processor 230. The crosstalk cancellation processor 230 performs crosstalk cancellation on the combined output of the subband spatial processor 210 and the crosstalk compensation processor 220 to generate the left enhancement channel AL and the right enhancement channel AR. Additional details regarding the spatial enhancement processor 210 are described below with reference to Figures 3-6.
[0019] The b-chain processor 240 includes a speaker matching processor 250 connected to a delay and gain processor 260. Specifically, the b-chain processor 240 is capable of adjusting the overall delay time of the difference between loudspeakers 110L and 110R and the listener's head, the difference in (perceived and intended) signal levels between loudspeakers 110L and 110R and the listener's head, and the difference in frequency response between loudspeakers 110L and 110R and the listener's head.
[0020] The speaker matching processor 250 receives the left enhancement channel AL and the right enhancement channel AR and performs speaker balancing for devices that do not provide a matched speaker pair, such as a speaker pair in a mobile device or other types of left / right speaker pairs. In embodiments, the speaker matching processor 250 applies equalization and gain or attenuation to each of the left enhancement channel AL and the right enhancement channel AR to provide a spectrally perceptually balanced stereo image from the perspective of an ideal listening sweet spot. The delay and gain processor 260 receives the output of the speaker matching processor 250 and applies equalization and gain or attenuation to each of the channels AL and AR to perform time alignment and further perceptually balance the spatial image from a particular listener's head position, given actual physical asymmetries in the rendering / listening system (e.g., off-center head position and / or unequal distance between loudspeakers and head). The processing applied by the speaker matching processor 250 and the delay and gain processor 260 may be performed in different orders. Further details regarding the b-chain processor 240 are described below with reference to Figure 7.
[0021] Examples of spatial enhancement processors Figure 3 is a schematic block diagram of a spatial enhancement processor 205 according to several embodiments. The spatial enhancement processor 205 spatially enhances an input audio signal and performs crosstalk cancellation on the spatially enhanced audio signal. To this end, the spatial enhancement processor 205 receives an input audio signal X, which includes a left input channel XL and a right input channel XR. In embodiments, the input audio signal X is provided from a source component of a digital bitstream (e.g., PCM data). The source component may be a computer, a digital audio player, an optical disc player (e.g., DVD, CD, Blu-ray, etc.), a digital audio streamer, or another source of a digital audio signal. The spatial enhancement processor 205 processes input channels XL and XR to generate an output audio signal A, which includes two output channels AL and AR. Output audio signal A is the spatially enhanced audio signal of the input audio signal X with crosstalk compensation and crosstalk cancellation. Although not shown in Figure 3, the spatial enhancement processor 205 may further include an amplifier that amplifies the output audio signal A from the crosstalk cancellation processor 230 and provides signal A to output devices that convert output channels AL and AR into sound, such as loudspeakers 110R and 110R.
[0022] The spatial enhancement processor 205 includes a subband spatial processor 210, a crosstalk compensation processor 220, a combiner 222, and a crosstalk cancellation processor 230. The spatial enhancement processor 205 performs crosstalk compensation and subband spatial processing on the input audio input channels XL and XR, combines the results of the subband spatial processing with the results of the crosstalk compensation, and then performs crosstalk cancellation on the combined signal.
[0023] The subband spatial processor 210 includes a spatial frequency band divider 310, a spatial frequency band processor 320, and a spatial frequency band combiner 330. The spatial frequency band divider 310 is connected to input channels XL and XR and the spatial frequency band processor 320. The spatial frequency band divider 310 receives the left input channel XL and the right input channel XR and processes the input channels into a spatial (or "side") component Ys and a non-spatial (or "mid") component Ym. For example, the spatial component Ys can be generated based on the difference between the left input channel XL and the right input channel XR. The non-spatial component Ym can be generated based on the sum of the left input channel XL and the right input channel XR. The spatial frequency band divider 310 provides the spatial component Ys and the non-spatial component Ym to the spatial frequency band processor 320.
[0024] The spatial frequency band processor 320 is connected to the spatial frequency band divider 310 and the spatial frequency band combiner 330. The spatial frequency band processor 320 receives spatial Ys and non-spatial components Ym from the spatial frequency band divider 310 and enhances the received signals. In particular, the spatial frequency band processor 320 generates an enhanced spatial component Es from the spatial component Ys and an enhanced non-spatial component Em from the non-spatial component Ym.
[0025] For example, the spatial frequency band processor 320 applies a subband gain to the spatial component Ys to generate an enhanced spatial component Es, and applies a subband gain to the non-spatial component Ym to generate an enhanced non-spatial component Em. In some embodiments, as an addition or alternative, the spatial frequency band processor 320 provides a subband delay to the spatial component Ys to generate an enhanced spatial component Es, and a subband delay to the non-spatial component Ym to generate an enhanced non-spatial component Em. The subband gains and / or delays can be different for different (e.g., n) subbands of the spatial component Ys and the non-spatial component Ym, or they can be the same (e.g., for two or more subbands). The spatial frequency band processor 320 adjusts the gains and / or delays of different subbands of the spatial component Ys and the non-spatial component Ym with respect to each other to generate an enhanced spatial component Es and an enhanced non-spatial component Em. Next, the spatial frequency band processor 320 provides the enhanced spatial component Es and the enhanced non-spatial component Em to the spatial frequency band combiner 330.
[0026] The spatial frequency band combiner 330 is connected to the spatial frequency band processor 320, which in turn is connected to the combiner 222. The spatial frequency band combiner 330 receives an enhanced spatial component Es and an enhanced non-spatial component Em from the spatial frequency band processor 320, and combines the enhanced spatial component Es and the enhanced non-spatial component Em into a left spatial enhancement channel EL and a right spatial enhancement channel ER. For example, the left spatial enhancement channel EL can be generated based on the sum of the enhanced spatial component Es and the enhanced non-spatial component Em, and the right spatial enhancement channel ER can be generated based on the difference between the enhanced non-spatial component Em and the enhanced spatial component Es. The spatial frequency band combiner 330 provides the left spatial enhancement channel EL and the right spatial enhancement channel ER to the combiner 222.
[0027] The crosstalk compensation processor 220 performs crosstalk compensation to compensate for spectral defects and artifacts in crosstalk cancellation. The crosstalk compensation processor 220 receives input channels XL and XR and performs processing to compensate for artifacts in subsequent crosstalk cancellation of the enhanced non-spatial component Em and enhanced spatial component Es, which are performed by the crosstalk cancellation processor 230. In some embodiments, the crosstalk compensation processor 220 may perform enhancement on the non-spatial component Xm and spatial component Xs by applying filters that generate a crosstalk compensation signal Z including a left crosstalk compensation channel ZL and a right crosstalk compensation channel ZR. In other embodiments, the crosstalk compensation processor 220 may perform enhancement only on the non-spatial component Xm.
[0028] The combiner 222 generates a left enhancement compensation channel TL by combining the left spatial enhancement channel EL with the left crosstalk compensation channel ZL, and generates a right enhancement compensation channel TR by combining the right spatial enhancement channel ER with the right crosstalk compensation channel ZR. The combiner 222 is connected to the crosstalk cancellation processor 230 and provides the left enhancement compensation channel TL and the right enhancement compensation channel TR to the crosstalk cancellation processor 230.
[0029] The crosstalk cancellation processor 230 receives the left enhancement compensation channel TL and the right enhancement compensation channel TR, performs crosstalk cancellation on channels TL and TR, and generates an output audio signal A that includes the left output channel OL and the right output channel OR.
[0030] Further details regarding the subband spatial processor 210 are described below with reference to Figure 4, further details regarding the crosstalk compensation processor 220 are described below with reference to Figure 5, and further details regarding the crosstalk cancellation processor 230 are described below with reference to Figure 6.
[0031] Figure 4 is a schematic block diagram of a subband spatial processor 210 according to several embodiments. The subband spatial processor 210 includes a spatial frequency band divider 310, a spatial frequency band processor 320, and a spatial frequency band combiner 330. The spatial frequency band divider 310 is connected to the spatial frequency band processor 320, and the spatial frequency band processor 320 is connected to the spatial frequency band combiner 330.
[0032] The spatial frequency band divider 310 includes an L / RM / S converter 402 that receives the left input channel XL and the right input channel XR and converts these inputs into a spatial component Xs and a non-spatial component Xm. The spatial component Xs may be generated by subtracting the left input channel XL and the right input channel XR. The non-spatial component Xm may be generated by adding the left input channel XL and the right input channel XR.
[0033] The spatial frequency band processor 320 receives a non-spatial component Xm and applies a set of subband filters to generate a non-spatial enhancement subband component Em. Furthermore, the spatial frequency band processor 320 receives a spatial subband component Xs and applies a set of subband filters to generate a non-spatial enhancement subband component Em. The subband filters can include various combinations of peak filters, notch filters, low-pass filters, high-pass filters, low-shelf filters, high-shelf filters, band-pass filters, band-stop filters, and / or all-pass filters.
[0034] In one embodiment, the spatial frequency band processor 320 includes subband filters for each of the n frequency subbands of the non-spatial component Xm, and subband filters for each of the n frequency subbands of the spatial component Xs. For example, for n=4, the spatial frequency band processor 320 includes a series of subband filters for the non-spatial component Xm, including a mid-equalization (EQ) filter 404(1) for subband (1), a mid-EQ filter 404(2) for subband (2), a mid-EQ filter 404(3) for subband (3), and a mid-EQ filter 404(4) for subband (4). Each mid-EQ filter 404 applies a filter to the frequency subband portion of the non-spatial component Xm to generate an enhanced non-spatial component Em.
[0035] Furthermore, the spatial frequency band processor 320 includes a series of subband filters for the frequency subbands of the spatial component Xs, including a side equalization (EQ) filter 406(1) for subband (1), a side EQ filter 406(2) for subband (2), a side EQ filter 406(3) for subband, and a side EQ filter 406(4) for subband. Each side EQ filter 406 applies the filter to the frequency subband portion of the spatial component Xs to generate an enhanced spatial component Es.
[0036] Each of the n frequency subbands relating to the non-spatial component Xm and spatial component Xs may correspond to a range of frequencies. For example, frequency subband (1) corresponds to 0–300 Hz, frequency subband (2) to 300–510 Hz, frequency subband (3) to 510–2700 Hz, and frequency subband (4) to 2700 Hz–Nyquist frequency. In some embodiments, the n frequency subbands are an integrated set of important bands. The important bands can be determined using a corpus of audible frequency samples from various musical genres. The long-term averaged energy ratio of the intermediate and side components in the critical band of the 24-Burk scale is determined from the samples. Then, continuous frequency bands with similar long-term averaged ratios are grouped to form a set of important bands. The range of the frequency subbands and the number of frequency subbands can be adjusted. In embodiments, each of the n frequency subbands may contain a set of important bands.
[0037] In the embodiment, the mid-EQ filter 404 or the side-EQ filter-406 may include a fourth-order filter (biquad filter) having a transfer function defined by Equation 1.
[0038]
number
[0039] However, z is a complex variable, and a0, a1, a2, b0, b1, and b2 are digital filter coefficients. The filter may be implemented using a direct form I topology defined by equation 2.
[0040]
number
[0041] Here, X is the input vector and Y is the output. Other topologies may have advantages for a particular processor, depending on the maximum word length and saturation behavior.
[0042] Next, it is possible to implement any second-order filter with real input and output values using the fourth order. To design a discrete-time filter, the continuous-time filter is designed and then converted to discrete-time via a bilinear transform. Furthermore, compensation for any resulting shifts in center frequency and bandwidth can be achieved using frequency distortion.
[0043] For example, a peak filter may include an S-plane transfer function defined by Equation 3.
[0044]
number
[0045] Here, S is a complex variable, A is the peak amplitude, and Q is the filter "quality" (which is canonically derived as follows):
[0046]
number
[0047] ) The digital filter coefficients are as follows:
[0048]
number
[0049] However, ω0 is the center frequency of the filter in radians and
[0050]
number
[0051] This is how it is represented.
[0052] The spatial frequency band combiner 330 receives the mid- and side components, applies gain to each component, and converts the mid- and side components into left and right channels. For example, the spatial frequency band combiner 330 receives an enhanced non-spatial component Em and an enhanced spatial component Es, and performs global mid-and-side gain before converting the enhanced non-spatial component Em and the enhanced spatial component Es into a left spatial enhancement channel EL and a right spatial enhancement channel ER.
[0053] Specifically, the spatial frequency band combiner 330 includes a global mid-gain 408, a global side-gain 410, and an M / SL / R converter 412 connected to the global mid-gain 408 and the global side-gain 410. The global mid-gain 408 receives an enhanced non-spatial component Em and applies its gain, while the global side-gain 410 receives an enhanced non-spatial component Es and applies its gain. The M / SL / R converter 412 receives the enhanced non-spatial component Em from the global mid-gain 408 and the enhanced spatial component Es from the global side-gain 410, and converts these inputs into a left spatial enhancement channel EL and a right spatial enhancement channel ER.
[0054] Figure 5 is a schematic block diagram of a crosstalk compensation processor 220 according to several embodiments. The crosstalk compensation processor 220 receives left and right input channels and generates left and right output channels by applying crosstalk compensation on the input channels. The crosstalk compensation processor 220 includes an L / RM / S converter 502, a mid-component processor 520, a side-component processor 530, and an M / SL / R converter 514.
[0055] When the crosstalk compensation processor 220 is part of audio system 202, 400, 500, or 504, the crosstalk compensation processor 220 receives input channels XL and XR and performs preprocessing to generate the left crosstalk compensation channel ZL and the right crosstalk compensation channel ZR. Channels ZL and ZR may be used to compensate for artifacts of crosstalk processing, such as crosstalk cancellation or simulation. The L / RM / S converter 502 receives the left input audio channel XL and the right input audio channel XR and generates the non-spatial component Xm and the spatial component Xs of the input channels XL and XR. Generally, the left and right channels may be added together to generate the non-spatial components of the left and right channels, and subtracted together to generate the spatial components of the left and right channels.
[0056] The mid-component processor 520 is equipped with multiple filters 540, such as m mid-filters 540(a), 540(b), and 540(m). Here, each of the m mid-filters 540 processes one of m frequency bands of a non-spatial component Xm and a spatial component Xs. The mid-component processor 520 generates a mid-crosstalk compensation channel Zm by processing the non-spatial component Xm. In an embodiment, the mid-filter 540 is configured using a frequency response plot of the non-spatial component Xm obtained through crosstalk processing via simulation. Furthermore, by analyzing the frequency response plot, spectral distortions such as peaks and troughs in the frequency response plot that occur as artifacts of crosstalk processing can be estimated beyond a preset threshold (e.g., 10 dB). These artifacts mainly result from the sum of the delayed and inverted contra-side signal and the corresponding ipsi-side signal in the crosstalk processing, thus effectively introducing a comb-filter-like frequency response into the final rendering result. The mid-crosstalk compensation channel Zm is generated by the mid-component processor 520 and can compensate for estimated peaks or troughs, where each of the m frequency bands corresponds to a peak or trough. Specifically, based on the specific delay, filtering frequency, and gain applied in the crosstalk processing, the peaks and troughs shift up and down in the frequency response, causing amplification or attenuation of energy in specific regions of the spectrum. Each mid-filter 540 can be configured to tune to one or more peaks and troughs.
[0057] The side component processor 530 includes multiple filters 550, such as m side filters 550(a), 550(b) to 550(m). The side component processor 530 generates a side crosstalk compensation channel Zs by processing the spatial component Xs. In embodiments, the frequency response plot of the spatial component Xs due to the crosstalk processing can be obtained by simulation. By analyzing the frequency response plot, spectral distortions such as peaks and troughs in the frequency response plot that occur as artifacts of the crosstalk processing can be estimated beyond a preset threshold (e.g., 10 dB). The side crosstalk compensation channel Zs can be generated by the side component processor 530 to compensate for the estimated peaks or troughs. Specifically, based on the specific delay, filtering frequency, and gain applied in the crosstalk processing, the peaks and troughs in the frequency response shift up and down, causing amplification or attenuation of energy in specific regions of the spectrum. Each side filter 550 can be configured to tune to one or more peaks and troughs. In some embodiments, the mid-component processor 520 and the side-component processor 530 may include different numbers of filters.
[0058] In the embodiment, the mid-filter 540 and the side filter 550 may include a fourth-order filter having a transfer function defined by Equation 4.
[0059]
number
[0060] Here, z is a complex variable, and a0, a1, a2, b0, b1, and b2 are digital filter coefficients. One way to implement such a filter is the direct form I topology defined in Equation 5.
[0061]
number
[0062] However, X is the input vector and Y is the output. Other topologies are used depending on the maximum word length and saturation behavior.
[0063] Subsequently, a biquad can be used to implement a second-order filter with real-value inputs and outputs. To design a discrete-time filter, a continuous-time filter is designed and then converted to discrete-time using a bilinear transform. Furthermore, the shift in center frequency and bandwidth can be compensated for using frequency distortion.
[0064] For example, a peak filter is defined by equation 6 and has a complex plane transfer function.
[0065]
number
[0066] However, s is a complex variable, A is the peak amplitude, Q is the filter "quality", and the digital filter coefficients are defined as follows:
[0067]
number
[0068] However, ω0 is the center frequency of the filter in radians and
[0069]
number
[0070] This is how it is represented.
[0071] Furthermore, the filter quality Q can be defined by Equation 7.
[0072]
number
[0073] however,
[0074]
number
[0075] is bandwidth, f c This is the center frequency.
[0076] The M / SL / R converter 514 receives the mid-crosstalk compensation channel Zm and the side-crosstalk compensation channel Zs, and generates the left-crosstalk compensation channel ZL and the right-crosstalk compensation channel ZR. Generally, the mid-channel and the side-channel may be added together to generate the left channels of the mid-component and side-component, and the mid-channel and the side-channel may be subtracted together to generate the right channels of the mid-component and side-component.
[0077] Figure 6 is a schematic block diagram of a crosstalk cancellation processor 230 according to several embodiments. The crosstalk cancellation processor 230 receives the left enhancement compensation channel TL and the right enhancement compensation channel TR from the combiner 222, performs crosstalk cancellation on channels TL and TR, and generates the left output channel AL and the right output channel AR.
[0078] The crosstalk cancellation processor 230 includes an in-out band divider 610, inverters 620 and 622, opposite-side estimators 630 and 640, combiners 650 and 652, and an in-out band combiner 660. These components divide the input channels TL and TR into in-band and out-of-band components, and perform crosstalk cancellation on the in-band components to generate output channels AL and AR.
[0079] By splitting the input audio signal T into different frequency band components and performing crosstalk cancellation on a selective component (such as an in-band component), crosstalk cancellation can be performed on a specific frequency band while eliminating degradation in other frequency bands. If crosstalk cancellation is performed without splitting the input audio signal T into different frequency bands, the audio signal after crosstalk cancellation may show significant attenuation or amplification of non-spatial and spatial components at low frequencies (such as below 350Hz), high frequencies (such as above 12000Hz), or both. By selectively performing crosstalk cancellation in the in-band (such as 250Hz to 14000Hz) where most of the influential spatial cues are present, it is possible to maintain a balanced overall energy across the entire spectrum of the mix, especially in non-spatial components.
[0080] The input / output band divider 610 separates the input channels TL and TR into in-band channels TL,In and TR,In, and out-band channels TL,Out and TR,Out, respectively. In particular, the input / output band divider 610 divides the left enhancement compensation channel TL into the left in-band channel TL,In and the left out-band channel TL,Out. Similarly, the input / output band divider 610 divides the right enhancement compensation channel TR into the right in-band channel TR,In and the right out-band channel TR,Out. Each in-band channel encompasses a portion of each input channel corresponding to a frequency range, such as 250Hz to 14kHz. The frequency band range can be adjusted according to speaker parameters, etc.
[0081] Inverter 620 and the opposite channel estimator 630 work together to generate a left opposite channel cancellation component SL to compensate for the opposite channel sound component caused by the left in-band channel TL,In. Similarly, inverter 622 and the opposite channel estimator 640 work together to generate a right opposite channel cancellation component SR to compensate for the opposite channel sound component caused by the right in-band channel TR,In.
[0082] In one approach, the inverter 620 receives the in-band channel TL,In, inverts the polarity of the received in-band channel TL,In to generate an inverted in-band channel TL,In'. The counter-side estimator 630 receives the inverted in-band channel TL,In' and, through filtering, extracts a portion of the inverted in-band channel TL,In' corresponding to the counter-side sound component. Since the filtering is performed on the inverted in-band channel TL,In', the portion extracted by the counter-side estimator 630 is the inverse of the portion of the in-band channel TL,In attributable to the counter-side sound component. Thus, the portion extracted by the counter-side estimator 630 becomes the left counter-cancellation component SL, which can be added to the opposite in-band channel TR,In to reduce the counter-side sound component attributable to the in-band channel TL,In. In some embodiments, the inverter 620 and the counter-side estimator 630 are implemented in a different order.
[0083] The inverter 622 and the opposite side estimator 640 perform similar operations with respect to the in-band channel TR,In to generate the right opposite side cancellation component SR. Therefore, a detailed description thereof is omitted in this specification for the sake of brevity.
[0084] In one exemplary implementation, the counter-side estimator 630 includes a filter 632, an amplifier 634, and a delay unit 636. The filter 632 receives an inverted input channel TL,In' and extracts a portion of the inverted in-band channel TL,In' corresponding to the counter-side sound component through a filtering function. An example of filter implementation is a Notch or Highshelf filter using a center frequency selected between 5000 and 10000 Hz and a Q selected between 0.5 and 1.0. The gain in decibels (GdB) may be derived from Equation 8. G dB = -3.0 - log 1.333 (D) Equation 8 However, D is the delay amount by the delay unit 636 in the sample, for example, a sampling rate of 48 KHz.
[0085] Another implementation method is a low-pass filter, where the corner frequency is selected in the range of 5000 to 10000 Hz, and Q is selected in the range of 0.5 to 1.0. Furthermore, the amplifier 634 amplifies the extraction part by the corresponding gain factor G L,In to generate the left-side cancellation component SL by delaying the amplified output from the amplifier 634 according to the delay function D of the delay unit 636. The opposite-side estimator 640 includes a filter 642, an amplifier 644, and a delay unit 646. This unit performs a similar operation in the inverted in-band channel T R,In ’ to generate the right-side cancellation component SR. In one example, the opposite-side estimators 630, 640 generate the left-side cancellation component SL and the right-side cancellation component SR according to the following equations. S L = D[G L,In * F[T L,In ’]] Equation 9 S R = D[G R,In * F[T R,In ’]] Equation 10 However, F[] is a filter function and D[] is a delay function.
[0086] The crosstalk cancellation setting can be determined by the speaker parameters. For example, according to the angle of the two speakers 110 with respect to the listener, the center frequency of the filter, the delay amount, the amplifier gain, and the filter gain can be determined. In some embodiments, the values between the speaker angles are used to interpolate other values.
[0087] Combiner 650 combines the right contraband cancellation component SR with the left in-band channel TL,In to generate the left in-band compensation channel UL, and combiner 652 combines the left contraband cancellation component SL with the right in-band channel TR,In to generate the right in-band compensation channel UR. In-out-band combiner 660 combines the left in-band compensation channel UL with the out-of-band channels TL,Out to generate the left output channel AL, and combines the right in-band compensation channel UR with the out-of-band channels TR,Out to generate the right output channel AR.
[0088] Therefore, the left output channel AL includes a right contra-side cancellation component SR that corresponds to the inverse of a portion of the in-band channel TR,In attributable to the contra-side sound, and the right output channel AR includes a left contra-side cancellation component SL that corresponds to the inverse of a portion of the in-band channel TL,In attributable to the contra-side sound. In this configuration, the wavefront of the ipsilateral sound component output from loudspeaker 110R corresponding to the right output channel AR that reaches the right ear can cancel the wavefront of the contra-side sound component output from loudspeaker 110L corresponding to the left output channel AL. Similarly, the wavefront of the ipsilateral sound component output from loudspeaker 110L corresponding to the left output channel AL that reaches the left ear can cancel the wavefront of the contra-side sound component output from loudspeaker 110R corresponding to the right output channel AR. Therefore, the contra-side sound component can be reduced to enhance spatial detectability.
[0089] Exemplary b-chain processor Figure 7 is a schematic block diagram of a b-chain processor 240 according to several embodiments. The b-chain processor 240 includes a speaker matching processor 250 and a delay and gain processor 260. The speaker matching processor 250 includes an N-band equalizer (EQ) 702 connected to the left amplifier 704 and the right amplifier 706. The delay and gain processor 260 includes a left delay 708 connected to the left amplifier 712 and a right delay 710 connected to the right amplifier 714.
[0090] As shown in Figures 1A-1E, assuming that the orientation of the listener 140 remains fixed toward the center of an ideal spatial image (e.g., a virtual lateral center of the sound field, given symmetry, matching, and equidistant loudspeakers), the transformation relationship between the ideal spatial image and the actually rendered spatial image can be explained based on (a) the overall time delay between one speaker and the listener 140 being different from that of the other speaker, (b) the (perceived and intended) signal levels between one speaker and the listener 140 being different from those of the other speaker, and (c) the frequency response between one speaker and the listener 140 being different from those of the other speaker.
[0091] The b-chain processor 240 corrects the aforementioned relative differences in delay, signal level, and frequency response, resulting in a nearly ideal spatial image reconstruction, as if the listener 140 (e.g., head position) and / or rendering system were ideally configured.
[0092] The b-chain processor 240 receives audio signal A, which includes left enhancement channel AL and right enhancement channel AR, as input from the spatial enhancement processor 205. The input to the b-chain processor 240 may include any stereo audio stream that has been transaurally processed for a given listener / speaker configuration under ideal conditions (as illustrated in Figure 1A). If audio signal A has no spatial asymmetry and no other anomalies are present in the system, the spatial enhancement processor 205 provides the listener 140 with a dramatically enhanced sound field. However, if asymmetry is present in the system, as described above and illustrated in Figures 1B-1E, the b-chain processor 240 may be applied to maintain an enhanced sound field under non-ideal conditions.
[0093] While an ideal listener / speaker configuration includes a pair of loudspeakers with the left and right speakers at the same distance from the listener's head, many real-world setups fail to meet these criteria, resulting in a flawed stereo listening experience. For example, a mobile device might include a front-facing earpiece loudspeaker with a limited bandwidth (e.g., 1000-8000Hz frequency response) and a micro-loudspeaker oriented orthogonally (downward or sideways) (e.g., 200-20000Hz frequency response). Here, the speaker system is mismatched in two ways: the performance characteristics of the audio drivers (e.g., signal level, frequency response, etc.) differ, and the time alignment with respect to the "ideal" listener position is mismatched due to the non-parallel orientation of the speakers. Another example is a listener using a stereo desktop speaker system who does not position either the loudspeakers or the speakers themselves in an ideal configuration (e.g., as shown in Figures 1B, 1C, or 1E). Therefore, the b-chain processor 240 supports the adjustment of the characteristics of each channel, addressing the associated system-specific asymmetries, and ultimately resulting in a more perceptually compelling transoral sound field.
[0094] After spatial enhancement processing or other processing has been applied to the stereo input signal X, which has been adjusted under the assumption of an ideally configured system (i.e., a sweet spot listener, matching, symmetrically placed loudspeakers, etc.), the speaker matching processor 250 provides practical loudspeaker balancing to devices that do not supply matched speaker pairs, as is the case in most mobile devices. The N-band EQ 702 of the speaker matching processor 250 receives the left enhancement channel AL and the right enhancement channel AR and applies equalization to channels AL and AR, respectively.
[0095] In embodiments, the N-band EQ702 provides various EQ filter types, such as low-shelf filters, high-shelf filters, band-pass filters, band-stop filters, peak-notch filters, low-pass filters, and high-pass filters. For example, if one loudspeaker in a stereo pair is at an angle away from the ideal listener sweet spot, that loudspeaker will exhibit significant high-frequency attenuation from the listener sweet spot. One or more bands of the N-band EQ702 can be applied to the loudspeaker channel to restore high-frequency energy when viewed from the sweet spot (e.g., via a high-shelf filter), achieving a matching closer to the characteristics of the other forward loudspeaker. In another scenario, if both loudspeakers face forward but one loudspeaker has significantly different frequency characteristics, EQ tuning can be applied to both the left and right channels to balance the spectrum between the two. Applying the above adjustments can be equivalent to "rotating" the desired speaker to match the orientation of the other forward-facing speaker. In embodiments, the N-band EQ702 includes filters for each of n bands that are processed independently. The number of bandwidths may vary. In some embodiments, the number of bandwidths corresponds to the subbands of the subband spatial processing.
[0096] In some embodiments, speaker asymmetry may be predefined for a particular set of speakers by known asymmetries used as a basis for selecting the parameters of the N-band EQ702. In another example, speaker asymmetry may be determined based on speaker testing, such as by using a test audio signal, recording the sound produced by the speaker from the signal, and analyzing the recorded sound.
[0097] The left amplifier 704 is connected to the N-band EQ 702 to receive the left channel, and the right amplifier 706 is connected to the N-band EQ 702 to receive the right channel. Amplifiers 704 and 706 address asymmetry in the loudspeaker's loudness and dynamic range capabilities by adjusting the output gain on one or both channels. This is particularly useful for balancing loudness offsets at the distance of the loudspeaker from the listening position and for balancing mismatched loudspeaker pairs with significantly different sound pressure level (SPL) output characteristics.
[0098] The delay and gain processor 260 receives the left and right output channels of the speaker matching processor 250 and applies time delay and gain or attenuation to one or more channels. To this end, the delay and gain processor 260 includes a left delay 708 that receives the left channel output from the speaker matching processor 250 and applies a time delay, and a left amplifier 712 that applies gain or attenuation to the left channel to produce a left output channel OR. Furthermore, the delay and gain processor 260 includes a right delay 710 that receives the right channel output from the speaker matching processor 250 and applies a time delay, and a right amplifier 714 that applies gain or attenuation to the right channel to produce a right output channel OR. As previously stated, the speaker matching processor 250 ignores time-based asymmetry present in the actual configuration, focusing on perceptually balancing the left / right spatial image from the perspective of an ideal listener “sweet spot” and providing balanced SPL and frequency response to each driver from that position. After this speaker matching is achieved, the delay and gain processor 260 time-aligns the spatial image from a specific listener's head position, given the actual physical asymmetries of the rendering / listening system (e.g., off-center head position and / or unequal speaker-to-head distances), and further balances the perceptual sound.
[0099] The delay and gain values applied by the delay and gain processor 260 may be set to accommodate static system configurations, such as a mobile phone using orthogonally oriented loudspeakers, or to accommodate listeners who are laterally offset from the ideal sweet spot in front of the speakers, such as a home theater soundbar.
[0100] Furthermore, the delay and gain values applied by the delay and gain processor 260 may be dynamically adjusted based on the changing spatial relationship between the listener's head and the loudspeaker, as may occur in game scenarios that use physical movement as an element of gameplay (e.g., position tracking using a depth camera, such as in games or artificial reality systems). In embodiments, the audio processing system includes a camera, a light sensor, a proximity sensor, or other suitable device used to determine the position of the listener's head relative to the speaker. The determined position of the user's head may be used to determine the delay and gain values of the delay and gain processor 260.
[0101] The audio analysis routine provides appropriate speaker delays and gains used to configure the b-chain processor 240, which can be reduced to a time-aligned and perceptually balanced left / right stereo image. In embodiments where measurable data cannot be obtained from such analysis methods, intuitive user manual control or automatic control via computer vision or other sensor inputs can be achieved using mappings as defined by equations 11 and 12 below.
[0102]
number
[0103]
number
[0104] However, delayDelta and delay are in milliseconds, and gain is in decibels. The column vectors for delay and gain are assumed to have the first component related to the left channel and the second component related to the right channel. Therefore,
[0105]
number
[0106] 0 indicates that the delay of the left speaker is greater than or equal to the delay of the right speaker, while delayDelta<0 indicates that the delay of the left speaker is less than the delay of the right speaker.
[0107] In some embodiments, instead of applying attenuation to a channel, the same amount of gain may be applied to the opposite channel, or to a combination of gain applied to one channel and attenuation applied to the other channel. For example, gain may be applied to the left channel rather than to the left channel's attenuation. For close-range listening, such as that occurring in mobile devices, desktop PCs and console games, and home theater scenarios, the difference in distance between the listener's position and each speaker is small enough that the SPL delta between the listener's position and each speaker is small enough that any of the above mappings would help successfully restore a transaural spatial image while maintaining an overall acceptable sound field size compared to an ideal listener / speaker configuration.
[0108] Exemplary audio system processing Figure 8 is a flowchart of a method 800 for processing an input audio signal according to several embodiments. The method 800 may have fewer or additional steps, and the steps may be performed in a different order.
[0109] The audio processing system 200 (e.g., a spatial enhancement processor 205) enhances the input audio signal to generate an enhanced signal 802. The enhancement may include spatial enhancement. For example, the spatial enhancement processor 205 applies subband spatial processing, crosstalk compensation processing, and crosstalk cancellation processing to an input audio signal X, which includes a left input channel XL and a right input channel XR, to generate an enhanced signal A, which includes a left enhancement channel AL and a right enhancement channel AR. Here, the audio processing system 200 applies spatial enhancement by gain-adjusting the mid (non-spatial) and side (spatial) subband components of the input audio signal X, and the enhanced signal A is referred to as the "spatially enhanced signal". The audio processing system 200 may perform other types of enhancement to generate the enhanced signal A.
[0110] The audio processing system 200 (for example, the N-band EQ 702 of the speaker matching processor 250 of the b-chain processor 240) applies N-band equalization to the enhancement signal A to adjust the asymmetry of the frequency response between the left and right speakers 804. The N-band EQ 702 may apply one or more filters to the left enhancement channel AL, the right enhancement channel AR, or both the left channel AL and the right channel AR. One or more filters applied to the left enhancement channel AL and / or the right enhancement channel AR balance the frequency response for the left and right speakers. In embodiments, balancing the frequency response may be used to adjust the rotational offset from the ideal angle of the left and right speakers. In embodiments, the N-band EQ 702 adjusts the asymmetry of the left and right speakers and determines the filter parameters for applying the N-band EQ based on the determined asymmetry.
[0111] The audio processing system 200 (e.g., left amplifier 704 and / or right amplifier 706) applies gain to at least one of the left enhancement channel AL and the right enhancement channel AR to adjust for asymmetry between the left and right speakers in terms of signal level. The applied gain may be a positive or negative gain (also called attenuation) to address asymmetry in the loudness and dynamic range capabilities of the speakers or in mismatched speaker pairs having different sound pressure level (SPL) output characteristics.
[0112] The audio processing system 200 (e.g., the delay and gain processor 260 of the b-chain processor 240) applies delay and gain to enhancement signal A to adjust the listening position 808. The listening position may include the user's position relative to the left and right speakers. The user refers to the listener of the speakers. The delay and gain time-align and further perceptually balance the spatial image output from the speaker matching processor 250 to the listener's position, given the actual physical asymmetry of the rendering / listen system (e.g., off-center head position and / or unequal distance between the loudspeaker and the head). For example, to the left enhancement channel AL, the left delay 708 may apply delay and the left amplifier 712 may apply gain. To the right enhancement channel AR, the right delay 710 may apply delay and the right amplifier 714 may apply gain. In the embodiment, the delay may be applied to one of the left enhancement channel AL or the right enhancement channel AR, and the gain may be applied to one of the left enhancement channel AL or the right enhancement channel AR.
[0113] The audio processing system 200 (for example, the delay and gain processor 260 of the b-chain processor 240) adjusts at least one of the delay and gain in response to changes in the listening position. For example, the user's spatial position relative to the left and right speakers may change. The audio processing system 200 monitors the listener's position over time, determines the gain and delay applied to the enhancement signal O based on the listener's position, and adjusts the delay and gain applied to the enhancement signal O in response to changes in the listener's position over time to generate the left output channel OL and the right output channel OR.
[0114] Various asymmetric adjustments may be performed in different orders. For example, adjustments for asymmetry in speaker characteristics (e.g., frequency response) may be performed before, after, or in relation to adjustments for asymmetry in listening position with respect to speaker position or orientation. The audio processing system determines the asymmetry between the left and right speakers in frequency response, time alignment, and listening position signal levels, and generates the left output channel of the left speaker and the right output channel of the right speaker by applying N-band equalization to the spatial enhancement signal to adjust the asymmetry in frequency response between the left and right speakers, by applying delay to the spatial enhancement signal to adjust the asymmetry in time alignment, and by applying gain to the spatial enhancement signal to adjust the asymmetry in signal levels.
[0115] In some embodiments, rather than applying multiple gains or delays to adjust for different causes of asymmetry (e.g., speaker characteristics or listening position), a single gain and a single delay are used to adjust for multiple types of asymmetry resulting from differences in gain or time delay between speakers and culminating in a favorable listening position. However, separating the processing for speaker asymmetry and listening position asymmetry to reduce processing needs can be beneficial. For example, if the frequency response of a speaker is known, the same filter value may be used for speaker tuning, while separate time delay and signal level adjustments are made for changes in listening position (e.g., user movement).
[0116] Figure 9 illustrates a less-than-ideal head position and mismatched loudspeakers according to several embodiments. The listener 140 is at different distances from the left speaker 910L and the right speaker 910R. Furthermore, the frequency and / or amplitude characteristics of speakers 910L and 910R are not equivalent. Figure 10A illustrates the frequency response of the left speaker 910L, and Figure 10B illustrates the frequency response of the right speaker 910R.
[0117] As shown in Figures 9, 10A and 10B, in order to correct the speaker asymmetry of speakers 910L and 910R and the position of listener 140 with respect to speakers 910L and 910R respectively, the components of the b-chain processor 240 may use the following configuration: The N-band EQ 702 may apply a high-shelf filter with a cutoff frequency of 4,500 Hz, a Q value of 0.7, and a slope of -6 dB to the left enhancement channel AL, and a high-shelf filter with a cutoff frequency of 6,000 Hz, a Q value of 0.5, and a slope of +3 dB to the right enhancement channel AR. The left delay 708 may apply a delay of 0 milliseconds, the right delay 710 may apply a delay of 0.27 milliseconds, the left amplifier 712 may apply a gain of 0 dB, and the right amplifier 714 may apply a gain of -0.40625 dB.
[0118] Exemplary computing system It should be noted that the systems and processes described herein may be embodied in embedded electronic circuits or electronic systems. Furthermore, the systems and processes may be embodied in computing systems that include one or more processing systems (e.g., digital signal processors), memory (e.g., programmed read-only memory or programmable solid-state memory), or other circuits such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs).
[0119] Figure 11 illustrates an example of a computer system 1100 according to one embodiment. An audio system 200 may be implemented on the system 1100. At least one processor 1102 connected to a chipset 1104 is illustrated. The chipset 1104 includes a memory controller hub 1120 and an I / O (input / output) controller hub 1122. Memory 1106 and a graphics adapter 1112 are connected to the memory controller hub 1120, and a display device 1118 is connected to the graphics adapter 1112. A storage device 1108, a keyboard 1110, a pointing device 1114, and a network adapter 1116 are connected to the I / O controller hub 1122. Other embodiments of the computer 1100 have different architectures. For example, according to some embodiments, memory 1106 is directly connected to the processor 1102.
[0120] The storage device 1108 includes one or more temporarily readable storage media, such as a hard drive, a compact disc read-only memory (CD-ROM), a DVD, or a solid-state memory device. Memory 1106 holds instructions and data used by the processor 1102. For example, memory 1106 may store instructions that, when executed by the processor 1102, cause or configure the processor 1102 to perform or perform functions described herein, such as method 800. The pointing device 1114 is used in conjunction with the keyboard 1110 to input data into the computer system 1100. The graphics adapter 1112 displays images and other information on the display device 1118. In embodiments, the display device 1118 includes touchscreen capabilities for receiving user input and selections. The network adapter 1116 connects the computer system 1100 to a network. Some embodiments of the computer 1100 have different and / or other components than those shown in Figure 11. For example, computer system 1100 may be a server without a display device, keyboard, and other components, or it may use other types of input devices.
[0121] Additional considerations The disclosed configuration may include several benefits and / or advantages. For example, the input signal may be output to mismatched loudspeakers while maintaining or enhancing the spatial sense of the sound field. A high-quality listening experience can be achieved even when the speakers are mismatched, and even when the listener is not in the ideal listening position relative to the speakers.
[0122] A reader of this disclosure will still recognize additional alternative embodiments of the principles disclosed herein. Therefore, while specific embodiments and applications are illustrated and described, it should be understood that the disclosed embodiments are not limited to the structures and components as disclosed herein. Various modifications, changes, and variations, which will be obvious to a person skilled in the art, may be made to the arrangement, operation, and details of the methods and apparatus disclosed herein without departing from the scope described herein.
[0123] The steps, operations, or processes described herein may be performed or implemented by one or more hardware or software modules, either alone or in combination with other devices. In one embodiment, the software module is implemented on a computer-readable medium containing computer program code (e.g., a non-temporary computer-readable medium) and can be executed by a computer processor to perform some or all of the steps, operations, or processes described herein. [Explanation of Symbols]
[0124] 110L Left Loudspeaker 110R Right Loudspeaker 200 Audio Processing Systems 910L Left Speaker 910R Right Speaker 1100 Computer System
Claims
1. A system for enhancing input audio signals to a first speaker and a second speaker, Determine the speaker asymmetry between the first speaker and the second speaker. Based on the determined speaker asymmetry, the mismatch in at least one of the amplitude characteristics, time characteristics, or frequency characteristics associated with the first speaker and the second speaker is determined. By applying at least one of delay and gain to the input audio signal, a first output channel for the first speaker and a second output channel for the second speaker are generated to adjust the mismatch. Processing circuit configured as follows A system characterized by having the following features.
2. The aforementioned processing circuit is To generate test audio signals for the first speaker and the second speaker, Recording the aforementioned test audio signal, Analyzing the recorded test audio signal and The system according to claim 1, further configured to determine the speaker asymmetry between the first speaker and the second speaker.
3. The system according to claim 1, characterized in that the speaker asymmetry is based on predefined asymmetric data for the first speaker and the second speaker.
4. The aforementioned processing circuit is Using one or more sensors, determine the spatial relationship between the listener's head location and the first and second speakers. The system according to claim 1, further configured to determine the speaker asymmetry between the first speaker and the second speaker.
5. The aforementioned processing circuit is Crosstalk cancellation for the aforementioned input audio signal, or Both crosstalk compensation and crosstalk cancellation for the input audio signal The system according to claim 1, further configured to apply the following:
6. The system according to claim 5, wherein the processing circuit is configured to apply the crosstalk cancellation, or both the crosstalk compensation and the crosstalk cancellation, to the input audio signal before applying the delay and the gain to the input audio signal.
7. The aforementioned processing circuit is Dividing the input audio signal into a set of frequency band components, Applying crosstalk cancellation to a subset of the aforementioned set of frequency band components This generates a spatial enhancement signal. The system according to claim 1, further characterized by being configured as follows.
8. The system according to claim 7, characterized in that the subset of the set of frequency band components includes the in-band components of the input audio signal.
9. The aforementioned processing circuit is A spatial enhancement signal is generated by adjusting the gain of the spatial and non-spatial components of the input audio signal. The system according to claim 1, further characterized by being configured as follows.
10. The aforementioned processing circuit is Using the left channel and the right channel of the input audio signal, spatial and non-spatial components are generated. A set of subband filters is applied to the non-spatial component, and the set of subband filters includes a subband filter corresponding to each frequency subband of the non-spatial and spatial components. The system according to claim 1, further characterized by being configured as follows.
11. The system according to claim 10, characterized in that at least one of the non-spatial component and the frequency subband of the spatial component is a critical band.
12. The aforementioned processing circuit is Using the left channel and the right channel of the input audio signal, spatial and non-spatial components are generated. Provides subband delay to the spatial component or the non-spatial component. The system according to claim 1, further characterized by being configured as follows.
13. The system according to claim 1, wherein the processing circuit is further configured to apply N-band equalization to the input audio signal by applying one or more filters to at least one of the left or right channels of the input audio signal, thereby adjusting for asymmetry in the frequency response of the first speaker and the second speaker.
14. The system according to claim 13, characterized in that one or more filters balance the frequency responses of the first speaker and the second speaker.
15. The one or more filters mentioned above are Low shelf filters and high shelf filters, A bandpass filter and A bandstop filter and Peak notch filter and, Low-pass filter and high-pass filter The system according to claim 13, characterized by including at least one of the following.
16. The system according to claim 1, characterized in that the processing circuit is configured to apply the delay or gain to the input audio signal by applying the delay or gain to one of the left or right channels of the input audio signal.
17. The system according to claim 1, wherein the processing circuit is further configured to adjust the gain of the spatial and non-spatial components of the input audio signal.
18. A method for enhancing input audio signals to a first speaker and a second speaker, wherein a processing circuit is used to enhance the input audio signals to the first speaker and the second speaker. Determining the speaker asymmetry between the first speaker and the second speaker, Based on the determined speaker asymmetry, a mismatch is determined in at least one of the amplitude characteristics, time characteristics, or frequency characteristics associated with the first speaker and the second speaker. By applying at least one of delay and gain to the input audio signal, a first output channel for the first speaker and a second output channel for the second speaker are generated to adjust the mismatch. A method characterized by comprising:
19. Determining the speaker asymmetry between the first speaker and the second speaker is: To generate test audio signals for the first speaker and the second speaker, Recording the aforementioned test audio signal, Analyzing the recorded test audio signal and The method according to 18, characterized by including the following:
20. The method according to 18, characterized in that the speaker asymmetry is based on predefined asymmetric data for the first speaker and the second speaker.
21. Determining the speaker asymmetry between the first speaker and the second speaker is: Using one or more sensors, determine the spatial relationship between the listener's head location and the first and second speakers. The method according to 18, characterized by including the following:
22. The aforementioned processing circuit, Crosstalk cancellation for the aforementioned input audio signal, or Both crosstalk compensation and crosstalk cancellation for the input audio signal The method of 18, further comprising applying
23. The method according to 22, characterized in that at least one of the crosstalk compensation or crosstalk cancellation for the input audio signal is applied before applying at least one of the delay and gain to the input audio signal.
24. The aforementioned processing circuit, Dividing the input audio signal into a set of frequency band components, Applying crosstalk cancellation to a subset of the aforementioned set of frequency band components This generates a spatial enhancement signal. The method according to 18, further comprising:
25. The method according to 24, characterized in that the subset of the set of frequency band components includes the in-band components of the input audio signal.
26. The aforementioned processing circuit, A spatial enhancement signal is generated by adjusting the gain of the spatial and non-spatial components of the input audio signal. The method according to 18, further comprising:
27. The aforementioned processing circuit is Using the left channel and the right channel of the input audio signal, spatial and non-spatial components are generated. A set of subband filters is applied to the non-spatial component, and the set of subband filters includes a subband filter corresponding to each frequency subband of the non-spatial and spatial components. The method according to 18, characterized in that it is further configured as follows.
28. The method according to 27, characterized in that at least one of the non-spatial component and the frequency subband of the spatial component is a critical band.
29. The aforementioned processing circuit, Using the left channel and the right channel of the input audio signal, spatial and non-spatial components are generated. To provide a subband delay for the spatial component or the non-spatial component. The method according to 18, further comprising:
30. The aforementioned processing circuit, Applying N-band equalization to the input audio signal to adjust the asymmetry in the frequency response of the first speaker and the second speaker, wherein applying N-band equalization means, Applying one or more filters to at least one of the left or right channels of the input audio signal. This further includes The method according to 18, further comprising:
31. The method according to 30, characterized in that the one or more filters balance the frequency responses of the first speaker and the second speaker.
32. The one or more filters mentioned above are Low shelf filters and high shelf filters, A bandpass filter and A bandstop filter and Peak notch filter and, Low-pass filter and high-pass filter The method according to 30, characterized by comprising at least one of the following.
33. The method according to 18, characterized in that applying the delay or gain to the input audio signal includes applying the delay or gain to one of the left or right channels of the input audio signal.
34. The method according to 18, further comprising adjusting the gain of the spatial and non-spatial components of the input audio signal using the processing circuit.
35. A non-temporary computer-readable medium characterized by storing a program that causes a processing circuit to perform the method described in any one of claims 18 to 34.