Multi-channel crosstalk processing

By performing subband space processing and crosstalk processing on multi-channel audio signals, the problem of losing the sense of sound field space on stereo speakers is solved, achieving a more realistic and rich listening experience.

CN120201361APending Publication Date: 2025-06-24BOOMCLOUD 360 INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510597597.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-10-10
Filing Date
2020-09-03
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When outputting multi-channel audio signals to stereo speakers, it is easy to lose the spatial sense of the sound field, affecting the listening experience.

Method used

By performing subband space processing and crosstalk processing on the multi-channel input audio signal, a stereo output signal suitable for the left speaker and the right speaker is generated, preserving or enhancing the sound field space sense of the audio signal.

Benefits of technology

It realizes the sound field space sense of retaining or enhancing the multi-channel audio signal on the stereo speakers, and enhances the authenticity and richness of the listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201361A_ABST
    Figure CN120201361A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to multi-channel crosstalk processing. The audio system processes a multi-channel input audio signal into stereo signals for left and right speakers while retaining a sense of space of the sound field of the input audio signal. A multi-channel input audio signal includes a first left-right channel pair including a left input channel and a right input channel and a second left-right channel pair including a left peripheral input channel and a right peripheral input channel. Sub-band spatial processing may be applied to the first and second left and right channel pairs. A first crosstalk process is applied to the first left and right channel pair to generate a first crosstalk processed channel. A second crosstalk process is applied to the second left and right channel pair to generate a second crosstalk processed channel. A left output channel and a right output channel are generated from the first and second crosstalk processed channels. Crosstalk processing may include crosstalk cancellation or crosstalk simulation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Division Case Explanation

[0002] This application is a divisional application of a patent application for invention titled "Multi-channel Crosstalk Processing" with a filing date of September 3, 2020, a national application number of 202080082388.8. Technical Field

[0003] Embodiments of the present disclosure generally relate to the field of audio signal processing and, more particularly, to spatially enhanced multi-channel audio. Background Art

[0004] Surround sound refers to the sound reproduction of an audio signal including multiple channels using speakers located around a listener. For example, 5.1 surround sound uses six channels for front speakers, left and right speakers, a subwoofer, and rear (or "surround") left and right speakers. In another example, 7.1 surround sound uses eight channels by splitting the rear left and right speakers of a 5.1 surround sound configuration into four separate speakers, such as left surround, right surround, left rear surround, and right rear surround speakers. The audio channels of a multi-channel audio signal can be associated with angular positions that correspond to the positions of the speakers that output the audio channels. Thus, when an audio signal is output to speakers at different positions, a multi-channel audio signal allows a listener to perceive a sense of space in the sound field. However, when a multi-channel audio signal for surround sound is output to stereo (e.g., left and right) speakers or headphones, the sense of space may be lost. Summary of the Invention

[0005] Embodiments relate to processing a multi-channel input audio signal (e.g., surround sound) into a stereo output signal for left and right speakers while preserving or enhancing the sense of space of the sound field of the multi-channel input audio signal. Among other things, the processing results in a listening experience whereby each channel of the audio signal is perceived as originating from the same or a similar direction as would occur when rendering the audio signal on a surround sound system (e.g., 5.1, 7.1, etc.).

[0006] In some example embodiments, a multi-channel input audio signal including a left input channel, a right input channel, a left peripheral input channel, and a right peripheral input channel is received. Sub-band spatial processing is performed on the left input channel, the right input channel, the left peripheral input channel, and the right peripheral input channel to create spatially enhanced channels. The sub-band spatial processing may include gain adjustment of the mid components and side sub-band components of the left input channel, the right input channel, the left peripheral input channel, and the right peripheral input channel. Crosstalk processing is performed on the spatially enhanced channels to create a crosstalk-processed left channel and a crosstalk-processed right channel. A left output channel is generated from the left crosstalk-processed channel, and a right output channel is generated from the right crosstalk-processed channel. The crosstalk processing may include crosstalk cancellation or crosstalk simulation.

[0007] The left peripheral channel and the right peripheral channel may include a left surround input channel and a right surround input channel, and / or a left rear surround input channel and a right rear surround input channel. The multi-channel input audio signal may further include a center channel and a low-frequency channel that may be combined with the output of the crosstalk processing.

[0008] In some embodiments, sub-band spatial processing is performed on each of a corresponding pair of left and right channels. For example, sub-band spatial processing may be performed by adjusting the gains of the mid sub-band components and side sub-band components of the left input channel and the right input channel, adjusting the gains of the mid sub-band components and side sub-band components of the left peripheral input channel and the right peripheral input channel, and combining the gain-adjusted mid sub-band components and gain-adjusted side sub-band components of the left input channel, the right input channel, the left peripheral input channel, and the right peripheral input channel into a left combined channel and a right combined channel. Crosstalk processing is performed on the left combined channel and the right combined channel to generate output channels.

[0009] In some embodiments, sub-band spatial processing is performed on the combined left and right channels. For example, the sub-band spatial processing may include combining the left input channel and the left peripheral input channel into a left combined channel, combining the right input channel and the right peripheral input channel into a right combined channel, and adjusting the gains of the mid sub-band components and side sub-band components of the left combined channel and the right combined channel to create a left spatially enhanced channel and a right spatially enhanced channel. Crosstalk processing is performed on the left spatially enhanced channel and the right spatially enhanced channel to generate output channels.

[0010] In some embodiments, a binaural filter is applied to at least a portion of the input channels. For example, a binaural filter is applied to the peripheral input channels to adjust the angular positions associated with the peripheral input channels. In some embodiments, a binaural filter is applied to any input channel suitable for adjusting the angular position associated with the input channels, including the left or right input channels.

[0011] Some embodiments may include a system for processing a multi-channel input audio signal. The system includes circuitry configured to: receive a multi-channel input audio signal including a plurality of left-right channel pairs, a first left-right channel pair of the plurality of left-right channel pairs including a left input channel and a right input channel, and a second left-right channel pair of the plurality of left-right channel pairs including a left peripheral input channel and a right peripheral input channel; apply a first crosstalk processing to the first left-right channel pair to generate a first crosstalk processed channel; apply a second crosstalk processing to the second left-right channel pair to generate a second crosstalk processed channel; and generate a left output channel and a right output channel from the first and second crosstalk processed channels.

[0012] In some embodiments, the circuitry is further configured to: apply a first subband spatial processing to the first left-right channel pair, the first subband spatial processing including gain adjustment of middle components and side components of the left input channel and the right input channel; and apply a second subband spatial processing to the second left-right channel pair, the second subband spatial processing including gain adjustment of middle components and side components of the left peripheral input channel and the right peripheral input channel.

[0013] Some embodiments may include a non-transitory computer-readable medium storing program code that, when executed by a processor, causes the processor to: receive a multi-channel input audio signal including a plurality of left-right channel pairs, a first left-right channel pair of the plurality of left-right channel pairs including a left input channel and a right input channel, and a second left-right channel pair of the plurality of left-right channel pairs including a left peripheral input channel and a right peripheral input channel; apply a first crosstalk processing to the first left-right channel pair to generate a first crosstalk processed channel; apply a second crosstalk processing to the second left-right channel pair to generate a second crosstalk processed channel; and generate a left output channel and a right output channel from the first and second crosstalk processed channels.

[0014] In some embodiments, the computer-readable medium further includes program code that causes the processor to perform the following operations: apply a first subband spatial processing to the first left-right channel pair, the first subband spatial processing including gain adjustment of middle components and side components of the left input channel and the right input channel; and apply a second subband spatial processing to the second left-right channel pair, the second subband spatial processing including gain adjustment of middle components and side components of the left peripheral input channel and the right peripheral input channel.

[0015] Some embodiments may include a method for processing a multichannel input audio signal. The method may include, by circuitry: receiving a multichannel input audio signal including a plurality of left-right channel pairs, a first left-right channel pair of the plurality of left-right channel pairs including a left input channel and a right input channel, and a second left-right channel pair of the plurality of left-right channel pairs including a left peripheral input channel and a right peripheral input channel; applying a first crosstalk process to the first left-right channel pair to generate a first crosstalk-processed channel; applying a second crosstalk process to the second left-right channel pair to generate a second crosstalk-processed channel; and generating a left output channel and a right output channel from the first and second crosstalk-processed channels

[0016] In some embodiments, the method further includes, by circuitry: applying a first subband spatial process to the first left-right channel pair, the first subband spatial process including gain adjustment of a mid component and a side component of the left input channel and the right input channel; applying a second subband spatial process to the second left-right channel pair, the second subband spatial process including gain adjustment of a mid component and a side component of the left peripheral input channel and the right peripheral input channel. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Illustrates an example of a surround sound stereo audio reproduction system according to one embodiment.

[0018] Figure 2 Illustrates an example of an audio system according to one embodiment.

[0019] Figure 3 Illustrates an example of a subband spatial processor according to one embodiment.

[0020] Figure 4 Illustrates an example of a crosstalk cancellation processor according to one embodiment.

[0021] Figure 5 Illustrates an example of a method for enhancing an audio signal using an audio system shown in Figure 2 .

[0022] Figure 6 Illustrates an example of an audio system according to one embodiment.

[0023] Figure 7 Illustrates an example of a method for enhancing an audio signal using an audio system shown in Figure 6 .

[0024] Figure 8 Illustrates an example of a computer system according to one embodiment.

[0025] Figure 9 Illustrates an example of an audio system according to one embodiment.

[0026] Figure 10 Illustrates an example of an audio system according to one embodiment.

[0027] Figure 11 Illustrates an example of a method for enhancing an audio signal using the Figure 9 or Figure 10 audio system shown in

[0028] Figure 12 Illustrates an example of a crosstalk simulation processor according to one embodiment. DETAILED DESCRIPTION

[0029] The features and advantages described in the specification are not all inclusive, and in particular, many additional features and advantages will be apparent to those of ordinary skill in the art in view of the drawings, specification, and claims. Further, it should be noted that the language used in the specification has been principally selected for readability and instructional purposes and may not have been selected to delineate or circumscribe the inventive subject matter.

[0030] The drawings (Figs.) and the following description relate only by way of illustration to the preferred embodiments. It should be noted that from the following discussion, alternative embodiments of the structures and methods disclosed herein will be readily recognized as viable alternatives that may be employed without departing from the principles of the present invention.

[0031] Reference will now be made in detail to several embodiments of the present invention, examples of which are illustrated in the drawings. Note that wherever feasible, like or identical reference numerals may be used in the drawings and may indicate like or identical functionality. The drawings depict embodiments for purposes of illustration only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods illustrated herein may be employed without departing from the principles described herein.

[0032] Example Surround Sound Stereo and Example Audio System

[0033] The audio systems discussed herein provide crosstalk processing and spatial enhancement for multi-channel surround sound audio signals output to stereo (e.g., left and right) speakers. The signal processing results in the preservation or enhancement of the spatial sense of the sound field encoded in the multi-channel surround sound audio signal. Among other things, the spatial sense achieved using a multi-speaker surround sound system is achieved using stereo speakers.

[0034] Figure 1Illustrated is an example of a surround sound stereo audio reproduction system 100 according to one embodiment. System 100 is an example of a 7.1 surround sound system that provides audio signal reproduction to a listener 140. System 100 includes a left speaker 110L, a right speaker 110R, a center speaker 115, a subwoofer 125, a left surround speaker 120L, a right surround speaker 120R, a left surround rear speaker 130L, and a right surround rear speaker 130R. The center speaker 115 and the subwoofer 125 may be positioned in front of the listener 140, which defines a forward axis of 0°. The left speaker 110L may be positioned at an angle between -20° and -30° relative to the forward axis, and the right speaker 110R may be positioned at an angle between 20° and 30° relative to the forward axis. The left surround speaker 120L may be positioned at an angle between -90° and -110° relative to the forward axis, and the right surround speaker 120R may be positioned at an angle between 90° and 110° relative to the forward axis. The left surround rear speaker 130L may be positioned at an angle between -135° and -150° relative to the forward axis, and the right surround rear speaker 130R may be positioned at an angle between 135° and 150° relative to the forward axis. System 100 may be configured to receive an audio signal including channels for each of the speakers 110, 115, 120, and 130, and the subwoofer 125. The multiple speakers and their positioning arrangement provide a sense of space in the sound field that can be perceived by the listener 140. As discussed in more detail below, the audio system may be configured to process a multi-channel input audio signal for the surround sound system 100 into an enhanced stereo signal for the left and right speakers (e.g., speakers 110L and 110R) that reproduces or simulates the sense of space in the sound field generated by the surround sound system 100 using the multi-channel audio signal.

[0035] Figure 2 Illustrated is an example of an audio system 200 according to one embodiment. The audio system 200 receives an input audio signal including a left input channel 201A, a right input channel 210B, a center input channel 210C, a low-frequency input channel 210D, a left surround input channel 210E, a right surround input channel 210F, a left surround rear input channel 210G, and a right surround rear input channel 210H.

[0036] The sound channels 210E, 210F, 210G, and 210H are examples of peripheral channels for surround speakers. The peripheral channels can include channels other than the left input channel and the right input channel. The peripheral channels can include channel pairs, such as left-right pairs, or front-back pairs, or other pair arrangements. For example, when an input audio signal is output by the surround stereo audio reproduction system 100, the left surround speaker 120L receives the left surround input channel 210E, the right surround speaker 120R receives the right surround input channel 210F, the left rear surround speaker 130L receives the left rear surround input channel 210G, and the right rear surround speaker 130R receives the right rear surround input channel 210H. In some embodiments, the input audio signal has fewer or more peripheral channels. For example, an audio input signal for a 5.1 surround sound system can include only two peripheral channels, such as the left and right surround input channels that can be output to the left and right surround speakers. Similarly, the left speaker 110L can receive the left input channel 210A, the right speaker 110R can receive the right input channel 210B, the center speaker 115 can receive the center input channel 210C, and the subwoofer 125 can receive the low-frequency input channel 210D. The input audio signal provides a sense of space for the sound field when output by the surround stereo audio reproduction system 100.

[0037] The audio system 200 receives an input audio signal and generates an output signal including a left output channel 290L and a right output channel 290R. The audio system 200 can combine the input channels of the input audio signal and can further provide enhancements such as subband spatial processing and crosstalk cancellation to generate the output audio signal. The left output channel 290L can be provided to the left speaker, and the right output channel 290R can be output to the right speaker. The output audio signal provides a sense of space for the sound field using the left and right speakers (e.g., the left speaker 110L and the right speaker 110R), which is typically achieved by outputting the input audio signal using a surround sound system including multiple (e.g., peripheral) speakers.

[0038] The audio system 200 includes gains 215A, 215B, 215C, 215D, 215E, 215F, 215G, and 215H, subband spatial processors 230A, 230B, and 230C, shelf filters 220, distributors 240, binaural filters 250A, 250B, 250C, and 250D, left channel combiners 260A, right channel combiners 260B, crosstalk cancellation processors 270, left channel combiners 260C, right channel combiners 260D, and output gains 280.

[0039] Each of gains 215A to 215H may receive a respective input channel 210A to 210H and may apply a gain to input channels 210A to 210H. Gains 215A to 215H may be different to adjust the gains of the input channels relative to each other, or may be the same. In some embodiments, a positive gain is applied to the left and right peripheral input channels 210E, 210F, 210G, and 210H, while a negative gain is applied to the central input channel 210C. For example, gain 215A may apply 0 dB gain, gain 215B may apply 0 dB gain, gain 215C may apply -3 dB gain, gain 215D may apply 0 dB gain, gain 215E may apply 3 dB gain, gain 215F may apply 3 dB gain, gain 215G may apply 3 dB gain, and gain 215H may apply 3 dB gain.

[0040] Gain 215A and gain 215B are coupled to subband spatial processor 230. Similarly, gains 215E and 215F are coupled to subband spatial processor 230B, and gains 215G and 215H are coupled to subband spatial processor 230C. Subband spatial processors 230A, 230B, and 230C each apply subband spatial processing to corresponding left and right channel pairs.

[0041] Each subband spatial processor 230 performs subband spatial processing on the left and right input channels by performing gain adjustment on the mid components and side subband components of the left and right input channels to generate left and right spatially enhanced channels. Subband spatial processor 230A performs subband spatial processing on the left and right input channels, while the other subband spatial processors 230B and 230C each perform subband spatial processing on corresponding left and right peripheral channels. Depending on the number of peripheral channels in the input audio signal, the audio system 200 may include more or fewer subband spatial processors. In some embodiments, channels without a left / right counterpart (such as central input channel 210C, low-frequency input channel 210D, or other types of channels such as rear center, top center, etc.) may bypass the SBS processing.

[0042] Subband spatial processor 230B is coupled to binaural filters 250A and 250B. Subband spatial processor 230B provides the left spatially enhanced channel to binaural filter 250A and provides the right spatially enhanced channel to binaural filter 250B. Similarly, subband spatial processor 230C is coupled to binaural filters 250C and 250D. Subband spatial processor 230C provides the left spatially enhanced channel to binaural filter 250C and provides the right spatially enhanced channel to binaural filter 250D. Additional details regarding subband spatial processor 230 are shown in Figure 3 and discussed below.

[0043] Each of the binaural filters 250A, 250B, 250C, and 250D applies a head-related transfer function (HRTF) that describes a target source location from which a listener is to perceive sound for an input channel. Each binaural filter receives the input channel and generates a left output channel and a right output channel by applying the HRTF that adjusts for an angular position associated with the input channel. The angular position can include an angle defined relative to the listener 140 in the X-Y "azimuth" plane, such as Figure 1As shown, and may also include an angle defined in the Z-axis, such as for an ambisonic signal or a channel-based format including signals intended to be rendered above or below the X-Y plane relative to the listener 140. For example, the binaural filter 250A may be configured to apply a filter based on the left surround input channel 210E associated with an angle between -90° and -110° relative to the forward axis of the left surround speaker 120L (defined in the X-Y plane). The binaural filter 250B may be configured to apply a filter based on the right surround input channel 210F associated with an angle between 90° and 110° relative to the forward axis of the left surround speaker 120L. The binaural filter 250C may be configured to apply a filter based on the left surround rear input channel 210G associated with an angle between -135° and -150° relative to the forward axis of the left surround rear speaker 130L. The binaural filter 250D may be configured to apply a filter based on the right surround rear input channel 210H associated with an angle between 135° and 150° relative to the forward axis of the right surround rear speaker 130R. In some embodiments, binaural processing may be completely bypassed to preserve inter-channel spectral uniformity. One or more of the binaural filters 250A, 250B, 250C, and 250D may be omitted from the audio system 200. However, the binaural filters 250A, 250B, 250C, and 250D may be used to enhance spatial imaging. In some embodiments, binaural filtering may be applied to channels other than the peripheral input channels. For example, the binaural filter may be applied to each of the left and right spatial enhancement channels output from the sub-band spatial processor 230A to adjust for different left and right output speaker positions. In another example, if the input audio signal includes channels associated with other speaker positions (i.e., overhead, rear center, etc.), then binaural processing may be applied to the other input channels. In this sense, binaural processing may be applied to one or more of the left input channel 210A, right input channel 210B, center input channel 210C, or low-frequency input channel 210D. In some embodiments, the HRTF is not applied, and one or more of the binaural filters 250A, 250B, 250C, and 250D may be bypassed or omitted from the system 200.

[0044] An example binaural filter may be defined by Equation 1:

[0045] S o (z) = H(θ, z)S i (z) Equation (1)

[0046] where S o and S i are the output and input signals, respectively. The parameter θ to Si and S o Encode the angles of each sound channel in o . The z value is an arbitrary complex number, and our solution for it is a function of the encoded frequency. Thus, H(θ,z) is a function of the angle θ and z, returning a transfer function that is itself a function of z, which can be selected or interpolated from a set of transfer functions that may be derived from an anthropometric database. In this notation, if multi-channel processing is desired, the angle θ and S and H(θ) as a function of z can be evaluated as vectors. In this case, each coefficient in S(z) and H(θ,z) corresponds to a different sound channel, and each coefficient in θ associates the angle with each sound channel.

[0047] In some embodiments, the input audio signal is an ambient stereo audio signal that defines a loudspeaker-independent representation of the sound field. The ambient audio signal can be decoded into a multi-channel audio signal for a surround sound system. The sound channels can be associated with loudspeaker positions at different locations, including positions above or below the listener. A binaural filter can be applied to each decoded input sound channel of the ambient audio signal to adjust for the associated position of the decoded input audio channel.

[0048] In some embodiments, binaural filtering is performed before sub-band spatial processing. For example, a binaural filter can be applied to one or more input sound channels suitable for adjusting for the angular position associated with the sound channel. For each pair of left input sound channels and right input sound channels, the left output sound channels of the binaural filter can be combined, and the right output sound channels of the binaural filter can be combined, and sub-band spatial processing can be applied to the combined left and right sound channels. In some embodiments, the binaural filter is applied to the central input sound channel 210C or the low-frequency input sound channel 210D. In some embodiments, the binaural filter is applied to each input sound channel except the low-frequency input sound channel 210D.

[0049] The left channel combiner 260A is coupled to the sub-band spatial processor 230A and the binaural filters 250A, 250B, 250C, and 250D. The left channel combiner 260A receives the left output sound channels of the sub-band spatial processor 230A and the binaural filters 250A, 250B, 250C, and 250D and combines these channels into a left combined channel. The right channel combiner 260B is also coupled to the sub-band spatial processor 230A and the binaural filters 250A, 250B, 250C, and 250D. The right channel combiner 260B receives the right output sound channels of the sub-band spatial processor 230A and the binaural filters 250A, 250B, 250C, and 250D and combines these channels into a right combined channel.

[0050] The crosstalk cancellation processor 270 receives a left input channel and a right input channel and performs crosstalk cancellation to generate a left crosstalk cancellation channel and a right crosstalk cancellation channel. The crosstalk cancellation processor is coupled to the left channel combiner 260A to receive the left combined channel, and is coupled to the right channel combiner 260B to receive the right combined channel. Here, the left combined channel and the right combined channel processed by the crosstalk cancellation processor 270 represent the mixed left corresponding input channel and right corresponding input channel. Additional details regarding the crosstalk cancellation processor 270 are shown in Figure 4 and discussed below.

[0051] The shelving filter 220 receives the center input channel 210C and applies a high-frequency shelving or peaking filter. The shelving filter 220 provides a "voice boost" on the center input channel 210C. In some embodiments, the shelving filter 220 is bypassed or omitted from the audio system 200. The shelving filter 220 can attenuate or amplify frequencies above the corner frequency. The shelving filter 220 is coupled to the left channel combiner 260C and the right channel combiner 260D. In some embodiments, the shelving filter 220 is defined by a 750 Hz corner frequency, +3 dB gain, and 0.8 Q factor. The shelving filter 220 generates a left center channel and a right center channel as outputs, such as by separating the center input channel into two separate left center channels and right center channels.

[0052] The splitter 240 receives the low-frequency input channel 210D and separates the low-frequency input channel 210D into a left low-frequency channel and a right low-frequency channel. The splitter 240 is coupled to the left channel combiner 260C and the right channel combiner 260D, and provides the left low-frequency channel to the left channel combiner 260C and the right low-frequency channel to the right channel combiner 260D.

[0053] The left channel combiner 260C is coupled to the crosstalk cancellation processor 270, the shelving filter 220, and the splitter 240. The left channel combiner 260C receives the left crosstalk channel from the crosstalk cancellation processor 270, the left center channel from the shelving filter 220, and the left low-frequency channel from the splitter 240, and combines these channels into a left output channel.

[0054] The right channel combiner 260D is coupled to the crosstalk cancellation processor 270, the shelving filter 220, and the splitter 240. The right channel combiner 260D receives the right crosstalk channel from the crosstalk cancellation processor 270, the right output channel from the shelving filter 220, and the right low-frequency channel from the splitter 240, and combines these channels into a right output channel.

[0055] In some embodiments, the left center channel from the overhead filter 220 and the left low-frequency channel from the distributor 240 are combined by the left channel combiner 260A with the left spatial enhancement channel from the sub-band spatial processor 230A and the left output channels from the binaural filters 250A, 250B, 250C, and 250D to generate a left combined channel. Similarly, the right output channel from the overhead filter 220 and the right low-frequency channel from the distributor 240 are combined by the right channel combiner 260B with the right spatial enhancement channel from the sub-band spatial processor 230A and the right output channels from the binaural filters 250A, 250B, 250C, and 250D to generate a right combined channel. The left combined channel and the right combined channel are input into the crosstalk cancellation processor 270. Here, the center and low-frequency channels receive crosstalk cancellation operations. The left channel combiner 260C and the right channel combiner 260D may be omitted. In some embodiments, one of the center or low-frequency channels receives a crosstalk cancellation operation.

[0056] The output gain 280 is coupled to the left channel combiner 260C and the right channel combiner 260D. The output gain 280 applies a gain to the left output channel from the left channel combiner 260C and applies a gain to the right output channel from the right channel combiner 260D. The output gain 280 may apply the same gain to the left output channel and the right output channel, or may apply different gains. The output gain 280 outputs the left output channel 290L and the right output channel 290R of the channels representing the output signal of the audio system 200.

[0057] Exemplary sub-band spatial processor

[0058] Figure 3 FIG. illustrates an example of a sub-band spatial processor 230 according to one embodiment. The sub-band spatial processor 230 is an example of the sub-band spatial processors 230A, 230B, or 230C of the audio system 200. The sub-band spatial processor 230 includes a spatial band allocator 340, a spatial band processor 345, and a spatial band combiner 350. The spatial band allocator 340 is coupled to the spatial band processor 345, and the spatial band processor 345 is coupled to the spatial band combiner 350.

[0059] The spatial band allocator 340 includes an L / R to M / S converter 312, which receives the left input channel X L and the right input channel X R , and converts these inputs into a spatial component X m and a non-spatial component X s . The spatial component X s can be generated by subtracting the left input channel X L from the right input channel X R . The non-spatial component Xm can be generated by adding the left input channel X L and the right input channel X R together.

[0060] The spatial band processor 345 receives the non-spatial component X m and applies a set of sub-band filters to generate an enhanced non-spatial sub-band component E m . The spatial band processor 345 also receives the spatial sub-band component X s and applies a set of sub-band filters to generate an enhanced non-spatial sub-band component E m . The sub-band filters can include various combinations of peak filters, notch filters, low-pass filters, high-pass filters, low-shelf filters, high-shelf filters, band-pass filters, band-stop filters, and / or all-pass filters.

[0061] In some embodiments, the spatial band processor 345 includes sub-band filters for each of the n frequency sub-bands of the non-spatial component X m and sub-band filters for each of the n frequency sub-bands of the spatial component X s . For example, for n = 4 sub-bands, the spatial band processor 345 includes a series of sub-band filters for the non-spatial component X m , including: an intermediate equalization (EQ) filter 362(1) for sub-band (1), an intermediate EQ filter 362(2) for sub-band (2), an intermediate EQ filter 362(3) for sub-band (3), and an intermediate EQ filter 362(4) for sub-band (4). Each intermediate EQ filter 362 applies a filter to the frequency sub-band portion of the non-spatial component X m to generate an enhanced non-spatial component E m .

[0062] The spatial band processor 345 also includes a series of sub-band filters for the frequency sub-bands of the spatial component X s , including a side equalization (EQ) filter 364(1) for sub-band (1), a side EQ filter 364(2) for sub-band (2), a side EQ filter 364(3) for sub-band (3), and a side EQ filter 364(4) for sub-band (4). Each side EQ filter 364 applies a filter to the frequency sub-band portion of the spatial component X s to generate an enhanced spatial component E s .

[0063] The non-spatial component X m and the spatial component X sEach of the n frequency subbands can correspond to a frequency range. For example, subband (1) can correspond to 0 to 300 Hz, subband (2) can correspond to 300 to 510 Hz, subband (3) can correspond to 510 to 2700 Hz, and subband (4) can correspond to 2700 Hz to the Nyquist frequency. In some embodiments, the n frequency subbands are a set of combined critical bands. A corpus of audio samples from multiple music genres can be used to determine the critical bands. The long-term average energy ratio of the mid-to-side components on 24 Bark scale critical bands was determined from the samples. Then, consecutive bands with similar long-term average ratios were grouped together to form a set of critical bands. The range of the frequency subbands and the number of frequency subbands can be adjustable.

[0064] In some embodiments, the mid EQ filter 362 or the side EQ filter 364 can include a biquadratic filter having a transfer function defined by Equation 2:

[0065]

[0066] where z is a complex variable. The filter can be implemented using a direct form I topology defined by Equation 3:

[0067]

[0068] where X is the input vector and Y is the output. Other topologies may be beneficial for certain processors, depending on their maximum word length and saturation behavior.

[0069] The biquadratic can then be used to implement any second-order filter with real-valued inputs and outputs. To design a discrete-time filter, a continuous-time filter is designed and transformed to discrete-time via bilinear transformation. Additionally, frequency warping can be used to compensate for any resulting offset in the center frequency and bandwidth.

[0070] For example, a peak filter can include an S-plane transfer function defined by Equation 4:

[0071]

[0072] where s is a complex variable, A is the magnitude of the peak, and Q is the filter "quality" (standard derivation: ). The digital filter coefficients are:

[0073] where ω0 is the center frequency of the filter in radians, and

[0074] The spatial band combiner 350 receives the center component and the side component, applies gains to each component, and converts the center component and the side component into a left channel and a right channel. For example, the spatial band combiner 350 receives the enhanced non-spatial component E m and the enhanced spatial component E s , and performs global center gain and side gain before converting the enhanced non-spatial component E m and the enhanced spatial component E s into a left spatially enhanced channel E L and a right spatially enhanced channel E R .

[0075] More specifically, the spatial band combiner 350 includes a global center gain 322, a global side gain 324, and an M / S to L / R converter 326 coupled to the global center gain 322 and the global side gain 324. The global center gain 322 receives the enhanced non-spatial component E m and applies a gain, and the global side gain 324 receives the enhanced spatial component E s and applies a gain. The M / S to L / R converter 326 receives the enhanced non-spatial component E m from the global center gain 322 and the enhanced spatial component E s from the global side gain 324, and converts these inputs into a left spatially enhanced channel E L and a right spatially enhanced channel E R .

[0076] Example crosstalk cancellation processor

[0077] Figure 4 Illustrates a crosstalk cancellation processor 270 according to an example embodiment. The crosstalk cancellation processor 270 receives a left channel (e.g., the left spatially enhanced channel E L ) as an input from the left channel combiner 260A and a right channel (e.g., the right spatially enhanced channel E R ) as an input from the right channel combiner 260B, and performs crosstalk cancellation on the left channel and the right channel to generate a left output channel O L and a right output channel O R .

[0078] The crosstalk cancellation processor 270 includes an in-out band allocator 410, inverters 420 and 422, a cross-side estimator 430 and 440, combiners 450 and 452, and an in-out band combiner 460. These components operate together to divide the input channels T L , T R into in-band components and out-of-band components, and perform crosstalk cancellation on the in-band components to generate output channels O L , OR 。

[0079] By dividing the input audio signal E into different frequency band components and performing crosstalk cancellation on selective components (e.g., in-band components), crosstalk cancellation can be performed for a specific frequency band while avoiding degradation in other frequency bands. If crosstalk cancellation is performed without dividing the input audio signal E into different frequency bands, the audio signal after such crosstalk cancellation may exhibit significant attenuation or amplification in low frequencies (e.g., below 350 Hz), high frequencies (e.g., above 12000 Hz), or both non-spatial and spatial components. By selectively performing crosstalk cancellation within the band where the vast majority of the influential spatial cues are located (e.g., between 250 Hz and 14000 Hz), a balanced overall energy can be maintained across the entire spectrum of the mix, especially in the non-spatial components.

[0080] The in-band / out-of-band splitter 410 splits the input channels E L 、E R into in-band channels E L,In 、E R,In and out-of-band channels E L,Out 、E R,Out respectively. In particular, the in-band / out-of-band splitter 410 divides the left enhanced compensation channel E L into a left in-band channel E L,In and a left out-of-band channel E L,Out . Similarly, the in-band / out-of-band splitter 410 separates the right enhanced compensation channel E R into a right in-band channel E R,In and a right out-of-band channel E R,Out . Each in-band channel may include a portion of the corresponding input channel corresponding to a frequency range that includes, for example, 250 Hz to 14 kHz. The frequency band range may be adjustable, for example, according to speaker parameters.

[0081] The inverter 420 and the crosstalk estimator 430 operate together to generate a left crosstalk cancellation component S L to compensate for the crosstalk sound component caused by the left in-band channel E L,In . Similarly, the inverter 422 and the crosstalk estimator 440 operate together to generate a right crosstalk cancellation component S R to compensate for the crosstalk sound component caused by the right in-band channel E R,In .

[0082] In one method, the inverter 420 receives the in-band channel E L,In , and inverts the polarity of the received in-band channel E L,In to generate an inverted in-band channel E L,In' . The crosstalk estimator 430 receives the inverted in-band channel EL,In' , and through filtering, the in-band channel E corresponding to the contralateral sound component is extracted L,In' as a part. Since the filtering is performed on the in-band channel E L,In' , the part extracted by the contralateral estimator 430 becomes the inversion of the part of the in-band channel E L,In attributed to the contralateral sound component. Therefore, the part extracted by the contralateral estimator 430 becomes the left contralateral cancellation component S L , which can be added to the corresponding in-band channel E R,In to reduce the contralateral sound component caused by the in-band channel E L,In . In some embodiments, the inverter 420 and the contralateral estimator 430 are implemented in a different order.

[0083] The inverter 422 and the contralateral estimator 440 perform similar operations on the in-band channel E R,In to generate the right contralateral cancellation component S R . Therefore, for the sake of brevity, its detailed description is omitted here.

[0084] In an exemplary implementation, the contralateral estimator 430 includes a filter 432, an amplifier 434, and a delay unit 436. The filter 432 receives the inverted input channel E L,In' , and through a filtering function, extracts a part of the in-band channel E L,In' corresponding to the contralateral sound component. An example filter implementation is a notch or shelving filter, which has a center frequency selected between 5000 and 10000 Hz and a Q selected between 0.5 and 1.0. The gain (G dB ) in decibels can be derived from Equation 5:

[0085] G dB = -3.0 - log 1.333 (D) Equation (5)

[0086] where D is the delay amount of the delay unit 1556A / B in samples, for example, at a sampling rate of 48 KHz. Another implementation is a low-pass filter, which has a corner frequency selected between 5000 and 10000 Hz and a Q selected between 0.5 and 1.0. In addition, the amplifier 434 amplifies the extracted part by a corresponding gain factor G L,In , and the delay unit 436 delays the amplified output from the amplifier 434 according to the delay function D to generate the left contralateral cancellation component S L . The contralateral estimator 440 includes a filter 442, an amplifier 444, and a delay unit 446, and the delay unit 446 performs a similar operation on the in-band channel E R,In' to generate the right contralateral cancellation component S R. In one example, the contralateral estimators 430, 440 generate a left contralateral cancellation component S according to the following equations L and S R :

[0087] S L = D[G L,In * F[E L,In '] Equation (6)

[0088] S R = D[G R,In * F[E R,In '] Equation (7)

[0089] where F[] is a filter function and D[] is a delay function.

[0090] The configuration of crosstalk cancellation can be determined by speaker parameters. In one example, the filter center frequency, delay amount, amplifier gain, and filter gain can be determined according to the angle formed between two output speakers of the output signal relative to the listener, or other characteristics of the speakers such as relative position, power, etc. In some embodiments, values between speaker angles are used to interpolate other values.

[0091] The combiner 450 combines the right contralateral cancellation component S R into the left in-band channel E L,In to generate a left in-band compensated channel U L , and the combiner 452 combines the left contralateral cancellation component SL into the right in-band channel E R,In to generate a right in-band compensated channel U R . The in-band / out-of-band combiner 460 combines the left in-band compensated channel U L with the out-of-band channel E L,Out to generate a left output channel O L , and combines the right in-band compensated channel U R with the out-of-band channel E R,Out to generate a right output channel O R .

[0092] Thus, the left output channel O L includes the right contralateral cancellation component S R,In corresponding to the inversion of a part of the contralateral sound attributable to the in-band channel T R , and the right output channel O R includes the left contralateral cancellation component S L,In corresponding to the inversion of a part of the contralateral sound attributable to the in-band channel T L . In this configuration, according to the right output channel O RThe wavefront of the ipsilateral sound component output from the right speaker (e.g., speaker 110R) reaching the right ear can cancel the wavefront of the contralateral sound component according to the left output channel O L The wavefront of the contralateral sound component output from the right speaker (e.g., speaker 110L). Similarly, according to the left output channel O L The wavefront of the ipsilateral sound component output from the left speaker reaching the left ear can cancel the wavefront of the contralateral sound component according to the right output channel O R The wavefront of the contralateral sound component output from the right speaker. Therefore, the contralateral sound component can be reduced to enhance spatial detectability.

[0093] Example of audio signal enhancement process

[0094] Figure 5 Illustrates an example of a method 500 for enhancing an audio signal using the audio system 200 shown in Figure 2 In some embodiments, method 500 may include different and / or additional steps, or some steps may be in a different order.

[0095] The audio system 200 receives 505 a multi-channel input audio signal. The multi-channel audio signal may be a surround audio signal including a left input channel, a right input channel, at least one left peripheral input channel, and at least one right peripheral input channel. The multi-channel audio signal may also include a central input channel 210C and a low-frequency input channel 210D. For example, the input audio signal may be for a 7.1 surround sound system including a left input channel 210A and a right input channel 210B, and peripheral channels including a left surround input channel 210E and a right surround input channel 210F, and a left surround back input channel 210G and a right surround back input channel 210H. In another example of an input audio signal for a 5.1 surround sound system, the peripheral channels may include a single left peripheral channel and a single right peripheral channel.

[0096] The audio system 200 (e.g., gains 215A to 215H) applies 510 gains to the channels of the multi-channel input audio signal. The gains 215A to 215H may vary to control the contribution of a particular input channel to the output signal generated by the audio system 200. In some embodiments, the central input channel 210C receives a negative gain while the peripheral input channels receive a positive gain.

[0097] The audio system 200 (e.g., sub-band spatial processor 230A) generates 515 a left spatial enhancement channel and a right spatial enhancement channel by performing sub-band spatial processing on the left input channel and the right input channel. For example, the sub-band spatial processor 230A generates the spatial enhancement channels by adjusting the gains of n sub-bands of the middle components and side components of the left input channel 210A and the right input channel 210B.

[0098] The audio system 200 (e.g., the sub-band spatial processors 230B and / or 230C) generates left and right spatially enhanced peripheral channels 520 by performing sub-band spatial processing on the left and right peripheral input channels. For example, the sub-band spatial processor 230B adjusts the gains of n sub-bands of the mid and side components of the left surround input channel 210E and the right surround input channel 210F to generate the left and right spatially enhanced peripheral channels. The sub-band spatial processor 230C adjusts the gains of n sub-bands of the mid and side components of the left rear surround input channel 210G and the right rear surround input channel 210H to generate the left and right spatially enhanced peripheral channels.

[0099] The audio system 200 (e.g., the binaural filters 250A to 250D) applies 525 the binaural filters to each of the left and right spatially enhanced peripheral channels. For example, the binaural filter 250A generates left and right output channels from the left spatially enhanced peripheral channel output from the sub-band spatial processor 230B by applying a head-related transfer function (HRTF). The binaural filter 250B generates left and right output channels from the spatially enhanced right channel output from the sub-band spatial processor 230B by applying the HRTF. The binaural filter 250C generates left and right output channels from the spatially enhanced left channel output from the sub-band spatial processor 230C by applying the HRTF. The binaural filter 250D generates left and right output channels from the spatially enhanced right channel output from the sub-band spatial processor 230C by applying the HRTF. In some embodiments, the binaural filtering is bypassed.

[0100] The audio system 200 (e.g., the overhead filter 220) applies 530 the overhead filter to the center input channel 210C. In some embodiments, a gain is applied to the center input channel 210C. Additionally, the overhead filter 220 separates the center input channel 210C into a left center channel and a right center channel.

[0101] The audio system 200 (e.g., the distributor 240) separates 535 the low-frequency input channel into a left low-frequency channel and a right low-frequency channel.

[0102] The audio system 200 (e.g., the left channel combiner 260A) combines 540 the left spatially enhanced channel from the sub-band spatial processor 230A and the left output channels of the binaural filters 250A, 250B, 250C, and 250D to generate a left combined channel. For example, the left spatially enhanced channel can be added to the left output channel.

[0103] An audio system 200 (e.g., right channel combiner 260B) combines 545 the right spatial enhancement channel from the subband spatial processor 230A and the right output channels of the binaural filters 250A, 250B, 250C, and 250D to generate a right combined channel. For example, the right spatial enhancement channel can be added to the right output channel.

[0104] The audio system 200 (e.g., crosstalk cancellation processor 270) performs 550 crosstalk cancellation on the left combined channel and the right combined channel to generate a left crosstalk cancellation channel and a right crosstalk cancellation channel.

[0105] The audio system 200 (e.g., left channel combiner 260C and right channel combiner 260D) combines 555 the left crosstalk cancellation channel from the crosstalk cancellation processor 270 with the left low-frequency channel from the distributor 240 and the left center channel from the overhead filter 220 to generate a left output channel, and combines the right crosstalk cancellation channel from the crosstalk cancellation processor 270 with the right low-frequency channel from the distributor 240 and the right center channel from the overhead filter 220 to generate a right output channel. Additionally, the audio system 200 (e.g., output gain 280) can apply a gain to each of the left output channel and the right output channel. The audio system 200 outputs an output audio signal including the left output channel and the right output channels 290L and 290R. Example audio systems and example audio processing procedures

[0106] Figure 6 An example of an audio system 600 according to one embodiment is illustrated. The audio system 600 can be similar to the audio system 200, but can differ from the audio system 200 at least in that the left input channel and the right input channel are combined with the left peripheral channel and the right peripheral channel before subband spatial processing in the audio system 600. Here, a single subband spatial processor and corresponding subband spatial processing steps can be used instead of, as shown for the audio system 200, separate subband spatial processors for the left and right channel pairs.

[0107] The audio system 600 receives an input audio signal. The input audio signal can include a left input channel 610A, a right input channel 610B, a center input channel 610C, a low-frequency input channel 610D, a left surround input channel 610E, a right surround input channel 610F, a left surround back input channel 610G, and a right surround back input channel 610H. The channels 610E, 610F, 610G, and 610H are examples of peripheral channels that can be provided to surround speakers. In some embodiments, the audio system 600 can receive and process an input audio signal having fewer or more channels.

[0108] The audio system 600 uses enhancements such as sub-band spatial processing and crosstalk cancellation of an input audio signal to generate an output signal including a left output channel 690L and a right output channel 690R. The left output channel 690L can be provided to a left speaker, and the right output channel 690R can be output to a right speaker. The output audio signal uses the left and right speakers (e.g., left speaker 110L and right speaker 110R) to provide a sense of space of a sound field associated with a surround sound input audio signal.

[0109] The audio system 600 includes gains 615A, 615B, 615C, 615D, 615E, 615F, 615G, and 615H, a shelving filter 620, a distributor 640, binaural filters 650A, 650B, 650C, and 650D, a left channel combiner 660A, a right channel combiner 660B, a sub-band spatial processor 630, a crosstalk cancellation processor 670, a left channel combiner 660C, a right channel combiner 660D, and an output gain 680.

[0110] Each of the gains 615A to 615H can receive a corresponding input channel 610A to 610H, and can apply a gain to the input channels 610A to 610H. The gains 615A to 615H can be different to adjust the gains of the input channels relative to each other, or can be the same. In some embodiments, a positive gain is applied to the left and right peripheral input channels 610E, 610F, 610G, and 610H, while a negative gain is applied to the central input channel 610C. For example, the gain 615A can apply 0 dB gain, the gain 615B can apply 0 dB gain, the gain 615C can apply -3 dB gain, the gain 615D can apply 0 dB gain, the gain 615E can apply 3 dB gain, the gain 615F can apply 3 dB gain, the gain 615G can apply 3 dB gain, and the gain 615H can apply 3 dB gain.

[0111] The gain 615A for the left input channel 610A is coupled to the left channel combiner 660A. The gain 615B for the right input channel 610B is coupled to the right channel combiner 660B. The gain 615C is coupled to the shelving filter 620. The gain 615D is coupled to the distributor 640. The gains 615E, 615F, 615G, and 615H for the peripheral input channels are each coupled to the binaural filters 650. In particular, the gain 615E is coupled to the binaural filter 650A, the gain 615F is coupled to the binaural filter 650B, the gain 615G is coupled to the binaural filter 650C, and the gain 615H is coupled to the binaural filter 650D.

[0112] Each of the binaural filters 650A, 650B, 650C, and 650D applies a head-related transfer function (HRTF) that describes a target source location from which a listener should perceive the sound of an input audio channel. Each binaural filter receives the input audio channel and generates a left output audio channel and a right output audio channel by applying the HRTF. The discussion of the binaural filters 250A, 250B, 250C, and 250D of the audio system 200 can be applied to the binaural filters 650A, 650B, 650C, and 650D. For example, each of the binaural filters 650A through 650D can apply an adjustment for an angular position associated with their respective input audio channels. In some embodiments, one or more of the binaural filters 650A through 650D can be bypassed or omitted from the audio system 600.

[0113] The left-channel combiner 660A is coupled to the gain 615A and the binaural filters 650A through 650D. The left-channel combiner 660A receives the left output audio channels of the binaural filters 650A through 650D and combines the left output audio channels with the output of the gain 615A. The right-channel combiner 660B is coupled to the gain 615B and the binaural filters 650A through 650D. The right-channel combiner 660B receives the right output audio channels of the binaural filters 650A through 650D and combines the right output audio channels with the output of the gain 615B.

[0114] In some embodiments, binaural filtering is performed after subband spatial processing. For example, the binaural filters can be applied to the left and right outputs that are adjusted for an angular position associated with the audio channels and that are suitable for the subband spatial processor 630. In some embodiments, the binaural filters are applied to the peripheral input audio channels, as Figure 6 shown. In some embodiments, the binaural filters are applied to the central input audio channel 610C or the low-frequency input audio channel 610D. In some embodiments, the binaural filters are applied to each input audio channel except for the low-frequency input audio channel 610D.

[0115] The sub-band spatial processor 630 performs sub-band spatial processing on the left and right input channels by adjusting the gains of the mid and side sub-band components of the left and right input channels to generate left and right spatially enhanced channels as outputs. The sub-band spatial processor 630 is coupled to the left channel combiner 660A to receive the left combined channel from the left channel combiner 660A and coupled to the right channel combiner 660B to receive the right combined channel from the right channel combiner 660B. Different from the sub-band spatial processors 230A, 230B, and 230C of the audio system 200 each processing corresponding left and right input channels, the sub-band spatial processor 630 processes the left and right channels after they are combined into the left and right combined channels. Thus, the audio system 600 can include only a single sub-band spatial processor 630. In some embodiments, Figure 3 the sub-band spatial processor 230 shown in

[0116] is an example of a single sub-band spatial processor 630. The crosstalk cancellation processor 670 performs crosstalk cancellation on the output of the sub-band spatial processor 630, which can represent the downmixed stereo signal of the input audio signal. The crosstalk cancellation processor 670 receives the left and right input channels from the sub-band spatial processor 630 and performs crosstalk cancellation to generate left and right crosstalk cancellation channels. The crosstalk cancellation processor 670 is coupled to the left channel combiner 260A and the right channel combiner 260B. In some embodiments, Figure 4 the crosstalk cancellation processor 270 shown in

[0117] is an example of the crosstalk cancellation processor 670. The shelving filter 620 receives the center input channel 610C and applies a high-frequency shelving or peaking filter. The shelving filter 620 provides "voice boost" on the center input channel 610C. In some embodiments, the shelving filter 620 is bypassed or omitted from the audio system 600. The shelving filter 620 can attenuate frequencies above the corner frequency. The shelving filter 620 is coupled to the left channel combiner 660C and the right channel combiner 660D. In some embodiments, the shelving filter 620 is defined by a 750 Hz corner frequency, +3 dB gain, and 0.8 Q factor. The shelving filter 620 generates left and right center channels as outputs.

[0118] The distributor 640 receives the low-frequency input channel 610D and separates the low-frequency input channel 610D into a left low-frequency channel and a right low-frequency channel. The distributor 640 is coupled to the left channel combiner 660C and the right channel combiner 660D, and provides the left low-frequency channel to the left channel combiner 660C and the right low-frequency channel to the right channel combiner 660D.

[0119] The left channel combiner 660C is coupled to the crosstalk cancellation processor 670, the overhead filter 620, and the distributor 640. The left channel combiner 660C receives the left crosstalk channel from the crosstalk cancellation processor 670, the left center channel from the overhead filter 620, and the left low-frequency channel from the distributor 640, and combines these channels into a left output channel.

[0120] The right channel combiner 660D is coupled to the crosstalk cancellation processor 670, the overhead filter 620, and the distributor 640. The right channel combiner 660D receives the right crosstalk channel from the crosstalk cancellation processor 670, the right center channel from the overhead filter 620, and the right low-frequency channel from the distributor 640, and combines these channels into a right output channel.

[0121] In some embodiments, the left center channel from the overhead filter 620 and the left low-frequency channel from the distributor 640 are combined by the left channel combiner 660A with the left output channels of the binaural filters 650A to 650D and the output of the gain 615A to generate a left combined channel. The right center channel from the overhead filter 620 and the right low-frequency channel from the distributor 640 are combined by the right channel combiner 660B with the right output channels of the binaural filters 650A to 650D and the output of the gain 615B to generate a right combined channel. The left combined channel and the right combined channel are input to the subband spatial processor 630 and the crosstalk cancellation processor 670. Here, the center and low-frequency channels receive subband spatial processing and crosstalk cancellation operations. The left channel combiner 660C and the right channel combiner 660D may be omitted. In some embodiments, one of the center or low-frequency channels receives subband spatial processing and crosstalk cancellation operations.

[0122] The output gain 680 is coupled to the left channel combiner 660C and the right channel combiner 660D. The output gain 680 applies a gain to the left output channel from the left channel combiner 660C and applies a gain to the right output channel from the right channel combiner 660D. The output gain 680 may apply the same gain to the left output channel and the right output channel, or may apply different gains. The output gain 680 outputs the left output channel 690L and the right output channel 690R of the channels representing the output signal of the audio system 600.

[0123] Figure 7 Illustrated is an example of a method 700 for enhancing an audio signal using the Figure 6 audio system 600 shown in. In some embodiments, the method 700 may include different and / or additional steps, or some steps may be in a different order.

[0124] Audio system 600 receives a multi-channel input audio signal 705. The input audio signal may include a left input channel 610A, a right input channel 610B, at least one left peripheral input channel, and at least one right peripheral input channel. The multi-channel audio signal may further include a center input channel 610C and a low-frequency input channel 610D.

[0125] The audio system 600 (e.g., gains 615A to 615H) applies a gain 710 to the channels of the multi-channel input audio signal. The gains 615A to 615H may vary to control the contribution of a particular input channel to the output signal generated by the audio system 600.

[0126] The audio system 600 (e.g., binaural filters 650A to 650D) applies a binaural filter 715 to each of the left and right peripheral channels. For example, binaural filter 650A generates a left output channel and a right output channel from the left surround input channel 610E by applying a head-related transfer function (HRTF). Binaural filter 650B generates a left output channel and a right output channel from the right surround input channel 610F by applying an HRTF. Binaural filter 650C generates a left output channel and a right output channel from the left surround rear input channel 610G by applying an HRTF. Binaural filter 650D generates a left output channel and a right output channel from the right surround rear input channel 610H by applying an HRTF.

[0127] The audio system 600 (e.g., overhead filter 620) applies an overhead filter 720 to the center input channel 610C. In some embodiments, a gain is applied to the center input channel 610C. Additionally, the overhead filter 620 separates the center input channel 610C into a left center channel and a right center channel.

[0128] The audio system 600 (e.g., splitter 640) separates 725 the low-frequency input channel into a left low-frequency channel and a right low-frequency channel.

[0129] The audio system 600 (e.g., left channel combiner 660A) combines 730 the left input channel 610A and the left output channels of the binaural filters 650A, 650B, 650C, and 650D to generate a left combined channel.

[0130] The audio system 600 (e.g., right channel combiner 660B) combines 735 the right input channel 610B and the right output channels of the binaural filters 650A, 650B, 650C, and 650D to generate a right combined channel.

[0131] An audio system 600 (e.g., a subband spatial processor 630) generates left and right spatially enhanced channels 740 by performing subband spatial processing on left and right combined channels. For example, the subband spatial processor 630 receives the left and right combined channels from a left channel combiner 660A and a right channel combiner 660B, and generates the spatially enhanced channels by adjusting the gains of n subbands of the middle and side components of the left and right combined channels.

[0132] An audio system 600 (e.g., a crosstalk cancellation processor 670) performs crosstalk cancellation 745 on the left and right spatially enhanced channels from the subband spatial processor 630 to generate left and right crosstalk-cancelled channels.

[0133] An audio system 600 (e.g., a left channel combiner 660C and a right channel combiner 660D) combines 750 the left crosstalk-cancelled channel from the crosstalk cancellation processor 670 with the left low-frequency channel from a distributor 640 and the left center channel from an overhead filter 620 to generate a left output channel, and combines the right crosstalk-cancelled channel from the crosstalk cancellation processor 670 with the right low-frequency channel from the distributor 640 and the right center channel from the overhead filter 620 to generate a right output channel. Additionally, an audio system 600 (e.g., an output gain 680) can apply a gain to each of the left and right output channels. The audio system 600 outputs an output audio signal including the left and right output channels 690L and 690R.

[0134] Note that the systems and processes described herein can be embodied in an embedded electronic circuit or electronic system. The systems and processes can also be embodied in a computing system that includes one or more processing systems (e.g., a digital signal processor) and memory (e.g., a programmed read-only memory or programmable solid-state memory), or in some other circuit device, such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA) circuit.

[0135] Figure 8FIG. illustrates an example of a computer system 800 in accordance with one embodiment. The computer system 800 is an example of circuitry implementing an audio system. At least one processor 802 is illustrated coupled to a chipset 804. The chipset 804 includes a memory controller hub 820 and an input / output (I / O) controller hub 822. A memory 806 and a graphics adapter 812 are coupled to the memory controller hub 820, and a display device 818 is coupled to the graphics adapter 812. A storage device 808, a keyboard 810, a pointing device 814, and a network adapter 816 are coupled to the I / O controller hub 822. Other embodiments of the computer 800 have different architectures. For example, in some embodiments, the memory 806 is directly coupled to the processor 802.

[0136] The storage device 808 includes one or more non-transitory computer-readable storage media, such as a hard disk drive, a compact disc read-only memory (CD-ROM), a DVD, or a solid-state memory device. The memory 806 holds instructions and data used by the processor 802. For example, the memory 806 may store instructions that, when executed by the processor 802, cause or configure the processor 802 to perform the methods discussed herein, such as method 500 or 700. The pointing device 814 is used in conjunction with the keyboard 810 to input data into the computer system 800. The graphics adapter 812 displays images and other information on the display device 818. In some embodiments, the display device 818 includes touchscreen capabilities for receiving user input and selections. The network adapter 816 couples the computer system 800 to a network. Some embodiments of the computer 800 have components different from and / or additional to those shown in Figure 8 Those shown. For example, the computer system 800 may be a server that lacks a display device, a keyboard, and other components.

[0137] The computer 800 is adapted to execute computer program modules to provide the functionality described herein. As used herein, the term "module" refers to computer program instructions and / or other logic used to provide the specified functionality. Thus, a module may be implemented in hardware, firmware, and / or software. In one embodiment, program modules formed of executable computer program instructions are stored on the storage device 808, loaded into the memory 806, and executed by the processor 802.

[0138] Other examples of circuitry that may implement an audio system may include application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), and the like.

[0139] Example Audio Systems and Example Audio Processing Procedures

[0140] Figure 9Illustrated is an example of an audio system 900 according to one embodiment. Audio system 900 is similar to audio system 200, except that crosstalk processing is performed on each left and right channel pair before combining into left output channel 990L and right output channel 990R. Applying crosstalk processing and sub-band spatial processing separately to each left and right channel pair provides an opportunity for a unique sub-band spatial processing and crosstalk processing configuration for each "virtual" speaker pair. For example, the sub-band spatial processing for a given left and right channel pair can be configured to apply more or less per-band emphasis to the spatial components in the signal, resulting in an increase or decrease in the perceived spatial "intensity" compared to other channel pairs. Similarly, for a given left and right channel pair, the crosstalk processing filter and delay parameters can be uniquely configured based on the binaural filtering applied to that channel pair to achieve the maximum perceived effect.

[0141] Audio system 900 receives an input audio signal, which includes a left input channel 910A, a right input channel 910B, a center input channel 910C, a low-frequency input channel 910D, a left surround input channel 910E, a right surround input channel 910F, a left surround back input channel 910G, and a right surround back input channel 910H. The left input channel 910A and the right input channel 910B form a left and right channel pair for the front speakers. The left surround input channel 910E and the right surround input channel 910F form another left and right channel pair, and the left surround back input channel 910G and the right surround back input channel 910H form another left and right channel pair. These other left and right channel pairs are peripheral left and right channel pairs. Audio system 900 performs one or more of sub-band spatial processing and crosstalk cancellation on each of the left and right channel pairs and combines the outputs into left output channel 990L and right output channel 990R.

[0142] Audio system 900 includes gains 915A, 915B, 915C, 915D, 915E, 915F, 915G, and 915H, binaural filters 950A, 950B, 950C, 950D, 950E, and 950F, sub-band spatial processors 930A, 930B, and 930C, crosstalk cancellation processors 970A, 970B, and 970C, an overhead filter 920, a distributor 940, a left channel combiner 960A, a right channel combiner 960B, and an output gain 980.

[0143] Each of the gains 915A to 915H can receive the corresponding input channels 910A to 910H and can apply a gain to the input channels 910A to 910H. The gains 915A to 915H can be different to adjust the gains of the input channels relative to each other, or can be the same.

[0144] The binaural filters are applied to the channels of the left and right channel pair. Gain 915A is coupled to binaural filter 950A, gain 915B is coupled to binaural filter 950B, gain 915E is coupled to binaural filter 950C, gain 915F is coupled to binaural filter 950D, gain 915G is coupled to binaural filter 950E, and gain 915H is coupled to binaural filter 950F. Each of the binaural filters 950A, 950B, 950C, 950D, 950E, and 950F applies a head-related transfer function (HRTF) that describes the target source location from which the listener is to perceive the sound of the input channels. Each binaural filter receives the input channels and generates a left output channel and a right output channel by applying the HRTF adjusted for the angular position associated with the input channels. The angular position can include an angle defined in the X-Y "azimuth" plane relative to the listener 140, as Figure 1 shown, and can also include an angle defined in the Z axis, such as for an ambient sound signal or a channel-based format that includes signals intended to be rendered above or below the X-Y plane relative to the listener 140.

[0145] For example, binaural filter 950A can apply the filter based on the left input channel 910A associated with an angle between -30° and -45° relative to the forward axis of the left speaker 110L. Binaural filter 950B can apply the filter based on the right input channel 910B associated with an angle between 30° and 45° relative to the forward axis of the right speaker 110R. Binaural filter 950C can apply the filter based on the left surround input channel 910E associated with an angle between -90° and -110° relative to the forward axis of the left surround speaker 120L. Binaural filter 950D can apply the filter based on the right surround input channel 910F associated with an angle between 90° and 110° relative to the forward axis of the right surround speaker 120R. Binaural filter 950E can apply the filter based on the left surround rear input channel 910G associated with -135° to -150° relative to the forward axis of the left surround rear speaker 130L. Binaural filter 950F can apply the filter based on the right surround rear input channel 910H associated with an angle between 135° and 150° relative to the forward axis of the right surround rear speaker 130R. Each of the binaural filters 950A through 950F generates a left channel and a right channel.

[0146] In some embodiments, binaural processing on the left and right input channels 910A and 910B can be bypassed. Here, the binaural filters 950A and 950B can be omitted from the audio system 900. In some embodiments, binaural processing can be completely bypassed to preserve inter-channel spectral uniformity. One or more of the binaural filters 950A, 950B, 950C, 950D, 950E, or 950F can be omitted from the audio system 900.

[0147] In some embodiments, the input audio signal is an ambient stereo audio signal that defines a loudspeaker-independent representation of the sound field. The ambient stereo audio signal can be decoded into a multi-channel audio signal for a surround sound system. The channels can be associated with loudspeaker positions at different locations, including positions above or below the listener. The binaural filter can be applied to each decoded input channel of the ambient audio signal to adjust for the associated position of the decoded input audio channel.

[0148] Each of the subband spatial processors 930 applies subband spatial processing to different left and right channel pairs. The subband spatial processor 930A is coupled to each of the binaural filters 950A and 950B. The subband spatial processor 930A receives the left channels from each of the binaural filters 950A and 950B, combines these left channels into a combined left channel, and applies subband spatial processing to the combined left channel. The subband spatial processor 930A receives the right channels from each of the binaural filters 950A and 950B, combines these right channels into a combined right channel, and applies subband spatial processing to the combined right input channel. The subband spatial processor 930A performs subband spatial processing on the left and right input channels by adjusting the gains of the mid and side subband components of the left and right input channels to generate a left spatially enhanced channel and a right spatially enhanced channel.

[0149] The subband spatial processor 930B is coupled to each of the binaural filters 950C and 950D. The subband spatial processor 930B receives the left channels from each of the binaural filters 950C and 950D, combines these left channels into a combined left channel, and applies subband spatial processing to the combined left channel. The subband spatial processor 930B receives the right channels from each of the binaural filters 950C and 950D, combines these right channels into a combined right channel, and applies subband spatial processing to the combined right channel. The subband spatial processor 930B performs subband spatial processing on the left and right input channels by adjusting the gains of the mid and side subband components of the left and right input channels to generate a left spatially enhanced channel and a right spatially enhanced channel.

[0150] The subband spatial processor 930C is coupled to each of the binaural filters 950E and 950F. The subband spatial processor 930C receives the left channels from each of the binaural filters 950E and 950F, combines these left channels into a combined left channel, and applies subband spatial processing to the combined left channel. The subband spatial processor 930C receives the right channels from each of the binaural filters 950E and 950F, combines these right channels into a combined right channel, and applies subband spatial processing to the combined right channel. The subband spatial processor 930C performs subband spatial processing on the left input channel and the right input channel by performing gain adjustment on the middle components and the side subband components of the left input channel and the right input channel to generate a left spatially enhanced channel and a right spatially enhanced channel.

[0151] Each of the crosstalk cancellation processors 970 applies crosstalk cancellation to different left and right channel pairs. The crosstalk cancellation processor 970A is coupled to the subband spatial processor 930A, the crosstalk cancellation processor 970B is coupled to the subband spatial processor 930B, and the crosstalk cancellation processor 970C is coupled to the subband spatial processor 930C.

[0152] The crosstalk cancellation processor 970A receives the left spatially enhanced channel and the right spatially enhanced channel from the subband spatial processor 930A, and applies crosstalk cancellation processing to the left spatially enhanced channel and the right spatially enhanced channel to generate a left output channel and a right output channel. These left output channels and right output channels correspond to the left and right channel pairs formed by the left input channel and the right input channels 910A and 910B after subband spatial processing and crosstalk cancellation.

[0153] The crosstalk cancellation processor 970B receives the left spatially enhanced channel and the right spatially enhanced channel from the subband spatial processor 930B, and applies crosstalk cancellation processing to the left spatially enhanced channel and the right spatially enhanced channel to generate a left output channel and a right output channel. These left output channels and right output channels correspond to the left and right channel pairs formed by the left surround input channel and the right surround input channels 910E and 910F after subband spatial processing and crosstalk cancellation.

[0154] The crosstalk cancellation processor 970C receives the left spatially enhanced channel and the right spatially enhanced channel from the subband spatial processor 930C, and applies crosstalk cancellation processing to the left spatially enhanced channel and the right spatially enhanced channel to generate a left output channel and a right output channel. These left output channels and right output channels correspond to the left and right channel pairs formed by the left and right surround post-input channels 910G and 910H after subband spatial processing and crosstalk cancellation.

[0155] The elevation filter 920 is coupled to the gain 915C. The elevation filter 920 receives the center input channel 910C and applies a high shelf or peak filter. The elevation filter 920 can attenuate or amplify frequencies above the corner frequency. In some embodiments, the elevation filter 920 is defined by a 750 Hz corner frequency, +3 dB gain, and 0.8 Q factor. The elevation filter 920 generates a left center channel and a right center channel as outputs, such as by splitting the center input channel into two separate left center and right center channels. In some embodiments, the elevation filter 920 is bypassed or omitted from the audio system 900.

[0156] The distributor 940 is coupled to the gain 915D. The distributor 940 receives the low frequency input channel 910D and separates the low frequency input channel 910D into a left low frequency channel and a right low frequency channel.

[0157] The left channel combiner 960A and the right channel combiner 960B are each coupled to the crosstalk cancellation processors 970A, 970B, 970C, the elevation filter 920, and the distributor 940. The left channel combiner 960A receives the left channels output from each of the crosstalk cancellation processors 970A, 970B, 970C, the elevation filter 920, and the distributor 940, and combines these left channels into a left output channel. The right channel combiner 960B receives the right channels output from each of the crosstalk cancellation processors 970A, 970B, 970C, the elevation filter 920, and the distributor 940, and combines these right channels into a right output channel.

[0158] The output gain 980 is coupled to the left channel combiner 960A and the right channel combiner 960B. The output gain 980 applies gain to the left output channel from the left channel combiner 960A and applies gain to the right output channel from the right channel combiner 960B. The output gain 980 can apply the same gain to the left output channel and the right output channel, or can apply different gains. The output gain 980 outputs the left output channel 990L and the right output channel 990R of the channels representing the output signal of the audio system 900.

[0159] Figure 10 An example of an audio system 1000 according to one embodiment is illustrated. The audio system 1000 is similar to the audio system 900, but differs from the audio system 900 at least in that a binaural filter is applied after subband spatial processing and before crosstalk cancellation processing on one or more of the left and right channel pairs.

[0160] The audio system 1000 includes gains 915A, 915B, 915C, 915D, 915E, 915F, 915G, and 915H, sub-band spatial processors 930A, 930B, and 930C, crosstalk cancellation processors 970A, 970B, and 970C, and an overhead filter 920, a distributor 940, a left-channel combiner 960A, a right-channel combiner 960B, and an output gain 980. The audio system 1000 also includes binaural filters 1050A, 1050B, 1050C, 1050D, 1050E, and 1050F.

[0161] The binaural filters 1050A and 1050B are coupled to the sub-band spatial processor 930A and the crosstalk cancellation processor 970A. The binaural filters 1050A and 1050B apply binaural filtering to a left and right channel pair including a left input channel 910A and a right input channel 910B after sub-band spatial processing and before crosstalk cancellation processing. In some embodiments, the binaural filters 1050A and 1050B can be bypassed or excluded from the audio system 1000.

[0162] The audio system 100 applies similar sub-band spatial processing, binaural filtering, and crosstalk cancellation processing to each peripheral left and right channel pair. For processing a left and right channel pair including a left surround input channel 910E and a right surround input channel 910F, the binaural filters 1050C and 1050D are coupled to the sub-band spatial processor 930B and the crosstalk cancellation processor 970B. For processing a left and right channel pair including a left surround rear input channel 910G and a right surround rear input channel 910H, the binaural filters 1050E and 1050F are coupled to the sub-band spatial processor 930C and the crosstalk cancellation processor 970C.

[0163] In some embodiments, the crosstalk cancellation processors 970A, 970B, and 970C can each be a crosstalk simulation processor. Instead of generating a crosstalk cancellation channel, the crosstalk simulation processor generates a crosstalk simulation channel with an additional crosstalk effect.

[0164] Figure 11 Illustrated is an example of a method 1100 for enhancing an audio signal using the audio system 900 shown in Figure 9 or the audio system 1000 shown in Figure 10 according to one embodiment. In some embodiments, the method 1100 can include different and / or additional steps, or some steps can be in a different order. The method 1100 is discussed in more detail below with reference to the audio system 900.

[0165] The audio system 900 receives 1105 a multi-channel input audio signal that includes a left and right channel pair. The multi-channel audio signal can be a surround audio signal that includes multiple left and right channel pairs. For example, a left input channel and a right input channel can form a first left and right channel pair, and at least one left surround input channel and at least one right surround input channel can form another left and right channel pair. The multi-channel input signal can include multiple left and right channel pairs for the surround input channels. For example, left surround input channels 910E and 910F form a surround pair, and left rear surround input channel 910G and right rear surround input channel 910H form a rear surround pair. The multi-channel audio signal can also include a center input channel and a low-frequency input channel.

[0166] The audio system 900 (e.g., gains 915A to 915H) applies 1110 a gain to the channels of the multi-channel input audio signal. The gains 915A to 915H can be varied to control the contribution of a particular input channel pair to the output signal generated by the audio system 900.

[0167] The audio system 900 (e.g., binaural filters 950A to 950F) applies 1115 a binaural filter to each of the left and right channel pairs of the multi-channel input audio signal. For each channel, the binaural filter is adjusted for an angular position associated with the channel. In some embodiments, the binaural filter is applied to the surround left and right channel pairs, but not to the left and right channel pair that includes the left input channel and the right input channel.

[0168] The audio system 900 (e.g., subband spatial processors 930A, 930B, and 930C) applies 1120 subband spatial processing to each left and right channel pair to generate spatially enhanced channels. For example, subband spatial processor 930A applies subband spatial processing to the left and right channel pair that includes left input channel 910A and right input channel 910B to generate a spatially enhanced channel. The subband spatial processing includes gain adjustments to the mid and side components of left input channel 910A and right input channel 910B.

[0169] Sub-band spatial processing is also applied to at least one of the left and right channel pairs for the surround channels. For example, the sub-band spatial processor 930B applies sub-band spatial processing to the left and right channel pair including the left surround input channel 910E and the right surround input channel 910F to generate spatially enhanced channels. The sub-band spatial processing includes gain adjustment of the mid and side components of the left surround input channel 910E and the right surround input channel 910F. The sub-band spatial processor 930C applies sub-band spatial processing to the left and right channel pair including the left rear surround input channel 910G and the right rear surround input channel 910H to create spatially enhanced channels. The sub-band spatial processing includes gain adjustment of the mid and side components of the left rear surround input channel 910G and the right rear surround input channel 910H. Thus, spatially enhanced channels are created for each of the left and right channel pairs.

[0170] In some embodiments, sub-band spatial processing for each left and right channel pair is performed before binaural filtering, as for the audio system 1000 Figure 10 shown. Here, each of the left and right spatially enhanced channels output from the sub-band spatial processors 930A, 930B, and 930C is input to the binaural filter.

[0171] The audio system 900 (e.g., the crosstalk cancellation processors 970A, 970B, and 970C) applies crosstalk processing 1125 to each left and right channel pair to generate crosstalk processed channels. The crosstalk processing may include crosstalk cancellation or crosstalk simulation. In the case of crosstalk cancellation, the crosstalk processed channels include crosstalk cancellation channels. In the case of crosstalk simulation, the crosstalk processed channels include crosstalk simulation channels. Crosstalk cancellation may be used for speaker output, and crosstalk simulation may be used for headphone output. For each left and right channel pair, the crosstalk processing may include applying a filter, a time delay, and a gain to at least one of the spatially enhanced channels to generate the crosstalk processed channels. In some embodiments, crosstalk processing may be performed on each left and right channel pair before sub-band spatial processing of each left and right channel pair.

[0172] The audio system 900 (e.g., the left channel combiner 960A and the right channel combiner 960B) generates 1130 the left output channel and the right output channel from the crosstalk processed channels. For example, the left channel combiner 960A combines the left channels of the crosstalk processed channels from each of the crosstalk cancellation processors 970A, 970B, and 970C to generate the left output channel, and the right channel combiner 960B combines the right channels of the crosstalk processed channels from each of the crosstalk cancellation processors 970A, 970B, and 970C to generate the right output channel.

[0173] The left channel combiner 960A can also combine the left channel with the left low - frequency channel and the left center channel to generate a left output channel. The right channel combiner 960B can also combine the right channel with the right low - frequency channel and the right center channel to generate a right output channel. The audio system 900 (e.g., the elevation filter 920) applies an elevation filter to the center input channel of the multi - channel input audio signal to generate a left center channel and a right center channel. The audio system 900 (e.g., the splitter 940) applies a separation that separates the low - frequency input channel into the center input channel of the multi - channel input audio signal to generate a left low - frequency channel and a right low - frequency channel.

[0174] Figure 12 An example of a crosstalk simulation processor 1200 according to one embodiment is illustrated. When the crosstalk processing is crosstalk simulation, the crosstalk simulation processor 1200 can be used in the audio system instead of the crosstalk cancellation processor. The crosstalk simulation processor 1200 can be used to provide a speaker - like listening experience on a head - mounted speaker.

[0175] The crosstalk simulation processor 1200 includes a left head - related low - pass filter 1202, a left head - related high - pass filter 1204, a left crosstalk delay 1210, and a left head - related gain 1224 to process the left channel (e.g., the left spatially enhanced channel E L ). The crosstalk simulation processor 1200 also includes a right head - related low - pass filter 1206, a right head - related high - pass filter 1208, a right crosstalk delay 1212, and a right head - related gain 1226 to process the right channel (e.g., the right spatially enhanced channel E R ).

[0176] The left head - related low - pass filter 1202 and the left head - related high - pass filter 1204 each apply a modulation that models the frequency response of the signal after passing through the listener's head. The left crosstalk delay 1210 applies a time delay that represents the inter - aural distance traveled by the contralateral sound component relative to the ipsilateral sound component. The frequency response can be generated based on empirical experiments to determine the frequency - dependent characteristics of the acoustic wave modulation of the listener's head. In some embodiments, the left crosstalk delay 1210 can be applied before the left head - related low - pass filter 1202 and the left head - related high - pass filter 1204. The left head - related gain 1224 applies a gain to generate a left crosstalk simulation channel O L .

[0177] The right HRTF low-pass filter 1206 and the right HRTF high-pass filter 1208 each apply a modulation that models the frequency response of the signal after passing through the listener's head. The right crosstalk delay 1212 applies a time delay that represents the interaural distance traversed by the contralateral sound component relative to the ipsilateral sound component. The frequency response can be generated based on empirical experiments to determine the frequency-dependent characteristics of the acoustic wave modulation of the listener's head. In some embodiments, the right crosstalk delay 1212 can be applied before the right HRTF low-pass filter 1206 and the right HRTF high-pass filter 1208. The right HRTF gain 1226 applies a gain to generate the right crosstalk simulation channel O L 。

[0178] The application of the HRTF low-pass filter, the HRTF high-pass filter, the crosstalk delay, and the HRTF gain for each of the left and right channels can be performed in a different order, and one or more of these stages can be skipped. Using both low-pass and high-pass filters simultaneously on the left and right channels may result in a more accurate model of the frequency response through the listener's head.

[0179] Other considerations

[0180] The disclosed configuration can include many benefits and / or advantages. For example, a multi-channel input signal can be output to stereo speakers while preserving or enhancing the spatial sense of the sound field. On mobile devices, soundbars, or smart speakers, for example, a high-quality listening experience can be obtained without the need for an expensive multi-speaker audio system.

[0181] After reading this disclosure, those skilled in the art will appreciate additional alternative embodiments of the principles disclosed herein. Thus, while specific embodiments and applications have been illustrated and described, it should be understood that the disclosed embodiments are not limited to the exact structures and components disclosed herein. Various modifications, changes, and variations to the arrangements, operations, and details of the methods and apparatuses disclosed herein will be apparent to those skilled in the art without departing from the scope described herein.

[0182] Any of the steps, operations, or processes described herein can be performed or implemented using one or more hardware or software modules alone or in combination with other devices. In one embodiment, the software module is implemented with a computer program product that includes a computer-readable medium (e.g., a non-transitory computer-readable medium) that includes computer program code that can be executed by a computer processor to perform any or all of the steps, operations, or processes described.

Claims

1. A system for processing an audio signal, comprising: circuit means configured to: receive an audio signal that defines a loudspeaker-independent representation of a sound field; decode the audio signal into a multi-channel audio signal comprising decoded channels, each decoded channel corresponding to a respective loudspeaker position having an angular position including an angle defined in a Z-axis, the Z-axis defining positions above and below an X-Y azimuth plane of a listening position; apply binaural processing to the decoded channels to generate binaurally processed channels, the binaural processing including adjusting a respective head-related transfer function (HRTF) of a respective angular position for each decoded channel, the respective angular position including the angle defined in the Z-axis of the decoded channel; and apply crosstalk simulation to the binaurally processed channels to generate a left output channel and a right output channel.

2. The system according to claim 1, wherein the audio signal comprises an ambisonic signal.

3. The system according to claim 1, wherein an angle defined in the Z-axis of at least one of the decoded channels defines a position above the listening position.

4. The system according to claim 1, wherein an angle defined in the Z-axis of at least one of the decoded channels defines a position below the listening position.

5. The system according to claim 1, wherein the decoded channels at least comprise a first channel pair and a second channel pair, wherein the first channel pair comprises a left channel and a right channel, and the second channel pair comprises a surround channel pair.

6. The system according to claim 5, wherein the binaural processing comprises: performing a first binaural processing on the first channel pair to generate a first binaurally processed channel pair, the first binaural processing including adjusting a respective HRTF of a respective angular position including the angle defined in the Z-axis of the decoded channel for each decoded channel of the first channel pair; and performing a second binaural processing on the second channel pair to generate a second binaurally processed channel pair, the second binaural processing including adjusting a respective HRTF of a respective angular position including the angle defined in the Z-axis of the decoded channel for each decoded channel of the second channel pair; and wherein the crosstalk simulation comprises: a first crosstalk simulation applied to the first binaurally processed channel pair to generate a first crosstalk simulated left channel and a first crosstalk simulated right channel; and a second crosstalk simulation applied to the second binaurally processed channel pair to generate a second crosstalk simulated left channel and a second crosstalk simulated right channel.

7. The system according to claim 6, wherein the left output channel is generated based on the first crosstalk simulated left channel and the second crosstalk simulated left channel, and the right output channel is generated based on the first crosstalk simulated right channel and the second crosstalk simulated right channel.

8. The system according to claim 1, wherein the decoded channels further comprise at least one of a top channel and a rear center channel.

9. The system according to claim 1, wherein the circuitry is further configured to filter the mid and side components of the left and right channel pairs of the decoded channels.

10. The system according to claim 1, wherein the circuitry is further configured to: apply a first sub-band spatial processing to a first left and right channel pair of the decoded channels, the first left and right channel pair of the decoded channels including a left input channel and a right input channel, the first sub-band spatial processing including gain adjustment of the mid and side components of the left input channel and the right input channel; and apply a second sub-band spatial processing to a second left and right channel pair of the decoded channels, the second left and right channel pair of the decoded channels including a left peripheral input channel and a right peripheral input channel, the second sub-band spatial processing including gain adjustment of the mid and side components of the left peripheral input channel and the right peripheral input channel.

11. A non-transitory computer-readable medium storing program code that, when executed by a processor, causes the processor to: receive an audio signal that defines a loudspeaker-independent representation of a sound field; decode the audio signal into a multi-channel audio signal including decoded channels, each decoded channel corresponding to a respective loudspeaker position having an angular position including an angle defined on a Z-axis that defines positions above and below an X-Y azimuthal plane of a listening position; apply binaural processing to the decoded channels to generate binaurally processed channels, the binaural processing including adjusting a respective head-related transfer function (HRTF) for a respective angular position of each decoded channel, the respective angular position including the angle defined in the Z-axis of the decoded channel; and apply crosstalk simulation to the binaurally processed channels to generate a left output channel and a right output channel.

12. The non-transitory computer-readable medium according to claim 11, wherein the audio signal includes an ambisonic signal.

13. The non-transitory computer-readable medium according to claim 11, wherein the angle defined in the Z-axis of at least one of the decoded channels defines a position above the listening position.

14. The non-transitory computer-readable medium according to claim 11, wherein the angle defined in the Z-axis of at least one of the decoded channels defines a position below the listening position.

15. The non-transitory computer-readable medium according to claim 11, wherein the decoded channels include at least a first channel pair and a second channel pair, wherein the first channel pair includes a left channel and a right channel, and the second channel pair includes a peripheral channel pair.

16. The non-transitory computer-readable medium according to claim 15, wherein the binaural processing includes: performing a first binaural processing on the first channel pair to generate a first pair of binaurally processed channels, the first binaural processing including adjusting a respective HRTF for a respective angular position including the angle defined in the Z-axis of each decoded channel of the first channel pair; and Perform second binaural processing on the second channel pair to generate a second binaurally processed channel pair, the second binaural processing including adjusting, for each decoded channel of the second channel pair, a respective HRTF of a respective angular position defining an angle included in the Z-axis of the decoded channel; And wherein the crosstalk simulation includes: a first crosstalk simulation applied to the first binaurally processed channel pair to generate a first crosstalk simulated left channel and a first crosstalk simulated right channel; And a second crosstalk simulation applied to the second binaurally processed channel pair to generate a second crosstalk simulated left channel and a second crosstalk simulated right channel.

17. The non-transitory computer-readable medium according to claim 16, wherein the program code, when executed by the processor, further causes the processor to generate the left output channel based on the first crosstalk simulated left channel and the second crosstalk simulated left channel, and to generate the right output channel based on the first crosstalk simulated right channel and the second crosstalk simulated right channel.

18. The non-transitory computer-readable medium according to claim 11, wherein the decoded channels further include at least one of an overhead channel and a posterior center channel.

19. The non-transitory computer-readable medium according to claim 11, wherein the program code, when executed by the processor, further causes the processor to filter an intermediate component and a side component of a left and right channel pair of the decoded channels.

20. The non-transitory computer-readable medium according to claim 11, wherein the program code, when executed by the processor, further causes the processor to: Apply first sub-band spatial processing to a first left and right channel pair of the decoded channels, the first left and right channel pair of the decoded channels including a left input channel and a right input channel, the first sub-band spatial processing including gain adjustment of an intermediate component and a side component of the left input channel and the right input channel; and Apply second sub-band spatial processing to a second left and right channel pair of the decoded channels, the second left and right channel pair of the decoded channels including a left peripheral input channel and a right peripheral input channel, the second sub-band spatial processing including gain adjustment of an intermediate component and a side component of the left peripheral input channel and the right peripheral input channel.

21. A method for processing a multi-channel input audio signal by a circuit device, comprising: Receiving an audio signal defining a loudspeaker-independent representation of a sound field; Decoding the audio signal into a multi-channel audio signal including decoded channels, each decoded channel corresponding to a respective loudspeaker position having an angular position including an angle defined in the Z-axis, the Z-axis defining positions above and below an X-Y azimuth plane of a listening position; Applying binaural processing to the decoded channels to generate binaurally processed channels, the binaural processing including adjusting, for each decoded channel, a respective head-related transfer function HRTF of a respective angular position defining an angle included in the Z-axis of the decoded channel; And Apply crosstalk simulation to the binaurally processed sound channels to generate a left output sound channel and a right output sound channel.

22. The method according to claim 21, wherein the audio signal includes an ambient stereo signal.

23. The method according to claim 21, wherein the angle defined in the Z-axis of at least one of the decoded sound channels defines a position above the listening position.

24. The method according to claim 21, wherein the angle defined in the Z-axis of at least one of the decoded sound channels defines a position below the listening position.

25. The method according to claim 21, wherein the decoded sound channels at least include a first sound channel pair and a second sound channel pair, wherein the first sound channel pair includes a left sound channel and a right sound channel, and the second sound channel pair includes a peripheral sound channel pair.

26. The method according to claim 25, wherein applying the binaural processing includes: Applying a first binaural processing to the first sound channel pair to generate a first binaurally processed sound channel pair, the first binaural processing including adjusting, for each decoded sound channel of the first sound channel pair, the corresponding HRTF of the corresponding angular position including the angle defined in the Z-axis of the decoded sound channel; and Applying a second binaural processing to the second sound channel pair to generate a second binaurally processed sound channel pair, the second binaural processing including adjusting, for each decoded sound channel of the second sound channel pair, the corresponding HRTF of the corresponding angular position including the angle defined in the Z-axis of the decoded sound channel; And wherein applying the crosstalk simulation includes: applying a first crosstalk simulation to the first binaurally processed sound channel pair to generate a first crosstalk simulated left sound channel and a first crosstalk simulated right sound channel, and applying a second crosstalk simulation to the second binaurally processed sound channel pair to generate a second crosstalk simulated left sound channel and a second crosstalk simulated right sound channel.

27. The method according to claim 26, further comprising: Generating the left output sound channel based on the first crosstalk simulated left sound channel and the second crosstalk simulated left sound channel, and generating the right output sound channel based on the first crosstalk simulated right sound channel and the second crosstalk simulated right sound channel.

28. The method according to claim 21, wherein the decoded sound channels further include at least one of a top sound channel and a rear center sound channel.

29. The method according to claim 21 further comprises: Filter the middle component and the side component of the left and right sound channel pairs of the decoded sound channels.

30. The method according to claim 21, further comprising: Applying a first sub-band spatial processing to a first left and right sound channel pair of the decoded sound channels, the first left and right sound channel pair of the decoded sound channels including a left input sound channel and a right input sound channel, the first sub-band spatial processing including gain adjustment of the middle component and the side component of the left input sound channel and the right input sound channel; And Applying a second sub-band spatial processing to a second left and right sound channel pair of the decoded sound channels, the second left and right sound channel pair of the decoded sound channels including a left peripheral input sound channel and a right peripheral input sound channel, the second sub-band spatial processing including gain adjustment of the middle component and the side component of the left peripheral input sound channel and the right peripheral input sound channel.