System and method for processing audio signal

By generating audio components with orthogonal spectrum and filtering, the isolation problem of intermediate components and side components in stereo audio signals is solved, and the precise processing and spatial enhancement effect of stereo audio signals is achieved.

CN120358447APending Publication Date: 2025-07-22BOOMCLOUD 360 INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510481943.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-08-03
Filing Date
2020-08-10
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art is difficult to effectively isolate and process intermediate and side components in stereo audio signals, resulting in limited range of possibility for audio processing.

Method used

Enhance audio content by generating super-intermediate components, super-side components, residual intermediate components and residual side components that are orthogonal to the spectrum, and filtering and signaling these components, including Fourier transform, gain adjustment, time delay, dynamic range processing, and machine learning-based style transfer.

Benefits of technology

Accurate processing of stereo audio signals is achieved, and the audio content can be adjusted without changing the spectrum energy in other parts of the sound field, providing spatial enhancement and vocal enhancement effects while reducing vocal loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358447A_ABST
    Figure CN120358447A_ABST
Patent Text Reader

Abstract

Systems and methods for processing an audio signal are provided. A system includes circuitry that generates a middle component and a side component from left and right channels of an audio signal. The circuitry generates a hyper-lateral component that includes a spectral energy that removes the intermediate component from the spectral energy of the lateral component. The circuitry filters the hyper-lateral components and generates a left output channel and a right output channel using the filtered hyper-lateral components.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a Chinese patent application with application number 202080085638.3, application date August 10, 2020, earliest priority date October 10, 2019, and invention title "Spectrum Orthogonal Audio Component Processing". Technical Field

[0002] This disclosure generally relates to audio processing, and more particularly to spatial audio processing. Background Art

[0003] Conceptually, the side (or "spatial") components of a left-right stereo signal can be considered as the parts in the left and right channels that include spatial information (i.e., the sounds that appear at any position to the left or right of the center of the sound field). Conversely, the middle (or "non-spatial") components of a left-right stereo signal can be considered as the parts in the left and right channels that include non-spatial information (i.e., the sounds that appear at the center of the sound field). Although the middle components contain the energy in the stereo signal that is perceived as non-spatial, they typically also have energy from elements in the stereo signal that are not perceptually located at the center of the sound field. Similarly, although the side components contain the energy in the stereo signal that is perceived as spatial, they typically also have energy from elements in the stereo signal that are perceptually located at the center of the sound field. To enhance the range of possibilities for processing audio, it is necessary to isolate and manipulate parts of the middle and side components that are "orthogonal" to each other in the frequency spectrum. Summary of the Invention

[0004] Embodiments relate to audio processing using spectrum orthogonal audio components, such as the super-middle component, super-side component, residual middle component, or residual side component of a stereo audio signal or other multi-channel audio signal. The super-middle component and the super-side component are orthogonal to each other in the frequency spectrum, and the residual middle component and the residual side component are orthogonal to each other in the frequency spectrum.

[0005] Some embodiments include a system for processing an audio signal. The system includes circuitry for generating middle and side components from the left and right channels of the audio signal. The circuitry generates a super-middle component that includes removing the spectral energy of the side component from the spectral energy of the middle component. The circuitry filters the super-middle component, such as to provide spatial cue processing, including panning or binaural processing, dynamic range processing, or other types of processing. The circuitry generates a left output channel and a right output channel using the filtered super-middle component.

[0006] In some embodiments, the circuitry applies a Fourier transform to the middle and side components to transform the middle and side components into the frequency domain. The circuitry generates the super-middle component by subtracting the magnitude of the side component in the frequency domain from the magnitude of the middle component in the frequency domain.

[0007] In some embodiments, the circuit device filters the super-intermediate component to adjust the gain or time delay of the sub-bands of the super-intermediate component. In some embodiments, the circuit device filters the super-intermediate component to apply dynamic range processing to the super-intermediate component. In some embodiments, the circuit device filters the super-intermediate component to adjust the frequency-dependent amplitude or frequency-dependent delay of the super-intermediate component. In some embodiments, the circuit device filters the super-intermediate component to apply machine learning-based style transfer, transformation, or resynthesis to the super-intermediate component.

[0008] In some embodiments, the circuit device generates a residual intermediate component including removing the spectral energy of the super-intermediate component from the spectral energy of the intermediate component, filters the residual intermediate component, and generates a left output channel and a right output channel using the filtered residual intermediate component.

[0009] In some embodiments, the circuit device filters the residual intermediate component to adjust the gain or time delay of the sub-bands of the residual intermediate component. In some embodiments, the circuit device filters the residual intermediate component to apply dynamic range processing to the residual intermediate component. In some embodiments, the circuit device filters the residual intermediate component to adjust the frequency-dependent amplitude or frequency-dependent delay of the residual intermediate component. In some embodiments, the circuit device filters the residual intermediate component to apply machine learning-based style transfer, transformation, or resynthesis to the residual intermediate component.

[0010] In some embodiments, the circuit device applies a Fourier transform to the intermediate component to transform the intermediate component into the frequency domain. The circuit device generates a residual intermediate component by subtracting the magnitude of the super-intermediate component in the frequency domain from the magnitude of the intermediate component in the frequency domain.

[0011] In some embodiments, the circuit device applies an inverse Fourier transform to the super-intermediate component to transform the super-intermediate component in the frequency domain into the time domain, generates a delayed intermediate component by time-delaying the intermediate component, generates a residual intermediate component by subtracting the super-intermediate component in the time domain from the delayed intermediate component in the time domain, filters the residual intermediate component, and generates a left output channel and a right output channel using the filtered residual intermediate component.

[0012] In some embodiments, the circuit device generates a super-side component including removing the spectral energy of the intermediate component from the spectral energy of the side component, filters the super-side component, and generates a left output channel and a right output channel using the filtered super-side component.

[0013] In some embodiments, the circuit device applies Fourier transform to the intermediate component and the side component to transform the intermediate component and the side component into the frequency domain. The circuit device generates a super side component by subtracting the magnitude of the intermediate component in the frequency domain from the magnitude of the side component in the frequency domain.

[0014] In some embodiments, the circuit device filters the super side component to adjust the gain or time delay of the subbands of the super side component. In some embodiments, the circuit device filters the super side component to apply dynamic range processing to the super side component. In some embodiments, the circuit device filters the super side component to adjust the frequency-dependent amplitude or frequency-dependent delay of the super side component. In some embodiments, the circuit device filters the super side component to apply machine learning-based style transfer, transformation, or resynthesis to the super side component.

[0015] In some embodiments, the circuit device generates a super side component that includes removing the spectral energy of the intermediate component from the spectral energy of the side component, generates a residual side component that includes removing the spectral energy of the super side component from the spectral energy of the side component, filters the residual side component, and generates a left output channel and a right output channel using the filtered residual side component.

[0016] In some embodiments, the circuit device filters the residual side component to adjust the gain or time delay of the subbands of the residual side component. In some embodiments, the circuit device filters the residual side component to apply dynamic range processing to the residual side component. In some embodiments, the circuit device filters the residual side component to adjust the frequency-dependent amplitude or frequency-dependent delay of the residual side component. In some embodiments, the circuit device filters the residual side component to apply machine learning-based style transfer, transformation, or resynthesis to the residual side component.

[0017] In some embodiments, the circuit device applies Fourier transform to the side component to transform the side component into the frequency domain. The circuit device generates a residual side component by subtracting the magnitude of the super side component in the frequency domain from the magnitude of the side component in the frequency domain

[0018] In some embodiments, the circuit device generates a super side component that includes removing the spectral energy of the intermediate component from the spectral energy of the side component, applies inverse Fourier transform to the super side component to transform the super intermediate component into the time domain, generates a delayed side component by time-delaying the side component, generates a residual side component by subtracting the super side component in the time domain from the delayed side component in the time domain, filters the residual side component, and generates a left output channel and a right output channel using the filtered residual side component.

[0019] Some embodiments include a non-transitory computer-readable medium including stored program code. The program code, when executed by at least one processor, configures the at least one processor to generate an intermediate component and a side component from left and right channels of an audio signal, generate a super-intermediate component including removing spectral energy of the side component from spectral energy of the intermediate component, filter the super-intermediate component, and generate a left output channel and a right output channel using the filtered super-intermediate component.

[0020] Some embodiments include a method for processing an audio signal by a circuit device. The method includes generating an intermediate component and a side component from left and right channels of an audio signal, generating a super-intermediate component including removing spectral energy of the side component from spectral energy of the intermediate component, filtering the super-intermediate component, and generating a left output channel and a right output channel using the filtered super-intermediate component. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The disclosed embodiments have other advantages and features that will be more readily apparent from the detailed description, the appended claims, and the drawings (or figures). A brief introduction to the drawings is provided below.

[0022] Figure 1 is a block diagram of an audio processing system according to one or more embodiments.

[0023] Figure 2A is a block diagram of an orthogonal component generator according to one or more embodiments.

[0024] Figure 2B is a block diagram of an orthogonal component generator according to one or more embodiments.

[0025] Figure 2C is a block diagram of an orthogonal component generator according to one or more embodiments.

[0026] Figure 3 is a block diagram of an orthogonal component processor according to one or more embodiments.

[0027] Figure 4 is a block diagram of a subband spatial processor according to one or more embodiments.

[0028] Figure 5 is a block diagram of a crosstalk compensation processor according to one or more embodiments.

[0029] Figure 6 is a block diagram of a crosstalk simulation processor according to one or more embodiments.

[0030] Figure 7 is a block diagram of a crosstalk cancellation processor according to one or more embodiments.

[0031] Figure 8Is a flowchart of a process for spatial processing using at least one of a super intermediate component, a residual intermediate component, a super side component, or a residual side component according to one or more embodiments.

[0032] Figure 9 Is a flowchart of a process for sub-band spatial processing and crosstalk compensation processing using at least one of a super intermediate component, a residual intermediate component, a super side component, or a residual side component according to one or more embodiments.

[0033] Figures 10 - 19 Is a graph depicting the spectral energy of the intermediate component and the side component of an example white noise signal according to one or more embodiments.

[0034] Figure 20 Is a block diagram of a computer system according to one or more embodiments. Detailed Description

[0035] The accompanying drawings and the following description relate only by way of illustration to preferred embodiments. It should be noted that from the following discussion, alternative embodiments of the structures and methods disclosed herein will readily be recognized as viable alternatives that may be employed without departing from the principles claimed.

[0036] Reference will now be made in detail to several embodiments, examples of which are illustrated in the accompanying drawings. Note that wherever feasible, similar or like reference numerals may be used in the drawings and may indicate similar or like functions. The drawings depict embodiments of the disclosed system (or method) for illustrative purposes only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods shown herein may be employed without departing from the principles described herein.

[0037] Embodiments relate to spatial audio processing using medial and lateral components that are orthogonal to each other spectrally. For example, an audio processing system generates a super-medial component or a super-lateral component, where the super-medial component isolates the portion of the medial component that corresponds only to the spectral energy present at the center of the sound field, and the super-lateral component isolates the portion of the lateral component that corresponds only to the spectral energy not present at the center of the sound field. The super-medial component includes removing the spectral energy of the lateral component from the spectral energy of the medial component, and the super-lateral component includes removing the spectral energy of the medial component from the spectral energy of the lateral component. The audio processing system can also generate a residual medial component and a residual lateral component, where the residual medial component corresponds to the spectral energy of the medial component with the super-medial component removed (e.g., by subtracting the spectral energy of the super-medial component from the spectral energy of the medial component), and the residual lateral component corresponds to the spectral energy of the lateral component with the super-medial component removed (e.g., by subtracting the spectral energy of the super-lateral component from the spectral energy of the lateral component). By isolating these orthogonal components and performing various types of audio processing using these components, the audio processing system can provide targeted audio content enhancement. The super-medial component represents the non-spatial (i.e., medial) spectral energy at the center of the sound field. For example, the non-spatial spectral energy at the center of the sound field can include dialogue in a movie or the main vocal content in music. Applying signal processing operations to the super-medial enables adjusting such audio content without changing the spectral energy present elsewhere in the sound field. For example, in some embodiments, a sound vocal can be partially and / or completely removed by applying a filter that reduces the spectral energy within the typical human vocal range to the super-medial component. In other embodiments, targeted vocal enhancement or effects can be applied to the vocal content by a filter that increases the energy within the typical human vocal range (e.g., via compression, reverb, and / or other audio processing techniques). The residual medial component represents the non-spatial spectral energy not at the center of the sound field. Applying signal processing techniques to the residual medial allows for similar transformations to occur orthogonally from the other components. For example, in some embodiments, to provide a spatial widening effect to the audio content with minimal change in the overall perceived gain and minimal loss of vocal presence, targeted spectral energy in the residual medial component can be partially and / or completely removed while increasing the spectral energy in the residual lateral component.

[0038] Example audio processing system

[0039] Figure 1is a block diagram of an audio processing system 100 according to one or more embodiments. The audio processing system 100 is a circuit device that processes an input audio signal to generate a spatially enhanced output audio signal. The input audio signal includes a left input channel 103 and a right input channel 105, and the output audio signal includes a left output channel 121 and a right output channel 123. The audio processing system 100 includes an L / R to M / S converter module 107, an orthogonal component generator module 113, an orthogonal component processor module 117, an M / S to L / R converter module 119, and a crosstalk processor module 141. In some embodiments, the audio processing system 100 includes a subset of the above components and / or additional components in addition to the above components. In some embodiments, the audio processing system 100 processes the input audio signal in an order different from Figure 1 that shown. For example, the audio processing system 100 may process the input audio using crosstalk processing before processing using the orthogonal component generator module 113 and the orthogonal component processor module 117.

[0040] The L / R to M / S converter module 107 receives the left input channel 103 and the right input channel 105, and generates an intermediate component 109 (e.g., a non-spatial component) and a side component 111 (e.g., a spatial component) from the input channels 103 and 105. In some embodiments, the intermediate component 109 is generated based on the sum of the left input channel 103 and the right input channel 105, and the side component 111 is generated based on the difference between the left input channel 103 and the right input channel 105. In some embodiments, a plurality of intermediate components and side components are generated from a multi-channel input audio signal (e.g., surround sound). Other L / R to M / S type transforms may be used to generate the intermediate component 109 and the side component 111.

[0041] The orthogonal component generator module 113 processes the intermediate component 109 and the side component 111 to generate at least one of the following: a super-intermediate component M1, a super-side component S1, a residual intermediate component M2, and a residual side component S2. The super-intermediate component M1 is the intermediate component 109 with the side component 111 removed. The super-side component S1 is the spectral energy of the side component 111 with the spectral energy of the intermediate component 109 removed. The residual intermediate component M2 is the spectral energy of the intermediate component 109 with the spectral energy of the super-intermediate component M1 removed. The residual side component S2 is the spectral energy of the side component 111 with the spectral energy of the super-side component S1 removed. In some embodiments, the audio processing system 100 generates the left output channel 121 and the right output channel 123 by processing at least one of the super-intermediate component M1, the super-side component S1, the residual intermediate component M2, and the residual side component S2. The orthogonal component generator module 113 is described further with respect to Figures 2A - 2C be described further.

[0042] The orthogonal component processor module 117 processes one or more of the super-middle component M1, the super-side component S1, the residual middle component M2, and / or the residual side component S2. The processing of the components M1, M2, S1, and S2 may include various types of filtering, such as spatial cue processing (e.g., translation based on amplitude or delay, binaural processing, etc.), dynamic range processing, machine learning-based processing, gain application, reverberation, adding audio effects, or other types of processing. In some embodiments, the orthogonal component processor module 117 uses the super-middle component M1, the super-side component S1, the residual middle component M2, and / or the residual side component S2 to perform sub-band spatial processing and / or crosstalk compensation processing to generate a processed middle component 131 and a processed side component 139. Sub-band spatial processing is processing performed on the frequency sub-bands of the middle and side components of an audio signal for spatially enhancing the audio signal. Crosstalk compensation processing is processing performed on an audio signal for adjusting spectral artifacts caused by crosstalk processing, such as crosstalk compensation for speakers or crosstalk simulation for headphones. The orthogonal component processor module 117 is further described with respect to Figure 3 further described.

[0043] The M / S to L / R converter module 119 receives the processed middle component 131 and the processed side component 139 and generates a processed left component 151 and a processed right component 159. In some embodiments, the processed left component 151 is generated based on the sum of the processed middle component 131 and the processed side component 139, and the processed right component 159 is generated based on the difference between the processed middle component 131 and the processed side component 139. Other M / S to L / R transformation types may be used to generate the processed left component 151 and the processed right component 159.

[0044] The crosstalk processor module 141 receives the processed left component 151 and the processed right component 159 and performs crosstalk processing thereon. Crosstalk processing includes, for example, crosstalk simulation or crosstalk cancellation. Crosstalk simulation is processing performed on an audio signal (e.g., output via headphones) for simulating the effect of speakers. Crosstalk cancellation is processing performed on an audio signal configured to be output via speakers for canceling crosstalk caused by the speakers. The crosstalk processor module 141 outputs a left output channel 121 and a right output channel 123.

[0045] Example orthogonal component generators

[0046] Figures 2A - 2C are block diagrams of orthogonal component generator modules 213, 223, and 243 according to one or more embodiments. The orthogonal component generator modules 213, 223, and 243 are examples of the orthogonal component generator module 113.

[0047] Reference Figure 2A , the orthogonal component generator module 213 includes a subtraction unit 205, a subtraction unit 209, a subtraction unit 215, and a subtraction unit 219. As described above, the orthogonal component generator module 113 receives an intermediate component 109 and a side component 111, and outputs one or more of a super-intermediate component M1, a super-side component S1, a residual intermediate component M2, and a residual side component S2.

[0048] The subtraction unit 205 removes the spectral energy of the side component 111 from the spectral energy of the intermediate component 109 to generate a super-intermediate component M1. For example, the subtraction unit 205 subtracts the magnitude of the side component 111 in the frequency domain from the magnitude of the intermediate component 109 in the frequency domain, without considering the phase, to generate the super-intermediate component M1. The frequency domain subtraction can be performed on a time domain signal using a Fourier transform to generate a frequency domain signal, followed by subtraction of the frequency domain signals. In other examples, the frequency domain subtraction can be performed in other ways, such as using a wavelet transform instead of a Fourier transform. The subtraction unit 209 generates a residual intermediate component M2 by removing the spectral energy of the super-intermediate component M1 from the spectral energy of the intermediate component 109. For example, the subtraction unit 209 subtracts the magnitude of the super-intermediate component M1 in the frequency domain from the magnitude of the intermediate component 109 in the frequency domain, without considering the phase, to generate the residual intermediate component M2. While subtracting the side from the intermediate in the time domain results in the original right channel of the signal, the above operations in the frequency domain isolate and distinguish between: a portion of the spectral energy of the intermediate component that is different from the spectral energy of the side component (referred to as M1, or super-intermediate), and a portion of the spectral energy of the intermediate component that is the same as the spectral energy of the side component (referred to as M2, or residual intermediate).

[0049] In some embodiments, when subtracting the spectral energy of the side component 111 from the spectral energy of the intermediate component 109 results in a negative value for the super-intermediate component M1 (e.g., for one or more intervals in the frequency domain), additional processing may be used. In some embodiments, when subtracting the spectral energy of the side component 111 from the spectral energy of the intermediate component 109 results in a negative value, the super-intermediate component M1 is clamped at 0. In some embodiments, the super-intermediate component M1 is wrapped around by using the absolute value of the negative value as the value of the super-intermediate component M1. When subtracting the spectral energy of the side component 111 from the spectral energy of the intermediate component 109 results in M1 being negative, other types of processing may be used. Similar additional processing, such as clamping at 0, wrapping around, or other processing, may be used when the subtraction result for generating the super-side component S1, the residual side component S2, or the residual intermediate component M2 is negative. Clamping the super-intermediate component M1 at 0 when the subtraction results in a negative value will ensure spectral orthogonality between M1 and the two side components. Similarly, clamping the super-side component S1 at 0 when the subtraction results in a negative value will ensure spectral orthogonality between S1 and the two intermediate components. By creating orthogonality between the super-intermediate and side components and their appropriate intermediate / side corresponding components (i.e., the side component for the super-intermediate, the intermediate component for the super-side), the derived residual intermediate M2 and residual side S2 components contain spectral energy that is not orthogonal to their appropriate intermediate / side corresponding components (i.e., that is common with them). That is, when clamping the super-intermediate at 0 and using that M1 component to derive the residual intermediate, a super-intermediate component with spectral energy not common with the side component and a residual intermediate component with spectral energy completely common with the side component are generated. The same relationship applies to the super-side and the residual side when clamping the super-side at 0. When applying frequency domain processing, there is typically a trade-off in the resolution between frequency and timing information. As the frequency resolution increases (i.e., as the FFT window size and the number of frequency intervals increase), the time resolution decreases, and vice versa. The above spectral subtraction occurs on a per-frequency interval basis, so in some cases, such as when removing vocal energy from the super-intermediate component, it is preferable to use a larger FFT window size (e.g., 8192 samples, which results in 4096 frequency intervals for a given real-valued input signal). Other cases may require higher time resolution and thus lower overall latency and lower frequency resolution (e.g., a 512-sample FFT window size, which results in 256 frequency intervals for a given real-valued input signal). In the latter case, the low frequency resolution of the intermediate and side can produce audible spectral artifacts when subtracted from each other to derive the super-intermediate M1 and super-side S1 components, because the spectral energy of each frequency interval is an average representation of the energy over too large a frequency range).In this case, taking the absolute value of the difference between the middle and the side when deriving the super-middle M1 or the super-side S1 can help mitigate perceptual artifacts by allowing each frequency bin to diverge from true orthogonality in the components. Instead of or in addition to returning 0, coefficients can be applied to the subtracted value to scale the value between 0 and 1, thus providing a method for interpolating between the following extremes: at one extreme (i.e., value of 1), complete orthogonality of the super and residual middle / side components; and at the other extreme (i.e., value of 0), the super-middle M1 and super-side S1 being the same as their corresponding original middle and side components.

[0050] The subtraction unit 215 removes the spectral energy of the middle component 109 in the frequency domain from the spectral energy of the side component 111 in the frequency domain, regardless of phase, to generate the super-side component S1. For example, the subtraction unit 215 subtracts the magnitude of the middle component 109 in the frequency domain from the magnitude of the side component 111 in the frequency domain, regardless of phase, to generate the super-side component S1. The subtraction unit 219 removes the spectral energy of the super-side component S1 from the spectral energy of the side component 111 to generate the residual side component S2. For example, the subtraction unit 219 subtracts the magnitude of the super-side component S1 in the frequency domain from the magnitude of the side component 111 in the frequency domain, regardless of phase, to generate the residual side component S2.

[0051] In Figure 2B the orthogonal component generator module 223 is similar to the orthogonal component generator module 213 in that it receives the middle component 109 and the side component 111 and generates the super-middle component M1, the residual middle component M2, the super-side component S1, and the residual side component S2. The orthogonal component generator module 223 differs from the orthogonal component generator module 213 in that the super-middle component M1 and the super-side component S1 are generated in the frequency domain and then these components are converted back to the time domain to generate the residual middle component M2 and the residual side component S2. The orthogonal component generator module 223 includes a forward FFT unit 220, a bandpass unit 222, a subtraction unit 224, a super-middle processor 225, an inverse FFT unit 226, a time delay unit 228, a subtraction unit 230, a forward FFT unit 232, a bandpass unit 234, a subtraction unit 236, a super-side processor 237, an inverse FFT unit 240, a time delay unit 242, and a subtraction unit 244.

[0052] The forward fast Fourier transform (FFT) unit 220 applies a forward FFT to the intermediate component 109 to transform the intermediate component 109 into the frequency domain. The transformed intermediate component 109 in the frequency domain includes magnitude and phase. The bandpass unit 222 applies a bandpass filter to the frequency domain intermediate component 109, where the bandpass filter specifies the frequencies in the super-intermediate component M1. For example, to isolate the typical human vocal range, the bandpass filter can specify frequencies between 300 and 8000 Hz. In another example, to remove audio content associated with the typical human vocal range, the bandpass filter can retain the lower frequencies (e.g., generated by a bass guitar or drums) and higher frequencies (e.g., generated by cymbals) in the super-intermediate component M1. In other embodiments, in addition to and / or instead of the bandpass filter applied by the bandpass unit 222, the quadrature component generator module 223 applies various other filters to the frequency domain intermediate component 109. In some embodiments, the quadrature component generator module 223 does not include the bandpass unit 222 and does not apply any filters to the frequency domain intermediate component 109. In the frequency domain, the subtraction unit 224 subtracts the side component 111 from the filtered intermediate component to generate the super-intermediate component M1. In other embodiments, in addition to and / or instead of the later processing applied to the super-intermediate component M1 performed by the quadrature component processor module (e.g., Figure 3 the quadrature component processor module), the quadrature component generator module 223 applies various audio enhancements to the frequency domain super-intermediate component M1. The super-intermediate processor 225 performs processing on the super-intermediate component M1 in the frequency domain before the super-intermediate component M1 is transformed into the time domain. This processing can include sub-band spatial processing and / or crosstalk compensation processing. In some embodiments, instead of and / or in addition to the processing that can be performed by the quadrature component processor module 117, the super-intermediate processor 225 performs processing on the super-intermediate component M1. The inverse FFT unit 226 applies an inverse FFT to the super-intermediate component M1 to transform the super-intermediate component M1 back into the time domain. The super-intermediate component M1 in the frequency domain includes the magnitude of M1 and the phase of the intermediate component 109, and the inverse FFT unit 226 transforms it into the time domain. The time delay unit 228 applies a time delay to the intermediate component 109 such that the intermediate component 109 and the super-intermediate component M1 arrive at the subtraction unit 230 simultaneously. The subtraction unit 230 subtracts the super-intermediate component M1 in the time domain from the time-delayed intermediate component 109 in the time domain to generate the residual intermediate component M2. In this example, the spectral energy of the super-intermediate component M1 is removed from the spectral energy of the intermediate component 109 using processing in the time domain.

[0053] The forward FFT unit 232 applies a forward FFT to the side component 111 to transform the side component 111 into the frequency domain. The transformed side component 111 in the frequency domain includes magnitude and phase. The band-pass unit 234 applies a band-pass filter to the frequency-domain side component 111. The band-pass filter specifies the frequencies in the super side component S1. In other embodiments, in addition to and / or instead of the band-pass filter, the quadrature component generator module 223 applies various other filters to the frequency-domain side component 111. In the frequency domain, the subtraction unit 236 subtracts the intermediate component 109 from the filtered side component 111 to generate the super side component S1. In other embodiments, in addition to and / or instead of the later processing applied to the super side component S1 performed by the quadrature component processor (e.g., Figure 3 the quadrature component processor module), the quadrature component generator module 223 applies various audio enhancements to the frequency-domain super side component S1. The super side processor 237 performs processing on the super side component S1 in the frequency domain before the super side component S1 is transformed into the time domain. This processing may include sub-band spatial processing and / or crosstalk compensation processing. In some embodiments, instead of and / or in addition to the processing that may be performed by the quadrature component processor module 117, the super side processor 237 performs processing on the super side component S1. The inverse FFT unit 240 applies an inverse FFT to the super side component S1 in the frequency domain to generate the super side component S1 in the time domain. The super side component S1 in the frequency domain includes the magnitude of S1 and the phase of the side component 111, and the inverse FFT unit 226 transforms it into the time domain. The time delay unit 242 applies a time delay to the side component 111 such that the side component 111 arrives at the subtraction unit 244 simultaneously with the super side component S1. The subtraction unit 244 then subtracts the super side component S1 in the time domain from the time-delayed side component 111 in the time domain to generate the residual side component S2. In this example, the spectral energy of the super side component S1 is removed from the spectral energy of the side component 111 using processing in the time domain.

[0054] In some embodiments, the super intermediate processor 225 and the super side processor 237 may be omitted if the processing performed by these components is performed by the quadrature component processor module 117.

[0055] In Figure 2CIn [the context], the orthogonal component generator module 245 is similar to the orthogonal component generator module 223 in that it receives the intermediate component 109 and the side component 111 and generates the super-intermediate component M1, the residual intermediate component M2, the super-side component S1, and the residual side component S2. The difference is that the orthogonal component generator module 245 generates each of the components M1, M2, S1, and S2 in the frequency domain and then converts these components to the time domain. The orthogonal component generator module 245 includes a forward FFT unit 247, a band-pass unit 249, a subtraction unit 251, a super-intermediate processor 252, a subtraction unit 253, a residual intermediate processor 254, an inverse FFT unit 255, an inverse FFT unit 257, a forward FFT unit 261, a band-pass unit 263, a subtraction unit 265, a super-side processor 266, a subtraction unit 267, a residual side processor 268, an inverse FFT unit 269, and an inverse FFT unit 271.

[0056] The forward FFT unit 247 applies a forward FFT to the intermediate component 109 to transform the intermediate component 109 into the frequency domain. The transformed intermediate component 109 in the frequency domain includes magnitude and phase. The forward FFT unit 261 applies a forward FFT to the side component 111 to transform the side component 111 into the frequency domain. The transformed side component 111 in the frequency domain includes magnitude and phase. The bandpass unit 249 applies a bandpass filter to the frequency domain intermediate component 109, and the bandpass filter specifies the frequency of the super-intermediate component M1. In some embodiments, in addition to and / or instead of the bandpass filter, the quadrature component generator module 245 applies various other filters to the frequency domain intermediate component 109. The subtraction unit 251 subtracts the frequency domain side component 111 from the frequency domain intermediate component 109 to generate the super-intermediate component M1 in the frequency domain. The super-intermediate processor 252 performs processing on the super-intermediate component M1 in the frequency domain before it is transformed into the time domain. In some embodiments, the super-intermediate processor 252 performs sub-band spatial processing and / or crosstalk compensation processing. In some embodiments, instead of and / or in addition to the processing that can be performed by the quadrature component processor module 117, the super-intermediate processor 252 performs processing on the super-intermediate component M1. The inverse FFT unit 257 applies an inverse FFT to the super-intermediate component M1 to transform it back into the time domain. The super-intermediate component M1 in the frequency domain includes the magnitude of M1 and the phase of the intermediate component 109, and the inverse FFT unit 257 transforms it into the time domain. The subtraction unit 253 subtracts the super-intermediate component M1 from the intermediate component 109 in the frequency domain to generate the residual intermediate component M2. The residual intermediate processor 254 performs processing on the residual intermediate component M2 in the frequency domain before it is transformed into the time domain. In some embodiments, the residual intermediate processor 254 performs sub-band spatial processing and / or crosstalk compensation processing on the residual intermediate component M2. In some embodiments, instead of and / or in addition to the processing that can be performed by the quadrature component processor module 117, the residual intermediate processor 254 performs processing on the residual intermediate component M2. The inverse FFT unit 255 applies an inverse FFT to transform the residual intermediate component M2 into the time domain. The residual intermediate component M2 in the frequency domain includes the magnitude of M2 and the phase of the intermediate component 109, and the inverse FFT unit 255 transforms it into the time domain.

[0057] The band - pass unit 263 applies a band - pass filter to the frequency - domain side component 111. The band - pass filter specifies the frequencies in the super - side component S1. In other embodiments, in addition to and / or instead of the band - pass filter, the quadrature - component generator module 245 applies various other filters to the frequency - domain side component 111. In the frequency domain, the subtraction unit 265 subtracts the intermediate component 109 from the filtered side component 111 to generate the super - side component S1. The super - side processor 266 performs processing on the super - side component S1 in the frequency domain before it is converted to the time domain. In some embodiments, the super - side processor 266 performs sub - band spatial processing and / or crosstalk compensation processing on the super - side component S1. In some embodiments, instead of and / or in addition to the processing that can be performed by the quadrature - component processor module 117, the super - side processor 266 performs processing on the super - side component S1. The inverse FFT unit 271 applies an inverse FFT to convert the super - side component S1 back to the time domain. The super - side component S1 in the frequency domain includes the magnitude of S1 and the phase of the side component 111, and the inverse FFT unit 271 converts it to the time domain. The subtraction unit 267 subtracts the super - side component S1 from the side component 111 in the frequency domain to generate the residual side component S2. The residual - side processor 268 performs processing on the residual - side component S2 in the frequency domain before it is converted to the time domain. In some embodiments, the residual - side processor 268 performs sub - band spatial processing and / or crosstalk compensation processing on the residual - side component S2. In some embodiments, instead of and / or in addition to the processing that can be performed by the quadrature - component processor module 117, the residual - side processor 268 performs processing on the residual - side component S2. The inverse FFT unit 269 applies an inverse FFT to the residual - side component S2 to convert it to the time domain. The residual - side component S2 in the frequency domain includes the magnitude of S2 and the phase of the side component 111, and the inverse FFT unit 269 converts it to the time domain.

[0058] In some embodiments, if the processing performed by the super - intermediate processor 252, the super - side processor 266, the residual - intermediate processor 254, or the residual - side processor 268 is performed by the quadrature - component processor module 117, these components can be omitted.

[0059] Example Quadrature - Component Processor

[0060] Figure 3Is a block diagram of an orthogonal component processor module 317 according to one or more embodiments. The orthogonal component processor module 317 is an example of the orthogonal component processor module 117. The orthogonal component processor module 317 may include a subband spatial processing and / or crosstalk compensation processing unit 320, an addition unit 325, and an addition unit 330. The orthogonal component processor module 317 performs subband spatial processing and / or crosstalk compensation processing on at least one of the super intermediate component M1, the residual intermediate component M2, the super side component S1, and the residual side component S2. As a result of the subband spatial processing and / or crosstalk compensation processing 320, the orthogonal component processor module 317 outputs at least one of the processed M1, the processed M2, the processed S1, and the processed S2. The addition unit 325 adds the processed M1 and the processed M2 to generate a processed intermediate component 131, and the addition unit 330 adds the processed S1 and the processed S2 to generate a processed side component 139.

[0061] In some embodiments, the orthogonal component processor module 317 performs subband spatial processing and / or crosstalk compensation processing 320 on at least one of the super intermediate component M1, the residual intermediate component M2, the super side component S1, and the residual side component S2 in the frequency domain to generate a processed intermediate component 131 and a processed side component 139 in the frequency domain. The orthogonal component generator module 113 may provide the components M1, M2, S1, or S2 in the frequency domain to the orthogonal component processor, where an inverse FFT is performed. After generating the processed intermediate component 131 and the processed side component 139, the orthogonal component processor module 317 may perform an inverse FFT on the processed intermediate component 131 and the processed side component 139 to convert these components back to the time domain. In some embodiments, the orthogonal component processor module 317 performs an inverse FFT on the processed M1, the processed M2, the processed S1, and the processed S1 to generate a processed intermediate component 131 and a processed side component 139 in the time domain.

[0062] An example of the orthogonal component processor module 317 is in Figure 4 and Figure 5is shown. In some embodiments, the quadrature component processor module 317 performs sub-band spatial processing and crosstalk compensation processing. The processing performed by the quadrature component processor module 317 is not limited to sub-band spatial processing or crosstalk compensation processing. Any type of spatial processing using the mid / side space can be performed by the quadrature component processor module 317, such as by using super-mid components instead of mid components or super-side components instead of side components. Some other types of processing can include gain application, amplitude- or delay-based translation, binaural processing, reverberation, dynamic range processing (such as compression and limiting), and other linear or non-linear audio processing techniques and effects, ranging from chorus or flanging to methods such as machine learning-based vocal or instrumental style transfer, transformation, or resynthesis, etc.

[0063] Exemplary sub-band spatial processor

[0064] Figure 4 is a block diagram of a sub-band spatial processor module 410 according to one or more embodiments. The sub-band spatial processor module 410 is an example of the quadrature component processor module 317. The sub-band spatial processor module 410 includes mid EQ filters 404(1), mid EQ filters 404(2), mid EQ filters 404(3), mid EQ filters 404(4), side EQ filters 406(1), side EQ filters 406(2), side EQ filters 406(3), and side EQ filters 406(4). In some embodiments, the sub-band spatial processor module 410 includes other components in addition to and / or instead of the components described herein.

[0065] The sub-band spatial processor module 410 receives the non-spatial component Y m and the spatial component Y s and performs gain adjustment on the sub-bands of one or more of these components to provide spatial enhancement. The non-spatial component Y m can be the super-mid component M1 or the residual mid component M2. The spatial component Y s can be the super-side component S1 or the residual side component S2.

[0066] The sub-band spatial processor module 410 receives the non-spatial component Y m and applies the mid EQ filters 404(1) to 404(4) to different sub-bands of Y m to generate an enhanced non-spatial component E m . The sub-band spatial processor module 410 also receives the spatial component Y s and applies the side EQ filters 406(1) to 406(4) to different sub-bands of Y s to generate an enhanced spatial component E s。The sub-band filters can include various combinations of peak filters, notch filters, low-pass filters, high-pass filters, low-shelf filters, high-shelf filters, band-pass filters, band-stop filters, and / or all-pass filters. The sub-band filters can also apply gain to the corresponding sub-bands. More specifically, the sub-band spatial processor module 410 includes sub-band filters for each of the n frequency sub-bands of the non-spatial component Y m and sub-band filters for each of the n sub-bands of the spatial component Y s For example, for n = 4 sub-bands, the sub-band spatial processor module 410 includes a series of sub-band filters for the non-spatial component Y m , including an intermediate equalization (EQ) filter 404(1) for sub-band (1), an intermediate EQ filter 404(2) for sub-band (2), an intermediate EQ filter 404(3) for sub-band (3), and an intermediate EQ filter 404(4) for sub-band (4). Each intermediate EQ filter 404 applies a filter to the frequency sub-band portion of the non-spatial component Y m to generate an enhanced non-spatial component E m .

[0067] The sub-band spatial processor module 410 also includes a series of sub-band filters for the frequency sub-bands of the spatial component Y s , including a side equalization (EQ) filter 406(1) for sub-band (1), a side EQ filter 406(2) for sub-band (2), a side EQ filter 406(3) for sub-band (3), and a side EQ filter 406(4) for sub-band (4). Each side EQ filter 406 applies a filter to the frequency sub-band portion of the spatial component Y s to generate an enhanced spatial component E s .

[0068] Each of the n frequency sub-bands of the non-spatial component Y m and the spatial component Y s can correspond to a certain range of frequencies. For example, frequency sub-band (1) can correspond to 0 to 300 Hz, frequency sub-band (2) can correspond to 300 to 510 Hz, frequency sub-band (3) can correspond to 510 to 2700 Hz, and frequency sub-band (4) can correspond to 2700 Hz to the Nyquist frequency. In some embodiments, the n frequency sub-bands are a combined set of critical bands. The critical bands can be determined using a corpus of audio samples from multiple music genres. The long-term average energy ratio of the mid-to-side components above 24 Bark scale critical bands is determined from the samples. Then consecutive frequency bands with similar long-term average ratios are combined together to form the set of critical bands. The range of the frequency sub-bands and the number of frequency sub-bands can be adjustable.

[0069] In some embodiments, the subband spatial processor module 410 processes the residual intermediate component M2 as a non-spatial component Y m , and uses one of the side component, the super side component S1, or the residual side component S2 as the spatial component Y s .

[0070] In some embodiments, the subband spatial processor module 410 processes one or more of the super intermediate component M1, the super side component S1, the residual intermediate component M2, and the residual side component S2. The filters applied to the subbands of each of these components can be different. Each of the super intermediate component M1 and the residual intermediate component M2 can be processed as discussed for the non-spatial component Y m . Each of the super side component S1 and the residual side component S2 can be processed as discussed for the spatial component Y s .

[0071] Example crosstalk compensation processor

[0072] Figure 5 is a block diagram of a crosstalk compensation processor module 510 according to one or more embodiments. The crosstalk compensation processor module 510 is an example of the orthogonal component processor module 317. The crosstalk compensation processor module 510 includes an intermediate component processor 520 and a side component processor 530. The crosstalk compensation processor module 510 receives the non-spatial component Y m and the spatial component Y s , and applies filters to one or more of these components to compensate for spectral defects caused by (e.g., subsequent or previous) crosstalk processing. The non-spatial component Y m can be the super intermediate component M1 or the residual intermediate component M2. The spatial component Y s can be the super side component S1 or the residual side component S2.

[0073] The crosstalk compensation processor module 510 receives the non-spatial component Y m and the intermediate component processor 520 applies a set of filters to generate an enhanced non-spatial crosstalk compensation component Z m . The crosstalk compensation processor module 510 also receives the spatial subband component Y s , and applies a set of filters in the side component processor 530 to generate an enhanced spatial subband component E s . The intermediate component processor 520 includes a plurality of filters 540, such as m intermediate filters 540(a), 540(b) to 540(m). Here, each of the m intermediate filters 540 processes one of the m frequency bands of the non-spatial component X m . The intermediate component processor 520 accordingly processes the non-spatial component X mto generate an intermediate crosstalk compensation channel Z m 。In some embodiments, the intermediate filter 540 is configured using a non-spatial X m frequency response map and crosstalk processing is performed through simulation. Additionally, by analyzing the frequency response map, any spectral defects that occur as artifacts of the crosstalk processing can be estimated, such as peaks or valleys in the frequency response map that exceed a predetermined threshold (e.g., 10 dB). These artifacts are mainly due to the addition of the delayed and possibly inverted opposite-side signal to its corresponding same-side signal during crosstalk processing, effectively introducing a comb-filter-like frequency response into the final rendering result. The intermediate crosstalk compensation channel Z m can be generated by the intermediate component processor 520 to compensate for the estimated peaks or valleys, where each of the m frequency bands corresponds to a peak or valley. Specifically, based on the specific delays, filtering frequencies, and gains applied during crosstalk processing, the peaks and valleys shift up and down in the frequency response, resulting in variable amplification and / or attenuation of the energy in specific regions of the spectrum. Each of the intermediate filters 540 can be configured to adjust for one or more of the peaks and valleys.

[0074] The side component processor 530 includes a plurality of filters 550, such as m side filters 550(a), 550(b) to 550(m). The side component processor 530 generates a side crosstalk compensation channel Z s by processing the spatial component X s 。In some embodiments, a frequency response map of the spatial X with crosstalk processing can be obtained through simulation. By analyzing the frequency response map, any spectral defects that occur as artifacts of the crosstalk processing can be estimated, such as peaks or valleys in the frequency response map that exceed a predetermined threshold (e.g., 10 dB). The side crosstalk compensation channel Z s can be generated by the side component processor 530 to compensate for the estimated peaks or valleys. Specifically, based on the specific delays, filtering frequencies, and gains applied during crosstalk processing, the peaks and valleys shift up and down in the frequency response, resulting in variable amplification and / or attenuation of the energy in specific regions of the spectrum. Each of the side filters 550 can be configured to adjust for one or more of the peaks and valleys. In some embodiments, the intermediate component processor 520 and the side component processor 530 can include different numbers of filters. s In some embodiments, the intermediate filter 540 and the side filter 550 can include biquadratic filters having a transfer function defined by Equation 1:

[0075]

[0076]

[0077] where z is a complex variable, and a0, a1, a2, b0, b1, and b2 are digital filter coefficients. One way to implement such a filter is the direct form I topology defined by Equation 2:

[0078]

[0079]

[0080] where X is the input vector and Y is the output. Other topologies can be used, depending on their maximum word length and saturation behavior. Then, a second-order filter with real-valued inputs and outputs can be implemented using a biquadratic. To design a discrete-time filter, a continuous-time filter is designed and then transformed to discrete-time via a bilinear transformation. Additionally, frequency warping can be used to compensate for the final offset of the center frequency and bandwidth.

[0081] For example, a peak filter can have an S-plane transfer function defined by Equation 3:

[0082]

[0083] where s is a complex variable, A is the amplitude of the peak, and Q is the filter "quality", and the digital filter coefficients are defined by:

[0084] b0 = 1 + αA

[0085] b1 = -2 * cos(ω0)

[0086] b2 = 1 - αA

[0087]

[0088] a1 = -2cos(ω0)

[0089]

[0090] where ω0 is the center frequency of the filter in radians and In addition, the filter quality Q can be defined by Equation 4:

[0091]

[0092] where Δf is the bandwidth and f c is the center frequency. The intermediate filter 540 is shown in series, and the side filter 550 is shown in series. In some embodiments, the intermediate filter 540 is applied in parallel to the intermediate component X m , and the side filter 540 is applied in parallel to the side component X s .

[0093] In some embodiments, the crosstalk compensation processor module 510 processes each of the super intermediate component M1, the super side component S1, the residual intermediate component M2, and the residual side component S2. The filters applied to each of these components can be different.

[0094] Example crosstalk processor

[0095] Figure 6 is a block diagram of a crosstalk simulation processor module 600 according to one or more embodiments. As described with respect to Figure 1 In some embodiments, the audio processing system 100 includes a crosstalk processor module 141 that applies crosstalk processing to the processed left component 151 and the processed right component 159. Crosstalk processing includes, for example, crosstalk simulation and crosstalk cancellation. In some embodiments, the crosstalk processor module 141 includes a crosstalk simulation processor module 600. The crosstalk simulation processor module 600 generates a cross-side sound component for output to the stereo headphones, thereby providing a speaker-like listening experience on the headphones. The left input channel X L can be the processed left component 151, and the right input channel X R can be the processed right component 159. In some embodiments, crosstalk simulation can be performed before orthogonal component processing.

[0096] The crosstalk simulation processor module 600 includes a left head shadow low-pass filter 602, a left head shadow high-pass filter 624, a left crosstalk delay 604, and a left head shadow gain 610 to process the left input channel X L . The crosstalk simulation processor module 600 further includes a right head shadow low-pass filter 606, a right head shadow high-pass filter 626, a right crosstalk delay 608, and a right head shadow gain 612 to process the right input channel X R . The left head shadow low-pass filter 602 and the left head shadow high-pass filter 624 apply a modulation to the left input channel X L that simulates the frequency response of the signal after passing through the listener's head. The output of the left head shadow high-pass filter 624 is provided to the left crosstalk delay 604, and the left crosstalk delay 604 applies a time delay. The time delay represents the transmural distance that the cross-side sound component travels relative to the same-side sound component. The left head shadow gain 610 applies a gain to the output of the left crosstalk delay 604 to generate the right-to-left analog channel W L .

[0097] Similarly, for the right input channel X R , the right head shadow low-pass filter 606 and the right head shadow high-pass filter 626 apply a modulation to the right input channel X RApply modulation that simulates the frequency response of a listener's head. The output of the right head-related high-pass filter 626 is provided to the right crosstalk delay 608, which applies a time delay. The right head-related gain 612 applies a gain to the output of the right crosstalk delay 608 to generate the right crosstalk analog channel W R .

[0098] Applying the head-related low-pass filter, the head-related high-pass filter, the crosstalk delay, and the head-related gain to each of the left and right channels can be performed in a different order.

[0099] Figure 7 is a block diagram of a crosstalk cancellation processor module 700 according to one or more embodiments. The crosstalk processor module 141 may include the crosstalk cancellation processor module 700. The crosstalk cancellation processor module 700 receives the left input channel X L and the right input channel X R , and performs crosstalk cancellation on the channels X L , X R to generate the left output channel O L and the right output channel O R . The left input channel X L may be the processed left component 151, and the right input channel X R may be the processed right component 159. In some embodiments, crosstalk cancellation may be performed before quadrature component processing.

[0100] The crosstalk cancellation processor module 700 includes an in-band / out-of-band divider 710, inverters 720 and 722, opposite-side estimators 730 and 740, combiners 750 and 752, and an in-band / out-of-band combiner 760. These components operate together to divide the input channels T L , T R into in-band components and out-of-band components, and perform crosstalk cancellation on the in-band components to generate the output channels O L , O R .

[0101] By dividing the input audio signal T into different frequency band components and by performing crosstalk cancellation on selective components (e.g., in-band components), crosstalk cancellation can be performed for a specific frequency band while avoiding degradation in other frequency bands. If crosstalk cancellation is performed without dividing the input audio signal T into different frequency bands, the audio signal after such crosstalk cancellation exhibits significant attenuation or amplification of non-spatial and spatial components at low frequencies (e.g., below 350 Hz), high frequencies (e.g., above 12000 Hz), or both. By performing crosstalk cancellation in the in-band (e.g., between 250 Hz and 14000 Hz) where the vast majority of the influential spatial cues are located, a balanced overall energy can be retained across the entire spectrum of the mix, especially in the non-spatial components.

[0102] The in-out divider 710 separates the input channels T L , T R into an in-band channel T L,In , T R,In and an out-of-band channel T L,Out , T R,Out . Specifically, the in-out divider 710 divides the left enhanced compensation channel T L into a left in-band channel T L,In and a left out-of-band channel T L,Out . Similarly, the in-out divider 710 separates the right enhanced compensation channel T R into a right in-band channel T R,In and a right out-of-band channel T R,Out . Each in-band channel may contain a portion of the corresponding input channel corresponding to a certain frequency range, which frequency range includes, for example, 250 Hz to 14 kHz. This frequency band range may be adjustable, for example, according to speaker parameters.

[0103] The inverter 720 and the crosstalk estimator 730 operate together to generate a left crosstalk cancellation component S L to compensate for the crosstalk component due to the left in-band channel T L,In . Similarly, the inverter 722 and the crosstalk estimator 740 operate together to generate a right crosstalk cancellation component S R to compensate for the crosstalk component due to the right in-band channel T R,In .

[0104] In one method, the inverter 720 receives the in-band channel T L,In and inverts the polarity of the received in-band channel T L,In to generate an inverted in-band channel T L,In '. The crosstalk estimator 730 receives the inverted in-band channel T L,In ' and extracts, through filtering, a portion of the inverted in-band channel T L,In ' corresponding to the crosstalk component. Since the filtering is performed on the inverted in-band channel T L,In ', the portion extracted by the crosstalk estimator 730 becomes the reciprocal of a portion of the in-band channel T L,In attributable to the crosstalk component. Thus, the portion extracted by the crosstalk estimator 730 becomes the left crosstalk cancellation component S L , which can be added to the corresponding in-band channel T R,In to reduce the crosstalk component due to the in-band channel T L,In . In some embodiments, the inverter 720 and the crosstalk estimator 730 are implemented in a different order.

[0105] The inverter 722 and the contralateral estimator 740 perform similar operations on the in-band channel T R,In to generate the right contralateral cancellation component S R . Therefore, for the sake of brevity, its detailed description is omitted here.

[0106] In one example implementation, the contralateral estimator 730 includes a filter 732, an amplifier 734, and a delay unit 736. The filter 732 receives the inverted input channel T L,In ', and extracts a part of the inverted in-band channel T L,In ' corresponding to the contralateral sound component through a filtering function. An example filter implementation is a notch or shelving filter, whose center frequency is selected between 5000 and 10000 Hz, and Q is selected between 0.5 and 1.0. The gain (G dB ) in decibels can be derived from Equation 5:

[0107] G dB =-3.0 - log 1.333 (D) Equation (5)

[0108] where D is the delay amount in samples of the delay units 736 and 646, for example, at a sampling rate of 48 KHz. An alternative implementation is a low-pass filter, where the corner frequency is selected between 5000 and 10000 Hz, and Q is selected between 0.5 and 1.0. Additionally, the amplifier 734 amplifies the extracted part by the corresponding gain coefficient G L,In , and the delay unit 736 delays the amplified output of the amplifier 734 according to the delay function D to generate the left contralateral cancellation component S L . The contralateral estimator 740 includes a filter 742, an amplifier 744, and a delay unit 746, and the delay unit 746 performs a similar operation on the inverted in-band channel T R,In ' to generate the right contralateral cancellation component S R . In one example, the contralateral estimators 730, 740 generate the left contralateral cancellation component S L and the right contralateral cancellation component S R according to the following equations:

[0109] S L =D[G L,In *F[T L,In ’]] Equation (6)

[0110] S R =D[G R,In *F[T R,In ’]] Equation (7)

[0111] where F[] is the filter function and D[] is the delay function.

[0112] The configuration for crosstalk cancellation can be determined by speaker parameters. In one example, the filter center frequency, delay amount, amplifier gain, and filter gain can be determined based on the angle formed between two speakers relative to a listener. In some embodiments, values between speaker angles are used to interpolate other values.

[0113] Combiner 750 combines the right cancellation component S R into the left in-band channel T L,In to generate the left in-band crosstalk channel U L , and combiner 752 combines the left cancellation component S L into the right in-band channel T R,In to generate the right in-band crosstalk channel U R . The in-out combiner 760 combines the left in-band crosstalk channel U L with the out-of-band channel T L,Out to generate the left output channel O L , and combines the right in-band crosstalk channel U R with the out-of-band channel T R,Out to generate the right output channel O R .

[0114] Thus, the left output channel O L includes the right cancellation component S R corresponding to the inversion of a part of the in-band channel T R attributable to the opposite-side sound, and the right output channel O R includes the left cancellation component S L,In corresponding to the inversion of a part of the in-band channel T L attributable to the opposite-side sound. In this configuration, the wavefront of the ipsilateral sound component output by the right speaker according to the right output channel O R arriving at the right ear can cancel the wavefront of the opposite-side sound component output by the left speaker according to the left output channel O L . Similarly, the wavefront of the ipsilateral sound component output by the left speaker according to the left output channel O L arriving at the left ear can cancel the wavefront of the opposite-side sound component output by the right speaker according to the right output channel O R . Thus, the opposite-side sound component can be reduced to enhance spatial detectability.

[0115] Orthogonal component spatial processing

[0116] Figure 8A flowchart of a process for spatial processing using at least one of a supermid, residual mid, superside, or residual side component according to one or more embodiments. The spatial processing may include methods such as gain application, amplitude or delay-based panning, binaural processing, reverberation, dynamic range processing (such as compression and limiting), linear or nonlinear audio processing techniques and effects, chorus effects, flanging effects, machine learning-based vocal or instrumental style transfer, transformation, or resynthesis. The process may be performed to provide spatially enhanced audio to a user's device. The process may include fewer or more steps, and the steps may be performed in a different order.

[0117] An audio processing system (e.g., audio processing system 100) receives 810 an input audio signal (e.g., left input channel 103 and right input channel 105). In some embodiments, the input audio signal may be a multi-channel audio signal including a plurality of left and right channel pairs. For the left and right input channels, each left and right channel pair may be processed as discussed herein.

[0118] The audio processing system generates 820 a non-spatial mid component (e.g., mid component 109) and a spatial side component (e.g., side component 111) from the input audio signal. In some embodiments, an L / R to M / S converter (e.g., L / R to M / S converter module 107) performs the conversion of the input audio signal to the mid and side components.

[0119] The audio processing system generates 830 at least one of a supermid component (e.g., supermid component M1), a superside component (e.g., superside component S1), a residual mid component (e.g., residual mid component M2), and a residual side component (e.g., residual side component S2). The audio processing system may generate at least one and / or all of the components listed above. The supermid component includes removing the spectral energy of the side component from the spectral energy of the mid component. The residual mid component includes removing the spectral energy of the supermid component from the spectral energy of the mid component. The superside component includes removing the spectral energy of the mid component from the spectral energy of the side component. The residual side component includes removing the spectral energy of the superside component from the spectral energy of the side component. The processing for generating M1, M2, S1, or S2 may be performed in the frequency domain or the time domain.

[0120] The audio processing system filters 840 at least one of the supermid component, the residual mid component, the superside component, and the residual side component to enhance the audio signal. The filtering may include spatial cue processing, such as by adjusting the frequency-dependent amplitude or frequency-dependent delay of the supermid component, the residual mid component, the superside component, or the residual side component. Some examples of spatial cue processing include amplitude or delay-based panning or binaural processing.

[0121] Filtering may include dynamic range processing, such as compression or limiting. For example, when a threshold level for compression is exceeded, the super intermediate component, residual intermediate component, super side component, or residual side component may be compressed according to a compression ratio. In another example, when a threshold level for limiting is exceeded, the super intermediate component, residual intermediate component, super side component, or residual side component may be limited to a maximum level.

[0122] Filtering may include machine learning-based alteration of the super intermediate component, residual intermediate component, super side component, or residual side component. Some examples include machine learning-based vocal or instrumental style transfer, transformation, or resynthesis.

[0123] Filtering of the super intermediate component, residual intermediate component, super side component, or residual side component may include gain application, reverb, and other linear or non-linear audio processing techniques and effects (chorus and / or flanging) or other types of processing. In some embodiments, filtering may include filtering for sub-band spatial processing and crosstalk compensation, as discussed in more detail below Figure 9 and discussed in more detail below.

[0124] Filtering may be performed in the frequency domain or the time domain. In some embodiments, the intermediate component and side component are converted from the time domain to the frequency domain, the super and / or residual components are generated in the frequency domain, filtering is performed in the frequency domain, and the filtered components are converted back to the time domain. In other embodiments, the super and / or residual components are converted to the time domain, and filtering is performed on these components in the time domain.

[0125] The audio processing system uses one or more of the filtered super / residual components to generate the 850 left output channel (e.g., left output channel 121) and the right output channel (e.g., right output channel 123). For example, the conversion from M / S to L / R may be performed using an intermediate component (e.g., processed intermediate component 131) or a side component (e.g., processed side component 139) generated from at least one of the filtered super intermediate component, filtered residual intermediate component, filtered super side component, or filtered residual side component. In another example, the filtered super intermediate component or the filtered residual intermediate component may be used as the intermediate component for the M / S to L / R conversion, or the filtered super side component or residual side component may be used as the side component for the M / S to L / R conversion.

[0126] Orthogonal Subband Spatial and Crosstalk Processing

[0127] Figure 9It is a flowchart of a process for subband spatial processing and crosstalk compensation processing using at least one of a super intermediate component, a residual intermediate component, a super side component, or a residual side component according to one or more embodiments. The crosstalk processing may include crosstalk cancellation or crosstalk simulation. Subband spatial processing may be performed to provide audio content with enhanced spatial detectability, such as by creating the sensation that sound is directed to a listener from a large area rather than a specific point in space corresponding to the speaker location (e.g., sound field enhancement), thereby providing a more immersive listening experience for the listener. Crosstalk simulation may be used for the audio output of headphones to simulate the speaker experience with crosstalk from the opposite side. Crosstalk cancellation may be used for the audio output to speakers to eliminate the effects of crosstalk interference. Crosstalk compensation may compensate for spectral defects caused by crosstalk cancellation or crosstalk simulation. The process may include fewer or more steps, and the steps may be performed in a different order. The super and residual intermediate / side components may be manipulated in different ways for different purposes. For example, in the case of crosstalk compensation, targeted subband filtering may be applied only to the super intermediate component M1 (where most of the vocal dialogue energy in many movie contents occurs) in an effort to eliminate spectral artifacts generated by crosstalk processing in only that component. In the case of sound field enhancement with or without crosstalk processing, targeted subband gain may be applied to the residual intermediate component M2 and the residual side component S2. For example, the residual intermediate component M2 may be attenuated, and the residual side component S2 may be amplified in reverse to increase the distance between these components from the perspective of gain (which can increase spatial detectability if done well) without causing a drastic overall change in the perceived loudness in the final L / R signal, while also avoiding attenuation of the super intermediate M1 component (e.g., the part of the signal that typically contains most of the vocal energy).

[0128] An audio processing system receives 910 an input audio signal, the input audio signal including a left channel and a right channel. In some embodiments, the input audio signal may be a multi-channel audio signal including a plurality of left and right channel pairs. For each left and right input channel pair, each left and right channel pair may be processed as discussed herein.

[0129] The audio processing system applies 920 crosstalk processing to the received input audio signal. The crosstalk processing includes at least one of crosstalk simulation and crosstalk cancellation.

[0130] In steps 930 to 960, the audio processing system performs crosstalk compensation for subband spatial processing and crosstalk processing using one or more of a super intermediate, a super side, a residual intermediate, or a residual side component. In some embodiments, the crosstalk processing may be performed after the processing in steps 930 to 960.

[0131] The audio processing system generates 930 intermediate components and side components from the (e.g., crosstalk-processed) audio signal.

[0132] The audio processing system generates at least one of a super intermediate component, a residual intermediate component, a super side component, and a residual side component. The audio processing system may generate at least one and / or all of the components listed above.

[0133] The audio processing system filters 950 the subbands of at least one of the super intermediate component, the residual intermediate component, the super side component, and the residual side component to apply subband spatial processing to the audio signal. Each subband may include a range of frequencies, such as may be defined by a set of critical bands. In some embodiments, the subband spatial processing further includes applying a time delay to the subbands of at least one of the super intermediate component, the residual intermediate component, the super side component, and the residual side component.

[0134] The audio processing system filters 960 at least one of the super intermediate component, the residual intermediate component, the super side component, and the residual side component to compensate for spectral defects from crosstalk processing of the input audio signal. The spectral defects may include peaks or valleys in the frequency response plots of the super intermediate component, the residual intermediate component, the super side component, or the residual side component that exceed a predetermined threshold (e.g., 10 dB) and appear as artifacts of the crosstalk processing. The spectral defects may be estimated spectral defects.

[0135] In some embodiments, the filtering of the spectral orthogonal components for subband spatial processing in step 950 and the crosstalk compensation in step 960 may be integrated into a single filtering operation for each spectral orthogonal component selected for filtering.

[0136] In some embodiments, the filtering of the super / residual intermediate / side components for subband spatial processing or crosstalk compensation may be performed in combination with filtering for other purposes, such as gain application, amplitude- or delay-based translation, binaural processing, reverberation, dynamic range processing (such as compression and limiting), linear or nonlinear audio processing techniques and effects, ranging from chorus and / or flanging, machine learning-based vocal or instrument style transfer, transformation or resynthesis methods, or other types of processing that use any one of the super intermediate component, the residual intermediate component, the super side component, and the residual side component.

[0137] The filtering may be performed in the frequency domain or the time domain. In some embodiments, the intermediate components and the side components are converted from the time domain to the frequency domain, the super and / or residual components are generated in the frequency domain, the filtering is performed in the frequency domain, and the filtered components are converted back to the time domain. In other embodiments, the super and / or residual components are converted to the time domain, and the filtering is performed on these components in the time domain.

[0138] The audio processing system generates a 970 left output channel and a right output channel from the filtered super-middle component. In some embodiments, the left output channel and the right output channel are additionally based on at least one of the filtered residual middle component, the filtered super-side component, and the filtered residual side component.

[0139] Example Orthogonal Component Audio Processing

[0140] Figures 10 - 19 is a graph depicting the spectral energy of the middle component and the side component of an example white noise signal according to one or more embodiments.

[0141] Figure 10 shows a graph of the white noise signal 1000 translated to the far left (hard left). The left and right white noise signals are converted into a middle component 1005 and a side component 1010 and translated to the far left using the constant power sine / cosine translation law. When the white noise signal is translated to the far left 1000, a user located between the left and right speaker pair will perceive the sound to be at and / or around the left speaker. The white noise signal (split into the left input channel and the right input channel of the white noise signal) can be converted into a middle component 1005 and a side component 1010 using an L / R to M / S converter module 107. As Figure 10 shown, when the white noise signal is translated to the far left 1000, the middle component 1005 and the side component 1010 have approximately equal energy. Similarly, when the white noise signal is translated to the far right ( Figure 10 not shown in), the middle component and the side component will have approximately equal energy.

[0142] Figure 11 shows a graph of the white noise signal 1100 translated to the mid-left. When the white noise signal is translated to the mid-left 1100 using the common constant power sine / cosine translation law, a user located between the left and right speaker pair will perceive the sound to be in the middle between the user's front and the left speaker. Figure 11 depicts the middle component 1105 and the side component 1110 of the white noise signal 1100 translated to the mid-left, and the white noise signal 1000 translated to the far left. Compared to the white noise signal 1000 translated to the far left, the middle component 1105 increases by approximately 3 dB, while the side component 1110 decreases by approximately 6 dB. When the white noise signal is translated to the mid-right, the middle component 1105 and the side component 1110 will have energy similar to that Figure 11 shown.

[0143] Figure 12 shows a graph of the white noise signal 1200 translated to the center. When the white noise signal is translated to the center 1200 using the common constant power sine / cosine translation law, a user located between the left and right speaker pair will perceive the sound to be in front of the user (e.g., between the left and right speakers). AsFigure 12 As shown, the white noise signal 1200 translated to the center only has an intermediate component 1205.

[0144] From Figure 10 , Figure 11 and Figure 12 in the above examples, it can be seen that although for sounds translated to the center as shown in Figure 12 , the intermediate component contains the only energy in the signal (i.e., the left and right channels are the same), in the case where the sound in the original L / R stream is usually perceived as off-center, such as Figure 10 and Figure 11 shown (i.e., sounds translated from the center to the left and right), there is also intermediate component energy.

[0145] It is worth noting that the above three scenarios representing the vast majority of L / R audio use cases do not include the scenario where the side contains the only energy. This only occurs when the left and right channels differ by 180 degrees (i.e., are phase-inverted), which is rare in stereo audio used for music and entertainment. Therefore, while the intermediate component is ubiquitous in almost all stereo left / right audio streams and also includes the only energy in the translated-to-center content, the side component exists in all content except the translated-to-center content and rarely (if at all) as the only energy in the signal.

[0146] Orthogonal component processing isolates and operates on the parts of the intermediate and side components that are "orthogonal" to each other spectrally. That is, using orthogonal component processing, a part of the intermediate component corresponding only to the energy present at the center of the sound field (i.e., the super-intermediate component) can be isolated, and similarly, a part of the side component corresponding only to the energy not present at the center of the sound field (i.e., the super-side component) can be isolated. Conceptually, the super-middle component is the energy corresponding to the thin sound column perceived at the center of the sound field, both for speakers and headphones. Additionally, using a simple scalar, the "thinness" of the column can be controlled to provide an interpolation space from super-intermediate to intermediate and from super-side to side. Furthermore, as a byproduct of deriving our super-intermediate / side component signals, the residual signals (e.g., residual intermediate and side components) can also be operated on, which are combined with the super-intermediate / super-side components to form the original complete intermediate and side components. Each of these four sub-components of intermediate and side can be independently processed in various ways, from simple gain scaling to multi-band equalization to custom and special effects.

[0147] Figures 13 to 19 Shows the orthogonal component processing of the white noise signal. Figure 13A graph showing a white noise signal 1305 translated to the center and band - passed between 20 and 100 Hz (e.g., using an 8 - th order Butterworth filter) and a white noise signal 1310 translated to the far left and band - passed between 5000 and 10000 Hz (e.g., using an 8 - th order Butterworth filter), and without orthogonal component processing. The graph depicts the intermediate components 1315 and side components 1320 of each of the translated white noise signals 1305 and 1310. The white noise signal 1305 translated to the center has energy only in its intermediate component 1315, while the white noise signal translated to the far left has an equal amount of energy in its intermediate component 1315 and side component 1320. This is similar to Figure 10 and Figure 12 the results shown.

[0148] Figure 14 Shows Figure 13 the translated white noise signals 1305 and 1310, where the energy of the side component 1320 is removed. The low - frequency band white noise signal 1305 translated to the center remains unchanged. The high - frequency band white noise signal 1310 translated to the far left now has zero side energy, while a portion of the energy represented by the intermediate component 1315 still remains. Even after removing the lateral energy, there is still non - center - translated energy in the intermediate signal, as shown by signal 1310.

[0149] Figure 15 Shows the translated white noise signals of Figure 13 using orthogonal component processing. Specifically, orthogonal component processing is used to isolate the super - intermediate component 1510 and remove the other energy of the audio signal. Here, the signal translated to the far left is removed, leaving only the signal 1500 translated to the center. This shows that the super - intermediate component 1510 isolates only the energy in the signal that occupies the very center of the sound field and nothing else.

[0150] Because the super - intermediate component of an audio signal can be isolated, the audio signal can be manipulated to control which elements of the original signal ultimately appear in the various M1 / M2 / S1 / S2 components. The scope of such pre - processing operations can range from simple amplitude and delay adjustments to more complex filtering techniques. These pre - processing operations can then be inverted subsequently to restore the original sound field.

[0151] Figure 16 Shows another embodiment of the translated white noise signal of Figure 13 using orthogonal component processing. The L / R audio signal is rotated in such a way that the high - frequency band white noise translated to the far left (e.g., as shown by the signal 1310 in Figure 13 is placed at the center of the sound field and the low - frequency band noise translated to the center (e.g., as shown by Figure 13away from the center as shown by the signal 1305 in []. Then, the white noise signal 1600 that was initially translated to the far left and band - passed between 5000 and 10000 Hz can be extracted by isolating the super - intermediate component 1610 of the rotated L / R signal and further processed.

[0152] Figure 17 shows the decorrelated white noise signal 1700. The input white noise signal 1700 can be a two - channel orthogonal white noise signal including a right - channel component 1710 and a left - channel component 1720. The figure also shows an intermediate component 1730 and a side component 1740 generated from the white noise signal. The spectral energy of the left - channel component 1720 matches the spectral energy of the right - channel component 1710, and the spectral energy of the intermediate component 1730 matches the spectral energy of the side component 1740. The signal levels of the intermediate component 1730 and the side component 1740 are approximately 3 dB lower than those of the right - channel component 1710 and the left - channel component 1720.

[0153] Figure 18 shows the intermediate component 1730 decomposed into a super - intermediate component 1810 and a residual intermediate component 1820. The intermediate component 1730 represents the non - spatial information of the input audio signal in the sound field. The super - intermediate component 1810 includes a sub - component of the non - spatial information directly found at the center of the sound field; the residual intermediate component 1820 is the residual non - spatial information. In a typical stereo audio signal, the super - intermediate component 1810 can include key features of the audio signal, such as dialogue or vocals. In Figure 18 the residual intermediate component 1820 is approximately 3 dB lower than the intermediate component 1730, while the super - intermediate component 1810 is approximately 8 - 9 dB lower than the intermediate component 1730.

[0154] Figure 19 shows the side component 1740 decomposed into a super - side component 1910 and a residual side component 1920. The side component 1740 represents the spatial information in the input audio signal in the sound field. The super - side component 1910 includes a sub - component of the spatial information found at the edge of the sound field; the residual side component 1920 is the residual spatial information. In a typical stereo audio signal, the residual side component 1920 includes key features generated by processing, such as the effects of binaural processing, panning techniques, reverberation, and / or decorrelation processing. As Figure 19 shown, the relationship between the side component 1740, the super - side component 1910, and the residual side component 1920 is similar to the relationship between the intermediate component 1730, the super - intermediate component 1810, and the residual side component 1820.

[0155] Computer architecture

[0156] Figure 20is a block diagram of a computer system 2000 in accordance with one or more embodiments. The computer system 2000 is an example of circuitry that implements an audio processing system. At least one processor 2002 is shown coupled to a chipset 2004. The chipset 2004 includes a memory controller hub 2020 and an input / output (I / O) controller hub 2022. A memory 2006 and a graphics adapter 2012 are coupled to the memory controller hub 2020, and a display device 2018 is coupled to the graphics adapter 2012. A storage device 1008, a keyboard 2010, a pointing device 2014, and a network adapter 2016 are coupled to the I / O controller hub 2022. The computer system 2000 can include various types of input or output devices. Other embodiments of the computer system 2000 have different architectures. For example, in some embodiments, the memory 2006 is directly coupled to the processor 2002.

[0157] The storage device 2008 includes one or more non-transitory computer-readable storage media, such as a hard disk drive, a compact disc read-only memory (CD-ROM), a DVD, or a solid-state memory device. The memory 2006 holds program code (composed of one or more instructions) and data used by the processor 2002. The program code can correspond to the processing aspects described in conjunction with Figures 1 - 19 the description.

[0158] The pointing device 2014 is used in combination with the keyboard 2010 to input data into the computer system 2000. The graphics adapter 2012 displays images and other information on the display device 2018. In some embodiments, the display device 2018 includes touch screen capabilities for receiving user input and selections. The network adapter 2016 couples the computer system 2000 to a network. Some embodiments of the computer system 2000 have components that are different and / or additional to those shown in Figure 20 the figure.

[0159] The circuitry can include one or more processors that execute program code stored in a non-transitory computer-readable medium, which when executed by the one or more processors configures the one or more processors to implement an audio processing system or a module of an audio processing system. Other examples of circuitry that implements an audio processing system or a module of an audio processing system can include integrated circuit devices, such as application-specific integrated circuit devices (ASICs), field-programmable gate arrays (FPGAs), or other types of computer circuitry.

[0160] Additional Notes

[0161] Example benefits and advantages of the disclosed configuration include dynamic audio enhancement resulting from an enhanced audio system adapting to the device and the associated audio rendering system, as well as other relevant information provided by the device OS, such as use case information (e.g., indicating that the audio signal is for music playback rather than gaming). The enhanced audio system can be integrated into the device (e.g., using a software development kit) or stored on a remote server for on-demand access. In this way, the device does not need to use storage or processing resources to maintain an audio enhancement system specific to its audio rendering system or audio rendering configuration. In some embodiments, the enhanced audio system is capable of making different levels of queries of the rendering system information, such that effective audio enhancement can be applied across different levels of available device-specific rendering information.

[0162] Throughout the specification, multiple instances may implement components, operations, or structures described as a single instance. Although individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed simultaneously, and there is no requirement that the operations be performed in the order shown. Structures and functions presented as separate components in example configurations may be implemented as a combined structure or component. Similarly, structures and functions presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein.

[0163] Certain embodiments are described herein as including logic or a plurality of components, modules, or mechanisms. A module may constitute a software module (e.g., code included on a machine-readable medium or in a transmitted signal) or a hardware module. A hardware module is a tangible unit capable of performing certain operations and may be configured or arranged in a certain manner. In example embodiments, one or more computer systems (e.g., stand-alone client or server computer systems) or one or more hardware modules of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or a portion of an application) as a hardware module for performing certain operations described herein.

[0164] The various operations of the example methods described herein may be performed, at least in part, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors may constitute processor-implemented modules that perform one or more operations or functions. In some example embodiments, the modules referred to herein may include processor-implemented modules.

[0165] Similarly, the methods described herein can be implemented, at least in part, by a processor. For example, at least some of the operations of a method can be performed by one or more processors or hardware modules implemented by a processor. The execution of certain operations can be distributed among one or more processors, not only residing within a single machine but also deployed across multiple machines. In some example embodiments, one or more processors can be located in a single location (e.g., in a home environment, an office environment, or as a server farm), while in other embodiments, the processors can be distributed across multiple locations.

[0166] Unless otherwise explicitly stated, discussions herein using terms such as "processing", "computing", "calculating", "determining", "presenting", "displaying", etc., can refer to actions or processes of a machine (e.g., a computer) that manipulates or transforms data represented as physical (electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.

[0167] As used herein, any reference to "an embodiment" or "embodiments" means that a particular element, feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The phrase "in an embodiment" appearing in various places in the specification is not necessarily all referring to the same embodiment.

[0168] Some embodiments may use the terms "coupled" and "connected" along with their derivatives to describe. It should be understood that these terms are not intended as synonyms for each other. For example, the term "connected" can be used in some embodiments to indicate that two or more elements are in direct physical or electrical contact with each other. In another example, the term "coupled" can be used in some embodiments to indicate that two or more elements are in direct physical or electrical contact. However, the term "coupled" can also mean that two or more elements are not in direct contact with each other but still cooperate or interact with each other. The embodiments are not limited to this context.

[0169] As used herein, the terms "comprises", "comprising", "includes", "including", "has", "having" or any other variation thereof are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, "or" refers to an inclusive or and not an exclusive or. For example, any one of the following satisfies condition A or B: A is true (or present) and B is false (or absent), A is false (or absent) and B is true (or present), A and B are both true (or present).

[0170] In addition, the articles "a" or "an" are used to describe elements and components of the embodiments herein. This is done merely for convenience and to give a general sense of the invention. This description should be understood to include one or at least one, and the singular also includes the plural, unless it is obvious that it has a different meaning.

[0171] Some portions of this specification describe embodiments in terms of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are commonly used by those skilled in the data processing arts to effectively convey the substance of their work to others skilled in the art. These operations, while described functionally, computationally, or logically, are understood to be implemented by a computer program, an equivalent circuit arrangement, microcode, or the like. Additionally, without loss of generality, it has sometimes proven convenient to refer to these operational arrangements as modules. The described operations and their associated modules can be embodied in software, firmware, hardware, or any combination thereof.

[0172] Any step, operation, or process described herein can be performed or implemented using one or more hardware or software modules, alone or in combination with other devices. In one embodiment, the software modules are implemented with a computer program product that includes a computer-readable medium containing computer program code that can be executed by a computer processor to perform any or all of the described steps, operations, or processes.

[0173] The embodiments may also relate to apparatus for performing the operations herein. The apparatus may be specially constructed for the required purposes, and / or it may comprise a general purpose computing device selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a non-transitory tangible computer readable storage medium, or in any type of suitable medium for storing electronic instructions, which may be coupled to a computer system bus. Further, any computing system referred to in this specification may include a single processor, or may be an architecture employing multiple processor designs to increase computing capability.

[0174] The embodiments may also relate to products produced by the computing processes described herein. Such products may include information produced by a computing process, where the information is stored on a non-transitory tangible computer readable storage medium and may include any embodiment of a computer program product or other data combinations described herein.

[0175] After reading this disclosure, those skilled in the art will appreciate additional alternative structural and functional designs of systems and processes for audio enhancement using device-specific metadata through the principles disclosed herein. Thus, while particular embodiments and applications have been illustrated and described, it should be understood that the disclosed embodiments are not limited to the precise construction and components disclosed herein. Various modifications, changes and variations apparent to those skilled in the art may be made to the arrangement, operation and details of the methods and apparatus disclosed herein without departing from the spirit and scope defined by the appended claims.

[0176] Finally, the language used in the specification has been principally selected for readability and guidance purposes, rather than to describe or limit the patent rights. Accordingly, it is intended that the scope of the patent rights not be limited by this detailed description, but rather by any claims issued on an application based hereon. Thus, the disclosure of embodiments is intended to illustrate rather than limit the scope of the patent rights set forth in the appended claims.

Claims

1. A system for processing an audio signal, comprising: circuit means configured to: generate an intermediate component and a side component from a left channel and a right channel of the audio signal; convert the intermediate component and the side component to a frequency domain; generate a super side component including removing the spectral energy of the intermediate component from the spectral energy of the side component by subtracting a magnitude of the intermediate component in the frequency domain from a magnitude of the side component in the frequency domain; filter the super side component; and generate a left output channel and a right output channel using the filtered super side component.

2. The system according to claim 1, wherein: the super side component isolates a part of the side component corresponding to spectral energy not at the center of the sound field.

3. The system according to claim 1, wherein: the circuit means is configured to apply a Fourier transform to the intermediate component and the side component to convert the intermediate component and the side component to the frequency domain.

4. The system according to claim 1, wherein the circuitry is configured to filter the super-side component to include: the circuit means is configured to perform at least one of gain adjustment or time delay on subbands of the super side component.

5. The system according to claim 1, wherein the circuitry is configured to filter the super-side component to include: the circuit means is configured to apply dynamic range processing to the super side component.

6. The system according to claim 1, wherein the circuitry is configured to filter the super-side component to include: the circuit means is configured to adjust a frequency-dependent amplitude or a frequency-dependent delay of the super side component.

7. The system according to claim 1, wherein the circuitry is configured to filter the super-side component to include: the circuit means is configured to apply machine learning-based style transfer, conversion, or resynthesis to the super side component.

8. The system according to claim 1, wherein the circuit means is further configured to: generate a residual side component including removing the spectral energy of the super side component from the spectral energy of the side component; filter the residual side component; and generate the left output channel and the right output channel using the filtered residual side component.

9. The system according to claim 8, wherein the circuit device is configured to generate the residual side component including: the circuit means is configured to subtract a magnitude of the super side component in the frequency domain from a magnitude of the side component in the frequency domain.

10. The system according to claim 8, wherein the circuit device is configured to generate a residual side component including: the circuit means is configured to generate a delayed side component by time-delaying the side component; generate the residual side component by subtracting the super side component from the delayed side component in the time domain.

11. The system according to claim 1, wherein the circuit means is further configured to: generate a super intermediate component including removing the spectral energy of the side component from the spectral energy of the intermediate component; filter the super intermediate component; and generate the left output channel and the right output channel using the filtered super intermediate component.

12. The system according to claim 11, wherein, the circuit means is configured to generate the super intermediate component including: the circuit means is configured to subtract a magnitude of the intermediate component in the frequency domain from a magnitude of the side component in the frequency domain.

13. The system according to claim 1, wherein the circuit means is further configured to: generate a super intermediate component including removing the spectral energy of the side component from the spectral energy of the intermediate component; generate a residual intermediate component including removing the spectral energy of the super intermediate component from the spectral energy of the intermediate component; filter the residual intermediate component; and generate the left output channel and the right output channel using the filtered residual intermediate component.

14. The system according to claim 13, wherein, The circuit device is configured to generate the residual intermediate component including: the circuit device is configured to subtract the magnitude of the super intermediate component in the frequency domain from the magnitude of the intermediate component in the frequency domain.

15. The system according to claim 13, wherein, The circuit device is configured to generate the residual intermediate component including that the circuit device is configured to: generate a delayed intermediate component by time-delaying the intermediate component; generate the residual intermediate component by subtracting the super intermediate component from the delayed intermediate component in the time domain.

16. The system according to claim 1, wherein the circuit device is further configured to apply crosstalk processing to the audio signal, and the crosstalk processing includes one of crosstalk cancellation or crosstalk simulation.

17. The system according to claim 16, wherein the circuit device is further configured to filter the side component to compensate for spectral defects caused by the crosstalk processing.

18. The system according to claim 16, wherein the circuit device is further configured to filter the super side component to compensate for spectral defects caused by the crosstalk processing.

19. A non-transitory computer-readable medium, including stored program code, the program code when executed by at least one processor configures the at least one processor to: generate an intermediate component and a side component from the left and right channels of an audio signal; convert the intermediate component and the side component to the frequency domain; generate a super side component including removing the spectral energy of the intermediate component from the spectral energy of the side component by subtracting the magnitude of the intermediate component in the frequency domain from the magnitude of the side component in the frequency domain; filter the super side component; and generate a left output channel and a right output channel using the filtered super side component.

20. The non-transitory computer-readable medium according to claim 19, wherein: the super side component isolates the part of the side component corresponding to the spectral energy not at the center of the sound field.

21. The non-transitory computer-readable medium according to claim 19, wherein the program code further configures the at least one processor to: apply a Fourier transform to the intermediate component and the side component to convert the intermediate component and the side component to the frequency domain.

22. The non-transitory computer-readable medium according to claim 19, wherein the program code configuring the at least one processor to filter the super side component further configures the at least one processor to: perform at least one of gain adjustment or time delay on subbands of the super side component.

23. The non-transitory computer-readable medium according to claim 19, wherein the program code configuring the at least one processor to filter the super side component further configures the at least one processor to: apply dynamic range processing to the super side component.

24. The non-transitory computer-readable medium according to claim 19, wherein the program code configuring the at least one processor to filter the super side component further configures the at least one processor to: adjust the frequency-dependent amplitude or frequency-dependent delay of the super side component.

25. The non-transitory computer-readable medium according to claim 19, wherein the program code configures the at least one processor to filter the super-side component and further configures the at least one processor to apply machine learning-based style transfer, transformation, or resynthesis to the super-side component.

26. The non-transitory computer-readable medium according to claim 19, wherein the program code further configures the at least one processor to: generate a residual side component, the residual side component including removing the spectral energy of the super-side component from the spectral energy of the side component; filter the residual side component; and generate the left output channel and the right output channel using the filtered residual side component.

27. The non-transitory computer-readable medium according to claim 26, wherein the program code configures the at least one processor to generate a residual side component and further configures the at least one processor to: subtract the magnitude of the super-side component in the frequency domain from the magnitude of the side component in the frequency domain.

28. The non-transitory computer-readable medium according to claim 26, wherein the program code configures the at least one processor to generate a residual side component and further configures the at least one processor to: generate a delayed side component by time-delaying the side component; generate the residual side component by subtracting the super-side component from the delayed side component in the time domain.

29. The non-transitory computer-readable medium according to claim 19, wherein the program code further configures the at least one processor to: generate a super-middle component, the super-middle component including the spectral energy of the side component removed from the spectral energy of the middle component; filter the super-middle component; and generate the left output channel and the right output channel using the filtered super-middle component.

30. The non-transitory computer-readable medium according to claim 29, wherein the program code configures the at least one processor to generate the super-middle component and further configures the at least one processor to: subtract the magnitude of the middle component in the frequency domain from the magnitude of the side component in the frequency domain.

31. The non-transitory computer-readable medium according to claim 19, wherein the program code further configures the at least one processor to: generate a super-middle component including removing the spectral energy of the side component from the spectral energy of the middle component; generate a residual middle component including removing the spectral energy of the super-middle component from the spectral energy of the middle component; filter the residual middle component; and generate the left output channel and the right output channel using the filtered residual middle component.

32. The non-transitory computer-readable medium according to claim 31, wherein the program code configures the at least one processor to generate the residual middle component and further configures the at least one processor to: subtract the magnitude of the super-middle component in the frequency domain from the magnitude of the middle component in the frequency domain.

33. The non-transitory computer-readable medium according to claim 31, wherein the program code configures the at least one processor to generate the residual intermediate component and further configures the at least one processor to: generate a delayed intermediate component by time-delaying the intermediate component; generate the residual intermediate component by subtracting the super intermediate component from the delayed intermediate component in the time domain.

34. The non-transitory computer-readable medium according to claim 19, wherein the circuitry is further configured to apply crosstalk processing to the audio signal, the crosstalk processing including one of crosstalk cancellation or crosstalk simulation.

35. The non-transitory computer-readable medium according to claim 24, wherein the program code further configures the at least one processor to filter the side component to compensate for spectral defects caused by the crosstalk processing.

36. The non-transitory computer-readable medium according to claim 34, wherein the program code further configures the at least one processor to filter the super side component to compensate for spectral defects caused by the crosstalk processing.

37. A method for processing an audio signal, comprising: By circuitry generate an intermediate component and a side component from a left channel and a right channel of the audio signal; convert the intermediate component and the side component to the frequency domain; generate a super side component including removing spectral energy of the intermediate component from spectral energy of the side component by subtracting a magnitude of the intermediate component in the frequency domain from a magnitude of the side component in the frequency domain; filter the super side component; and generate a left output channel and a right output channel using the filtered super side component.

38. The method according to claim 37, wherein the super side component isolates a portion of the side component corresponding to spectral energy not at the center of the sound field.

39. The method according to claim 37, further comprising: applying a Fourier transform to the intermediate component and the side component to convert the intermediate component and the side component to the frequency domain.

40. The method according to claim 37, wherein filtering the super side component comprises: performing at least one of gain adjustment or time delay on subbands of the super side component.

41. The method according to claim 37, wherein filtering the super-side component comprises: applying dynamic range processing to the super side component.

42. The method according to claim 37, wherein filtering the super side component comprises: adjusting frequency-dependent amplitude or frequency-dependent delay of the super side component.

43. The system according to claim 37, wherein filtering the super lateral component comprises: applying machine learning-based style transfer, transformation, or resynthesis to the super side component.

44. The method according to claim 37, further comprising: generate a residual side component by the circuitry, the residual side component including removing spectral energy of the super side component from spectral energy of the side component; filter the residual side component; and generate the left output channel and the right output channel using the filtered residual side component.

45. The method according to claim 44, wherein generating the residual side component comprises: Subtracting a magnitude of the super side component in the frequency domain from a magnitude of the side component in the frequency domain.

46. The method according to claim 44, wherein generating the residual side component includes: generating a delayed side component by time-delaying the side component; generating the residual side component by subtracting the super side component from the delayed side component in the time domain.

47. The method according to claim 37, further comprising by the circuitry: Generate a super-intermediate component that includes spectral energy of a side component removed from spectral energy of a mid-component; Filter the super-intermediate component; and Generate the left output channel and the right output channel using the filtered super-intermediate component.

48. The method according to claim 47, wherein generating the super intermediate component comprises: Subtract a magnitude of the mid-component in the frequency domain from a magnitude of the side component in the frequency domain.

49. The method according to claim 37, further comprising, by the circuitry: Generate a super-intermediate component that includes spectral energy of the side component removed from spectral energy of the mid-component; Generate a residual mid-component that includes spectral energy of the super-intermediate component removed from spectral energy of the mid-component; Filter the residual mid-component; And Generate the left output channel and the right output channel using the filtered residual mid-component.

50. The method according to claim 49, wherein generating the residual intermediate component comprises: Subtract a magnitude of the super-intermediate component in the frequency domain from a magnitude of the mid-component in the frequency domain.

51. The method according to claim 49, wherein generating the residual mid-component includes: Generate a delayed mid-component by applying a time delay to the mid-component; Generate the residual mid-component by subtracting the super-intermediate component from the delayed mid-component in the time domain.

52. The method according to claim 37, further comprising, by the circuitry: Apply crosstalk processing to the audio signal, the crosstalk processing including one of crosstalk cancellation or crosstalk simulation.

53. The method according to claim 52, further comprising, by the circuitry: Filter the side component to compensate for spectral defects caused by the crosstalk processing.

54. The method according to claim 52, further comprising, by the circuitry: Filter the super-side component to compensate for spectral defects caused by the crosstalk processing.