Spectral quadrature audio component processing

By generating and operating audio components that are orthogonal to the spectrum, the limitations of processing audio signals in the prior art are solved, and a wider processing possibility and a more immersive auditory experience of the audio signals are achieved.

CN114830693BActive Publication Date: 2025-05-06BOOMCLOUD 360 INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080085638.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-08-03
Filing Date
2020-08-10
Publication Date
2025-05-06
Estimated Expiration
2040-08-10

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process spatial information in audio signals, especially in terms of enhancing the range of possibilities for processing audio.

Method used

By generating and operating super-intermediate components, super-side components, residual intermediate components and residual side components orthogonal to the spectrum, these components are utilized for filtering and other types of processing to provide spatial cues and audio content enhancement.

Benefits of technology

A wider processing possibilities for audio signals are achieved, enabling the ability to adjust audio content without changing the spectrum energy of other parts of the sound field, providing a more immersive auditory experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114830693B_ABST
    Figure CN114830693B_ABST
Patent Text Reader

Abstract

A system for processing an audio signal using spectrally orthogonal sound components. The system includes a circuit device that generates a mid component and a side component from a left channel and a right channel of the audio signal. The circuit device generates a super-mid component, the super-mid component including spectral energy of the side component removed from spectral energy of the mid component. The circuit device filters the super-mid component, such as to provide spatial cue processing, including panning or binaural processing, dynamic range processing, or other types of processing. The circuit device generates a left output channel and a right output channel using the filtered super-mid component.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to audio processing, and more particularly to spatial audio processing. Background Art

[0002] Conceptually, the side (or "spatial") components of a left and right stereo signal can be thought of as portions of the left and right channels that include spatial information (i.e., sounds in the stereo signal that appear anywhere to the left or right of the center of the sound field). Conversely, the middle (or "non-spatial") components of a left and right stereo signal can be thought of as portions of the left and right channels that include non-spatial information (i.e., sounds in the stereo signal that appear in the center of the sound field). While the middle component contains energy in the stereo signal that is perceived as non-spatial, it typically also has energy from elements of the stereo signal that are not perceptually located in the center of the sound field. Similarly, while the side component contains energy in the stereo signal that is perceived as spatial, it typically also has energy from elements of the stereo signal that are perceptually located in the center of the sound field. In order to enhance the range of possibilities for processing audio, it is necessary to isolate and manipulate portions of the middle and side components that are spectrally "orthogonal" to each other. Summary of the invention

[0003] Embodiments relate to audio processing using spectrally orthogonal audio components, such as a super-mid component, a super-side component, a residual mid component, or a residual side component of a stereo audio signal or other multi-channel audio signal. The super-mid component and the super-side component are spectrally orthogonal to each other, and the residual mid component and the residual side component are spectrally orthogonal to each other.

[0004] Some embodiments include a system for processing an audio signal. The system includes a circuit device that generates a mid component and a side component from a left channel and a right channel of the audio signal. The circuit device generates a super-mid component that includes spectral energy of the side component removed from spectral energy of the mid component. The circuit device filters the super-mid component, such as to provide spatial cue processing, including panning or binaural processing, dynamic range processing, or other types of processing. The circuit device generates a left output channel and a right output channel using the filtered super-mid component.

[0005] In some embodiments, the circuit device applies a Fourier transform to the mid component and the side component to convert the mid component and the side component to the frequency domain. The circuit device generates a super mid component by subtracting a magnitude of the side component in the frequency domain from a magnitude of the mid component in the frequency domain.

[0006] In some embodiments, the circuit device filters the super-middle component to gain adjust or time delay a sub-band of the super-middle component. In some embodiments, the circuit device filters the super-middle component to apply dynamic range processing to the super-middle component. In some embodiments, the circuit device filters the super-middle component to adjust the frequency-dependent amplitude or frequency-dependent delay of the super-middle component. In some embodiments, the circuit device filters the super-middle component to apply machine learning-based style transfer, conversion, or resynthesis to the super-middle component.

[0007] In some embodiments, the circuit arrangement generates a residual mid component comprising removing spectral energy of the super mid component from spectral energy of the mid component, filters the residual mid component, and generates left and right output channels using the filtered residual mid component.

[0008] In some embodiments, the circuit device filters the residual intermediate component to gain adjust or time delay a subband of the residual intermediate component. In some embodiments, the circuit device filters the residual intermediate component to apply dynamic range processing to the residual intermediate component. In some embodiments, the circuit device filters the residual intermediate component to adjust the frequency-dependent amplitude or frequency-dependent delay of the residual intermediate component. In some embodiments, the circuit device filters the residual intermediate component to apply machine learning-based style transfer, conversion, or resynthesis to the residual intermediate component.

[0009] In some embodiments, the circuit arrangement applies a Fourier transform to the intermediate component to convert the intermediate component to the frequency domain.The circuit arrangement generates a residual intermediate component by subtracting a magnitude of the super intermediate component in the frequency domain from a magnitude of the intermediate component in the frequency domain.

[0010] In some embodiments, the circuit device applies an inverse Fourier transform to the super intermediate component to convert the super intermediate component in the frequency domain to the time domain, generates a delayed intermediate component by time delaying the intermediate component, generates a residual intermediate component by subtracting the super intermediate component in the time domain from the delayed intermediate component in the time domain, filters the residual intermediate component, and generates a left output channel and a right output channel using the filtered residual intermediate component.

[0011] In some embodiments, the circuit arrangement generates a super-side component comprising removing spectral energy of the mid component from spectral energy of the side component, filters the super-side component, and generates left and right output channels using the filtered super-side component.

[0012] In some embodiments, the circuit device applies a Fourier transform to the mid component and the side component to convert the mid component and the side component to the frequency domain. The circuit device generates the super-side component by subtracting the magnitude of the mid component in the frequency domain from the magnitude of the side component in the frequency domain.

[0013] In some embodiments, the circuit device filters the super-side component to gain adjust or time delay a subband of the super-side component. In some embodiments, the circuit device filters the super-side component to apply dynamic range processing to the super-side component. In some embodiments, the circuit device filters the super-side component to adjust the frequency-dependent amplitude or frequency-dependent delay of the super-side component. In some embodiments, the circuit device filters the super-side component to apply machine learning-based style transfer, conversion, or resynthesis to the super-side component.

[0014] In some embodiments, the circuit device generates a super-side component including spectral energy of the mid component removed from spectral energy of the side component, generates a residual side component including spectral energy of the super-side component removed from spectral energy of the side component, filters the residual side component, and generates a left output channel and a right output channel using the filtered residual side component.

[0015] In some embodiments, the circuit device filters the residual side component to gain adjust or time delay a subband of the residual side component. In some embodiments, the circuit device filters the residual side component to apply dynamic range processing to the residual side component. In some embodiments, the circuit device filters the residual side component to adjust the frequency-dependent amplitude or frequency-dependent delay of the residual side component. In some embodiments, the circuit device filters the residual side component to apply machine learning-based style transfer, conversion, or resynthesis to the residual side component.

[0016] In some embodiments, the circuit device applies a Fourier transform to the side component to convert the side component to the frequency domain. The circuit device generates a residual side component by subtracting the magnitude of the super side component in the frequency domain from the magnitude of the side component in the frequency domain.

[0017] In some embodiments, the circuit device generates a super-side component including removing spectral energy of the middle component from spectral energy of the side component, applying an inverse Fourier transform to the super-side component to convert the super-middle component to a time domain, generating a delayed side component by time-delaying the side component, generating a residual side component by subtracting the super-side component in a time domain from the delayed side component in the time domain, filtering the residual side component, and generating a left output channel and a right output channel using the filtered residual side component.

[0018] Some embodiments include a non-transitory computer-readable medium including stored program code. The program code, when executed by at least one processor, configures the at least one processor to generate a mid component and a side component from left and right channels of an audio signal, generate a super-mid component including spectral energy of the side component removed from spectral energy of the mid component, filter the super-mid component, and generate a left output channel and a right output channel using the filtered super-mid component.

[0019] Some embodiments include a method for processing an audio signal by a circuit device. The method includes generating a mid component and a side component from left and right channels of the audio signal, generating a super mid component including spectral energy of the side component removed from spectral energy of the mid component, filtering the super mid component, and generating a left output channel and a right output channel using the filtered super mid component. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The disclosed embodiments have other advantages and features that will be more readily apparent from the detailed description, the appended claims, and the accompanying drawings (or figures).The following is a brief description of the drawings.

[0021] Figure 1 is a block diagram of an audio processing system according to one or more embodiments.

[0022] Figure 2A is a block diagram of a quadrature component generator according to one or more embodiments.

[0023] Figure 2B is a block diagram of a quadrature component generator according to one or more embodiments.

[0024] Figure 2C is a block diagram of a quadrature component generator according to one or more embodiments.

[0025] Figure 3 is a block diagram of a quadrature component processor according to one or more embodiments.

[0026] Figure 4 is a block diagram of a sub-band spatial processor according to one or more embodiments.

[0027] Figure 5 is a block diagram of a crosstalk compensation processor according to one or more embodiments.

[0028] Figure 6 is a block diagram of a crosstalk simulation processor according to one or more embodiments.

[0029] Figure 7 is a block diagram of a crosstalk cancellation processor according to one or more embodiments.

[0030] Figure 8is a flow chart of a process for spatial processing using at least one of a super mid component, a residual mid component, a super side component, or a residual side component according to one or more embodiments.

[0031] Fig. 9 is a flow chart of a process for performing sub-band spatial processing and crosstalk compensation processing using at least one of a super middle component, a residual middle component, a super side component, or a residual side component according to one or more embodiments.

[0032] Figure 10-Figure 19 is a graph depicting spectral energies of mid and side components of an example white noise signal in accordance with one or more embodiments.

[0033] Fig. 20 is a block diagram of a computer system according to one or more embodiments. DETAILED DESCRIPTION

[0034] The drawings and the following description relate to the preferred embodiments by way of illustration only. It should be noted that from the following discussion, alternative embodiments of the structures and methods disclosed herein will readily be identified as viable alternatives that may be employed without departing from the principles claimed.

[0035] Reference will now be made in detail to several embodiments, examples of which are shown in the accompanying drawings. Note that, wherever feasible, similar or like reference numerals may be used in the accompanying drawings and may indicate similar or like functions. The accompanying drawings depict embodiments of the disclosed systems (or methods) for illustrative purposes only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods shown herein may be employed without departing from the principles described herein.

[0036] Embodiments relate to spatial audio processing using mid and side components that are spectrally orthogonal to each other. For example, an audio processing system generates a super-mid component or a super-side component, the super-mid component isolating a portion of the mid component corresponding only to spectral energy present in the center of the sound field, and the super-side component isolating a portion of the side component corresponding only to spectral energy not present in the center of the sound field. The super-mid component includes the spectral energy of the side component removed from the spectral energy of the mid component, and the super-side component includes the spectral energy of the mid component removed from the spectral energy of the side component. The audio processing system can also generate a residual mid component and a residual side component, the residual mid component corresponding to the spectral energy of the mid component with the super-mid component removed (e.g., by subtracting the spectral energy of the super-mid component from the spectral energy of the mid component), and the residual side component corresponding to the spectral energy of the side component with the super-mid component removed (e.g., by subtracting the spectral energy of the super-side component from the spectral energy of the side component). By isolating these orthogonal components and performing various types of audio processing using these components, the audio processing system can provide targeted audio content enhancement. The super-mid component represents non-spatial (i.e., mid) spectral energy in the center of the sound field. For example, the non-spatial spectral energy in the center of the sound field may include the main vocal content in the dialogue or music of a movie. Applying signal processing operations to the super-middle enables such audio content to be adjusted without changing the spectral energy present elsewhere in the sound field. For example, in some embodiments, the vocals of the sound can be partially and / or completely removed by applying a filter that reduces the spectral energy within the typical human vocal range to the super-middle component. In other embodiments, targeted vocal enhancement or effects can be applied to the vocal content by increasing the filter of the energy within the typical human vocal range (e.g., via compression, reverberation and / or other audio processing techniques). The residual middle component represents the non-spatial spectral energy that is not in the center of the sound field. Applying signal processing techniques to the residual middle allows similar transformations to occur orthogonally from other components. For example, in some embodiments, in order to provide a spatial widening effect to the audio content with minimal changes in the overall perceived gain and minimal loss of vocal presence, the targeted spectral energy in the residual middle component can be partially and / or completely removed while increasing the spectral energy in the residual side component.

[0037] Example Audio Processing System

[0038] Figure 11 is a block diagram of an audio processing system 100 according to one or more embodiments. The audio processing system 100 is a circuit device that processes an input audio signal to generate a spatially enhanced output audio signal. The input audio signal includes a left input channel 103 and a right input channel 105, and the output audio signal includes a left output channel 121 and a right output channel 123. The audio processing system 100 includes an L / R to M / S converter module 107, an orthogonal component generator module 113, an orthogonal component processor module 117, an M / S to L / R converter module 119, and a crosstalk processor module 141. In some embodiments, the audio processing system 100 includes a subset of the above components and / or additional components in addition to the above components. In some embodiments, the audio processing system 100 is configured to be different from the above components. Figure 1 For example, the audio processing system 100 may process the input audio using crosstalk processing before processing using the quadrature component generator module 113 and the quadrature component processor module 117 .

[0039] The L / R to M / S converter module 107 receives the left input channel 103 and the right input channel 105, and generates a middle component 109 (e.g., a non-spatial component) and a side component 111 (e.g., a spatial component) from the input channels 103 and 105. In some embodiments, the middle component 109 is generated based on the sum of the left input channel 103 and the right input channel 105, and the side component 111 is generated based on the difference between the left input channel 103 and the right input channel 105. In some embodiments, several middle components and side components are generated from a multi-channel input audio signal (e.g., surround sound). Other L / R to M / S types of transformations may be used to generate the middle component 109 and the side component 111.

[0040] The orthogonal component generator module 113 processes the mid component 109 and the side component 111 to generate at least one of the following: a super mid component M1, a super side component S1, a residual mid component M2, and a residual side component S2. The super mid component M1 is the mid component 109 with the side component 111 removed. The super side component S1 is the spectral energy of the side component 111 with the spectral energy of the mid component 109 removed. The residual mid component M2 is the spectral energy of the mid component 109 with the spectral energy of the super mid component M1 removed. The residual side component S2 is the spectral energy of the side component 111 with the spectral energy of the super side component S1 removed. In some embodiments, the audio processing system 100 generates a left output channel 121 and a right output channel 123 by processing at least one of the super mid component M1, the super side component S1, the residual mid component M2, and the residual side component S2. The orthogonal component generator module 113 is related to Figure 2A-2C Further description.

[0041] The quadrature component processor module 117 processes one or more of the super-middle component M1, the super-side component S1, the residual middle component M2, and / or the residual side component S2. The processing of the components M1, M2, S1, and S2 may include various types of filtering, such as spatial cue processing (e.g., amplitude or delay-based panning, binaural processing, etc.), dynamic range processing, machine learning-based processing, gain application, reverberation, adding audio effects, or other types of processing. In some embodiments, the quadrature component processor module 117 uses the super-middle component M1, the super-side component S1, the residual middle component M2, and / or the residual side component S2 to perform sub-band spatial processing and / or crosstalk compensation processing to generate a processed middle component 131 and a processed side component 139. Sub-band spatial processing is processing performed on frequency sub-bands of the middle and side components of the audio signal to spatially enhance the audio signal. Crosstalk compensation processing is processing performed on the audio signal to adjust spectral artifacts caused by crosstalk processing, such as crosstalk compensation of a speaker or crosstalk simulation of a headphone. Quadrature component processor module 117 Figure 3 Further description.

[0042] The M / S to L / R converter module 119 receives the processed middle component 131 and the processed side component 139 and generates a processed left component 151 and a processed right component 159. In some embodiments, the processed left component 151 is generated based on the sum of the processed middle component 131 and the processed side component 139, and the processed right component 159 is generated based on the difference between the processed middle component 131 and the processed side component 139. Other M / S to L / R transform types may be used to generate the processed left component 151 and the processed right component 159.

[0043] The crosstalk processor module 141 receives the processed left component 151 and the processed right component 159 and performs crosstalk processing thereon. Crosstalk processing includes, for example, crosstalk simulation or crosstalk cancellation. Crosstalk simulation is a process performed on an audio signal (e.g., output via headphones) to simulate the effect of a loudspeaker. Crosstalk cancellation is a process performed on an audio signal configured to be output via a loudspeaker to cancel the crosstalk caused by the loudspeaker. The crosstalk processor module 141 outputs a left output channel 121 and a right output channel 123.

[0044] Example Quadrature Component Generator

[0045] Figure 2A-2C 2 and 3 are block diagrams of orthogonal component generator modules 213, 223, and 243, respectively, according to one or more embodiments. Orthogonal component generator modules 213, 223, and 243 are examples of orthogonal component generator module 113.

[0046] refer to Figure 2A , the orthogonal component generator module 213 includes a subtraction unit 205, a subtraction unit 209, a subtraction unit 215, and a subtraction unit 219. As described above, the orthogonal component generator module 113 receives the middle component 109 and the side component 111, and outputs one or more of the super middle component M1, the super side component S1, the residual middle component M2, and the residual side component S2.

[0047] The subtraction unit 205 removes the spectral energy of the side component 111 from the spectral energy of the middle component 109 to generate a super middle component M1. For example, the subtraction unit 205 subtracts the size of the side component 111 in the frequency domain from the size of the middle component 109 in the frequency domain, while ignoring the phase, to generate the super middle component M1. Frequency domain subtraction can be performed on a time domain signal using a Fourier transform to generate a frequency domain signal, followed by subtraction of the frequency domain signal. In other examples, frequency domain subtraction can be performed in other ways, such as using a wavelet transform instead of a Fourier transform. The subtraction unit 209 generates a residual middle component M2 by removing the spectral energy of the super middle component M1 from the spectral energy of the middle component 109. For example, the subtraction unit 209 subtracts the size of the super middle component M1 in the frequency domain from the size of the middle component 109 in the frequency domain, while ignoring the phase, to generate a residual middle component M2. While subtracting the side from the middle in the time domain results in the original right channel of the signal, the above operation in the frequency domain isolates and distinguishes between a portion of the spectral energy of the mid component that is different from the spectral energy of the side component (called M1, or super middle), and a portion of the spectral energy of the mid component that is the same as the spectral energy of the side component (called M2, or residual middle).

[0048] In some embodiments, additional processing may be used when the spectral energy of the side component 111 is subtracted from the spectral energy of the middle component 109 to obtain a negative value of the super-middle component M1 (e.g., for one or more of the intervals in the frequency domain). In some embodiments, when the spectral energy of the side component 111 is subtracted from the spectral energy of the middle component 109 to obtain a negative value, the super-middle component M1 is clamped at a value of 0. In some embodiments, the super-middle component M1 is wrapped around by taking the absolute value of the negative value as the value of the super-middle component M1. Other types of processing may be used when the spectral energy of the side component 111 is subtracted from the spectral energy of the middle component 109 to obtain a negative value of M1. Similar additional processing, such as clamping at 0, wrapping around, or other processing, may be used when the subtraction result of generating the super-side component S1, the residual side component S2, or the residual middle component M2 is negative. When the subtraction results in a negative value, clamping the super-middle component M1 at 0 will ensure spectral orthogonality between M1 and the two side components. Similarly, when the subtraction results in a negative value, clamping the super-side component S1 at 0 will ensure spectral orthogonality between S1 and the two middle components. By creating orthogonality between the super-middle and side components and their appropriate middle / side counterparts (i.e., side components for the super-middle, middle components for the super-side), the derived residual middle M2 and residual side S2 components contain spectral energy that is not orthogonal to (i.e., shared with) their appropriate middle / side counterparts. That is, when the super-middle is clamped at 0 and the residual middle is derived using the M1 component, a super-middle component whose spectral energy is not shared with the side component and a residual middle component whose spectral energy is completely shared with the side component are generated. When the super-side is clamped to 0, the same relationship applies to the super-side and the residual side. When applying frequency domain processing, it is usually necessary to make a trade-off in resolution between frequency and timing information. As the frequency resolution increases (i.e., as the FFT window size and the number of frequency bins increase), the time resolution decreases, and vice versa. The spectral subtraction described above occurs on a per frequency bin basis, so in some cases, such as when removing vocal energy from the super-mid component, it is preferable to use a larger FFT window size (e.g., 8192 samples, resulting in 4096 frequency bins given a real-valued input signal). Other cases may require a higher time resolution and therefore a lower overall delay and a lower frequency resolution (e.g., a 512 sample FFT window size, resulting in 256 frequency bins given a real-valued input signal). In the latter case, the low-frequency resolution of the mid and side can produce audible spectral artifacts when subtracted from each other to derive the super-mid M1 and super-side S1 components, because the spectral energy of each frequency bin is an average representation of the energy over an excessively large frequency range.In this case, taking the absolute value of the difference between the middle and side when deriving the super-middle M1 or super-side S1 can help mitigate perceptual artifacts by allowing each frequency bin to diverge from the true orthogonality in the components. In addition to or in lieu of wrapping around zero, a coefficient can be applied to the subtraction value, scaling the value between 0 and 1, thus providing a method for interpolating between the following extremes: at one extreme (i.e., a value of 1), complete orthogonality of the super and residual middle / side components; and at the other extreme (i.e., a value of 0), super-middle M1 and super-side S1 that are identical to their corresponding original middle and side components.

[0049] The subtraction unit 215 removes the spectral energy of the middle component 109 in the frequency domain from the spectral energy of the side component 111 in the frequency domain without considering the phase to generate the super-side component S1. For example, the subtraction unit 215 subtracts the size of the middle component 109 in the frequency domain from the size of the side component 111 in the frequency domain without considering the phase to generate the super-side component S1. The subtraction unit 219 removes the spectral energy of the super-side component S1 from the spectral energy of the side component 111 to generate the residual side component S2. For example, the subtraction unit 219 subtracts the size of the super-side component S1 in the frequency domain from the size of the side component 111 in the frequency domain without considering the phase to generate the residual side component S2.

[0050] exist Figure 2B , the orthogonal component generator module 223 is similar to the orthogonal component generator module 213 in that it receives the middle component 109 and the side component 111 and generates a super middle component M1, a residual middle component M2, a super side component S1, and a residual side component S2. The orthogonal component generator module 223 is different from the orthogonal component generator module 213 in that the super middle component M1 and the super side component S1 are generated in the frequency domain and then converted back to the time domain to generate the residual middle component M2 and the residual side component S2. The orthogonal component generator module 223 includes a forward FFT unit 220, a passband unit 222, a subtraction unit 224, a super middle processor 225, an inverse FFT unit 226, a time delay unit 228, a subtraction unit 230, a forward FFT unit 232, a passband unit 234, a subtraction unit 236, a super side processor 237, an inverse FFT unit 240, a time delay unit 242, and a subtraction unit 244.

[0051] The forward fast Fourier transform (FFT) unit 220 applies a forward FFT to the intermediate component 109 to convert the intermediate component 109 to the frequency domain. The converted intermediate component 109 in the frequency domain includes magnitude and phase. The bandpass unit 222 applies a bandpass filter to the frequency domain intermediate component 109, wherein the bandpass filter specifies the frequency in the super intermediate component M1. For example, in order to isolate the typical human vocal range, the bandpass filter can specify the frequency between 300 and 8000 Hz. In another example, in order to remove the audio content associated with the typical human vocal range, the bandpass filter can keep the lower frequencies (e.g., generated by a bass guitar or drum) and the higher frequencies (e.g., generated by a cymbal) in the super intermediate component M1. In other embodiments, in addition to and / or in place of the bandpass filter applied by the bandpass unit 222, the orthogonal component generator module 223 applies various other filters to the frequency domain intermediate component 109. In some embodiments, the orthogonal component generator module 223 does not include the bandpass unit 222 and does not apply any filter to the frequency domain intermediate component 109. In the frequency domain, the subtraction unit 224 subtracts the side component 111 from the filtered middle component to generate the super middle component M1. In other embodiments, in addition to and / or in lieu of the orthogonal component processor module (e.g., Figure 3 The orthogonal component generator module 223 applies various audio enhancements to the frequency domain super-intermediate component M1 after the later processing applied to the super-intermediate component M1 is performed by the orthogonal component processor module 117. The super-intermediate processor 225 performs processing on the super-intermediate component M1 in the frequency domain before it is converted to the time domain. The processing may include sub-band spatial processing and / or crosstalk compensation processing. In some embodiments, instead of and / or in addition to the processing that can be performed by the orthogonal component processor module 117, the super-intermediate processor 225 performs processing on the super-intermediate component M1. The inverse FFT unit 226 applies an inverse FFT to the super-intermediate component M1 to convert the super-intermediate component M1 back to the time domain. The super-intermediate component M1 in the frequency domain includes the magnitude of M1 and the phase of the intermediate component 109, which the inverse FFT unit 226 converts to the time domain. The time delay unit 228 applies a time delay to the intermediate component 109 so that the intermediate component 109 and the super-intermediate component M1 arrive at the subtraction unit 230 at the same time. The subtraction unit 230 subtracts the super middle component M1 in the time domain from the time delayed middle component 109 in the time domain to generate a residual middle component M2. In this example, the spectral energy of the super middle component M1 is removed from the spectral energy of the middle component 109 using processing in the time domain.

[0052] The forward FFT unit 232 applies a forward FFT to the side component 111 to convert the side component 111 to the frequency domain. The converted side component 111 in the frequency domain includes magnitude and phase. The bandpass unit 234 applies a bandpass filter to the frequency domain side component 111. The bandpass filter specifies the frequency in the super side component S1. In other embodiments, in addition to and / or in place of the bandpass filter, the orthogonal component generator module 223 applies various other filters to the frequency domain side component 111. In the frequency domain, the subtraction unit 236 subtracts the intermediate component 109 from the filtered side component 111 to generate the super side component S1. In other embodiments, in addition to and / or in place of the orthogonal component processor (e.g., Figure 3 The orthogonal component processor module of the orthogonal component generator module 223 applies various audio enhancements to the frequency domain super-side component S1. The super-side processor 237 performs processing on the super-side component S1 in the frequency domain before it is converted to the time domain. The processing may include sub-band spatial processing and / or crosstalk compensation processing. In some embodiments, instead of and / or in addition to the processing that can be performed by the orthogonal component processor module 117, the super-side processor 237 performs processing on the super-side component S1. The inverse FFT unit 240 applies an inverse FFT to the super-side component S1 in the frequency domain to generate the super-side component S1 in the time domain. The super-side component S1 in the frequency domain includes the magnitude of S1 and the phase of the side component 111, and the inverse FFT unit 226 converts it to the time domain. The time delay unit 242 time delays the side component 111 so that the side component 111 arrives at the subtraction unit 244 at the same time as the super-side component S1. The subtraction unit 244 then subtracts the super side component S1 in the time domain from the time delayed side component 111 in the time domain to generate a residual side component S2. In this example, the spectral energy of the super side component S1 is removed from the spectral energy of the side component 111 using processing in the time domain.

[0053] In some embodiments, super middle processor 225 and super side processor 237 may be omitted if the processing performed by these components is performed by quadrature component processor module 117 .

[0054] exist Figure 2C, the orthogonal component generator module 245 is similar to the orthogonal component generator module 223 in that it receives the middle component 109 and the side component 111 and generates a super middle component M1, a residual middle component M2, a super side component S1 and a residual side component S2, but the difference is that the orthogonal component generator module 245 generates each of the components M1, M2, S1 and S2 in the frequency domain and then converts these components to the time domain. The orthogonal component generator module 245 includes a forward FFT unit 247, a bandpass unit 249, a subtraction unit 251, a super middle processor 252, a subtraction unit 253, a residual middle processor 254, an inverse FFT unit 255, an inverse FFT unit 257, a forward FFT unit 261, a bandpass unit 263, a subtraction unit 265, a super side processor 266, a subtraction unit 267, a residual side processor 268, an inverse FFT unit 269 and an inverse FFT unit 271.

[0055] The forward FFT unit 247 applies a forward FFT to the intermediate component 109 to convert the intermediate component 109 to the frequency domain. The converted intermediate component 109 in the frequency domain includes magnitude and phase. The forward FFT unit 261 applies a forward FFT to the side component 111 to convert the side component 111 to the frequency domain. The converted side component 111 in the frequency domain includes magnitude and phase. The bandpass unit 249 applies a bandpass filter to the frequency domain intermediate component 109, and the bandpass filter specifies the frequency of the super intermediate component M1. In some embodiments, in addition to and / or in place of the bandpass filter, the orthogonal component generator module 245 applies various other filters to the frequency domain intermediate component 109. The subtraction unit 251 subtracts the frequency domain side component 111 from the frequency domain intermediate component 109 to generate the super intermediate component M1 in the frequency domain. The super intermediate processor 252 performs processing on the super intermediate component M1 in the frequency domain before converting it to the time domain. In some embodiments, the super intermediate processor 252 performs sub-band spatial processing and / or crosstalk compensation processing. In some embodiments, the super intermediate processor 252 performs processing on the super intermediate component M1 instead of and / or in addition to the processing that can be performed by the orthogonal component processor module 117. The inverse FFT unit 257 applies an inverse FFT to the super intermediate component M1 to convert it back to the time domain. The super intermediate component M1 in the frequency domain includes the magnitude of M1 and the phase of the intermediate component 109, which the inverse FFT unit 257 converts to the time domain. The subtraction unit 253 subtracts the super intermediate component M1 from the intermediate component 109 in the frequency domain to generate a residual intermediate component M2. The residual intermediate processor 254 performs processing on the residual intermediate component M2 in the frequency domain before converting it to the time domain. In some embodiments, the residual intermediate processor 254 performs sub-band spatial processing and / or crosstalk compensation processing on the residual intermediate component M2. In some embodiments, the residual intermediate processor 254 performs processing on the residual intermediate component M2 instead of and / or in addition to the processing that can be performed by the orthogonal component processor module 117. The inverse FFT unit 255 applies an inverse FFT to convert the residual intermediate component M2 to the time domain. The residual intermediate component M2 in the frequency domain includes the magnitude of M2 and the phase of the intermediate component 109, which is converted into the time domain by the inverse FFT unit 255.

[0056] The bandpass unit 263 applies a bandpass filter to the frequency domain side component 111. The bandpass filter specifies the frequencies in the super side component S1. In other embodiments, in addition to and / or in place of the bandpass filter, the orthogonal component generator module 245 applies various other filters to the frequency domain side component 111. In the frequency domain, the subtraction unit 265 subtracts the intermediate component 109 from the filtered side component 111 to generate the super side component S1. The super side processor 266 performs processing on the super side component S1 in the frequency domain before converting it to the time domain. In some embodiments, the super side processor 266 performs sub-band spatial processing and / or crosstalk compensation processing on the super side component S1. In some embodiments, instead of and / or in addition to the processing that can be performed by the orthogonal component processor module 117, the super side processor 266 performs processing on the super side component S1. The inverse FFT unit 271 applies an inverse FFT to convert the super side component S1 back to the time domain. The super side component S1 in the frequency domain includes the magnitude of S1 and the phase of the side component 111, and the inverse FFT unit 271 converts it to the time domain. The subtraction unit 267 subtracts the super side component S1 from the side component 111 in the frequency domain to generate a residual side component S2. The residual side processor 268 performs processing on the residual side component S2 in the frequency domain before converting it to the time domain. In some embodiments, the residual side processor 268 performs subband spatial processing and / or crosstalk compensation processing on the residual side component S2. In some embodiments, the residual side processor 268 performs processing on the residual side component S2 instead of and / or in addition to the processing that can be performed by the orthogonal component processor module 117. The inverse FFT unit 269 applies an inverse FFT to the residual side component S2 to convert it to the time domain. The residual side component S2 in the frequency domain includes the magnitude of S2 and the phase of the side component 111, and the inverse FFT unit 269 converts it to the time domain.

[0057] In some embodiments, if the processing performed by the super intermediate processor 252, the super side processor 266, the residual intermediate processor 254, or the residual side processor 268 is performed by the orthogonal component processor module 117, these components may be omitted.

[0058] Example Quadrature Component Processor

[0059] Figure 33 is a block diagram of a quadrature component processor module 317 according to one or more embodiments. The quadrature component processor module 317 is an example of the quadrature component processor module 117. The quadrature component processor module 317 may include a subband spatial processing and / or crosstalk compensation processing unit 320, an addition unit 325, and an addition unit 330. The quadrature component processor module 317 performs subband spatial processing and / or crosstalk compensation processing on at least one of the super middle component M1, the residual middle component M2, the super side component S1, and the residual side component S2. As a result of the subband spatial processing and / or crosstalk compensation processing 320, the quadrature component processor module 317 outputs at least one of the processed M1, the processed M2, the processed S1, and the processed S2. The addition unit 325 adds the processed M1 and the processed M2 to generate the processed middle component 131, and the addition unit 330 adds the processed S1 and the processed S2 to generate the processed side component 139.

[0060] In some embodiments, the quadrature component processor module 317 performs sub-band spatial processing and / or crosstalk compensation processing 320 on at least one of the super middle component M1, the residual middle component M2, the super side component S1, and the residual side component S2 in the frequency domain to generate the processed middle component 131 and the processed side component 139 in the frequency domain. The quadrature component generator module 113 may provide the component M1, M2, S1, or S2 in the frequency domain to the quadrature component processor, where an inverse FFT is performed. After generating the processed middle component 131 and the processed side component 139, the quadrature component processor module 317 may perform an inverse FFT on the processed middle component 131 and the processed side component 139 to convert the components back to the time domain. In some embodiments, the quadrature component processor module 317 performs an inverse FFT on the processed M1, the processed M2, the processed S1, and the processed S1 to generate the processed middle component 131 and the processed side component 139 in the time domain.

[0061] An example of a quadrature component processor module 317 is shown in Figure 4 and Figure 5. In some embodiments, the quadrature component processor module 317 performs sub-band spatial processing and crosstalk compensation processing. The processing performed by the quadrature component processor module 317 is not limited to sub-band spatial processing or crosstalk compensation processing. Any type of spatial processing using mid / side space can be performed by the quadrature component processor module 317, such as by using a super-mid component instead of a mid component or using a super-side component instead of a side component. Some other types of processing may include gain application, amplitude or delay-based panning, binaural processing, reverberation, dynamic range processing (such as compression and limiting), and other linear or nonlinear audio processing techniques and effects, ranging from chorus or flanging to machine learning-based vocal or instrumental style transfer, conversion or resynthesis.

[0062] Example Subband Spatial Processor

[0063] Figure 4 4 is a block diagram of a subband spatial processor module 410 according to one or more embodiments. The subband spatial processor module 410 is an example of the quadrature component processor module 317. The subband spatial processor module 410 includes a middle EQ filter 404(1), a middle EQ filter 404(2), a middle EQ filter 404(3), a middle EQ filter 404(4), a side EQ filter 406(1), a side EQ filter 406(2), a side EQ filter 406(3), and a side EQ filter 406(4). In some embodiments, the subband spatial processor module 410 includes other components in addition to and / or in place of the components described herein.

[0064] The sub-band spatial processor module 410 receives the non-spatial component Y m and the spatial component Y s And the subbands of one or more of these components are gain-adjusted to provide spatial enhancement. Non-spatial component Y m It can be the super middle component M1 or the residual middle component M2. Spatial component Y s It can be the super side component S1 or the residual side component S2.

[0065] The sub-band spatial processor module 410 receives the non-spatial component Y m and applying intermediate EQ filters 404(1) to 404(4) to Y m to generate the enhanced non-spatial component E m The subband spatial processor module 410 also receives the spatial component Y s and applying side EQ filters 406(1) to 406(4) to Y s to generate enhanced spatial components E sThe subband filters may include various combinations of peak filters, notch filters, low pass filters, high pass filters, low shelf filters, high shelf filters, band pass filters, band stop filters, and / or all pass filters. The subband filters may also apply gains to the corresponding subbands. More specifically, the subband spatial processor module 410 includes a processor for the non-spatial component Y m The subband filter for each of the n frequency subbands and for the spatial component Y s For example, for n=4 subbands, the subband spatial processor module 410 includes a subband filter for each of the n subbands of Y. m , including an intermediate equalization (EQ) filter 404(1) for subband (1), an intermediate EQ filter 404(2) for subband (2), an intermediate EQ filter 404(3) for subband (3), and an intermediate EQ filter 404(4) for subband (4). Each intermediate EQ filter 404 applies a filter to the non-spatial component Y m The frequency subband part of the enhanced non-spatial component E m .

[0066] The sub-band spatial processor module 410 also includes a processor for the spatial component Y s A series of subband filters for frequency subbands of Y include a side equalization (EQ) filter 406 (1) for subband (1), a side EQ filter 406 (2) for subband (2), a side EQ filter 406 (3) for subband (3), and a side EQ filter 406 (4) for subband (4). Each side EQ filter 406 applies a filter to the spatial component Y s The frequency subband part of the enhanced spatial component E s .

[0067] Non-spatial component Y m and the spatial component Y s Each of the n frequency subbands may correspond to a range of frequencies. For example, frequency subband (1) may correspond to 0 to 300 Hz, frequency subband (2) may correspond to 300 to 510 Hz, frequency subband (3) may correspond to 510 to 2700 Hz, and frequency subband (4) may correspond to 2700 Hz to the Nyquist frequency. In some embodiments, the n frequency subbands are a combined set of critical bands. The critical bands may be determined using a corpus of audio samples from a variety of music genres. The long-term average energy ratio of the mid-to-side components over 24 Bark scale critical bands is determined from the samples. Continuous frequency bands with similar long-term average ratios are then combined together to form the set of critical bands. The range of the frequency subbands and the number of frequency subbands may be adjustable.

[0068] In some embodiments, the sub-band spatial processor module 410 processes the residual intermediate component M2 into a non-spatial component Y m , and use one of the side component, super side component S1 or residual side component S2 as the spatial component Y s .

[0069] In some embodiments, the sub-band spatial processor module 410 processes one or more of the super-mid-component M1, the super-side component S1, the residual mid-component M2, and the residual side component S2. The filters applied to the sub-bands of each of these components may be different. The super-mid-component M1 and the residual mid-component M2 may each be processed as for the non-spatial component Y m The super-side component S1 and the residual-side component S2 can each be processed as for the spatial component Y s Processed as discussed.

[0070] Example Crosstalk Compensation Processor

[0071] Figure 5 5 is a block diagram of a crosstalk compensation processor module 510 according to one or more embodiments. The crosstalk compensation processor module 510 is an example of the orthogonal component processor module 317. The crosstalk compensation processor module 510 includes a middle component processor 520 and a side component processor 530. The crosstalk compensation processor module 510 receives the non-spatial component Y m and the spatial component Y s , and a filter is applied to one or more of these components to compensate for spectral defects caused by (e.g., subsequent or previous) crosstalk processing. Non-spatial component Y m It can be the super middle component M1 or the residual middle component M2. Spatial component Y s It can be the super side component S1 or the residual side component S2.

[0072] The crosstalk compensation processor module 510 receives the non-spatial component Y m And the intermediate component processor 520 applies a set of filters to generate an enhanced non-spatial crosstalk compensation component Z m The crosstalk compensation processor module 510 also receives the spatial sub-band components Y s , and a set of filters are applied in the side component processor 530 to generate enhanced spatial subband components E s The intermediate component processor 520 includes a plurality of filters 540, such as m intermediate filters 540(a), 540(b) to 540(m). Here, each of the m intermediate filters 540 processes the non-spatial component X. m The intermediate component processor 520 accordingly processes the non-spatial component X mTo generate the intermediate crosstalk compensation channel Z m In some embodiments, the intermediate filter 540 uses a non-spatial X m The frequency response graph is configured and the crosstalk processing is performed by simulation. In addition, by analyzing the frequency response graph, any spectral defects that appear as artifacts of the crosstalk processing, such as peaks or valleys in the frequency response graph that exceed a predetermined threshold (e.g., 10dB), can be estimated. These artifacts are mainly the result of the delayed and possibly inverted contra-side signal being added to its corresponding ipsi-side signal in the crosstalk processing, effectively introducing a comb filter-like frequency response to the final rendering result. Intermediate crosstalk compensation channel Z m Can be generated by the intermediate component processor 520 to compensate for the estimated peaks or valleys, where each of the m frequency bands corresponds to a peak or valley. Specifically, based on the specific delay, filter frequency and gain applied in the crosstalk processing, the peaks and valleys move up and down in the frequency response, resulting in variable amplification and / or attenuation of energy in specific areas of the spectrum. Each of the intermediate filters 540 can be configured to adjust for one or more of the peaks and valleys.

[0073] The side component processor 530 includes a plurality of filters 550, such as m side filters 550(a), 550(b) to 550(m). The side component processor 530 processes the spatial component X s To generate the side crosstalk compensation channel Z s In some embodiments, the space X with crosstalk processing can be obtained by simulation s By analyzing the frequency response graph, any spectral defects that appear as artifacts of the crosstalk processing, such as peaks or valleys in the frequency response graph that exceed a predetermined threshold (e.g., 10 dB), can be estimated. s Can be generated by the side component processor 530 to compensate for the estimated peak or valley. Specifically, based on the specific delay, filter frequency and gain applied in the crosstalk processing, the peaks and valleys move up and down in the frequency response, resulting in variable amplification and / or attenuation of energy in a specific area of ​​the spectrum. Each of the side filters 550 can be configured to adjust for one or more of the peaks and valleys. In some embodiments, the intermediate component processor 520 and the side component processor 530 can include different numbers of filters.

[0074] In some embodiments, middle filter 540 and side filter 550 may include biquad filters having a transfer function defined by Equation 1:

[0075]

[0076] Where z is a complex variable and a0, a1, a2, b0, b1, and b2 are digital filter coefficients. One way to implement this filter is the direct form I topology defined by Equation 2:

[0077]

[0078] Where X is the input vector and Y is the output. Other topologies can be used, depending on their maximum word length and saturation behavior. A second-order filter with real-valued input and output can then be implemented using a biquad. To design a discrete-time filter, a continuous-time filter is designed and then transformed to discrete time via a bilinear transform. Additionally, frequency warping can be used to compensate for the resulting shift in center frequency and bandwidth.

[0079] For example, a peaking filter may have an S-plane transfer function defined by Equation 3:

[0080]

[0081] Where s is a complex variable, A is the amplitude of the peak, and Q is the filter "quality", the digital filter coefficients are defined by:

[0082] b0=1+αA

[0083] b1=-2*cos(ω0)

[0084] b2=1-αA

[0085]

[0086] a1=-2cos(ω0)

[0087]

[0088] where ω0 is the center frequency of the filter in radians and Furthermore, the filter quality Q can be defined by Equation 4:

[0089]

[0090] where Δf is the bandwidth and f c is the center frequency. The middle filter 540 is shown in series and the side filter 550 is shown in series. In some embodiments, the middle filter 540 is applied in parallel to the middle component X m , and the side filter 540 is applied in parallel to the side component X s .

[0091] In some embodiments, the crosstalk compensation processor module 510 processes each of the super middle component M1, the super side component S1, the residual middle component M2, and the residual side component S2. The filters applied to each of these components may be different.

[0092] Example Crosstalk Processor

[0093] Figure 6 is a block diagram of a crosstalk simulation processor module 600 according to one or more embodiments. Figure 1 As described, in some embodiments, the audio processing system 100 includes a crosstalk processor module 141, which applies crosstalk processing to the processed left component 151 and the processed right component 159. The crosstalk processing includes, for example, crosstalk simulation and crosstalk cancellation. In some embodiments, the crosstalk processor module 141 includes a crosstalk simulation processor module 600. The crosstalk simulation processor module 600 generates a contralateral sound component for output to a stereo headset, thereby providing a speaker-like listening experience on the headset. The left input channel X L It can be the processed left component 151, the right input channel X R This may be the processed right component 159. In some embodiments, crosstalk simulation may be performed prior to quadrature component processing.

[0094] The crosstalk simulation processor module 600 includes a left head shadow low pass filter 602, a left head shadow high pass filter 624, a left crosstalk delay 604, and a left head shadow gain 610 to process the left input channel X. L The crosstalk simulation processor module 600 also includes a right head shadow low pass filter 606, a right head shadow high pass filter 626, a right crosstalk delay 608 and a right head shadow gain 612 to process the right input channel X R The left head shadow low pass filter 602 and the left head shadow high pass filter 624 are used to filter the left input channel X. L A modulation is applied which simulates the frequency response of the signal after passing through the listener's head. The output of the left head shadow high pass filter 624 is provided to the left crosstalk delay 604 which applies a time delay. The time delay represents the transmural distance that the contralateral sound component has traveled relative to the ipsilateral sound component. The left head shadow gain 610 applies a gain to the output of the left crosstalk delay 604 to generate the right and left analog channels W. L .

[0095] Similarly, for the right input channel X R , the right head shadow low pass filter 606 and the right head shadow high pass filter 626 are applied to the right input channel X RA modulation is applied that simulates the frequency response of the listener's head. The output of the right head shadow high pass filter 626 is provided to the right crosstalk delay 608, which applies a time delay. The right head shadow gain 612 applies a gain to the output of the right crosstalk delay 608 to generate a right crosstalk simulated channel W R .

[0096] Applying the head shadow low pass filter, head shadow high pass filter, crosstalk delay, and head shadow gain to each of the left and right channels may be performed in a different order.

[0097] Figure 7 1 is a block diagram of a crosstalk cancellation processor module 700 according to one or more embodiments. The crosstalk processor module 141 may include the crosstalk cancellation processor module 700. The crosstalk cancellation processor module 700 receives the left input channel X L and right input channel X R , and for channel X L , X R Crosstalk cancellation is performed to generate the left output channel O L and right output channel O R . Left input channel X L It can be the processed left component 151, and the right input channel X R This may be the processed right component 159. In some embodiments, crosstalk cancellation may be performed prior to quadrature component processing.

[0098] The crosstalk cancellation processor module 700 includes an in-band and out-band divider 710, inverters 720 and 722, opposite-side estimators 730 and 740, combiners 750 and 752, and an in-band and out-band combiner 760. These components operate together to divide the input channels T L 、T R The output channel is divided into an in-band component and an out-of-band component, and crosstalk cancellation is performed on the in-band component to generate an output channel O L , O R .

[0099] By dividing the input audio signal T into different frequency band components and by performing crosstalk cancellation on selective components (e.g., in-band components), crosstalk cancellation can be performed for specific frequency bands while avoiding degradation in other frequency bands. If crosstalk cancellation is performed without dividing the input audio signal T into different frequency bands, the audio signal after such crosstalk cancellation exhibits significant attenuation or amplification of non-spatial and spatial components in low frequencies (e.g., below 350 Hz), high frequencies (e.g., above 12,000 Hz), or both. By performing crosstalk cancellation in the in-band where most of the influential spatial cues are located (e.g., between 250 Hz and 14,000 Hz), balanced overall energy can be retained throughout the entire frequency spectrum of the mixture, especially in the non-spatial components.

[0100] The in-band and out-band divider 710 divides the input channel T L , T R Separate into in-band channels T L,In , T R,In and out-of-band channel T L,Out , T R,Out Specifically, the in-band and out-band divider 710 divides the left enhanced compensation channel T L Divided into left in-band channel T L,In and left out-band channel T L,Out Similarly, the in-band and out-band divider 710 divides the right enhanced compensation channel T R Separated into right in-band channel T R,In and right out-of-band channel T R,Out Each in-band channel may contain a portion of the corresponding input channel corresponding to a frequency range, which may include, for example, 250 Hz to 14 kHz. The frequency band range may be adjustable, for example according to loudspeaker parameters.

[0101] The inverter 720 and the opposite side estimator 730 operate together to generate the left opposite side cancellation component S L , to compensate for the left in-band channel T L,In Similarly, the inverter 722 and the opposite side estimator 740 operate together to generate the right opposite side cancellation component S R , to compensate for the right in-band channel T R,In The contralateral sound component caused by

[0102] In one approach, the inverter 720 receives the in-band channel T L,In And the received in-band channel T L,In The polarity of the channel is reversed to generate an inverted in-band channel T L,In '. The contralateral estimator 730 receives the inverted in-band channel T L,In ', and extract the in-band channel T corresponding to the opposite side sound component by filtering L,In ' part. Because the filtering is done on the in-band channel T L,In ' is performed, so the portion extracted by the contralateral estimator 730 becomes the in-band channel T attributed to the contralateral sound component L,In Therefore, the part extracted by the opposite side estimator 730 becomes the left opposite side cancellation component S L , which can be added to the corresponding in-band channel T R,In To reduce the in-band channel T L,In In some embodiments, the inverter 720 and the contralateral estimator 730 are implemented in different orders.

[0103] The inverter 722 and the contralateral estimator 740 are used to calculate the in-band channel T R,In A similar operation is performed to generate the right contralateral cancellation component S R Therefore, for the sake of brevity, its detailed description is omitted here.

[0104] In one example implementation, the contralateral estimator 730 includes a filter 732, an amplifier 734, and a delay unit 736. The filter 732 receives an inverted input channel T L,In ', and extract the inverse in-band channel T corresponding to the contralateral sound component through the filter function L,In ' part. Example filter implementations are notch or shelf filters with center frequencies selected between 5000 and 10000 Hz and Qs selected between 0.5 and 1.0. The gain (G in decibels) dB ) can be derived from Equation 5:

[0105] G dB =-3.0-log 1.333 (D) Equation (5)

[0106] Where D is the amount of delay in samples by delay units 736 and 646, for example, at a sampling rate of 48 kHz. An alternative implementation is a low pass filter where the corner frequency is selected between 5000 and 10000 Hz and Q is selected between 0.5 and 1.0. In addition, amplifier 734 amplifies the extracted portion by a corresponding gain factor G L,In , and the delay unit 736 delays the amplified output of the amplifier 734 according to the delay function D to generate the left opposite side cancellation component S L The contralateral estimator 740 includes a filter 742, an amplifier 744, and a delay unit 746. The delay unit 746 performs a reverse phase operation on the in-band channel T. R,In 'Perform similar operations to generate the right contralateral cancellation component S R In one example, the contralateral estimators 730, 740 generate the left contralateral cancellation component S according to the following equation: L and the right contralateral elimination component S R :

[0107] S L =D[G L,In *F[T L,In ']] Equation (6)

[0108] S R =D[G R,In *F[T R,In ']] Equation (7)

[0109] Where F[] is the filter function and D[] is the delay function.

[0110] The configuration of the crosstalk cancellation can be determined by the speaker parameters. In one example, the filter center frequency, delay amount, amplifier gain, and filter gain can be determined based on the angle between the two speakers relative to the listener. In some embodiments, the values ​​between the speaker angles are used to interpolate other values.

[0111] The combiner 750 cancels the right-side component S R Combined to the left in-band channel T L,In To generate the left in-band crosstalk channel U L , and the combiner 752 cancels the left-side component S L Combined to right in-band channel T R,In To generate the right in-band crosstalk channel U R The in-band and out-band combiner 760 combines the left in-band crosstalk channel U L With out-of-band channel T L,Out Combined to generate the left output channel O L , and the right in-band crosstalk channel U R With out-of-band channel T R,Out Combined to generate the right output channel O R .

[0112] Therefore, the left output channel O L Includes in-band channel T attributed to the measured sound R The opposite side of the corresponding right side cancellation component S R , and the right output channel O R Includes in-band channel T attributed to the contralateral sound L,In The inverted phase of a part of the corresponding left side cancellation component S L In this configuration, the sound reaching the right ear is output by the right speaker according to the right output channel O. R The wavefront of the sound component on the same side of the output can be cancelled by the left speaker according to the left output channel O L Similarly, the wavefront of the contralateral sound component reaching the left ear is output by the left speaker according to the left output channel O L The wavefront of the sound component on the same side of the output can be cancelled by the right speaker according to the right output channel O R The wavefront of the contralateral sound component is outputted. Therefore, the contralateral sound component can be reduced to enhance spatial detectability.

[0113] Orthogonal component space processing

[0114] Figure 8is a flow chart of a process for spatial processing using at least one of a super-middle, residual-middle, super-side, or residual-side component according to one or more embodiments. Spatial processing may include gain application, amplitude or delay-based panning, binaural processing, reverberation, dynamic range processing (such as compression and limiting), linear or nonlinear audio processing techniques and effects, chorus effects, flanger effects, machine learning-based vocal or instrumental style transfer, conversion or resynthesis, and other methods. The process may be performed to provide spatially enhanced audio to a user's device. The process may include fewer or more steps, and the steps may be performed in a different order.

[0115] An audio processing system (e.g., audio processing system 100) receives 810 an input audio signal (e.g., left input channel 103 and right input channel 105). In some embodiments, the input audio signal may be a multi-channel audio signal including a plurality of left and right channel pairs. For left and right input channels, each left and right channel pair may be processed as discussed herein.

[0116] The audio processing system generates 820 a non-spatial mid component (e.g., mid component 109) and a spatial side component (e.g., side component 111) from an input audio signal. In some embodiments, an L / R to M / S converter (e.g., L / R to M / S converter module 107) performs conversion of the input audio signal into mid and side components.

[0117] The audio processing system generates 830 at least one of a super-middle component (e.g., a super-middle component M1), a super-side component (e.g., a super-side component S1), a residual middle component (e.g., a residual middle component M2), and a residual side component (e.g., a residual side component S2). The audio processing system may generate at least one component and / or all of the components listed above. The super-middle component includes removing the spectral energy of the side component from the spectral energy of the middle component. The residual middle component includes removing the spectral energy of the super-middle component from the spectral energy of the middle component. The super-side component includes removing the spectral energy of the middle component from the spectral energy of the side component. The residual side component includes removing the spectral energy of the super-side component from the spectral energy of the side component. The processing for generating M1, M2, S1, or S2 may be performed in the frequency domain or the time domain.

[0118] The audio processing system filters 840 at least one of the super-mid component, the residual mid component, the super-side component, and the residual side component to enhance the audio signal. The filtering may include spatial cue processing, such as by adjusting the frequency-dependent amplitude or frequency-dependent delay of the super-mid component, the residual mid component, the super-side component, or the residual side component. Some examples of spatial cue processing include amplitude- or delay-based panning or binaural processing.

[0119] Filtering may include dynamic range processing, such as compression or limiting. For example, when a threshold level for compression is exceeded, the super middle component, residual middle component, super side component, or residual side component may be compressed according to a compression ratio. In another example, when a threshold level for limiting is exceeded, the super middle component, residual middle component, super side component, or residual side component may be limited to a maximum level.

[0120] The filtering may include machine learning based changes to the super-mid component, the residual mid component, the super-side component, or the residual side component. Some examples include machine learning based vocal or instrumental style transfer, conversion, or resynthesis.

[0121] The filtering of the super-mid component, residual mid component, super-side component or residual side component may include the application of gain, reverberation, and other linear or non-linear audio processing techniques and effects (chorus and / or flanger) or other types of processing. In some embodiments, the filtering may include filtering for sub-band spatial processing and crosstalk compensation, as described below in conjunction with Fig. 9 discussed in more detail.

[0122] Filtering can be performed in the frequency domain or the time domain. In some embodiments, the mid and side components are converted from the time domain to the frequency domain, super and / or residual components are generated in the frequency domain, filtering is performed in the frequency domain, and the filtered components are converted to the time domain. In other embodiments, the super and / or residual components are converted to the time domain, and filtering is performed on these components in the time domain.

[0123] The audio processing system generates 850 a left output channel (e.g., left output channel 121) and a right output channel (e.g., right output channel 123) using one or more of the filtered super / residual components. For example, the conversion from M / S to L / R may be performed using a mid component (e.g., processed mid component 131) or a side component (e.g., processed side component 139) generated from at least one of the filtered super mid component, the filtered residual mid component, the filtered super side component, or the filtered residual side component. In another example, the filtered super mid component or the filtered residual mid component may be used as a mid component for the M / S to L / R conversion, or the filtered super side component or the residual side component may be used as a side component for the M / S to L / R conversion.

[0124] Orthogonal component sub-band space and crosstalk processing

[0125] Fig. 9The present invention is a flowchart of a process for performing sub-band spatial processing and crosstalk compensation processing using at least one of a super-mid component, a residual mid component, a super-side component, or a residual side component according to one or more embodiments. The crosstalk processing may include crosstalk cancellation or crosstalk simulation. The sub-band spatial processing may be performed to provide audio content with enhanced spatial detectability, such as by creating a sense that the sound is directed to the listener from a large area rather than a specific point in the space corresponding to the speaker location (e.g., sound field enhancement), thereby bringing a more immersive listening experience to the listener. The crosstalk simulation may be used for the audio output of the headphones to simulate the speaker experience with contralateral crosstalk. The crosstalk cancellation may be used for the audio output to the speaker to eliminate the effects of crosstalk interference. The crosstalk compensation may compensate for the spectral defects caused by the crosstalk cancellation or crosstalk simulation. The process may include fewer or more steps, and the steps may be performed in different orders. The super and residual mid / side components may be manipulated in different ways for different purposes. For example, in the case of crosstalk compensation, targeted subband filtering may be applied only to the super-mid component M1 (where most of the vocal dialogue energy in much movie content occurs) in an effort to eliminate spectral artifacts resulting from the crosstalk processing in only that component. In the case of sound field enhancement with or without crosstalk processing, targeted subband gains may be applied to the residual mid component M2 and the residual side component S2. For example, the residual mid component M2 may be attenuated and the residual side component S2 may be inversely amplified to increase the distance between these components from a gain perspective (which may increase spatial detectability if done well) without producing a drastic overall change in perceived loudness in the final L / R signal, while also avoiding attenuation of the super-mid M1 component (e.g., the portion of the signal that typically contains most of the vocal energy).

[0126] The audio processing system receives 910 an input audio signal, the input audio signal comprising a left channel and a right channel. In some embodiments, the input audio signal may be a multi-channel audio signal comprising a plurality of left and right channel pairs. For the left and right input channels, each left and right channel pair may be processed as discussed herein.

[0127] The audio processing system applies 920 crosstalk processing to the received input audio signal. The crosstalk processing includes at least one of crosstalk simulation and crosstalk cancellation.

[0128] In steps 930 to 960, the audio processing system performs crosstalk compensation of subband spatial processing and crosstalk processing using one or more of the super mid, super side, residual mid or residual side components. In some embodiments, crosstalk processing may be performed after processing in steps 930 to 960.

[0129] The audio processing system generates 930 a mid component and a side component from the (eg, crosstalk processed) audio signal.

[0130] The audio processing system generates 940 at least one of a super mid component, a residual mid component, a super side component, and a residual side component. The audio processing system may generate at least one and / or all of the components listed above.

[0131] The audio processing system filters 950 subbands of at least one of the super-mid component, the residual mid component, the super-side component, and the residual side component to apply subband spatial processing to the audio signal. Each subband can include a range of frequencies, such as can be defined by a set of critical bands. In some embodiments, the subband spatial processing also includes time delaying subbands of at least one of the super-mid component, the residual mid component, the super-side component, and the residual side component.

[0132] The audio processing system filters 960 at least one of the super-mid component, the residual mid component, the super-side component, and the residual side component to compensate for spectral defects from the crosstalk processing of the input audio signal. The spectral defects may include peaks or valleys in a frequency response graph of the super-mid component, the residual mid component, the super-side component, or the residual side component that exceed a predetermined threshold (e.g., 10 dB) that appear as artifacts of the crosstalk processing. The spectral defects may be estimated spectral defects.

[0133] In some embodiments, the filtering of spectral quadrature components for sub-band spatial processing in step 950 and the crosstalk compensation in step 960 may be integrated into a single filtering operation for each spectral quadrature component selected for filtering.

[0134] In some embodiments, filtering of super / residual mid / side components for sub-band spatial processing or crosstalk compensation may be performed in conjunction with filtering for other purposes, such as gain application, amplitude or delay based panning, binaural processing, reverberation, dynamic range processing (such as compression and limiting), linear or non-linear audio processing techniques and effects ranging from chorus and / or flanging, machine learning based vocal or instrumental style transfer, conversion or resynthesis methods, or other types of processing using any of the super-middle component, residual mid component, super-side component, and residual side component.

[0135] Filtering can be performed in the frequency domain or the time domain. In some embodiments, the mid and side components are converted from the time domain to the frequency domain, super and / or residual components are generated in the frequency domain, filtering is performed in the frequency domain, and the filtered components are converted to the time domain. In other embodiments, the super and / or residual components are converted to the time domain, and filtering is performed on these components in the time domain.

[0136] The audio processing system generates 970 left and right output channels from the filtered super-mid component.In some embodiments, the left and right output channels are additionally based on at least one of the filtered residual mid component, the filtered super-side component, and the filtered residual side component.

[0137] Example Quadrature Component Audio Processing

[0138] Figure 10-Figure 19 is a graph depicting spectral energies of mid and side components of an example white noise signal in accordance with one or more embodiments.

[0139] Fig.10 A diagram of a white noise signal 1000 panned to the hard left is shown. The left and right white noise signals are converted to a middle component 1005 and a side component 1010 using a constant power sine / cosine panning law and panned to the hard left. When the white noise signal is panned to the hard left 1000, a user located between the left and right speaker pairs will perceive the sound as occurring at and / or around the left speaker. The white noise signal (split into a left input channel and a right input channel of the white noise signal) can be converted to a middle component 1005 and a side component 1010 using an L / R to M / S converter module 107. Fig.10 As shown, when the white noise signal is panned to the far left 1000, the middle component 1005 and the side component 1010 have approximately equal energy. Similarly, when the white noise signal is panned to the far right ( Fig.10 (not shown), the middle component and the side components will have approximately equal energy.

[0140] Fig.11 A graph of a white noise signal panned to center left 1100 is shown. When the white noise signal is panned to center left 1100 using the common constant power sine / cosine panning law, a user positioned between the left and right speaker pair will perceive the sound to occur midway between the user's front and the left speaker. Fig.11 Depicted are the middle component 1105 and the side component 1110 of the white noise signal 1100 panned to the center left, and the white noise signal 1000 panned to the far left. Compared to the white noise signal 1000 panned to the far left, the middle component 1105 increases by about 3dB, while the side component 1110 decreases by about 6dB. When the white noise signal is panned to the center right, the middle component 1105 and the side component 1110 will have the same Fig.11 Similar energies as shown.

[0141] Fig.12 1 shows a graph of a white noise signal 1200 panned to the center. When the white noise signal is panned to the center 1200 using the common constant power sine / cosine panning law, a user positioned between the left and right speaker pair will perceive the sound as appearing in front of the user (e.g., between the left and right speakers). Fig.12 As shown, the white noise signal 1200 shifted to the center has only the center component 1205 .

[0142] from Fig.10 , Fig.11 and Fig.12 In the above example, it can be seen that although for Fig.12 The sound panned to the center is shown, the center component contains the only energy in the signal (i.e., the same for both left and right channels), where the sound in the original L / R stream is usually perceived as off-center, as shown in Fig.10 and Fig.11 As shown (i.e., a sound with the center panned to the left or right), there is also intermediate component energy.

[0143] It is worth noting that the above three scenarios, which represent the vast majority of L / R audio use cases, do not include scenarios where the side contains unique energy. This situation only occurs when the left and right channels are 180 degrees out of phase (i.e., inverted in sign), which is rare in two-channel audio used for music and entertainment. Therefore, while the mid component is ubiquitous in almost all two-channel left / right audio streams and also includes unique energy panned into center content, the side component is present in everything except panned into center content and is rarely, if ever, the only energy in the signal.

[0144] Quadrature component processing isolates and operates on portions of the mid and side components that are spectrally "orthogonal" to each other. That is, using quadrature component processing, a portion of the mid component corresponding only to energy present in the center of the sound field (i.e., the super-mid component) can be isolated, and likewise a portion of the side component corresponding only to energy not present in the center of the sound field (i.e., the super-side component) can be isolated. Conceptually, the super-mid component is the energy corresponding to the thin column of sound perceived at the center of the sound field, both for speakers and headphones. Furthermore, using a simple scalar, the degree of "thinness" of this column can be controlled to provide interpolation space from super-mid to mid and from super-side to side. Furthermore, as a byproduct of deriving our super-mid / side component signals, operations can also be performed on the residual signal (e.g., residual mid and side components), which is combined with the super-mid / super-side components to form the original complete mid and side components. Each of these four sub-components of the mid and side can be processed independently through a variety of operations, from simple gain staging to multi-band equalizers to custom and special effects.

[0145] Figures 13 to 19 The quadrature component processing of a white noise signal is shown. Fig.13A graph of a white noise signal 1305 shifted to the center and band-passed between 20 and 100 Hz (e.g., using an 8th order Butterworth filter) and a white noise signal 1310 shifted to the far left and band-passed between 5000 and 10000 Hz (e.g., using an 8th order Butterworth filter) is shown, and no quadrature component processing is performed. The graph depicts a mid-component 1315 and a side component 1320 of each of the shifted white noise signals 1305 and 1310. The white noise signal 1305 shifted to the center has energy only in its mid-component 1315, while the white noise signal shifted to the far left has an equal amount of energy in its mid-component 1315 and side components 1320. This is similar to Fig.10 and Fig.12 Results shown.

[0146] Fig.14 Shows Fig.13 1305 and 1310, where the energy of the side component 1320 has been removed. The low-band white noise signal 1305, which was panned to the center, has not changed. The high-band white noise signal 1310, which was panned to the far left, now has zero side energy, while a portion of the energy represented by the middle component 1315 is still present. Even with the side energy removed, there is still energy in the middle signal that is not panned to the center, as shown in signal 1310.

[0147] Fig.15 shows the use of orthogonal component processing Fig.13 1500. Specifically, quadrature component processing is used to isolate the super-middle component 1510 and remove the other energy of the audio signal. Here, the signal panned to the far left is removed, leaving only the signal panned to the middle 1500. This shows that the super-middle component 1510 isolates only the energy in the signal that occupies the very center of the sound field, and nothing else.

[0148] Because the super-mid components of the audio signal can be isolated, the audio signal can be manipulated to control which elements of the original signal end up in the various M1 / M2 / S1 / S2 components. Such pre-processing operations can range from simple amplitude and delay adjustments to more complex filtering techniques. These pre-processing operations can then be subsequently inverted to restore the original sound field.

[0149] Fig.16 shows the use of orthogonal component processing Fig.13 Another embodiment of a panned white noise signal. The L / R audio signal is rotated in such a way that the high-frequency band white noise (e.g., Fig.13 1310 in the center of the sound field and panning the low-frequency noise to the center (e.g., Fig.13The white noise signal 1600, which is initially shifted to the far left and band-passed between 5000 and 10000 Hz, can then be extracted and further processed by isolating the super-middle component 1610 of the rotated L / R signal.

[0150] Fig.17 1700 is shown. The input white noise signal 1700 may be a two-channel orthogonal white noise signal including a right channel component 1710 and a left channel component 1720. The figure also shows a middle component 1730 and a side component 1740 generated from the white noise signal. The spectral energy of the left channel component 1720 matches the spectral energy of the right channel component 1710, and the spectral energy of the middle component 1730 matches the spectral energy of the side component 1740. Compared with the right channel component 1710 and the left channel component 1720, the signal levels of the middle component 1730 and the side component 1740 are about 3 dB lower.

[0151] Fig.18 The mid-component 1730 is shown decomposed into a super-mid-component 1810 and a residual mid-component 1820. The mid-component 1730 represents the non-spatial information of the input audio signal in the sound field. The super-mid-component 1810 includes a sub-component of non-spatial information found directly in the center of the sound field; the residual mid-component 1820 is the residual non-spatial information. In a typical stereo audio signal, the super-mid-component 1810 may include key features of the audio signal, such as dialogue or vocals. Fig.18 , the residual middle component 1820 is about 3 dB lower than the middle component 1730, while the super middle component 1810 is about 8-9 dB lower than the middle component 1730.

[0152] Fig.19 The side component 1740 is shown decomposed into a super side component 1910 and a residual side component 1920. The side component 1740 represents the spatial information in the input audio signal in the sound field. The super side component 1910 includes subcomponents of spatial information found at the edges of the sound field; the residual side component 1920 is the residual spatial information. In a typical stereo audio signal, the residual side component 1920 includes key features resulting from processing such as the effects of binaural processing, panning techniques, reverberation, and / or decorrelation processing. Fig.19 As shown, the relationship between the side component 1740 , the super side component 1910 , and the residual side component 1920 is similar to the relationship between the middle component 1730 , the super middle component 1810 , and the residual side component 1820 .

[0153] Computer architecture

[0154] Fig. 202000 is a block diagram of a computer system 2000 according to one or more embodiments. The computer system 2000 is an example of a circuit device that implements an audio processing system. At least one processor 2002 coupled to a chipset 2004 is shown. The chipset 2004 includes a memory controller hub 2020 and an input / output (I / O) controller hub 2022. The memory 2006 and the graphics adapter 2012 are coupled to the memory controller hub 2020, and the display device 2018 is coupled to the graphics adapter 2012. The storage device 1008, the keyboard 2010, the pointing device 2014, and the network adapter 2016 are coupled to the I / O controller hub 2022. The computer system 2000 may include various types of input or output devices. Other embodiments of the computer system 2000 have different architectures. For example, in some embodiments, the memory 2006 is directly coupled to the processor 2002.

[0155] The storage device 2008 includes one or more non-transitory computer-readable storage media, such as a hard drive, a compact disk read-only memory (CD-ROM), a DVD, or a solid-state memory device. The memory 2006 stores program code (composed of one or more instructions) and data used by the processor 2002. The program code may correspond to a combination of Figure 1-Figure 19 Processing aspects of the description.

[0156] Pointing device 2014 is used in conjunction with keyboard 2010 to enter data into computer system 2000. Graphics adapter 2012 displays images and other information on display device 2018. In some embodiments, display device 2018 includes touch screen capabilities for receiving user input and selections. Network adapter 2016 couples computer system 2000 to a network. Some embodiments of computer system 2000 have Fig. 20 The components shown in FIG. 7 are different and / or other components.

[0157] The circuit device may include one or more processors that execute program code stored in a non-transitory computer readable medium, which, when executed by the one or more processors, configures the one or more processors to implement an audio processing system or a module of an audio processing system. Other examples of circuit devices that implement an audio processing system or a module of an audio processing system may include integrated circuit devices, such as application specific integrated circuit devices (ASICs), field programmable gate arrays (FPGAs), or other types of computer circuit devices.

[0158] Additional considerations

[0159] Example benefits and advantages of the disclosed configuration include dynamic audio enhancements due to the enhanced audio system adapting to the device and the associated audio rendering system, and other relevant information provided by the device OS, such as use case information (e.g., indicating that the audio signal is used for music playback rather than gaming). The enhanced audio system can be integrated into the device (e.g., using a software development kit) or stored on a remote server for on-demand access. In this way, the device does not need to use storage or processing resources to maintain an audio enhancement system specific to its audio rendering system or audio rendering configuration. In some embodiments, the enhanced audio system is able to query the rendering system information at different levels, so that effective audio enhancement can be applied across different levels of available device-specific rendering information.

[0160] Throughout this specification, multiple instances can implement components, operations or structures described as single instances. Although the individual operations of one or more methods are illustrated and described as separate operations, one or more individual operations can be performed simultaneously, and there is no requirement that these operations be performed in the order shown. The structure and function presented as separate components in the example configuration can be implemented as a combined structure or component. Similarly, the structure and function presented as a single component can be implemented as a separate component. These and other changes, modifications, additions and improvements fall within the scope of the subject matter herein.

[0161] Certain embodiments are described herein as including logic or multiple components, modules, or mechanisms. Modules may constitute software modules (e.g., code contained on a machine-readable medium or in a transmission signal) or hardware modules. A hardware module is a tangible unit capable of performing certain operations and may be configured or arranged in a certain manner. In an example embodiment, one or more computer systems (e.g., independent client or server computer systems) or one or more hardware modules of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or application portion) as a hardware module for performing certain operations described herein.

[0162] The various operations of the example methods described herein may be performed at least in part by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Such processors, whether temporarily configured or permanently configured, may constitute processor-implemented modules that are used to perform one or more operations or functions. In some example embodiments, the modules mentioned herein may include processor-implemented modules.

[0163] Similarly, the methods described herein may be implemented at least in part by a processor. For example, at least some of the operations of a method may be performed by one or more processors or processor-implemented hardware modules. The execution of certain operations may be distributed between one or more processors, not only residing in a single machine, but also deployed on multiple machines. In some example embodiments, one or more processors may be located in a single location (e.g., in a home environment, an office environment, or as a server farm), while in other embodiments, the processors may be distributed in multiple locations.

[0164] Unless expressly stated otherwise, discussions herein using terms such as "processing," "computing," "calculating," "determining," "presenting," "displaying," and the like may refer to the action or process of a machine (e.g., a computer) that manipulates or transforms data represented as physical (electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.

[0165] As used herein, any reference to "one embodiment" or "an embodiment" means that a particular element, feature, structure, or characteristic described in conjunction with the embodiment is included in at least one embodiment. The phrase "in one embodiment" appearing in various places in the specification does not necessarily refer to the same embodiment.

[0166] Some embodiments may be described using the expressions "coupled" and "connected" along with their derivatives. It should be understood that these terms are not intended to be synonymous with each other. For example, the term "connected" may be used to describe some embodiments to indicate that two or more elements are in direct physical or electrical contact with each other. In another example, the term "coupled" may be used to describe some embodiments to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term "coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other. The embodiments are not limited to this context.

[0167] As used herein, the terms "comprises," "comprising," "includes," "including," "has," "having," or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that includes a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. In addition, unless expressly stated to the contrary, "or" refers to an inclusive or and not an exclusive or. For example, any of the following satisfies condition A or B: A is true (or exists) and B is false (or does not exist), A is false (or does not exist) and B is true (or exists), and both A and B are true (or exist).

[0168] In addition, "a" or "an" is used to describe elements and components of the embodiments herein. This is done merely for convenience and to give a general sense of the invention. The description should be understood to include one or at least one, and the singular also includes the plural unless it is obvious that it has another meaning.

[0169] Some parts of this specification describe embodiments in terms of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are often used by technicians in the field of data processing to effectively convey the essence of their work to other technicians in the field. Although these operations are described functionally, computationally, or logically, they are understood to be implemented by computer programs or equivalent circuit devices, microcodes, etc. In addition, without loss of generality, it is sometimes convenient to refer to these operational arrangements as modules. The described operations and their associated modules can be embodied in software, firmware, hardware, or any combination thereof.

[0170] Any steps, operations or processes described herein may be performed or implemented using one or more hardware or software modules, alone or in combination with other devices. In one embodiment, the software module is implemented with a computer program product, which includes a computer-readable medium containing computer program code, which may be executed by a computer processor to perform any or all of the steps, operations or processes described.

[0171] Embodiments may also relate to apparatus for performing the operations herein. The apparatus may be specially constructed for the desired purpose, and / or it may include a general-purpose computing device selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a non-transitory tangible computer-readable storage medium, or in any type of suitable medium for storing electronic instructions, which may be coupled to a computer system bus. In addition, any computing system mentioned in this specification may include a single processor, or may be an architecture that employs multiple processor designs to increase computing power.

[0172] Embodiments may also relate to products produced by the computing processes described herein. Such products may include information produced by the computing processes, wherein the information is stored on a non-transitory tangible computer-readable storage medium and may include any embodiment of a computer program product or other data combination described herein.

[0173] After reading this disclosure, those skilled in the art will appreciate additional alternative structural and functional designs for systems and processes for audio enhancement using device-specific metadata through the principles disclosed herein. Therefore, although specific embodiments and applications have been illustrated and described, it should be understood that the disclosed embodiments are not limited to the precise construction and components disclosed herein. Various modifications, changes and variations that will be apparent to those skilled in the art may be made to the arrangement, operation and details of the methods and apparatus disclosed herein without departing from the spirit and scope defined by the appended claims.

[0174] Finally, the language used in the specification is selected primarily for readability and instructional purposes, and not for the purpose of describing or limiting the patent rights. Therefore, it is intended that the scope of the patent rights be limited not by this detailed description, but by any claims issued on an application based thereon. Therefore, the disclosure of the embodiments is intended to illustrate, not to limit, the scope of the patent rights set forth in the appended claims.

Claims

1. A system for processing an audio signal, comprising: A circuit arrangement is configured to: generating a mid component and a side component from a left channel and a right channel of the audio signal; converting the middle component and the side component into a frequency domain; generating a super-middle component including removing spectral energy of the side component from spectral energy of the middle component by subtracting the magnitude of the side component in the frequency domain from the magnitude of the middle component in the frequency domain; filtering the super intermediate component; as well as A left output channel and a right output channel are generated using the filtered super middle component.

2. The system of claim 1, wherein: The super mid-component isolates a portion of the mid-component corresponding to spectral energy present in the center of the sound field.

3. The system of claim 1 , wherein the circuit arrangement is configured to filter the super intermediate component comprises: The circuit arrangement is configured to at least one of gain adjust or time delay a sub-band of the super intermediate component.

4. The system of claim 1 , wherein the circuit arrangement is configured to filter the super intermediate component comprises: The circuit arrangement is configured to apply dynamic range processing to the super intermediate component.

5. The system of claim 1 , wherein the circuit arrangement is configured to filter the super intermediate component comprises: The circuit arrangement is configured to adjust a frequency-dependent amplitude or a frequency-dependent delay of the super intermediate component.

6. The system of claim 1 , wherein the circuit arrangement is configured to filter the super intermediate component comprises: The circuit arrangement is configured to apply a machine learning based style transfer, translation or resynthesis to the super intermediate component.

7. The system of claim 1 , wherein the circuit device is further configured to: generating a residual intermediate component, the residual intermediate component comprising removing spectral energy of the super intermediate component from the spectral energy of the intermediate component; filtering the residual intermediate component; and The left output channel and the right output channel are generated using the filtered residual intermediate component.

8. The system of claim 7, wherein the circuit arrangement is configured to filter the residual intermediate component comprises: The circuit arrangement is configured to at least one of gain adjust or time delay a subband of the residual intermediate component.

9. The system of claim 7, wherein the circuit arrangement is configured to filter the residual intermediate component comprises: The circuit arrangement is configured to apply a dynamic range processing to the residual intermediate component.

10. The system of claim 7, wherein the circuit arrangement is configured to filter the residual intermediate component comprises: The circuit arrangement is configured to adjust a frequency-dependent amplitude or a frequency-dependent delay of the residual intermediate component.

11. The system of claim 7, wherein the circuit arrangement is configured to filter the residual intermediate component comprises: The circuit arrangement is configured to apply a machine learning based style transfer, conversion or resynthesis to the residual intermediate component.

12. The system of claim 7, wherein: The circuit arrangement is further configured to apply a Fourier transform to the intermediate components to convert the intermediate components to the frequency domain; and The circuit arrangement being configured to generate the residual middle component comprising removing spectral energy of the super middle component from the spectral energy of the middle component comprises: the circuit arrangement being configured to subtract a magnitude of the super middle component in the frequency domain from a magnitude of the middle component in the frequency domain.

13. The system of claim 1 , wherein the circuit device is further configured to: applying an inverse Fourier transform to the super intermediate component to convert the super intermediate component in the frequency domain to the time domain; generating a delayed intermediate component by time-delaying the intermediate component; generating a residual intermediate component by subtracting the super intermediate component in the time domain from the delayed intermediate component in the time domain; filtering the residual intermediate component; and The left output channel and the right output channel are generated using the filtered residual intermediate component.

14. The system of claim 1, wherein the circuit device is further configured to: generating a super side component comprising removing the spectral energy of the mid component from the spectral energy of the side component; filtering the super-side component; as well as The left output channel and the right output channel are generated using the filtered super-side component.

15. The system of claim 14, wherein: The circuit arrangement is further configured to apply a Fourier transform to the mid-component and the side component to convert the mid-component and the side component to a frequency domain; and The circuit arrangement being configured to generate the super-side component comprising removing the spectral energy of the middle component from the spectral energy of the side component comprises: the circuit arrangement being configured to subtract a magnitude of the middle component in the frequency domain from a magnitude of the side component in the frequency domain.

16. The system of claim 14, wherein the circuit arrangement is configured to filter the super-side component comprises: The circuit arrangement is configured to at least one of gain adjust or time delay a sub-band of the super-side component.

17. The system of claim 14, wherein the circuit arrangement is configured to filter the super-side component comprises: The circuit arrangement is configured to apply dynamic range processing to the super-side component.

18. The system of claim 14, wherein the circuit arrangement is configured to filter the super-side component comprises: The circuit arrangement is configured to adjust a frequency-dependent amplitude or a frequency-dependent delay of the super-side component.

19. The system of claim 14, wherein the circuit arrangement is configured to filter the super-side component comprises: The circuit arrangement is configured to apply a machine learning based style transfer, conversion or resynthesis to the super-side component.

20. The system of claim 1, wherein the circuit arrangement is further configured to: generating a super side component comprising removing the spectral energy of the mid component from the spectral energy of the side component; generating a residual side component comprising removing spectral energy of the super-side component from the spectral energy of the side component; filtering the residual side component; as well as The left output channel and the right output channel are generated using the filtered residual side component.

21. The system of claim 20, wherein the circuit device is configured to filter the residual side component comprises: The circuit arrangement is configured to at least one of gain adjust or time delay a subband of the residual side component.

22. The system of claim 20, wherein the circuit device is configured to filter the residual side component comprises: The circuit arrangement is configured to apply a dynamic range processing to the residual-side component.

23. The system of claim 20, wherein the circuit device is configured to filter the residual side component comprises: The circuit arrangement is configured to adjust a frequency-dependent amplitude or a frequency-dependent delay of the residual-side component.

24. The system of claim 20, wherein the circuit device is configured to filter the residual side component comprises: The circuit arrangement is configured to apply a machine learning based style transfer, conversion or resynthesis to the residual side component.

25. The system of claim 20, wherein: The circuit arrangement is further configured to apply a Fourier transform to the side component to convert the side component to a frequency domain; and The circuit arrangement being configured to generate the residual side component comprising removing the spectral energy of the super-side component from the spectral energy of the side component comprises: the circuit arrangement being configured to subtract a magnitude of the super-side component in the frequency domain from a magnitude of the side component in the frequency domain.

26. The system of claim 1, wherein the circuit arrangement is further configured to: generating a super side component comprising removing the spectral energy of the mid component from the spectral energy of the side component; applying an inverse Fourier transform to the super-side component to convert the super-middle component in the frequency domain to the time domain; generating a delayed side component by time-delaying the side component; generating a residual side component by subtracting the super side component in the time domain from the delayed side component in the time domain; filtering the residual side component; as well as The left output channel and the right output channel are generated using the filtered residual side component.

27. A non-transitory computer readable medium comprising stored program code, which when executed by at least one processor configures the at least one processor to: generating a mid component and a side component from a left channel and a right channel of an audio signal; converting the middle component and the side component into a frequency domain; generating a super-middle component including removing spectral energy of the side component from spectral energy of the middle component by subtracting the magnitude of the side component in the frequency domain from the magnitude of the middle component in the frequency domain; filtering the super intermediate component; as well as A left output channel and a right output channel are generated using the filtered super middle component.

28. The non-transitory computer readable medium of claim 27, wherein: The super mid-component isolates a portion of the mid-component corresponding to spectral energy present in the center of the sound field.

29. The non-transitory computer-readable medium of claim 27, wherein the program code that configures the at least one processor to filter the super intermediate component further configures the at least one processor to at least one of gain adjust or time delay a subband of the super intermediate component.

30. The non-transitory computer readable medium of claim 27, wherein the program code that configures the at least one processor to filter the super intermediate component further configures the at least one processor to apply dynamic range processing to the super intermediate component.

31. The non-transitory computer-readable medium of claim 27, wherein the program code that configures the at least one processor to filter the super intermediate component further configures the at least one processor to adjust a frequency-dependent amplitude or a frequency-dependent delay of the super intermediate component.

32. The non-transitory computer-readable medium of claim 27, wherein the program code that configures the at least one processor to filter the super-intermediate component further configures the at least one processor to apply machine learning-based style transfer, conversion, or resynthesis to the super-intermediate component.

33. The non-transitory computer readable medium of claim 27, wherein the program code further configures the at least one processor to: generating a residual mid component comprising removing spectral energy of the super mid component from the spectral energy of the mid component; filtering the residual intermediate component; and The left output channel and the right output channel are generated using the filtered residual intermediate component.

34. The non-transitory computer readable medium of claim 33, wherein the program code that configures the at least one processor to filter the residual intermediate component further configures the at least one processor to at least one of gain adjust or time delay a subband of the residual intermediate component.

35. The non-transitory computer readable medium of claim 33, wherein the program code that configures the at least one processor to filter the residual intermediate component further configures the at least one processor to apply dynamic range processing to the residual intermediate component.

36. The non-transitory computer readable medium of claim 33, wherein the program code that configures the at least one processor to filter the residual intermediate component further configures the at least one processor to adjust a frequency-dependent amplitude or a frequency-dependent delay of the residual intermediate component.

37. A non-transitory computer-readable medium according to claim 33, wherein the program code that configures the at least one processor to filter the residual intermediate component also configures the at least one processor to: apply machine learning-based style transfer, conversion, or resynthesis to the residual intermediate component.

38. The non-transitory computer readable medium of claim 33, wherein: The program code further configures the at least one processor to: apply a Fourier transform to the intermediate components to convert the intermediate components to a frequency domain; The program code configuring the at least one processor to generate the residual intermediate component comprising removing the spectral energy of the super intermediate component from the spectral energy of the intermediate component further configures the at least one processor to: subtract the size of the super intermediate component in the frequency domain from the size of the intermediate component in the frequency domain.

39. The non-transitory computer readable medium of claim 27, wherein the program code further configures the at least one processor to: applying an inverse Fourier transform to the super intermediate component to convert the super intermediate component in the frequency domain to the time domain; generating a delayed intermediate component by time-delaying the intermediate component; generating a residual intermediate component by subtracting the super intermediate component in the time domain from the delayed intermediate component in the time domain; filtering the residual intermediate component; and The left output channel and the right output channel are generated using the filtered residual intermediate component.

40. The non-transitory computer readable medium of claim 27, wherein the program code further configures the at least one processor to: generating a super side component comprising removing the spectral energy of the mid component from the spectral energy of the side component; filtering the super-side component; as well as The left output channel and the right output channel are generated using the filtered super-side component.

41. The non-transitory computer readable medium of claim 40, wherein: The program code further configures the at least one processor to apply a Fourier transform to the mid-component and the side component to convert the mid-component and the side component to a frequency domain; as well as The program code configuring the at least one processor to generate the super-side component comprising removing the spectral energy of the middle component from the spectral energy of the side component further configures the at least one processor to: subtract the magnitude of the middle component in the frequency domain from the magnitude of the side component in the frequency domain.

42. The non-transitory computer readable medium of claim 40, wherein the program code configuring the at least one processor to filter the super-side component comprises: Program code configured to configure the at least one processor to at least one of gain adjust or time delay a subband of the super-side component.

43. The non-transitory computer readable medium of claim 40, wherein the program code configuring the at least one processor to filter the super-side component comprises: Program code configuring the at least one processor to apply dynamic range processing to the super-side component.

44. The non-transitory computer readable medium of claim 40, wherein the program code configuring the at least one processor to filter the super-side component comprises: Program code configuring the at least one processor to adjust a frequency-dependent amplitude or a frequency-dependent delay of the super-side component.

45. The non-transitory computer readable medium of claim 40, wherein the program code configuring the at least one processor to filter the super-side component comprises: Program code configuring the at least one processor to apply machine learning based style transfer, translation, or resynthesis to the super-side component.

46. ​​The non-transitory computer readable medium of claim 27, wherein the program code further configures the at least one processor to: generating a super side component comprising removing the spectral energy of the mid component from the spectral energy of the side component; generating a residual side component comprising removing spectral energy of the super-side component from the spectral energy of the side component; filtering the residual side component; as well as The left output channel and the right output channel are generated using the filtered residual side component.

47. The non-transitory computer-readable medium of claim 46, wherein the program code that configures the at least one processor to filter the residual side component further configures the at least one processor to at least one of gain adjust or time delay a subband of the residual side component.

48. The non-transitory computer readable medium of claim 46, wherein the program code that configures the at least one processor to filter the residual side component further configures the at least one processor to apply dynamic range processing to the residual side component.

49. The non-transitory computer readable medium of claim 46, wherein the program code that configures the at least one processor to filter the residual side component further configures the at least one processor to adjust a frequency-dependent amplitude or a frequency-dependent delay of the residual side component.

50. The non-transitory computer-readable medium of claim 46, wherein the program code that configures the at least one processor to filter the residual side component further configures the at least one processor to apply machine learning-based style transfer, conversion, or resynthesis to the residual side component.

51. The non-transitory computer readable medium of claim 46, wherein: The program code further configures the at least one processor to apply a Fourier transform to the side component to convert the side component to a frequency domain; as well as The program code configuring the at least one processor to generate the residual side component comprising removing the spectral energy of the super-side component from the spectral energy of the side component further configures the at least one processor to: subtract the magnitude of the super-side component in the frequency domain from the magnitude of the side component in the frequency domain.

52. The non-transitory computer readable medium of claim 27, wherein the program code further configures the at least one processor to: generating a super side component comprising removing the spectral energy of the mid component from the spectral energy of the side component; applying an inverse Fourier transform to the super-side component to convert the super-middle component in the frequency domain to the time domain; generating a delayed side component by time-delaying the side component; generating a residual side component by subtracting the super side component in the time domain from the delayed side component in the time domain; filtering the residual side component; as well as The left output channel and the right output channel are generated using the filtered residual side component.

53. A method for processing an audio signal, comprising: generating a mid component and a side component from a left channel and a right channel of an audio signal; converting the middle component and the side component into a frequency domain; generating a super-middle component including removing spectral energy of the side component from spectral energy of the middle component by subtracting the magnitude of the side component in the frequency domain from the magnitude of the middle component in the frequency domain; filtering the super intermediate component; as well as A left output channel and a right output channel are generated using the filtered super middle component.

54. The method of claim 53, wherein: The super mid-component isolates a portion of the mid-component corresponding to spectral energy present in the center of the sound field.

55. The method of claim 53, wherein filtering the super intermediate component comprises: At least one of gain adjustment and time delay is performed on the sub-band of the super intermediate component.

56. The method of claim 53, wherein filtering the super intermediate component comprises: Dynamic range processing is applied to the super intermediate component.

57. The method of claim 53, wherein filtering the super intermediate component comprises: A frequency-dependent amplitude or a frequency-dependent delay of the super intermediate component is adjusted.

58. The method of claim 53, wherein filtering the super intermediate component comprises: Applying machine learning based style transfer, translation, or resynthesis to the super intermediate component.

59. The method of claim 53, further comprising, by the circuit device: generating a residual mid component comprising removing spectral energy of the super mid component from the spectral energy of the mid component; filtering the residual intermediate component; and The left output channel and the right output channel are generated using the filtered residual intermediate component.

60. The method of claim 59, wherein filtering the residual intermediate component comprises: At least one of gain adjustment and time delay is performed on the subband of the residual intermediate component.

61. The method of claim 59, wherein filtering the residual intermediate component comprises: Dynamic range processing is applied to the residual intermediate components.

62. The method of claim 59, wherein filtering the residual intermediate component comprises: A frequency-dependent amplitude or a frequency-dependent delay of the residual intermediate component is adjusted.

63. The method of claim 59, wherein filtering the residual intermediate component comprises: Applying machine learning based style transfer, translation, or resynthesis to the residual intermediate components.

64. The method of claim 59, wherein: The method further comprises: applying a Fourier transform to the intermediate components to convert the intermediate components to the frequency domain; and Generating the residual middle component comprising removing spectral energy of the super middle component from the spectral energy of the middle component comprises subtracting a magnitude of the super middle component in the frequency domain from a magnitude of the middle component in the frequency domain.

65. The method of claim 53, further comprising, by the circuit device: applying an inverse Fourier transform to the super intermediate component to convert the super intermediate component in the frequency domain to the time domain; generating a delayed intermediate component by time-delaying the intermediate component; generating a residual intermediate component by subtracting the super intermediate component in the time domain from the delayed intermediate component in the time domain; filtering the residual intermediate component; and The left output channel and the right output channel are generated using the filtered residual intermediate component.

66. The method of claim 53, further comprising, by the circuit device: generating a super side component comprising removing the spectral energy of the mid component from the spectral energy of the side component; filtering the super-side component; as well as The left output channel and the right output channel are generated using the filtered super-side component.

67. The method of claim 66, wherein: The method further comprises: applying a Fourier transform to the mid-component and the side component to convert the mid-component and the side component to a frequency domain; as well as Generating the super-side component including removing the spectral energy of the middle component from the spectral energy of the side component includes subtracting a magnitude of the middle component in the frequency domain from a magnitude of the side component in the frequency domain.

68. The method of claim 66, wherein filtering the super-side component comprises: At least one of gain adjustment and time delay is performed on the subband of the super-side component.

69. The method of claim 66, wherein filtering the super-side component comprises: Dynamic range processing is applied to the super-side component.

70. The method of claim 66, wherein filtering the super-side component comprises: The frequency-dependent amplitude or frequency-dependent delay of the super-side component is adjusted.

71. The method of claim 66, wherein filtering the super-side component comprises: Applying machine learning based style transfer, translation or resynthesis to the super-side component.

72. The method of claim 53, further comprising: generating a super side component comprising removing the spectral energy of the mid component from the spectral energy of the side component; generating a residual side component comprising removing spectral energy of the super-side component from the spectral energy of the side component; filtering the residual side component; as well as The left output channel and the right output channel are generated using the filtered residual side component.

73. The method of claim 72, wherein filtering the residual side component further comprises: At least one of gain adjustment and time delay is performed on the subband of the residual side component.

74. The method of claim 72, wherein filtering the residual side component further comprises: Dynamic range processing is applied to the residual side component.

75. The method of claim 72, wherein filtering the residual side component further comprises: The frequency-dependent amplitude or frequency-dependent delay of the residual side component is adjusted.

76. The method of claim 72, wherein filtering the residual side component further comprises: Applying machine learning based style transfer, conversion or resynthesis to the residual side component.

77. The method of claim 72, wherein: The method further includes: applying a Fourier transform to the side component to convert the side component to a frequency domain; and Generating the residual side component including removing the spectral energy of the super-side component from the spectral energy of the side component further includes subtracting a magnitude of the super-side component in the frequency domain from a magnitude of the side component in the frequency domain.

78. The method of claim 53, further comprising: generating a super side component comprising removing the spectral energy of the mid component from the spectral energy of the side component; applying an inverse Fourier transform to the super-side component to convert the super-middle component in the frequency domain to the time domain; generating a delayed side component by time-delaying the side component; generating a residual side component by subtracting the super side component in the time domain from the delayed side component in the time domain; filtering the residual side component; as well as The left output channel and the right output channel are generated using the filtered residual side component.

Citation Information

Patent Citations

  • Apparatus and method for sound stage enhancement

    CN108293165A

  • Processing of audio channels

    US20120076307A1