Sub-band space processing using spectrally orthogonal audio component and crosstalk processing

By generating spectrally orthogonal components and applying subband spatial processing and crosstalk compensation, the system enhances audio processing capabilities and corrects spectral defects, enabling targeted audio enhancements and adjustments.

JP2025163239APending Publication Date: 2025-10-28BOOMCLOUD 360 INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025132685
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-08-03
Filing Date
2025-08-07
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing audio processing technologies struggle to effectively separate and manipulate spectrally orthogonal components of stereo signals, limiting the range of audio processing possibilities and introducing spectral defects from crosstalk processing.

Method used

The system generates hyper-mid, hyper-side, residual mid, and residual side components from left and right audio channels, applying subband spatial processing and crosstalk compensation to enhance spatial detectability and compensate for spectral defects.

Benefits of technology

This approach allows targeted enhancements and adjustments to audio content, such as vocal content removal or spatialization effects, with minimal overall gain change and vocal presence loss, while addressing spectral imperfections from crosstalk processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025163239000001_ABST
    Figure 2025163239000001_ABST
Patent Text Reader

Abstract

To expand a range of the possibility when processing audio.SOLUTION: A system includes a circuitry formed to generate a mid component and a side component from a left channel and a right channel of an audio signal, generate a hyper mid component by decorrelating the mid component from the side component, and generate a left output channel and a right output channel by using the hyper mid component.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE DISCLOSURE This disclosure relates generally to audio processing, and more particularly to spatial audio processing. [Background technology]

[0002] Conceptually, the side (or "spatial") components of a left-right stereo signal can be thought of as those portions of the left and right channels that contain spatial information (i.e., sounds in the stereo signal that appear somewhere to the left or right of center in the sound field). Conversely, the mid (or "non-spatial") components of a left-right stereo signal can be thought of as those portions of the left and right channels that contain non-spatial information (i.e., sounds in the stereo signal that appear in the center of the sound field). The mid component contains energy in the stereo signal that is perceived as non-spatial, but it also has energy from elements in the stereo signal that are not generally perceptually located in the center of the sound field. Similarly, the side component contains energy in the stereo signal that is perceived as spatial, but it also has energy from elements in the stereo signal that are generally perceptually located in the center of the sound field. To expand the range of possibilities in processing audio, it is desirable to separate and manipulate portions of the mid and side components that are spectrally "orthogonal" to one another. Summary of the Invention

[0003] Embodiments relate to spatial audio processing using spectrally orthogonal audio components, such as hyper-mid, hyper-side, residual mid, or residual side components, of a stereo or other multi-channel audio signal. The spatial audio processing may include subband spatial processing to enhance spatial detectability or crosstalk compensation processing to compensate for spectral defects resulting from crosstalk processing applied to the audio signal. The crosstalk processing may include crosstalk cancellation or crosstalk simulation.

[0004] Some embodiments include a system for processing an audio signal. The system includes a circuit that generates a mid component and a side component from left and right channels of the audio signal. The circuit generates a hyper-mid component that includes spectral energy of the side component removed from the spectral energy of the mid component, and generates a residual mid component that includes spectral energy of the hyper-mid component removed from the spectral energy of the mid component. The circuit filters subbands of the residual mid component, such as to apply subband spatial processing. The circuit generates a left output channel and a right output channel using the filtered subbands of the residual mid component.

[0005] In some embodiments, each of the subbands of the residual mid component includes a set of critical bands.

[0006] In some embodiments, the circuit generates a hyperside component that includes the spectral energy of the mid component removed from the spectral energy of the side component, generates a residual side component that includes the spectral energy of the hyperside component removed from the spectral energy of the side component, filters subbands of the residual side component, and generates a left output channel and a right output channel using the filtered subbands of the residual side component.

[0007] In some embodiments, the circuit generates a hyperside component that includes the spectral energy of the mid component removed from the spectral energy of the side component, filters a subband of the hyperside component, and generates a left output channel and a right output channel using the filtered subband of the hyperside component.

[0008] In some embodiments, the circuit filters the subbands of the side component and generates a left output channel and a right output channel using the filtered subbands of the side component.

[0009] In some embodiments, the circuit filters sub-bands of the hyper-mid component and generates a left output channel and a right output channel using the filtered sub-bands of the hyper-mid component.

[0010] In some embodiments, the circuit applies crosstalk processing to the audio signal. The crosstalk processing includes crosstalk cancellation or crosstalk simulation. In some embodiments, the circuit filters the hyper-mid component to compensate for spectral imperfections caused by the crosstalk processing. In some embodiments, the circuit filters the residual mid component to compensate for spectral imperfections caused by the crosstalk processing. In some embodiments, the circuit filters the mid component to compensate for spectral imperfections caused by the crosstalk processing. In some embodiments, the circuit generates a hyper-side component including the spectral energy of the mid component removed from the spectral energy of the side component, and filters the hyper-side component to compensate for spectral imperfections caused by the crosstalk processing. In some embodiments, the circuit generates a residual side component including the spectral energy of the hyper-side component removed from the spectral energy of the side component, and filters the residual side component to compensate for spectral imperfections caused by the crosstalk processing. In some embodiments, the circuit filters the side component to compensate for spectral imperfections caused by the crosstalk processing.

[0011] Some embodiments include a non-transitory computer-readable medium including stored program code that, when executed by at least one processor, configures the at least one processor to generate a mid component and a side component from left and right channels of the audio signal, generate a hyper-mid component that includes spectral energy of the side component removed from the spectral energy of the mid component, generate a residual mid component that includes spectral energy of the hyper-mid component removed from the spectral energy of the mid component, filter sub-bands of the residual mid component, and generate a left output channel and a right output channel using the filtered sub-bands of the residual mid component.

[0012] Some embodiments include a method performed by a circuit, the method including generating a mid component and a side component from left and right channels of an audio signal, generating a hyper-mid component including spectral energy of the side component removed from the spectral energy of the mid component, generating a residual mid component including spectral energy of the hyper-mid component removed from the spectral energy of the mid component, filtering subbands of the residual mid component, and generating a left output channel and a right output channel using the filtered subbands of the residual mid component. [Brief explanation of the drawings]

[0013] The disclosed embodiments have other advantages and features that will become more readily apparent from the detailed description, the appended claims, and the accompanying figures (or drawings), a brief introduction of which is as follows: [Figure 1] FIG. 1 is a block diagram of an audio processing system according to one or more embodiments. [Figure 2A] FIG. 2A is a block diagram of a quadrature component generator according to one or more embodiments. [Figure 2B] FIG. 2B is a block diagram of a quadrature component generator according to one or more embodiments. [Figure 2C] FIG. 2C is a block diagram of a quadrature component generator according to one or more embodiments. [Figure 3] FIG. 3 is a block diagram of a quadrature component processor according to one or more embodiments. [Figure 4] FIG. 4 is a block diagram of a subband spatial processor according to one or more embodiments. [Figure 5] FIG. 5 is a block diagram of a crosstalk compensation processor according to one or more embodiments. [Figure 6] FIG. 6 is a block diagram of a crosstalk simulation processor according to one or more embodiments. [Figure 7] FIG. 7 is a block diagram of a crosstalk cancellation processor according to one or more embodiments. [Figure 8] FIG. 8 is a flowchart of a process for spatial processing using at least one of a hypermid component, a residual mid component, a hyperside component, or a residual side component, according to one or more embodiments. [Figure 9] FIG. 9 is a flowchart of a process for subband spatial processing and compensation for crosstalk using at least one of a hypermid component, a residual mid component, a hyperside component, or a residual side component, in accordance with one or more embodiments. [Figure 10] FIG. 10 is a plot illustrating the spectral energy of the mid and side components of an exemplary white noise signal, in accordance with one or more embodiments. [Figure 11] FIG. 11 is a plot illustrating the spectral energy of the mid and side components of an exemplary white noise signal, in accordance with one or more embodiments. [Figure 12]FIG. 12 is a plot illustrating the spectral energy of the mid and side components of an exemplary white noise signal, in accordance with one or more embodiments. [Figure 13] FIG. 13 is a plot illustrating the spectral energy of the mid and side components of an exemplary white noise signal, in accordance with one or more embodiments. [Figure 14] FIG. 14 is a plot illustrating the spectral energy of the mid and side components of an exemplary white noise signal, in accordance with one or more embodiments. [Figure 15] FIG. 15 is a plot illustrating the spectral energy of the mid and side components of an exemplary white noise signal, in accordance with one or more embodiments. [Figure 16] FIG. 16 is a plot illustrating the spectral energy of the mid and side components of an exemplary white noise signal, in accordance with one or more embodiments. [Figure 17] FIG. 17 is a plot illustrating the spectral energy of the mid and side components of an exemplary white noise signal, in accordance with one or more embodiments. [Figure 18] FIG. 18 is a plot illustrating the spectral energy of the mid and side components of an exemplary white noise signal, in accordance with one or more embodiments. [Figure 19] FIG. 19 is a plot illustrating the spectral energy of the mid and side components of an exemplary white noise signal, in accordance with one or more embodiments. [Figure 20] FIG. 20 is a block diagram of a computer system according to one or more embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0014] The figures and the following description relate to preferred embodiments by way of example only. It should be noted from the following description that alternative embodiments of the structures and methods disclosed herein will be readily recognized as viable alternatives that may be utilized without departing from the principles of what is claimed.

[0015] Reference will now be made in detail to several embodiments, examples of which are illustrated in the accompanying figures. It is noted that wherever practicable, like or similar reference numerals may be used in the figures and may indicate like or similar functionality. The figures depict embodiments of the disclosed system (or method) for illustrative purposes only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods illustrated herein may be utilized without departing from the principles described herein.

[0016] Embodiments relate to spatial audio processing using mid and side components that are spectrally orthogonal to each other. For example, an audio processing system may generate a hyper-mid component, which isolates a portion of the mid component that corresponds only to spectral energy present in the center of the sound field, or a hyper-side component, which isolates a portion of the side component that corresponds only to spectral energy not present in the center of the sound field. The hyper-mid component includes the spectral energy of the side component removed from the spectral energy of the mid component, and the hyper-side component includes the spectral energy of the mid component removed from the spectral energy of the side component. The audio processing system may also generate a residual mid component, which corresponds to the spectral energy of the mid component with the hyper-mid component removed (e.g., by subtracting the spectral energy of the hyper-mid component from the spectral energy of the mid component), or a residual side component, which corresponds to the spectral energy of the side component with the hyper-side component removed (e.g., by subtracting the spectral energy of the hyper-side component from the spectral energy of the side component). By isolating these orthogonal components and using them to perform various types of audio processing, the audio processing system can provide targeted enhancements of audio content. The hyper-mid components represent non-spatial (i.e., mid) spectral energy in the center of the sound field. For example, non-spatial spectral energy in the center of the sound field may contain dialogue in a movie or primary vocal content in music. Applying signal processing operations to the hyper-mids allows for adjustment of such audio content without altering the spectral energy present elsewhere in the sound field. For example, in some embodiments, vocal content may be partially and / or completely removed by applying a filter to the hyper-mid components that reduces spectral energy in the typical human vocal range.In other embodiments, targeted enhancements or effects to vocal content may be applied by filters that increase energy in the typical human vocal range (e.g., via compression, reverb, and / or other audio processing techniques). The residual mid component represents non-spatial spectral energy that is not in the center of the sound field. Applying signal processing techniques to the residual mid component allows similar transformations to be performed orthogonally to other components. For example, in some embodiments, targeted spectral energy in the residual mid component may be partially and / or completely removed, while spectral energy in the residual side components is increased, to provide a spatialization effect to the audio content with minimal change in overall perceived gain and minimal loss of vocal presence.

[0017] Exemplary Audio Processing System FIG. 1 is a block diagram of an audio processing system 100 according to one or more embodiments. Audio processing system 100 is a circuit that processes an input audio signal to generate a spatially enhanced output audio signal. The input audio signal includes a left input channel 103 and a right input channel 105, and the output audio signal includes a left output channel 121 and a right output channel 123. Audio processing system 100 includes an L / R to M / S converter module 107, a quadrature component generator module 113, a quadrature component processor module 117, an M / S to L / R converter module 119, and a crosstalk processor module 141. In some embodiments, audio processing system 100 includes a subset of the components described above and / or additional components to those described above. In some embodiments, audio processing system 100 processes the input audio signal in a different order than that shown in FIG. 1. For example, the audio processing system 100 may process the input audio with crosstalk processing prior to processing using the quadrature component generator module 113 and the quadrature component processor module 117 .

[0018] The L / R to M / S converter module 107 receives the left input channel 103 and the right input channel 105 and generates a mid component 109 (e.g., a non-spatial component) and a side component 111 (e.g., a spatial component) from the input channels 103 and 105. In some embodiments, the mid component 109 is generated based on the sum of the left input channel 103 and the right input channel 105, and the side component 111 is generated based on the difference between the left input channel 103 and the right input channel 105. In some embodiments, several mid and side components are generated from a multi-channel input audio signal (e.g., surround sound). Other L / R to M / S type conversions can be used to generate the mid component 109 and the side component 111.

[0019] The quadrature component generator module 113 processes the mid component 109 and the side component 111 to generate at least one of a hyper-mid component M1, a hyper-side component S1, a residual mid component M2, and a residual side component S2. The hyper-mid component M1 is the mid component 109 with the side component 111 removed. The hyper-side component S1 is the spectral energy of the side component 111 with the spectral energy of the mid component 109 removed. The residual mid component M2 is the spectral energy of the mid component 109 with the spectral energy of the hyper-mid component M1 removed. The residual side component S2 is the spectral energy of the side component 111 with the spectral energy of the hyper-side component S1 removed. In some embodiments, the audio processing system 100 generates the left output channel 121 and the right output channel 123 by processing at least one of the hyper-mid component M1, hyper-side component S1, residual mid component M2, and residual side component S2. The quadrature component generator module 113 is further described with respect to Figures 2A-2C.

[0020] The quadrature component processor module 117 processes one or more of the hyper-mid component M1, hyper-side component S1, residual mid component M2, and / or residual side component S2. The processing on components M1, M2, S1, and S2 may include various types of filtering, such as spatial cue processing (e.g., amplitude or delay-based panning, binaural processing, etc.), dynamic range processing, machine learning-based processing, gain application, reverberation, audio effect addition, or other types of processing. In some embodiments, the quadrature component processor module 117 performs subband spatial processing and / or crosstalk compensation processing using the hyper-mid component M1, hyper-side component S1, residual mid component M2, and / or residual side component S2 to generate the processed mid component 131 and the processed side component 139. Subband spatial processing is processing performed on frequency subbands of the mid and side components of an audio signal to spatially enhance the audio signal. Crosstalk compensation processing is processing performed on audio signals to adjust for spectral artifacts caused by crosstalk processing, such as crosstalk compensation for loudspeakers or crosstalk simulation for headphones. The quadrature component processor module 117 is further described with respect to FIG.

[0021] M / S to L / R converter module 119 receives processed mid component 131 and processed side component 139 and generates processed left component 151 and processed right component 159. In some embodiments, processed left component 151 is generated based on the sum of processed mid component 131 and processed side component 139, and processed right component 159 is generated based on the difference between processed mid component 131 and processed side component 139. Other M / S to L / R type converters may be used to generate processed left component 151 and processed right component 159.

[0022] The crosstalk processor module 141 receives the processed left component 151 and the processed right component 159 and performs crosstalk processing thereon. Crosstalk processing includes, for example, crosstalk simulation or crosstalk cancellation. Crosstalk simulation is processing performed on an audio signal (e.g., output through headphones) to simulate the effect of loudspeakers. Crosstalk cancellation is processing performed on an audio signal configured to be output through loudspeakers to remove crosstalk caused by the loudspeakers. The crosstalk processor module 141 outputs a left output channel 121 and a right output channel 123.

[0023] Exemplary Quadrature Component Generator 2A-2C are block diagrams of quadrature component generator modules 213, 223, and 245, respectively, according to one or more embodiments. Quadrature component generator modules 213, 223, and 245 are examples of quadrature component generator module 113.

[0024] 2A , the quadrature component generator module 213 includes a subtraction unit 205, a subtraction unit 209, a subtraction unit 215, and a subtraction unit 219. As described above, the quadrature component generator module 113 receives the mid component 109 and the side component 111, and outputs one or more of a hyper-mid component M1, a hyper-side component S1, a residual mid component M2, and a residual side component S2.

[0025] The subtraction unit 205 removes the spectral energy of the side component 111 from the spectral energy of the mid component 109 to generate the hyper-mid component M1. For example, the subtraction unit 205 subtracts the magnitude of the side component 111 in the frequency domain from the magnitude of the mid component 109 in the frequency domain, leaving only the phase, to generate the hyper-mid component M1. The subtraction in the frequency domain may be performed using a Fourier transform on the time-domain signal to generate the signal in the frequency domain, followed by subtraction of the signals in the frequency domain. In other examples, the subtraction in the frequency domain may be performed in other ways, such as using a wavelet transform instead of a Fourier transform. The subtraction unit 209 generates the residual mid component M2 by removing the spectral energy of the hyper-mid component M1 from the spectral energy of the mid component 109. For example, the subtraction unit 209 subtracts the magnitude of the hyper-mid component M1 in the frequency domain from the magnitude of the mid component 109 in the frequency domain, leaving only the phase, to generate the residual mid component M2. While subtracting the side from the mid in the time domain yields the original right channel of the signal, the above operation in the frequency domain separates and distinguishes between the portion of the spectral energy of the mid component that differs from the spectral energy of the side component (called M1, or hypermid) and the portion of the spectral energy of the mid component that is the same as the spectral energy of the side component (called M2, or residual mid).

[0026] In some embodiments, when subtracting the spectral energy of the side component 111 from the spectral energy of the mid component 109 results in a negative value for the hyper-mid component M1 (e.g., for one or more of the bins in the frequency domain), additional processing may be used. In some embodiments, when subtracting the spectral energy of the side component 111 from the spectral energy of the mid component 109 results in a negative value, the hyper-mid component M1 is clamped to a zero value. In some embodiments, the hyper-mid component M1 is wrapped around by taking the absolute value of the negative value as the value of the hyper-mid component M1. When subtracting the spectral energy of the side component 111 from the spectral energy of the mid component 109 results in a negative value for M1, other types of processing may be used. When the subtraction that produces the hyper-side component S1, the residual side component S2, or the residual mid component M2 results in a negative value, similar additional processing, such as clamping to zero, wraparound, or other processing, may be used. Fixing the hypermid component M1 to 0 ensures spectral orthogonality between M1 and both side components when the subtraction results in a negative value. Similarly, fixing the hyperside component S1 to 0 ensures spectral orthogonality between S1 and both mid components when the subtraction results in a negative value. The residual mid M2 and residual side S2 components derived by creating orthogonality between the hypermid and hyperside components and their appropriate mid / side counterparts (i.e., side components for hypermid and mid components for hyperside) contain spectral energy that is not orthogonal to (i.e., common to) their appropriate mid / side counterparts. That is, applying a fix of 0 to the hypermid to derive the residual mid and using the M1 component produces a hypermid component that has no spectral energy in common with the side component and a residual mid component that has sufficient spectral energy in common with the side component. When the hyperside is fixed to 0, the same relationship applies to the hyperside and residual side.When applying frequency-domain processing, there is generally a resolution trade-off between frequency and timing information. As frequency resolution increases (i.e., as the FFT window size and number of frequency bins increase), time resolution decreases, and vice versa. The spectral subtraction described above is performed per frequency bin, so having a large FFT window size (e.g., 8192 samples, resulting in 4096 frequency bins, assuming a real-valued input signal) may be preferable in some situations, such as when removing vocal energy from hyper-mid components. Other situations may require greater time resolution, and therefore lower overall latency, and lower frequency resolution (e.g., an FFT window size of 512 samples, resulting in 256 frequency bins, assuming a real-valued input signal). In the latter case, low mid and side frequency resolution may produce audible spectral artifacts when subtracted from each other to derive the hyper-mid M1 and hyper-side S1 components, because the spectral energy of each frequency bin is an average representation of energy over too large a frequency range. In this case, taking the absolute value of the difference between mid and side when deriving hypermid M1 or hyperside S1 can help reduce perceptual artifacts by allowing for frequency bin-by-frequency bin deviations from true orthogonality in the components. In addition to, or instead of, wrapping around to 0, a factor may be applied to the subtraction value to scale it between 0 and 1, thus providing a method of interpolation between perfectly orthogonal hyper and residual mid / side components at one pole (i.e., a value of 1) and hypermid M1 and hyperside S1 that are identical to the corresponding original mid and side components at the other pole (i.e., a value of 0).

[0027] The subtraction unit 215 removes the spectral energy of the mid component 109 in the frequency domain from the spectral energy of the side component 111 in the frequency domain, while retaining only the phase, to generate the hyper side component S1. For example, the subtraction unit 215 subtracts the magnitude of the mid component 109 in the frequency domain from the magnitude of the side component 111 in the frequency domain, while retaining only the phase, to generate the hyper side component S1. The subtraction unit 219 removes the spectral energy of the hyper side component S1 from the spectral energy of the side component 111 to generate the residual side component S2. For example, the subtraction unit 219 subtracts the magnitude of the hyper side component S1 in the frequency domain from the magnitude of the side component 111 in the frequency domain, while retaining only the phase, to generate the residual side component S2.

[0028] 2B, quadrature component generator module 223 is similar to quadrature component generator module 213 in that it receives mid component 109 and side component 111 and generates hypermid component M1, residual mid component M2, hyperside component S1, and residual side component S2. Quadrature component generator module 223 differs from quadrature generator module 213 by generating hypermid component M1 and hyperside component S1 in the frequency domain and then transforming these components back to the time domain to generate residual mid component M2 and residual side component S2. The quadrature component generator module 223 includes a forward FFT unit 220, a bandpass unit 222, a subtraction unit 224, a hypermid processor 225, an inverse FFT unit 226, a time delay unit 228, a subtraction unit 230, a forward FFT unit 232, a bandpass unit 234, a subtraction unit 236, a hyperside processor 237, an inverse FFT unit 240, a time delay unit 242, and a subtraction unit 244.

[0029] The forward fast Fourier transform (FFT) unit 220 applies a forward FFT to the mid component 109, transforming it into the frequency domain. The transformed mid component 109 in the frequency domain includes magnitude and phase. The bandpass unit 222 applies a bandpass filter to the frequency-domain mid component 109, where the bandpass filter specifies frequencies in the hyper-mid component M1. For example, to isolate a typical human vocal range, the bandpass filter may specify frequencies between 300 Hz and 8000 Hz. In another example, to remove audio content associated with a typical human vocal range, the bandpass filter may preserve lower frequencies (e.g., generated by a bass guitar or drums) and higher frequencies (e.g., generated by a cymbal) in the hyper-mid component M1. In other embodiments, the quadrature component generator module 223 applies various other filters to the frequency-domain mid component 109 in addition to and / or instead of the bandpass filter applied by the bandpass unit 222. In some embodiments, the quadrature component generator module 223 does not include a band-pass unit 222 and does not apply any filter to the frequency-domain mid component 109. In the frequency domain, a subtraction unit 224 subtracts the side component 111 from the filtered mid component to generate the hyper-mid component M1. In other embodiments, in addition to and / or instead of subsequent processing applied to the hyper-mid component M1, such as that performed by a quadrature component processor module (e.g., the quadrature component processor module of FIG. 3), the quadrature component generator module 223 applies various audio enhancements to the frequency-domain hyper-mid component M1. The hyper-mid processor 225 performs processing on the hyper-mid component M1 in the frequency domain before transforming it to the time domain. The processing may include subband spatial processing and / or crosstalk compensation processing.In some embodiments, the hyper-mid processor 225 performs processing on the hyper-mid component M1 instead of and / or in addition to processing that may be performed by the quadrature component processor module 117. The inverse FFT unit 226 applies an inverse FFT to the hyper-mid component M1 and transforms it back to the time domain. The hyper-mid component M1 in the frequency domain contains the magnitude of M1 and the phase of the mid component 109, which the inverse FFT unit 226 transforms back to the time domain. The time delay unit 228 applies a time delay to the mid component 109 so that the mid component 109 and the hyper-mid component M1 arrive at the subtraction unit 230 at the same time. The subtraction unit 230 subtracts the hyper-mid component M1 in the time domain from the time-delayed mid component 109 in the time domain to generate a residual mid component M2. In this example, the spectral energy of the hyper-mid component M1 is removed from the spectral energy of the mid component 109 using processing in the time domain.

[0030] The forward FFT unit 232 applies a forward FFT to the side component 111 to transform it into the frequency domain. The transformed side component 111 in the frequency domain includes magnitude and phase. The band-pass unit 234 applies a band-pass filter to the frequency-domain side component 111. The band-pass filter specifies the frequency in the hyper-side component S1. In other embodiments, the quadrature component generator module 223 applies various other filters to the frequency-domain side component 111 in addition to and / or instead of the band-pass filter. In the frequency domain, the subtraction unit 236 subtracts the mid component 109 from the filtered side component 111 to generate the hyper-side component S1. In other embodiments, the quadrature component generator module 223 applies various audio enhancements to the hyper-side component S1 in the frequency domain in addition to and / or instead of subsequent processing applied to the hyper-side component S1, such as performed by a quadrature component processor (e.g., the quadrature component processor module of FIG. 3 ). The hyperside processor 237 performs processing on the hyperside component S1 in the frequency domain before transforming it to the time domain. The processing may include subband spatial processing and / or crosstalk compensation processing. In some embodiments, the hyperside processor 237 performs processing on the hyperside component S1 instead of and / or in addition to processing that may be performed by the quadrature component processor module 117. The inverse FFT unit 240 applies an inverse FFT to the hyperside component S1 in the frequency domain to generate the hyperside component S1 in the time domain. The hyperside component S1 in the frequency domain contains the magnitude of S1 and the phase of the side component 111, which the inverse FFT unit 226 transforms to the time domain. The time delay unit 242 time-delays the side component 111 so that it arrives at the subtraction unit 244 at the same time as the hyperside component S1. Then, a subtraction unit 244 subtracts the hyper side component S1 in the time domain from the time-delayed side component 111 in the time domain to generate a residual side component S2.In this example, the spectral energy of the hyperside component S1 is removed from the spectral energy of the side component 111 using processing in the time domain.

[0031] In some embodiments, the hyper mid processor 225 and hyper side processor 237 may be omitted if the processing performed by these components is performed by the orthogonal component processor module 117 .

[0032] In FIG. 2C, the quadrature component generator module 245 is similar to the quadrature component generator module 223 in that it receives the mid component 109 and the side component 111 and generates the hyper-mid component M1, the residual mid component M2, the hyper-side component S1, and the residual side component S2, except that the quadrature component generator module 245 generates each of the components M1, M2, S1, and S2 in the frequency domain and then converts these components to the time domain. The quadrature component generator module 245 includes a forward FFT unit 247, a band-pass unit 249, a subtraction unit 251, a hyper mid processor 252, a subtraction unit 253, a residual mid processor 254, an inverse FFT unit 255, an inverse FFT unit 257, a forward FFT unit 261, a band-pass unit 263, a subtraction unit 265, a hyper mid processor 266, a subtraction unit 267, a residual mid processor 268, an inverse FFT unit 269, and an inverse FFT unit 271.

[0033] The forward FFT unit 247 applies a forward FFT to the mid component 109, transforming it into the frequency domain. The transformed mid component 109 in the frequency domain includes a magnitude and a phase. The forward FFT unit 261 applies a forward FFT to the side component 111, transforming it into the frequency domain. The transformed side component 111 in the frequency domain includes a magnitude and a phase. The bandpass unit 249 applies a bandpass filter to the frequency domain mid component 109, which specifies the frequency of the hyper-mid component M1. In some embodiments, the quadrature component generator module 245 applies various other filters to the frequency domain mid component 109 in addition to and / or instead of the bandpass filter. The subtraction unit 251 subtracts the frequency domain side component 111 from the frequency domain mid component 109 to generate the hyper-mid component M1 in the frequency domain. The hyper-mid processor 252 performs processing on the hyper-mid component M1 in the frequency domain before transforming it to the time domain. In some embodiments, the hyper-mid processor 252 performs subband spatial processing and / or crosstalk compensation processing. In some embodiments, the hyper-mid processor 252 performs processing on the hyper-mid component M1 instead of and / or in addition to processing that may be performed by the quadrature component processor module 117. The inverse FFT unit 257 applies an inverse FFT to the hyper-mid component M1 and transforms it back to the time domain. The hyper-mid component M1 in the frequency domain includes the magnitude of M1 and the phase of the mid component 109, which the inverse FFT unit 257 transforms back to the time domain. The subtraction unit 253 subtracts the hyper-mid component M1 from the mid component 109 in the frequency domain to generate a residual mid component M2. The residual mid processor 254 performs processing on the residual mid component M2 in the frequency domain before transforming it to the time domain. In some embodiments, the residual mid processor 254 performs subband spatial processing and / or crosstalk compensation processing on the residual mid component M2.In some embodiments, the residual mid processor 254 performs processing on the residual mid component M2 instead of and / or in addition to processing that may be performed by the quadrature component processor module 117. The inverse FFT unit 255 applies an inverse FFT to transform the residual mid component M2 into the time domain. The residual mid component M2 in the frequency domain includes the magnitude of M2 and the phase of the mid component 109, which the inverse FFT unit 255 transforms into the time domain.

[0034] The bandpass unit 263 applies a bandpass filter to the frequency-domain side component 111. The bandpass filter specifies the frequency in the hyperside component S1. In other embodiments, the quadrature component generator module 245 applies various other filters to the frequency-domain side component 111 in addition to and / or instead of the bandpass filter. In the frequency domain, the subtraction unit 265 subtracts the mid component 109 from the filtered side component 111 to generate the hyperside component S1. The hyperside processor 266 performs processing on the hyperside component S1 in the frequency domain before transforming it to the time domain. In some embodiments, the hyperside processor 266 performs subband spatial processing and / or crosstalk compensation processing on the hyperside component S1. In some embodiments, the hyperside processor 266 performs processing on the hyperside component S1 instead of and / or in addition to processing that may be performed by the quadrature component processor module 117. The inverse FFT unit 271 applies an inverse FFT to transform the hyperside component S1 back to the time domain. The hyper-side component S1 in the frequency domain includes the magnitude of S1 and the phase of the side component 111, which the inverse FFT unit 271 transforms to the time domain. The subtraction unit 267 subtracts the hyper-side component S1 from the side component 111 in the frequency domain to generate a residual side component S2. The residual side processor 268 performs processing on the residual side component S2 in the frequency domain before transforming it to the time domain. In some embodiments, the residual side processor 268 performs subband spatial processing and / or crosstalk compensation processing on the residual side component S2. In some embodiments, the residual side processor 268 performs processing on the residual side component S2 instead of and / or in addition to processing that may be performed by the quadrature component processor module 117. The inverse FFT unit 269 applies an inverse FFT to the residual side component S2 and transforms it to the time domain.The residual side component S2 in the frequency domain contains the magnitude of S2 and the phase of the side component 111, which the inverse FFT unit 269 transforms into the time domain.

[0035] In some embodiments, the hyper mid processor 252, hyper side processor 266, residual mid processor 254, or residual side processor 268 may be omitted if the processing performed by these components is performed by the quadrature component processor module 117.

[0036] Exemplary Quadrature Component Processor 3 is a block diagram of a quadrature component processor module 317 according to one or more embodiments. The quadrature component processor module 317 is an example of the quadrature component processor module 117. The quadrature component processor module 317 may include a subband spatial processing and / or crosstalk compensation processing unit 320, a summing unit 325, and a summing unit 330. The quadrature component processor module 317 performs subband spatial processing and / or crosstalk compensation processing on at least one of the hyper-mid component M1, the residual mid component M2, the hyper-side component S1, and the residual side component S2. As a result of the subband spatial processing and / or crosstalk compensation processing 320, the quadrature component processor module 317 outputs at least one of processed M1, processed M2, processed S1, and processed S2. Summing unit 325 adds processed M1 and processed M2 to generate processed mid component 131, and summing unit 330 adds processed S1 and processed S2 to generate processed side component 139.

[0037] In some embodiments, the quadrature component processor module 317 performs subband spatial processing and / or crosstalk compensation processing 320 on at least one of the hyper-mid component M1, the residual mid component M2, the hyper-side component S1, and the residual side component S2 in the frequency domain to generate the processed mid component 131 and the processed side component 139 in the frequency domain. The quadrature component generator module 113 may provide the frequency-domain components M1, M2, S1, or S2 to a quadrature component processor, which performs an inverse FFT. After generating the processed mid component 131 and the processed side component 139, the quadrature component processor module 317 may perform an inverse FFT on the processed mid component 131 and the processed side component 139 to transform these components back to the time domain. In some embodiments, the quadrature component processor module 317 performs an inverse FFT on the processed M1, the processed M2, the processed S1, and the processed S2 to generate the processed mid component 131 and the processed side component 139 in the time domain.

[0038] Examples of the quadrature component processor module 317 are shown in Figures 4 and 5. In some embodiments, the quadrature component processor module 317 performs both subband spatial processing and crosstalk compensation processing. The processing performed by the quadrature component processor module 317 is not limited to subband spatial processing or crosstalk compensation processing. Any type of spatial processing using mid / side space, such as by using hyper-mid components instead of mid components or hyper-side components instead of side components, may be performed by the quadrature component processor module 317. Some other types of processing may include gain application, amplitude or delay-based panning, binaural processing, reverberation, dynamic range processing such as compression and limiting, and other linear or nonlinear audio processing techniques and effects, ranging from chorus or flanging to machine learning-based approaches to vocal or instrumental style transfer, transformation, or resynthesis.

[0039] Exemplary Subband Spatial Processor 4 is a block diagram of a subband spatial processor module 410 according to one or more embodiments. The subband spatial processor module 410 is an example of the quadrature component processor module 317. The subband spatial processor module 410 includes a mid EQ filter 404(1), a mid EQ filter 404(2), a mid EQ filter 404(3), a mid EQ filter 404(4), a side EQ filter 406(1), a side EQ filter 406(2), a side EQ filter 406(3), and a side EQ filter 406(4). In some embodiments, the subband spatial processor module 410 includes components in addition to and / or instead of those described herein.

[0040] The subband spatial processor module 410 processes the non-spatial components Y m and the spatial component Y s and gain-adjusting one or more subbands of these components to provide spatial enhancement. m can be the hypermid component M1 or the residual mid component M2. s can be the hyperside component S1, or the residual side component S2.

[0041] The subband spatial processor module 410 processes the non-spatial components Y m and receives the enhanced non-spatial component E m To generate Y m The subband spatial processor module 410 also applies mid EQ filters 404(1) to 404(4) to the different subbands of the spatial component Y s and receives the enhanced spatial component E s To generate Y sThe subband spatial processor module 410 applies side EQ filters 406(1) through 406(4) to different subbands of the non-spatial component Y. The subband filters may include various combinations of peak filters, notch filters, low-pass filters, high-pass filters, low-shelf filters, high-shelf filters, band-pass filters, band-stop filters, and / or all-pass filters. The subband filters may also apply a gain to each subband. More specifically, the subband spatial processor module 410 processes the non-spatial component Y. m and a subband filter for each of the n frequency subbands of the spatial component Y s For example, for n=4 subbands, the subband spatial processor module 410 generates a non-spatial component Y , including a mid equalization (EQ) filter 404(1) for subband(1), a mid EQ filter 404(2) for subband(2), a mid EQ filter 404(3) for subband(3), and a mid EQ filter 404(4) for subband(4). m Each mid EQ filter 404 includes a series of subband filters for the enhanced non-spatial component E m To generate the non-spatial component Y m The filter is applied to the frequency subband portion of

[0042] The subband spatial processor module 410 processes the spatial component Y , which includes a side equalization (EQ) filter 406(1) for subband(1), a side EQ filter 406(2) for subband(2), a side EQ filter 406(3) for subband(3), and a side EQ filter 406(4) for subband(4). s Each side EQ filter 406 further includes a series of subband filters for the frequency subbands E s To generate the spatial component Y s The filter is applied to the frequency subband portion of

[0043] non-spatial component Ym and the spatial component Y s Each of the n frequency subbands may correspond to a frequency range. For example, frequency subband (1) may correspond to 0 Hz to 300 Hz, frequency subband (2) may correspond to 300 Hz to 510 Hz, frequency subband (3) may correspond to 510 Hz to 2700 Hz, and frequency subband (4) may correspond to 2700 Hz to the Nyquist frequency. In some embodiments, the n frequency subbands are a concatenated set of critical bands. The critical bands may be determined using a corpus of audio samples from a wide variety of musical genres. The long-term average energy ratio of mid components to side components across the 24 Bark scale critical bands is determined from the samples. Contiguous frequency bands with similar long-term average ratios are then grouped together to form the set of critical bands. The range of the frequency subbands and the number of frequency subbands may be adjustable.

[0044] In some embodiments, the subband spatial processor module 410 converts the residual mid component M2 into the non-spatial component Y m and one of the side component, hyperside component S1, or residual side component S2 is processed as the spatial component Y s Use as.

[0045] In some embodiments, the subband spatial processor module 410 processes one or more of the hypermid component M1, hyperside component S1, residual mid component M2, and residual side component S2. The filters applied to the subbands of each of these components may be different. The hypermid component M1 and residual mid component M2 each process a non-spatial component Y m The hyper-side component S1 and the residual side component S2 can be processed as described for the spatial component Y s can be processed as described for

[0046] Exemplary Crosstalk Compensation Processor 5 is a block diagram of a crosstalk compensation processor module 510 in accordance with one or more embodiments. The crosstalk compensation processor module 510 is an example of the quadrature component processor module 317. The crosstalk compensation processor module 510 includes a mid component processor 520 and a side component processor 530. The crosstalk compensation processor module 510 processes the non-spatial component Y m and the spatial component Y s and apply a filter to one or more of these components to compensate for spectral defects caused by (e.g., subsequent or preceding) crosstalk processing. m can be the hypermid component M1 or the residual mid component M2. s can be the hyperside component S1, or the residual side component S2.

[0047] The crosstalk compensation processor module 510 calculates the non-spatial component Y m and the mid component processor 520 receives the enhanced non-spatial crosstalk compensated component Z m The crosstalk compensation processor module 510 also applies a set of filters to generate the spatial subband components Y s and receiving the enhanced spatial subband components E s A set of filters is applied in the side component processor 530 to generate the non-spatial component Y . The mid component processor 520 includes multiple filters 540, such as m mid filters 540(a), 540(b) through 540(m), where each of the m mid filters 540 filters the non-spatial component Y . m The mid component processor 520 processes one of the m frequency bands of the non-spatial component Y m By processing the mid-crosstalk compensation channel Z m In some embodiments, the mid-filter 540 generates a non-spatial Y mThe frequency response plot is constructed using the frequency response plot of the crosstalk compensation channel Z. Additionally, by analyzing the frequency response plot, any spectral defects, such as peaks or troughs in the frequency response plot that exceed a predetermined threshold (e.g., 10 dB), that occur as artifacts of the crosstalk processing can be estimated. These artifacts result primarily from the summation of delayed and possibly inverted contralateral signals with their corresponding ipsilateral signals in the crosstalk processing, thereby effectively introducing a comb filter-like frequency response into the final rendered result. To compensate for the estimated peaks or troughs, a mid-crosstalk compensation channel Z m can be generated by the mid component processor 520, with each of the m frequency bands corresponding to a peak or trough. Specifically, based on the specific delays, filtering frequencies, and gains applied in the crosstalk processing, the peaks and troughs shift up or down in the frequency response, causing variable amplification and / or attenuation of energy in specific regions of the spectrum. Each of the mid filters 540 can be configured to adjust one or more of the peaks and troughs.

[0048] The side component processor 530 includes multiple filters 550, such as m side filters 550(a), 550(b) through 550(m). The side component processor 530 processes the spatial component Y s By processing the side crosstalk compensation channel Z s In some embodiments, a spatial Y s A frequency response plot of can be obtained through simulation. By analyzing the frequency response plot, any spectral defects, such as peaks or troughs in the frequency response plot that exceed a predetermined threshold (e.g., 10 dB), that occur as artifacts of crosstalk processing can be estimated. To compensate for the estimated peaks or troughs, a side crosstalk compensation channel Z scan be generated by the side component processor 530. Specifically, based on the particular delays, filtering frequencies, and gains applied in the crosstalk processing, the peaks and troughs shift up or down in the frequency response, causing variable amplification and / or attenuation of energy in particular regions of the spectrum. Each of the side filters 550 can be configured to adjust one or more of the peaks and troughs. In some embodiments, the mid component processor 520 and the side component processor 530 can include different numbers of filters.

[0049] In some embodiments, the mid-filter 540 and the side-filter 550 may include biquad filters having a transfer function defined by Equation 1:

[0050]

number

[0051] where z is a complex variable and a0, a1, a2, b0, b1, and b2 are digital filter coefficients. One way to implement such a filter is the direct form I topology, as defined in Equation 2.

[0052]

number

[0053] where X is the input vector and Y is the output. Other topologies can be used depending on their maximum word length and saturation behavior. A biquad can then be used to implement a second-order filter with real-valued inputs and outputs. To design a discrete-time filter, a continuous-time filter is designed and then transformed to discrete time via a bilinear transformation. Furthermore, the resulting shift in center frequency and bandwidth can be compensated for using frequency warping.

[0054] For example, the peaking filter may have an S-plane transfer function defined by Equation 3:

[0055]

number

[0056] where s is a complex variable, A is the amplitude of the peak, Q is the filter "quality", and the digital filter coefficients are

[0057]

number

[0058] is defined by

[0059] where ω is the center frequency of the filter in radians,

[0060]

number

[0061] Furthermore, the filter quality Q can be defined by Equation 4:

[0062]

number

[0063] where Δf is the bandwidth and fc is the center frequency. The mid filter 540 is shown as being in series and the side filter 550 is shown as being in series. In some embodiments, the mid filter 540 filters the mid component Y m are applied in parallel to the side components Y s are applied in parallel to

[0064] In some embodiments, the crosstalk compensation processor module 510 processes each of the hyper-mid component M1, hyper-side component S1, residual mid component M2, and residual side component S2. The filters applied to each of these components may be different.

[0065] Exemplary Crosstalk Processor 6 is a block diagram of a crosstalk simulation processor module 600 according to one or more embodiments. As discussed with respect to FIG. 1, in some embodiments, the audio processing system 100 includes a crosstalk processor module 141 that applies crosstalk processing to the processed left component 151 and the processed right component 159. The crosstalk processing includes, for example, crosstalk simulation and crosstalk cancellation. In some embodiments, the crosstalk processor module 141 includes a crosstalk simulation processor module 600. The crosstalk simulation processor module 600 generates contralateral sound components for output to stereo headphones, thereby providing a loudspeaker-like listening experience in the headphones. Left input channel X L may be the processed left component 151 and the right input channel X R may be the processed right component 159. In some embodiments, a crosstalk simulation may be performed before the quadrature component processing.

[0066] The crosstalk simulation processor module 600 is L , includes a left head shadow low pass filter 602, a left head shadow high pass filter 624, a left crosstalk delay 604, and a left head shadow gain 610. The crosstalk simulation processor module 600 processes the right input channel X RThe left head shadow low pass filter 602 and the left head shadow high pass filter 624 apply a modulation that models the frequency response of the signal after passing through the listener's head to the left input channel X. L The output of the left head shadow high pass filter 624 is provided to the left crosstalk delay 604, which applies a time delay representing the transaural distance traversed by the contralateral sound component relative to the ipsilateral sound component. The left head shadow gain 610 is applied to the right and left simulation channels W L A gain is applied to the output of the left crosstalk delay 604 to generate

[0067] Right Input Channel X R Similarly, the right head shadow low-pass filter 606 and the right head shadow high-pass filter 626 apply modulation that models the frequency response of the listener's head to the right input channel X R The output of the right head shadow high pass filter 626 is provided to the right crosstalk delay 608, which applies a time delay. The right head shadow gain 612 is applied to the right crosstalk simulation channel W R A gain is applied to the output of the right crosstalk delay 608 to generate

[0068] The application of the head shadow low pass filter, head shadow high pass filter, crosstalk delay, and head shadow gain to each of the left and right channels may be performed in different orders.

[0069] 7 is a block diagram of a crosstalk cancellation processor module 700 according to one or more embodiments. The crosstalk processor module 141 may include the crosstalk cancellation processor module 700. The crosstalk cancellation processor module 700 is configured to process the left input channel XL and right input channel X R and receives the left output channel O L and the right output channel O R and channel X to generate L , X R Perform crosstalk cancellation on the left input channel X L may be the processed left component 151 and the right input channel X R may be the processed right component 159. In some embodiments, crosstalk cancellation may be performed before the quadrature component processing.

[0070] The crosstalk cancellation processor module 700 includes an in-out-band divider 710, inverters 720 and 722, contralateral estimators 730 and 740, combiners 750, 752, and an in-out-band combiner 760. These components are used to divide the input channel T L , T R into in-band and out-of-band components and output channels O L , O R , which operate together to perform crosstalk cancellation on the in-band components to generate

[0071] By dividing the input audio signal T into different frequency band components and performing crosstalk cancellation on selected components (e.g., in-band components), crosstalk cancellation can be performed on specific frequency bands while avoiding degradation in other frequency bands. If crosstalk cancellation were performed without dividing the input audio signal T into different frequency bands, the audio signal after such crosstalk cancellation may exhibit significant attenuation or amplification of non-spatial and spatial components at low frequencies (e.g., below 350 Hz), higher frequencies (e.g., above 12000 Hz), or both. By selectively performing crosstalk cancellation on the in-band (e.g., between 250 Hz and 14000 Hz), where most of the influential spatial cues reside, a balanced overall energy across the spectrum in the mix, especially in the non-spatial components, can be maintained.

[0072] The in-out band divider 710 divides the input channel T L , T R respectively for the in-band channel T L,In ,T R,In and out-of-band channel T L,Out ,T R,Out In particular, the in-out band splitter 710 splits the enhanced left compensation channel T L Left in-band channel T L,In and left out-of-band channel T L,Out Similarly, the in-out band divider 710 divides the enhanced right compensation channel T R Right in-band channel T R,In and right out-of-band channel T R,Out Each in-band channel may contain a portion of the respective input channel corresponding to a frequency range, for example, including 250 Hz to 14 kHz. The range of the frequency band may be adjustable, for example, according to speaker parameters.

[0073] The inverter 720 and the contralateral estimator 730 generate the left in-band channel T L,In To compensate for the contralateral sound component due to L Similarly, inverter 722 and contralateral estimator 740 operate together to generate the right in-band channel T R,In To compensate for the contralateral sound component due to R They work together to produce

[0074] In one approach, inverter 720 is connected to in-band channel T L,In Receives the inverted in-band channel T L,In ', the received in-band channel T L,In The contralateral estimator 730 inverts the polarity of the inverted in-band channel T L,In ' and, through filtering, the inverted in-band channel T corresponding to the contralateral sound component. L,In The filtering is performed by extracting the inverted in-band channel T L,In ', the portion extracted by the contralateral estimator 730 is attributed to the contralateral sound component, the in-band channel T L,In Therefore, the part extracted by the contralateral estimator 730 is the left contralateral cancellation component S L This means that the in-band channel T L,In To reduce the contralateral sound component caused by the corresponding in-band channel T R,In In some embodiments, the inverter 720 and the contralateral estimator 730 are performed in a different order.

[0075] The inverter 722 and the contralateral estimator 740 generate the right contralateral cancellation component S R To generate the in-band channel T R,In, and therefore a detailed description thereof will be omitted herein for the sake of brevity.

[0076] In one exemplary implementation, the contralateral estimator 730 includes a filter 732, an amplifier 734, and a delay unit 736. The filter 732 receives the inverted input channel T L,In ', and through filtering, the inverted in-band channel T corresponding to the contralateral sound component is generated. L,In An exemplary filter implementation is a notch or high shelf filter with a center frequency selected between 5000 Hz and 10000 Hz and a Q selected between 0.5 and 1.0. The gain (G dB ) can be derived from Equation 5. G dB =-3.0-log 1.333 (D) Equation (5)

[0077] where D is the delay by delay units 736 and 646 in samples, for example at a sampling rate of 48 KHz. An alternative implementation is a low pass filter with a corner frequency selected between 5000 Hz and 10000 Hz and a Q selected between 0.5 and 1.0. Furthermore, amplifier 734 applies the extracted portion with a corresponding gain factor G L,In and delay unit 736 amplifies the left contralateral cancellation component S L The amplified output from amplifier 734 is delayed according to a delay function D to generate the right contralateral cancellation component S R The inverted in-band channel T R,In ', an amplifier 744, and a delay unit 776. In one example, the contralateral estimators 730, 740 calculate the left contralateral cancellation component S according to the following equation: L and the right contralateral cancellation component S R and generate. S L =D[GL,In *F[T L,In ']] Formula (6) S R =D[G R,In *F[T R,In ']] Formula (7)

[0078] where F[] is the filter function and D[] is the delay function.

[0079] The crosstalk cancellation configuration can be determined by speaker parameters. In one example, the filter center frequency, delay, amplifier gain, and filter gain can be determined according to the angle formed between the two speakers with respect to the listener. In some embodiments, values ​​between speaker angles are used to interpolate other values.

[0080] Combiner 750 combines the left in-band crosstalk channel U L To generate the right contralateral cancellation component S R Left in-band channel T L,In and combiner 752 combines the right in-band crosstalk channel U R To generate the left contralateral cancellation component S L Right in-band channel T R,In The in-out band combiner 760 couples the left output channel O L To generate the left in-band crosstalk channel U L out-of-band channel T L,Out Combined with the right output channel R To generate the right in-band crosstalk channel U R out-of-band channel T R,Out and combine.

[0081] Therefore, the left output channel O L is attributed to the contralateral sound, in-band channel T R,In The right contralateral cancellation component S corresponds to the inversion of the RRight output channel O R is attributed to the contralateral sound, in-band channel T L,In The left contralateral cancellation component S corresponds to the inversion of the L In this configuration, the right output channel O R According to L Similarly, the wavefront of the left output channel O arriving at the left ear can be cancelled according to L According to R The wavefront of the contralateral sound component output by the right loudspeaker can be cancelled according to: Thus, the contralateral sound component can be reduced to increase spatial detectability.

[0082] Cartesian component spatial processing 8 is a flowchart of a process for spatial processing using at least one of a hyper-mid component, a residual mid component, a hyper-side component, or a residual side component, according to one or more embodiments. The spatial processing may include, among other things, gain application, amplitude or delay-based panning, binaural processing, reverberation, dynamic range processing such as compression and limiting, linear or nonlinear audio processing techniques and effects, chorus effects, flanging effects, vocal or instrumental style transfer, transformation, or machine learning-based approaches to resynthesis. The process may be performed to provide spatially enhanced audio to a user's device. The process may include fewer or additional steps, and the steps may be performed in a different order.

[0083] An audio processing system (e.g., audio processing system 100) receives 810 an input audio signal (e.g., left input channel 103 and right input channel 105). In some embodiments, the input audio signal may be a multi-channel audio signal including multiple left-right channel pairs. Each left-right channel pair may be processed as described herein for the left and right input channels.

[0084] The audio processing system generates 820 a non-spatial mid component (e.g., mid component 109) and a spatial side component (e.g., side component 111) from the input audio signal. In some embodiments, an L / R to M / S converter (e.g., L / R to M / S converter module 107) performs the conversion of the input audio signal into the mid and side components.

[0085] The audio processing system generates 830 at least one of a hyper-mid component (e.g., hyper-mid component M1), a hyper-side component (e.g., hyper-side component S1), a residual mid component (e.g., residual mid component M2), and a residual side component (e.g., residual side component S2). The audio processing system may generate at least one and / or all of the components listed above. The hyper-mid component includes the spectral energy of the side component removed from the spectral energy of the mid component. The residual mid component includes the spectral energy of the hyper-mid component removed from the spectral energy of the side component. The residual side component includes the spectral energy of the hyper-side component removed from the spectral energy of the side component. The processing used to generate M1, M2, S1, or S2 may be performed in the frequency domain or the time domain.

[0086] The audio processing system filters 840 at least one of the hyper-mid component, the residual mid component, the hyper-side component, and the residual side component to enhance the audio signal. The filtering may include spatial cue processing, such as by adjusting a frequency-dependent amplitude or frequency-dependent delay of the hyper-mid component, the residual mid component, the hyper-side component, or the residual side component. Some examples of spatial cue processing include amplitude- or delay-based panning or binaural processing.

[0087] The filtering may include dynamic range processing, such as compression or limiting. For example, the hypermid, residual mid, hyperside, or residual side components may be compressed according to a compression ratio when a threshold level for compression is exceeded. In another example, the hypermid, residual mid, hyperside, or residual side components may be limited to a maximum level when a threshold level for limiting is exceeded.

[0088] The filtering may include machine learning-based modifications to the hyper-mid, residual mid, hyper-side, or residual side components. Some examples include machine learning-based vocal or instrumental style transfer, transformation, or resynthesis.

[0089] The filtering of the hyper-mid, residual-mid, hyper-side, or residual-side components may include gain application, reverberation, and other linear or non-linear audio processing techniques and effects ranging from chorus and / or flanging, or other types of processing. In some embodiments, the filtering may include filtering for subband spatial processing and crosstalk compensation, as described in more detail below in connection with FIG. 9.

[0090] The filtering may be performed in the frequency domain or the time domain. In some embodiments, the mid and side components are transformed from the time domain to the frequency domain, the hyper and / or residual components are generated in the frequency domain, filtering is performed in the frequency domain, and the filtered components are transformed into the time domain. In other embodiments, the hyper and / or residual components are transformed into the time domain, and filtering is performed on these components in the time domain.

[0091] The audio processing system generates a left output channel (e.g., left output channel 121) and a right output channel (e.g., right output channel 123) using one or more of the filtered hyper / residual components 850. For example, the M / S to L / R conversion may be performed using a mid component (e.g., processed mid component 131) or a side component (e.g., processed side component 139) generated from at least one of the filtered hyper mid component, the filtered residual mid component, the filtered hyper side component, or the filtered residual side component. In another example, the filtered hyper mid component or the filtered residual mid component may be used as the mid component for the M / S to L / R conversion, or the filtered hyper side component or the residual side component may be used as the side component for the M / S to L / R conversion.

[0092] Quadrature Component Subband Spatial and Crosstalk Processing 9 is a flowchart of a process for subband spatial processing and crosstalk compensation using at least one of a hyper-mid component, a residual mid component, a hyper-side component, or a residual side component, according to one or more embodiments. The crosstalk processing may include crosstalk cancellation or crosstalk simulation. The subband spatial processing provides enhanced spatial detectability to the audio content, such as by creating the perception that sound is directed to the listener from a wide area rather than a specific point in space corresponding to the location of the loudspeakers (e.g., sound field enhancement), thereby creating a more immersive listening experience for the listener. Crosstalk simulation may be used on audio output to headphones to simulate a loudspeaker experience with contralateral crosstalk. Crosstalk cancellation may be used on audio output to loudspeakers to remove the effects of crosstalk interference. Crosstalk compensation compensates for spectral defects caused by crosstalk cancellation or crosstalk simulation. The process may include fewer or additional steps, and the steps may be performed in a different order. The hyper- and residual mid / side components can be manipulated in different ways for different purposes. For example, in the case of crosstalk compensation, targeted subband filtering is applied only to the hyper-mid component M1 (where most of the vocal dialogue energy in many movie content occurs) in an effort to remove spectral artifacts resulting from crosstalk processing in just that component. In the case of sound field enhancement, with or without crosstalk processing, targeted subband gains can be applied to the residual mid component M2 and the residual side component S2.For example, the residual mid component M2 may be attenuated and the residual side component S2 may be conversely amplified, without causing a dramatic overall change in perceived loudness in the final L / R signal, while also avoiding attenuation in the hyper-mid M1 component (e.g., that part of the signal that often contains the majority of the vocal energy), increasing the distance between these components in terms of gain (which, if done well, can increase spatial detectability).

[0093] The audio processing system receives 910 an input audio signal, the input audio signal including a left channel and a right channel. In some embodiments, the input audio signal may be a multi-channel audio signal including multiple left-right channel pairs. Each left-right channel pair may be processed as described herein for the left input channel and the right input channel.

[0094] The audio processing system applies crosstalk processing to the received input audio signal 920. The crosstalk processing includes at least one of crosstalk simulation and crosstalk cancellation.

[0095] In steps 930 through 960, the audio processing system performs subband spatial processing and crosstalk compensation for crosstalk processing using one or more of the hyper-mid component, the hyper-side component, the residual mid component, or the residual side component. In some embodiments, the crosstalk processing may be performed after the processing in steps 930 through 960.

[0096] The audio processing system generates 930 a mid component and a side component from the (eg, crosstalk processed) audio signal.

[0097] The audio processing system generates at least one of a hyper-mid component, a residual mid component, a hyper-side component, and a residual side component 940. The audio processing system may generate at least one and / or all of the components listed above.

[0098] The audio processing system filters at least one subband of the hyper-mid component, the residual mid component, the hyper-side component, and the residual side component 950 to apply subband spatial processing to the audio signal. Each subband may include a range of frequencies, such as may be defined by a set of critical bands. In some embodiments, the subband spatial processing further includes time-delaying at least one subband of the hyper-mid component, the residual mid component, the hyper-side component, and the residual side component.

[0099] The audio processing system filters 960 at least one of the hyper-mid component, the residual mid component, the hyper-side component, and the residual side component to compensate for spectral defects from crosstalk processing of the input audio signal. The spectral defects may include peaks or troughs in a frequency response plot of the hyper-mid component, the residual mid component, the hyper-side component, or the residual side component that exceed a predetermined threshold (e.g., 10 dB) that occur as artifacts of crosstalk processing. The spectral defects may be estimated spectral defects.

[0100] In some embodiments, the filtering of spectral orthogonal components for subband spatial processing in step 950 and the crosstalk compensation in step 960 may be combined into a single filtering operation for each spectral orthogonal component selected for filtering.

[0101] In some embodiments, filtering of the hyper / residual mid / side components for subband spatial processing or crosstalk compensation may be performed in conjunction with filtering for other purposes, such as gain application, amplitude or delay-based panning, binaural processing, reverberation, dynamic range processing such as compression and limiting, linear or non-linear audio processing techniques and effects ranging from chorus and / or flanging, machine learning-based approaches to vocal or instrumental style transfer, transformation, or resynthesis, or other types of processing using any of the hyper-mid, residual mid, hyper-side, and residual side components.

[0102] The filtering may be performed in the frequency domain or the time domain. In some embodiments, the mid and side components are transformed from the time domain to the frequency domain, the hyper and / or residual components are generated in the frequency domain, filtering is performed in the frequency domain, and the filtered components are transformed to the time domain. In other embodiments, the hyper and / or residual components are transformed to the time domain, and filtering is performed on these components in the time domain.

[0103] The audio processing system generates a left output channel and a right output channel from the filtered hyper-mid component 970. In some embodiments, the left output channel and the right output channel are additionally based on at least one of the filtered residual mid component, the filtered hyper-side component, and the filtered residual side component.

[0104] Exemplary Quadrature Component Audio Processing 10-19 are plots illustrating the spectral energy of the mid and side components of an exemplary white noise signal, according to one or more embodiments.

[0105] FIG. 10 illustrates a plot of a white noise signal panned hard left 1000. The left-right white noise signal is converted into a mid component 1005 and a side component 1010 and panned hard left using a constant-power sine / cosine pan-law. When the white noise signal is panned hard left 1000, a user positioned between a pair of left and right loudspeakers perceives the sound as appearing at and / or around the left loudspeaker. The white noise signal split into the left and right input channels of the white noise signal can be converted into a mid component 1005 and a side component 1010 using the L / R to M / S converter module 107. As shown in FIG. 10, when the white noise signal is panned hard left 1000, both the mid component 1005 and the side component 1010 have approximately equal amounts of energy. Similarly, when a white noise signal is subjected to heavy panning to the right (not shown in FIG. 10), the mid and side components have approximately equal amounts of energy.

[0106] FIG. 11 illustrates a plot of a white noise signal panned center left 1100. When the white noise signal is panned center left 1100 using a common constant-power sine / cosine pan law, a user positioned between a pair of left and right loudspeakers perceives the sound as appearing halfway between the user's front and the left loudspeaker. FIG. 11 shows the mid component 1105 and side component 1110 of the center-left panned white noise signal 1100 and the white noise signal 1000 panned hard to the left. Compared to the hard-left panned white noise signal 1000, the mid component 1105 increases by approximately 3 dB, while the side component 1110 decreases by approximately 6 dB. When the white noise signal is panned center right, the mid component 1105 and side component 1110 have energies similar to those shown in FIG. 11.

[0107] 12 illustrates a plot of a white noise signal panned center 1200. When the white noise signal is panned center 1200 using a common constant-power sine / cosine pan law, a user positioned between a pair of left and right loudspeakers perceives the sound as appearing in front of the user (e.g., between the left and right loudspeakers). As shown in FIG. 12, the center-panned white noise signal 1200 has only a mid component 1205.

[0108] From the above examples in Figures 10, 11, and 12, it can be seen that the mid component only contains energy for sounds panned to the center in the signal (i.e., the left and right channels are identical), as shown in Figure 12, but in scenarios where the sounds in the original L / R stream are generally perceived as off-center (i.e., as sounds panned to the left or right of center), as shown in Figures 10 and 11, mid component energy is also present.

[0109] In particular, the above three scenarios, which represent the vast majority of L / R audio use cases, do not encompass the scenario in which the Side constitutes the only energy. This is only the case when the left and right channels are 180 degrees out of phase (i.e., sign-inverted), which is rare in two-channel audio for music and entertainment. Thus, the Mid component is ubiquitous in virtually all two-channel left / right audio streams and constitutes the only energy in center-panned content, while the Side component is present in all but center-panned content and rarely, if ever, serves as the only energy in the signal.

[0110] Quadrature component processing separates and manipulates portions of the mid and side components that are spectrally "orthogonal" to one another. That is, using quadrature component processing, the portion of the mid component that corresponds only to energy present in the center of the sound field (i.e., the hyper-mid component) can be isolated, and similarly, the portion of the side component that corresponds only to energy not present in the center of the sound field (i.e., the hyper-side component) can be isolated. Conceptually, the hyper-mid component is the energy that corresponds to the thin pillar of sound perceived in the center of the sound field, whether through loudspeakers or headphones. Furthermore, using a simple scalar, it is possible to control how "thin" this pillar is, providing an interpolation space from hyper-mid to mid and hyper-side to side. Furthermore, as a by-product of deriving our hyper-mid / side component signals, it is also possible to manipulate the residual signals (e.g., residual mid and residual side components) that combine together with the hyper-mid or hyper-side components to form the original complete mid and side components. Each of these four sub-components of Mid and Side can be processed independently using all manner of manipulation, from simple gain staging, to multi-band EQ, to custom and idiosyncratic effects.

[0111] 13-19 illustrate quadrature processing of a white noise signal. FIG. 13 illustrates plots of a white noise signal 1305 panned center and bandpassed between 20 Hz and 100 Hz (e.g., using an 8th-order Butterworth filter) and a white noise signal 1310 panned hard to the left and bandpassed between 5000 Hz and 10000 Hz (e.g., using an 8th-order Butterworth filter) without quadrature processing. The plots show the mid component 1315 and the side component 1320 for each of the panned white noise signals 1305 and 1310. The center-panned white noise signal 1305 has energy only in its mid component 1315, while the white noise signal panned hard to the left has equal amounts of energy in its mid component 1315 and its side component 1320. This is similar to the results shown in FIGS. 10 and 12.

[0112] Figure 14 illustrates the panned white noise signals 1305 and 1310 of Figure 13 with the energy of the side component 1320 removed. The center-panned low band of the white noise of signal 1305 remains unchanged. The high band panned hard to the left of the white noise of signal 1310 now has zero side energy, but a portion of the energy represented by the mid component 1315 is still present. Even though the side energy has been removed, there is still some non-center-panned energy present in the mid signal, as shown by signal 1310.

[0113] Figure 15 illustrates the panned white noise signal 1500 of Figure 13 using quadrature component processing. In particular, quadrature component processing is used to isolate the hyper-mid component 1510 and remove other energy in the audio signal. Here, the signal panned heavily to the left has been removed, leaving only the center-panned signal 1500. This shows that the hyper-mid component 1510 is an isolation of only the energy in the signal occupying the center of the sound field, and not anything else.

[0114] Because it is possible to isolate the hyper-mid components of an audio signal, the audio signal can be manipulated to control which elements of the original signal become the various M1 / M2 / S1 / S2 components. These preprocessing operations can range from simple amplitude and delay adjustments to more complex filtering techniques. These preprocessing operations can then be reversed to restore the original sound field.

[0115] Figure 16 illustrates another embodiment of the panned white noise signal 1600 of Figure 13 using quadrature component processing. The L / R audio signal is rotated in such a way as to center the high-band white noise panned hard to the left (e.g., as shown by signal 1310 in Figure 13) in the sound field and shift the center-panned low-band noise (e.g., as shown by signal 1305 in Figure 13) farther from the center. The white noise signal, originally panned hard to the left and bandpassed between 5000 Hz and 10000 Hz 1600, can then be extracted and further processed by isolating the hyper-mid component 1610 of the rotated L / R signal.

[0116] 17 shows a decorrelated white noise signal 1700. The input white noise signal 1700 may be a two-channel quadrature white noise signal including a right channel component 1710 and a left channel component 1720. The plot also shows a mid component 1730 and a side component 1740 generated from the white noise signal. The spectral energy of the left channel component 1720 matches that of the right channel component 1710, and the spectral energy of the mid component 1730 matches that of the side component 1740. The mid component 1730 and the side component 1740 have signal levels approximately 3 dB lower than the right channel component 1710 and the left channel component 1720.

[0117] FIG. 18 illustrates a mid component 1730 decomposed into a hyper-mid component 1810 and a residual mid component 1820. The mid component 1730 represents the non-spatial information of the input audio signal in the sound field. The hyper-mid component 1810 contains subcomponents of non-spatial information found directly in the center of the sound field, while the residual mid component 1820 is the remaining non-spatial information. In a typical stereo audio signal, the hyper-mid component 1810 may contain the main features of the audio signal, such as dialogue or vocals. In FIG. 18, the residual mid component 1820 is approximately 3 dB lower than the mid component 1730, and the hyper-mid component 1810 is approximately 8-9 dB lower than the mid component 1730.

[0118] FIG. 19 illustrates the side components 1740 decomposed into hyper-side components 1910 and residual side components 1920. The side components 1740 represent spatial information in the input audio signal in the sound field. The hyper-side components 1910 include sub-components of spatial information found at the edges of the sound field, while the residual side components 1920 are the remaining spatial information. In a typical stereo audio signal, the residual side components 1920 include key features resulting from processing, such as the effects of binaural processing, panning techniques, reverberation, and / or decorrelation processes. As shown in FIG. 19 , the relationship between the side components 1740, hyper-side components 1910, and residual side components 1920 is similar to that between the mid components 1730, hyper-mid components 1810, and residual side components 1820.

[0119] Computing Machine Architecture FIG. 20 is a block diagram of a computer system 2000 according to one or more embodiments. The computer system 2000 is an example of a circuit for implementing an audio processing system. Illustrated is at least one processor 2002 coupled to a chipset 2004. The chipset 2004 includes a memory controller hub 2020 and an input / output (I / O) controller hub 2022. A memory 2006 and a graphics adapter 2012 are coupled to the memory controller hub 2020, and a display device 2018 is coupled to the graphics adapter 2012. A storage device 2008, a keyboard 2010, a pointing device 2014, and a network adapter 2016 are coupled to the I / O controller hub 2022. The computer system 2000 may include various types of input or output devices. Other embodiments of the computer system 2000 have different architectures. For example, the memory 2006 is coupled directly to the processor 2002 in some embodiments.

[0120] The storage device 2008 includes one or more non-transitory computer-readable storage media, such as a hard drive, a compact disc read-only memory (CD-ROM), a DVD, or a solid-state memory device. The memory 2006 holds program code (comprising one or more instructions) and data used by the processor 2002. The program code may correspond to the processing aspects described with reference to Figures 1-19.

[0121] A pointing device 2014 is used in combination with the keyboard 2010 to input data into the computer system 2000. The graphics adapter 2012 displays images and other information on a display device 2018. In some embodiments, the display device 2018 includes touch screen capabilities for receiving user inputs and selections. The network adapter 2016 couples the computer system 2000 to a network. Some embodiments of the computer system 2000 have different and / or other components than those shown in FIG. 20 .

[0122] The circuitry may include one or more processors executing program code stored on a non-transitory computer-readable medium, the program code, when executed by the one or more processors, configuring the one or more processors to implement an audio processing system or a module of an audio processing system. Other examples of circuitry implementing an audio processing system or a module of an audio processing system may include integrated circuits, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other types of computer circuitry.

[0123] Additional Considerations Exemplary benefits and advantages of the disclosed configurations include dynamic audio enhancements resulting from an enhanced audio system adapted to the device and associated audio rendering system, as well as other relevant information made available by the device OS, such as use-case information (e.g., indicating that an audio signal is used for music playback, but not for gaming). The enhanced audio system may be integrated into the device (e.g., using a software development kit) or stored on a remote server for on-demand access. In this manner, the device need not devote storage or processing resources to maintaining an audio enhancement system specific to its audio rendering system or audio rendering configuration. In some embodiments, the enhanced audio system allows for various levels of querying of rendering system information, so that effective audio enhancements can be applied across various levels of available device-specific rendering information.

[0124] Throughout this specification, multiple instances may implement components, operations, or structures described as a single instance. While individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed simultaneously, and nothing requires that the operations be performed in the order illustrated. Structures and functionality presented as separate components in exemplary configurations may be implemented as combined structures or components. Similarly, structures and functionality presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements are included within the scope of the subject matter herein.

[0125] Certain embodiments are described herein as including logic or numerous components, modules, or mechanisms. The modules may constitute either software modules (e.g., code embodied on a machine-readable medium or in a transmission signal) or hardware modules. A hardware module is a tangible unit capable of performing certain operations and may be configured or arranged in a certain manner. In exemplary embodiments, one or more computer systems (e.g., standalone, client, or server computer systems), or one or more hardware modules of a computer system (e.g., a processor or group of processors), may be configured by software (e.g., an application or application portion) as hardware modules that operate to perform certain operations as described herein.

[0126] Various operations of the example methods described herein may be performed, at least in part, by one or more processors that are temporarily or permanently configured (e.g., by software) to perform the associated operations. Whether temporarily or permanently configured, such processors may constitute processor-implemented modules that operate to perform one or more operations or functions. Modules referred to herein may, in some example embodiments, constitute processor-implemented modules.

[0127] Similarly, the methods described herein may be at least partially processor-implemented. For example, at least some of the operations of the methods may be performed by one or more processors or processor-implemented hardware modules. Execution of some of the operations may be distributed among one or more processors located within a single machine, as well as deployed across several machines. In some exemplary embodiments, one or more processors may be located at a single location (e.g., in a home environment, an office environment, or as a server farm), while in other embodiments, the processors may be distributed across several locations.

[0128] Unless otherwise indicated, descriptions herein using words such as "processing," "calculating," "computing," "determining," "presenting," or "displaying" may refer to machine (e.g., computer) actions or processes that manipulate or transform data represented as physical (e.g., electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.

[0129] As used herein, any reference to "one embodiment" or "an embodiment" means that a particular element, feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment. The appearances of the phrase "in one embodiment" in various places within this specification are not necessarily all referring to the same embodiment.

[0130] Some embodiments may be described using the terms "coupled" and "connected," along with their derivatives. It should be understood that these terms are not intended as synonyms for each other. For example, some embodiments may be described using the term "connected," which indicates that two or more elements are in direct physical or electrical contact with each other. In another example, some embodiments may be described using the term "coupled," which indicates that two or more elements are in direct physical or electrical contact with each other. However, the term "coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other. The embodiments are not limited in this context.

[0131] As used herein, the terms "comprise," "comprising," "includes," "including," "has," "having," or any other variation thereof, are intended to include a non-exclusive inclusion. For example, a process, method, article, or apparatus that includes a list of elements is not necessarily limited to only those elements and may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Furthermore, unless expressly stated to the contrary, "or" refers to an inclusive "or," not an exclusive "or." For example, a condition A or B can be satisfied by any one of the following: A being true (or present) and B being false (or absent); A being false (or absent) and B being true (or present); and both A and B being true (or present).

[0132] Additionally, the use of "a" or "an" is utilized to describe elements and components of embodiments herein. This is done merely for convenience and to give a general sense of the invention. This description should be read to include one or at least one, and the singular also includes the plural unless it is clear that this is not intended.

[0133] Some portions of this description will describe embodiments in terms of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are commonly used by those skilled in the data processing arts to effectively convey the substance of their work to others skilled in the art. These operations, while described functionally, computationally, or logically, will be understood to be implemented by computer programs or equivalent electrical circuits, microcode, or the like. Further, it has proven convenient at times to refer to these arrangements of operations as modules, without loss of generality. The described operations and their associated modules may be embodied in software, firmware, hardware, or any combination thereof.

[0134] Any of the steps, operations, or processes described herein may be performed or implemented using one or more hardware or software modules, alone or in combination with other devices. In one embodiment, the software modules are implemented in a computer program product, including a computer-readable medium containing computer program code that can be executed by a computer processor to perform each and every step, operation, or process described.

[0135] Embodiments may also relate to an apparatus for performing the operations herein. This apparatus may be specially constructed for the required purposes and / or it may comprise a general-purpose computing device selectively activated or reconfigured by a computer program stored within the computer. Such a computer program may be stored on a non-transitory, tangible, computer-readable storage medium that may be coupled to a computer system bus, or any type of medium suitable for storing electronic instructions. Furthermore, any computing system referred to herein may include a single processor or may be an architecture utilizing a multiple-processor design to increase computing power.

[0136] Embodiments may also relate to products produced by the computing processes described herein. Such products may include information resulting from the computing processes, which information may be stored on a non-transitory, tangible, computer-readable storage medium, and may include any embodiment of a computer program product or other data combination described herein.

[0137] Upon reading this disclosure, those skilled in the art will appreciate further and additional alternative structural and functional designs for systems and processes for audio enhancement using device-specific metadata through the principles disclosed herein. Accordingly, while particular embodiments and applications have been illustrated and described, it should be understood that the disclosed embodiments are not limited to the precise structure and components disclosed herein. Various modifications, changes, and variations apparent to those skilled in the art may be made in the arrangement, operation, and details of the methods and apparatus disclosed herein without departing from the spirit and scope, as defined in the appended claims.

[0138] Finally, the language used herein has been selected primarily for ease of reading and instructional purposes, and may not have been selected to recite or define patent rights. Accordingly, it is intended that the scope of patent rights be limited not by this detailed description, but rather by any claims issued in an application based on this specification. Accordingly, the disclosure of the embodiments is intended to illustrate, but not limit, the scope of patent rights, which is defined in the following claims.

Claims

1. generating a mid component and a side component from the left and right channels of the audio signal; generating a hyper-side component by separating a portion of the spectral energy of the side component that is different from the spectral energy of the mid component; generating a left output channel and a right output channel using the hyperside component; A system comprising a circuit configured as follows.

2. 2. The system of claim 1, wherein the circuitry is configured to separate a portion of the spectral energy of the side component that differs from the spectral energy of the mid component by removing the magnitude of the side component in the frequency domain from the spectral energy of the mid component in the frequency domain.

3. The circuit comprises: generating a residual side component by separating a portion of the spectral energy of the side component that differs from the spectral energy of the hyperside component; The system of claim 1 , further configured: the left and right output channels are further generated using the residual side components.

4. The circuit comprises: The system of claim 3 , further configured to generate the residual side component by removing the hyperside component from the side component in the time domain.

5. the circuit filters the subbands of the hyperside component; The system of claim 1 , further configured such that the left and right output channels are generated using filtered subbands of the hyperside component.

6. The system of claim 5 , wherein each of the subbands of the hyperside component includes a set of critical bands.

7. The circuit comprises: generating a hyper-mid component by separating a portion of the spectral energy of the mid component that is different from the spectral energy of the side component; generating a residual mid component by further separating a portion of the spectral energy of the mid component that is different from the spectral energy of the hyper-mid component; The system of claim 1 , further configured to generate left and right output channels using the hyper-mid component and the residual mid component.

8. The system of claim 1 , wherein the circuitry is further configured to apply crosstalk processing to the audio signal, the crosstalk processing including one of crosstalk cancellation or crosstalk simulation.

9. 9. The system of claim 8, wherein the circuitry is further configured to filter the hyperside component to compensate for spectral defects caused by the crosstalk process.

10. the circuit generates a residual side component by isolating a portion of the spectral energy of the side component that is different from the spectral energy of the hyper side component; The system of claim 8 , further configured to filter the residual side components to compensate for spectral imperfections caused by the crosstalk processing.

11. The system of claim 8 , wherein the circuitry is further configured to filter the side components to compensate for spectral defects caused by the crosstalk process.

12. The circuit comprises: generating a hyper-mid component by isolating a portion of the spectral energy of the mid component that differs from the spectral energy of the side component; generating a residual mid component by further isolating portions of the spectral energy of the mid component that differ from the spectral energy of the hyper-mid component; The system of claim 8 , further configured to filter the hyper-mid or residual mid components to compensate for spectral defects caused by the crosstalk processing.

13. The system of claim 8 , wherein the circuitry is further configured to filter the mid component to compensate for spectral defects caused by the crosstalk processing.

14. A non-transitory computer-readable medium having stored thereon program code, the program code, when executed by at least one processor, generating a mid component and a side component from the left and right channels of the audio signal; generating a hyper-side component by separating a portion of the spectral energy of the side component that is different from the spectral energy of the mid component; A non-transitory computer-readable medium configuring the at least one processor to generate a left output channel and a right output channel using the hyperside component.

15. 15. The non-transitory computer-readable medium of claim 14, wherein the program code configures the at least one processor to isolate a portion of the spectral energy of the side component that is different from the spectral energy of the mid component by removing a magnitude of the side component in the frequency domain from a spectral energy of the mid component in the frequency domain.

16. The program code generating a residual side component by separating a portion of the spectral energy of the side component that differs from the spectral energy of the hyperside component; 15. The non-transitory computer-readable medium of claim 14, further configuring the at least one processor such that the left and right output channels are further generated using the residual side component.

17. The program code 17. The non-transitory computer-readable medium of claim 16, further configuring the at least one processor to: generate the residual side component by removing the hyperside component from the side component in the time domain.

18. The program code further configuring the at least one processor to filter subbands of the hyperside component; 15. The non-transitory computer-readable medium of claim 14, wherein the left and right output channels are generated using filtered subbands of the hyperside component.

19. 20. The non-transitory computer-readable medium of claim 18, wherein each of the sub-bands of the hyperside component includes a set of critical bands.

20. The program code generating a hyper-mid component by separating a portion of the spectral energy of the mid component that is different from the spectral energy of the side component; generating a residual mid component by further separating a portion of the spectral energy of the mid component that is different from the spectral energy of the hyper-mid component; 15. The non-transitory computer-readable medium of claim 14, further configuring the at least one processor to generate left and right output channels using the hyper-mid component and the residual mid component.

21. 15. The non-transitory computer-readable medium of claim 14, wherein the program code further configures the at least one processor to apply crosstalk processing to the audio signal, the crosstalk processing including one of crosstalk cancellation or crosstalk simulation.

22. 22. The non-transitory computer-readable medium of claim 21 , wherein the program code further configures the at least one processor to filter the hyperside component to compensate for spectral defects caused by the crosstalk processing.

23. the program code generating a residual side component by isolating a portion of the spectral energy of the side component that differs from the spectral energy of the hyper side component; 22. The non-transitory computer-readable medium of claim 21, further configuring the at least one processor to filter the residual side components to compensate for spectral imperfections caused by the crosstalk processing.

24. 22. The non-transitory computer-readable medium of claim 21 , wherein the program code further configures the at least one processor to filter the side components to compensate for spectral imperfections caused by the crosstalk processing.

25. The program code generating a hyper-mid component by isolating a portion of the spectral energy of the mid component that differs from the spectral energy of the side component; generating a residual mid component by further isolating portions of the spectral energy of the mid component that differ from the spectral energy of the hyper-mid component; 22. The non-transitory computer-readable medium of claim 21, further configuring the at least one processor to filter the hyper-mid component or the residual mid component to compensate for spectral imperfections caused by the crosstalk processing.

26. 22. The non-transitory computer-readable medium of claim 21 , wherein the program code further configures the at least one processor to filter the mid component to compensate for spectral imperfections caused by the crosstalk processing.

27. By the circuit, generating a mid component and a side component from a left channel and a right channel of an audio signal; generating a hyper-side component by separating a portion of the spectral energy of the side component that differs from the spectral energy of the mid component; generating a left output channel and a right output channel using the hyperside component; A method comprising:

28. 28. The method of claim 27, wherein separating the portion of the spectral energy of the side component that differs from the spectral energy of the mid component comprises removing the magnitude of the side component in the frequency domain from the spectral energy of the mid component in the frequency domain.

29. By the circuit, generating a residual side component by separating a portion of the spectral energy of the side component that is different from the spectral energy of the hyperside component; 28. The method of claim 27, wherein the left and right output channels are further generated using the residual side components.

30. 30. The method of claim 29, wherein generating the residual side component comprises removing the hyperside component from the side component in the time domain.

31. further comprising the step of filtering, by a circuit, the subbands of the hyperside component; The method of claim 27, wherein the left and right output channels are generated using filtered subbands of the hyperside component.

32. The method of claim 31 , wherein each of the sub-bands of the hyperside component includes a set of critical bands.

33. By the circuit, generating a hyper-mid component by isolating a portion of the spectral energy of the mid component that is different from the spectral energy of the side component; generating a residual mid component by further separating a portion of the spectral energy of the mid component that is different from the spectral energy of the hyper-mid component; generating left and right output channels using the hyper-mid and residual mid components; 28. The method of claim 27, further comprising:

34. 28. The method of claim 27, further comprising applying, by the circuitry, crosstalk processing to the audio signal, the crosstalk processing comprising one of crosstalk cancellation or crosstalk simulation.

35. 35. The method of claim 34, further comprising filtering the hyperside component to compensate for spectral defects caused by the crosstalk processing by the circuitry.

36. By the circuit, generating a residual side component by isolating a portion of the spectral energy of the side component that differs from the spectral energy of the hyperside component; 35. The method of claim 34, further comprising filtering the residual side components to compensate for spectral imperfections caused by the crosstalk processing.

37. 35. The method of claim 34, further comprising filtering the side components to compensate for spectral defects caused by the crosstalk processing by the circuitry.

38. By the circuit, generating a hyper-mid component by isolating the portion of the spectral energy of the mid component that is different from the spectral energy of the side component; generating a residual mid component by further isolating portions of the spectral energy of the mid component that differ from the spectral energy of the hyper-mid component; filtering the hyper-mid or residual mid components to compensate for spectral defects caused by the crosstalk processing; 35. The method of claim 34, further comprising:

39. 35. The method of claim 34, further comprising filtering the mid component to compensate for spectral defects caused by the crosstalk processing by the circuitry.