Allpass network system for constrained colorless decorrelation
The audio system uses a single-input multi-output all-pass filter to constrain the sum of upmixed channels, addressing the challenge of infinite attenuation in monaural-to-stereo upmixing and ensuring quality compliance.
Patent Information
- Application Number
- JP2025055561
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-02-19
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-02-17
AI Technical Summary
Existing audio systems face challenges in upmixing monaural audio to stereo without causing infinite attenuation when endpoint devices lack independent channels, leading to failures in decorrelation techniques like phase inversion.
An audio system that uses a single-input multi-output all-pass filter to generate multiple channels by constraining the sum of these channels with a target amplitude response, defined by amplitude and frequency relationships, to ensure compliance with minimum quality requirements.
The system effectively decorrelates monaural audio into multiple channels while preserving spectral intensity, allowing for adjustable perceptual transformations and ensuring compatibility with monaural presentations.
Smart Images

Figure 2025098219000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to audio processing, and more particularly to decorrelation of audio content. (
[0001] )
Background Art
[0002] Cross - reference to Related Applications This application claims the benefit of priority of U.S. Patent Application No. 17 / 180,643, filed on February 19, 2021, entitled "Global Pass - Through Network System for Constrained Colorless Decorrelation", which is incorporated herein by reference in its entirety.
[0001]
[0003] The channels of audio data can be up - mixed into multiple channels. For example, a content provider may desire up - mixing from monaural to stereo, but an endpoint device may not be able to provide two independent channels and, alternatively, may add the stereo channels together. When addition occurs at the endpoint, decorrelation techniques such as phase inversion or reverb - based effects can fail. The failure state that can occur when using phase inversion can result in the output being attenuated to infinity. Therefore, it is desirable to constrain the worst - case results of up - mixing so that the sum of the up - mixed channels exceeds the minimum quality requirements.
[0002]
Summary of the Invention
[0004] Some embodiments include a method for generating a plurality of channels from a monaural channel. The method includes determining, by a processing circuit, a target amplitude response that defines one or more constraints on the sum of the plurality of channels. The target amplitude response is defined by a relationship between the amplitude value of the sum and the frequency value of the sum. The method further includes determining a transfer function of a single-input multi-output all-pass filter based on the target amplitude response, and determining all-pass filter coefficients based on the transfer function. The method further includes processing the monaural channel using the all-pass filter coefficients to generate a plurality of channels [
[0003] ].
[0005] Some embodiments include a system for generating a plurality of channels from a monaural channel. The system includes one or more computing devices configured to determine a target amplitude response that defines one or more constraints on the sum of the plurality of channels. The target amplitude response is defined by a relationship between the amplitude value of the sum and the frequency value of the sum. One or more computers determine a transfer function of a single-input multi-output all-pass filter based on the target amplitude response. One or more computers determine all-pass filter coefficients based on the transfer function and process the monaural channel using the all-pass filter coefficients to generate a plurality of channels [
[0004] ].
[0006] Some embodiments include a non-transitory computer-readable medium storing instructions for generating a plurality of channels from a monaural channel. When executed by at least one processor, the instructions configure the processor to determine a target amplitude response that defines one or more constraints on the sum of the plurality of channels, the target amplitude response being defined by a relationship between the amplitude value of the sum and the frequency value of the sum, determine a transfer function of a single-input multi-output all-pass filter based on the target amplitude response, determine all-pass filter coefficients based on the transfer function, and process the monaural channel using the all-pass filter coefficients to generate a plurality of channels [
[0005] ].
Brief Description of the Drawings
[0007]
Figure 1
[0006] .
Figure 2
[0007] .
Figure 3
[0008] .
Figure 4A
[0009] .
Figure 4B
[0010] .
Figure 4C
[0011] .
Figure 4D
[0012] .
Figure 4E
[0013] .
Figure 5
[0014] .
[0008] The drawings illustrate various embodiments for illustrative purposes only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the configurations and methods illustrated herein may be employed without departing from the principles disclosed herein
[0015] .
DETAILED DESCRIPTION OF THE INVENTION
[0009] The drawings (figures) and the following description relate to preferred embodiments for illustrative purposes only. It should be noted that from the following description, alternative embodiments of the configurations and methods disclosed herein will be readily recognized as viable alternatives that may be employed without departing from the principles of the claims
[0016] .
[0010] Next, several embodiments will be referred to in detail, and examples thereof are shown in the accompanying drawings. In all cases where possible, similar or like reference numerals may be used in the drawings and may indicate similar or like functions. The drawings show embodiments of the disclosed system (or method) for illustrative purposes only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods shown herein may be employed without departing from the principles described herein
[0017] .
[0011] Embodiments relate to an audio system that provides monaural presentation compatibility for decorrelating a monaural channel into a plurality of channels. The audio system uses colorless decorrelation of the audio, subject to constraints, to achieve compatibility with monaural presentations. The audio system constrains the worst-case result of upmixing such that the sum of the upmixed channels meets or exceeds a minimum quality requirement. These quality requirements or constraints may be defined by a target amplitude response as a function of frequency. Decorrelation refers to changing the channels of audio data such that the psychoacoustic extent (or "spread") of the audio data can increase when played back on two or more speakers. Colorless refers to the preservation of the spectral intensity of the input audio data in the individual output channels. The audio system uses decorrelation for upmixing. The audio system constructs an all-pass filter according to a target amplitude response and applies the all-pass filter to the monaural channel to generate a plurality of output channels. The filters used for decorrelation are colorless and perceptually expand the extent of the sound field of the monaural audio. These filters allow the user to specify constraints regarding attenuation and coloration that may occur due to the unexpected sum of two or more decorrelated versions of the monaural signal
[0018] 。
[0012] Advantages of constrained colorless decorrelation include the ability to adjust the type and degree of perceptual transformation of the sum output. In the adjustment that may be defined by the target amplitude response, considerations such as the characteristics of the presentation device, the expected content of the audio data, the listener's perceptual capabilities according to the situation, or the minimum quality requirements for monaural presentation compatibility may be made
[0019] 。
[0013] The audio system Figure 1 is a block diagram of an audio system 100 according to some embodiments. The audio system 100 provides decorrelation to a plurality of channels of a monaural channel. The system 100 includes an amplitude response module 102, an all-pass filter configuration module 104, and an all-pass filter module 106. The system 100 processes a monaural input channel x(t) to generate a plurality of output channels, for example, channel y a (t) provided to speaker 110a, and channel y b (t) provided to speaker 110b. Although two output channels are shown, the system 100 may generate any number of output channels (each denoted as channel y(t)). The system 100 may be a computing device such as a music player, a speaker, a smart speaker, a smartphone, a wearable device, a tablet, a laptop, a desktop, etc.
[0020] 。
[0014] The amplitude response module 102 determines a target amplitude response that defines one or more constraints on the sum of the output channels y(t). The target amplitude response is defined by the relationship between the amplitude value of the sum of the channels and the frequency value of the sum of the channels, such as the amplitude as a function of frequency. One or more constraints on the sum of the channels may include a target broadband attenuation, a target subband attenuation, a critical point, or filter characteristics. The amplitude response module 102 may receive data 114 and the monaural channel x(t) and use those inputs to determine the target amplitude response. The data 114 may include information such as the characteristics of a presentation device (e.g., one or more speakers), the expected content of the audio data, the listener's perceptual capabilities according to the situation, or the minimum quality requirements for monaural presentation compatibility.
[0021] 。
[0015] The target broadband attenuation is a constraint on the maximum attenuation of the total amplitude for all frequencies. The target sub-band attenuation is a constraint on the maximum attenuation of the total amplitude for the frequency range defined by the sub-band. The target amplitude response may include one or more target sub-band attenuation values for each of the different sub-bands of the sum.
[0022] 。
[0016] A critical point is a constraint on the curvature of the target amplitude response of the filter and is described as a frequency value at which the sum of the gains becomes a predefined value such as -3 dB or -∞ dB. The placement of this point can have an overall effect on the curvature of the target amplitude response. An example of a critical point corresponds to the frequency at which the target amplitude response becomes -∞ dB. Since the behavior of the target amplitude response is to nullify the signal at frequencies close to this point, the critical point is a null point. Another example of a critical point corresponds to the frequency at which the target amplitude response becomes -3 dB. Since the behavior of the target amplitude response for the sum and difference of the channels intersects at this point, this critical point is an intersection point.
[0023] 。
[0017] The filter characteristics are constraints on how the sum is filtered. Examples of filter characteristics include high-pass filter characteristics, low-pass filter characteristics, band-pass filter characteristics, or band-reject characteristics. The filter characteristics describe the resulting shape of the sum as if it were the result of equalization filtering. Equalization filtering can be described in terms of which frequencies can pass through the filter or which frequencies are blocked. Thus, low-pass characteristics allow frequencies below the inflection point to pass and attenuate frequencies above the inflection point. High-pass characteristics are the opposite, allowing frequencies above the inflection point to pass and attenuating frequencies below the inflection point. Band-pass characteristics allow frequencies in the band near the inflection point to pass and attenuate other frequencies. Band-reject characteristics block frequencies in the band near the inflection point and allow other frequencies to pass.
[0024] 。
[0018] The target amplitude response may define more constraints for the sum than for a single one. For example, the target amplitude response may define constraints on the filter characteristics of the sum output of the critical point and the all-pass filter. In other examples, the target amplitude response may also define constraints on the target broadband attenuation, the critical point, and the filter characteristics. Although described as independent constraints, in most regions of the parameter space, the constraints may be interdependent. This result may be brought about due to the system being non-linear with respect to phase. To address this, additional high-level descriptors of the target amplitude response, which are non-linear functions of the target amplitude response parameters, may be devised
[0025] 。
[0019] Based on the target amplitude response received from the amplitude response module 102, the filter configuration module 104 determines the characteristics of a single-input multi-output all-pass filter. In particular, the filter configuration module determines the transfer function of the all-pass filter based on the target amplitude response and determines the coefficients of the all-pass filter based on the transfer function. The all-pass filter is a decorrelation filter, and the decorrelation filter is constrained by the target amplitude response and applied to the monaural input channel x(t) to output channels y a (t) and y b (t) is generated
[0026] 。
[0020] The all-pass filter may include different configurations and parameters based on the constraints defined by the target amplitude response. The decorrelation filter that constrains the target broadband attenuation of the channel sum has the advantage of (e.g., completely) preserving the spectral components. Such a filter may be useful when no assumptions are made about either the input channels or the audio presentation device regarding the prioritization of specific spectral bands. The transfer function of the all-pass filter is defined as a constant function at a level specified by the value θ for each of the output channels
[0027] 。
[0021] To configure or create the filter, the filter configuration module 104 determines a pair of orthogonal all-pass filters using a continuous-time prototype according to Equation 1.
[0022]
Equation
[0023] The all-pass filter provides constraints on the 90-degree phase relationship between two output signals and the combined amplitude relationship between the input signal and both output signals, but does not guarantee the phase relationship between the input (mono) signal and any of the two (stereo) output signals
[0029] .
[0024] The discrete form of Η(x(t)) is represented by Η2(x(t)) and is defined by the action on the mono signal x(t). The result is a two-dimensional vector as defined by Equation 2.
[0025]
Equation
[0026] The filter configuration module 104 determines a 2×2 orthogonal rotation matrix according to Equation 3.
[0027]
Equation
[0028] Here, θ defines the rotation angle
[0031] .
[0029] The filter configuration module 104 determines a projection onto one dimension as defined by Equation 4.
[0030]
Equation
[0031] Their product is concatenated to the right of a second 2×1 dimensional projection as defined by Equation 5.
[0032]
Number
[0033] The filter configured by the filter configuration module 104 can thus be defined by Equation 6.
[0034]
Number
[0035] As defined by Equation 6, the all-pass filter allows the phase angle rotation of one output channel with respect to other channels
[0034] .
[0036] The multi-output of the all-pass filter is not limited to two output channels. In some embodiments, the system 100 generates more than two output channels from a monaural input channel. The all-pass filter can be generalized to N channels by defining the rotation and projection operations O N (θ).
[0037]
Number
[0038] Here, θ is an (N-1)-dimensional vector of the rotation angle. Then, this operation can be substituted into the equation, and the resulting N-dimensional output vector will include each decorrelated version of the input. The all-pass filter allows the constraint of the broadband attenuation of the sum, unlike the case where it uses phase-inverting decorrelation that is essentially unconstrained because the broadband attenuation of the sum is +∞ dB
[0035] .
[0039] Here, α b The broadband attenuation of the sum in the case of N = 2, represented as, can be determined as follows.
[0040] [Number]
[0041] As a result of the channels used in the sum differing only in the phase term, the attenuation constraint α b is accurate. To define the target amplitude response including the broadband attenuation constant, Equation 9 can be solved for θ.
[0042] [Number]
[0043] Using Equation 9, the all-pass filter A b (x(t), θ) can be parameterized by the constraint on the sum's broadband attenuation. In the context of a presentation, the parameter θ obtained from this equation will maximize the perceived spatial extent of the output. α b is defined by the minimum allowable sum gain factor, so if the perceived width exceeds the requirements for a particular use case, a value of θ that results in a larger gain factor can be selected
[0038] .
[0044] When N > 2, the more generalized form of Equation 8 is defined by Equation 10.
[0045] [Number]
[0046] Equation 10 can be applied as a constraint while selecting the value of θ
[0039] .
[0047] A b (x(t), θ)'s coefficients are determined by the orthogonal filter networks Η2(x(t))1 and Η2(x(t))2, and the angle θ as follows.
[0048]
Number
[0049] Here, the orthogonal filter coefficients β h1 and β h2 depend on the implementation of the orthogonal filter itself
[0040] .
[0050] In some embodiments, a decorrelation filter that restricts the spectral sub-band region of attenuation in the sum is desirable when some change in timbre in the sum is allowed. By relaxing the constraint that the sum must be completely colorless, the spatial range can be further expanded beyond the range possible with a filter such as A b (x(t),θ). The resulting target amplitude response is relaxed from a constant function to a polynomial, and the characteristics of the polynomial can be parameterized using a control similar to the control used when defining the filter for equalization
[0041] .
[0051] In some embodiments, system 100 uses the time-domain specifications of an all-pass filter. For example, a first-order all-pass filter may be defined by Equation 12.
[0052]
Number
[0053] Here, β is a filter coefficient in the range from -1 to +1. The implementation of the filter may be defined by Equation 13.
[0054]
Number
[0055] The transfer function of this filter is the differential phase shift from one output to another output
[0056]
Number
[0057] It is represented by. This differential phase shift is a function of the angular frequency ω, as defined by Equation 14.
[0058]
Number
[0059] Here, the target amplitude response can be derived by substituting θ in Equation 9. As defined in Equations 15 and 16, the gain α of the sum
[0060]
Number
[0061]
Number
[0062]
Number
[0063] By normalizing the target amplitude response to 0 dB, this critical point can correspond to the parameter f which can be the -3 dB point. c
[0044] .
[0064] In some embodiments, the target amplitude response can define constraints on the attenuation of the broadband and sub-bands. For all possible values that the filter coefficient β f can take, this system always behaves like a low-pass filter in the sum. This is due to the term x(t - 1) which is not scaled by β f
[0045] .
[0065] Af (x(t), β) is combined with A b (x(t), θ) to realize more adaptable constraint functions. Formally, as defined by Equation 17, two filters are combined.
[0066]
Number
[0067] Here, γ f :{0, 1} and γ b :{0, 1} are Boolean parameters that bypass the first-order all-pass filter subsystems A f (x(t), β) and A b (x(t), θ), respectively. These parameters allow for adding an additional specific subspace of parameters for the combination of the two parameter spaces as defined by Equation 17 in the case of γ f = γ b = 1
[0046] .[[]]
[0068] The angular frequency ω defined in Equation (15) c becomes the critical point where the target amplitude response asymptotically approaches -∞.
[0069]
Number
[0070] Here, ψ is derived from the higher-order parameter 0 < θ bf < 1 / 2 and Γ: {0.1} via Equation (19).
[0071]
Number
[0072] The parameter θ bf allows for controlling the filter characteristics regarding the inflection point f c 0 < θbf <In the case of 1 / 4, the characteristic is low-pass, and it becomes null at f c , and the spectral slope of the target amplitude function is smoothly interpolated from a suitable low frequency to a flat state as θ bf increases. For 1 / 4 < θ bf < 1 / 2, as θ bf increases, the characteristic is smoothly interpolated from a flat state that is null at f c to high-pass. When θ bf = 1 / 4, the target amplitude function becomes a pure band-stop that is null at f c
[0048] .
[0073] The parameter Γ is a boolean value that places the target amplitude function determined by f c and θ bf either in the sum (i.e., L + R) or the difference (i.e., L - R) of two channels. Due to the all-pass constraint on both outputs to the filter network, the operation of Γ switches between complementary target amplitude responses
[0049] .
[0074] The coefficient β bf and the coefficient β ab both sets are used in the calculation of the overall system coefficient β abf . This provides the composition operation in Equation (17). In the coefficient space, the composition of two linear filters is equivalent to the multiplication of two polynomials. Considering this, the coefficient β abf obtained directly from the definition of the combined system in (17) can be described as follows.
[0075]
Number
[0076] Here, the symbol ★ is used to indicate the multiplication of polynomial coefficients
[0050] .
[0077] In some embodiments, system 100 uses frequency domain specifications for an all-pass filter. For example, filter configuration module 104 uses mathematical equations in the form of Equation 9 to generate a vectorized target amplitude response of K narrowband attenuation constraints α ≡ α1, α2, ···, α K from the vectorized target amplitude response, K phase angles θ ≡ θ1, θ2, ···, θ K of the vectorized transfer function may be determined
[0051] .
[0078] The phase angle vector θ generates a finite impulse response filter as defined by Equation 21.
[0079]
Number
[0080] Here, DFT -1 means the inverse discrete Fourier transform and j ≡ √-1. And the vector B n (θ) of 2(K - 1) FIR filter coefficients may be applied to x(t) as defined by Equation 22.
[0081]
Number
[0082] Here,
[0083]
Number
[0084] means the convolution operation
[0053] .
[0085] Equations 21 and 22 provide effective means for constraining the target amplitude response, but their implementation often relies on relatively high-order FIR filters resulting from the inverse DFT operation. This may not be suitable for systems with resource constraints. In such cases, a low-order infinite impulse response (IIR) implementation as described in relation to Equation 16 may be used
[0054] .
[0086] The all-pass filter module 106 applies the all-pass filter configured by the filter configuration module 104 to the monaural channel x(t) to output channels y a (t) and y b (t). The application of the all-pass filter to channel x(t) may be implemented as defined by Equation 6, 11, 15, or 17. The all-pass filter module 106 provides individual output channels to individual speakers, such as channel y a (t) to speaker 110a and channel y b (t) to speaker 110b
[0055] .
[0087] FIG. 2 is a block diagram of a computing system environment 200 according to some embodiments. The computing system 200 may include an audio system 202, which may include one or more computing devices (e.g., servers) connected to user devices 210a and 210b via a network 208. The audio system 202 provides audio content to user devices 210a and 210b (also referred to as user devices 210) via the network 208. The network 208 facilitates communication between the system 202 and the user devices 210. The network 208 may include various types of networks including the Internet
[0056] .
[0088] The audio system 202 includes one or more processors 204 and a computer-readable medium 206. The one or more processors 204 execute program modules that cause the one or more processors to perform functions such as generating multi-output channels from a monaural channel. The processor 204 can include one or more central processing units (CPUs), graphics processing units (GPUs), controllers, state machines, other types of processing circuits, or combinations of one or more of them. The processor 204 can further include, among other things, program modules, operating system data, and local memory
[0057] .
[0089] The computer-readable medium 206 is a non-transitory storage medium that stores program code for the amplitude response module 102, the filter configuration module 104, the all-pass filter module 106, and the channel summing module 212. The all-pass filter module 106 is configured by the amplitude response module 102 and the filter configuration module 104 to generate multi-output channels from a monaural channel. The system 202 provides the multi-output channels to a user device 210a having a multi-speaker 214 that renders individual output channels
[0058] .
[0090] The channel summing module 212 generates a monaural output channel by summing the multi-output channels generated by the omnipass filter module 106. The system 202 provides the monaural output channel to a user device 210b having a single speaker 216 that renders the monaural output channel. In some embodiments, the channel summing module 212 is disposed in the user device 210b. The audio system 202 provides the multi-output channels to a user device 210b that converts the multi-channels to monaural channels for the speaker 216. The user device 210 presents audio content to the user. The user device 210 may be a user computing device such as a music player, a smart speaker, a smartphone, a wearable device, a tablet, a laptop, a desktop, etc. [
[0059] ].
[0091] Processing Example FIG. 3 is a flowchart of a process 300 for generating a plurality of channels from a monaural channel according to some embodiments. The process shown in FIG. 3 may be performed by components of an audio system (e.g., system 100 or 202). In other embodiments, other entities may perform some or all of the steps in FIG. 3. Embodiments may include different and / or additional steps, or may perform each step in a different order [
[0060] ].
[0092] The audio system determines 305 a target amplitude response that defines one or more constraints on the sum of the multi-channels generated from the monaural channel. The one or more constraints on the sum of the multi-channels may include a target broadband attenuation, a target sub-band attenuation, a critical point, or filter characteristics. The critical point may be an inflection point at 3 dB. The filter characteristics may include one of a high-pass filter characteristic, a low-pass filter characteristic, a band-pass characteristic, or a band-stop characteristic [
[0061] ].
[0093] One or more constraints may be determined based on characteristics of the presentation device (e.g., speaker frequency response, speaker location), the expected content of the audio data, the listener's perceptual capabilities according to the situation, or minimum quality requirements for monaural presentation compatibility. For example, if the speaker cannot adequately reproduce frequencies below 200 Hz, the audio system may effectively conceal the attenuation region of the target amplitude response below that frequency. Similarly, if the expected audio content is conversation, the audio system may select a target amplitude response that affects only frequencies other than those required for intelligibility. If the listener obtains audible cues from other sources according to the situation, such as another speaker array in the location, the audio system may determine a target amplitude response that is complementary to those simultaneous cues
[0062] .
[0094] The audio system determines 310 the transfer function of a single-input multi-output all-pass filter based on the target amplitude response. The transfer function defines the relative phase angle rotation of the output channels. The transfer function describes the influence that the filter network gives to the input for each individual output in terms of the phase angle rotation as a function of frequency
[0063] .
[0095] The audio system determines 315 the coefficients of the all-pass filter based on the transfer function. These coefficients are selected in an optimal way for the type of constraints and the chosen implementation and applied to the input audio stream. Some examples of coefficient sets are defined by equations 11, 16, 18, 20, and 21. In some embodiments, determining the coefficients of the all-pass filter based on the transfer function includes using the inverse discrete Fourier transform (idft). In this case, the coefficient set may be determined as defined by equation 21. In some embodiments, determining the coefficients of the all-pass filter based on the transfer function includes using a phase vocoder. In this case, the coefficient set may be determined as defined by equation 21, except that it is applied in the frequency domain before resynthesizing the time-domain data
[0064] .
[0096] The audio system 320 processes the monaural channel using the coefficients of the all-pass filter to generate a plurality of channels. If the system is operating in the time domain using an IIR implementation such as in equations 11, 16, 18, and 20, the coefficients may scale the appropriate feedback and feedforward delays. If an FIR implementation such as equation 21 is used, only the resulting feedforward delay may be used. If the coefficients are determined and applied in the frequency domain, the coefficients may be applied to the spectral data before resynthesis by complex multiplication. The audio system may provide the plurality of output channels to a presentation device such as a user device connected to the audio system via a network. In some embodiments, if the presentation device has only a single speaker, the audio system synthesizes the plurality of channels into a monaural output channel and provides the monaural output channel to the presentation device
[0065] .
[0097] FIG. 4A is a diagram showing an example of a target amplitude response including target broadband attenuation according to some embodiments. The multi-channel sum 402 generated from the monaural channel and the multi-channel difference 404 are shown. Constraints on the target amplitude response are applied to the sum, while the difference can be adapted to maintain all-pass characteristics. In this example, the target broadband attenuation over all frequencies is -6 dB
[0066] .
[0098] FIG. 4B is a diagram showing an example of a target amplitude response including a critical point according to some embodiments. The multi-channel sum 406 generated from the monaural channel and the multi-channel difference 408 are shown. The critical point includes a -3 dB critical point (e.g., intersection point) at 1 kHz
[0067] .
[0099] FIG. 4C is a diagram showing an example of a target amplitude response including a critical point according to some embodiments. The multi-channel sum 410 generated from the monaural channel and the multi-channel difference 412 are shown. The critical point includes a -∞ dB critical point (e.g., null) at 1 kHz
[0068] .
[0100] FIG. 4D is a diagram showing an example of a target amplitude response including a critical point and high-pass filter characteristics according to some embodiments. The multi-channel sum 414 generated from the monaural channel and the multi-channel difference 416 are shown. The -∞ dB critical point is at 1 kHz and there are high-pass filter characteristics
[0069] .
[0101] FIG. 4E is a diagram showing an example of a target amplitude response including a critical point and low-pass filter characteristics according to some embodiments. The multi-channel sum 418 generated from the monaural channel and the multi-channel difference 420 are shown. The -∞ dB critical point is at 1 kHz and there are low-pass filter characteristics
[0070] .
[0102] FIG. 5 is a block diagram of a computer 500 according to some embodiments. The computer 500 is an example of a computing device that includes circuitry for executing an audio system such as the audio system 100 or 202. At least one processor 502 coupled to a chipset 504 is illustrated. The chipset 504 includes a memory controller hub 520 and an input / output (I / O) controller hub 522. A memory 506 and a graphics adapter 512 are coupled to the memory controller hub 520, and a display device 518 is coupled to the graphics adapter 512. A storage device 508, a keyboard 510, a pointing device 514, and a network adapter 516 are coupled to the I / O controller hub 522. The computer 500 may include various types of input or output devices. Other embodiments of the computer 500 have different architectures. For example, the memory 506 is directly coupled to the processor 502 in some embodiments [
[0071] ].
[0103] The storage device 508 includes one or more non-transitory computer-readable storage media such as a hard drive, a compact disc read-only memory (CD-ROM), a DVD, or a solid state memory device. The memory 506 holds program code (consisting of one or more instructions) and data used by the processor 502. The program code may correspond to the processing modes described with reference to FIGS. 1 through 3 [
[0072] ].
[0104] The pointing device 514 is used in combination with the keyboard 510 to input data into the computer system 500. The graphics adapter 512 displays images and other information on the display device 518. In some embodiments, the display device 518 includes a touch screen capable of receiving user input and selections. The network adapter 516 couples the computer system 500 to a network. Some embodiments of the computer 500 have different and / or other components than those shown in FIG. 5
[0073] .
[0105] The circuit may include one or more processors that execute program code stored on a non-transitory computer-readable medium, and the program code configures the one or more processors to execute an audio system or a module of the audio system when executed by the one or more processors. Other examples of circuits that execute an audio system or a module of the audio system may include integrated circuits such as application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other types of computer circuits
[0074] .
[0106] Additional Considerations Examples of the merits and advantages of the disclosed configuration include dynamic audio enhancement by adapting an enhanced audio system to a device and related audio rendering system, and other related information made available by the device OS, such as use case information (e.g., indicating that an audio signal is used for music playback rather than for gaming purposes). The enhanced audio system is integrated into the device (e.g., using a software development kit) or stored on a remote server for on-demand access. In this way, the device does not need to expend storage or processing resources on maintaining an audio enhancement system specific to its audio rendering system or audio rendering configuration. In some embodiments, the enhanced audio system enables varying levels of queries of the rendering system information so as to be able to apply effective audio enhancement across various levels of available device-specific rendering information [
[0075] ].
[0107] Throughout this specification, multiple instances may implement a component, operation, or structure described as a single instance. Individual operations in one or more methods are illustrated and described as separate operations, but one or more of the individual operations may be performed simultaneously and need not be performed in the order illustrated. Structures and functions shown as separate components in the exemplary configurations may be implemented as a combined structure or component. Similarly, structures and functions shown as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements are within the scope of the subject matter of this specification [
[0076] ].
[0108] Certain embodiments are described herein as including logic or many components, modules, or mechanisms. A module can comprise either a software module (e.g., code embodied in a machine-readable medium or a transmission signal) or a hardware module. A hardware module is a tangible unit capable of performing certain operations and can be configured or arranged in a particular manner. In an exemplary embodiment, one or more computer systems (e.g., a stand-alone, client, or server computer system), or one or more hardware modules of a computer system (e.g., a processor or group of processors), can be configured by software (e.g., an application or an application portion) as a hardware module that operates to perform certain operations as described herein
[0077] .
[0109] The various operations of the exemplary methods described herein can be at least partially executed by one or more processors, temporarily configured (e.g., by software) or permanently configured, to perform the relevant operations. Whether temporarily or permanently configured, such processors can constitute processor-implemented modules that operate to perform one or more operations or functions. The modules referred to herein can, in some exemplary embodiments, constitute processor-implemented modules
[0078] .
[0110] Similarly, the methods described herein may be at least partially processor-implemented. For example, at least some of the operations of the method may be performed by one or more processors or processor-implemented hardware modules. Execution of certain operations may be distributed among one or more processors, and may reside not only within a single machine, but also be deployed across many machines. In some exemplary embodiments, the processor or group of processors may be located in a single location (e.g., within a home environment, an office environment, or as a server farm), while in other embodiments, the processors may be distributed across many locations [
[0079] ].
[0111] Unless otherwise noted, the descriptions herein using terms such as "processing," "computing," "calculating," "determining," "presenting," "displaying," etc., may refer to operations or processes of a machine, e.g., a computer. A machine may manipulate or transform data represented as physical (e.g., electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information [
[0080] ].
[0112] Any reference herein to "one embodiment" or "an embodiment" means that a particular element, feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Although the phrase "in one embodiment" may appear in various places herein, it is not necessarily all referring to the same embodiment [
[0081] ].
[0113] Some embodiments may be described using the terms "coupled" and "connected" along with their derivatives. It should be understood that these terms are not intended to be synonyms of each other. For example, in some embodiments, the term "connected" may be used to describe two or more elements that are in direct physical or electrical contact with each other. In another example, in some embodiments, the term "coupled" may be used to describe two or more elements that are in direct physical or electrical contact with each other. However, the term "coupled" may also mean that two or more elements are not in direct contact with each other but still cooperate or interact with each other. Embodiments are not limited in this context.
[0114] As used herein, the terms "comprises," "comprising," "includes," "including," "has," "having," or other variations are intended to cover non-exclusive inclusion. For example, a process, method, article, or apparatus that consists of a list of elements is not necessarily limited to only those elements, but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, "or" is inclusive and not exclusive. For example, the condition A or the condition B is satisfied by any of the following: A is true (or present) and B is false (or absent), A is false (or absent) and B is true (or present), and both A and B are true (or present).
[0115] Furthermore, the use of "a" or "an" is employed to describe the elements and components of the embodiments herein. This is merely for convenience and is done to give a general meaning to the present invention. This specification should be construed to include one or at least one, unless it is apparent that it has a different meaning, and the singular form also includes the plural form
[0084] .
[0116] Some portions of this specification describe embodiments from the perspective of algorithms and symbolic representations of operations on information. These descriptions and representations of algorithms are commonly used by those skilled in the data processing arts and are used to effectively convey the essence of their work to those skilled in the art. These operations are described functionally, computationally, or logically, but are understood to be implemented by a computer program or equivalent electrical circuits, microcode, etc. Furthermore, it has been shown to be sometimes convenient, without loss of generality, to refer to the arrangement of these operations as modules. The operations described above and the modules associated with them can be embodied by software, firmware, hardware, or any combination thereof
[0085] .
[0117] Any step, operation, or process described herein can be implemented or carried out by one or more hardware modules or software modules, either alone or in combination with other devices. In one embodiment, a software module is implemented by a computer program product that includes a computer-readable medium containing computer program code, and the computer program product can be executed by a computer processor that executes any or all of the steps, operations, or processes described above
[0086] .
[0118] Embodiments may also relate to an apparatus for performing the operations herein. This apparatus may be specially configured for the required purpose and / or may comprise a general-purpose computing device selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored on a non-transitory tangible computer-readable storage medium that may be coupled to a computer system bus or any type of medium suitable for storing electronic instructions. Further, the computing systems referred to herein may include a single processor or may be an architecture that employs a multi-processor design to enhance computing power
[0087] .
[0119] Embodiments may also relate to a product produced by the computing processes described herein. Such a product may include information obtained from the computing process, which may be stored on a non-transitory tangible computer-readable storage medium and may include any embodiment of the computer program product or other data combinations described herein
[0088] .
[0120] Upon reading this disclosure, those skilled in the art will appreciate additional alternative structural and functional designs for systems and processes for audio content decorrelation according to the principles disclosed herein. Accordingly, while specific embodiments and applications have been illustrated and described, the disclosed embodiments are to be understood as not limited to the detailed structures and components disclosed herein. Various modifications, changes, and variations will be apparent to those skilled in the art and may be made in the arrangement, operation, and details of the methods and apparatuses disclosed herein without departing from the spirit and scope defined in the appended claims
[0089] .
[0121] Finally, the language used in this specification has been principally selected for readability and for purposes of explanation, and not for purposes of defining the outer limits or contours of the patent rights. Accordingly, the scope of the patent rights is not intended to be limited by this detailed description, but rather is intended to be defined by the claims of the patent issued on the basis of this application. Thus, the disclosure of the embodiments is intended to illustrate, but not to limit, the scope of the patent rights as defined by the claims set forth below [
[0090] ].
Claims
1. 1. A system for generating multiple channels from a mono channel, comprising: one or more computing devices, the computing devices comprising: Setting one or more coefficients of a single-input multiple-output all-pass filter; and A system configured to process the mono channel with the all-pass filter to generate a plurality of channels, the all-pass filter configured based on the one or more coefficients.
2. The system of claim 1 , wherein the one or more coefficients are determined based on a single-input multiple-output transfer function.
3. The system of claim 2 , wherein the transfer function is determined based on a target magnitude response, the target magnitude response being defined by a relationship between magnitude values and frequency values.
4. 4. The system of claim 3, wherein the target magnitude response is determined by defining one or more constraints on a sum of the multiple channels, the target magnitude response being defined by a relationship between an amplitude value of the sum and a frequency value of the sum.
5. The system of claim 4 , wherein the one or more constraints include a target broadband attenuation for the sum of the plurality of channels.
6. The system of claim 4 , wherein the one or more constraints include a target subband attenuation for the sum of the multiple channels.
7. The system of claim 4 , wherein the one or more constraints include a critical point that defines a curvature of the target magnitude response.
8. The system of claim 7 , wherein the critical point defines a frequency at which the target amplitude response is −3 dB.
9. The system of claim 7 , wherein the critical point defines a frequency at which the target amplitude response is −∞ dB.
10. The system of claim 4 , wherein the one or more constraints include a filter characteristic in the sum of the multiple channels.
11. The filter characteristics are High pass filter characteristics, Low pass filter characteristics, Bandpass filter characteristics, or Band Rejection Filter Characteristics The system of claim 10, comprising one of:
12. The system of claim 4 , wherein the one or more constraints include critical points and filter characteristics.
13. The system of claim 4 , wherein the one or more constraints include a target broadband attenuation, a critical point, and a filter characteristic.
14. 3. The system of claim 2, wherein the one or more computing devices configured to determine the one or more coefficients of the all-pass filter based on the transfer function comprise the one or more computing devices configured to use an inverse discrete Fourier transform (idft).
15. 3. The system of claim 2, wherein the one or more computing devices configured to determine the one or more coefficients of the all-pass filter based on the transfer function comprise the one or more computing devices configured to use a phase vocoder.
16. 3. The system of claim 2, wherein the transfer function defines a rotation of a first phase angle of a first channel in the plurality of channels relative to a second phase angle of a second channel in the plurality of channels.
17. The system of claim 1 , wherein the one or more computing devices are further configured to combine the multiple channels into a mono output channel.
18. The system of claim 1 , wherein the one or more computing devices are further configured to provide the plurality of channels to a user device over a network.
19. 1. A method for generating multiple channels from a mono channel, comprising the steps of: By circuit, Setting one or more coefficients of a single-input multiple-output all-pass filter; processing the mono channel with an all-pass filter to generate a plurality of channels, the all-pass filter being configured based on the one or more coefficients; and A method comprising:
20. The method of claim 19 , wherein the one or more coefficients are determined based on a single-input multiple-output transfer function.
21. The method of claim 20 , wherein the transfer function is determined based on a target magnitude response, the target magnitude response being defined by a relationship between amplitude values and frequency values.
22. 22. The method of claim 21 , wherein the target magnitude response is determined by defining one or more constraints on a sum of the multiple channels, the target magnitude response being defined by a relationship between an amplitude value of the sum and a frequency value of the sum.
23. The method of claim 22 , wherein the one or more constraints include a target broadband attenuation for the sum of the plurality of channels.
24. The method of claim 22 , wherein the one or more constraints include a target subband attenuation for the sum of the multiple channels.
25. The method of claim 22 , wherein the one or more constraints include critical points that define a curvature of the target magnitude response.
26. The method of claim 25, wherein the critical point defines a frequency at which the target amplitude response is -3 dB.
27. The method of claim 25, wherein the critical point defines a frequency at which the target amplitude response is −∞ dB.
28. The method of claim 22 , wherein the one or more constraints include a filter characteristic in the sum of the multiple channels.
29. The filter characteristics are High pass filter characteristics, Low pass filter characteristics, Bandpass filter characteristics, or Band Rejection Filter Characteristics 30. The method of claim 28, comprising one of:
30. The method of claim 22 , wherein the one or more constraints include critical points and filter characteristics.
31. The method of claim 22 , wherein the one or more constraints include a target broadband attenuation, a critical point, and a filter characteristic.
32. 21. The method of claim 20, wherein determining the one or more coefficients of the all-pass filter based on the transfer function includes using an inverse discrete Fourier transform (idft).
33. The method of claim 20 , wherein determining the one or more coefficients of the all-pass filter based on the transfer function includes using a phase vocoder.
34. 21. The method of claim 20, wherein the transfer function defines a rotation of a first phase angle of a first channel in the plurality of channels relative to a second phase angle of a second channel in the plurality of channels.
35. 20. The method of claim 19, further comprising combining, by the circuitry, the multiple channels into a mono output channel.
36. 20. The method of claim 19, further comprising providing, by the circuitry, the plurality of channels to a user device over a network.
37. 1. A non-transitory computer readable medium having stored thereon instructions for generating multiple channels from a mono channel, the instructions, when executed by at least one processor, causing the at least one processor to: Setting one or more coefficients of a single-input multiple-output all-pass filter; and 23. A non-transitory computer-readable medium operative to process the mono channel with the all-pass filter to generate a plurality of channels, the all-pass filter configured based on the one or more coefficients.
38. 40. The non-transitory computer-readable medium of claim 37, wherein the one or more coefficients are determined based on a single-input multiple-output transfer function.
39. 40. The non-transitory computer-readable medium of claim 38, wherein the transfer function is determined based on a target magnitude response, the target magnitude response defined by a relationship between amplitude values and frequency values.
40. 40. The non-transitory computer-readable medium of claim 39, wherein the target magnitude response is determined by defining one or more constraints on a sum of the multiple channels, the target magnitude response being defined by a relationship between an amplitude value of the sum and a frequency value of the sum.
41. 41. The non-transitory computer-readable medium of claim 40, wherein the one or more constraints include a target broadband attenuation for the sum of the plurality of channels.
42. 41. The non-transitory computer-readable medium of claim 40, wherein the one or more constraints include a target subband attenuation for the sum of the multiple channels.
43. 41. The non-transitory computer-readable medium of claim 40, wherein the one or more constraints include a critical point that defines a curvature of the target magnitude response.
44. 44. The non-transitory computer readable medium of claim 43, wherein the critical point defines a frequency at which the target amplitude response is -3 dB.
45. 44. The non-transitory computer readable medium of claim 43, wherein the critical point defines a frequency at which the target amplitude response is -∞ dB.
46. 41. The non-transitory computer-readable medium of claim 40, wherein the one or more constraints include a filter characteristic in the sum of the multiple channels.
47. The filter characteristics are High pass filter characteristics, Low pass filter characteristics, Bandpass filter characteristics, or Band Rejection Filter Characteristics 47. The non-transitory computer readable medium of claim 46, comprising one of:
48. 41. The non-transitory computer-readable medium of claim 40, wherein the one or more constraints include critical points and filter characteristics.
49. 41. The non-transitory computer-readable medium of claim 40, wherein the one or more constraints include a target broadband attenuation, a critical point, and a filter characteristic.
50. 40. The non-transitory computer-readable medium of claim 38, wherein the instructions configuring the at least one processor to determine the one or more coefficients of the all-pass filter based on the transfer function operate the at least one processor to use an inverse discrete Fourier transform (idft).
51. 40. The non-transitory computer-readable medium of claim 38, wherein the instructions operating the at least one processor to determine the one or more coefficients of the all-pass filter based on the transfer function operate the at least one processor to use a phase vocoder.
52. 40. The non-transitory computer readable medium of claim 38, wherein the transfer function defines a rotation of a first phase angle of a first channel in the multiple channels relative to a second phase angle of a second channel in the multiple channels.
53. 40. The non-transitory computer-readable medium of claim 37, wherein the instructions further operate the at least one processor to combine the multiple channels into a mono output channel.
54. 40. The non-transitory computer-readable medium of claim 37, wherein the instructions further operate the at least one processor to provide the plurality of channels to a user device over a network.
Citation Information
Patent Citations
Three-dimension stereophonic sound generating system
JP1997284900A
Pseudo-stereophonic conversion device
JP1999262098A
Acoustic reproducing device
JP1999289598A
Mono-stereo conversion device, audio reproduction system using such device, and mono-stereo conversion method
JP2000504526A
Apparatus and method for synthesizing pseudo-stereophonic output from monaural input
JP2002528020A