Adaptive filter bank using scale-dependent non-linearity for psychoacoustic frequency range extension
Through the circuit system generating and processing of orthogonal components, a psychological acoustic impression of frequencies exceeding the bandwidth of the physical driver is achieved, and the problem of difficulty in rendering high frequencies by speakers in the prior art is solved, thereby achieving a high-quality listening experience.
Patent Information
- Application Number
- CN202280048258.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-09-09
- Filing Date
- 2022-07-14
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-07-14
AI Technical Summary
The bandwidth of existing acoustic actuators such as speakers and headphones is usually limited to the bandwidth subdomain of human auditory systems, making it difficult to effectively render frequencies beyond the bandwidth of physical drivers.
The circuit system is adopted to generate orthogonal components, and to apply positive transformation to rotate its spectrum from the standard basis to the rotating basis, isolate the components at the target frequency, and generate weighted phase-coherent harmonic spectrum orthogonal components through nonlinear processing, and finally rotate its spectrum back to the standard basis by inverse transformation, combining the output channels to provide to the speaker.
A psychological acoustic impression of frequencies beyond the physical driver bandwidth is achieved, allowing low-cost speakers to provide a high-quality listening experience without the need for hardware modifications to the speakers.
Smart Images

Figure CN117616780B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 222,370, filed Jul. 15, 2021, and U.S. Application No. 17 / 471,012, filed Sep. 9, 2021, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present disclosure generally relates to audio processing and, more particularly, to creating the impression of frequencies beyond the bandwidth of a physical driver. Background Art
[0004] The bandwidth of speakers, headphones, and other acoustic actuators is typically limited to a sub-domain of the bandwidth of the human auditory system. This is typically a problem in the low-frequency region of the audible spectrum (about 18 Hz to 250 Hz). It is desirable to modify an audio signal to create the impression of frequencies beyond the bandwidth of a physical driver. Summary of the Invention
[0005] Some embodiments include a system that includes circuitry (e.g., one or more processors) that provides psychoacoustic frequency range extension for a speaker. The circuitry generates orthogonal components from an audio channel, the orthogonal components defining an orthogonal representation of the audio channel, and generates rotated spectral orthogonal components by applying a forward transform that rotates the spectrum of the orthogonal components from a standard basis to a rotated basis. In the rotated basis, the circuitry isolates components at a target frequency in the rotated spectral orthogonal components and generates weighted phase-coherent harmonic spectral orthogonal components by applying a non-linearity having scale-dependence that complies with a constraint to the isolated components. The circuitry generates harmonic spectral components by applying an inverse transform that rotates the spectrum of the weighted phase-coherent harmonic spectral orthogonal components from the rotated basis to the standard basis. The circuitry combines the harmonic spectral components with frequencies of the audio channel outside of the target frequency to generate an output channel and provides the output channel to the speaker.
[0006] In some embodiments, the non-linearity includes a weighted mixture of component non-linearities. The constraints each include a constraint on gain correction for an input to which a corresponding component non-linearity is applied.
[0007] In some embodiments, the non-linearity includes a weighted sum of Chebyshev polynomials of the first kind, the magnitude of which is selectively factored out in compliance with a constraint.
[0008] In some embodiments, the circuitry is further configured to generate a plurality of harmonic spectral components. Each harmonic spectral component is generated using a different frequency band of the audio channel. The circuitry is configured to generate the output channel by combining the plurality of harmonic spectral components.
[0009] In some embodiments, the circuitry is configured to generate multiple harmonic spectral components in series, where each downstream harmonic spectral component is generated using the residue of an upstream harmonic spectral component as an input.
[0010] In some embodiments, the circuitry is configured to generate multiple harmonic spectral components in parallel.
[0011] In some embodiments, the circuitry is further configured to apply an odd non-linearity to the harmonic spectral components.
[0012] In some embodiments, the harmonic spectral components include frequencies different from the target frequency of the audio channel and create a psychoacoustic impression of the target frequency when rendered by a speaker.
[0013] In some embodiments, the forward transform rotates the spectrum of the orthogonal components such that the target frequency is mapped to 0 Hz. The inverse transform rotates the spectrum of the weighted phase-coherent harmonic spectral orthogonal components such that 0 Hz is mapped to the target frequency.
[0014] In some embodiments, the target frequency includes a frequency between 18 Hz and 250 Hz.
[0015] In some embodiments, the circuitry is further configured to determine the target frequency based on the reproducible range of the speaker, a reduction in the power consumption of the speaker, or an increased lifespan of the speaker.
[0016] In some embodiments, the speaker is a component of a mobile device.
[0017] In some embodiments, the circuitry is further configured to isolate the components at the target amplitude using a gate function. In some embodiments, the circuitry is further configured to apply a smoothing function to the isolated components.
[0018] Some embodiments include a method. The method includes, by a circuitry: generating orthogonal components from an audio channel, the orthogonal components defining an orthogonal representation of the audio channel; generating rotated spectral orthogonal components by applying a forward transform that rotates the spectrum of the orthogonal components from a standard basis to a rotated basis; in the rotated basis: isolating the components at the target frequency in the rotated spectral orthogonal components; and generating weighted phase-coherent harmonic spectral orthogonal components by applying a non-linearity with scale-dependence that complies with a constraint to the isolated components; generating harmonic spectral components by applying an inverse transform that rotates the spectrum of the weighted phase-coherent harmonic spectral orthogonal components from the rotated basis to the standard basis; combining the harmonic spectral components with the frequencies of the audio channel outside of the target frequency to generate an output channel; and providing the output channel to a speaker.
[0019] Some embodiments include a non-transitory computer-readable medium including stored instructions that, when executed by at least one processor, configure the at least one processor to: generate an orthogonal component from an audio channel, the orthogonal component defining an orthogonal representation of the audio channel; generate a rotated spectral orthogonal component by applying a forward transform that rotates the spectrum of the orthogonal component from a standard basis to a rotated basis; in the rotated basis: isolate the component at a target frequency in the rotated spectral orthogonal component; and generate a weighted phase-coherent harmonic spectral orthogonal component by applying a non-linearity having a scale-dependence that complies with a constraint to the isolated component; generate a harmonic spectral component by applying an inverse transform that rotates the spectrum of the weighted phase-coherent harmonic spectral orthogonal component from the rotated basis to the standard basis; combine the harmonic spectral component with frequencies of the audio channel outside the target frequency to generate an output channel; and provide the output channel to a speaker. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a block diagram of an audio system according to some embodiments.
[0021] Figure 2 is a block diagram of a harmonic processing module according to some embodiments.
[0022] Figure 3 is a block diagram of a forward transform module according to some embodiments.
[0023] Figure 4 is a block diagram of a coefficient arithmetic unit module according to some embodiments.
[0024] Figure 5 is a block diagram of an inverse transform module according to some embodiments.
[0025] Figure 6 is a block diagram of a combiner module according to some embodiments.
[0026] Figure 7 is a block diagram of a filter bank module according to some embodiments.
[0027] Figure 8 is a flowchart of a process for psychoacoustic frequency range extension according to some embodiments.
[0028] Figure 9 is a block diagram of a computer according to some embodiments.
[0029] The drawings depict various embodiments for illustrative purposes only. Those skilled in the art will readily recognize from the following discussion that alternative embodiments of the structures and methods shown herein may be employed without departing from the principles described herein. DETAILED DESCRIPTION
[0030] The accompanying drawings and the following description relate only by way of illustration to preferred embodiments. It should be noted that alternative embodiments of the structures and methods disclosed herein will readily be recognized as being viable alternatives that may be employed without departing from the principles claimed herein.
[0031] Reference will now be made in detail to several embodiments, examples of which are illustrated in the accompanying drawings. Note that wherever practicable, similar or like reference numerals may be used in the drawings and may indicate similar or like functionality. The drawings depict embodiments of the disclosed system (or method) for illustrative purposes only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods shown herein may be employed without departing from the principles described herein.
[0032] Embodiments relate to providing psychoacoustic frequency range extension. Since the human auditory system responds to cues in a non-linear fashion, psychoacoustic phenomena can be exploited to create virtual stimuli where physical stimuli are not feasible. An audio system may include circuitry that provides an adaptive non-linear filter bank that uses highly tunable non-linearity with scale-dependence that complies with constraints. The non-linearity is used to generate weighted phase-coherent harmonic spectra from one or more sub-bands of an audio channel. The non-linearity may include a weighted mixture of component non-linearities. The constraints may each include constraints on gain correction of the inputs applied to the respective component non-linearities. Independent constraints may be applied to each of the component non-linearities that define the sum of the non-linearity, which allows for selective spectral animation among a selected subset of the generated harmonics. This allows for a more natural effect that successfully generalizes the content. Additionally, it reduces the perceived salience of intermodulation artifacts, which may allow for a smaller number of filters to be employed with a wider bandwidth. In some embodiments, the non-linearity includes a weighted sum of Chebyshev polynomials of the first kind, the magnitudes of which are selectively factored out while complying with the constraints. When the frequency of a sub-band exceeds the bandwidth of a physical driver, a sub-band impression is created for the phase-coherent harmonic spectrum of one or more sub-bands.
[0033] In some embodiments, an adaptive non-linear filter bank may include a plurality of harmonic processors. Each harmonic processor includes a non-linear filter that analyzes a target sub-band within an audio signal and re-synthesizes the data of the sub-band using a configurable spectral transform. The harmonic processors each generate harmonic spectral components using different frequency bands of an audio channel, and these harmonic spectral components are combined to generate an output channel. The harmonic spectral components may be generated in parallel or serially. In the serial case, each downstream harmonic spectral component uses the residual of the upstream harmonic spectral component as an input. The parallel case, while conceptually simple, can occasionally cause difficult tuning processes, such as when the parallel design does not limit the power spectrum of what is being analyzed. By utilizing a serial architecture, where subsequent filters only act on the residual of the input signal, the total spectral power is preserved at the input of the filter bank. The result is a filter bank whose constituent filters are not subject to constructive interference.
[0034] Advantages of frequency range extension include allowing (e.g., low-quality) speakers that cannot render certain frequencies to produce a psychoacoustic impression of those frequencies. Thus, low-cost speakers (such as those commonly found on mobile devices) can provide a high-quality listening experience. Psychoacoustic frequency range extension is achieved by processing the audio signal, such as by the processing circuitry found in a mobile device, and does not require hardware modification to the speaker. Frequency range extension and frequency response improvement can also be useful for improving the power consumption characteristics and service life of a speaker driver when achieved without resorting to increasing the amount of physical energy in sub-optimal sub-bands.
[0035] Audio processing system
[0036] Figure 1 is a block diagram of an audio system 100 according to some embodiments. The audio system 100 uses a non-linear filter bank module 120 to provide frequency range extension for a speaker 110. The system 100 includes a filter bank module 120, and the filter bank module 120 includes harmonic processing modules 104(1), 104(2), 104(3), and 104(4), an all-pass filter network module 122, and a combiner module 106. Some embodiments of the audio system 100 may include components different from those described herein.
[0037] The filter bank module 120 uses highly tunable non-linearity with scale-dependence that complies with constraints to generate phase-coherent harmonic spectra from an audio channel a(t). In some embodiments, the harmonic processing modules 104 may be connected in parallel, as shown. Some embodiments may include a serial implementation of the filter bank module, where the residual of each upstream harmonic processing module is passed to a downstream harmonic processing module. Combine Figure 7Discuss the serial implementation in more detail. The system 100 generates an output channel o(t) that is provided to the speaker 110 for rendering. The harmonic processing modules 104(1) to 104(4) of the filter bank module 120 provide psychoacoustic frequency range extension beyond the physical bandwidth of the speaker 110 for the audio channel a(t).
[0038] The filter bank module 120 includes a plurality of harmonic processing modules 104(n) that generate harmonic spectral components h(t)(n). In some embodiments, each of the harmonic processing modules 104(1) to 104(4) analyzes the entire audio channel a(t) and synthesizes the corresponding harmonic spectral components h(t)(1) to h(t)(4). In some embodiments, each harmonic processing module may analyze different target subbands of the audio channel. Each harmonic spectral component h(t)(n) is a phase-coherent spectral transform of the data in a(t). Each harmonic spectral component h(t)(n) has a weighted phase-coherent harmonic spectrum that includes frequencies different from the data frequencies in the corresponding target subband of a(t) and that creates a psychoacoustic impression of the frequencies of the corresponding target subband when output by the speaker 110. One or more of the harmonic processing modules 104(n) may be selected to generate the harmonic spectral component h(t)(n) to provide psychoacoustic frequency range extension for the speaker 110. In some embodiments, the selection of the target subbands may be based on the capabilities of the speaker 110, such as the frequency response of the speaker 110. For example, if the speaker 110 cannot effectively render low frequencies of sound, the harmonic processing module 104 may be configured to target the frequency subband components corresponding to the low frequencies, and these may be converted into the harmonic spectral components h(t)(n). The audio system 100 may include one or more harmonic processing modules 104. Additional details regarding the harmonic processing module 104 are discussed in conjunction with Figures 2 to 5 are discussed.
[0039] The all-pass filter network module 122 generates a filtered audio channel a(t) to ensure coherence between the audio channel a(t) and the output of the filter bank module 120. The all-pass filter network 122 compensates for the phase change caused by the application of the harmonic processing module 104(n) by applying a matching phase change to the input signal a(t). This allows for a coherent summation between a signal that is perceptually indistinguishable from a(t) but has a manipulated phase and the harmonic spectral components h(t)(n) generated by the filter bank module 120.
[0040] The combiner module 106 generates an output channel o(t) by combining the filtered audio channel a(t) from the all-pass filter network module 122 and one or more harmonic spectral components h(t)(n) from the filter bank module 120. The combiner module 106 provides the output channel o(t) to the speaker 110. In some embodiments, the combiner module 106 performs additional processing on the summed harmonic spectral components h(t)(n), as discussed in more detail in conjunction with Figure 6 as discussed in more detail.
[0041] Figure 2 is a block diagram of the harmonic processing module 104 according to some embodiments. The harmonic processing module 104 provides a non-linear filter that analyzes the audio channel and resynthesizes the data of the target sub-band using a configurable spectral transform. The harmonic processing module 104 includes an all-pass network module 202, a forward transform module 204, a coefficient arithmetic module 206, and an inverse transform module 208. The all-pass network module 202 applies a pair of phase transforms to the audio channel x(t) to generate orthogonal components. The forward transform module 204 applies a forward transform to the orthogonal components, which rotates the entire spectrum such that the selected frequency is mapped to 0 Hz to generate a rotated spectral orthogonal component. The shift of the selected frequency to 0 Hz is referred to as a change from the standard basis to the rotated basis. The selected frequency can be the center frequency of the target sub-band or other frequencies. The coefficient arithmetic module 206 performs operations in the rotated basis, including selectively filtering the data based on frequency, amplitude, or phase, and generating weighted phase-coherent harmonic spectral orthogonal components by applying a non-linearity to the isolated components with scale-dependence that conforms to the constraints. The inverse transform module 208 applies an inverse transform to rotate the spectrum of the weighted phase-coherent rotated spectral orthogonal component such that 0 Hz is mapped to the selected frequency to generate the harmonic spectral component The shift of 0 Hz to the selected frequency is referred to as a change from the rotated basis to the standard basis. The harmonic spectral component may include frequencies different from the target sub-band of the audio channel x(t), but produces a psychoacoustic impression of the frequencies of the target sub-band of the audio channel x(t) when rendered by the speaker.
[0042] In some embodiments, the audio component x(t) input to the harmonic processing module 104 can be a sub-band component a(t)(n). In this example, the selective filtering for selecting the target frequency performed by the coefficient arithmetic module 206 can be skipped.
[0043] The all-pass network 202 converts the audio channel x(t) into a vector y(t) including orthogonal components y 1 (t) and y 2 (t). The orthogonal components y 1 (t) and y 2(t) includes a 90° phase relationship. Orthogonal component y 1 (t) and y 2 (t), and the input signal x(t) include a uniform amplitude relationship for all frequencies. The real-valued input signal x(t) is tuned to orthogonal values by a pair of matched all-pass filters H1 and H2. This operation can be defined via a continuous-time prototype as shown in Equation 1:
[0044]
[0045] Some embodiments will not necessarily guarantee the phase relationship between the input (mono) signal and either of the two (stereo) orthogonal components y 1 (t) and y 2 (t), but generate orthogonal components y 1 (t) and y 2 (t) that include a 90° phase relationship, and orthogonal components y 1 (t) and y 2 (t) and the input signal x(t) that include a uniform amplitude relationship for all frequencies.
[0046] Figure 3 is a block diagram of the positive converter module 204 according to some embodiments. The positive converter module 204 includes a rotation matrix module 302 and a matrix multiplier 304. The positive converter module 204 receives the orthogonal components y 1 (t) and y 2 (t) and applies a positive transform to generate a vector u(t) that includes the rotated spectral orthogonal components u 1 (t) and u 2 (t). This transform is applied by generating a time-varying rotation matrix via the rotation matrix module 302 and applying it to the orthogonal components via the matrix multiplier 304, thereby producing the rotated spectral orthogonal components u(t). The vector u(t) is a frequency-shifted form of the spectrum of the audio signal x(t) and defines a coefficient space in which each u at different times t is defined as a rotated spectral orthogonal component. The coefficients defined by the vector u(t) are the result of rotating the spectrum of x(t) such that the desired center frequency θc is now at 0 Hz.
[0047] The positive transform can be applied as a time-varying 2D rotation on the orthogonal signals, as defined by Equation 2:
[0048] u[t] = H 1 (x[t])R 2 (-θ c t) (2)
[0049] where H1 is an all-pass filter, the rotation R 2 (-θ cThe angular frequency of t) is θc and is defined by Equation 3:
[0050]
[0051] Equations 2 and 3 involve iterative calls to trigonometric functions. Over intervals where θc is constant, the forward transform can be computed by recursive 2D rotation rather than iterative calls to trigonometric functions. When this optimization strategy is used, calls to sin and cos are made only when θc is initialized or changed. This optimization recursively defines each matrix R 2 (-θ c t) as the successive powers of an infinitesimal rotation matrix, i.e.: R 2 (-θ c (t + 1)) ≡ R 2 (-θ c t)R 2 (-θ c ). Since multiplying two 2×2 matrices together is a highly optimized computation on most architectures, this definition may provide a performance advantage compared to the iterative calls to trigonometric functions presented in Equation 3, although it is equivalent.
[0052] Figure 4 is a block diagram of coefficient arithmetic unit module 206 according to some embodiments. Coefficient arithmetic unit module 206 includes filter module 402, amplitude module 404, gate module 406, divider arithmetic units 408 and 410, harmonic generator module 412, multiplier arithmetic units 414 and 416, and max module 420. Coefficient arithmetic unit module 206 generates a rotated spectrum including weighted phase-coherent rotated spectral orthogonal components 1 (t) and u 2 (t) of vector u(t) including rotated spectral orthogonal components and
[0053] In some embodiments, filter module 402 is a two-channel low-pass filter. In this case, harmonic processing module 104 is configured to perform a spectral transform on a target subband centered at θc at a bandwidth that is twice the cut-off frequency of filter module 402. Filter module 402 may apply low-pass filter F(x), which results in a tunable band-pass filter after inverse transform. In this case, the cut-off frequency of F(x) corresponds to half of the bandwidth of the analysis region of the non-linear filter.
[0054] Amplitude module 404 determines the length of a 2D vector, which is used as a measure of the instantaneous amplitude and can be selectively factored out from the filtered signal vector using divider arithmetic units 408 and 410. For example, divider arithmetic unit 408 may divide u of u(t)1 (t) component performs division, and the divider 410 can divide the u of u(t). 2 (t) component performs division. The constraint on scale dependence defined by the max() function in Equation 9 is applied by the maximum module 420, which effectively constrains the actions of dividers 408 and 410. In some embodiments, the amplitude can be factored out regardless of scale to allow the harmonic generator module 412 to provide harmonics based on signals whose relationships are scale-independent.
[0055] The harmonic generator module 412 generates a nonlinearity that includes a weighted sum of component non-linearities. The non-linearity provides a harmonic spectrum for a target sub-band based on rotation-based spectral orthogonal components. For example, the harmonic generator module 412 generates component non-linearities for different harmonics, applies a weight a n to the component non-linearities, and generates a non-linearity as the sum of the weighted component non-linearities.
[0056] Then the amplitude provided by the amplitude module 404 is used again, this time passed through the gate module 406. The gate module 406 generates an envelope whose instantaneous slope is limited by the slewing (sle) limiter 418. The resulting slew-limited envelope is then applied to the output of the harmonic generator module 412 via multipliers 414 and 416. For example, multiplier 416 can multiply the u of u(t). 1 (t) component, and multiplier 414 can multiply the u of u(t). 2 (t) component. The non-linearity defined by the sum of the weighted harmonics is multiplied by the time-varying envelope to generate a rotating spectrum
[0057] The coefficients of u(t) can be represented in polar coordinates using Equation 4:
[0058]
[0059] where the term ||u(t)|| is the instantaneous amplitude of the coefficient signal, and ∠u(t) is the instantaneous phase. These terms can now be manipulated before the inverse transform stage.
[0060] The coefficients defined by u(t) are selectively filtered based on their instantaneous amplitudes. The filtering can include a gate function applied by the gate module 406 and a slew-limiting filter applied by the slew limiter 418. The gate function based on a threshold n can be defined by Equation 5:
[0061]
[0062] The case where x≥n gives rise to a retention coefficient, while the case where x<n gives rise to the removal of the coefficient. In some embodiments, the case where x < n may alternatively give rise to an attenuation of the coefficient rather than a complete removal of the coefficient. Since the gate function operates based on an estimate of the instantaneous amplitude, it is generally faster than a gate response based on the real-valued amplitude and has fewer artifacts.
[0063] Time-domain smoothing can be achieved via a slew-rate limiting filter to further adjust the envelope characteristics of the response of the non-linear filter. A slew-rate limiting filter is a non-linear filter that saturates the maximum (positive) and minimum (negative) slopes of a function. Various types of slew-rate limiting filters or elements can be used, such as a non-linear filter with independent control of the positive and negative saturation points, labeled below as S(x). Applying slew-rate limiting to the output of the gate function results in a time-varying envelope: S(G(||u[t]||)). This can be used to shape the envelope of the coefficients.
[0064] To generate a phase-coherent harmonic spectrum, the harmonic generator module 412 can use Chebyshev polynomials of the first kind defined by Equation 6:
[0065] T n (x) = cos(n cos -1 (x)) (6)
[0066] These polynomials provide a controlled generation of harmonics by summing their outputs, as defined by Equation 7 or 8 for scale-independent non-linearity:
[0067]
[0068] Or equivalently:
[0069]
[0070] where a n = [a0,a1,a2...aN] are the harmonic weights applied to each harmonic n of the phase-coherent harmonic spectrum, and N is the highest generated harmonic. In both representations of Equation 7 and 8, the non-linearity (e.g., defined by the summation result) is independent of the input scale. This can prevent the output spectrum from varying with the input loudness, allowing only the variation determined by the spectral weights a. The weights are typically arranged as a decaying series, mimicking the harmonic series of naturally occurring sounds to which the human auditory system is accustomed. This series of weights is independent of the scale of the incoming audio channel.
[0071] Although equivalent, Equation 7 has the advantage of allowing direct manipulation of the output phase, while Equation 8 omits potentially expensive trigonometric functions and operates only on the amplitude.
[0072] In Equations 7 and 8, the non-linear output spectrum does not vary as a function of the input coefficient magnitude ||u(t)||. While this results in a tightly controlled and predictable non-linearity, this uniformity can produce textures that sound unnatural in some cases. This eerie effect is particularly evident on certain input content, such as spoken and sung voices, and is exacerbated if low-frequency content is also present.
[0073] For example, movie content often includes low-frequency effects (LFE) content in conjunction with dialogue. This LFE content is precisely the type of content we want to reproduce using this technology; however, the resulting intermodulation distortion can affect the clarity and realism of the sound.
[0074] To address this issue, varying degrees of control can be applied to each component non-linearity, allowing the resulting harmonic mixing to be (e.g., to some extent) animated in response to the input content. The degree to which the incoming amplitude is clipped to a uniform level will determine the degree of spectral stability. When the amplitude is below uniformity, the harmonic contributions of the non-linear components will include a mixture of lower integer harmonics. Even polynomials will generate a mixture of even integer harmonics, while odd polynomials will generate a mixture of odd integer harmonics.
[0075] Since the instantaneous amplitude calculation is applied directly in Equation 8, we can simply modify the algorithm to apply constraints to it, as defined by Equation 9:
[0076]
[0077] where b n = [b0, b1, b2... bN] defines a minimum constraint for the amplitude correction factor defined for each harmonic n of the phase-coherent harmonic spectrum, and N is the highest generated harmonic. For each harmonic n, the amplitude correction factor max(||u(t)||, b n ) defines a constraint on the gain correction applied to the input u(t) of the component non-linearity, as defined by Equation 10: n )
[0078]
[0079] Thus, the non-linearity is defined as in Equation 11:
[0080]
[0081] includes a weighted (e.g., by a n ) mixture of component non-linearities for different harmonics (n = 0 to N), where the component non-linearity is defined by Equation 10.
[0082] For u(t) amplitudes below b n the signal amplitude for correction is allowed to fluctuate. For u(t) amplitudes above b n the harmonic content is defined as the sum of the harmonics corresponding to the order of the polynomial, as in the case of all possible amplitudes for Equation 8. At the amplitude of u(t) between b and 0, the higher harmonic content generally decreases with decreasing amplitude, however for higher order polynomial mixing this relationship may be more complex than simply monotonic.
[0083] For example, a transfer function including the third Chebyshev polynomial as defined by Equation 12:
[0084] T 3 (x) = 4x 3 - 3x (12)
[0085] When x is a unit amplitude cosine wave, the following pure third harmonic (and -∞ dB of the first harmonic) is generated, as defined by Equation 13:
[0086] T 3 (cos(x)) = cos(3x) (13)
[0087] But when x is replaced by a -6 dB amplitude cosine wave, a mixture of harmonics is generated, as defined by Equation 14:
[0088]
[0089] Or, put simply, the third harmonic is -18 dB and the first (fundamental) harmonic is +1 dB. This mixing also demonstrates the singularity of the harmonics generated by all components. In addition, the first harmonic has been amplified relative to the input, resulting in a positive dB value.
[0090] When applied to a -12 dB cosine wave, the same transfer function creates the result as defined by Equation 15:
[0091]
[0092] which includes a non-monotonic behavior of decreasing third harmonic and first harmonic.
[0093] By constraining the degree of spectral clipping, the algorithm can better summarize the overall content. In addition, fewer frequency bands may need to be calculated because any intermodulation effects are less perceptually present.
[0094] Intermodulation effects are a typical byproduct of applying a non - linear transfer function to a signal with more than one frequency. Typically, these intermodulation effects include frequencies that are the sum and difference of the input signal frequencies. Left unconstrained, these intermodulation effects are given additional weight and stability. By constraining the spectral clipping function, the resulting spectrum is more unstable and places more emphasis on the dominant frequencies rather than the intermodulation effects.
[0095] Therefore, expanding the frequency range via constrained spectral clipping can achieve a similar effect with fewer individual non - linear filters than using an unconstrained approach. This can lead to an improvement in computational efficiency. Additionally, the reduction in parameters may also lead to an algorithm that is easier to tune because the interaction between many filters can sometimes be difficult to manage.
[0096] As shown in Equation 14, processing the third Chebyshev polynomial applied to a cosine with an amplitude of - 6dB can cause amplification rather than being degraded to attenuation. This fact, combined with the relatively non - intuitive behavior of harmonic mixing, can cause clipping if not carefully avoided. In some embodiments, odd non - linearities can be applied to the harmonic spectral components generated by the filter bank module 120 to manage the resulting dynamics, as discussed in more detail in Figure 6 More detailed discussion.
[0097] Figure 5 is a block diagram of the inverse converter module 208 according to some embodiments. The inverse converter module 208 includes a rotation matrix module 502, a matrix multiplier 504, a projection operator 506, and a matrix transpose operator 508. The inverse converter module 208 generates harmonic spectral components from a rotation spectrum that includes phase - coherent rotated spectral orthogonal components and The rotation matrix module 502 generates the same rotation matrix as the rotation matrix generated by the matrix module 302. The rotation matrix generated by the rotation matrix module 502 is transposed by the matrix transpose operator 508 and applied by the matrix multiplier 504 to the incoming 2D vector of the phase - coherent rotated spectral orthogonal components generates harmonic spectral components and The resulting 2D vector is projected by the projection operator 506 onto a single dimension. To perform the inverse transformation from the rotated basis back to the standard basis, the output spectrum is shifted so that 0Hz returns to its original position θc, as defined by Equation 16:
[0098]
[0099]
[0100] where P is a projection from a two-dimensional real coefficient space to a single dimension, as defined by Equation 17:
[0101]
[0102] Since the forward transform R 2 (-θ c t) includes an orthogonal rotation, the inverse transform is the transpose. This algebraic structure allows caching of the forward transform matrix and simply inverting it by changing the order in which the coefficients are multiplied. It is in this sense that Figure 3 the rotation matrix module 302 in Figure 5 and the rotation matrix module 502 in are considered the same. The harmonic spectral component
[0103] Figure 6 is an example of the harmonic spectral component h(t)(n) and can thus be the response of a non-linear filter in a larger filter bank.
[0104] is a block diagram of the combiner module 106 according to some embodiments. The combiner module 106 performs further processing on the harmonic spectral components h(t)(n) from the filter bank module 120, combines the harmonic spectral components h(t)(n) to generate a combined component z(t), performs further processing on the combined component z(t), and combines the combined component z(t) with the filtered audio channel a(t) from the all-pass filter network module 122 to generate an output channel o(t).
[0105] For the constrained non-linearity defined in Equation 10, the large variations in the output level that may occur suggest taking additional measures to limit the instantaneous peak level. In the harmonic spectral component h(t)(n) (or as defined by Equation 16 ) After the creation of (), the component processor 602(n) applies a non-linearity to the signal, constraining it to the range (-1, 1). This non-linearity can be an odd non-linearity, such as the sigmoid function. This non-linearity typically preserves the sign and gradually slopes towards either extreme of the range. The hyperbolic tangent with a scaling factor is an example of such a function, as defined by Equation 18:
[0106]
[0107] When employed to reduce peaks, this non-linearity can also add odd harmonics to the harmonic spectral component h(t)(n). These odd harmonics will be in phase with the harmonics of the harmonic spectral component h(t)(n). The odd harmonics at this stage will transform the change in the overall amplitude into a change in timbre, in a manner that conforms to common human auditory cues for loudness.
[0108] When combined with a peak limiter, the peak limiting threshold can be set to a small amount below the threshold in Equation 18, such that the harmonic characteristics of the limiting function are dominated by the more perceptually meaningful hyperbolic tangent rather than the sharp corners of the peak limiter.
[0109] In some embodiments, one or more of the component processors 602(n) can attenuate (e.g., using independent tuning) their respective harmonic spectral components h(t)(n) to achieve the desired non-linear characteristics for the combined component z(t).
[0110] The harmonic spectral component combiner 604 combines the harmonic spectral components h(t)(n), such as the harmonic spectral components h(t)(1) to h(t)(n), to generate the combined component z(t).
[0111] The combined component processing module 606 processes the combined component z(t). The combined component processing module 606 can also apply various types of processing, such as high-pass filtering, dynamic range processing (e.g., limiting or compression), etc.
[0112] The output combiner 608 combines the combined component z(t) with the filtered audio channel a(t) from the all-pass filter network module 122 to generate the output channel o(t). In some embodiments, the output combiner 608 can attenuate the filtered audio channel a(t) or the combined component z(t) before combination.
[0113] Figure 7FIG. 0 is a block diagram of a filter bank module 700 according to some embodiments. The filter bank module 700 is an embodiment of the filter bank module 120. The filter bank module 700 uses a serial implementation where each downstream harmonic spectral component is generated using the residue of the upstream harmonic spectral component as input. Although the construction of a filter bank module with independent filters for parallel applications is relatively straightforward, tuning such a filter bank module can be a complex task. This difficulty is the result of the loss of power spectrum conservation. In practice, filter bank tuning with problematic power spectrum conservation typically gives the impression of short delays or comb filters at low frequencies, disturbing the listener's ability to determine timing. This occurs because the envelope hitting the low-frequency content usually drops in both amplitude and fundamental frequency simultaneously. Thus, the discontinuity in the power spectrum causes multiple transients to be perceived where there was only one transient before.
[0114] In the serial paradigm, each filter of the filter bank module 700 forks the signal between the frequency band to be analyzed and the residue of the incoming content. This is done by replacing the low-pass filter F(x) with a 2-band splitting network. Note that in some cases, this can be simply achieved by subtracting the low-pass signal from the wideband signal right before the low-pass operation. Subsequent filters operate only on the residue high-pass signal, omitting the spectral data previously acted upon by the upstream filters. As a result, the total spectral energy analyzed by the filter bank module 700 is the same as the total spectral energy at the input.
[0115] Just as in the parallel case, each serial filter uses independent forward and inverse transforms. This can be achieved in multiple ways. In a first example, the forward and inverse transforms of each filter are applied before moving on to the forward and inverse transforms of the downstream filter, and so on. In a second example, a pyramid algorithm is used where the coordinates for the forward transform of subsequent filters are transformed, which includes calculating the transformation matrix using the difference between the frequency shift θcn-1 of the upstream filter and the frequency shift θcn of the next filter. After all the forward transforms are applied, the inverse transforms can be applied in the reverse order, starting from the most downstream filter and moving up serially. This allows caching of the frequency increments between the forward and inverse steps.
[0116] The filter bank module 700 uses a pyramid algorithm for the forward and inverse transforms. In this example, the audio channel a(t) has N subbands that are processed serially, from subband 1 to subband N. Blocks op1 718, op2 734, and opM 752 perform coefficient operations on the first, second, and Nth subbands respectively. Each of op1 718, op2 734, and opM 752 can perform the coefficient operations discussed herein for the coefficient arithmetic unit module 206.
[0117] Blocks R704, R720, and R736 each perform the multiplication of the 2-D signal on the right side by the time-varying rotation matrix R, as discussed herein for the rotation matrix module 302. Block H702 represents the orthogonal filter operation described in Equation 1, where blocks H and R together perform the operation defined by Equation 2. 2 Blocks F706, F708, F722, F724, F740, and F742 each perform the low-pass filter operation F(x), such as that discussed herein for the filter module 402.
[0118] Blocks *(-1)710, *(-1)712, *(-1)726, *(-1)728, *(-1)744, and *(-1)746 invert the received input. Blocks +714, +716, +730, +732, +748, +750, +774, and +776 combine the received inputs to generate an output.
[0119] Blocks R754, R756, R762, R766, R764, and R772 perform the inverse transformation of the R block. For example, blocks R704 and R772 and R766 use a rotation of –(θc1t). Blocks R720 and R764 and R762 use a rotation of –(θc2 - θc1)t. Blocks R736 and R754 and R756 use a rotation of –(θcN – θc(N-1))t.
[0120] Block R -1 754, R -1 756, R -1 762, R -1 766, R -1 764 and R -1 772 perform the inverse transformation of the R block. For example, block R704 and R772 and R766 use a rotation of –(θc1t). Blocks R720 and R764 and R762 use a rotation of –(θc2 - θc1)t. Blocks R736 and R754 and R756 use a rotation of –(θcN – θc(N-1))t. -1 、772 and R -1 766 use a rotation of –(θc1t). Blocks R720 and R -1 、764 and R -1 762 use a rotation of –(θc2 - θc1)t. Blocks R736 and R -1 、754 and R -1 756 use a rotation of –(θcN – θc(N-1))t.
[0121] Block P778 performs the 1-D projection operation described in Equation 17.
[0122] Note the use of the difference between adjacent values of θcn, rather than the angular frequency θc. For certain choices of θcn, the pyramid algorithm can provide a more computationally efficient implementation by limiting the number of times the rotation R 2 (-θ c t) is calculated. A particularly computationally efficient choice for the distribution of θcn is linear (where the difference between θc for adjacent filters remains constant), thus completely minimizing the recalculation of R 2 (-θ c t) since the matrices will be identical to each other.
[0123] The final residual contains data that is not affected by the entire filter bank, eliminating the possibility of constructive or destructive interference between the affected and unaffected signals. The transfer function of this residual signal will perfectly match the analysis region of the filter bank. This does not necessarily mean a perfect reconstruction of the power spectrum of the output signal, as coefficient operations may cause modifications to the dynamic behavior or the synthesis of entirely new content. In many cases, this final residual can be completely discarded, and the output of H702 can be used to mix the unaffected content back into the final summation.
[0124] The filter bank module 700 uses the residuals of the upstream harmonic spectral components as inputs to generate each downstream harmonic spectral component. In this case, a filter bank topology containing M non - linear filters can be described as a serial architecture. Thus, the non - linear filters can be defined by an index m with values ranging from 1 to M. For example, blocks +714 and +716 output the residuals of the first - order harmonic spectral component (e.g., m = 1), which are used to generate the second - order harmonic spectral component (e.g., m = 2). Here, the residuals of the first - order harmonic spectral component refer to the parts in the audio channel that are filtered out by blocks F706 and F708 and thus not processed by block Op1 718. These residual parts are generated by inverting the filtered parts by blocks *(-1)710 and *(-1)712 and adding the inverted filtered parts to the filtered parts by blocks +714 and +716. Further downstream processing works in a similar manner. For example, blocks +730 and +732 output the residuals of the second - order harmonic spectral component, which are used to generate the third - order harmonic spectral component (e.g., m = 3), and so on.
[0125] Example process
[0126] Figure 8 is a flowchart of a process 800 for psychoacoustic frequency range extension according to some embodiments. Figure 8 The process shown can be performed by components of an audio system (e.g., audio system 100). In other embodiments, other entities can perform Figure 8 some or all of the steps. Embodiments can include different and / or additional steps, or perform these steps in a different order.
[0127] The audio system generates 805 orthogonal components that define an orthogonal representation of the audio channel. The audio channel can be a channel of a multi - channel audio signal, such as the left or right channel of a stereo audio signal. The orthogonal components include a 90° phase relationship. The orthogonal components and the audio channel include a unity amplitude relationship for all frequencies. In some embodiments, a real - valued input signal is tuned to orthogonal values by a matched pair of all - pass filters.
[0128] The audio system generates 810 rotated spectral orthogonal components by applying a forward transform that rotates the spectrum of the quadrature components (e.g., the entire spectrum) from a standard basis to a rotated basis. The standard basis refers to the frequencies of the input audio channels before rotation. The rotation can cause the target frequency to be mapped to 0 Hz. The target frequency can be the center of the analysis region of a harmonic processing module, such as the center frequency of a target subband for psychoacoustic range extension. The forward transform can be computed using iterative calls to trigonometric functions as defined by Equation 3 or using an equivalent recursive 2D rotation.
[0129] The audio system isolates the components of the 815 rotated spectral orthogonal components at the target frequency and target amplitude. Isolating these components can be performed in the rotated basis. For example, the target frequency can be isolated using a filter F(x), where x includes components defined by u(t). In some embodiments, the filter removes frequencies above a threshold, and this has the effect of symmetrically isolating the target subband that is twice the threshold across the center frequency θc to which the forward transform is tuned. In some embodiments, the audio system determines the target frequency based on factors such as the reproducible range of a speaker, a reduction in the power consumption of the speaker, or an increased service life of the speaker.
[0130] The audio system can also isolate the components at the target amplitude from the rotated spectral orthogonal components, such as by using a gating function. The gating function can be configured to discard unwanted information in a subband or to preserve the amplitude envelope. The gating function can also include a slew rate limiting filter or a similar smoothing function.
[0131] The audio system generates 820 weighted phase-coherent harmonic spectral orthogonal components by applying a scale-dependent non-linearity that complies with a constraint to the isolated components. The weighted phase-coherent rotated spectral orthogonal components can be generated in the rotated basis. This rotated basis is well-suited for the generation of a designer spectrum because it represents the standard basis signal as a 2-dimensional vector and because it concentrates the target frequency near zero. The vector can then be further decomposed into polar coordinates, as shown in Equation 4, which is similar to computing the magnitude and argument of a single bin in a short-time Fourier transform (STFT), a natural descriptor of information about a particular frequency. Compared to the STFT representation, this implementation has several distinct advantages. First, the bin information is computed only as needed, rather than for the entire spectrum. Another advantage is that the result is computed at the time resolution required to correctly represent transient data. Additionally, the filter (which operates similar to a window function in STFT techniques) can be conveniently tuned for the purpose of separating the target spectral content from its residue, and in the case of multiple harmonic processing modules, can have non-uniform tuning.
[0132] The nonlinearity, which functions mainly to generate a phase-coherent spectrum given the phase information in the orthogonal components of a rotational spectrum, can have scale dependence that complies with a constraint as defined by Equation 11. The nonlinearity includes a weighted mixture of component non-linearities, each component non-linearity being defined by Equation 10 and corresponding to a different harmonic n. The application of the nonlinearity to the isolated components is defined by Equation 9. For each harmonic n, the amplitude correction factor max(||u(t)||, b n ) defines the constraint on the gain correction of the input u(t) applied to the component non-linearity. The scale refers to the amplitude of the input component u(t), as defined by ||u(t)||, representing the energy present in the signal at time t. Different harmonics n can include different minimum value constraints b n . For example, lower harmonics (e.g., fundamental harmonic n = 1) can be unconstrained (e.g., b n = 0), while higher harmonics can be more constrained with higher values of b n .
[0133] The nonlinearity itself can include a weighted sum of Chebyshev polynomials of the first kind, where the amplitudes are selectively factored out subject to the constraint. Each component non-linearity of the nonlinearity can be weighted by a predefined harmonic weight a n , as defined by Equation 9.
[0134] The audio system generates 625 harmonic spectral components by applying an inverse transform that rotates the spectrum of the weighted phase-coherent rotational spectrum orthogonal components from the rotational basis to the standard basis. The inverse transform can rotate the spectrum such that 0 Hz is mapped to the target frequency. The harmonic spectral components include frequencies different from the target frequency but create a psychoacoustic impression of the target frequency when rendered by a loudspeaker. The frequencies of the harmonic spectral components can be within the bandwidth of the loudspeaker, while the sub-band frequencies can be outside the bandwidth of the loudspeaker. In some embodiments, the sub-band frequencies are lower than the frequencies of the harmonic spectral components. In some embodiments, the sub-band frequencies include frequencies between 18 Hz and 250 Hz. In some embodiments, the target sub-band or frequency can be within the reproducible range of the loudspeaker but can have been selected for application-specific reasons, e.g., to reduce the power consumption of the audio system or to increase the service life of the loudspeaker.
[0135] The audio system combines harmonic spectral components with frequencies 830 of the audio channel outside the target frequency to generate an output channel and provides 835 the output channel to a speaker. In some embodiments, the audio system generates an output channel by combining harmonic spectral components with the original audio channel and provides the output channel to a speaker. In some embodiments, the audio system filters the audio channel or other sub-band components of the audio channel (e.g., excluding the (multiple) sub-band components used for frequency range extension) to ensure that the audio channel or other sub-band components are coherent with the harmonic spectral components, and combines the filtered audio channel or other sub-band components with the harmonic spectral components to generate an output channel for the speaker. In some embodiments, the combination of the filtered or original audio channel and the harmonic spectral components can be further processed, such as using equalization, compression, etc., to generate an output channel for the speaker.
[0136] In steps 805 to 825, harmonic spectral components are generated for the frequency bands of the audio channel. In some embodiments, multiple harmonic spectral components are generated and combined 830, where each of the harmonic spectral components is generated using a different frequency band of the audio channel. The output channel can be generated by combining the frequencies of the audio channel outside the target frequency of the harmonic spectral components. The harmonic spectral components can be generated in parallel or serially. For the serial case, each downstream harmonic spectral component can be generated using the residue of the upstream harmonic spectral component as an input. In some embodiments, different speakers can have different available bandwidths or frequency responses. For example, a mobile device (e.g., a mobile phone) can include unbalanced speakers. Different sub-band components can be used for frequency range extension of different speakers.
[0137] Example computer
[0138] Figure 9 is a block diagram of a computer 900 according to some embodiments. The computer 900 is an example of a circuit system that implements an audio system and its components, such as the audio system 100 or the filter bank module 120 or the filter bank module 700. At least one processor 902 coupled to a chipset 904 is shown. The chipset 904 includes a memory controller hub 920 and an input / output (I / O) controller hub 922. A memory 906 and a graphics adapter 912 are coupled to the memory controller hub 920, and a display device 918 is coupled to the graphics adapter 912. A storage device 908, a keyboard 910, a pointing device 914, and a network adapter 916 are coupled to the I / O controller hub 922. The computer 900 can include various types of input or output devices. Other embodiments of the computer 900 have different architectures. For example, in some embodiments, the memory 906 is directly coupled to the processor 902.
[0139] The storage device 908 includes one or more non-transitory computer-readable storage media, such as a hard disk drive, a compact disc read-only memory (CD-ROM), a DVD, or a solid-state memory device. The memory 906 holds program code (including one or more instructions) and data used by the processor 902. The program code may correspond to the processing aspects described with reference to Figures 1 to 8 the description.
[0140] The pointing device 914 is used in combination with the keyboard 910 to input data into the computer system 900. The graphics adapter 912 displays images and other information on the display device 918. In some embodiments, the display device 918 includes touchscreen capabilities for receiving user input and selections. The network adapter 916 couples the computer system 900 to a network. Some embodiments of the computer 900 have different and / or other components than those shown in Figure 9 those shown.
[0141] The circuitry may include one or more processors that execute program code stored on a non-transitory computer-readable medium, which, when executed by the one or more processors, configures the one or more processors to implement an audio processing system or a module of an audio processing system. Other examples of circuitry that implements an audio processing system or a module of an audio processing system may include integrated circuits, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other types of computer circuitry.
[0142] Additional Notes
[0143] Example benefits and advantages of the disclosed configurations include allowing speakers to effectively render (e.g., lower) frequencies that are beyond the physical capabilities of the speakers. By processing the audio signals as discussed herein, the rendered sound creates the impression of frequencies that are outside the bandwidth of the physical drivers.
[0144] Throughout this specification, multiple instances may implement components, operations, or structures described as a single instance. Although individual operations of one or more methods are shown and described as separate operations, one or more of the given operations may be executed concurrently, and the operations need not be executed in the order shown. Structures and functions presented as separate components in example configurations may be implemented as a combined structure or component. Similarly, structures and functions presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein.
[0145] Certain embodiments are described herein as including logic or a plurality of components, modules, blocks, or mechanisms. A module can be a software module (e.g., code embodied on a machine-readable medium or in a transmitted signal) or a hardware module. A hardware module is a tangible unit capable of performing certain operations and can be configured or arranged in a certain manner. In an example embodiment, one or more computer systems (e.g., standalone client or server computer systems) or one or more hardware modules of a computer system (e.g., a processor or a group of processors) can be configured by software (e.g., an application or a portion of an application) to operate as a hardware module that performs certain operations described herein.
[0146] The various operations of the example methods described herein can be performed, at least in part, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors can constitute processor-implemented modules that operate to perform one or more operations or functions. In some example embodiments, the modules referred to herein can include processor-implemented modules.
[0147] Similarly, the methods described herein can be at least in part processor-implemented. For example, at least some operations of a method can be performed by one or more processors or processor-implemented hardware modules. Execution of certain operations can be distributed among one or more processors, not only residing within a single machine but also deployed across multiple machines. In some example embodiments, one or more processors can be located in a single location (e.g., within a home environment, an office environment, or as a server farm), while in other embodiments, the processors can be distributed across multiple locations.
[0148] Unless otherwise specifically stated, discussions herein using terms such as "processing," "computing," "calculating," "determining," "presenting," "displaying," etc., can refer to actions or processes of a machine (e.g., a computer) that manipulates or transforms data represented as physical (e.g., electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.
[0149] As used herein, any reference to "one embodiment" or "an embodiment" means that a particular element, feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The appearances of the phrase "in one embodiment" in various places in the specification are not necessarily all referring to the same embodiment.
[0150] Some embodiments may be described using the terms "coupled" and "connected" and their derivatives. It should be understood that these terms are not intended as synonyms for each other. For example, some embodiments may be described using the term "connected" to indicate that two or more elements are in direct physical or electrical contact with each other. In another example, some embodiments may be described using the term "coupled" to indicate that two or more elements are in direct physical or electrical contact. However, the term "coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other. The embodiments are not limited to this context.
[0151] As used herein, the terms "comprises", "comprising", "includes", "including", "has", "having" or any other variation thereof are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, "or" refers to an inclusive or rather than an exclusive or. For example, condition A or B is satisfied by any one of the following: A is true (or present) and B is false (or not present), A is false (or not present) and B is true (or present), and both A and B are true (or present).
[0152] In addition, the articles "a" or "an" are used to describe elements and components of the embodiments herein. Doing so is merely for convenience and to give a general sense of the invention. This description should be understood to include one or at least one, and the singular also includes the plural unless there is an obvious contrary meaning.
[0153] Some portions of this description describe embodiments in terms of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are commonly used by those skilled in the data processing arts to effectively convey the substance of their work to others skilled in the art. While functionally, computationally, or logically described, these operations are to be understood as being implemented by a computer program or equivalent circuitry, microcode, etc. Additionally, it is sometimes convenient to refer to the arrangement of these operations as a module, without loss of generality. The described operations and their associated modules can be embodied in software, firmware, hardware, or any combination thereof.
[0154] Any step, operation, or process described herein can be performed or implemented singly or in combination with other devices using one or more hardware or software modules. In one embodiment, the software module is implemented using a computer program product that includes a computer-readable medium containing computer program code that can be executed by a computer processor to perform any or all of the described steps, operations, or processes.
[0155] Embodiments may also relate to apparatus for performing the operations herein. The apparatus may be specially constructed for the required purposes, and / or it may comprise a general-purpose computing device selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a non-transitory tangible computer-readable storage medium or any type of medium suitable for storing electronic instructions, which may be coupled to a computer system bus. Additionally, any computing system referred to in this specification may include a single processor or may be an architecture employing a multi-processor design to increase computing power.
[0156] Embodiments may also relate to products produced by the computing processes described herein. Such products may include information produced by the computing process, where the information is stored on a non-transitory, tangible computer-readable storage medium and may include any embodiment of a computer program product or other data combinations described herein.
[0157] After reading this disclosure, those skilled in the art will appreciate additional alternative structural and functional designs for systems and processes based on the principles disclosed herein. Thus, while specific embodiments and applications have been shown and described, it should be understood that the disclosed embodiments are not limited to the precise structures and components disclosed herein. Various modifications, changes, and variations that are obvious to those skilled in the art can be made to the arrangement, operation, and details of the methods and apparatus disclosed herein without departing from the spirit and scope defined in the appended claims.
[0158] Finally, the language used in this specification has been principally selected for readability and guidance purposes and may not have been selected to delineate or circumscribe the scope of the patent rights. Accordingly, it is intended that the scope of the patent rights not be limited by this detailed description, but rather by any claims issued on an application based hereon. Thus, the disclosure of the embodiments is intended to illustrate rather than limit the scope of the patent rights set forth in the appended claims.
Claims
1. A system for psychoacoustic frequency range extension, comprising: A circuit system configured to: Generate in-phase and quadrature components from an audio channel, the in-phase and quadrature components defining an in-phase and quadrature representation of the audio channel; Generate rotated spectral in-phase and quadrature components by applying a forward transform that rotates the spectrum of the in-phase and quadrature components from a standard basis to a rotated basis; In the rotated basis: Isolate the components at a target frequency in the rotated spectral in-phase and quadrature components; And Generate weighted phase-coherent harmonic spectral in-phase and quadrature components by applying a non-linear function to the isolated components, the non-linear function having a scale-dependence that complies with constraints; Generate harmonic spectral components by applying an inverse transform that rotates the spectrum of the weighted phase-coherent harmonic spectral in-phase and quadrature components from the rotated basis to the standard basis; Combine the harmonic spectral components with the frequencies of the audio channel outside the target frequency to generate an output channel; And Provide the output channel to a loudspeaker.
2. The system according to claim 1, wherein: The non-linear function comprises a weighted mixture of component non-linear functions; The constraints each comprise a constraint on gain correction for the input applied to the respective component non-linear function.
3. The system according to claim 2, wherein the non-linear function comprises a weighted sum of Chebyshev polynomials of the first kind, the magnitude of which is selectively extracted in compliance with the constraints.
4. The system according to claim 1, wherein the circuit system is further configured to generate a plurality of harmonic spectral components, each harmonic spectral component being generated using a different frequency band of the audio channel, and wherein the circuit system is configured to generate the output channel by combining the plurality of harmonic spectral components.
5. The system according to claim 4, wherein the circuit system is configured to generate the plurality of harmonic spectral components serially, wherein each downstream harmonic spectral component uses the residue of the upstream harmonic spectral component as an input.
6. The system according to claim 4, wherein the circuit system is configured to generate the plurality of harmonic spectral components in parallel.
7. The system according to claim 1, wherein the circuit system is further configured to apply an odd non-linear function to the harmonic spectral components.
8. The system according to claim 1, wherein the harmonic spectral components include frequencies different from the target frequency of the audio channel and create a psychoacoustic impression of the target frequency when rendered by the loudspeaker.
9. The system according to claim 1, wherein: The forward transform rotates the spectrum of the in-phase and quadrature components such that the target frequency is mapped to 0 Hz; and The inverse transform rotates the spectrum of the weighted phase-coherent harmonic spectral in-phase and quadrature components such that 0 Hz is mapped to the target frequency.
10. The system according to claim 1, wherein the target frequency includes a frequency between 18 Hz and 250 Hz.
11. The system according to claim 1, wherein the circuit system is further configured to determine the target frequency based on at least one of: The reproducible range of the loudspeaker; a reduction in power consumption of the loudspeaker; or an increased service life of the loudspeaker.
12. The system according to claim 1, wherein the loudspeaker is a component of a mobile device.
13. The system according to claim 1, wherein the circuitry is further configured to isolate the component at a target amplitude using a gate function.
14. The system according to claim 1, wherein the circuitry is further configured to apply a smoothing function to the isolated component.
15. A non-transitory computer-readable medium comprising stored instructions that, when executed by at least one processor, configure the at least one processor to: generate orthogonal components from an audio channel, the orthogonal components defining an orthogonal representation of the audio channel; generate rotated spectral orthogonal components by applying a forward transform that rotates the spectrum of the orthogonal components from a standard basis to a rotated basis; in the rotated basis: isolate the component at a target frequency in the rotated spectral orthogonal components ; and generate weighted phase-coherent harmonic spectral orthogonal components by applying a non-linear function to the isolated component, the non-linear function having a scale dependence that complies with a constraint; generate harmonic spectral components by applying an inverse transform that rotates the spectrum of the weighted phase-coherent harmonic spectral orthogonal components from the rotated basis to the standard basis; combine the harmonic spectral components with frequencies of the audio channel outside the target frequency to generate an output channel; and provide the output channel to a loudspeaker.
16. The non-transitory computer-readable medium according to claim 15, wherein: the non-linear function comprises a weighted mixture of component non-linear functions; each of the constraints comprises a constraint on gain correction applied to the input of a corresponding component non-linear function.
17. The non-transitory computer-readable medium according to claim 16, wherein the non-linear function comprises a weighted sum of Chebyshev polynomials of the first kind, the magnitude of which is selectively factored out in compliance with the constraint.
18. The non-transitory computer-readable medium according to claim 15, wherein: the instructions further configure the at least one processor to generate a plurality of harmonic spectral components, each harmonic spectral component being generated using a different frequency band of the audio channel; the output channel is generated by combining the plurality of harmonic spectral components; and the plurality of harmonic spectral components are generated serially, wherein each downstream harmonic spectral component uses the residue of an upstream harmonic spectral component as an input.
19. The non-transitory computer-readable medium according to claim 15, wherein the instructions further configure the at least one processor to apply an odd non-linear function to the harmonic spectral components.
20. A method for psychoacoustic frequency range extension, comprising, by circuitry: generate orthogonal components from an audio channel, the orthogonal components defining an orthogonal representation of the audio channel; generate rotated spectral orthogonal components by applying a forward transform that rotates the spectrum of the orthogonal components from a standard basis to a rotated basis; in the rotated basis: Isolate the component at the target frequency in the rotated spectral orthogonal component ; and Generate a weighted phase coherent harmonic spectral orthogonal component by applying a non - linear function to the isolated component, the non - linear function having a scale - dependence that complies with the constraints; Generate a harmonic spectral component by applying an inverse transform that rotates the spectrum of the weighted phase coherent harmonic spectral orthogonal component from the rotated basis to the standard basis; Combine the harmonic spectral component with the frequencies of the audio channel outside the target frequency to generate an output channel; and Provide the output channel to a speaker.
21. The method according to claim 20, wherein: The non - linear function includes a weighted mixture of component non - linear functions; The constraints each include a constraint on the gain correction applied to the input of the corresponding component non - linear function.
22. The method according to claim 21, wherein the non - linear function includes a weighted sum of Chebyshev polynomials of the first kind, the magnitudes of which are selectively factored out in compliance with the constraints.
23. The method according to claim 20, further comprising generating a plurality of harmonic spectral components by the circuit system, each harmonic spectral component being generated using a different frequency band of the audio channel, and wherein: The output channel is generated by combining the plurality of harmonic spectral components; and The plurality of harmonic spectral components are generated serially, wherein each downstream harmonic spectral component uses the residue of the upstream harmonic spectral component as an input.
24. The method according to claim 20, further comprising applying an odd non - linear function to the harmonic spectral component by the circuit system.
Citation Information
Patent Citations
Apparatus and method for canceling acoustic echoes including non-linear distortions in loudspeaker telephones
CN1176034A
Constrained nonlinear parameter estimation for robust nonlinear loudspeaker modeling for the purpose of smart limiting
US20190200146A1