Virtual Bass Enhancement based on Source Separation
By segmenting audio signals into distinct acoustic sources and applying virtual bass enhancement independently, the method reduces distortion and complexity, achieving effective bass enhancement for small speakers.
Patent Information
- Application Number
- JP2024065364
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-04-15
- Filing Date
- 2024-04-15
- Publication Date
- 2025-09-17
- Estimated Expiration
- 2044-04-15
AI Technical Summary
Existing virtual bass enhancement algorithms suffer from intermodulation distortion and computational complexity, particularly in time-domain methods, while frequency-domain methods lose temporal resolution and introduce smearing effects.
The method involves segmenting an audio signal into distinct acoustic sources using neural networks, applying virtual bass enhancement to each segment independently, and combining the processed segments to generate an enhanced audio signal, utilizing time-domain algorithms to handle transients and frequency-domain methods for tonal components.
This approach reduces intermodulation distortion and computational complexity, while maintaining effective bass enhancement by allowing independent processing of each acoustic source, enhancing bass perception without signal distortion.
Smart Images

Figure 0007741234000003 
Figure 0007741234000004 
Figure 0007741234000005
Abstract
Description
[Technical Field]
[0001] The present invention relates to the field of audio signal processing, and in particular to a method and apparatus for improving the audio characteristics of a speaker in the bass or low frequency range. [Background technology]
[0002] Due to physical limitations, small speakers are characterized by poor acoustic response, especially at low frequencies. Typical small speakers, such as those found in portable electronic devices such as smartphones and laptops, exhibit a cutoff frequency of approximately 150 Hz for electrodynamic speakers and approximately 300 Hz for piezoelectric speakers. This results in poor reproduction of audio signals in the low-frequency range, which is typically understood to be in the 20 Hz to 300 Hz range below the cutoff frequency.
[0003] Common approaches based on linear filters such as equalizers can damage the transducer, cause unwanted distortion, and ultimately fail to solve the problem.
[0004] This problem has been addressed through two main approaches: first, new transducers have been developed that overcome the physical limitations of their operation according to the device design; second, signal processing algorithms have been developed to enhance the acoustic performance of the transducers. In the latter approach, a class of digital signal processing algorithms is known as virtual bass enhancement (VBE).
[0005] VBE dates back to the 90s, when the idea of psychoacoustic effects began to be explored. In particular, several algorithms known in the prior art are based on the so-called missing fundamental phenomenon. According to this effect, the human brain is able to perceive low frequencies as present due to the periodicity of higher harmonics, even when these frequencies are not physically reproduced. In other words, the human brain is able to reconstruct the missing fundamental from the higher harmonics.
[0006] Over the past few decades, different VBE algorithms have been proposed, which can be divided into two main categories: time-domain and frequency-domain methods.
[0007] Time-domain methods are simple, lightweight, and perform well on transients. In this method, the low end is typically extracted from the audio track using a crossover network. A nonlinear device (NLD) is then applied to generate harmonics, and finally the harmonically enhanced track is weighted and added to a high-pass version of the original signal to produce a bass-enhanced audio track.
[0008] The time domain VBE algorithm is known from, for example, the following Non-Patent Documents 1 to 5.
[0009] Frequency domain approaches, on the other hand, are based on phase vocoders and tend to perform better on tonal components than on transients. In this approach, a pitch shift is typically applied to map frequencies originally below the transducer cutoff frequency to higher regions of the frequency spectrum. The newly introduced harmonics are then weighted according to a frequency envelope or equal loudness contour.
[0010] Finally, hybrid methods have been proposed to combine the advantages of the two approaches. These methods aim to apply time-domain methods to the transients and frequency-domain methods to the tonal parts of the audio track. This is usually achieved by applying such a decomposition in the frequency domain. Hybrid methods often have the disadvantage of being computationally expensive, making them inapplicable to real-time scenarios.
[0011] The frequency domain and combined VBE algorithms are known from, for example, the following Non-Patent Documents 6 to 8. [Prior art documents] [Non-patent literature]
[0012] [Non-Patent Document 1] E. Larsen and RM Aarts, Audio Bandwidth Extension: Application of Psychoacoustics, Signal Processing and Loudspeaker Design. John Wiley and Sons, Ltd, 2004 [Non-patent document 2] D. Ben-Tzur, “The effect of the maxxbass 1 psychoacoustic bass enhancement system on loudspeaker design,” in Proceedings of the 106th Audio Engineering Society Convention, 5 1999 [Non-patent document 3] N. Oo, W.-S. Gan, and M. O. J. Hawksford, “Perceptually-motivated objective grading of nonlinear processing in virtual-bass systems,” Journal of the Audio Engineering Society, vol. 59, pp. 804-824, 12 2011 [Non-Patent Document 4] N. Oo and W.-S. Gan, “Harmonic analysis of nonlinear devices for virtual bass system,” in Proc. Int. Conf. Audio, Language, and Image Processing, 8 2008, pp. 279-284 [Non-Patent Document 5] R. Giampiccolo, A. Bernardini, and A. Sarti, “A Time-Domain Virtual Bass Enhancement Circuital Model for Real-Time Music Applications,”IEEE 24th International Workshop on Multimedia Signal Processing (MMSP), Shanghai, China, 26-28 September 2022 [Non-Patent Document 6] M. R. Bai and W.-C. Lin, “Synthesis and implementation of virtual bass system with a phase-vocoder approach,” Journal of the Audio Engineering Society, vol. 54, pp. 1077-1091, 2006 [Non-Patent Document 7] E. Moliner, J. Ramo, and V. Valimaki, “Virtual bass system with fuzzy separation of tones and transients,” in Proceedings of the 23rd International Conference on Digital Audio Effects (DAFx2020), 9 2020 [Non-patent document 8] AJ Hill and MOJ Hawksford, “A hybrid virtual bass system for optimized steady-state and transient performance,” in Proceedings of the 2nd Computer Science and Electronic Engineering Conference (CEEC), 9 2010, pp. 1-6 Summary of the Invention [Problem to be solved by the invention]
[0013] However, known time-domain techniques suffer from intermodulation distortion (IMD): in particular, feeding a low-pass version of the original audio track (e.g., a polyphonic mixture of instruments) into a nonlinear function such as NLD will generate harmonics of several frequency components at once, inevitably producing unpleasant inharmonic distortions.
[0014] Frequency domain methods, on the other hand, suffer from the smearing effect caused by frame-by-frame processing, which results in a loss of temporal resolution and thus adversely affects the perception of transients and onsets. Furthermore, while they offer improved control over harmonic generation, they usually come at the expense of a higher computational load.
[0015] Therefore, there is a need for an improved virtual bass enhancement algorithm that can make a listener perceive lower low frequency sounds than the physical capabilities of the loudspeaker without distorting the signal and with manageable computational complexity. [Means for solving the problem]
[0016] In general, the present invention is based on the observation that previously known virtual bass enhancement techniques can be improved by applying them to selected portions of the input audio signal rather than to the entire audio signal, and more specifically, the present invention is based on how such selected portions are extracted from the input audio signal.
[0017] In particular, contrary to previously known methods that operate on a portion of the input audio signal (e.g., a low-pass signal introduced by a crossover network), the present invention applies VBE to a segmented music stem. In other words, the present invention relates to applying VBE to a segmented sound source or group of such sound sources that share a common sound production mechanism, such as, for example, multiple vocal lines, percussion, or a string ensemble. In other words, the present invention is characterized by applying VBE to each portion independently of the others. As will become clear from the following description, this can be advantageously obtained by using a music separation model to extract such components. In this way, intermodulation distortions can be avoided.
[0018] Additionally, different pre- and post-processing stages can be added to the signal processing pipeline to further improve bass enhancement.
[0019] Accordingly, an embodiment may relate to a virtual bass enhancement device for enhancing virtual bass in an input audio signal, comprising: a separation unit configured to extract from the input audio signal at least one audio channel corresponding to an acoustic source or a group of acoustic sources in the input audio signal; at least one virtual bass enhancement unit configured to generate harmonics for enhancing bass perception in the audio channel; and at least one adder configured to add the harmonics to the input audio signal to generate the enhanced audio signal.
[0020] In some embodiments, the separator may include at least one neural network trained to extract at least one audio channel from the input audio signal.
[0021] In some embodiments, the separator may include multiple neural networks trained to respectively extract multiple audio channels from the input audio signal.
[0022] In some embodiments, the virtual bass enhancement device may further comprise at least one filter unit configured to filter the at least one audio channel and output the at least one filtered audio channel, and the at least one virtual bass enhancement unit may be configured to generate harmonics for enhanced bass perception in the filtered audio channel.
[0023] In some embodiments, the virtual bass enhancer may be a time-domain virtual bass enhancer.
[0024] In some embodiments, at least one filter section may be a linear phase digital filter or a zero phase digital filter.
[0025] In some embodiments, the virtual bass enhancement device may further comprise at least one subtractor configured to subtract the at least one filtered audio channel from the input audio signal.
[0026] In some embodiments, the at least one virtual bass enhancement section may include a normalization section, a non-linear device, and an amplification section.
[0027] In some embodiments, the at least one virtual bass enhancement unit may be configured to implement at least a function f(x) having consecutive first and second derivatives with values less than 1 in the interval (0,1].
[0028] In some embodiments, the at least one virtual bass enhancer may be configured to implement at least the function f(x)=tanh(kx), where k is a predetermined value, preferably equal to and / or greater than 1.
[0029] In some embodiments, the at least one virtual bass enhancement unit may be configured to implement at least a function f(x) according to equation (2).
number
[0030] In some embodiments, the virtual bass enhancement device may further comprise a high-pass filter that receives the enhanced audio signal as an input and outputs a filtered enhanced audio signal, and a peak normalization unit and a loudness normalization unit that operate on the filtered enhanced audio signal.
[0031] In some embodiments, the virtual bass enhancement device may be configured for use with a transducer having a cutoff frequency, and the high-pass filter may have a cutoff frequency that corresponds to the cutoff frequency of the transducer.
[0032] In some embodiments, the audio source includes either drums, vocals, or instruments. [Brief explanation of the drawings]
[0033] [Figure 1] FIG. 1 shows a schematic diagram of a virtual bass enhancement device 1000 . [Figure 2] FIG. 2 shows a schematic diagram of a virtual bass enhancement device 2000 that differs from the virtual bass enhancement device 1000 by further comprising a plurality of filter sections 2410 to 241N. [Figure 3] FIG. 3 shows a schematic representation of a virtual bass enhancement device 3000 that differs from the virtual bass enhancement device 2000 by further comprising a plurality of subtractors 3510-351N. [Figure 4] FIG. 4 shows a schematic representation of a virtual bass enhancement unit 4210. [Figure 5] FIG. 5 shows a schematic representation of a virtual bass enhancement device 5000 that differs from any of the virtual bass enhancement devices 1000, 2000, 3000 by further comprising post-processing elements 5610, 5620, 5630. DETAILED DESCRIPTION OF THE INVENTION
[0034] 1 schematically illustrates a virtual bass enhancement device 1000. The virtual bass enhancement device 1000 is generally configured to enhance the virtual bass of an input audio signal IN. Those skilled in the art will appreciate that the input audio signal IN may be an analog or digital signal, and elements such as filters described below may be configured accordingly.
[0035] The input audio signal IN is typically the result of combining multiple acoustic sources in a single audio signal. For example, a band including drums, bass, guitar, and vocals may be recorded on an audio track as a result of combining these acoustic sources. Therefore, in the context of this application, the term acoustic source may be understood to correspond to, for example, a physical or synthesized instrument or voice.
[0036] In preferred embodiments, the acoustic source may include drums, vocals, or a musical instrument. In particularly preferred embodiments, the acoustic source may include drums. In particularly preferred embodiments, the acoustic source may include any musical instrument whose spectral energy is mostly located at frequencies below 500 Hz, preferably below 250 Hz. Alternatively or additionally, in particularly preferred embodiments, the acoustic source may include any musical instrument whose peak radiation frequency is located at frequencies below 500 Hz, preferably below 250 Hz. The peak radiation frequency may be understood as the radiation frequency whose amplitude is highest. Alternatively or additionally, in particularly preferred embodiments, the acoustic source may include any musical instrument having a fundamental frequency, and more preferably, any musical instrument whose dominant fundamental frequency is located at frequencies below 500 Hz, preferably below 250 Hz. Here, the dominant fundamental frequency may be understood as the frequency with the highest amplitude when multiple fundamental frequencies are included.
[0037] As will become apparent below, in contrast to the prior art in which VBE processing is applied either to the entire input audio signal IN or to components resulting from the output of various filters, the present invention provides an innovative aspect of separating at least one acoustic source from the input audio signal IN and then applying VBE processing to the resulting separated at least one acoustic source.
[0038] To this end, the virtual bass enhancement device 1000 includes a separation unit 1100. The separation unit 1100 is generally configured to extract at least one audio channel 1110-111N from the input audio signal IN. The audio channel 1110-111N may correspond to a single audio source in the input audio signal IN, such as a drum or bass guitar, or a group of audio sources, such as the entire drum and cymbals of a drum set. Separating the single audio source in a given audio channel 1110-111N allows for greater flexibility and granularity in signal processing, particularly VBE processing, and allows it to be applied to a specific audio source. Conversely, if multiple audio sources are included in a single audio channel, such as a drum and bass, the granularity may be reduced, but the computational burden is reduced.
[0039] It will be apparent to those skilled in the art that several methods are available for separating the input audio signal IN into multiple audio channels. In the preferred embodiment below, the separator 1100 is described using one or more trained neural networks, but it will be apparent that the present invention is not limited to this.
[0040] The virtual bass enhancement device 1000 includes at least one virtual bass enhancement unit 1210-121N, preferably a time-domain virtual bass enhancement unit 1210-121N, although the present invention is not limited thereto and a frequency-domain virtual bass enhancement unit may alternatively be used. The at least one virtual bass enhancement unit 1210-121N is configured to generate harmonics for enhancing bass perception in the audio channels 1110-111N. Preferably, the number of virtual bass enhancement units 1210-121N corresponds to the number of audio channels 1110-111N or is less than the number of audio channels 1110-111N if generation of harmonics for virtual bass enhancement is desired for only some of the audio channels 1110-111N.
[0041] It will be apparent to those skilled in the art that several methods are available for generating harmonics for the purpose of enhancing or improving the bass characteristics of a signal. Even when limited to time-domain VBE algorithms, several such algorithms are available. It is apparent that any of them may be employed in the present invention unless otherwise indicated or unless technically inconsistent with other factors.
[0042] The virtual bass enhancement device 1000 further comprises at least one adder 1310-131N configured to add harmonics to the input audio signal IN to generate an enhanced audio signal OUT. Preferably, the number of adders 1310-131N corresponds to the number of audio channels 1110-111N. In this way, the enhanced audio signal OUT may include the various audio channels 1110-111N after one or more have been processed via a VBE algorithm.
[0043] In other words, the embodiment of Figure 1 splits or separates the input audio signal IN into multiple audio channels corresponding to different acoustic sources, applies VBE processing to at least one of the audio channels, and then recombines the audio channels to obtain the enhanced audio signal OUT, preferably all of the audio channels are combined to obtain the enhanced audio signal OUT.
[0044] This approach advantageously allows for the independent generation of harmonics from the VBE processing to be avoided, resulting in objectionable inharmonic distortion (IMD), for a given acoustic source, or group of acoustic sources that are known to produce an acceptable level of inharmonic distortion (IMD) when applied together to the VBE processing.
[0045] This approach therefore overcomes one of the main drawbacks of known VBE algorithms, particularly the time-domain based approaches, while retaining all the advantages of known VBE algorithms, in particular their low computational requirements and their ability to operate on transients.
[0046] As mentioned above, various methods are known to those skilled in the art for separating an audio signal into multiple channels based on respective acoustic or instrumental sources. In a preferred embodiment of the present application, the separator 1100 may include at least one neural network trained to extract at least one audio channel 1110-111N from the input signal IN.
[0047] This approach is particularly advantageous because neural networks have been found to be particularly effective at correctly separating different acoustic sources into different respective channels.
[0048] Furthermore, it has been found that while a single neural network can be trained to recognize and segment multiple acoustic sources, segmentation of various acoustic sources can be successfully performed by multiple neural networks, each trained to recognize and segment one or more acoustic sources. Thus, in some embodiments, separator 1100 may include multiple neural networks trained to respectively extract multiple audio channels 1110-111N from input signal IN. Preferably, each of the multiple neural networks, and more preferably each of all neural networks, may be trained to recognize and segment a single corresponding acoustic source.
[0049] In this way, it is advantageously possible to train one neural network for each audio channel 1110-111N: one for vocals, another for drums, another for bass, etc. This has been found to be particularly advantageous because the type of training required to recognize one acoustic source, such as vocals, is often different from the type required to recognize another acoustic source, such as drums.
[0050] In preferred embodiments, more channels are preferable than fewer. Indeed, dividing the input signal IN into more channels generally allows for better control over the processing applied to each individual instrument, audio source, or stem. In principle, the number of channels in existing separation models is limited only by the availability of training data, and the separator is not inherently limited by a particular set of instruments. However, not all instruments contain significant energy in the low-end or bass portion of the frequency spectrum. Therefore, those instruments or audio sources may be assigned to a single "other" channel with little or no impact on the proposed system.
[0051] FIG. 2 shows a schematic diagram of a virtual bass enhancement device 2000 that differs from the virtual bass enhancement device 1000 by further comprising a plurality of filter sections 2410 to 241N.
[0052] In particular, the virtual bass enhancement device 2000 further comprises at least one filter unit 2410-241N configured to filter at least one audio channel 1110-111N, respectively, and to output the respective filtered audio channel 1110-111N, and a corresponding virtual bass enhancement unit 1210-121N may then be configured to operate on the respective filtered audio channel 1110-111N, for example, to generate harmonics for enhanced bass perception.
[0053] This particular approach allows harmonics to be generated for particular portions of the audio channels 1110-111N, such as those more associated with bass. The filter sections 2410-241N can be configured according to the sonic characteristics of the particular channel. For example, in a preferred embodiment, a low-pass filter may be used to extract the low end from the drum channel audio. Alternatively, or in addition, flat transfer functions may be used for the other channels, i.e., no filters may be applied. This can be advantageous, since in some cases it may be desirable to operate each channel across the entire frequency spectrum. It will be apparent that a similar effect can be achieved by removing the filter sections 2410-241N.
[0054] In some preferred embodiments, the virtual bass enhancers 1210-121N and 4210 are time-domain virtual bass enhancers 1210-121N and 4210. As noted above, frequency-domain virtual bass enhancers may also be used as the time-domain virtual bass enhancers 1210-121N and 4210. In some further embodiments, some of the virtual bass enhancers 1210-121N and 4210 may be time-domain based and some may be frequency-domain based.
[0055] Preferably, at least one of the filter sections 2410-241N, and preferably a majority of them, and even more preferably all of them, are linear phase or zero phase digital filters, which is particularly advantageous as it avoids introducing phase distortions that could change the shape of the waveform and thereby disturb the results of downstream operations, such as adding or subtracting signals as they pass through the adders 1310-131N.
[0056] FIG. 3 shows schematically a virtual bass enhancement device 3000 which differs from the virtual bass enhancement device 2000 by further comprising at least one subtractor 3510-351N.
[0057] In particular, at least one subtractor 3510-351N may be configured to subtract at least one filtered audio channel 1110-111N from the input audio signal IN, and preferably the unfiltered audio channel 1110-111N or the filtered audio channel 1110-111N processed by the respective virtual bass enhancement unit is added back to the input audio signal IN.
[0058] This approach advantageously makes it possible to avoid double consideration of the filtered audio channels 1110-111N, and in embodiments in which the filter units 2410-241N are low-pass filters, in particular, low frequency parts of the audio signal may be used in the VBE processing, while not themselves being included in the enhanced audio signal OUT.
[0059] In the embodiments described so far, it has been explained that in principle any known virtual bass enhancement algorithm may be used for the virtual bass enhancement units 1210-121N, and preferably a time domain algorithm may be used. In addition, Figure 4 schematically shows a virtual bass enhancement unit 4210 that can implement any of the virtual bass enhancement units 1210-121N.
[0060] 4, the virtual bass enhancement unit 4210 includes either a normalization unit 4211, a non-linear device 4212, or an amplification unit 4213. It will be apparent that any of these elements can be implemented without the others.
[0061] The purpose of the normalizer 4211 is generally to normalize the signal, preferably within a given time window. This is advantageous because it improves the operation of the non-linear device (NLD) 4212. In particular, the NLD 4212 can generate fewer harmonics in number and / or amplitude when the input signal spans a range of values where non-linear behavior is less pronounced, also known as the quasi-linear region. Conversely, signals spanning a range where the non-linear behavior of the NLD is more pronounced experience greater harmonic enhancement.
[0062] Therefore, in a preferred embodiment, the normalizer 4211 may be configured to normalize the signal generally so that the normalized signal has values that are not limited to the quasi-linear region of the NLD 4212, preferably so that the normalized signal has values outside the quasi-linear region of the NLD 4212.
[0063] In a preferred embodiment, this is achieved by normalizing the input signal on a frame-by-frame basis. For example, in a digital implementation, normalization may be applied to all digital samples contained in a time window of length M that slides over the input signal as new samples are processed, and possibly one sample at a time. That is, normalization may be performed on samples in a time sliding window of predetermined duration. This determines a time-varying normalization whose parameters are updated over time by depending on the past M samples. It will be apparent that multiple normalization algorithms may be employed within a given window.
[0064] In some preferred embodiments, the normalizer 4211 may be performed by adaptive rescaling. For example, in a possible implementation, the normalizer 4211 may be configured to divide an input sample by the maximum absolute value of the past M samples, so that the extreme values of the short-term signal within the window are ±1. The normalizer 4211 may further be configured to multiply each sample thus obtained by a predetermined positive value to ensure that the signal takes on values outside a desired range, in particular the quasi-linear region of the NLD 4212.
[0065] In some further preferred embodiments, the normalization applied to the current window may also depend, at least in part, on the rescaling parameters of the previous window. For example, an exponential moving average update rule may be used. In this way, the normalized intensity, i.e., the harmonic enhancement by the next NLD, does not change rapidly in one window compared to the previous window. In preferred embodiments, any of the above time windows may be several seconds long, for example, at least 2 seconds.
[0066] The nonlinear device 4212 may be configured, for example in a digital implementation, to perform a nonlinear function f(x) (preferably instantaneous) that takes as input samples of the signal x[k], e.g., the output of the normalizer 4211, and outputs processed samples y[k], where y[k]=f(x[k]), where k is a time index. In a preferred embodiment, when the nonlinear function is instantaneous, each sample in x[k] is processed independently of the others for all k.
[0067] Various formulations for the nonlinear function f(x) can be implemented, and several are known in the art. In addition, NLD functions with continuous first and second derivatives less than 1 in the interval (0,1) have been found to perform better for devices characterized by low cutoff frequencies, such as electrodynamic loudspeakers.
[0068] A particularly advantageous example of an NLD function f(x) is tanh(kx), where k is a predetermined value, preferably equal to and / or greater than 1.
[0069] The above formulations have also been found to be particularly effective when applied to small speakers. At the same time, they have also been found to be highly suitable for piezoelectric transducers. In particular, certain implementations using tanh have been found to lead to stronger bass enhancement and are preferred for piezoelectric transducers. This is because piezoelectric transducers tend to have higher cutoff frequencies. By employing an NLD function f(x) characterized by double-sided saturation behavior, more harmonics can be generated, improving the VBE effect compared to electrodynamic speakers. Herein, the term "double-sided" may be understood to refer to the saturation of both the positive and negative half-waves. The use of tanh(x) has been found to provide a particularly favorable tradeoff between perceived bass enhancement and audible distortion when high-amplitude signals are fed. This is likely due to tanh(x)'s resemblance to typical saturation units, such as symmetrical diode clipping and / or overdrive units commonly found in music processing.
[0070] As a further example of f(x), the formulation according to equation (3) below has been found to be particularly effective.
number
[0071] This implementation has been found to be particularly advantageous because it avoids extremely unequal weighting of the positive and negative half waveforms. In general, unequal weighting of the positive and negative half waveforms is not, in itself, negative. In fact, as far as VBE is concerned, asymmetric functions are often preferable. However, various conventional NLD functions have very unequal effects on the positive and negative half waveforms, with one gaining more of an effect than the other. In contrast, one advantage of the NLD function described above is that it does not disproportionately amplify one half waveform relative to the other.
[0072] A further advantage is that the NLD function is asymmetric, generating both even and odd harmonics. In fact, the missing fundamental phenomenon is more effectively induced when both even and odd harmonics are present.
[0073] The amplifier 4213 may be configured to multiply its input, e.g. the output of the NLD 4212, by a predetermined gain value to adjust the amplitude, thereby controlling the level of the processed audio component.
[0074] In some preferred embodiments, the gain value may be a function of a normalization parameter in the normalizer 4211. In a preferred implementation, the gain value may be configured to reduce the level of signals that have undergone more pronounced harmonic generation compared to signals in previous and / or subsequent time windows.
[0075] It is therefore clear from the above that the present invention allows multiple acoustic sources to be extracted from the input signal IN independently of one another, with the result that these acoustic sources are processed separately and therefore differently from one another, thereby providing a greater degree of modularity, granularity and control than in the prior art.
[0076] Such granularity can be applied anywhere in the signal processing chain. For example, the ideal filter units 2410-241N can be configured appropriately depending on the spectral content of a given channel 1110-111N. Alternatively or additionally, the normalization can be configured to operate independently on various signals that may differ from each other, allowing the normalizer unit 4211 to operate differently for different acoustic sources.
[0077] Furthermore, depending on the music track under consideration, the division of the acoustic source into independent channels allows the nonlinear device 4212 to apply the functionality of a given NLD function to a given channel and the functionality of a different NLD function to another channel. In particular, the NLD characteristics can be selected with respect to the transducer for which the VBE system is designed. More specifically, the NLD characteristics can be selected from among the nonlinearities that are best suited to the acoustic source in a given channel, for example, with respect to the number, type, amplitude, and energy of introduced harmonics.
[0078] Additionally, the amplifier section 4213 may also operate differently for different channels, allowing the present invention to adjust and tune the bass enhancement on a per channel basis.
[0079] For example, in a preferred embodiment, the drum channel may be applied with a low pass filter, and another filter section 241N may be configured to have a unitary flat transfer function. This is also preferred for audio sources typically associated with the perception of low frequencies, such as the bass channel, which the inventors have found to be advantageously processed in their entirety in some embodiments, instead of focusing on sub-bands of the spectrum.
[0080] Thus, the separation of the input signal IN into multiple channels allows the designer of the virtual bass enhancement device to configure a VBE algorithm tailored to the perceptual and tonal characteristics of each sound source independently of the others, which not only has the benefit of reducing IMD, but also allows for greater flexibility and granularity in selecting the particular VBE algorithm in the virtual bass enhancement units 1210-121N or 4210 that is best suited to each sound source.
[0081] 5 shows a schematic diagram of a virtual bass enhancement device 5000 that differs from any of the virtual bass enhancement devices 1000, 2000, 3000 by further comprising post-processing elements 5610, 5620, 5630 for the enhanced audio signal OUT. It will be apparent that although described together in a single embodiment, any of these elements may be implemented separately from one another and / or any combination of these elements may be implemented.
[0082] In particular, the virtual bass enhancement device 5000, as shown, includes any of the elements of the virtual bass enhancement devices described above for generating an enhanced audio signal OUT, and a high-pass filter 5610 that takes the enhanced audio signal OUT as an input and outputs a filtered enhanced audio signal.
[0083] In a preferred embodiment, the virtual bass enhancement device is configured for use with a transducer having a cutoff frequency. The high-pass filter 5610 is configured to have a cutoff frequency, preferably equal to the cutoff frequency of the transducer. This is particularly advantageous because the frequencies removed by the high-pass filter 5610 cannot be properly reproduced by the speaker, at least at an adequate level and / or without significant distortion. Additionally, removing these frequency components is advantageous because they do not affect subsequent normalization. The cutoff frequency of a transducer, especially a small transducer, can be understood as the frequency below which the transducer cannot operate normally. This can be understood as the minimum output frequency.
[0084] The virtual bass enhancement device 5000 further includes a peak normalizer 5620 and / or a loudness normalizer 5630 that operate on the filtered enhanced audio signal, as shown. These elements may employ any known algorithms for peak and loudness normalization.
[0085] Any of these normalizations, especially combinations of them, avoid clipping and / or keep the perceived loudness of the original track, and also avoid sudden energy spikes that could damage speakers.
[0086] In particular, one of the purposes of VBE is to introduce newly generated harmonics into the signal, providing the listener with the sensation of unplayed low-frequency frequencies. This results in an increase in the signal's energy. When driving loudspeakers, especially small ones, high-energy signals can stress the transducer's mechanical components, eventually causing damage or failure. Damage can also be caused by the high distortion produced by a simple boost of low-frequency amplitude, such as that obtained by additive equalization. This becomes even more problematic when the energy increase occurs in bursts.
[0087] An example would be a kick drum in a music track. A kick drum hit contains a lot of low end and generates a corresponding amount of harmonic content. Including these harmonics in the reinforced signal increases the energy every time the kick drum is played, causing sudden mechanical surges that can damage speakers.
[0088] Applying peak and loudness normalization addresses this issue and may be implemented as an embodiment of the present invention, however, if no filtering of the signal is performed, the overall volume may decrease because the normalization process takes into account energy in frequency ranges that the transducer cannot reproduce.
[0089] Therefore, in a preferred embodiment, energy in frequency bands that the transducer cannot reproduce is advantageously removed by high-pass filter 5610. The resulting signal is then normalized by peak and / or loudness normalization. This prevents the transducer from being driven by an excessively high amplitude signal. At the same time, because the signal is filtered by high-pass filter 5610, the normalization in peak normalization unit 5620 and / or loudness normalization unit 5630 is not affected by those frequency components. This allows the present invention to achieve a larger signal than would be possible without high-pass filter 5610.
[0090] In a further preferred embodiment, the normalization in the peak normalization unit 5620 and / or the loudness normalization unit 5630 may be configured such that the sound pressure level (SPL) of the normalized audio signal is the same as the SPL of the input audio signal IN.
[0091] In the above embodiments, it can be understood that the virtual bass enhancement device can be implemented by physical components and / or software. Thus, in a purely software implementation, the virtual bass enhancement device may include a processing unit (processor) and a storage unit (memory), and the storage unit may include instructions that cause the processing unit to perform any of the above-mentioned components or elements.
[0092] Thus, various embodiments have been described in terms of how they can be implemented to provide improved virtual bass enhancement. While specific features have been described in each of the various embodiments, it will be apparent to those skilled in the art that one or more features of any embodiment may be combined with one or more features of any other embodiment, and in particular may be combined separately from the remaining features of each embodiment. [Explanation of symbols]
[0093] 1000 Virtual Bass Enhancer 1100 Separation section 1110~111N audio channels 1210~121N Virtual Bass Enhancement Unit 1310~131N Adder IN Input audio signal OUT Reinforced audio signal 2000 Virtual Bass Enhancer 2410~241N Filter section 3000 Virtual Bass Enhancer 3510~351N Subtractor 4210 Virtual Bass Enhancement Unit 4211 Normalization section 4212 Nonlinear Devices 4213 Amplifier 5000 Virtual Bass Enhancer 5610 High-pass filter 5620 Peak normalization section 5630 Loudness normalization unit
Claims
1. A virtual bass enhancement device (1000, 2000, 3000, 5000) for enhancing virtual bass in an input audio signal (IN), comprising: a separation unit (1100) configured to extract from the input audio signal (IN) at least one audio channel (1110-111N) corresponding to an acoustic source or a group of acoustic sources in the input audio signal (IN); at least one virtual bass enhancement unit (1210-121N, 4210) configured to generate harmonics for enhanced bass perception in said audio channels (1110-111N); at least one adder (1310-131N) configured to add the harmonics to the input audio signal (IN) to generate an enhanced audio signal (OUT); and at least one filter unit (2410-241N) configured to filter at least one of the audio channels (1110-111N) and output the at least one filtered audio channel (1110-111N); at least one of the virtual bass enhancement units (1210-121N, 4210) is configured to generate the harmonics for enhancing the bass perception in the filtered audio channels (1110-111N); The virtual bass enhancement device (1000, 2000, 3000) further comprises at least one subtractor (3510-351N) configured to subtract at least one filtered audio channel from the input audio signal (IN).
2. The virtual bass enhancement device (2000, 3000, 5000) of claim 1, wherein the virtual bass enhancement unit (1210-121N, 4210) is a time-domain virtual bass enhancement unit (1210-121N, 4210).
3. The virtual bass enhancement device (2000, 3000, 5000) of claim 1, wherein at least one of the filter sections (2410-241N) is a linear-phase digital filter or a zero-phase digital filter.
4. 2. The virtual bass enhancement device (1000, 2000, 3000, 5000) of claim 1, wherein the at least one virtual bass enhancement unit (4210) includes a normalization unit (4211), a non-linear device (4212), and an amplification unit (4213).
5. a high-pass filter (5610) that receives as input the enhanced audio signal (OUT) and outputs a filtered enhanced audio signal; 10. The virtual bass enhancement device (5000) of claim 1, further comprising a peak normalizer (5620) and a loudness normalizer (5630) that operate on the filtered enhanced audio signal.
6. configured for use with a transducer having a cutoff frequency; 6. The virtual bass enhancement device (5000) of claim 5, wherein the high-pass filter (5610) has a cutoff frequency corresponding to a cutoff frequency of the transducer.
7. 10. The virtual bass enhancement device (1000, 2000, 3000, 5000) of claim 1, wherein the sound source comprises one of a drum, a vocal, or a musical instrument.
Citation Information
Patent Citations
Audio data processing method and electronic equipment
CN114299976A
Audio signal processing device, control method and program for audio signal processing device
JP2015195432A
Multi-channel decomposition and harmonic synthesis
WO2021154211A1
Signal processing device, signal processing method, and signal processing program
WO2021161543A1