A bass enhancement system based on bass delay and virtual bass

By generating and combining bass harmonics with a delayed version of the input bass, the method enhances the bass effect in small speakers, addressing the limitations of existing virtual bass algorithms.

WO2026155967A1PCT designated stage Publication Date: 2026-07-23DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2026-01-12
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing virtual bass algorithms fail to provide a satisfying bass effect, particularly in small speakers, as they do not effectively enhance the 'punch' of bass frequencies.

Method used

A method involving generating bass harmonics and a delayed version of the input bass portion, combined with the original bass portion to enhance the bass effect, using techniques like saturation, distortion, and waveshaping, and controlling the delay through a smoothing factor to ensure precise enhancement.

Benefits of technology

The method enhances the perceived bass effect by adding higher frequencies and extending the duration of bass impulses, providing a stronger and more impactful sound experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2026010971_23072026_PF_FP_ABST
    Figure US2026010971_23072026_PF_FP_ABST
Patent Text Reader

Abstract

The disclosure relates to a method of bass processing for an input audio signal, the method comprising: generating bass harmonics based on an input bass portion included in the input audio signal; obtaining an additional bass portion including a delayed version of at least part of the input bass portion; combining the additional bass portion with the input bass portion; and mixing said combination into the generated bass harmonics.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] D25001W001

[0002] A BASS ENHANCEMENT SYSTEM BASED ON BASS DELAY AND VIRTUAL BASS

[0003] Cross-Reference to Related Applications

[0004] This application claims the priority benefit of International Patent Application PCT / CN2025 / 072288 filed January 14, 2025, the contents of which are hereby incorporated by reference in its’ entirety.

[0005] Technical Field

[0006] The present disclosure relates to techniques for audio processing, in particular bass processing for an audio signal such as techniques for providing bass enhancement for the audio signal.

[0007] Background

[0008] Virtual bass algorithm is widely used, especially for miniature and small speakers that are incapable of playing back bass. With the development of electroacoustic technology, the low-frequency responses of many small size speakers on phones and tablets do not attenuate so rapidly, making playing bass back possible. However, only applying virtual bass stills fails to provide satisfying bass effect to a listener, for example fails to show the “punch” effect of the bass, such as a big drum.

[0009] It is therefore an object of the present disclosure to enhance the bass or bass effect for an audio signal to be played back to a listener.

[0010] Summary

[0011] In view of the above need, the present disclosure provides methods, as well as corresponding apparatus, computer programs, and computer-readable storage media, having the features of respective independent claims.

[0012] One aspect of the present disclosure relates to a method of bass processing for an input audio signal. The method may comprise a step of generating bass harmonics based on an input bass portion included in the input audio signal. The method may further comprise a step of obtaining an additional bass portion including a delayed version of at least part of the inputD25001W001

[0013] bass portion. The method may further comprise a step of combining the additional bass portion with the input bass portion. The method may further comprise a step of mixing said combination into the generated bass harmonics.

[0014] Bass, as well understood by a person skilled in the art, refers generally to the low-frequency part of an audio signal. For example, in the frequency domain a bass portion of an audio signal may have a frequency range of 20 to 250 Hz. As another example, a bass portion of an audio signal may have a frequency range of 100 to 200 Hz. Generally, a bass portion of an audio signal includes slower oscillations in the waveform of the audio signal and is perceived by a listener as deep, low-pitched, weighty sound, for example like kick drum thump, or low synth tones and etc. A bass portion of an audio signal may be observed on the audio signal’s waveform or spectrum.

[0015] In the present disclosure, for an input audio signal, the proposed method provides processing of an input bass portion included in the input audio signal. The processed bass includes at least one of the following parts: the input bass portion, bass harmonics generated based on the input bass portion, and an additional bass portion including a delayed version of at least part of the input bass portion.

[0016] Therein, the generated bass harmonics may add one or more frequency components to the original frequency or frequencies or frequency range of the input bass portion, so as to enable the listener to better perceive the bass effect that would originally not be well heard for example due to that the original input bass portion is of a rather low frequency. For example, the generated bass harmonics may add integer multiples of the original frequency or frequencies or frequency range of the input bass portion, providing higher frequencies for the listener to perceive. Different manners of implementing the generation of bass harmonics may be provided, including for example saturation / soft clipping, distortion / clipping, waveshaping, subharmonic synth and etc. For another example, bass harmonics may be generated based on virtual bass techniques. For capability-limited devices, attenuation of the original bass frequency or frequencies or frequency range may be provided to protect the loudspeaker. The delayed version of the at least part of the input bass portion provides further bass enhancement. As the same part of the input bass portion may be played again within a pre-set period of time of delay after the original input bass portion, thus providing an additional bass portion which includes this delayed version of the original input bass portion, the bass effect or the time length of the bass or total energy of the processed bass may be enhanced andD25001W001

[0017] extended in comparison to the original input bass portion. The listener thus perceives a stronger bass effect.

[0018] Therefore, by for example combining at least two of the input bass portion, the generated bass harmonics, and the additional bass portion, the listener is provided with enhanced bass effect. The combination may be provided in the frequency domain or in the time domain. In some examples, whether to delay said at least part of the input bass portion may be based on a smoothing factor. The method may further comprise a step of setting the smoothing factor to a predefined start value when detecting said at least part of the input bass portion to be transient. The method may further comprise a step of updating the smoothing factor by multiplying its current value with a decay coefficient having a value smaller than 1 when detecting said at least part of the input bass portion to be non-transient.

[0019] An input bass portion may start with a transient bass usually carrying a burst of energy. The duration of a transient bass at the beginning of an input bass portion may vary depending on the specific input audio signal. In the present disclosure, for the beginning of the input bass portion or for a transient bass detected in the input bass portion, the smoothing factor may be set to the predefined start value. This smoothing factor may then be reduced by multiplying with the decay coefficient for non-transient bass included in the input bass portion, wherein non-transient bass generally includes steady bass usually of a longer period of time compared with the transient bass. The smoothing factor having an updatable value thereby provides an indicator for determining when or whether to delay at least part of the input bass portion. In some examples, the method may comprise a step of determining to delay said at least part of the input bass portion when the updated smoothing factor is greater than a smoothing threshold value.

[0020] As the smoothing factor may get updated or reduced in value when detecting non-transient bass that follows the transient bass, the smaller the value of the smoothing factor is, the more distant in time the non-transient bass to the transient bass is. Therefore, non-transient bass that is further away from the beginning of the input bass portion may not need to be delayed, whereas non-transient bass that is closer to the beginning of the input bass portion may be delayed to be added to the original input bass portion. Avoiding delaying distant non-transient bass ensures that only the more relevant part of the input bass portion is enhanced. In some examples, the method may comprise a step of repeatedly performing said determination for consecutive parts of the input bass portion until the updated smoothingD25001W001

[0021] factor becomes smaller than the smoothing threshold value. The method may comprise a step of outputting directly remaining parts included in the input bass portion once the updated smoothing factor is determined to be smaller than the smoothing threshold value.

[0022] By repeatedly performing said determination, and by setting a smoothing threshold value for regulating when to stop the determination or the delaying of the bass, the bass enhancement can be flexibly and precisely controlled. In particular, said predefined start value and / or said decay coefficient and / or said smoothing threshold value may be configurable, so that the configuration of the duration of to-be-delayed or enhanced bass part can be flexibly and precisely controlled. As a result, it can be effectively controlled which part of the input bass portion is to be delayed and combined for the bass enhancement.

[0023] In some examples, the method may further comprise a step of storing the additional bass portion in a buffer. Further, the method may comprise a step of aligning the buffered additional bass portion with the input bass portion. The method may additionally comprise a step of combining the buffered additional bass portion after alignment with the input bass portion.

[0024] In this implementation manner, for the combination of the delayed version of at least part of the input bass portion and the original at least part of the input bass portion, the delayed version or the additional bass portion may be stored in a buffer, wherein the buffer aligns itself with the original input bass portion, so that the delayed “later” components stored in the buffer are combined with the “earlier” components in the input bass portion.

[0025] In some examples, said combination may be based on the smoothing factor, a weighted combination of the additional bass portion, and the input bass portion, wherein a weighting factor for the additional bass portion has a value of the set or updated smoothing factor, and wherein a weighting factor for the input bass portion has a value complementary to the set or updated smoothing factor (e.g., 1 minus the set or updated smoothing factor).

[0026] Thereby, the additional bass portion and the input bass portion may be combined in a weighted manner. As the smoothing factor may be consecutively updated for consecutive parts of the input bass portion, the respective parts may be combined with the delayed version of the input bass portion, each part being provided with a corresponding updated smoothing factor. This provides dynamic bass enhancement. In particular, as the smoothing factor may be reduced for non-transient bass that is more distant from the transient bass, the weight in said combination for the delayed version may accordingly get reduced as well, whereas forD25001W001

[0027] non-transient bass that is closer to the transient bass, more weight is provided to the delayed version. As the bass gets weaker in the development of the input bass portion starting from the transient bass to the consecutive non-transient bass, the bass enhancement needs also to be reduced as the value of smoothing factors gets reduced, to the point where the delay is completely stopped and the remaining part or parts of the input bass portion or input audio signal is directly output without bass enhancement.

[0028] In some examples, the method may further comprise a step of determining a frame of the input bass portion to be transient when energy of all samples over all frequencies included in a frequency-domain representation of the frame is higher than an energy threshold value. As a transient bass usually carried a burst of energy, the presence of a transient bass may be detected by evaluating the energy of a part of the input bass portion, which may be performed frame-wise. In the present disclosure, a frame may be of a pre-set time length depending on the scenario, which is not limited.

[0029] In some examples, the method may further comprise a step of generating a high-frequency portion based on the input audio signal. The method may further comprise a step of mixing the additional bass portion, the input bass portion, and the generated bass harmonics with the generated high-frequency portion for providing an output signal. The enhanced bass may be added or mixed into the original high-frequency portion of the input audio signal, so as to provide an output signal to be played out to the listener.

[0030] In some examples, a time length of delay may be configurable, wherein the time length of delay is preferably between 20ms and 60ms and more preferably 40ms.

[0031] In some examples, the method may be performed in either time domain or frequency domain. In some examples, the predefined start value may be 1.

[0032] In some examples, the decay coefficient may be between 0.9 and 0.98, and preferably 0.94. In some examples, the smoothing threshold value may be between 0.0095 and 0.0015, preferably 0.001, times the predefined start value.

[0033] Another aspect of the present disclosure relates to an apparatus, comprising a processor and a memory coupled to the processor, wherein the processor may be adapted to carry out the method according to the above aspect and its related embodiments.

[0034] According to another aspect, an apparatus is provided. The apparatus may include one or more processors and a memory coupled thereto and storing instructions for the one or moreD25001W001

[0035] processors. The one or more processors may be configured to perform the methods or method steps outlined throughout the present disclosure. This apparatus may relate to an encoder, encoding apparatus, or encoding system, to a transcoder, transcoding apparatus, or transcoding system, or to a decoder, decoding apparatus, or decoding system, as the case may be. The encoder, encoding apparatus, or encoding system may relate to or be part of a wearable device. For example, the encoder may be part of a streaming service.

[0036] According to a further aspect, a computer program is described. The computer program may comprise executable instructions for performing the methods or method steps outlined throughout the present disclosure when executed by a computing device (e.g., one or more processors).

[0037] According to another aspect, a computer-readable storage medium is described. The storage medium may store a computer program adapted for execution on a computing device (e.g., one or more processors) and for performing the methods or method steps outlined throughout the present disclosure when carried out on the computing device.

[0038] It should be noted that the methods and apparatus including its preferred embodiments as outlined in the present disclosure may be used stand-alone or in combination with the other methods and apparatus disclosed in this document. Furthermore, all aspects of the methods and apparatus outlined in the present disclosure may be arbitrarily combined. In particular, the features of the claims may be combined with one another in an arbitrary manner.

[0039] It will be appreciated that apparatus features and method steps may be interchanged in many ways. In particular, the details of the disclosed method(s) can be realized by the corresponding apparatus, and vice versa, as the skilled person will appreciate. Moreover, any of the above statements made with respect to the method(s) (and, e.g., their steps) are understood to likewise apply to the corresponding apparatus (and, e.g., their blocks, stages, units), and vice versa.

[0040] Brief Description of the Drawings

[0041] Example embodiments of the disclosure are explained below with reference to the accompanying drawings, wherein

[0042] Fig. 1 schematically illustrates an example of a method of bass processing for an input audio signal according to embodiments of the present disclosure;D25001W001

[0043] Fig. 2 schematically illustrates an example of a bass enhancement system according to embodiments of the present disclosure;

[0044] Fig. 3 schematically illustrates an example of a method of bass delay according to embodiments of the present disclosure;

[0045] Fig. 4 schematically illustrates an example of a waveform of a transient bass according to embodiments of the present disclosure;

[0046] Fig. 5 schematically illustrates an example of a delayed and current frame mixing process according to embodiments of the present disclosure;

[0047] Fig. 6 schematically illustrates an example of a comparison of an original transient bass and a transient bass processed by the proposed bass enhancement method or system according to embodiments of the present disclosure;

[0048] Fig. 7 schematically illustrates an example of an apparatus for performing the methods or method steps outlined throughout the present disclosure according to embodiments of the present disclosure; and

[0049] Fig. 8 schematically illustrates a schematic block diagram of an example electronic device or architecture suitable for implementing example embodiments of the present disclosure.

[0050] Detailed Description

[0051] The present disclosure provides a bass enhancement system and method for providing bass delay, in particular by utilizing temporal summation to enhance loudness perception of the bass of an audio signal. For example, the tone impulse of 100ms duration produces a loudness which is about twice as large as the loudness of the 10ms impulse. Thus, the present disclosure proposes that the bass impulse may be played back just behind its harmonics, leading to extension of the effective bass duration, and a more “punching” drum can be perceived. In some examples, by specifically controlling the time between the bass harmonics and the bass to be less than a precedence effect threshold, echo artifacts can be avoided. Fig. 1 schematically illustrates an example of a method of bass processing for an input audio signal according to embodiments of the present disclosure.

[0052] The method comprises a step SI 10 of generating bass harmonics based on an input bass portion included in the input audio signal.D25001W001

[0053] The method further comprises a step SI 20 of obtaining an additional bass portion including a delayed version of at least part of the input bass portion.

[0054] The method further comprises a step S130 of combining the additional bass portion with the input bass portion.

[0055] The method further comprises a step SI 40 of mixing said combination into the generated bass harmonics.

[0056] Explanations concerning the aforementioned method steps have been provided above, which is not repeated here.

[0057] Fig. 2 schematically illustrates an example of a diagram of a bass enhancement system according to embodiments of the present disclosure. The system may be configured to perform the method according to Fig. 1. In particular, Fig. 2 provides an example wherein the method of Fig. 1 may be performed frame-wise.

[0058] An input audio may include one or more frames of audio data. The length of a frame may be configured depending on the scenario, which is not limited in the present disclosure. An input frame in the example of Fig. 2 is converted to a frequency representation. Bass harmonics of the input frame are generated in a virtual bass module. Of course, other technique for generation of bass harmonics may be applied instead. Meanwhile, high-frequency part of the input frame is mixed with the generated bass harmonics. Bass part of the input frame is also mixed with the high-frequency part and the generated bass harmonics after being processed by the “Delay and Mixing Logic Control” and the “Delay Frame and Frame Mixing” blocks in Fig- 2. The delay time may be configurable.

[0059] The boundary of bass in the present disclosure may be 100 to 200Hz. For example, at least one of the following may have a frequency range of 100 to 200Hz: the original bass part of the input frame, the generated bass harmonics of the input frame, and the delayed bass part. An example of a bass delay method is shown in Fig. 3, which may be performed by the “Delay and Mixing Logic Control” and the “Delay Frame and Frame Mixing” blocks shown in Fig. 2.

[0060] Transient Detection

[0061] In the bass delay method in Fig. 3, when receiving a bass frame, a transient detection may be applied to decide whether this frame is a transient bass. A transient bass usually refers to a drum. An example waveform of a transient bass is shown in Fig. 4.D25001W001

[0062] The bass frame may be denoted as Xm(n, k), where m denotes the frame index, n denotes the sample index in this frame and k denotes the frequency index. The short-term energy of a frame is:

[0063]

[0064] where N denotes the sample number and K denotes the frequency number.

[0065] In order for determining whether a bass frame is a transient bass, the long-term energy may be calculated from the short-term energy:

[0066] EL(m) = (1 — a)EL(m — 1) + aE(m) (2) and compared with the short-term energy. In equation (2), the parameter a determines the speed of response to signal changes, and may be adjusted for different use cases.

[0067] When the short-term energy exceeds an energy threshold value, the bass frame is determined as a transient frame:

[0068] E(m) > pEL(m) (3) For example, 14dB may be applied as the energy threshold, and the corresponding p is calculated by:

[0069] 20log10(j)) = 14 dB (4) so that p is 5.

[0070] The transient detection may be implemented on all kinds of time-frequency transforms, including but not limited to Fourier transform and Quadrature Mirror Filter (QMF) bank. When the algorithm is based on QMF bank, transient detection may be performed in the lowest filter bank. In this case the bass frame may be denoted as Xm(n), and K equals 1 in equation (1).

[0071] Using short-term energy and long-term energy provides an example a real-time detection of transient bass. Of course other techniques for transient bass or frame detection may be applied instead.D25001W001

[0072] Delay and mixing logic control

[0073] In the bass delay method in Fig. 3, a smoothing factor smooth_fact may be set or initialized to a predefined start value, for example when a transient bass is detected. Alternatively, the smoothing factor smooth_fact may be initialized to for the first frame of an input bass portion, for example when an input bass frame is the first frame in the input bass portion. The predefined start value may for example be 1.

[0074] After initialization, the smoothing factor smooth_fact may be updated by being multiplied with a decay coefficient. In an example, the updating may be performed when the bass frame is not transient. The updating of smooth_fact may be provided as follows:

[0075] smooth_fact = smooth_fact * decay_coef (5) The set or updated smooth_fact is then used in the next step. In the above equation, the decay_coef is used for gradually attenuating smooth_fact. A value of decay_coef may be less than 1. When choosing a value of the decay _coef, the frame length and the bass interval may be considered. For example, decay _coef may have a value of 0.94.

[0076]

[0077] In the bass delay method in Fig. 3, whether to delay a bass frame may be regulated by the smooth_fact. For example, when smooth_fact is less than a smooth threshold value smooth_thre_value, it means that the current frame is far away from a transient bass and unnecessary to be delayed. Therefore, this current frame will be output directly. A value of smooth_thre_value is preferred to be set to a small number, for example 0.001.

[0078] When smooth_f act is larger than smooth_thre_value, it means that the current frame is a transient frame or closely behind a transient frame. In a case where the onset of a drum is detected as a transient, the whole drum (including but not limited to said onset of the drum) may need to be delayed until the whole drum ends. Therefore, as long as smooth_fact is large thansmooth_thre_value, the current frame is determined still as a part of a drum, and thus need to be delayed. The delay time may be defined by the user. Preferably the delay time may not exceed 40ms to guarantee the precedence effect. In other words, if the delay time is too long, the drum will be discontinuous. The delay time may be configurable or defined by the listener.

[0079] The delayed frame may be stored in a delayed-frame buffer. To smoothly transit from a “delay state” to a “no-delay state”, the output bass frame Xoutn) may be provided based onD25001W001

[0080] a mixture of the current frame Xcur(n), and the time-aligned frame XbUf(n) stored in the delayed-frame buffer, and the mixing ratio may be controlled by the smooth_fact, as follows:

[0081]

[0082] Fig. 5 schematically illustrates an example of a delayed and current frame mixing process according to embodiments of the present disclosure, using the example of delaying the transient bass by 1 frame.

[0083] The delayed-frame buffer is initialized to be zero. When a frame is detected as a transient bass, it may be delayed 1 frame and stored in the delayed-frame buffer, as the Fig. 5 shows. The delayed-frame buffer aligns in time with the input bass frame. For example, in Fig. 5, the third input bass frame Xcur(3) is delay 1 frame, which the delayed frame is stored in the buffer as Xftu^(4). This delayed frame XbUf(4~) in the buffer then is added with the fourth input bass frame Xcur(4) to provide the output bass frame Xout(4), all aligning in time with the fourth output bass frame.

[0084] In the example of Fig. 5, the second output frame is almost zero, as there is nothing in the delayed-frame buffer currently. The output of the next frame is dominantly the transient bass frame. As the value of smooth_fact decreases, the input frames are continued to be delayed; more signal in the current frame will be mixed into the output while less signal in the delayed frame will be mixed. When smooth_fact is near smooth_thre_value, almost all the output is the signal in the current frame. No artifact will happen when smooth_fact is less than smooth_thre_value as the output signal will all be the current frame. The above process implements a perfect transition from a “delay state” to a “no-delay state”.

[0085] The bass delay method may be implemented by a time-domain non-linear phase filter. It has a positive group delay around the bass frequency, and the group delay at the other frequencies are 0. This can also achieve the effect that only bass is delayed.

[0086] Fig. 6 schematically illustrates an example of a comparison of an original transient bass and a transient bass processed by the proposed bass enhancement method or system according to embodiments of the present disclosure. After delaying the transient bass and mixing with the generated harmonics, the transient duration time extends. Listeners will perceive a louder drum.D25001W001

[0087] Fig. 7 schematically illustrates an example of an apparatus for performing the methods or method steps outlined throughout the present disclosure according to embodiments of the present disclosure.

[0088] The apparatus 700 comprises a processor 710 (or multiple processors) and a memory 720 coupled to the processor. The memory 720 may store instructions for execution by the processor 710. Processor 710 may be adapted to implement the apparatus described throughout the disclosure and / or to perform methods (e.g., methods of encoding or decoding, methods of generating or modifying bitstreams) described throughout the disclosure. The apparatus may receive inputs 730 (e.g., audio signals, bitstreams, etc.) and generate outputs 740 (e.g., audio signals, bitstreams, etc.) as described throughout the disclosure. Accordingly, the apparatus may relate to any of an encoding apparatus or a decoding apparatus, as the case may be.

[0089] The present disclosure further relates to programs (e.g., computer programs) comprising instructions that, when executed by a processor (or multiple processors), cause the processor (or multiple processors) to carry out any of the methods described throughout the disclosure, and to computer-readable storage media storing such programs.

[0090] Fig- 8 shows a schematic block diagram of an example electronic device or architecture 200 (e.g., an apparatus 200) suitable for implementing example embodiments of the present disclosure. Architecture 200 includes but is not limited to servers and client devices, systems, modules and methods as described in reference to Figs. 1-6. As shown, the architecture 200 includes central processing unit (CPU) 201 which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 202 or a program loaded from, for example, storage unit 208 to random access memory (RAM) 203. The CPU 201 may be, for example, an electronic processor 201, which may include one or more processor cores, and in some examples the processor 201 may be multiple processors. In RAM 203, the data used when CPU 201 performs the various processes is also stored, as required. CPU 201, ROM 202 and RAM 203 are connected to one another via bus 204. Input / output (I / O) interface 205 is also connected to bus 204.

[0091] The following components are connected to I / O interface 205: input unit 206, that may include a keyboard, a mouse, or the like; output unit 207 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 208 including a hardD25001W001

[0092] disk, or another suitable storage device; and communication unit 209 which may include a network interface card such as a network card (e.g., wired or wireless).

[0093] In some implementations, input unit 206 includes one or more microphones in different positions (depending on the host device) enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).

[0094] In some implementations, output unit 207 include systems with various number of speakers. Output unit 207 (depending on the capabilities of the host device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).

[0095] In some embodiments, communication unit 209 is configured to communicate with other devices (e.g., via a network). Drive 210 is also connected to I / O interface 205, as required. Removable medium 211, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 210, so that a computer program read therefrom is installed into storage unit 208, as required. A person skilled in the art would understand that although apparatus 200 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alteration all fall within the scope of the present disclosure.

[0096] In accordance with example embodiments of the present disclosure, the processes described above may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program including program code for performing methods. In such embodiments, the computer program may be downloaded and mounted from the network via the communication unit 209, and / or installed from the removable medium 211, as shown in FIG. 7.

[0097] Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic or any combination thereof. For example, the units discussed above can be executed by control circuitry (e.g., CPU 201 in combination with other components of FIG. 7), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, a processor and / or other computing device(s), whichD25001W001

[0098] may include control circuitry. While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques, or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.

[0099] Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.

[0100] In the context of the disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may be non-transitory and may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random-access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0101] Computer program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to one or more processors of a general-purpose computer, special purpose computer, or other programmable data processing apparatus that has control circuitry, such that the program codes, when executed by one or more processors of the computer or other programmable data processing apparatus, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a computer, partly on the computer, as a stand-alone software package, partly on theD25001W001

[0102] computer and partly on a remote computer or entirely on the remote computer or server or distributed over one or more remote computers and / or servers.

[0103] Enumerated Example Embodiments

[0104] Various Aspects and implementations of the invention may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims.

[0105] EEE Al. A method of bass processing for an input audio signal, the method comprising: generating bass harmonics based on an input bass portion included in the input audio signal;

[0106] obtaining an additional bass portion including a delayed version of at least part of the input bass portion;

[0107] combining the additional bass portion with the input bass portion; and

[0108] mixing said combination into the generated bass harmonics.

[0109] EEE A2. The method according to EEE Al, wherein whether to delay said at least part of the input bass portion is based on a smoothing factor; and

[0110] wherein the method further comprises:

[0111] when detecting said at least part of the input bass portion to be transient, setting the smoothing factor to a predefined start value; and

[0112] when detecting said at least part of the input bass portion to be non-transient, updating the smoothing factor by multiplying its current value with a decay coefficient having a value smaller than 1.

[0113] EEE A3. The method according to EEE A2, further comprising:

[0114] determining to delay said at least part of the input bass portion when the updated smoothing factor is greater than a smoothing threshold value.

[0115] EEE A4. The method according to EEE A3, further comprising:

[0116] repeatedly performing said determination for consecutive parts of the input bass portion until the updated smoothing factor becomes smaller than the smoothing threshold value; and

[0117] outputting directly remaining parts included in the input bass portion once the updated smoothing factor is determined to be smaller than the smoothing threshold value.

[0118] EEE A5. The method according to any one of EEEs Al to A4, further comprising:

[0119] storing the additional bass portion in a buffer;

[0120] aligning the buffered additional bass portion with the input bass portion; andD25001W001

[0121] combining the buffered additional bass portion after alignment with the input bass portion.

[0122] EEE A6. The method according to any one of EEEs A2 to A5, wherein said combination is based on the smoothing factor, a weighted combination of the additional bass portion, and the input bass portion, wherein a weighting factor for the additional bass portion has a value of the set or updated smoothing factor, and wherein a weighting factor for the input bass portion has a value complementary to the set or updated smoothing factor.

[0123] EEE A7. The method according to any one of EEEs A2 to A6, further comprising:

[0124] determining a frame of the input bass portion to be transient when energy of all samples over all frequencies included in a frequency-domain representation of the frame is higher than an energy threshold value.

[0125] EEE A8. The method according to any one of EEEs Al to A7, further comprising:

[0126] generating a high-frequency portion based on the input audio signal; and

[0127] mixing the additional bass portion, the input bass portion, and the generated bass harmonics with the generated high-frequency portion for providing an output signal.

[0128] EEE A9. The method according to any one of EEEs Al to A8, wherein a time length of delay is configurable, wherein the time length of delay is preferably between 20ms and 60ms and more preferably 40ms.

[0129] EEE A10. The method according to any one of EEEs Al to A9, wherein the method is performed in either time domain or frequency domain.

[0130] EEE All. The method according to any one of EEEs A2 to A10, wherein the predefined start value is 1.

[0131] EEE A12. The method according to any one of EEEs A2 to All, wherein the decay coefficient is between 0.9 and 0.98, and preferably 0.94.

[0132] EEE A13. The method according to any one of EEEs A3 to A12, wherein the smoothing threshold value is between 0.0095 and 0.0015, preferably 0.001, times the predefined start value.

[0133] EEE A14. An apparatus comprising a processor and a memory coupled to the processor, and storing instructions for the processor, wherein the processor is adapted to carry out the method according to any one of EEEs Al to A13.D25001W001 EEE A15. A program comprising instructions that, when executed by a processor, cause the processor to implement the method according to any one of EEEs Al to A13.

[0134] EEE A16. A computer-readable storage medium storing the program of EEE A13.

[0135] EEE Bl. A bass enhancement system, comprising:

[0136] implementing time-frequency transform of the input signal and frequency -time transform of the output signal;

[0137] generating bass harmonics;

[0138] delaying bass;

[0139] mixing the generated harmonics and the delayed bass into the output signal.

[0140] EEE B2. The method of delaying bass in EEE Bl, comprising:

[0141] detecting transient bass;

[0142] implementing a smooth transition from the “delay state” to the “no-delay” state.

[0143] EEE B3. The method of detecting transient bass in EEE Bl or B2, comprising:

[0144] calculating the long-term energy and the short-term energy of the signal; comparing the long-term and the short-term energy and deciding whether this frame is a transient.

[0145] EEE B4. The method of implementing a smoothing transition in EEE B2, comprising: controlling the delay logic by a smoothing factor;

[0146] controlling the mixing logic by a smoothing factor.

Claims

D25001W001CLAIMS1. A method of bass processing for an input audio signal, the method comprising: generating bass harmonics based on an input bass portion included in the input audio signal;obtaining an additional bass portion including a delayed version of at least part of the input bass portion;combining the additional bass portion with the input bass portion; andmixing said combination into the generated bass harmonics.

2. The method according to claim 1, wherein whether to delay said at least part of the input bass portion is based on a smoothing factor; andwherein the method further comprises:when detecting said at least part of the input bass portion to be transient, setting the smoothing factor to a predefined start value; andwhen detecting said at least part of the input bass portion to be non-transient, updating the smoothing factor by multiplying its current value with a decay coefficient having a value smaller than 1.

3. The method according to claim 2, further comprising:determining to delay said at least part of the input bass portion when the updated smoothing factor is greater than a smoothing threshold value.

4. The method according to claim 3, further comprising:repeatedly performing said determination for consecutive parts of the input bass portion until the updated smoothing factor becomes smaller than the smoothing threshold value; andoutputting directly remaining parts included in the input bass portion once the updated smoothing factor is determined to be smaller than the smoothing threshold value.

5. The method according to any one of claims 1 to 4, further comprising:storing the additional bass portion in a buffer;aligning the buffered additional bass portion with the input bass portion; andD25001W001combining the buffered additional bass portion after alignment with the input bass portion.

6. The method according to any one of claims 2 to 5, wherein said combination is based on the smoothing factor, a weighted combination of the additional bass portion, and the input bass portion, wherein a weighting factor for the additional bass portion has a value of the set or updated smoothing factor, and wherein a weighting factor for the input bass portion has a value complementary to the set or updated smoothing factor.

7. The method according to any one of claims 2 to 6, further comprising: determining a frame of the input bass portion to be transient when energy of all samples over all frequencies included in a frequency-domain representation of the frame is higher than an energy threshold value.

8. The method according to any one of claims 1 to 7, further comprising: generating a high-frequency portion based on the input audio signal; andmixing the additional bass portion, the input bass portion, and the generated bass harmonics with the generated high-frequency portion for providing an output signal.

9. The method according to any one of claims 1 to 8, wherein a time length of delay is configurable, wherein the time length of delay is preferably between 20ms and 60ms and more preferably 40ms.

10. The method according to any one of claims 1 to 9, wherein the method is performed in either time domain or frequency domain.

11. The method according to any one of claims 2 to 10, wherein the predefined start value is 1.

12. The method according to any one of claims 2 to 11, wherein the decay coefficient is between 0.9 and 0.98, and preferably 0.94.D25001W00113. The method according to any one of claims 3 to 12, wherein the smoothing threshold value is between 0.0095 and 0.0015, preferably 0.001, times the predefined start value.

14. An apparatus comprising a processor and a memory coupled to the processor, and storing instructions for the processor, wherein the processor is adapted to carry out the method according to any one of claims 1 to 13.

15. A program comprising instructions that, when executed by a processor, cause the processor to implement the method according to any one of claims 1 to 13.

16. A computer-readable storage medium storing the program of claim 13.