Transient bass signal energy optimization

The system optimizes transient bass signals by decomposing audio into frequency bands, analyzing their characteristics, and adjusting decay rates and energy, addressing the temporal shortcomings of existing methods to enhance bass signals effectively.

WO2025155516A1PCT designated stage expired Publication Date: 2025-07-24DTS INC(US)
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/011494
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-18
Filing Date
2025-01-14
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing bass enhancement methods fail to account for the complete temporal characteristics of bass signals, leading to suboptimal enhancement of transient bass signals, which are short and sharp bursts of low-frequency sounds.

Method used

A system that decomposes audio signals into multiple frequency bands using a time or frequency domain filter bank, analyzes transient bass signals, and adjusts their decay rates and energy using dynamic threshold compression techniques to optimize punch sounds.

Benefits of technology

The system effectively enhances transient bass signals by maintaining the naturalness of the original sound while increasing their impact, preserving the balance and characteristics of the audio signal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025011494_24072025_PF_FP_ABST
    Figure US2025011494_24072025_PF_FP_ABST
Patent Text Reader

Abstract

Generally disclosed herein is a system and method for optimizing a transient bass signal. The system may be configured to detect the bass signal sounds from the received input sound signal by decomposing the signal into multiple frequency bands based on creating a time domain or frequency domain filter bank. The system may then be configured to conduct analyses of the relevant features of the detected signals to determine the transient bass signal from the bass signal sounds. The system may be configured to calculate the parameters in real time to determine a degree of optimization for the transient bass signal. The system may be configured to compress the transient bass signal and adjust its decay rate. The system may be configured to perform the above processing for each of the multiple frequency bands. A whole enhanced sound signal can be synthesized using the processed frequency bands.
Need to check novelty before this filing date? Find Prior Art

Description

TRANSIENT BASS SIGNAL ENERGY OPTIMIZATIONBACKGROUND

[0001] A bass enhancement method may refer to a technique used in audio processing to artificially create or increase the level of low-frequency content of a sound signal. Example methods may involve manipulating the existing low-frequency spectrum to make the sound signal sound fuller and more impactful. Many bass enhancement methods typically focus on extending low-frequency ranges and amplifying the amplitude of the low-frequency regions of the reproduced audio signals. Some of these techniques use the “virtual bass” technique, involving the addition / generation of harmonics corresponding to the bass fundamental frequency to create the “missing fundamental” phenomenon. Some related methods may address the dynamic range control of the bass signal without accounting for the signal's complete temporal characteristics.BRIEF SUMMARY

[0002] Generally disclosed herein is a mechanism for enhancing and optimizing transient bass signals. The transient bass signal may refer to a short and sharp burst of low-frequency sounds that occur at the very beginning of a bass note or a percussive bass sound. The transient bass signal can be detected by decomposing the input audio signal into multiple frequency bands, achieved through the creation of a time domain or frequency domain filter bank, and by analyzing the features of each band parameters in realtime. By analyzing these features, the system can determine the desired energy optimization of the transient bass signal. The detected transient bass signal can be adjusted by changing attack time and / or magnitude of energy using a dynamic threshold compression technique according to the determined desired level of optimization of the transient bass signal. The technology can further optimize the sustain and decay times of the transient bass signal by dynamically adjusting the decay rates of each frequency band signal.

[0003] An aspect of the disclosure provides a system for optimizing a transient bass signal. The system may comprise memory and one or more processors configured to receive audio signals, divide the received audio signals into a plurality of frequency bands, detect a transient bass signal from each of the plurality of the frequency bands, determine the degree of optimization for the detected transient bass signal, compress the detected transient bass signal using a dynamic threshold compression technique, enhance the detected transient bass signal by adjusting a decay rate of the compressed transient bass signal, and synthesize the enhanced transient bass signal of each frequency band of the plurality of the frequency bands.

[0004] In some examples, the degree of the optimization is determined by a binary classification, wherein the binary classification is determined based on whether the detected transient bass signal is a punch bass sound or non-punch bass sound.

[0005] In some examples, the degree of the optimization is determined by using a classification of a punch bass sound included in the detected transient bass signal.

[0006] In some examples, the punch bass sound includes a hard punch, a medium punch or a light punch based on one or more parameters.

[0007] In some examples, the one or more parameters include at least one or more of peak power, decay time, attach time or sustain energy.

[0008] In some examples, the one or more processors are configured to adjust the one or more parameters.

[0009] In some examples, the one or more processors are further configured to control an onset time of the transient bass signal.

[0010] In some examples, the one or more processors are further configured to use stationary bass enhancement techniques in conjunction with the dynamic threshold compression technique.

[0011] In some examples, the stationary bass enhancement techniques include a bass equalization filtering.

[0012] In some examples, the one or more processors are further configured to classify the audio signals based on a genre of the received audio signals.

[0013] In some examples, the one or more processors are further configured to utilize a pre-stored bass profile information to enhance the transient bass signal.

[0014] In some examples, the one or more processors are further configured to utilize meta data of audio content.

[0015] In some examples, the meta data includes content categories related to movies, music, or gaming.

[0016] In some examples, the meta data related to the music includes genre information of the music, wherein the genre information includes classical, pop, rock, hip-hop, electronics, or jazz music.

[0017] In some examples, the meta data related to the gaming includes game genre information, wherein the game genre information includes first-person shooting games, action games, adventure games, rollplaying games or racing games.

[0018] In some examples, the one or more processors are further configured to decrease the energy of the transient bass signal by using the dynamic threshold compression technique and the adjustment of the decay rate.

[0019] In some examples, the one or more processors are further configured to separate the transient bass signal from a sound mix by using an audio object extraction or separation technique and channel separation technique used for muti-channel contents.

[0020] In some examples, the one or more processors are further configured to enable a user to control the transient bass signal for personal preferences.

[0021] Another aspect of the disclosure provides a method for optimizing a transient bass signal. The method may comprise: receiving, by one or more processors, audio signals; dividing, by the one or more processors, the received audio signals into a plurality of frequency bands; detecting, by the one or more processors, a transient bass signal from each of the plurality of the frequency bands; determining, by the one or more processors, the degree of optimization for the detected transient bass signal; compressing, by the one or more processors, the detected transient bass signal using a dynamic threshold compression technique; enhancing, by the one or more processors, the detected transient bass signal by adjusting a decayrate of the compressed transient bass signal; and synthesizing, by the one or more processors, the enhanced transient bass signal of each frequency band of the plurality of the frequency bands.

[0022] Yet another aspect of the disclosure provides a non-transitory machine-readable medium comprising machine-readable instruction encoded thereon for performing a method of optimizing a transient bass signal, the method comprising: receiving audio signals; dividing the received audio signals into a plurality of frequency bands; detecting a transient bass signal from each of the plurality of the frequency bands; determining the degree of optimization for the detected transient bass signal; compressing the detected transient bass signal using a dynamic threshold compression technique; enhancing the detected transient bass signal by adjusting a decay rate of the compressed transient bass signal; and synthesizing the enhanced transient bass signal of each frequency band of the plurality of the frequency bands.BRIEF DESCRIPTION OF THE DRAWINGS

[0023] FIGS. 1 A-C depict graphs illustrating energy of various transient bass signals according to aspects of the disclosure.

[0024] FIG. 2 depicts a graph illustrating a principal component analysis (PCA) on transient bass signals with various audio features according to aspects of the disclosure.

[0025] FIG. 3 depicts a block diagram illustrating an example transient bass signal optimization system according to aspects of the disclosure.

[0026] FIG. 4 depicts a flow diagram illustrating a transient bass signal optimization process according to aspects of the disclosure.

[0027] FIG. 5 depicts a graph illustrating peak envelops of a transient bass signal using dynamic threshold processing and static threshold processing according to aspects of the disclosure.

[0028] FIGS. 6A-6C depicts a graph illustrating a signal level and a power level of the original transient bass signal and the optimized transient bass signal according to aspects of the disclosure.

[0029] FIG. 7 depicts a flow diagram illustrating an example method for optimizing a transient bass signal according to aspects of the disclosure.DETAILED DESCRIPTION

[0030] Generally disclosed herein is a system and method for optimizing a transient bass signal. The system may be configured to detect the bass signal sounds from the received input sound signal by decomposing the signal into multiple frequency bands based on creating a time domain or frequency domain filter bank. The system may then be configured to conduct analyses of the relevant features of the detected signals to determine the transient bass signal from the bass signal sounds. The system may be configured to calculate the parameters in real time to determine a degree of optimization for the transient bass signal. The system may be configured to compress the transient bass signal and adjust its decay rate. The system may be configured to perform the above processing for each of the multiple frequency bands. A whole enhanced sound signal can be synthesized using the processed frequency bands.

[0031] According to some examples, the system may be configured to maintain the naturalness of the original sound signal without altering the main characteristics of the original sound signal. For example,the system may be configured to differentiate a subtle kick drum sound in jazz music from a kick drum sound in Rock music or a gunshot sound in video games and movies can be differentiated from similar sharp sounds in other genres or media contents. The system may be configured to apply an adequate amount of processing based on the input sound signal characteristics to enhance or optimize the punch sound such that the original sound signal balance may be preserved.

[0032] According to some examples, the system may be configured to optimize the transient bass signal by controlling or adjusting its onset time. The system may also be configured to use stationary bass enhancement techniques in conjunction with changing the decay rate of the compressed transient signal.

[0033] According to some examples, the system may further optimize the transient bass energy based on the utilization of pre-defined presets. For example, music, movie, and game genres can represent the characteristics of the sound components within the content. The pre-defined preset can be formulated and stored by the system based on an analysis of both transient and non-transient aspects of bass sound signals across various content genres or categories. Since the available content genres may be limited, the categories may not fully capture the full characteristics of bass sound in the content.

[0034] FIGS. 1 A-C depict graphs illustrating energy of various transient bass signals. The transient signal optimization system (“system”) may be configured to receive input audio signals and detect bass sound signals. FIG.1A depicts a graph illustrating various bass sound signals 102A-F that can be classified as “hard” punch sounds. The X-axis represents the time of each bass sound signal and the y-axis represents the sound energy level in dB. For example, as represented by bass sound signals 102A-F, the hard “punch” sounds may have characteristics of fast onset time, balanced sustain time and relatively fast decay time. Attack / onset time may refer to the time taken for the rise of the level from zero level to the peak energy level. Sustain time may refer to the time the sound signal maintains its level until the sound signal is released from the peak level. Release / decay time may refer to the time taken for the sound signal to decay from sustain level to zero level. For example, bass sound signal 102A may have characteristics of 0.1 second of the onset time, 0.1 second of the sustain time, and 0.05 second of the decay time. FIG. IB may represent various samples of bass sound signals 104A-E that can be classified as “medium” punch sounds. For example, the bass sound signal 104A may have a longer onset time of approximately 0.15 seconds compared to that of the bass sound signal 102A. The sustain time of the bass sound signal 104A is more than 0.1 second and decay time is approximately 0.4 seconds. FIG. 1C may represent various bass sound signals 106A-C that may be classified as “light” punch sounds. The “light” punch sounds may have much slower onset and decay times compared to the “hard” punch sounds depicted in FIG. 1A. For example, bass sound signal 106 A has 0.2 seconds of the onset time and the decay time is approximately 1 second.

[0035] According to some examples, the system may be configured to receive an input audio signal and detect transient bass sound signals. The system may be configured to analyze each transient bass sound signal’s characteristics and classify each transient bass sound signal into one of the above “light”, “medium” and “hard” punches to achieve an objective assessment of each transient bass sound signal to be optimized. According to each categorized group, the system may be configured to perform additional analyses to discern further parameters that can be used to optimize each transient bass sound signal.

[0036] FIG. 2 depicts a graph illustrating a principal component analysis (PCA) on transient bass signals with various audio features. PCA may refer to a linear dimensional reduction technique that can be used for visualization and data preprocessing. For example, each feature of the audio signal may be linearly transformed onto a new coordinate system such as PCI, PC2, PC3, and PC4 (not shown) as depicted in FIG.2 Each coordinate may capture the largest variation in certain parameters such as release time, peak power sustain energy, or attack time, in each coordinate direction. The contribution of each feature to the PC coordinates can be calculated by projecting the feature to each PC coordinate. In the example of FIG.2, the system may be configured to determine that group 202 shares similar data values of the release time, peak power, sustain energy, and attack time. The system may also be configured to determine that a group 204 or 206 may be formed based on the above four parameters. The system may be configured to analyze each bass sound signal using different parameters and determine the effective parameters for enhancement and optimization purposes. For example, the four parameters: peak power, release time, attack time and sustain energy may be used as important features as the PCA demonstrates that all four parameters are suitable to differentiate each cluster of the punch sound groups 202, 204, and 206.

[0037] FIG. 3 depicts a block diagram illustrating an example transient bass signal optimization system. User computing device 312 and server computing device 315 can be communicatively coupled to one or more storage devices 330 over a network 360. The storage device(s) 330 can be a combination of volatile and non-volatile memory and can be at the same or different physical locations than the computing devices 312, 315. For example, the storage device(s) 330 can include any type of non-transitory computer-readable medium capable of storing information, such as a hard drive, solid-state drive, tape drive, optical storage, memory card, ROM, RAM, DVD, CD-ROM, write-capable, and read-only memories.

[0038] The server computing device 315 can include one or more processors 313 and memory 314. Memory 314 can store information accessible by processor(s) 313, including instructions 321 that can be executed by processor(s) 313. Memory 314 can also include data 323 that can be retrieved, manipulated, or stored by the processor(s) 313. Memory 314 can be a type of non-transitory computer-readable medium capable of storing information accessible by the processor(s) 313, such as volatile and non-volatile memory. The processor(s) 313 can include one or more central processing units (CPUs), graphic processing units (GPUs), field-programmable gate arrays (FPGAs), and / or application-specific integrated circuits (ASICs), such as tensor processing units (TPUs).

[0039] Instructions 321 can include one or more instructions that when executed by the processor! s) 313, cause one or more processors to perform actions defined by the instructions. Instructions 321 can be stored in object code format for direct processing by the processor(s) 313, or in other formats including interpretable scripts or collections of independent source code modules that are interpreted on demand or compiled in advance. Instructions 321 can include instructions for implementing processes consistent with aspects of this disclosure. Such processes can be executed using the processor(s) 313, and / or using other processors remotely located from the server computing device 315.

[0040] The data 323 can be retrieved, stored, or modified by processor(s) 313 in accordance with instructions 321. Data 323 can be stored in computer registers, in a relational or non-relational database asa table having a plurality of different fields and records, or as JSON, YAML, proto, or XML documents. Data 323 can also be formatted in a computer-readable format such as, but not limited to, binary values, ASCII, or Unicode. Moreover, data 323 can include information sufficient to identify relevant information, such as numbers, descriptive text, proprietary codes, pointers, references to data stored in other memories, including other network locations, or information that is used by a function to calculate relevant data.

[0041] User computing device 312 can also be configured similarly to the server computing device 315, with one or more processors 316, memory 317, instructions 318, and data 319. The user computing device 312 can also include a user output 326, and a user input 324. The user input 324 can include any appropriate mechanism or technique for receiving input from a user, such as a keyboard, mouse, mechanical actuators, soft actuators, touch screens, and microphones. User computing device 312 may interact with server computing device 315 to decompose input audio signals into multiple frequency bands and analyze the feature of each band parameter in real time in accordance with the present disclosure. Server computing device 315 may store information relating to metadata about characteristics or features of bass sound signals based on music or media content genres.

[0042] Server computing device 315 can be configured to transmit data to the user computing device 312. The user output 326 can also be used for displaying an interface between the user computing device 312 and the server computing device 315. User computing device 312 may interact with any known stereo systems. User computing device 312 may control individual optimization degree of each analyzed bass sound signal according to user preferences. The user output 326 can alternatively or additionally include one or more speakers, transducers, or other audio outputs, a haptic interface, or other tactile feedback that provides non-visual and non-audible information to the platform user of the user computing device 312.

[0043] Although FIG. 3 illustrates the processors 313, 316 and the memories 314, 317 as being within the computing devices 315, 312, components described in this specification, including the processors 313, 316 and the memories 314, 317 can include multiple processors and memories that can operate in different physical locations and not within the same computing device. For example, some of instructions 321, 318, and data 323, 319 can be stored on a removable SD card and others within a read-only computer chip. Some or all of the instructions and data can be stored in a location physically remote from, yet still accessible by, processors 313, 316. Similarly, processors 313, and 316 can include a collection of processors that can perform concurrent and / or sequential operations. Computing devices 315, and 312 can each include one or more internal clocks providing timing information, which can be used for time measurement for operations and programs run by computing devices 315, and 312.

[0044] The server computing device 315 can be configured to receive requests to process data from the user computing device 312. For example, environment 300 can be part of a computing platform configured to provide a variety of services to users, through various user interfaces and / or APIs exposing the platform services.

[0045] Devices 312, and 315 can be capable of direct and indirect communication over network 360. Devices 312, and 315 can set up listening sockets that may accept an initiating connection for sending and receiving information. The network 360 itself can include various configurations and protocols includingthe Internet, World Wide Web, intranets, virtual private networks, wide area networks, local networks, and private networks using communication protocols proprietary to one or more companies. Network 360 can support a variety of short and long-range connections. The network 360, in addition, or alternatively, can also support wired connections between devices 312, and 315, including over various types of Ethernet connection.

[0046] Although a single server computing device 315 and user computing device 312 are shown in FIG. 3, it is understood that the aspects of the disclosure can be implemented according to a variety of different configurations and quantities of computing devices, including in paradigms for sequential or parallel processing, or over a distributed network of multiple devices. In some implementations, aspects of the disclosure can be performed on a single device, and any combination thereof.

[0047] FIG. 4 depicts a flow diagram illustrating a transient bass signal optimization process. According to block 402, the system may be configured to analyze the received audio signal using filter banks. The system may decompose the received audio signal into multiple frequency bands through the creation of a time domain or frequency domain filter bank. Each filter bank may include an array of filters that can be used to separate the received audio signal into multiple components. The system may be configured to use the filter bank to analyze the signal or components in each frequency band and its sub-bands. According to block 404, the system may be configured to detect transient bass signals based on each band’s onset time and / or energy level. The system may be configured to calculate the attack time, sustain time, and decay rate of each band.

[0048] According to block 406, the system may be configured to determine whether each band is a transient bass signal or a non-transient bass signal as the system performs the calculation of the above parameters. If the system determines that one band of the multiple bands is not a transient bass signal, the system may be configured not to perform the above calculations.

[0049] According to block 408, based on the calculated attack time, sustain time, and decay rate, the system may be configured to adjust the amplitude of each band. The amplitude may refer to a gain of the sound signal. It may refer to the amount of amplification applied to an audio signal measured as a difference between the input and output levels of the sound signal. The system may be configured to receive the user control parameters at block 410. The user control parameters may include the user’ s historical data related to the user's preferred gain level. According to the historical data, the system may be configured to automatically adjust the gain level of each band.

[0050] According to block 412, the system may be configured to further adjust the above band with the dynamically adjusted amplitude by using a dynamic threshold compression technique and by changing the decay rate of the band. Unlike conventional audio signal compressor operation, the dynamic threshold compression technique may be used to adjust its threshold based on the signal level over a given period. Therefore, the system may be configured to maintain consistent sustain characteristics and energy envelopes of the signal, regardless of the transient bass component signal level. The system may also be configured to adjust the decay rate to provide further optimization. With the decay control, the output soundsignal may only maintain “punch” characteristic but rather be perceived as a simple gain boost or an abrupt decay change following the gain boost.

[0051] According to block 414, the system may be configured to repeat the above processes for each of the multiple bands and synthesize to output the adjusted audio signal using the same filter bank used at block 402.

[0052] FIG. 5 depicts a graph illustrating peak envelopes of a transient bass signal using dynamic threshold processing and static threshold processing. The dynamic threshold compression technique may be used to preserve the amount of the processing gain irrespective of the sound signal energy level. A high level of headroom for adjusting the band amplitude can be obtained in an increase in energy within the transient part. For example, signal 502A may represent the original sound signal and 502B may represent the sound signal adjusted using the dynamic threshold compression technique. The signal 502C may represent the signal with adjusted amplitude without a compression technique. Signal 502B may represent 8dB increase of energy level constantly through each portion of the signal while signal 502C may display a larger increase in the first portion of the original sound signal but less increase rate applied to the later portion of the signal. Similarly, signal 504 A may represent a different original sound signal, and signal 504B represents the signal adjusted using the dynamic threshold technique. Similarly to signal 502B, signal 504B may represent the adjusted signal with the compressor that adjusts the amplification according to dynamically changing threshold within the threshold or 8dB or less constantly through every portion of the signal while signal 504C represents the signal adjusted using a static threshold compression technique which applies the fixed threshold of 1 IdB to every portion of the original signal 504A. If the compression threshold remains fixed, the system may be configured to only increase the sustained energy of the sound signal exceeding the compressor threshold while the signal sound below the compressor threshold may only increase the amplitude rather than extending the sustain time.

[0053] FIGS. 6A-6C depicts a graph illustrating a signal level and a power level of the original transient bass signal and the optimized transient bass signal. FIG. 6A represents time domain signal levels of the original signal 602A and processed signal 602B. The X-axis represents time, and the y-axis represents the amplitude of the signal. The system may be configured to increase the amplitude level and also decaying portions of the signal after the time passes 0.5 or 1 second. The system may be configured to maintain the similar envelope shape of the original signal 602A. An envelope may refer to how a sound changes over time. FIG. 6B may represent another example of the original sound signal 604B with the process signal 604B. FIG. 6C shows signal power changes over time using a log scale. The original signal 604A may be processed using the dynamic threshold compression technique and decay rate adjustment to output the processed signal 606B. The overall envelope shape of the original signal 604A may be preserved but the magnitude of the amplification of each portion may vary. For example, at Tl, the amplitude increase is larger than the amplitude increase at T2.

[0054] FIG. 7 depicts a flow diagram illustrating an example method for optimizing a transient bass signal. According to block 702, the system may be configured to receive audio signals. The audio signals may include music, background sound of any type of media content, dialogues within the media content,etc. According to block 704, the system may be configured to divide the received audio signals into a plurality of frequency bands. The system may use a filter bank to divide the audio signal.

[0055] According to block 706, the system may he configured to detect a transient bass signal from each of the plurality of the frequency bands. The system may be configured to analyze each frequency band to determine certain parameters. The system may be configured to determine whether each frequency band is a transient bass sound signal or a non-transient bass sound signal. If it is a non-transient bass sound signal, the system may be configured to stop processing the non-transient bass sound signal.

[0056] According to block 708, the system may be configured to determine a degree of optimization for the detected transient bass signal. The degree of the optimization may be determined by using the classification of transient bass signal sounds. The system may be configured to determine whether each transient bass sound signal corresponds to a hard punch, a medium punch or a light punch sound. The system may be configured to classify the above punch sounds using certain parameters such as peak power, decay time, attack time or sustain energy level.

[0057] According to block 710, the system may be configured to compress the detected transient bass signal using a dynamic threshold compression technique. The system may use a different range of thresholds based on the parameters and classifications determined above.

[0058] According to block 712, the system may be configured to enhance the detected transient bass signal by adjusting the decay rate of the compressed transient bass signal. The system may be configured to change the decay rate based on the pre-stored user historical data or user’s manual commands.

[0059] According to some examples, the system may further optimize the transient bass energy based on the utilization of pre-defined presets. For example, music, movie, and game genres can represent the characteristics of the sound components within the content. The pre-defined preset can be formulated and stored by the system based on an analysis of both transient and non-transient aspects of bass sound signals across various content genres or categories. Since the available content genres may be limited, the categories may not fully capture the full characteristics of bass sound in the content. The system may be configured to create the bass profiles or categories through offline analysis, thereby generating metadata embedded in the content or stored in a separate database. The system may be able to customize the processing of each sound signal on a content-by-content basis or based on the newly defined profiles, rather than being constrained by a group of existing content genres.

[0060] According to block 714, the system may be configured to synthesize the enhanced transient bass signal of each frequency band of the plurality of the frequency bands. The system may be configured to adjust and optimize each detected transient bass sound signal and combine them to generate an output sound signal for the user.

[0061] The transient signal optimization system described herein is beneficial at least in that it provides an audio processing technique that artificially creates or increases the quality of the bass sound signal of the input sound signal by enhancing the “punch” sound, thereby transforming the sound signal to soundfuller and more impactful by dynamically adjusting the original amplitude using a dynamic threshold compression technique and by changing the decay rate of each band of the bass sound signal.

[0062] Aspects of this disclosure can be implemented in digital circuits, computer-readable storage media, as one or more computer programs, or a combination of one or more of the foregoing. The computer- readable storage media can be non-transitory, e.g., as one or more instructions executable by a cloud computing platform and stored on a tangible storage device.

[0063] In this specification, the phrase “configured to” is used in different contexts related to computer systems, hardware, or part of a computer program, engine, or module. When a system is said to be configured to perform one or more operations, this means that the system has appropriate software, firmware, and / or hardware installed on the system that, when in operation, causes the system to perform the one or more operations. When some hardware is said to be configured to perform one or more operations, this means that the hardware includes one or more circuits that, when in operation, receive input and generate output according to the input and corresponding to the one or more operations. When a computer program, engine, or module is said to be configured to perform one or more operations, this means that the computer program includes one or more program instructions, that when executed by one or more computers, causes the one or more computers to perform the one or more operations.

[0064] Although the technology herein has been described with reference to particular examples, it is to be understood that these examples are merely illustrative of the principles and applications of the present technology. It is therefore to be understood that numerous modifications may be made and that other arrangements may be devised without departing from the spirit and scope of the present technology as defined by the appended claims.

[0065] Unless otherwise stated, the foregoing alternative examples are not mutually exclusive, but may be implemented in various combinations to achieve unique advantages. As these and other variations and combinations of the features discussed above can be utilized without departing from the subject matter defined by the claims, the foregoing description should be taken by way of illustration rather than by way of limitation of the subject matter defined by the claims. In addition, the provision of the examples described herein, as well as clauses phrased as "such as," "including" and the like, should not be interpreted as limiting the subject matter of the claims to the specific examples; rather, the examples are intended to illustrate only one of many possible implementations. Further, the same reference numbers in different drawings can identify the same or similar elements.

Claims

CLAIMS1. A system for optimizing a transient bass signal, the system comprising: memory; and one or more processors configured to: receive audio signals; divide the received audio signals into a plurality of frequency bands; detect a transient bass signal from each of the plurality of the frequency bands; determine the degree of optimization for the detected transient bass signal; compress the detected transient bass signal using a dynamic threshold compression technique; enhance the detected transient bass signal by adjusting a decay rate of the compressed transient bass signal; and synthesize the enhanced transient bass signal of each frequency band of the plurality of the frequency bands.

2. The system of claim 1, wherein the degree of the optimization is determined by a binary classification, wherein the binary classification is determined based on the whether the detected transient bass signal is a punch bass sound or non-punch bass sound.

3. The system of claim 1, wherein the degree of the optimization is determined by using a classification of a punch bass sound included in the detected transient bass signal.

4. The system of claim 3, wherein the punch bass sound includes a hard punch, a medium punch or a light punch based on one or more parameters.

5. The system of claim 4, wherein the one or more parameters include at least one or more of peak power, decay time, attach time or sustain energy.

6. The system of claim 3, wherein the one or more processors are configured to adjust the one or more parameters.

7. The system of claim 1, wherein the one or more processors are further configured to control an onset time of the transient bass signal.

8. The system of claim 1, wherein the one or more processors arc further configured to use stationary bass enhancement techniques in conjunction with the dynamic threshold compression technique.

9. The system of claim 8, wherein the stationary bass enhancement techniques include a bass equalization filtering.

10. The system of claim 1, wherein the one or more processors are further configured to classify the audio signals based on a genre of the received audio signals.

11. The system of claim 1, wherein the one or more processors are further configured to utilize a prestored bass profile information to enhance the transient bass signal.

12. The system of claim 1, wherein the one or more processors are further configured to utilize meta data of audio content.

13. The system of claim 12, wherein the meta data includes content categories related to movies, music, or gaming.

14. The system of claim 13, wherein the meta data related to the music includes genre information of the music, wherein the genre information includes classical, pop, rock, hip-hop, electronics, or jazz music.

15. The system of claim 13, wherein the meta data related to the gaming includes game genre information, wherein the game genre information includes first-person shooting games, action games, adventure games, roll-playing games or racing games.

16. The system of claim 1, wherein the one or more processors are further configured to decrease the energy of the transient bass signal by using the dynamic threshold compression technique and the adjustment of the decay rate.

17. The system of claim 1 , wherein the one or more processors are further configured to separate the transient bass signal from a sound mix by using an audio object extraction or separation technique and channel separation technique used for muti-channel contents.

18. The system of claim 1, wherein the one or more processors are further configured to enable a user to control the transient bass signal for personal preferences.

19. A method for optimizing a transient bass signal, the method comprising: receiving, by one or more processors, audio signals; dividing, by the one or more processors, the received audio signals into a plurality of frequency bands;detecting, by the one or more processors, a transient bass signal from each of the plurality of the frequency bands; determining, by the one or more processors, the degree of optimization for the detected transient bass signal; compressing, by the one or more processors, the detected transient bass signal using a dynamic threshold compression technique; enhancing, by the one or more processors, the detected transient bass signal by adjusting a decay rate of the compressed transient bass signal; and synthesizing, by the one or more processors, the enhanced transient bass signal of each frequency band of the plurality of the frequency bands.

20. A non-transitory machine-readable medium comprising machine-readable instruction encoded thereon for performing a method of optimizing a transient bass signal, the method comprising: receiving audio signals; dividing the received audio signals into a plurality of frequency bands; detecting a transient bass signal from each of the plurality of the frequency bands; determining the degree of optimization for the detected transient bass signal; compressing the detected transient bass signal using a dynamic threshold compression technique; enhancing the detected transient bass signal by adjusting a decay rate of the compressed transient bass signal; and synthesizing the enhanced transient bass signal of each frequency band of the plurality of the frequency bands.

Citation Information

Patent Citations

  • Signal processing device, signal processing method, and program therefor

    US20090052695A1

  • Adaptive bass processing system

    US20150146890A1

  • System and method for bass enhancement

    US9307323B2