Audio boosting method using band segmentation
By employing a frequency band segmentation and weighted processing method for audio enhancement, the problems of impairment perception and real-time application in existing stereo audio processing technologies are solved, achieving real-time audio enhancement that improves audio quality while maintaining stereo perception.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2026-04-07
AI Technical Summary
Existing audio enhancement methods are prone to impairing stereo perception when processing stereo audio and are not suitable for real-time applications. Furthermore, existing AI-based methods consume a lot of computational resources, resulting in high latency, which limits their application in real-time scenarios.
The stereo signal is divided into low-frequency and mid-high-frequency components using frequency band segmentation technology, and processed separately. An audio boosting algorithm is used to boost the mid-high-frequency components, and the low-frequency and boosted mid-high-frequency components are combined by weighting and equalization to generate a boosted stereo signal.
While maintaining stereo perception, it effectively improves audio quality, is suitable for real-time applications, reduces processing latency, and achieves efficient audio enhancement.
Smart Images

Figure CN121811897A_ABST
Abstract
Description
[0001] Cross-referencing
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 704,081, filed October 7, 2024, the contents of which are incorporated herein by reference. [Background Technology]
[0003] Audio recording equipment often captures the desired audio subject (such as a person speaking or singing) and unwanted secondary sounds (such as background noise or interference). Traditional audio enhancement methods attempt to isolate and enhance the desired audio subject while suppressing unwanted sounds.
[0004] Existing audio enhancement solutions typically process the entire audio signal as a unit. While this approach is effective at reducing noise, it often compromises the stereo perception of the original recording. Furthermore, these methods usually require processing the entire audio file, making them unsuitable for real-time applications such as telephone calls or live recordings.
[0005] Recent advancements in artificial intelligence and machine learning have enabled more sophisticated audio enhancement features. However, these advanced algorithms typically require substantial computational resources and introduce significant latency, limiting their practical application in real-time scenarios. Therefore, there is a need for an audio enhancement system capable of overcoming these issues. [Summary of the Invention]
[0006] One embodiment provides a real-time audio boosting method executed by one or more processors. The method includes receiving an audio input signal associated with a first channel and a second channel; band-splitting the audio input signal to generate a set of first frequency components associated with the first channel and the second channel, and a set of second frequency components associated with the first channel and the second channel; processing the set of second frequency components using an audio boosting operation to generate a set of boosted second frequency components; and combining the set of first frequency components with the set of boosted second frequency components to generate a boosted audio signal associated with the first channel and the second channel.
[0007] One embodiment provides an apparatus for real-time audio boosting. The apparatus includes one or more processors configured to receive audio input signals associated with a first channel and a second channel, to band-divide the audio input signals to generate a set of first frequency components associated with the first channel and the second channel, and a set of second frequency components associated with the first channel and the second channel, to process the set of second frequency components using an audio boosting operation to generate a set of boosted second frequency components, and to combine the set of first frequency components with the set of boosted second frequency components to generate a boosted audio signal associated with the first channel and the second channel.
[0008] These and other objectives of the invention will undoubtedly become apparent to those skilled in the art upon reading the preferred embodiments described below in detail. [Attached Image Description]
[0009] Figure 1 It demonstrates a scenario where audio enhancement functionality is used in a real recording environment.
[0010] Figure 2 An audio enhancement device according to an embodiment is shown.
[0011] Figure 3 Showing Figure 2 A detailed view of the audio boosting section of the audio boosting device.
[0012] Figure 4 Another audio enhancement device according to an embodiment is shown.
[0013] Figure 5 Showing Figure 4 A detailed view of the audio boosting section of the audio boosting device.
[0014] Figure 6 Showing Figure 4 Detailed implementation of the audio main enhancement operation module.
[0015] Figure 7 A flowchart illustrating a real-time audio enhancement method according to an embodiment is provided.
[0016] Figure 8 A block diagram of an audio recording apparatus 800 according to an embodiment is shown.
Detailed Implementation Methods
[0017] This disclosure provides detailed descriptions of various embodiments. While specific implementation details are provided herein to facilitate a thorough understanding of this disclosure, it will be apparent to those skilled in the art that the invention can be practiced without following all such details. In some cases, exhaustive descriptions of well-established methods, procedures, components, and circuits have been omitted to avoid obscuring this disclosure. It should be understood that technical features described individually with respect to a single figure may be implemented alone or in combination with other features as described in this specification.
[0018] Figure 1 The demonstration showcases audio enhancement functionality in a real-world recording environment. The audio recording device 100 can function as a central hub for receiving and processing multiple audio inputs, whether it's a smartphone, professional recording equipment, or any device equipped with audio capture capabilities. The recording device 100 incorporates microphone technology to capture incoming audio signals from various sources in the environment.
[0019] The recording environment includes multiple sound sources, which the audio recording device 100 should process. The primary audio subject might be someone singing or speaking, representing the main focus of the recording and the sound that needs to be amplified. This is typically positioned as the central element in the recording setup. Additionally, secondary sound sources exist in the environment, represented by two separate graphical representations, producing potentially interfering audio inputs.
[0020] Environmental factors also play a role in the scene, with background noise being an important consideration during recording. The ambient noise, represented by the speaker icon in the diagram, adds another layer of complexity to the audio capture and enhancement process. All these different audio sources—primary sound, secondary sound, and background noise—converge simultaneously onto the recording equipment.
[0021] The practical application of this technology can be illustrated by the example of a concert recording scenario. In this case, the audio recording device 100 can strive to preserve and enhance the singer's voice while effectively removing or reducing unwanted elements such as the audience's singing and ambient noise. A comparison between the original recording and the enhanced version, displayed via speaker icons, demonstrates the technology's ability to significantly improve audio quality while maintaining the authenticity of the main audio elements.
[0022] Figure 2An audio enhancement device 200 according to an embodiment of the present invention is shown. The audio enhancement device 200 may be partially or wholly integrated into the audio recording device 100. The audio enhancement device 200 receives a stereo audio input signal 201, including a left channel signal 202 and a right channel signal 203. This stereo input represents any audio source containing spatial information across the two channels, such as music recordings, live performances, or voice recordings. The stereo input is then processed by a frequency band splitter 230 within a frequency band processing section 210, which separates each channel into different frequency components for specialized processing.
[0023] The frequency band splitter 230 divides the audio into four distinct frequency bands: a low-frequency component 211 for the left channel, a low-frequency component 212 for the right channel, a mid-high frequency component 213 for the left channel, and a mid-high frequency component 214 for the right channel. This separation is crucial because different frequency ranges may contain different types of audio information. The low-frequency components 211 and 212 typically contain fundamental tone and bass information, which is essential for maintaining spatial awareness and is therefore preserved to retain stereo information and routed directly to the mixer and equalization (EQ) block 240.
[0024] Mid-high frequency components 213 and 214, which typically include most of the secondary sounds and noise, undergo additional processing. These components are first combined by a weighted summing block 260 to generate a mixed mono mid-high frequency component 215. The weighting process ensures that the combination of these frequencies maintains proper balance and phase relationships. It is then fed into the audio body boosting section 220 within the audio body boosting algorithm block 250. The audio body boosting section 220 includes the audio boosting algorithm block 250 and a mixer and EQ block 240, which, if necessary, use complex signal processing to maintain stereo perception while improving audio quality.
[0025] The audio main boosting algorithm block 250 uses advanced signal processing technology to process the mixed mono mid-high frequency components 215 to perform one or more audio main boosting operations to identify and boost the main audio components while reducing unwanted secondary sounds and noise. The processed signal is then output to the mixer and EQ block 240. The mixer and EQ block 240 performs the critical task of recombining all processed signals and outputs boosted audio signals 221 and 222 associated with the left and right channel components, ensuring that the spatial characteristics of the original recording are preserved while improving audio quality.
[0026] It should be noted that the frequency division between the low-frequency and mid-high-frequency components represents a carefully considered threshold in audio processing. The boundary at approximately 300Hz marks a significant transition point in the audio spectrum, where different characteristics of sound become prominent. The low-frequency components, occupying the spectrum below 300Hz, contain fundamental tones that contribute significantly to spatial perception, warmth, and depth. These frequencies include bass instruments, fundamental vocal harmonics, and room acoustics, and are crucial for maintaining a natural stereo imaging.
[0027] Mid-to-high frequency components, occupying the spectrum above 300Hz, carry different but equally important acoustic information. This range includes the harmonic content of most musical instruments, vocal production, and many secondary sounds and ambient noises that need to be enhanced or reduced. The 300Hz threshold was chosen based on extensive analysis of human auditory perception and the typical distribution of musical and vocal content in the spectrum.
[0028] This 300Hz frequency division point aligns with important psychoacoustic principles. Below this frequency, human hearing relies more on phase relationships and temporal differences between ears for spatial localization, making these frequencies crucial for maintaining stereo perception. Above 300Hz, the ear increasingly uses intensity differences for spatial localization, allowing for more aggressive processing without compromising the overall stereo image.
[0029] Choosing 300Hz as an approximate dividing point also considered practical implementation aspects in audio processing systems. This frequency strikes an optimal balance between preserving sufficient low-frequency content to retain spatial information and allowing for effective enhancement of most unwanted sounds in the mid-to-high frequency range. This division enables the system to apply different processing strategies to each range, optimizing the enhancement of desired audio elements while preserving the natural characteristics of the recording.
[0030] Furthermore, this frequency threshold acknowledges the different behaviors of sound within these ranges. Low frequencies tend to be more omnidirectional and less affected by room acoustics, while frequencies above 300Hz become increasingly directional and more susceptible to environmental influences. This natural behavior of sound waves affects the contribution of each frequency range to the overall audio experience and guides the processing approach for each band.
[0031] The entire process aims to enhance the main audio elements while preserving the spatial characteristics of the original stereo recording, effectively managing the enhancement of the desired audio content and the maintenance of stereo quality in the final output. This careful balance between enhancement and preservation ensures that the processed audio achieves exceptional clarity and focus on the main audio elements while maintaining its natural stereo image.
[0032] Figure 3A detailed view of the audio boosting section 220 of the audio boosting device 200 according to an embodiment of the present invention is shown. This section receives three main inputs: a left channel low-frequency component 211, a right channel low-frequency component 212, and a mixed mono mid-high frequency component 215 from the frequency band processing section.
[0033] The mixed mono mid-high frequency component 215 is first processed by the audio main boosting algorithm block 250, which applies a specialized boosting technique to produce a boosted mono mid-high frequency component 217. This boosted signal represents a processed version of the mid-high frequency content, in which most of the secondary sounds and noise have been processed.
[0034] Within the mixer and equalizer module 260, a connector 262 combines the boosted mono mid-high frequency component 217 with unprocessed low-frequency components 211 and 212, respectively associated with the left and right channels. The combined signal then passes through dedicated equalization stages: a left-channel equalizer 264 and a right-channel equalizer 266. These equalizers maintain proper frequency balance and ensure that the boosted audio retains a natural character similar to the original recording.
[0035] The final output of the audio boost section 220 includes boosted audio signals 221 and 222, associated with the left and right channels respectively, representing fully processed stereo audio signals. Boosted audio signals 221 and 222 provide enhanced clarity and focus on the main audio subjects by selectively processing different frequency bands while preserving the spatial characteristics of the original stereo recording.
[0036] The entire signal chain within the main audio enhancement section 220 is designed to carefully balance enhancing the desired audio content with preserving the original stereo image, ensuring that the final output maintains both improved audio quality and natural stereo characteristics.
[0037] Figure 4 An audio enhancement device 400 is illustrated according to an embodiment of the present invention. The audio enhancement device 400 may be partially or wholly integrated into an audio recording device 100. The audio enhancement device 400 receives an audio input signal 401 comprising a left channel 402 and a right channel 403. This embodiment is specifically designed for audio enhancement operations requiring full-bandwidth information for computation, such as advanced AI-based processing systems or neural networks that analyze the complete audio spectrum to make intelligent enhancement decisions.
[0038] The signal path is divided into two parallel processes within the frequency band processing section 410. The first path passes the stereo signal through the frequency band splitter 430, extracting the low-frequency components: the left channel low-frequency component 411 and the right channel low-frequency component 412. These low-frequency components 411 and 412 are preserved to maintain stereo information and are directly routed to the mixer and equalizer module 460 within the audio main boost section 420. This preservation is crucial because the low-frequency components contain essential spatial information and contribute significantly to the listener's perception of the stereo image.
[0039] The second path processes the complete stereo signal via a weighted summing block 440, combining the components associated with the left and right channels to generate a mixed mono full-band component 415. This mono signal includes the full spectrum required for modern audio enhancement operations, such as those based on AI or neural networks, which may require full-band information for optimal performance. The weighted summing process ensures that the mono signal maintains proper phase relationships and energy distribution across the spectrum, preventing the loss of critical audio information during the stereo-to-mono conversion.
[0040] Within the main audio boost section 420, the mixed mono full-band component 415 is processed by the main audio boost algorithm block 450, which applies sophisticated boosting techniques or operations to the full-frequency content. These techniques may include advanced signal processing methods such as spectral analysis, machine learning-based feature extraction, and intelligent noise reduction. The boosted signal is then fed into the mixer and equalizer module 460, where it is combined with the preserved stereo low-frequency components 411 and 412.
[0041] The mixer and equalizer module 460 performs the critical task of recombining the boosted mono signal with the retained stereo low-frequency components. This module carefully balances relative levels and frequency response to ensure a natural transition between the processed and unprocessed portions of the spectrum. The module outputs boosted audio signals 421 and 422, associated with the left and right channels, respectively. This alternative architecture ensures that the boost operation has access to full-frequency information while maintaining stereo perception by retaining and properly mixing the low-frequency components.
[0042] This approach represents a significant advancement over traditional enhancement methods, allowing for sophisticated full-band processing while preserving the spatial characteristics crucial for an immersive listening experience. The architecture effectively balances the competing demands of advanced signal processing and natural stereo reproduction.
[0043] Figure 5 A detailed view of the main audio boost section 420 of the audio boost device 400 is shown. This section processes three input signals: a left channel low-frequency component 411, a right channel low-frequency component 412, and a mixed mono full-band component 415 containing the full spectrum.
[0044] The mixed mono full-band component 415 is first processed by the audio main boosting algorithm block 450, which applies advanced boosting techniques or operations to produce a boosted mono full-band component 417. This signal contains a boosted version of the full spectrum, in which the main audio main has been boosted while unwanted elements have been reduced or removed.
[0045] Within the mixer and equalizer sections, the weighted sum block 462 plays a crucial role in combining the boosted mono full-band component 417 with the retained low-frequency components 411 and 412. For audio boosting operations utilizing full-frequency information, the original stereo low-frequency component is mixed with the boosted mono low-frequency component using a weighted sum method to maintain proper balance and spatial characteristics.
[0046] The combined signal then passes through dedicated equalization stages: left channel equalizer 464 and right channel equalizer 466. These equalizers ensure that the final output maintains a natural frequency response while preserving the stereo image of the original recording. The equalizers can be adjusted to match the frequency characteristics of the original audio, helping to maintain a consistent and natural sound.
[0047] The output of the main audio enhancement section 420 includes enhanced audio signals 221 and 222, associated with the left and right channels respectively, representing the fully processed stereo audio signals. This architecture demonstrates how to effectively combine full-band processing with stereo preservation techniques to achieve enhanced audio quality and preserved spatial characteristics in the final output.
[0048] Figure 6 A detailed implementation of the audio boosting algorithm block 450 according to an embodiment of the present invention is shown. This block processes the mixed mono full-band signal 415 through a series of complex processing stages to produce a boosted mono full-band output 417. This architecture represents an advanced audio boosting method that combines artificial intelligence with traditional signal processing techniques.
[0049] The first stage employs an AI secondary sound detector 452 to analyze the input signal to identify and isolate secondary sounds and noise in the audio stream. This AI-based detection system uses advanced pattern recognition and machine learning techniques to distinguish between primary audio content and unwanted secondary sounds and noise. The AI secondary sound detector 452 outputs identified secondary sound components 416, representing unwanted audio elements that need to be removed or reduced from the original signal.
[0050] The removal section 454 comprises three consecutive processing blocks that work together to effectively eliminate detected secondary sounds and noise while preserving the quality of the main audio content. The first block is the subtraction block 455, which performs a precise spectral subtraction process to remove identified secondary sound components 416 from the original signal. This subtraction must be carefully calibrated to maintain the integrity of the main audio content while effectively removing unwanted elements.
[0051] To prevent overprocessing and preserve natural sound characteristics, the signal then passes through a maximum reduction threshold block 456. This critical component ensures that the reduction level applied to any part of the signal does not exceed a predetermined threshold, preventing artifacts or unnatural sound characteristics that could result from overprocessing. The threshold is carefully calibrated to balance effective noise reduction with natural sound preservation.
[0052] The third stage in the removal chain is an energy-based smoothing filter 457, which employs sophisticated filtering techniques to eliminate any potential discontinuities or artifacts that may be introduced during the subtraction process. This filter analyzes the energy distribution across the spectrum to ensure a smooth transition and process natural sound quality in the audio.
[0053] The processed signal appears as an enhanced mono full-band output 417, representing the original audio, in which secondary sounds and noise are effectively reduced while maintaining the integrity and naturalness of the main audio elements. The enhanced output retains the full spectrum required for high-quality audio reproduction while eliminating unnecessary elements that may affect the listening experience.
[0054] The entire process reflects a careful balance between aggressive noise reduction and preservation of natural sonic characteristics, ensuring that the enhanced output maintains high fidelity while effectively removing unwanted audio elements. This sophisticated processing chain demonstrates the power of combining artificial intelligence with traditional signal processing techniques to achieve superior audio enhancement.
[0055] The advantage of this implementation lies in its intelligent and systematic audio enhancement approach. By combining AI-based detection with a carefully controlled removal process, the system achieves more accurate and natural sound effects than traditional methods. The three-stage removal process ensures noise reduction without introducing artifacts or compromising key audio quality, while energy-based smoothing maintains the natural flow and continuity of the sound. This makes the system particularly effective in real-world applications where maintaining audio quality is just as important as removing unwanted noise.
[0056] Figure 7A flowchart illustrating a method 700 for real-time audio enhancement according to an embodiment of the present invention is shown. Method 700 corresponds to the above description and includes a systematic approach to process stereo audio signals while preserving spatial characteristics and improving audio quality. Method 700 includes the following steps:
[0057] S702: Receives audio input signals associated with the first and second channels;
[0058] S704: Band-segment the audio input signal to generate a set of first frequency components associated with the first channel and the second channel, and a set of second frequency components associated with the first channel and the second channel.
[0059] S706: Process the group of second frequency components using an audio boosting operation to generate a group of boosted second frequency components; and
[0060] S708: Combine the first frequency component of the group with the second frequency component of the group to generate a boosted audio signal associated with the first channel and the second channel.
[0061] The first step (S702) involves receiving audio input signals associated with the first and second channels. This stereo input signal represents the original audio content to be boosted, such as a music recording, a live performance, or any other stereo audio source. The two channels typically correspond to the left and right channels in a conventional stereo setup and contain spatial information for creating a stereo image.
[0062] In the next step (S704), the method performs frequency band segmentation on the audio input signal. This crucial step generates two distinct sets of frequency components: a first set of frequency components and a second set of frequency components, both associated with the first and second channels. The first frequency components typically represent low-frequency content, containing fundamental tones and important spatial information. The second frequency components encompass the mid-to-high frequency range, where most secondary sounds and noise are typically present. In some embodiments, the second frequency components may include full-band frequencies relevant to the described audio processing.
[0063] The third step (S706) focuses on processing the set of second frequency components using an audio boosting operation to generate a set of boosted second frequency components. This processing step applies sophisticated signal analysis and boosting techniques to identify and reduce unwanted secondary sounds and noise while preserving and boosting the main audio content. The boosting operation can employ various techniques or operations, such as AI-based detection, spectral subtraction, and adaptive filtering, to achieve optimal results.
[0064] In the final step (S708), the method combines the first set of frequency components with the second set of boosted frequency components to generate a boosted audio signal associated with the first and second channels. This combination process carefully balances the preservation of low-frequency spatial information with the boosted mid-to-high-frequency content, ensuring that the final output maintains a natural stereo imaging while benefiting from the boosting process.
[0065] This method represents a sophisticated real-time audio upscaling approach that addresses the dual challenges of maintaining stereo perception and achieving effective audio upscaling. By processing different frequency bands separately and intelligently recombining them, this method ensures that spatial information is preserved while allowing for effective upscaling of audio content.
[0066] Figure 8 A block diagram of an audio recording apparatus 800 according to an embodiment of the present invention is shown. The apparatus includes several key components interconnected to enable audio capture, processing, and user interaction. The audio recording apparatus 800 includes a processor 810, which acts as a central processing unit, coordinating the operations between all other components and performing audio boosting operations. The processor 800 can handle multiple tasks simultaneously: it can manage real-time audio signal processing, perform boosting operations, coordinate data flow between components, and respond to user input. The processor 800 is capable of handling the complex calculations required for frequency band segmentation, AI-based detection, and signal boosting while maintaining real-time performance.
[0067] Processor 810 is connected to microphone 820, which captures audio from the environment and converts it into an audio input signal. Microphone 820 represents the main input interface of the device. It can be designed to capture stereo audio from the environment in high fidelity and convert it into an analog or digital signal that processor 800 can handle. Microphone 820 may include advanced features such as noise cancellation, directional pickup patterns, and high-quality analog-to-digital conversion to ensure optimal audio capture.
[0068] The audio recording device 800 also includes a memory 830 coupled to the processor 810 for storing recorded audio data and processing operations. This memory may include volatile and non-volatile memory to handle temporary processing needs and long-term data storage. For example, the memory 830 may be a combination of random access memory (RAM) and flash memory to meet real-time processing needs and permanent data retention.
[0069] For user interaction, the device is equipped with a user interface (UI) 840, allowing users to control recording settings and boost parameters. UI 840 provides access to various recording settings, boost parameters, and processing options. UI 840 can be designed to be intuitive while also providing advanced control for professional users who require fine-tuning of boost parameters. Furthermore, the UI may include physical buttons, touch-sensitive controls, or other input mechanisms, depending on the specific implementation.
[0070] Display 850 provides visual feedback and status information to the user. Display 850 can be implemented as an LCD screen, LED array, or other visual output device depending on the specific requirements of the application. Both UI 840 and display 850 are coupled to processor 810 to enable interaction between the user and audio recording device 800.
[0071] This hardware architecture is specifically designed to support the complex audio enhancement features described in the foregoing embodiments. It provides the computational power required for real-time band splitting and enhancement while maintaining responsive user control and efficient data management. This design demonstrates a careful balance between processing power, user accessibility, and practical functionality, making it suitable for real-world recording scenarios.
[0072] This invention represents a significant advancement in audio enhancement technology through its innovative real-time signal processing and stereo preservation methods. Its core technology enables complex non-real-time operations in real-time applications, a breakthrough difficult to achieve in traditional audio processing systems. This is achieved through an innovative dual-path processing architecture that significantly reduces processing latency while maintaining high-quality enhancement capabilities. The system's ability to operate in real-time makes it particularly valuable in applications requiring instant processing, such as live phone calls, concert recordings, and broadcasting.
[0073] A key advantage of the invention lies in its ability to preserve the spatial characteristics of stereo recordings while performing the boosting operation. The system employs fine-grained frequency band management, particularly preserving low-frequency stereo information crucial for spatial perception. This preservation is achieved through sophisticated processing techniques that ensure the boosting process does not disrupt the natural stereo image of the original recording, thus maintaining its original spatial characteristics while improving audio quality.
[0074] This technology combines intelligent processing methods, offering significant advantages over traditional enhancement methods. By utilizing an artificial intelligence (AI)-based detection system, the invention can accurately identify and isolate secondary sounds for removal. A sophisticated three-stage removal process, including a maximum reduction threshold, ensures enhancement is completed without over-processing the audio signal. This careful balance effectively removes unwanted elements while preserving the natural characteristics of the sound, which is crucial for maintaining audio fidelity.
[0075] The invention's flexible architecture is another significant advantage. The system can adapt to both traditional and AI-based boosting operations and support full-band processing when needed. This adaptability makes it suitable for various types of audio content and can be implemented in different recording devices and scenarios. The versatility of the architecture ensures that it can meet diverse audio boosting needs while maintaining consistent performance.
[0076] Quality control mechanisms are implemented throughout the processing chain to ensure optimal results. The system employs energy-based smoothing filters to prevent artifacts, uses a weighted sum method to achieve balanced signal combination, and includes a dedicated equalization (EQ) stage to maintain a natural frequency response. This precise control over the boosting process prevents audio signal degradation while ensuring effective boosting. This comprehensive quality control approach demonstrates how this invention successfully combines advanced processing capabilities with usability in practical applications.
[0077] The terminology used in the various embodiments described herein is intended to describe particular embodiments and should not be construed as limiting. In this specification and the appended claims, the singular forms “a,” “an,” and “the” are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0078] It should be understood that the term "and / or" as used in this specification is intended to cover any and all possible combinations of one or more of the associated listed items. Furthermore, it should be noted that the terms "comprising," "including," "comprises," and / or "including" as used in this specification indicate the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0079] In the context of this disclosure, the terms “coupled,” “connected,” “in connection,” “electrically connected,” and similar expressions are used interchangeably to broadly describe the state of an electrical or electronic connection. Furthermore, when one entity “communicates” with another entity (or entities), it means that it electrically transmits and / or receives information signals to / from the other entity, regardless of whether these signals contain image / voice information or data / control information, and regardless of whether the signal type is analog or digital. It should be noted that such communication can be wired or wireless. The use of these terms is intended to cover all forms of electrical or electronic connection relevant to the described embodiments.
[0080] Ordinal identifiers such as "first," "second," etc., used in the specification and claims are used to distinguish multiple instances of elements with similar names. These identifiers do not imply any inherent order, priority, or chronological order of the manufacturing process or the functional relationship between elements, but are only used to uniquely identify and distinguish different instances of elements with a common name or description.
[0081] The directional terms used in the embodiments, such as up, down, left, right, upper side, lower side, front, or rear, are merely for reference to the accompanying drawings. Therefore, the directional terms used in this disclosure are for illustrative purposes only and are not intended to limit the scope of this disclosure. It should be noted that the elements specifically described or marked may exist in various forms for those skilled in the art.
[0082] The approximate and degree terms used in this specification and appended claims, such as “approximately,” “about,” “usually,” “substantially,” “almost,” “regarding,” etc., are intended to account for variations in precision, manufacturing tolerances, measurement accuracy, environmental conditions, and inherent material properties that may affect the described features or properties. These variations may range from ±20% in a broader range of applications to ±10%, ±5%, ±3%, ±2%, ±1%, or ±0.5% in more precise implementations. In any given context, the specific degree of variation covered by these approximate terms is determined by the nature of the described components, relationships, or parameters, the technical requirements of a particular embodiment, and the understanding of someone skilled in the art.
[0083] This interpretation of terminology is intended to ensure clarity and consistency in the specification and claims, and should not be construed as limiting the scope of the disclosed embodiments or appended claims.
[0084] Various exemplary components, logic, logic blocks, modules, circuits, operations, and algorithmic processes related to the disclosed embodiments can be implemented as electronic hardware, firmware, software, or a combination of hardware, firmware, and software, including the structures disclosed in this specification and their structural equivalents. The interchangeability of hardware, firmware, and software has been generally described, categorized by function, and illustrated in the various exemplary components, blocks, modules, circuits, and processes described above. Whether a function is implemented via hardware, firmware, or software depends on the specific application and design constraints of the overall system.
[0085] The hardware and data processing devices used to implement the various exemplary components, logic, logic blocks, modules, and circuits described herein may include, but are not limited to, one or a combination of the following: general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), other programmable logic devices (PLDs), discrete gate or transistor logic, discrete hardware components, or any suitable combination thereof. This hardware and devices shall be configured to perform the functions described herein.
[0086] A general-purpose processor may include, but is not limited to, a microprocessor, or alternatively, any conventional processor, controller, microcontroller, or state machine. In some implementations, the processor may be implemented through a combination of computing devices. These combinations may include, for example, a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other configuration suitable for the intended application.
[0087] It should be understood that in some embodiments, specific processes, operations, or methods may be performed by circuitry designed specifically for a particular function. Such function-specific circuitry may be optimized to improve performance, efficiency, or other task-related metrics. The choice of specific hardware implementation should be determined based on the specific requirements of the application, which may include, but are not limited to, performance specifications, power consumption limitations, cost considerations, and size constraints.
[0088] In some respects, the subject matter described herein can be implemented as software. Specifically, the various functions of the disclosed components, or the steps of the methods, operations, processes, or algorithms described herein, can be implemented as one or more modules in one or more computer programs. These computer programs may include non-transitory processor-executable or computer-executable instructions encoded on one or more tangible processor-readable or computer-readable storage media. Such instructions are configured to be executed by a data processing device that includes components of the device described herein. The aforementioned storage media may include, but are not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage, or other magnetic storage devices, or any medium capable of storing program code in the form of instructions or data structures. It should be understood that combinations of the aforementioned storage media are also within the scope of the computer-readable storage media disclosed herein.
[0089] Various modifications to the embodiments described in this disclosure will likely be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments without departing from the spirit or scope of this disclosure. Therefore, the claims should not be limited to the embodiments shown herein, but should be given the broadest scope consistent with this disclosure, the principles and novel features revealed herein.
[0090] In some implementations, embodiments may include the disclosed features and may optionally include additional features not explicitly described herein. Conversely, alternative implementations may be characterized by the substantial or complete absence of undisclosed elements. For the avoidance of doubt, it should be understood that in some embodiments, undisclosed elements may be intentionally omitted, partially or entirely, without departing from the scope of the invention. The omission of such undisclosed elements should not be construed as limiting the breadth of the claimed subject matter, provided that explicitly disclosed features are present in the embodiments.
[0091] Furthermore, the various features described herein in the context of individual embodiments can also be implemented in combination in a single implementation. Conversely, the various features described herein in the context of a single implementation can also be implemented individually or in any suitable sub-combination in multiple embodiments. Thus, although the features described above may be described as functioning in a particular combination, or even initially claimed to be so, in some cases one or more features may be removed from the combination, and the claimed combination may refer to a sub-combination or a variation of a sub-combination.
[0092] Operations depicted in a specific order in a diagram should not be interpreted as a requirement to strictly follow that order in practice, nor should they imply that all shown operations must be performed to achieve the desired result. Illustrative flowcharts may represent example processes, but it should be understood that additional operations, not shown, may be inserted at various points in the depicted sequence. Such additional operations may occur before, after, simultaneously with, or between any of the shown operations.
[0093] Furthermore, it should be understood that the various figures and component diagrams provided and discussed herein are for illustrative purposes only and are not drawn to scale. These visual representations are intended to facilitate understanding of the described embodiments and should not be construed as precise technical drawings or to limit the scope of the invention to the specific arrangements depicted.
[0094] In some implementations, multitasking and parallel processing may prove advantageous. Furthermore, although various system components are described as separate entities in some embodiments, this separation should not be construed as mandatory for all embodiments. The program components and systems described herein may be integrated into a single software package or distributed across multiple software packages, depending on specific implementation requirements.
[0095] It should be noted that other embodiments beyond the explicitly described embodiments fall within the scope of the appended claims. In some cases, the actions specified in the claims can be performed in a different order than they are presented, while still achieving the desired result. This flexibility in the order of execution is an inherent aspect of the claimed process and should be considered within the scope of the invention.
[0096] Those skilled in the art will readily observe that numerous modifications and alterations can be made to the apparatus and methods while retaining the teachings of the invention. Therefore, the foregoing disclosure should be interpreted only by the limits of the appended claims.
Claims
1. A real-time audio enhancement method executed by one or more processors, comprising: Receive audio input signals associated with the first and second channels; The audio input signal is band-segmented to generate a set of first frequency components associated with the first channel and the second channel, and a set of second frequency components associated with the first channel and the second channel. The group of second frequency components is processed using an audio boosting operation to generate a group of boosted second frequency components; as well as The first frequency component is combined with the second frequency component to generate a boosted audio signal associated with the first and second channels.
2. The method of claim 1, wherein processing the group of second frequency components using the audio boost operation comprises: The set of second frequency components associated with the first channel is combined with the set of second frequency components associated with the second channel, and a weighted sum is used to generate a set of mixed mono second frequency components. as well as Apply the audio boost operation to the group of mixed mono second frequency components to generate the group of boosted second frequency components.
3. The method of claim 2, wherein applying the audio boost operation to the group of mixed mono second frequency components comprises: A set of secondary sound components is detected from the second frequency component of the mixed mono channel; as well as Reduce the group of secondary sound components by the maximum reduction threshold.
4. The method of claim 3, wherein reducing the group of secondary sound components comprises: Subtract the secondary sound components from the group of mixed mono second frequency components to generate a group of subtracted mixed mono second frequency components. as well as An energy-based smoothing filter is applied to the subtracted mixed mono second frequency component of the group.
5. The method of claim 1, wherein combining the first set of frequency components with the second set of boosted frequency components includes adjusting the equalization (EQ) of the first channel and the second channel according to the audio input signal.
6. The method of claim 1, further comprising the processor outputting the boosted audio signal associated with the first channel and the second channel.
7. The method of claim 1, wherein the first frequency components comprise low-frequency components with frequencies approximately below 300 Hz.
8. The method of claim 1, wherein the second frequency component comprises a high-frequency component with a frequency approximately higher than 300 Hz.
9. The method of claim 1, wherein the set of second frequency components includes approximately full-band frequencies associated with the audio input signal.
10. A device for real-time audio enhancement, comprising: One or more processors that receive audio input signals associated with the first and second channels; The audio input signal is band-segmented to generate a set of first frequency components associated with the first channel and the second channel, and a set of second frequency components associated with the first channel and the second channel. The group of second frequency components is processed using an audio boosting operation to generate a group of boosted second frequency components; as well as The first frequency component is combined with the second frequency component to generate a boosted audio signal associated with the first and second channels.
11. The device of claim 10, wherein the one or more processors are further configured to: The set of second frequency components associated with the first channel is combined with the set of second frequency components associated with the second channel, and a weighted sum is used to generate a set of mixed mono second frequency components; and Apply the audio boost operation to the group of mixed mono second frequency components to generate the group of boosted second frequency components.
12. The device of claim 11, wherein the one or more processors are further configured to: Detect a set of secondary sound components from the group of mixed mono second frequency components; and Reduce the group of secondary sound components by the maximum reduction threshold.
13. The device of claim 12, wherein the one or more processors are further configured to: Subtract the secondary sound components from the group of mixed mono second frequency components to generate a group of subtracted mixed mono second frequency components; and An energy-based smoothing filter is applied to the subtracted mixed mono second frequency component of the group.
14. The device of claim 10, wherein, in order to combine the first set of frequency components with the second set of boosted frequency components, the processor is configured to adjust the equalization (EQ) of the first channel and the second channel according to the audio input signal.
15. The apparatus of claim 10, wherein the one or more processors are further configured to output the boosted audio signal associated with the first channel and the second channel.
16. The apparatus of claim 10, wherein the first frequency component set includes low-frequency components with frequencies approximately below 300 Hz.
17. The apparatus of claim 10, wherein the second frequency component set includes high-frequency components with frequencies approximately higher than 300 Hz.
18. The apparatus of claim 10, wherein the second set of frequency components includes approximately full-band frequencies associated with the audio input signal.