Audio processing method, system, device and medium based on multi-order frequency adjustment

By employing a multi-stage frequency adjustment audio processing method, the problem of insufficient audio continuity and smoothness in existing technologies has been solved, achieving continuity and smoothness in the audio frequency curve and enhancing the auditory experience.

CN121565185BActive Publication Date: 2026-07-28深圳市鸿宇光电有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Existing audio processing methods affect the continuity and smoothness of audio during frequency division, noise reduction, and resynthesis, resulting in output audio that does not meet the user's auditory requirements.

Method used

The audio processing method employs multi-order frequency adjustment. After frequency division, noise reduction, and synthesis of the initial audio, at least two EQ adjustment modules for multi-order frequency adjustment are used to adjust the frequency based on the target cavity characteristic data of the sound-producing device. This optimizes the frequency balance of the audio, enhances coherence and smoothness, and removes residual noise and bumps.

Benefits of technology

It effectively avoids artificial effects at the edges of frequency bands, making the frequency curve of the final output audio more continuous and smooth, which meets the user's listening requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121565185B_ABST
    Figure CN121565185B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of audio processing based on multi-order frequency adjustment, and discloses an audio processing method, system, device and medium based on multi-order frequency adjustment, wherein the method comprises the following steps: acquiring initial audio; performing frequency division on the initial audio to obtain initial frequency band data, wherein the initial frequency band data comprises low-frequency data, medium-frequency data and high-frequency data; performing denoising processing on the initial frequency band data to obtain target frequency band data; performing audio synthesis on the target frequency band data to obtain audio to be adjusted; and based on target cavity characteristic data of a sound production device, at least two EQ adjustment modules for multi-order frequency adjustment are adopted to perform frequency adjustment on the audio to be adjusted to obtain target audio suitable for the sound production device. Therefore, the artificial effect of the frequency band edge is effectively avoided, the frequency curve of the finally output audio is more continuous and smooth, and the audio is more in line with the hearing requirements of users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing technology based on multi-order frequency adjustment, and in particular to an audio processing method, system, device and medium based on multi-order frequency adjustment. Background Technology

[0002] In the field of audio processing technology, common audio processing methods include frequency division of the initial audio into three frequency bands: low frequency, mid frequency, and high frequency. Then, each frequency band is denoised separately to eliminate noise. Finally, the denoised frequency band data is synthesized into a complete audio file. This frequency division, denoising, and recombination architecture is effective in improving audio quality. However, because denoising is performed independently on absolutely divided frequency bands, the denoising algorithm can produce unpredictable artificial effects at the edges of each frequency band (especially near the cutoff frequency). For example, noise within a frequency band may be removed, but unnatural "cliffs" or "tails" may remain at the frequency division point, affecting the continuity and smoothness of the audio and causing the final output audio to fail to meet the user's listening requirements. Summary of the Invention

[0003] Based on this, it is necessary to address the technical problem that the architecture of frequency division, noise reduction and resynthesis in the existing technology affects the continuity and smoothness of audio, resulting in the final output audio not meeting the user's auditory requirements. In this regard, a data processing technology field is proposed, which in particular relates to an audio processing method, system, device and medium based on multi-order frequency adjustment.

[0004] In a first aspect, an audio processing method based on multi-order frequency adjustment is provided, the method comprising: Get the initial audio; The initial audio is divided into frequencies to obtain initial frequency band data, wherein the initial frequency band data includes: low frequency data, mid frequency data and high frequency data; The initial frequency band data is denoised to obtain the target frequency band data; The target frequency band data is used for audio synthesis to obtain the audio to be adjusted; Based on the target cavity characteristic data of the sound-generating device, at least two EQ adjustment modules for multi-level frequency adjustment are used to adjust the frequency of the audio to be adjusted, so as to obtain the target audio suitable for the sound-generating device.

[0005] Secondly, an audio processing system based on multi-order frequency adjustment is provided, the system comprising: an audio processing terminal and a sound-generating device; The audio processing terminal is configured to implement the audio processing method based on multi-order frequency adjustment as described in the first aspect; The sound-generating device is used to play the target audio output by the audio processing terminal.

[0006] Thirdly, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the above-described audio processing method based on multi-order frequency adjustment.

[0007] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described audio processing method based on multi-order frequency adjustment.

[0008] This application discloses an audio processing method, system, device, and medium based on multi-order frequency adjustment. The method involves acquiring initial audio, dividing the initial audio into frequencies to obtain initial frequency band data, which includes low-frequency, mid-frequency, and high-frequency data. Noise reduction processing is then performed on the initial frequency band data to obtain target frequency band data. Audio synthesis is then performed on the target frequency band data to obtain the audio to be adjusted. Based on the target cavity characteristic data of the sound-producing device, at least two EQ adjustment modules for multi-order frequency adjustment are used to adjust the frequency of the audio to be adjusted, resulting in a target audio suitable for the sound-producing device. In other words, after dividing, denoising, and synthesizing the initial audio to obtain the audio to be adjusted, this application uses at least two EQ adjustment modules for multi-order frequency adjustment based on the target cavity characteristic data of the sound-producing device. Multi-order frequency adjustment can progressively optimize the frequency balance of the audio, enhance coherence and smoothness, and remove residual spikes and bumps, thereby effectively avoiding artificial effects at frequency band edges. This results in a more continuous and smoother frequency curve in the final output audio, making the audio more in line with the user's auditory requirements. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] in: Figure 1 This is an application environment diagram of an audio processing method based on multi-order frequency adjustment in one embodiment; Figure 2 This is a flowchart of an audio processing method based on multi-order frequency adjustment in one embodiment; Figure 3 This is a block diagram of an audio processing system based on multi-order frequency adjustment in one embodiment; Figure 4 This is a structural block diagram of a computer device in one embodiment; Figure 5 This is another structural block diagram of a computer device in one embodiment. Detailed Implementation

[0011] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0012] The audio processing method based on multi-order frequency adjustment provided in this invention can be applied to, for example... Figure 1 In the application environment, client 110 communicates with server 120 via the network.

[0013] Server 120 can receive initial audio through client 110. Server 120 is configured to implement the audio processing method based on multi-order frequency adjustment of this application, specifically including: acquiring initial audio; dividing the initial audio into frequencies to obtain initial frequency band data, wherein the initial frequency band data includes: low-frequency data, mid-frequency data, and high-frequency data; performing noise reduction processing on the initial frequency band data to obtain target frequency band data; performing audio synthesis on the target frequency band data to obtain audio to be adjusted; and using at least two EQ adjustment modules for multi-order frequency adjustment based on the target cavity characteristic data of the sound-producing device to adjust the frequency of the audio to be adjusted, thereby obtaining target audio suitable for the sound-producing device.

[0014] Client 110 is configured to implement the audio processing method based on multi-order frequency adjustment of this application, specifically including: acquiring initial audio; dividing the initial audio into frequencies to obtain initial frequency band data, wherein the initial frequency band data includes: low-frequency data, mid-frequency data, and high-frequency data; performing noise reduction processing on the initial frequency band data to obtain target frequency band data; performing audio synthesis on the target frequency band data to obtain audio to be adjusted; and using at least two EQ adjustment modules for multi-order frequency adjustment based on the target cavity characteristic data of the sound-generating device to adjust the frequency of the audio to be adjusted, thereby obtaining target audio suitable for the sound-generating device.

[0015] After dividing, denoising, and synthesizing the initial audio to obtain the audio to be adjusted, this application uses at least two EQ adjustment modules for multi-stage frequency adjustment based on the target cavity characteristic data of the sound-producing device. Multi-stage frequency adjustment can gradually optimize the frequency balance of the audio, enhance coherence and smoothness, and remove residual spikes and bumps, thereby effectively avoiding artificial effects at the frequency band edges, making the frequency curve of the final output audio more continuous and smooth, and making the audio more in line with the user's listening requirements.

[0016] The client 110 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server 120 can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0017] It is understood that the electronic device (i.e., client 110) is selected from the processing module integrated in the sound-emitting device, the sound-emitting device itself with data processing capabilities, or a smart terminal that has both data processing and audio playback functions. The electronic device may also be server 120.

[0018] The sound-generating device includes a loudspeaker unit (the core of the electro-acoustic conversion), a housing, and a sound guide hole. The loudspeaker unit includes a diaphragm, a voice coil, and a magnetic circuit system. The housing defines a acoustic cavity, and its internal structure and materials determine the cavity characteristics of the device. The sound guide hole communicates with the acoustic cavity and is used to regulate airflow and balance internal and external air pressure. Furthermore, the device can integrate the aforementioned audio processing system or an interface module connected to it, thereby receiving the target audio signal or its corresponding signal after multi-order frequency adjustment, and converting it into a continuous, smooth, high-fidelity sound that highly matches the physical characteristics of the sound-generating device through the loudspeaker unit.

[0019] In this application, audio refers to all audible sound wave physical signals, including speech, music, environmental noise, etc.

[0020] Cavity characteristics refer to the sum of the decisive influences of the physical structure of a sound-producing device (such as a loudspeaker or headphones) on its acoustic performance, especially its frequency response characteristics. Essentially, it is an inherent "acoustic fingerprint" of the device that determines how audio signals are converted into the final sound.

[0021] Specifically, cavity characteristics include, but are not limited to, acoustic cavity structure and volume, sound tube / vent design, diaphragm material and suspension system, and shell material and damping.

[0022] Acoustic cavity structure and volume: The shape, size and cavity partition of the space inside the device housing, these factors work together with the speaker unit to produce resonance or cancellation at specific frequencies (especially low frequencies).

[0023] Sound guide tube / vent hole design: a structure used to balance the air pressure inside and outside the cavity. Its length, diameter and shape will significantly adjust the extension and texture of low frequencies.

[0024] Diaphragm material and suspension system: The material and structure of the core components of the speaker unit directly affect its inertia, rigidity and damping, which determine the transient response and distortion in the mid and high frequencies.

[0025] Shell material and damping: The materials used to manufacture the shell (such as plastic, metal, wood) and the sound-absorbing materials inside it will affect the intensity of cavity resonance and the suppression of unnecessary vibrations.

[0026] The target cavity characteristic data is a set of data describing the inherent acoustic properties of the sound-generating device.

[0027] When using a three-order frequency adjustment, the "multi-order frequency adjustment EQ module" specifically includes: based on the "target cavity characteristics" (such as frequency response curve), the first EQ adjustment module can attenuate the inherent resonant peaks of the device and enhance its recessed frequency bands to achieve "frequency balance"; the second EQ adjustment module can smoothly transition between different frequency bands, eliminate the "cliff effect" that may be caused by frequency division noise reduction, and enhance "coherence and smoothness"; the third EQ adjustment module, with a finer bandwidth, eliminates the minor "gaps and bumps" that may remain after the first two rounds of adjustment, achieving purity across the entire frequency band.

[0028] Optionally, the target cavity characteristic data includes, but is not limited to, the frequency response curve, main resonant frequency and its quality factor (Q value) of the sound-generating device, and the harmonic distortion curve under specified output conditions.

[0029] Optionally, the target cavity characteristic data includes, but is not limited to, the frequency response curve, main resonant frequency and its quality factor (Q value), harmonic distortion curve under specified output conditions, and sound pressure level sensitivity of the sound-generating device.

[0030] A frequency response curve is a curve showing how the output sound pressure level of a sound-generating device changes with frequency under specified conditions.

[0031] Harmonic distortion curves describe the characteristics of harmonic distortion. Harmonic distortion characteristics are the percentage of total harmonic distortion or specific order (such as second and third) harmonic distortion produced by a device at a given frequency and sound pressure level.

[0032] The quality factor is a parameter that describes the sharpness of resonance; a high Q value indicates a sharp resonance peak, while a low Q value indicates a smooth resonance.

[0033] Sound pressure level sensitivity is the sound pressure level produced by a device at a specified distance under a given input power.

[0034] The frequency response curve is the basis for the EQ adjustment module to perform multi-stage frequency adjustment.

[0035] The quality factor describes the sharpness of the resonance.

[0036] EQ adjustment modules (especially the first and second orders) need to perform precise attenuation or smoothing of the resonant peaks based on the resonant frequency and its quality factor in order to eliminate "hum" or "box sound".

[0037] In multi-order frequency adjustment, pre-attenuation can be performed at frequencies where distortion is likely to be high based on the harmonic distortion curve, in order to reduce distortion during actual playback and thus "remove residual glitches in the audio".

[0038] Sound pressure level sensitivity affects the gain structure of the entire audio link. This application can adjust the gain reference of each EQ adjustment module based on this parameter to ensure that the loudness of the final output "target audio" is appropriate.

[0039] EQ specifically refers to a DSP-based digital equalization processing technology that features dynamically configurable parameters and operates in a multi-level cascading manner.

[0040] EQ adjustment, short for "equalizer adjustment", refers to the technical process of selectively enhancing or weakening the intensity (gain) of specific frequency components in an audio signal through electronic circuits or digital algorithms, thereby changing the timbre of the audio, compensating for defects, or optimizing the overall listening experience.

[0041] In its implementation, the EQ adjustment module is a parametric equalizer based on digital signal processing (DSP). It precisely controls specific frequency components of the input audio by constructing one or more second-order IIR (Infinite Impulse Response) filter cores in the digital domain. These cores dynamically generate filter coefficients based on the center frequency, gain, and Q value, for example, using direct type II or dual second-order filters. Each filter core receives the center frequency, gain, and quality factor parameters from the target adjustment configuration and dynamically generates corresponding filter coefficients, thereby achieving filtering operations that boost or attenuate at a specified frequency point according to the required bandwidth. Multiple such filter cores can work in parallel or cascaded to form a complete EQ adjustment module to perform multi-order, multi-frequency adjustment tasks.

[0042] The EQ adjustment module is implemented in the form of DSP (Digital Signal Processing) algorithms, such as running a dual second-order filter core on an embedded system or a general-purpose processor.

[0043] The present invention will now be described in detail through specific embodiments.

[0044] Please see Figure 2 As shown, Figure 2 A flowchart illustrating an audio processing method based on multi-order frequency adjustment provided in an embodiment of the present invention includes the following steps: S1: Get the initial audio; The initial audio is audio that needs to be written based on multi-order frequency adjustment.

[0045] Optionally, the initial audio may be derived from voice signals captured in real time by a microphone, music data in audio files (such as WAV or MP3 formats), audio streams transmitted via network streaming media, or pre-recorded audio read from storage media (such as hard drives or memory).

[0046] Optionally, the initial audio can be obtained by amplifying (e.g., linear amplification, adjusting the signal amplitude) the voice signal collected in real time by the microphone, music data in audio files (such as WAV and MP3 formats), or audio streams transmitted via network streaming media.

[0047] S2: Divide the initial audio into frequencies to obtain initial frequency band data, wherein the initial frequency band data includes: low frequency data, mid frequency data and high frequency data; Low frequencies: typically refer to 20Hz-250Hz, including 20Hz and 250Hz. They are responsible for the foundation, thickness, and power of the sound. Examples include the impact of a kick drum, the rhythm of a bass, and the rumble of the environment.

[0048] Mid-frequency range: typically refers to 250Hz-4kHz, excluding 250Hz but including 4kHz. This is the most sensitive area for the human ear, containing the core fundamental tone and main harmonics of human voices and most musical instruments, determining the clarity, fullness, and distinctiveness of the sound.

[0049] High frequencies: typically refer to 4kHz-20kHz, excluding 4kHz but including 20kHz. They are responsible for the detail, clarity, and airiness of the sound. Examples include the crisp sound of cymbals, the sibilance (s, t) in vocals, and the overtones of a violin.

[0050] "Frequency division" refers to the process of dividing audio into multiple sub-bands according to frequency range, aiming to separate the energy distribution of different frequency bands. "Initial frequency band data" is the data set of each frequency sub-band obtained after frequency division, which typically includes low-frequency data (corresponding to the fundamental frequency and formants), mid-frequency data (corresponding to the main speech and instrument frequencies), and high-frequency data (corresponding to harmonics and detail components). Frequency division helps to process different frequency bands independently to optimize audio quality.

[0051] Low-frequency data refers to low-frequency related data in the initial audio. Mid-frequency data refers to mid-frequency related data in the initial audio. High-frequency data refers to high-frequency related data in the initial audio.

[0052] Frequency division can be achieved through digital filter banks (such as Butterworth filters and Chebyshev filters) or frequency domain transformations (such as Fast Fourier Transform). For example: using a bandpass digital filter bank: design low-pass, bandpass, and high-pass filters to extract each frequency band; based on frequency domain analysis: convert the audio to the frequency domain using FFT (Fast Fourier Transform), extract the coefficients of a specific frequency range, and then restore it to the time-domain sub-band signal using inverse FFT (Inverse Fourier Transform).

[0053] Specifically, low-frequency data is low-frequency related data extracted from the initial audio through a low-pass filter; mid-frequency data is mid-frequency related data extracted from the initial audio through a band-pass filter; and high-frequency data is high-frequency related data extracted from the initial audio through a high-pass filter.

[0054] S3: Perform noise reduction processing on the initial frequency band data to obtain the target frequency band data; Noise reduction processing refers to the technical process of suppressing or eliminating noise components from the initial frequency band data, with the aim of improving the signal-to-noise ratio and auditory clarity.

[0055] The target frequency band data is the clean frequency band data obtained after noise reduction, which retains the effective components of the original audio while reducing environmental noise, circuit noise, or artifacts introduced by the algorithm.

[0056] Specifically, for low-frequency data, adaptive filters (such as the LMS algorithm) are used to suppress low-frequency hum; for mid-frequency data, spectral subtraction is used: for mid-frequency data, FFT is used to transform to the frequency domain, the noise power spectrum is estimated (e.g., from the silent segment), and the noise component is subtracted from the signal power spectrum, and then restored by IFFT; for high-frequency data, wavelet denoising is applied: wavelet decomposition is performed on high-frequency data, thresholding is performed on detail coefficients (e.g., soft thresholding), and high-frequency noise is removed after reconstruction.

[0057] The denoising parameters used in the denoising process (such as threshold and filter order) can be dynamically adjusted according to the frequency band characteristics. The processed data for each frequency band is the target frequency band data, which is stored as an intermediate result.

[0058] It is understandable that the target frequency band data includes: denoised low-frequency data, denoised mid-frequency data, and denoised high-frequency data.

[0059] S4: Perform audio synthesis on the target frequency band data to obtain the audio to be adjusted; Audio synthesis refers to the process of merging the denoised data of each frequency band (denoised low-frequency data, denoised mid-frequency data, and denoised high-frequency data) back into a complete audio signal, aiming to restore the temporal continuity and frequency band consistency of audio.

[0060] Audio synthesis methods include additive synthesis, overlapping addition, filter combination, or frequency domain synthesis. For example, denoised low-frequency data, denoised mid-frequency data, and denoised high-frequency data can be weighted and superimposed to synthesize a single audio stream.

[0061] The audio to be adjusted is the synthesized audio that has not yet undergone frequency adjustment, and it serves as the input for multi-order frequency adjustment.

[0062] Specifically, the data for each frequency band in the target frequency band data is first time-aligned (e.g., based on timestamps or frame indexes).

[0063] Subsequently, the merging is performed using a synthesis filter bank or an overlap-addition method: if the frequency division uses a filter bank, complementary synthesis filters (such as a cosine modulation filter bank) are used to superimpose the data of each frequency band in the aligned target frequency band data to ensure a smooth transition at the frequency band edges; if the frequency division uses FFT, the data of each frequency band in the aligned target frequency band data are merged into a complete spectrum, and then converted into a time domain signal by IFFT.

[0064] During audio synthesis, window functions (such as the Hanning window) can be applied to reduce boundary effects. The synthesized audio is the audio to be adjusted and is cached in memory for subsequent adjustments.

[0065] S5: Based on the target cavity characteristic data of the sound-generating device, at least two EQ adjustment modules for multi-level frequency adjustment are used to adjust the frequency of the audio to be adjusted to obtain the target audio suitable for the sound-generating device.

[0066] Understandably, after obtaining the audio to be adjusted, the target adjustment configuration is first determined based on the target cavity characteristic data, and then the multi-stage frequency adjustment steps of S5 are executed.

[0067] Multi-order frequency adjustment refers to using multiple cascaded EQ adjustment modules to progressively optimize audio frequency balance, coherence, and detail. The order of multi-order frequency adjustments can be adjusted according to the configuration, for example, performing coherence adjustment first and then balance adjustment, or it may not be limited to three orders.

[0068] Specifically, the target cavity characteristic data of the sound-generating device is first loaded and parsed into EQ parameters (such as center frequency, gain value, and Q value). The Q value refers to the quality factor. Subsequently, at least two EQ adjustment modules are cascaded for processing.

[0069] The first EQ adjustment module: Based on the first adjustment configuration, it performs first-order frequency adjustment on the audio to be adjusted to compensate for frequency balance. For example, it attenuates the gain at the resonant frequency (e.g., -3dB) and increases the gain in the recessed frequency band (e.g., +2dB) to output the first audio.

[0070] The second EQ adjustment module: Based on the second adjustment configuration, it performs second-order frequency adjustment on the first audio, enhancing coherence and smoothness. For example, a wide-bandwidth EQ is set in the frequency transition region (such as the low-frequency-mid-frequency boundary) to smooth the frequency curve and output the second audio.

[0071] The third EQ adjustment module: Based on the third adjustment configuration, it performs third-order frequency adjustment on the second audio to remove residual glitches and spikes. For example, it scans for subtle peaks across the entire frequency range and applies a high-Q notch filter to attenuate them, outputting the target audio.

[0072] Residual spikes and bumps: These refer to local frequency anomalies remaining after the first two levels of adjustment, such as peaks caused by harmonic distortion or resonance. For example, a subtle peak at 3kHz caused by resonance of the sound-generating device, or artificial artifacts introduced by the noise reduction algorithm.

[0073] The parameters of the EQ adjustment module are dynamically determined based on target cavity characteristic data, scenario description data (such as usage environment) and computing resource consumption data.

[0074] It is understandable that the number of EQ adjustment modules can be 2, 3, 4, 5, 6, or more than 6, and there is no limit to this.

[0075] In other words, this application achieves physical optimization of the audio frequency curve through multi-stage frequency adjustment and cavity characteristic compensation.

[0076] Understandably, after obtaining the target audio, the target audio is converted into an analog signal, and the sound-producing device is controlled according to the analog signal to play the sound.

[0077] This embodiment acquires initial audio, divides it into frequencies to obtain initial frequency band data, which includes low-frequency, mid-frequency, and high-frequency data. The initial frequency band data is then denoised to obtain target frequency band data. This target frequency band data is then synthesized to obtain the audio to be adjusted. Based on the target cavity characteristic data of the sound-producing device, at least two EQ adjustment modules for multi-level frequency adjustment are used to adjust the frequency of the audio to be adjusted, resulting in the target audio suitable for the sound-producing device. In other words, after dividing, denoising, and synthesizing the initial audio to obtain the audio to be adjusted, this application uses at least two EQ adjustment modules for multi-level frequency adjustment based on the target cavity characteristic data of the sound-producing device. Multi-level frequency adjustment can progressively optimize the frequency balance of the audio, enhance coherence and smoothness, and remove residual spikes and bumps, thereby effectively avoiding artificial effects at frequency band edges. This results in a more continuous and smoother frequency curve in the final output audio, making the audio more in line with the user's auditory requirements.

[0078] In one embodiment, the number of EQ adjustment modules is three. The step of using at least two EQ adjustment modules for multi-order frequency adjustment to adjust the frequency of the audio to be adjusted based on the target cavity characteristic data of the sound-producing device to obtain the target audio suitable for the sound-producing device includes: S511: Based on the first adjustment configuration in the target adjustment configuration, the first EQ adjustment module is used to perform first-order frequency adjustment on the audio to be adjusted to obtain the first audio. Each EQ adjustment module has an adjustment configuration in the target adjustment configuration.

[0079] Understandably, the target adjustment configuration is determined before performing step S511.

[0080] First-order adjustment configuration: This refers to the set of parameters used for first-order frequency adjustment, including center frequency, gain value, and quality factor (Q value), etc., which aim to optimize the frequency balance of the audio, that is, to compensate for the inherent frequency response imbalance of the sound-producing device (such as formants or recessed frequency bands). For example, a center frequency of 200Hz, a gain value of -3dB, and a Q value of 2 are used to attenuate low-frequency formants; or a center frequency of 1kHz, a gain value of +2dB, and a Q value of 1 are used to boost the mid-frequency recessed range.

[0081] Specifically, the first adjustment configuration in the target adjustment configuration is loaded, and parameters such as center frequency, gain value, and Q value are resolved. The first EQ adjustment module dynamically generates the corresponding IIR filter coefficients (filter coefficients of an infinite impulse response digital filter) based on the resolved parameters. For example, using a dual second-order filter design formula, the filter coefficients (such as the numerator and denominator coefficients of a direct type II structure) are calculated based on the center frequency, gain, and Q value.

[0082] The audio to be adjusted is input into the first EQ adjustment module, and the audio is processed in real time using the filter corresponding to the first EQ adjustment module. Specifically, time-domain convolution or frequency-domain filtering algorithms are used to adjust the gain at a specified frequency point (also known as the adjustment frequency point) (such as attenuating formants or boosting concave frequency bands); the processed audio is used as the first audio output and stored in memory.

[0083] S512: Based on the second adjustment configuration in the target adjustment configuration, the second EQ adjustment module is used to perform second-order frequency adjustment on the first audio to obtain the second audio. Second adjustment configuration: refers to the set of parameters used for second-order frequency adjustment, including at least two adjustment frequencies (different from the adjustment frequencies of the first adjustment configuration), which are designed to enhance the coherence and smoothness of the audio, i.e. reduce artificial effects at the edges of frequency band transitions (such as "cliffs" or "tails").

[0084] Second audio: refers to the intermediate audio signal after second-order frequency adjustment, which has a smoother frequency band transition and improved continuity.

[0085] Example of the second adjustment configuration: For example, the adjustment frequency points include 500Hz and 2kHz, with a wide bandwidth (low Q value, such as 0.7) and a gain of +1dB, used to smooth the low-frequency-mid-frequency boundary; or the adjustment frequency point is 4kHz, with a gain of -2dB, used to suppress high-frequency harshness.

[0086] Specifically, the second EQ adjustment module loads the second adjustment configuration and generates the corresponding filter coefficients. The second EQ adjustment module applies multi-frequency equalization processing to the first audio, for example, by using multiple parallel IIR filter cores to perform wide-bandwidth gain adjustment at specified frequency points to smooth frequency band transitions; the processed audio is then output as the second audio.

[0087] S513: Based on the third adjustment configuration in the target adjustment configuration, the third EQ adjustment module is used to perform third-order frequency adjustment on the second audio to obtain the target audio; The third adjustment configuration refers to the set of parameters used for third-order frequency adjustment. The adjustment frequency points cover the entire sound frequency range (20Hz-20kHz), and the frequency points are different from those of the first and second adjustment configurations. It aims to remove residual spikes and bumps in the audio, that is, subtle distortion or resonance peaks.

[0088] Third adjustment configuration example: For example, adjust the frequency points to 8kHz and 12kHz, with a high Q value (e.g., 10) and a gain of -1dB to attenuate high-frequency glitches; or set multiple high-Q notch filters after full-band scanning.

[0089] Specifically, the third EQ adjustment module loads the third adjustment configuration and generates high-Q filter coefficients. This third EQ adjustment module performs a full-band scan and processing of the second audio signal, for example, using a notch filter to attenuate peaks at specific frequencies or using a multi-band equalizer to fine-tune the full-band response; the processed signal is then output as the target audio signal. In implementation, adaptive filtering technology can be used to adjust parameters in real time to ensure a high degree of match between the output audio and the cavity characteristics of the sound-producing device.

[0090] Specifically, the target adjustment configuration is determined based on the target audio characteristics of the initial audio, the target cavity characteristic data, the scene description data of the sound-generating device, and the computational resource consumption data.

[0091] Target audio characteristics refer to the key feature parameters obtained after analyzing the initial audio, used to guide frequency adjustment. Target audio characteristics include: audio type, spectral characteristics, dynamic range, or loudness information. Audio type: such as "speech," "classical music," "rock music," "podcast." Spectral characteristics: such as energy concentrated mainly in the mid-low frequencies (indicating a need for enhanced clarity), or rich high-frequency overtones (indicating a need for fine-tuning to avoid harshness). Dynamic range: such as large (classical music) or small (compressed pop music). Loudness information: overall average loudness or instantaneous peak loudness.

[0092] Scene description data describes information about the environment and usage scenario during audio playback. Scene description data includes: ambient noise, usage patterns, or user preferences. Ambient noise: such as "quiet indoors" (<30dB), "noisy street" (70dB), "in a moving car" (60dB). Usage patterns: such as "music listening," "movie watching," "gaming," "voice calls." User preferences: such as user-preset "bass enhancement" or "voice prominence" modes.

[0093] Computational resource consumption data represents parameters that characterize the current available processing power of the system. This data includes: CPU / GPU / DSP load, remaining memory, and power budget. CPU / GPU / DSP load: The current processor utilization rate. Remaining memory: The amount of available RAM. Power budget: The maximum computational intensity allowed by the device's current power state or thermal limitations.

[0094] A multi-factor adaptive decision engine (rule engine or lightweight machine learning model) is constructed, which is based on a predefined "if-then-else" rule base. The decision engine acquires and parses target audio characteristics (such as type and spectrum), target cavity characteristic data (such as frequency response curve), scene description data (such as ambient noise), and computational resource consumption data (such as CPU load). First, based on the target audio characteristics (such as music type) and scene description data (such as usage environment), a basic configuration template is selected from the preset library. Then, based on the target cavity characteristic data, precise compensatory parameter corrections are made for the inherent frequency response defects of the sound-producing device (such as formants and depressions). Next, the auditory fine-tuning (such as equal loudness compensation) is performed by combining information such as ambient noise in the scene description data. Finally, the overall configuration is dynamically optimized according to the computational resource consumption data. When resources are tight, the adjustment order or the number of filters is simplified to reduce the load, thereby outputting a customized target adjustment configuration that achieves the optimal balance between sound quality, adaptability, and efficiency.

[0095] This embodiment constitutes the core process of multi-stage frequency adjustment. By cascading three EQ adjustment modules, the frequency balance, coherence, and purity of the audio are gradually optimized. Each step is based on dynamically configured parameters, combined with the physical characteristics of the sound-producing device and the usage scenario, to ensure that the frequency curve of the final output target audio is continuous and smooth, effectively avoiding human effects and improving the listening experience.

[0096] In one embodiment, the step of determining the target adjustment configuration based on target audio characteristics, target cavity characteristic data, scene description data of the sound-generating device, and computational resource consumption data includes: S611: Determine the first adjustment configuration based on the frequency adjustment response curves corresponding to the target audio characteristics, the scene description data, and the target cavity characteristic data, wherein the first adjustment configuration is used to optimize the frequency balance of the audio. Frequency response curve: This refers to the predicted audio output frequency response curve calculated through software simulation based on the target cavity characteristic data (physical model) of the sound-producing device and a set of candidate EQ settings. It is not a fixed ideal target, but a computational tool used to simulate and evaluate the final acoustic effects of different EQ configurations on a specific sound-producing device.

[0097] The frequency adjustment response curve is obtained by an acoustic simulation engine. This engine takes target cavity characteristic data (such as the impedance curve, resonant frequency, and cabinet transfer function of the loudspeaker unit) and candidate EQ adjustment configurations (i.e., the set of filter parameters to be evaluated) as input. It simulates the sound pressure level-frequency curve by solving physical equations (such as equivalent circuit models) or by using numerical methods (such as finite element analysis), and finally outputs a predicted sound pressure level-frequency curve, i.e., the frequency adjustment response curve.

[0098] Specifically, the decision engine initially generates one or more candidate first adjustment configurations based on the target audio characteristics and scene description data (e.g., several different attenuation schemes for low-frequency resonance); each candidate configuration is input into the acoustic simulation engine along with the target cavity characteristic data. The engine calculates a corresponding frequency adjustment response curve (i.e., a predicted curve) for each candidate configuration; these predicted curves are compared with a pre-stored "target frequency curve" that conforms to the ideal listening experience (e.g., calculating the fit or error between the curves). The candidate configuration that makes the predicted curve closest to the ideal target curve is selected as the final determined first adjustment configuration.

[0099] The acoustic simulation engine is a digital signal processing algorithm based on an equivalent circuit model. It runs at the audio processing end, taking as input candidate EQ adjustment configurations (i.e., a set of filter parameters) and target cavity characteristic data of the sound-generating device (this data is pre-stored in the device in the form of Thiele-Small parameters of the speaker unit, cabinet volume, sound tube dimensions, etc.). The core algorithm of the engine models the physical speaker system as an equivalent network composed of resistors, capacitors, and inductors, where the component values ​​are mapped from the cavity characteristic data. When a candidate EQ configuration (which can be considered a transfer function) is input, the engine calculates the comprehensive frequency response from the electrical signal input to the sound pressure output by cascading convolution of the EQ transfer function with the equivalent circuit transfer function of the speaker system in the digital domain. This calculation process is completed quickly by solving linear system equations or using frequency domain multiplication, ultimately outputting a predicted frequency adjustment response curve (i.e., a curve of sound pressure level changing with frequency), which the decision engine uses to evaluate the acoustic effects of different EQ configurations.

[0100] S612: Determine the second adjustment configuration based on auditory sensitivity data, adjustment frequency points in the first adjustment configuration, target audio characteristics, scene description data, frequency adjustment response curve, and computational resource consumption data. The second adjustment configuration includes at least two adjustment frequency points, and the adjustment frequency points in the second adjustment configuration are different from those in the first adjustment configuration. The second adjustment configuration is used to enhance the coherence and smoothness of the audio. Auditory sensitivity data: This data describes the sensitivity of the human ear to sounds at different frequencies and is an important basis for optimizing hearing, such as equal-loudness curves. Auditory sensitivity data is integrated into the decision engine in the form of pre-stored data tables or mathematical models.

[0101] Specifically, the decision engine reads the adjustment frequencies in the first adjustment configuration to ensure that the frequencies selected in the second adjustment configuration are different from and do not overlap with them. The decision engine analyzes the frequency adjustment response curve, paying particular attention to the smoothness of the transition between frequency bands (such as low-frequency to mid-frequency, mid-frequency to high-frequency) after the first configuration correction. In these transition regions, filters with wide bandwidth (low Q value) and gradual gain changes are set. Combining auditory sensitivity data (such as appropriately boosting low and high frequencies at low volume listening), the smooth transition configuration is fine-tuned, for example, by increasing the low-frequency and high-frequency gain in low-volume environments. Finally, based on computational resource consumption data, it decides how many smoothing filters to ultimately enable, or to merge adjustment points when resources are limited.

[0102] S613: Determine the third adjustment configuration based on the adjustment frequency points in the first adjustment configuration, the adjustment frequency points in the second adjustment configuration, the target audio characteristics, the scene description data, the frequency adjustment response curve, and the computational resource consumption data. The adjustment frequency points in the third adjustment configuration are different from those in the first and second adjustment configurations. The adjustment frequency points in the third adjustment configuration cover the entire sound frequency band. The third adjustment configuration is used to remove residual glitches and bumps in the audio.

[0103] Residual glitches and bumps: These refer to distortion peaks or resonances in audio with a very narrow frequency range, usually introduced by device harmonic distortion or denoising algorithms.

[0104] Specifically, the system ensures that the frequency points selected in the third adjustment configuration are different from all frequency points in the first and second adjustment configurations to avoid adjustment conflicts and duplication. Within the entire frequency band (20Hz-20kHz) excluding already occupied frequencies, a more refined algorithm is used to analyze the frequency adjustment response curve or harmonic distortion curve, identifying residual, uncorrected sharp peaks (glitches) and valleys (bumps in this context refer to bulges on the frequency response curve). A high-Q, narrow-bandwidth filter parameter is configured for each identified glitch / bump. Based on computational resource consumption data, the system prioritizes processing the most significant or most sensitive frequencies, and performs more comprehensive processing only when resources permit.

[0105] In this embodiment, before physical playback, software simulation (S611) is used to pre-predict the final acoustic result (i.e., frequency response curve) after the interaction between different EQ configurations and cavity characteristics. This allows the system to intelligently select or generate an adjustment configuration that is not only reasonable at the electrical signal level but also synergizes well with the physical characteristics of the sound-producing device. Based on this, S612 and S613 then perform smooth transitions and fine-tuning in sequence, ensuring that each level of adjustment is based on an accurate prediction of the final output result. This fundamentally avoids blind adjustment, resulting in an unprecedented level of matching between the output target audio and the sound-producing device, achieving the best listening experience with a continuously smooth frequency curve and no distortion.

[0106] In one embodiment, the step of using a second EQ adjustment module to perform second-order frequency adjustment on the first audio based on the second adjustment configuration in the target adjustment configuration to obtain the second audio includes: S5121: Evaluate the adjustment effect based on the evaluation index corresponding to the first audio and the first adjustment configuration to obtain a first result; Evaluation metrics: These are pre-defined parameters or algorithms used to quantify the effectiveness of frequency regulation. Examples include frequency response smoothness, signal-to-noise ratio (SNR), or auditory perception score. Smoothness is measured using the standard deviation of the frequency response curve, or an auditory perception score is calculated using a perception evaluation model. Evaluation metrics are stored in the system's configuration library as predefined business rules or mathematical models. Different initial regulation configurations may be associated with different core evaluation metrics (e.g., for a configuration used to suppress resonance, the core metric is the attenuation of the resonant peak).

[0107] First result: refers to one or more quantitative evaluation scores or status indicators output after comparing the first audio with the evaluation index, used to characterize the actual effect of the first-order adjustment.

[0108] Specifically, the evaluation module performs real-time spectral analysis on the first audio signal (e.g., using a Short-Time Fourier Transform, STFT) to obtain its current frequency response characteristics. The evaluation module then calls an evaluation index algorithm corresponding to the first adjustment configuration. For example, if the first adjustment configuration aims to smooth low frequencies, it calculates the frequency response smoothness index for the 80Hz-300Hz frequency band. The calculated value is compared with a preset threshold or ideal value to generate a quantized first result. This first result indicates in which aspects the first-order adjustment has met expectations and in which aspects it is insufficient.

[0109] The evaluation module is implemented in software as a real-time spectrum analysis and index calculation unit. After each EQ adjustment (e.g., after obtaining the first audio signal), the module extracts a segment of the processed audio signal and performs a Short-Time Fourier Transform (STFT) to obtain its current frequency domain representation. Subsequently, the module calls the evaluation index algorithm corresponding to the adjustment target of that level for calculation. For example, for the first-level "frequency balance" optimization, the algorithm calculates the standard deviation of the frequency response within the 80Hz-300Hz band to quantify smoothness; for the second-level "coherence," it calculates the cross-correlation of the spectrum in the transition region of adjacent frequency bands (such as low frequency and mid frequency). These calculated quantized values ​​(i.e., "first result," "second result") are compared with preset ideal thresholds or target curves to generate an evaluation conclusion containing labels such as "pass," "fail," or "degree of optimization required." This conclusion is sent as a feedback signal to the decision engine to trigger the dynamic update rules for the next-level adjustment configuration.

[0110] S5122: Update the second adjustment configuration based on the first result; Specifically, the first result is analyzed to identify problem areas requiring further optimization (such as "insufficient smoothness in frequency band X") and areas that have been well-processed; an "if-then" type optimization rule is invoked. For example, IF "low-frequency smoothness" < threshold THEN adds a smoothing filter with Q=1.0 to the transition frequency band (such as 150Hz) in the second configuration. Based on the rule output, the engine dynamically modifies the parameters in the second adjustment configuration.

[0111] S5123: Based on the updated second adjustment configuration, the second EQ adjustment module is used to perform second-order frequency adjustment on the first audio to obtain the second audio; Specifically, the second EQ adjustment module loads the updated second adjustment configuration and generates the corresponding filter coefficients. The second EQ adjustment module applies multi-frequency equalization processing to the first audio, for example, using multiple parallel IIR filter cores to perform wide-bandwidth gain adjustment at specified frequency points to smooth frequency band transitions; the processed audio is then output as the second audio.

[0112] The step of using a third EQ adjustment module to perform third-order frequency adjustment on the second audio based on the third adjustment configuration in the target adjustment configuration to obtain the target audio includes: S5131: Evaluate the adjustment effect based on the evaluation index corresponding to the second audio and the second adjustment configuration to obtain a second result; Specifically, the evaluation module performs real-time spectral analysis on the second audio (e.g., using a Short-Time Fourier Transform, STFT) to obtain its current frequency response characteristics. The evaluation module then calls an evaluation index algorithm corresponding to the second adjustment configuration to generate a quantized second result for the second audio. This second result indicates in which aspects the second-order adjustment has met expectations and in which aspects it is insufficient.

[0113] S5132: Update the third adjustment configuration based on the second result; Specifically, the second result is analyzed to identify problem areas that need further optimization and areas that have been well processed; "if-then" type optimization rules are invoked; and based on the rule output, the engine dynamically modifies the parameters in the third adjustment configuration.

[0114] S5133: Based on the updated third adjustment configuration, the third EQ adjustment module is used to perform third-order frequency adjustment on the second audio to obtain the target audio.

[0115] Specifically, the third EQ adjustment module loads the updated third adjustment configuration and generates high-Q filter coefficients. This third EQ adjustment module performs a full-band scan and processing of the second audio signal, for example, using a notch filter to attenuate peaks at specific frequencies or using a multi-band equalizer to fine-tune the full-band response; the processed signal is then output as the target audio signal. In implementation, adaptive filtering technology can be used to adjust parameters in real time to ensure a high degree of match between the output audio and the cavity characteristics of the sound-producing device.

[0116] The closed-loop feedback and dynamic update mechanism in this embodiment enables intelligent and adaptive multi-stage frequency adjustment. By immediately evaluating the effect after each adjustment (S5121, S5131) and using the evaluation results (first result, second result) as input to dynamically optimize the adjustment configuration for the next stage (S5122, S5132), subsequent EQ adjustment modules (second and third) can perform targeted compensation and fine-tuning based on the actual effect of the previous stage processing. This "evaluation-feedback-update-execution" closed-loop control effectively overcomes the blindness of static adjustment strategies, can correct deviations in real time, and accurately eliminate residual problems from the previous stage, thereby significantly improving the overall collaborative efficiency of multi-stage adjustments and the coherence, smoothness, and purity of the final output audio.

[0117] In one embodiment, the step of determining the target adjustment configuration based on the target audio characteristics of the initial audio, the target cavity characteristic data, the scene description data of the sound-generating device, and the computational resource consumption data includes: S621: Based on the ambient noise data corresponding to the sound-generating device, the power consumption configuration corresponding to the sound-generating device, the target audio characteristics, the target cavity characteristic data, the scene description data of the sound-generating device, and the computing resource consumption data, determine the target adjustment configuration.

[0118] Ambient noise data: refers to the acoustic characteristics of the background noise in the environment where the sound-generating equipment is located, typically including its total sound pressure level (SPL) and spectral distribution (i.e., noise energy at different frequencies). Active acquisition: Real-time ambient sound is collected through the microphone built into the sound-generating equipment or its associated smart terminal (such as a mobile phone), and then the noise spectrum is obtained through Fast Fourier Transform (FFT) analysis. Indirect inference: Typical ambient noise model data for that scenario is retrieved from a pre-set database based on scene description data (such as the user-selected "subway mode").

[0119] Power Consumption Configuration: Refers to the constraints or strategies set for the current or expected power consumption of the audio processing system, typically related to the power status of the sound-generating device, its heat dissipation capacity, or user-defined energy-saving modes. System Status Reading: Reads the current battery level, remaining battery life, or currently active power consumption mode (such as "energy-saving mode") from the power management unit (PMU) or operating system of the sound-generating device. Policy Library Query: Based on the read system status, queries a predefined "power consumption-performance" policy library to determine the current audio processing power consumption limit or optimization strategy to be adopted.

[0120] Specifically, the decision engine receives all input data, including: ambient noise data (real-time spectrum), power consumption configuration (current power consumption strategy), target audio characteristics (such as music genre), target cavity characteristic data (device acoustic fingerprint), scene description data (such as "outdoor sports"), and computational resource consumption data (such as CPU load). The engine prioritizes processing ambient noise data. Its core logic is to perform "equal loudness compensation" and "masking effect avoidance." For example, in a noisy environment, it automatically increases the gain of low frequencies (enhancing dynamics) and high frequencies (enhancing clarity) in the target adjustment configuration to overcome the auditory masking of these frequency bands by ambient noise. The engine processes power consumption configuration simultaneously. If it is currently in low-power mode, the engine will adopt a "performance degradation" strategy: it will simplify all third-order (first, second, and third) adjustment configurations, for example, by reducing the total number of filters used, merging multiple narrowband adjustments with wide-bandwidth filters, or directly shutting down the most computationally complex third-order fine adjustment module. This aims to reduce the computational load on the DSP, thereby reducing the overall system power consumption and extending device battery life. Ultimately, the engine integrates and balances the environmental adaptation strategy and power consumption constraint strategy with the original configuration logic based on cavity characteristics, audio characteristics and other data, and outputs an optimized target adjustment configuration that can ensure good listening experience in the current noise environment and meet the current power consumption budget of the device.

[0121] This embodiment introduces two key variables—ambient noise data and power consumption configuration—enhancing the target adjustment configuration determination process from "optimization based on a fixed scenario" to "comprehensive optimization that dynamically adapts to the environment and conditions." The system proactively senses the acoustic characteristics of the playback environment and automatically compensates for audio noise levels, maintaining clarity and fullness even in noisy conditions. Simultaneously, it responds in real-time to device power and consumption constraints, intelligently simplifying the processing flow to reduce energy consumption while ensuring basic listening quality. This combination ensures that the audio processing system provides optimal, high-quality audio output that balances with device battery life, regardless of whether the environment is quiet indoors or noisy outdoors, or whether the battery is fully charged or in energy-saving mode. This effectively improves the intelligence and adaptability of the user experience.

[0122] In one embodiment, the step of using at least two EQ adjustment modules for multi-order frequency adjustment to adjust the frequency of the audio to be adjusted, based on the target cavity characteristic data of the sound-generating device, to obtain a target audio suitable for the sound-generating device, includes: S521: Based on the target cavity characteristic data and the channel characteristics corresponding to the left channel, at least two EQ adjustment modules for multi-level frequency adjustment are used to adjust the frequency of the left channel of the audio to be adjusted to obtain the left channel audio. The characteristics of the left channel refer to the unique properties associated specifically with the left channel sound-producing device (such as the left loudspeaker or left headphone unit) in a stereo or multi-channel system. This includes not only its inherent target cavity characteristic data, but also its positional characteristics in the sound field (such as its angle and distance from the listener) and subtle performance differences from the right channel due to manufacturing tolerances.

[0123] The characteristics of the left channel include: independent cavity data, position and acoustic parameters, and crosstalk compensation parameters. Independent cavity data: The left channel speaker has its own measured frequency response curve, which may show a slight peak at 8kHz, while the right channel may not have this peak at the same frequency. Position and acoustic parameters: In a home theater system, this refers to the azimuth and distance of the left channel speaker relative to the sweet spot (optimal listening position), and the calculated sound arrival time and intensity attenuation. Crosstalk compensation parameters: These parameters are used to compensate for the small amount of right channel signal crosstalk to the left channel.

[0124] The methods for obtaining the channel characteristics corresponding to the left channel include: inherent characteristics (such as independent cavity data) are measured independently for each channel unit at the factory and the data is stored separately; positional characteristics are automatically measured and calculated through the acoustic calibration process during the initial system setup (such as the user holding a microphone and playing a test tone with the system).

[0125] Specifically, based on the target cavity characteristic data (left) and the corresponding channel characteristics (such as position data) of the left channel, a separate target adjustment configuration is calculated for the left channel. This configuration is calculated separately from the right channel. At least two EQ adjustment modules (e.g., module L1, module L2, module L3) are instantiated for the left channel, forming a multi-order frequency adjustment link dedicated to the left channel. The left channel signal in the audio to be adjusted is separated and input into the left channel processing link. This link performs first-order, second-order, and third-order frequency adjustments step by step according to the left channel target adjustment configuration.

[0126] S522: Based on the target cavity characteristic data and the channel characteristics corresponding to the right channel, at least two EQ adjustment modules for multi-level frequency adjustment are used to adjust the frequency of the right channel of the audio to be adjusted to obtain the right channel audio. The characteristics of the right channel: corresponding to the concept of the left channel, it refers to the unique attributes specifically associated with the right channel sound-producing device.

[0127] Specifically, the implementation logic of this step is completely symmetrical to that of S521, but the operation object is the right channel. An independent processing link (modules R1, R2, R3) is established for the right channel. The right channel target adjustment configuration, which is independently calculated based on the target cavity characteristic data (right) and the corresponding channel characteristics of the right channel, is used to process the right channel signal in the audio to be adjusted, and finally outputs the right channel audio.

[0128] S523: Based on the left channel audio and the right channel audio, obtain the target audio suitable for the sound-generating device.

[0129] Specifically, ensure that the processed left and right channel audio are completely synchronized in time to avoid image drift caused by processing delays. The synchronized left and right channel audio are then reassembled in a stereo format; for example, in the digital domain, two mono signal streams are repackaged into a standard stereo PCM (Pulse Code Modulation) data stream. This reassembled, complete audio, containing independently optimized left and right channel audio, is the final target audio suitable for the entire sound system.

[0130] This embodiment achieves a leap from "overall system optimization" to "precise matching of individual devices" through the independent adjustment technology for the left and right channels implemented in steps S521 to S523. By independently adjusting the frequencies of the left and right channel sound-emitting devices according to their unique cavity characteristics, positions, and other channel characteristics, it can accurately compensate for the inherent frequency response defects of each individual device and those caused by different positions in the sound field. This processing method eliminates problems such as blurred sound image positioning and unbalanced timbre caused by inconsistent performance of the left and right channel devices or asymmetrical acoustic environment. As a result, the final synthesized target audio can construct a stereo sound field with highly consistent frequency response, extremely accurate sound image positioning, and significantly enhanced spatial sense and immersion, effectively improving the realism and presence of high-fidelity stereo reproduction.

[0131] Existing technologies employ an architecture that absolutely divides frequency bands and performs independent noise reduction before resynthesis, which has three inherent drawbacks: First, at the frequency band boundaries, the non-ideal characteristics of the filters easily produce fixed amplitude dips or bulges, resulting in an uneven frequency response in the synthesized audio. Second, for broadband transient signals (such as drum beats), the inability to perfectly align the phases of each frequency band disrupts the transient response, making the sound loose and muffled. Third, the inconsistent artificial effects introduced by independent noise reduction at the edges of each frequency band (such as "cliffs" or "tails") produce a rough "patchwork" feel during synthesis. These problems collectively result in a noticeable sense of disjointedness and unnaturalness in the final output audio, severely impacting the listening experience.

[0132] In one embodiment, the initial frequency band data further includes: low-frequency data, mid-frequency data, high-frequency data, mid-low-frequency data, and mid-high-frequency data; The step of performing audio synthesis on the target frequency band data to obtain the audio to be adjusted includes: S41: Extract the audio features of the initial audio, and generate a fusion auxiliary table based on the audio features, wherein the audio features include zero-crossing rate, energy and spectral centroid; Specifically, the low-to-mid frequency data is low-to-mid frequency related data extracted from the initial audio; the mid-to-high frequency data is mid-to-high frequency related data extracted from the initial audio.

[0133] The first part of the low-frequency data overlaps with part of the low-frequency data, and the second part of the low-frequency data overlaps with part of the medium-frequency data.

[0134] The first part of the mid-to-high frequency data overlaps with a portion of the mid-frequency data, and the second part of the mid-to-high frequency data overlaps with a portion of the high-frequency data.

[0135] Zero-crossing rate: The number of times an audio signal waveform crosses the zero axis per unit time. A high zero-crossing rate usually indicates unvoiced sounds or transient components. Energy: The sum of the squares of the amplitude of the audio signal per unit time, characterizing the signal strength. Spectral centroid: The center frequency point of the spectral energy distribution; a high value indicates high brightness.

[0136] Fusion Auxiliary Table: This refers to a data structure aligned with the audio timeline, recording the specific values ​​of the audio features at each time frame or key time point. It provides dynamic, audio content-related decision-making support for subsequent fusion processing. The fusion auxiliary table is a data sequence, for example: [timestamp: 0.1s, zero-crossing rate: 250, energy: 0.8, spectral centroid: 4500Hz], [timestamp: 0.2s, zero-crossing rate: 50, energy: 0.2, spectral centroid: 800Hz].

[0137] Specifically, the initial audio signal is divided into consecutive short frames, and a window function (such as the Hanning window) is applied to each frame to reduce spectral leakage. For each frame of audio data obtained from the division, its zero-crossing rate, energy, and spectral centroid are calculated in parallel or serially. A time-indexed data structure is created, using the start time or center time of each frame as a timestamp and binding it to the three characteristic values ​​of that frame, ultimately generating a complete fusion auxiliary table.

[0138] S42: Based on the first fusion weight configuration corresponding to the target audio characteristics of the initial audio, perform fusion processing of low-frequency overlapping data according to the fusion auxiliary table, the low-frequency data and the mid-low frequency data to obtain the first fusion data; Fusion weight configuration refers to a set of rules or functions that define how to dynamically adjust the mixing ratio of data from two adjacent frequency bands based on audio characteristics in overlapping frequency band regions. Fusion weight configuration is stored in the system as predefined business logic or mathematical models. Different target audio characteristics (such as "speech" or "music") correspond to different default weight configuration strategies.

[0139] A fusion strategy library is pre-stored. This library pre-defines a corresponding fusion weight configuration (i.e., first, second, third, and fourth configurations) for each target audio characteristic (such as "voice," "classical music," "rock music," etc.). Based on the target audio characteristics obtained from analyzing the initial audio, the matching complete configuration is retrieved from this strategy library and applied to the corresponding fusion processing step.

[0140] The first fusion weight configuration refers to the set of parameters and rules used to guide the fusion processing of overlapping areas between low-frequency and mid-to-low-frequency data. Its core strategy focuses on smooth energy transition and preserving the low-frequency richness of transient signals.

[0141] Fusion processing: refers to the process of weighting and superimposing the signals of two frequency bands within a preset overlapping frequency band of adjacent frequency bands according to the weights in the fusion weight configuration, so as to generate a new signal with a smooth transition.

[0142] Specifically, Step 1: Determine the overlapping frequency range of the low-frequency data and the mid-low-frequency data (this range is actually low frequency). Extract frequency band data A from the low-frequency data based on this range, and extract frequency band data B from the mid-low-frequency data based on this range. Step 2: For each time frame within the overlapping region, query the fusion auxiliary table to obtain the audio features (zero-crossing rate, energy, spectral centroid) of that frame. Input the audio features into the rules in the first fusion weight configuration to calculate the respective mixing weights (Weight_A, Weight_B) of frequency band data A and frequency band data B at that moment, ensuring that Weight_A + Weight_B = 1. Step 3: Multiply the signals from frequency band data A and frequency band data B in the overlapping band within the same time frame by the calculated Weight_A and Weight_B respectively, and then add the results to obtain the fused signal of that frame in the overlapping band. Repeat steps 2 and 3 for all time frames to finally generate the first fused data covering the entire overlapping frequency band and time length.

[0143] Example of configuring the first fusion weight: Rule 1: IF Energy_t is greater than the threshold TH_high AND Spectral_Centroid_t is less than the threshold TH_low THEN [determined as a strong low-frequency part, such as a bass drum] -> set low-frequency weight = 0.8, mid-low frequency weight = 0.2; Rule 2: IF ZCR_t is greater than threshold TH_zcr AND Energy_t is less than threshold TH_medium THEN [determined as transient or friction noise] -> Set low-frequency weight = 0.3, mid-low frequency weight = 0.7 to retain transient details that are more likely to exist in mid-low frequencies; Default rule: ELSE -> Set low frequency weight = 0.5, mid-low frequency weight = 0.5 to achieve a smooth transition.

[0144] ZCR_t is the zero-crossing rate, Energy_t is the energy, and Spectral_Centroid_t is the spectral centroid.

[0145] S43: Based on the second fusion weight configuration corresponding to the target audio characteristics, perform fusion processing of the mid-frequency overlapping data according to the fusion auxiliary table, the mid-frequency data and the mid-low frequency data to obtain the second fusion data; The second fusion weight configuration refers to the set of parameters and rules used to guide the fusion processing of overlapping areas between mid-frequency and low-frequency data. Its core strategy focuses on the coherence and clarity of vocal and basic instrument timbres.

[0146] Specifically, Step 1: Determine the overlapping frequency range of the intermediate frequency (IF) data and the mid-low frequency (MSF) data (this range is actually the IF). Based on this range, extract data from the IF data to obtain frequency band data C, and extract data from the MSF data to obtain frequency band data D. Step 2: For each time frame within the overlapping region, query the fusion auxiliary table to obtain the audio features (zero-crossing rate, energy, spectral centroid) of that frame. Input the audio features into the rules in the second fusion weight configuration to calculate the respective mixing weights (Weight_C, Weight_D) of frequency band data C and frequency band data D at that moment, ensuring that Weight_C + Weight_D = 1. Step 3: Multiply the signals from frequency band data C and frequency band data D in the overlapping band within the same time frame by the calculated Weight_C and Weight_D, respectively, and then add the results to obtain the fused signal of that frame in the overlapping band. Repeat steps 2 and 3 for all time frames to finally generate second fused data covering the entire overlapping frequency band and time length.

[0147] S44: Based on the third fusion weight configuration corresponding to the target audio characteristics, perform fusion processing of the intermediate frequency overlapping data according to the fusion auxiliary table, the intermediate frequency data and the intermediate high frequency data to obtain the third fusion data; The third fusion weight configuration refers to the set of parameters and rules used to guide the fusion processing of overlapping areas between mid-frequency and mid-high-frequency data. Its core strategy focuses on the natural representation of instrument overtones, details, and brightness.

[0148] Specifically, Step 1: Determine the overlapping frequency range of the intermediate frequency (IF) data and the mid-to-high frequency (MCF) data (this range is actually the IF). Based on this range, extract data from the IF data to obtain frequency band data E, and extract data from the MCF data to obtain frequency band data F. Step 2: For each time frame within the overlapping region, query the fusion auxiliary table to obtain the audio features (zero-crossing rate, energy, spectral centroid) of that frame. Input the audio features into the rules in the third fusion weight configuration to calculate the respective mixing weights (Weight_E, Weight_F) of frequency band data E and frequency band data F at that moment, ensuring that Weight_E + Weight_F = 1. Step 3: Multiply the signals from frequency band data E and frequency band data F in the overlapping band within the same time frame by the calculated Weight_E and Weight_F respectively, and then add the results to obtain the fused signal of that frame in the overlapping band. Repeat steps 2 and 3 for all time frames to finally generate third fused data covering the entire overlapping frequency band and time length.

[0149] S45: Based on the fourth fusion weight configuration corresponding to the target audio characteristics, perform high-frequency overlapping data fusion processing according to the fusion auxiliary table, the high-frequency data and the mid-high frequency data to obtain the fourth fusion data; The fourth fusion weight configuration refers to the set of parameters and rules used to guide the fusion processing of overlapping areas between high-frequency and mid-to-high-frequency data. Its core strategy focuses on preserving the atmospheric feel and details of ultra-high frequencies while suppressing harsh noise.

[0150] Specifically, Step 1: Determine the overlapping frequency range of the high-frequency data and the mid-high-frequency data (this range is actually the mid-frequency). Based on this range, extract frequency band data G from the high-frequency data and frequency band data H from the mid-high-frequency data. Step 2: For each time frame within the overlapping region, query the fusion auxiliary table to obtain the audio features (zero-crossing rate, energy, spectral centroid) of that frame. Input the audio features into the rules in the fourth fusion weight configuration to calculate the respective mixing weights (Weight_G, Weight_H) of frequency band data G and frequency band data H at that moment, ensuring that Weight_G + Weight_H = 1. Step 3: Multiply the signals from frequency band data G and frequency band data H in the overlapping band within the same time frame by the calculated Weight_G and Weight_H respectively, and then add the results to obtain the fused signal of that frame in the overlapping band. Repeat steps 2 and 3 for all time frames to finally generate fourth fused data covering the entire overlapping frequency band and time length.

[0151] S46: Audio synthesis is performed based on the low-frequency data, the mid-low-frequency data, the mid-frequency data, the first fusion data, the second fusion data, the third fusion data, and the fourth fusion data to obtain the audio to be adjusted.

[0152] Specifically, the original frequency band data (the low-frequency data, the mid-low-frequency data, and the mid-frequency data) and all newly generated fused data (the first to fourth fused data) are precisely aligned on the time and frequency axes. Frequency components in each original frequency band that are not in any overlapping region are directly retained. Frequency components in each original frequency band data that are in overlapping regions are replaced with the corresponding fused data. For example, in the overlapping band of low frequency and mid-low frequency, the original low-frequency data or mid-low-frequency data is no longer used; instead, the first fused data is used. The retained non-overlapping region data and the replaced fused data are superimposed in the frequency domain or time domain to synthesize a complete, full-frequency band audio signal, which is the audio to be adjusted.

[0153] Understandably, noise reduction primarily eliminates noise, while fusion processing ensures a smooth transition between frequency bands.

[0154] This embodiment extracts audio features and generates a fusion auxiliary table (S41), enabling the system to perceive dynamic changes in audio (such as transient moments) in real time. Then, within a preset frequency band overlap area, fusion weights are dynamically calculated based on these features, and signal fusion is performed (S42-S45). This process achieves an adaptive and smooth transition of amplitude response at frequency band boundaries, completely eliminating fixed amplitude distortion. Simultaneously, the dynamic weighting mechanism ensures natural coordination of energy and phase between frequency bands in the transient signal, restoring the integrity of the transient. Finally, the fused data replaces the original overlap area data for synthesis (S46), erasing the artificial traces left at the edges by independent processing of each frequency band (especially noise reduction). These three aspects work synergistically to ultimately result in a highly continuous and smooth frequency curve in the synthesized audio to be adjusted, fundamentally eliminating the "discontinuity" and "patchwork" feel, laying a high-quality foundation for subsequent multi-order frequency adjustment.

[0155] The audio processing method based on multi-order frequency adjustment proposed in this application is mainly applied to scenarios with high requirements for immersive sound experience, such as home audio-visual, personal entertainment, virtual reality, and in-vehicle sound field.

[0156] In home theaters or high-end audio-visual rooms, users seek an immersive and realistic experience, as if they were actually there. This system, based on the cavity characteristics and spatial layout of different speakers, precisely optimizes the frequency curves of each channel's audio through multi-stage frequency adjustment. For example, it enhances the depth and purity of the subwoofer, making explosions more impactful; simultaneously, it smooths the frequency response of the mid-high frequency speakers, resulting in clear and natural dialogue, rich instrumental details, and seamless transitions between frequency bands, effectively avoiding harshness or muddiness. This creates a balanced, full, and highly layered immersive sound field, providing users with an auditory feast comparable to a professional cinema.

[0157] Immersive sound is crucial when using headphones for music enjoyment or gaming. Based on the target cavity characteristics of the headphones (sound-producing devices), multi-level frequency adjustments can be made to the audio. This not only compensates for the inherent limitations of headphones in certain frequency bands but also provides refined processing for environmental sound effects, directional footsteps, and background music in games, enhancing the sense of sound localization and spatiality. When enjoying high-fidelity music, it can reproduce the atmosphere and details of the recording session, making the listener feel as if they are in a concert hall, achieving a deeply immersive auditory experience.

[0158] In virtual reality (VR) or augmented reality (AR) devices, immersive sound is a crucial element in creating a sense of realism. This application adapts to the unique cavity characteristics of compact speakers in VR devices (sound-emitting devices), and through multi-level frequency adjustment, eliminates audio glitches and resonances that may arise due to the physical limitations of the device, ensuring a smooth and natural sound across the entire audio frequency range, from the impact of low frequencies to the crispness of high frequencies. This allows for precise synchronization between sound and visual changes in the virtual world, resulting in a more realistic sense of direction and distance, greatly enhancing the user's sense of presence and immersion in the metaverse or simulation training.

[0159] Within a car audio system, the confined space and complex interior materials of the vehicle cabin easily create irregular sound fields and standing waves. This application intelligently adjusts the audio signal based on the characteristics of speakers in different locations within the vehicle and the overall acoustic environment of the cabin (target cavity characteristic data). It can specifically smooth out frequency response spikes and dips caused by sound wave reflection and interference, purifying sound quality and creating a balanced, clear, and dynamic immersive sound field for the driver or all passengers, providing a surround sound experience as if they were on stage, whether listening to symphonies or pop music. Understandably, in a car setting, multi-order frequency adjustment compensates for acoustic defects within the cabin.

[0160] Please see Figure 3 As shown, in one embodiment, an audio processing system based on multi-order frequency adjustment is provided, the system comprising: an audio processing terminal 901 and a sound-generating device 902; The audio processing terminal 901 is configured to implement the audio processing method based on multi-order frequency adjustment as described in the first aspect; The sound-generating device is used to play the target audio output by the audio processing terminal.

[0161] The steps of the audio processing method based on multi-order frequency adjustment include: Get the initial audio; The initial audio is divided into frequencies to obtain initial frequency band data, wherein the initial frequency band data includes: low frequency data, mid frequency data and high frequency data; The initial frequency band data is denoised to obtain the target frequency band data; The target frequency band data is used for audio synthesis to obtain the audio to be adjusted; Based on the target cavity characteristic data of the sound-generating device 902, at least two EQ adjustment modules for multi-level frequency adjustment are used to adjust the frequency of the audio to be adjusted, so as to obtain the target audio suitable for the sound-generating device 902.

[0162] This embodiment acquires initial audio, divides it into frequencies to obtain initial frequency band data, which includes low-frequency, mid-frequency, and high-frequency data. The initial frequency band data is then denoised to obtain target frequency band data. This target frequency band data is then synthesized to obtain the audio to be adjusted. Based on the target cavity characteristic data of the sound-producing device, at least two EQ adjustment modules for multi-level frequency adjustment are used to adjust the frequency of the audio to be adjusted, resulting in the target audio suitable for the sound-producing device. In other words, after dividing, denoising, and synthesizing the initial audio to obtain the audio to be adjusted, this application uses at least two EQ adjustment modules for multi-level frequency adjustment based on the target cavity characteristic data of the sound-producing device. Multi-level frequency adjustment can progressively optimize the frequency balance of the audio, enhance coherence and smoothness, and remove residual spikes and bumps, thereby effectively avoiding artificial effects at frequency band edges. This results in a more continuous and smoother frequency curve in the final output audio, making the audio more in line with the user's auditory requirements.

[0163] In one embodiment, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, performs the following steps: Get the initial audio; The initial audio is divided into frequencies to obtain initial frequency band data, wherein the initial frequency band data includes: low frequency data, mid frequency data and high frequency data; The initial frequency band data is denoised to obtain the target frequency band data; The target frequency band data is used for audio synthesis to obtain the audio to be adjusted; Based on the target cavity characteristic data of the sound-generating device, at least two EQ adjustment modules for multi-level frequency adjustment are used to adjust the frequency of the audio to be adjusted, so as to obtain the target audio suitable for the sound-generating device.

[0164] It is understood that electronic devices can be integrated as modules into sound-generating devices, or sound-generating devices can be used as electronic devices as described in this application, or devices with computing capabilities can be used as sound-generating devices.

[0165] This embodiment acquires initial audio, divides it into frequencies to obtain initial frequency band data, which includes low-frequency, mid-frequency, and high-frequency data. The initial frequency band data is then denoised to obtain target frequency band data. This target frequency band data is then synthesized to obtain the audio to be adjusted. Based on the target cavity characteristic data of the sound-producing device, at least two EQ adjustment modules for multi-level frequency adjustment are used to adjust the frequency of the audio to be adjusted, resulting in the target audio suitable for the sound-producing device. In other words, after dividing, denoising, and synthesizing the initial audio to obtain the audio to be adjusted, this application uses at least two EQ adjustment modules for multi-level frequency adjustment based on the target cavity characteristic data of the sound-producing device. Multi-level frequency adjustment can progressively optimize the frequency balance of the audio, enhance coherence and smoothness, and remove residual spikes and bumps, thereby effectively avoiding artificial effects at frequency band edges. This results in a more continuous and smoother frequency curve in the final output audio, making the audio more in line with the user's auditory requirements.

[0166] In one embodiment, the electronic device of this application is a computer device, which may be a server 120, and its internal structure diagram may be as follows. Figure 4 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external client 110 via a network connection. When the computer program is executed by the processor, it implements the functions or steps of an audio processing method based on multi-order frequency adjustment on the server side 120.

[0167] In one embodiment, a computer device is provided, which may be a client 110, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input system connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of an audio processing method based on multi-order frequency adjustment on the client side 110.

[0168] In one embodiment, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program that, when executed by a processor, performs the following steps: Get the initial audio; The initial audio is divided into frequencies to obtain initial frequency band data, wherein the initial frequency band data includes: low frequency data, mid frequency data and high frequency data; The initial frequency band data is denoised to obtain the target frequency band data; The target frequency band data is used for audio synthesis to obtain the audio to be adjusted; Based on the target cavity characteristic data of the sound-generating device, at least two EQ adjustment modules for multi-level frequency adjustment are used to adjust the frequency of the audio to be adjusted, so as to obtain the target audio suitable for the sound-generating device.

[0169] This embodiment acquires initial audio, divides it into frequencies to obtain initial frequency band data, which includes low-frequency, mid-frequency, and high-frequency data. The initial frequency band data is then denoised to obtain target frequency band data. This target frequency band data is then synthesized to obtain the audio to be adjusted. Based on the target cavity characteristic data of the sound-producing device, at least two EQ adjustment modules for multi-level frequency adjustment are used to adjust the frequency of the audio to be adjusted, resulting in the target audio suitable for the sound-producing device. In other words, after dividing, denoising, and synthesizing the initial audio to obtain the audio to be adjusted, this application uses at least two EQ adjustment modules for multi-level frequency adjustment based on the target cavity characteristic data of the sound-producing device. Multi-level frequency adjustment can progressively optimize the frequency balance of the audio, enhance coherence and smoothness, and remove residual spikes and bumps, thereby effectively avoiding artificial effects at frequency band edges. This results in a more continuous and smoother frequency curve in the final output audio, making the audio more in line with the user's auditory requirements.

[0170] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions of the server 120 side and the client 110 side in the aforementioned method embodiments. To avoid repetition, they will not be described one by one here.

[0171] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0172] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.

[0173] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for audio processing based on multi-order frequency adjustment, characterized in that, The method includes: Get the initial audio; The initial audio is divided into frequencies to obtain initial frequency band data, wherein the initial frequency band data includes: low frequency data, mid frequency data and high frequency data; The initial frequency band data is denoised to obtain the target frequency band data; The target frequency band data is used for audio synthesis to obtain the audio to be adjusted; Based on the target cavity characteristic data of the sound-generating device, an EQ adjustment module with three multi-stage frequency adjustment is used to adjust the frequency of the audio to be adjusted, so as to obtain the target audio suitable for the sound-generating device. The step of using a three-stage multi-frequency EQ adjustment module to adjust the frequency of the audio to be adjusted, based on the target cavity characteristic data of the sound-generating device, to obtain the target audio suitable for the sound-generating device includes: a) Based on the first adjustment configuration in the target adjustment configuration, the first EQ adjustment module is used to perform first-order frequency adjustment on the audio to be adjusted to obtain the first audio; Based on the second adjustment configuration in the target adjustment configuration, the second EQ adjustment module is used to perform second-order frequency adjustment on the first audio to obtain the second audio. Based on the third adjustment configuration in the target adjustment configuration, the third EQ adjustment module is used to perform third-order frequency adjustment on the second audio to obtain the target audio. or b) Determine the target adjustment configuration based on the target audio characteristics of the initial audio, the target cavity characteristic data, the scene description data of the sound-generating device, and the computational resource consumption data; Based on the target cavity characteristic data and the channel characteristics corresponding to the left channel, at least two EQ adjustment modules for multi-order frequency adjustment are used to adjust the frequency of the left channel of the audio to be adjusted, so as to obtain the left channel audio. Based on the target cavity characteristic data and the channel characteristics corresponding to the right channel, at least two EQ adjustment modules for multi-order frequency adjustment are used to adjust the frequency of the right channel of the audio to be adjusted, so as to obtain the right channel audio. Based on the left channel audio and the right channel audio, a target audio suitable for the sound-generating device is obtained.

2. The audio processing method based on multi-order frequency adjustment according to claim 1, characterized in that, The step of determining the target adjustment configuration based on the target audio characteristics, the target cavity characteristic data, the scene description data of the sound-generating device, and the computational resource consumption data includes: Based on the frequency adjustment response curves corresponding to the target audio characteristics, the scene description data, and the target cavity characteristic data, the first adjustment configuration is determined, wherein the first adjustment configuration is used to optimize the frequency balance of the audio. Based on auditory sensitivity data, the adjustment frequency points in the first adjustment configuration, the target audio characteristics, the scene description data, the frequency adjustment response curve, and the computational resource consumption data, the second adjustment configuration is determined. The second adjustment configuration includes at least two adjustment frequency points, and the adjustment frequency points in the second adjustment configuration are different from those in the first adjustment configuration. The second adjustment configuration is used to enhance the coherence and smoothness of the audio. The third adjustment configuration is determined based on the adjustment frequency points in the first adjustment configuration, the adjustment frequency points in the second adjustment configuration, the target audio characteristics, the scene description data, the frequency adjustment response curve, and the computational resource consumption data. The adjustment frequency points in the third adjustment configuration are different from those in the first and second adjustment configurations. The adjustment frequency points in the third adjustment configuration cover the entire sound frequency band. The third adjustment configuration is used to remove residual glitches and bumps in the audio.

3. The audio processing method based on multi-order frequency adjustment according to claim 1, characterized in that, The step of using a second EQ adjustment module to perform second-order frequency adjustment on the first audio based on the second adjustment configuration in the target adjustment configuration to obtain the second audio includes: The adjustment effect is evaluated based on the evaluation indicators corresponding to the first audio and the first adjustment configuration to obtain a first result; Based on the first result, update the second adjustment configuration; Based on the updated second adjustment configuration, the second EQ adjustment module is used to perform second-order frequency adjustment on the first audio to obtain the second audio. The step of using a third EQ adjustment module to perform third-order frequency adjustment on the second audio based on the third adjustment configuration in the target adjustment configuration to obtain the target audio includes: The adjustment effect is evaluated based on the evaluation indicators corresponding to the second audio and the second adjustment configuration to obtain a second result; Based on the second result, update the third adjustment configuration; Based on the updated third adjustment configuration, the third EQ adjustment module is used to perform third-order frequency adjustment on the second audio to obtain the target audio.

4. The audio processing method based on multi-order frequency adjustment according to claim 1, characterized in that, The step of determining the target adjustment configuration based on the target audio characteristics of the initial audio, the target cavity characteristic data, the scene description data of the sound-generating device, and the computational resource consumption data includes: Based on the ambient noise data corresponding to the sound-generating device, the power consumption configuration corresponding to the sound-generating device, the target audio characteristics, the target cavity characteristic data, the scene description data of the sound-generating device, and the computing resource consumption data, the target adjustment configuration is determined.

5. The audio processing method based on multi-order frequency adjustment according to claim 1, characterized in that, The initial frequency band data further includes: low-to-mid frequency data and mid-to-high frequency data. The step of synthesizing the target frequency band data to obtain the audio to be adjusted includes: The audio features of the initial audio are extracted, and a fusion auxiliary table is generated based on the audio features, wherein the audio features include zero-crossing rate, energy, and spectral centroid. Based on the first fusion weight configuration corresponding to the target audio characteristics of the initial audio, the low-frequency data and the mid-low frequency data are fused according to the fusion auxiliary table, the low-frequency data and the mid-low frequency data to obtain the first fused data; Based on the second fusion weight configuration corresponding to the target audio characteristics, the mid-frequency data and the mid-low frequency data are fused according to the fusion auxiliary table, the mid-frequency data and the mid-low frequency data to obtain the second fused data; Based on the third fusion weight configuration corresponding to the target audio characteristics, the mid-frequency overlapping data is fused according to the fusion auxiliary table, the mid-frequency data and the mid-high frequency data to obtain the third fused data; Based on the fourth fusion weight configuration corresponding to the target audio characteristics, the high-frequency overlapping data is fused according to the fusion auxiliary table, the high-frequency data and the mid-high frequency data to obtain the fourth fused data; The audio to be adjusted is obtained by synthesizing the low-frequency data, the mid-low-frequency data, the mid-frequency data, the first fusion data, the second fusion data, the third fusion data, and the fourth fusion data.

6. An audio processing system based on multi-order frequency adjustment, characterized in that, The system includes: an audio processing terminal and a sound-generating device; The audio processing terminal is configured to implement the audio processing method based on multi-order frequency adjustment as described in any one of claims 1 to 5; The sound-generating device is used to play the target audio output by the audio processing terminal.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the audio processing method based on multi-order frequency adjustment as described in any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the audio processing method based on multi-order frequency adjustment as described in any one of claims 1 to 5.