Audio processing method and system based on auditory model optimization

By optimizing audio processing through multi-scale frequency band decomposition and auditory masking effect calculation model, the problem of ignoring auditory characteristics in traditional methods is solved, and the subjective listening experience of high-fidelity, naturalness and clarity of audio signals is improved.

CN121531272APending Publication Date: 2026-02-13中央广播电视总台 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511692122.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Traditional audio processing methods fail to fully consider the physiological and psychological characteristics of the human auditory system, resulting in discrepancies between the processing results and subjective listening experience, and are prone to introducing intermodulation distortion and harmonic distortion.

Method used

Multi-scale frequency band decomposition is used to generate a set of sub-band signals that conform to the critical frequency band characteristics of the human ear. Dynamic gain adjustment and nonlinear distortion suppression are performed by combining an auditory masking effect calculation model. The time-domain audio signal is then reconstructed through a frequency band synthesis model.

Benefits of technology

The audio processing has been optimized to significantly improve subjective listening experience while enhancing physical performance, meeting the needs of high-level audio applications and reducing perceptible distortion and loss of signal dynamic range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121531272A_ABST
    Figure CN121531272A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audio signal processing, and discloses an audio processing method and system based on auditory model optimization. The method comprises the following steps: performing multi-scale frequency band decomposition processing on an input audio signal to generate a sub-band signal set covering different frequency ranges; the sub-band signal set is input into an auditory masking effect calculation model, and the model calculates masking threshold distribution of each sub-band signal according to frequency sensitivity characteristics of a human auditory system. And performing dynamic gain adjustment processing on the sub-band signal set based on the calculated masking threshold distribution to generate a gain-optimized sub-band signal set. And performing nonlinear distortion suppression processing on the sub-band signal set after gain optimization to generate a sub-band signal set after distortion suppression. And inputting the sub-band signal set after distortion suppression into a frequency band synthesis model to reconstruct a complete time domain audio signal. According to the invention, audio optimization processing conforming to human ear perception characteristics is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of audio signal processing, in particular to an audio processing method and system based on hearing model optimization. BACKGROUND

[0002] Audio signal processing technology has a wide range of applications in the fields of communication, multimedia, hearing aid devices, etc. Traditional audio processing methods are mainly based on the physical characteristics of the signal itself, such as frequency equalization, dynamic range compression, etc., and the parameter adjustment thereof depends on the experience of the operator or fixed presets. Although such methods can improve the audio quality to a certain extent, they do not fully consider the physiological and psychological characteristics of the human auditory system. The perception of sound by the human ear is not linear, and there are complex characteristics such as masking effect, frequency selectivity, etc. After the audio signal is processed by the traditional method, although the physical indicators may be optimized, the actual listening experience may not be improved, and even perceptible distortion or fatigue may occur.

[0003] The auditory masking effect refers to the fact that the presence of a stronger sound (masking sound) will raise the perception threshold of the auditory system for a weaker sound (masked sound) that exists at the same time. Traditional audio enhancement or compression algorithms often use the same processing strategy for all frequency bands, ignoring the perceptual non-uniformity caused by the masking effect. For example, near the strong energy frequency band, the human ear may not be able to perceive even if there is some noise or distortion; while in the quiet frequency band, a small processing flaw may be perceived. The existing technology lacks quantitative modeling and utilization of such auditory characteristics, resulting in a deviation between the processing result and the subjective listening experience.

[0004] In the reconstruction process of the audio signal, gain adjustment and nonlinear processing of the subband signal can easily introduce intermodulation distortion, harmonic distortion, etc. artifacts. Traditional methods often use simple clipping or filtering to suppress distortion, but these measures may cause loss of signal dynamic range or loss of high-frequency details. How to effectively enhance the audio signal while maximizing the suppression of perceptible distortion is a long-standing challenge in the field of audio processing. With the development of computational auditory scene analysis technology, processing based on the auditory model has gradually attracted attention, but its effective application to the actual audio processing flow still needs to solve key problems such as multi-scale analysis, dynamic parameter adjustment, and distortion control. SUMMARY

[0005] The purpose of the present application is to provide an audio processing method and system based on hearing model optimization to solve the problems raised in the background art.

[0006] To achieve the above purpose, the present application provides an audio processing method based on hearing model optimization, which comprises: performing multi-scale frequency band decomposition processing on the input audio signal to generate a set of subband signals containing different frequency ranges; inputting the sub-band signal set into an auditory masking effect calculation model, and calculating a masking threshold distribution of each sub-band signal according to a frequency sensitivity characteristic of a human auditory system; performing dynamic gain adjustment processing on the sub-band signal set based on the masking threshold distribution, to generate a gain-optimized sub-band signal set; performing non-linear distortion suppression processing on the gain-optimized sub-band signal set, to generate a distortion-suppressed sub-band signal set; inputting the distortion-suppressed sub-band signal set into a frequency band synthesis model, to reconstruct into a complete time-domain audio signal.

[0007] Preferably, the multi-scale frequency band decomposition processing on the input audio signal to generate a sub-band signal set containing different frequency ranges comprises: performing frequency band division on the audio signal by using a non-uniform filter bank, to generate an initial sub-band signal set; performing phase alignment processing on the initial sub-band signal set, to eliminate phase distortion introduced by the filter bank; performing frequency band energy normalization processing on the initial sub-band signal set according to a preset frequency band energy distribution threshold, to generate the sub-band signal set.

[0008] Preferably, the inputting of the sub-band signal set into an auditory masking effect calculation model to calculate a masking threshold distribution of each sub-band signal according to a frequency sensitivity characteristic of a human auditory system comprises: extracting an instantaneous energy distribution and a spectral envelope feature of each sub-band signal in the sub-band signal set; calling a pre-trained auditory critical band analysis model to perform inter-band masking effect calculation on the instantaneous energy distribution and the spectral envelope feature; generating a masking threshold distribution of each sub-band signal in different time windows according to the inter-band masking effect calculation result.

[0009] Preferably, the dynamic gain adjustment processing on the sub-band signal set based on the masking threshold distribution to generate a gain-optimized sub-band signal set comprises: determining a dynamic gain adjustment range of each sub-band signal according to the masking threshold distribution; performing real-time gain adjustment on the sub-band signal set in the dynamic gain adjustment range by using an adaptive gain control algorithm; performing energy smoothing processing on the gain-adjusted sub-band signal, to eliminate transient distortion introduced by gain mutation, to generate the gain-optimized sub-band signal set.

[0010] Preferably, the nonlinear distortion suppression processing on the gain-optimized subband signal set generates a distortion-suppressed subband signal set, including: detecting transient overload signal segments in the gain-optimized subband signal set; calling an auditory model-based distortion suppressor to perform dynamic clipping processing on the transient overload signal segments; performing harmonic compensation processing on the clipped subband signals to generate the distortion-suppressed subband signal set.

[0011] Preferably, the detecting transient overload signal segments in the gain-optimized subband signal set includes: calculating the ratio of instantaneous peak energy to long-term average energy of each subband signal; according to a preset overload detection threshold, screening out signal segments exceeding the overload detection threshold as the transient overload signal segments.

[0012] Preferably, the calling an auditory model-based distortion suppressor to perform dynamic clipping processing on the transient overload signal segments includes: determining a dynamic clipping threshold according to the energy distribution characteristics of the transient overload signal segments; using a soft clipping algorithm to perform smooth attenuation processing on signal amplitudes exceeding the dynamic clipping threshold.

[0013] Preferably, the performing harmonic compensation processing on the clipped subband signals includes: extracting harmonic distortion components of the clipped subband signals; performing selective compensation processing on the harmonic distortion components according to the masking threshold distribution.

[0014] Preferably, the inputting the distortion-suppressed subband signal set into a frequency band synthesis model to reconstruct a complete time-domain audio signal includes: performing phase consistency correction processing on the distortion-suppressed subband signal set; using an overlap-add method to synthesize the corrected subband signals into the complete time-domain audio signal.

[0015] Preferably, the present application further includes an auditory model-optimized audio processing system, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor, when executing the computer program, implements the steps of the above-mentioned auditory model-optimized audio processing method.

[0016] Compared with the prior art, the present application has the following beneficial effects: The application decomposes the audio signal into a sub-band set conforming to the critical frequency band characteristics of the human ear through multi-scale frequency band decomposition, laying a foundation for subsequent fine processing based on the hearing model. This decomposition method is closer to the frequency analysis mechanism of the auditory system, enabling the processing process to perform differentiated operations according to the perceptual importance of different frequency bands, thereby improving the scientificity and effectiveness of the processing.

[0017] The hearing masking effect calculation model is introduced to quantify the hearing threshold characteristics of the human ear at different frequencies and intensities. This method is no longer based on simple gain control of signal physical energy, but considers the masking relationship between sounds, allowing greater processing tolerance in frequency bands with strong masking sounds, and adopting a more cautious processing strategy in perceptually sensitive quiet bands, thereby optimizing the allocation of signal processing resources.

[0018] Based on the dynamic gain adjustment of the masking threshold distribution, the balance between audio enhancement and auditory comfort is achieved. The amplitude and method of gain adjustment dynamically change according to the real-time calculation of the masking threshold, improving the audibility of weak signals while avoiding the harshness caused by excessive enhancement or masking useful signals, making the processed audio more natural and clear in subjective listening.

[0019] A special nonlinear distortion suppression processing link effectively controls the auditory flaws that may be introduced in the signal processing process. Dynamic gain adjustment and other processes may produce nonlinear distortion, and the method reduces the impact of these distortions on the listening experience through targeted distortion suppression algorithms, ensuring the signal quality of the final output audio, and maintaining high fidelity characteristics while improving the listening experience.

[0020] Finally, the time-domain audio signal conforming to the auditory characteristics is reconstructed through the frequency band synthesis model, completing the complete closed loop from analysis to processing. The entire processing process is guided by the core of human auditory perception characteristics, so that the processing result is not only optimized in physical indicators, but more importantly, the subjective listening experience is significantly improved, meeting the needs of high-level audio applications. This method provides an effective technical means for audio quality optimization in the fields of communication, audio production, and hearing aid devices. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 A schematic diagram of multi-scale frequency band decomposition; Figure 2 A sub-flowchart of multi-scale frequency band decomposition processing; Figure 3 A sub-flowchart of hearing masking effect calculation; Figure 4 A dynamic gain adjustment effect comparison chart. DETAILED DESCRIPTION

[0022] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described, obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0023] Please refer to Figure 1 The present application provides an audio processing method and system based on an optimized hearing model, which includes: efficiently processing audio signals by simulating the characteristics of the human auditory system. The overall implementation scheme is as follows: performing multi-scale frequency band decomposition processing on the input audio signal to generate a sub-band signal set containing different frequency ranges; inputting the sub-band signal set into a hearing masking effect calculation model to calculate the masking threshold distribution of each sub-band signal according to the frequency sensitivity characteristics of the human auditory system; performing dynamic gain adjustment processing on the sub-band signal set based on the masking threshold distribution to generate a gain-optimized sub-band signal set; performing nonlinear distortion suppression processing on the gain-optimized sub-band signal set to generate a distortion-suppressed sub-band signal set; inputting the distortion-suppressed sub-band signal set into a frequency band synthesis model to reconstruct into a complete time-domain audio signal. This scheme fully utilizes the physiological basis of the hearing model to ensure that the audio processing minimizes the perceptual distortion while improving the quality.

[0024] The multi-scale frequency band decomposition processing divides the audio signal into multiple sub-bands using frequency division technology, and these sub-bands cover the sensitive frequency bands of human hearing, such as low, medium and high frequency regions. The decomposition process is based on a non-uniform division strategy to match the frequency response characteristics of the cochlea. The generation of the sub-band signal set is realized by a digital filter bank, and the filter design considers frequency overlap and phase consistency. The hearing masking effect calculation model integrates psychoacoustic parameters such as absolute hearing threshold and frequency masking curve, and the model estimates the masking threshold by real-time analysis of the time-frequency characteristics of the sub-band signal. The dynamic gain adjustment processing dynamically adjusts the gain of each sub-band according to the masking threshold, and the adjustment process uses an adaptive algorithm to avoid introducing audible noise. The nonlinear distortion suppression processing addresses the overload problem that may be caused by gain adjustment, and uses clipping and compensation mechanisms to suppress distortion. The frequency band synthesis model fuses the processed sub-band signals into time-domain signals through inverse transformation, and the synthesis process ensures phase continuity and energy conservation.

[0025] Embodiment 1: Please refer to Figure 2, the non-uniform filter bank is designed to strictly follow the frequency perceptual characteristics of the human auditory system, with the center frequency distribution set according to the Bark scale or equivalent rectangular bandwidth scale. The frequency band coverage of the non-uniform filter bank is set to the complete hearing range of 20 Hz to 20 kHz, and the implementation of the non-uniform filter bank uses a digital filter structure including an infinite impulse response filter or a finite impulse response filter. The number of filters of the non-uniform filter bank is determined according to the specific needs of the target audio application, and the frequency band division process of the non-uniform filter bank is completed through multi-rate signal processing technology. After the necessary sampling rate conversion, the input audio signal is convolved with the impulse response of the non-uniform filter bank, and the convolution operation can be performed in the time domain or the frequency domain to adapt to different computational efficiency requirements. The initial sub-band signal set contains a series of sub-band signals, each sub-band signal carrying time-domain information in a specific frequency range, and the generation process of the initial sub-band signal set ensures that the frequency resolution matches the auditory critical band. The design parameters of the non-uniform filter bank, including passband ripple and stopband attenuation, are determined through an optimization algorithm, with the goal of minimizing spectral leakage phenomena, and the frequency response characteristics of the non-uniform filter bank need to be verified through simulation or experiment to confirm that the isolation between sub-bands meets the processing requirements.

[0026] Phase alignment processing is performed on the initial sub-band signal set to eliminate the phase distortion introduced by the filter bank, and the phase alignment processing is implemented using an all-pass filter network or a group delay equalization method. The calculation of the phase alignment processing is based on detailed analysis of the phase response of each sub-band signal, and the phase response analysis is completed through Hilbert transform or analytical signal technology. The core goal of the phase alignment processing is to keep all sub-band signals accurately synchronized on the time axis, and the phase alignment processing is completed by adjusting the phase delay of each sub-band signal. The phase delay compensation operation is implemented using digital filtering or phase rotation methods, and the initial sub-band signal set after the phase alignment processing will exhibit consistent group delay characteristics. The actual effect of the phase alignment processing needs to be evaluated through cross-correlation function analysis or phase difference measurement to ensure that there is no phase cancellation problem in the subsequent synthesis stage, and the additional signal delay that may be introduced by the phase alignment processing is compensated through the buffer management mechanism of the system, thereby maintaining real-time processing performance.

[0027] The initial subband signal set is subjected to frequency band energy normalization processing according to a preset frequency band energy distribution threshold to generate a final subband signal set for subsequent processing. The frequency band energy distribution threshold is set based on an analysis of long-term energy statistical characteristics of the audio signal. The calculation process of the frequency band energy distribution threshold uses a sliding window analysis technique, and the length of the sliding window is dynamically adjusted according to the non-stationary degree of the signal. The frequency band energy normalization processing is implemented by accurately scaling the amplitude of each subband signal, and the scaling coefficient is determined depending on the real-time ratio of the subband signal energy to the frequency band energy distribution threshold. The frequency band energy normalization processing aims to balance the energy levels of each subband, avoiding the influence of auditory perception due to excessive or insufficient energy in certain frequency bands. The frequency band energy normalization processing can be operated on a logarithmic scale or a linear scale, and the specific choice is determined according to the actual needs of the audio dynamic range. The subband signal set after the frequency band energy normalization processing will be used as the input basis for the auditory masking effect calculation model, and the energy change needs to be continuously monitored during the frequency band energy normalization processing to prevent unnecessary distortion. The related parameters of the frequency band energy normalization processing, including attack time and release time, need to be configured according to specific application scenarios, and the output subband signal set of the frequency band energy normalization processing must preserve the perceptual characteristics of the original audio signal.

[0028] The specific implementation details of the non-uniform filter bank include the design process of the filter coefficients, which can be generated by a window function method or an optimization algorithm. The degree of frequency band overlap of the non-uniform filter bank needs to be controlled within a reasonable range allowed by auditory perception, and the computational complexity of the non-uniform filter bank can be reduced by using a polyphase filter structure or a fast convolution algorithm. The non-uniform filter bank needs to use a ring buffer for efficient data management when processing real-time audio streams, and the frequency response calibration of the non-uniform filter bank needs to be completed by inputting standard test signals. The output initial subband signal set of the non-uniform filter bank usually needs to be subjected to time-domain framing processing to facilitate subsequent module processing, and the frame length and overlap settings of the initial subband signal set must be matched with the analysis window of the auditory model.

[0029] The technical points of the phase alignment processing involve accurate measurement of phase consistency, which is achieved by calculating the cross-band phase correlation. The all-pass filter used in the phase alignment processing must have a flat amplitude response characteristic, and the design of the all-pass filter needs to ensure that the phase adjustment process does not introduce amplitude distortion. The phase alignment processing needs to independently apply phase correction to each subband signal in the initial subband signal set, and the phase correction amount is determined by an iterative optimization algorithm. The subband signals after the phase alignment processing should exhibit a coordinated phase relationship in the frequency domain, and the final effect of the phase alignment processing needs to be verified for perceptual transparency through professional auditory tests.

[0030] The implementation process of the band energy normalization processing includes a signal energy estimation step, and the energy estimation can be implemented by using a square law detector or an absolute value averager. The preset value of the band energy distribution threshold needs to be derived based on the principle of a psychoacoustic model, and the band energy normalization processing performs frame-by-frame energy adjustment on the initial subband signal set. The energy adjustment curve needs to be kept smooth to avoid audible mutations, and the gain control mechanism of the band energy normalization processing adopts a dynamic range control principle. The subband signal set output by the band energy normalization processing will be directly used in the masking threshold calculation module, and the parameters of the band energy normalization processing, including the threshold ratio and the compression ratio, need to be optimized through a series of hearing experiments.

[0031] The overall process of the multi-scale band decomposition processing integrates three core steps of non-uniform filter bank, phase alignment processing and band energy normalization processing, and the hardware platform of the multi-scale band decomposition processing needs to support digital signal processor or general processor running. The software implementation of the multi-scale band decomposition processing adopts the modular programming idea to enhance the maintainability, and the real-time requirement of the multi-scale band decomposition processing is guaranteed through the pipeline processing architecture. The performance evaluation of the multi-scale band decomposition processing needs to be based on the method combining objective index measurement and subjective auditory test, and the typical application scenarios of the multi-scale band decomposition processing include audio enhancement system and audio coding system.

[0032] The frequency band division strategy of the non-uniform filter bank fully considers the sensitivity characteristics of the human auditory system, and the non-uniform filter bank sets a higher frequency resolution in the low frequency region and a lower frequency resolution in the high frequency region. The boundary frequency of the non-uniform filter bank is accurately determined through the auditory critical band mapping relationship, and the engineering implementation of the non-uniform filter bank can call the optimized signal processing library function or develop custom code. The test verification of the non-uniform filter bank needs to be completed by inputting the sinusoidal sweep signal and analyzing its frequency response.

[0033] The specific algorithm selection of the phase alignment processing depends on the computational resource constraints of the system, and the real-time implementation version of the phase alignment processing can use a simplified model with lower computational complexity. The accuracy requirement of the phase alignment processing is determined by the quality requirement of the final audio synthesis signal, and the phase alignment processing corrects the phase deviation of the initial subband signal set in a statistical sense. The actual effect of the phase alignment processing needs to be measured by evaluating the seamless degree of the synthesis output signal.

[0034] The adaptive mechanism of the band energy normalization processing can dynamically adjust the energy threshold according to the type of the input signal, and the band energy normalization processing adopts a higher energy threshold for the speech signal to retain its intelligibility characteristics. The band energy normalization processing adopts a lower energy threshold for the music signal to enhance its dynamic range performance, and the specific implementation of the band energy normalization processing integrates energy detection and gain control closed loop. The subband signal set output by the band energy normalization processing is usually formatted as a multi-dimensional data array for use by subsequent processing modules.

[0035] The innovation of the multi-scale band decomposition processing method lies in the simulation of the auditory band processing mechanism, and the advantage of the multi-scale band decomposition processing method lies in the ability to improve the perceptual quality of audio processing. The implementation details of the multi-scale band decomposition processing cover multiple aspects such as filter design, phase management and energy control, and the parameter configuration of the multi-scale band decomposition processing can be realized through a standardized configuration file. The multi-scale band decomposition processing architecture has good scalability to support multi-channel audio processing requirements.

[0036] Embodiment 2: Please refer to Figure 3 The instantaneous energy distribution and spectral envelope features of each subband signal in the subband signal set are extracted, the instantaneous energy distribution is calculated using short-time energy analysis technology, and a time window function such as the Hanning window is used for the analysis window, and the window length is set to 10-50 milliseconds to match the time resolution of the auditory system. The instantaneous energy distribution is obtained by calculating the sum of squares of signal samples in each analysis window, and the energy value sequence constitutes the time-varying energy profile of the subband signal. The spectral envelope features are extracted using a spectral analysis method such as linear predictive coding, which obtains prediction coefficients by solving the Yule-Walker equation, and the prediction coefficients represent the formant structure and spectral shape of the subband signal. The spectral analysis obtains the cepstrum coefficients by inverse Fourier transform of the log power spectrum, and the cepstrum coefficients describe the smooth envelope of the spectrum. The instantaneous energy distribution and spectral envelope features are organized in matrix form, with the matrix rows corresponding to the subband index and the matrix columns corresponding to the time frame index.

[0037] The pre-trained auditory critical band analysis model is based on Bark scale, which is a kind of auditory frequency scale. The parameters of the pre-trained auditory critical band analysis model are trained by psychoacoustic experimental data. The input of the model is the time-frequency feature of the sub-band signal, and the output of the model is the masking threshold of each critical band. The inter-band masking effect calculation simulates the phenomena of simultaneous masking and temporal masking. The simultaneous masking calculation considers the energy interaction of adjacent sub-bands in the frequency domain, and the temporal masking calculation analyzes the masking effect of the preceding signal on the subsequent signal in the time domain. The masking effect calculation uses convolution operation as a mathematical operation method, and the convolution kernel shape is designed according to the auditory masking data. The calculation process generates the masking threshold distribution of each sub-band at different time points, and the masking threshold distribution represents the minimum sound pressure level that can be perceived by the auditory system to the sub-band signal.

[0038] The masking threshold distribution is represented in the form of a time-varying function, and the function value reflects the change of auditory sensitivity with frequency and time. The generation process of the masking threshold distribution includes threshold smoothing processing, which uses a moving average filter as a digital filter to eliminate the sudden components of the threshold curve. The masking threshold distribution data are stored in a ring buffer for real-time reading and use by the dynamic gain adjustment module.

[0039] Based on the masking threshold distribution, the dynamic gain adjustment range of each sub-band signal is determined. The calculation of the dynamic gain adjustment range is realized by comparing the difference between the sub-band signal energy and the masking threshold. The difference calculation uses the decibel scale, and the positive and negative of the difference determines the gain adjustment direction, and the size of the difference determines the gain adjustment amplitude. The dynamic gain adjustment range sets upper and lower boundaries. The upper boundary prevents distortion caused by excessive gain, and the lower boundary avoids loss of signal components caused by too small gain. The dynamic gain adjustment range is adaptively configured according to the signal type. A narrower adjustment range is used for speech signals to maintain speech intelligibility, and a wider adjustment range is used for music signals to enhance dynamic expression.

[0040] An adaptive gain control algorithm is used to perform real-time gain adjustment on the set of sub-band signals within the dynamic gain adjustment range. The adaptive gain control algorithm is based on the proportional-integral-derivative control principle of control theory. The proportional-integral-derivative control algorithm adjusts the gain coefficient through the error signal, which is the difference between the actual energy and the target energy. The gain adjustment process is performed frame by frame, and each frame of signal is multiplied by the corresponding gain coefficient. The gain coefficient is smoothly transitioned to avoid step changes. Real-time gain adjustment is realized using a digital multiplier structure. The multiplier input is the sub-band signal sample and the gain coefficient, and the multiplier output is the adjusted signal sample.

[0041] The energy-smoothing processing of the gain-adjusted subband signals eliminates the transient distortion introduced by the gain mutation, and the energy-smoothing processing uses a first-order infinite impulse response low-pass filter as a signal processor. The cutoff frequency of the filter is set according to the signal characteristics, and a higher cutoff frequency is used for voice signals to retain the transient characteristics, and a lower cutoff frequency is used for music signals to enhance the smoothness. The energy-smoothing processing is performed in the time domain, and each subband signal is processed independently. The smoothed signal maintains the overall shape of the original waveform. The signal output by the energy-smoothing processing constitutes a gain-optimized subband signal set, and the energy distribution of the gain-optimized subband signal set matches the auditory masking characteristics.

[0042] The training data of the pre-trained auditory critical band analysis model comes from large-scale psychoacoustic experiments that measure the masking threshold under different frequencies and intensities. The pre-trained auditory critical band analysis model uses a neural network structure as a computational model, and the neural network structure uses a multi-layer perceptron to learn the masking law. The model updating mechanism allows online learning to adapt to individual auditory differences, and the model updating is realized through a gradient descent algorithm as an optimization algorithm. Specifically, the model updating mechanism realizes adaptive learning of individual auditory differences by collecting real-time user feedback data, and the user feedback data comes from the subjective evaluation of the user on the processing result or the objective physiological signal measurement. The model updating mechanism converts the user feedback data into the parameter adjustment amount of the auditory critical band analysis model, and the parameter adjustment amount is calculated by the gradient descent algorithm. The gradient descent algorithm takes the difference between the model output and the expected output as the loss function, and calculates the gradient direction of the model parameters through the backpropagation algorithm. The model updating mechanism adjusts the weight matrix of the auditory critical band analysis model according to the negative gradient direction, and the adjustment amplitude of the weight matrix is controlled by the learning rate parameter. The learning rate parameter is dynamically adjusted according to the model convergence state, and a larger learning rate is used in the initial stage to speed up convergence, and a smaller learning rate is used in the later stage to improve accuracy. The model updating process uses a small batch training method, and each update uses a group of continuous time frames of data for gradient calculation. The model updating mechanism sets a verification link to prevent overfitting, and the verification link uses a reserved data set to evaluate the model generalization ability. The updated auditory critical band analysis model maintains the original architecture and only the parameter values change. The model updating mechanism supports continuous learning and incremental learning, and new knowledge is continuously integrated into the existing model without affecting the existing performance. The model updating process is executed in a background thread and does not affect the real-time performance of the front-end audio processing. The model updating mechanism includes a rollback function that automatically restores to the previous stable version when the update causes performance degradation. The trigger condition of the model updating mechanism is based on data quality evaluation, and the update process is started only when the user feedback data reaches a certain confidence level.

[0043] The simultaneous masking calculation in the inter-band masking effect calculation uses the spread function method, which describes the variation of the masking threshold with the frequency offset. The provisional masking calculation is divided into forward masking and backward masking. The forward masking calculation considers the influence of the masking sound on the subsequent test sound, and the backward masking calculation considers the influence of the test sound on the preceding masking sound. The time constant of the provisional masking is determined through auditory perception experiments. The forward masking time constant is about 20 milliseconds to 200 milliseconds, and the backward masking time constant is about 1 millisecond to 10 milliseconds. The safety margin of the dynamic gain adjustment range is set to 3 decibels to 6 decibels. The safety margin prevents the signal from being audible due to the error in the estimation of the masking threshold. The upper and lower boundaries of the dynamic gain adjustment range are dynamically adjusted according to the signal history statistics. The upper boundary refers to the recent signal peak level, and the lower boundary refers to the background noise level.

[0044] The update rate of the dynamic gain adjustment range is related to the non-stationary degree of the signal. The non-stationary signal uses a fast update strategy, and the stationary signal uses a slow update strategy. The parameters of the adaptive gain control algorithm include the attack time and the release time. The attack time controls the gain reduction speed, and the release time controls the gain recovery speed. The attack time is set to 1 millisecond to 5 milliseconds for fast response to signal mutations, and the release time is set to 50 milliseconds to 500 milliseconds for smooth gain variation. The implementation of the adaptive gain control algorithm uses a digital filter structure, and the attack time and the release time are configured through filter coefficients.

[0045] The filter design of the energy smoothing process considers the linear phase characteristic. The linear phase filter keeps the signal waveform undistorted. The filtering operation of the energy smoothing process uses a cross-fading strategy in the frame overlap area. The weighted average of the signal in the frame overlap area realizes seamless connection. The signal after the energy smoothing process is subjected to peak limiting check. The peak limiter prevents overload distortion introduced by the smoothing process.

[0046] The gain-optimized sub-band signal set is sent to the non-linear distortion suppression module for processing. The time-frequency characteristics of the gain-optimized sub-band signal set meet the requirements of auditory perception. The entire auditory masking effect calculation and dynamic gain adjustment processing flow is executed in real time on a digital signal processor, and the processing delay is controlled to be below the auditory perception threshold. The system performance is evaluated through objective indicators and subjective listening evaluation. The objective indicators include signal-to-noise ratio and total harmonic distortion, and the subjective listening evaluation uses double-blind listening tests.

[0047] Referring to Figure 4, the upper subgraph of the figure presents the normalized energy comparison of each subband before and after gain adjustment, which intuitively reflects the role of the adaptive gain control algorithm: different subbands obtain differentiated gain according to the auditory masking threshold, so that the energy distribution is more in line with the human ear perception characteristics; the lower subgraph shows the dynamic gain factor of each subband, which reflects the implementation details of determining the gain adjustment range according to the masking threshold distribution. The gain factor is dynamically allocated around 1.0, which not only ensures the accurate adjustment of the subbands that need to be optimized, but also avoids the risk of introducing transient distortion caused by gain mutation through reasonable gain amplitude, laying a foundation for subsequent energy smoothing processing.

[0048] In embodiment 3, the transient overload signal segments in the gain-optimized subband signal set are detected, and the detection process is based on the signal energy analysis method and peak statistical characteristics. The ratio of the instantaneous peak energy to the long-term average energy of each subband signal is calculated, the instantaneous peak energy is calculated by squaring the maximum amplitude value in the sliding time window, and the sliding time window length is set to 5 milliseconds corresponding to the duration of the transient component in the audio signal. The long-term average energy is calculated by a first-order recursive filter, and the filter time constant is set to 100 milliseconds to capture the stable energy level of the signal. The ratio calculation uses the decibel scale to represent the degree of dynamic range change, and the decibel value calculation formula is: In embodiment 3, the transient overload signal segments in the gain-optimized subband signal set are detected, and the detection process is based on the signal energy analysis method and peak statistical characteristics. The ratio of the instantaneous peak energy to the long-term average energy of each subband signal is calculated, the instantaneous peak energy is calculated by squaring the maximum amplitude value in the sliding time window, and the sliding time window length is set to 5 milliseconds corresponding to the duration of the transient component in the audio signal. The long-term average energy is calculated by a first-order recursive filter, and the filter time constant is set to 100 milliseconds to capture the stable energy level of the signal. The ratio calculation uses the decibel scale to represent the degree of dynamic range change, and the decibel value calculation formula is: The decibel ratio of the instantaneous peak energy to the long-term average energy is represented as The instantaneous peak energy is represented as The long-term average energy is represented. The signal segment exceeding the overload detection threshold is selected as the transient overload signal segment according to the preset overload detection threshold, and the overload detection threshold is determined as 6 decibels to 12 decibels according to an auditory distortion perception experiment. The overload detection threshold is dynamically adjusted according to the signal frequency characteristics, and a lower threshold is used for high-frequency signals to protect the auditory sensitive frequency band, and a higher threshold is used for low-frequency signals to retain the energy sense. The detection result is marked as a time region sequence, and each region contains a start sample point and an end sample point. An auditory model-based distortion suppressor is called to perform dynamic amplitude limiting processing on the transient overload signal segment, and the auditory model-based distortion suppressor integrates a frequency weighting function and auditory masking characteristics. The core processing module of the auditory model-based distortion suppressor includes a psychoacoustic model and a nonlinear processing unit, and the psychoacoustic model calculates the acceptable distortion threshold under the current signal state. The dynamic amplitude limiting threshold is determined according to the energy distribution characteristics of the transient overload signal segment, and the energy distribution characteristics are obtained by short-time spectral analysis including spectral centroid, spectral roll-off rate and other parameters. The dynamic amplitude limiting threshold is related to the signal frequency component, and an equal-loudness curve is used for weighting to make the amplitude limiting processing conform to the auditory sensitivity characteristics. The dynamic amplitude limiting threshold changes adaptively over time, and a lower threshold is used in the signal sharp change stage to provide protection, and a higher threshold is used in the stable stage to retain the dynamic range. A soft limiting algorithm is used to perform smooth attenuation processing on the signal amplitude exceeding the dynamic amplitude limiting threshold, and the soft limiting algorithm uses a continuous nonlinear function to realize amplitude compression. The input and output characteristic curves of the soft limiting function remain linear below the threshold, and above the threshold, a smooth transition curve is used to gradually approach the maximum output value. The implementation of the soft limiting algorithm uses a segmented polynomial function, and the polynomial coefficients are adjusted according to the signal sampling rate to ensure stability. The processing process maintains linear phase response to avoid introducing phase distortion, and each sample point is processed independently to ensure real-time performance. The attack time and release time parameters of the soft limiting algorithm are set according to the signal characteristics, the attack time controls the gain drop speed, which is usually 0.1 milliseconds to 1 millisecond, and the release time controls the gain recovery speed, which is usually 10 milliseconds to 50 milliseconds.

[0049] The sub-band signal after amplitude limiting processing is subjected to harmonic compensation processing, which aims to repair the high-order harmonic loss caused by the amplitude limiting process. The harmonic distortion component of the sub-band signal after amplitude limiting processing is extracted, and the extraction method uses adaptive filtering technology to separate the difference component of the original signal and the amplitude limiting signal. The harmonic distortion component analysis includes amplitude spectrum and phase spectrum characteristics, and focuses on the energy distribution of odd harmonics and even harmonics. According to the masking threshold distribution, the harmonic distortion component is selectively compensated, and the selective compensation is based on the auditory perception characteristics to only enhance the audible harmonic component. The compensation processing adopts frequency domain multiplication operation, and applies gain coefficient at the harmonic frequency point, and the gain coefficient is determined by the masking threshold of the corresponding frequency. The compensated harmonic component and the amplitude limiting signal are synthesized again, and the phase alignment is paid attention to in the synthesis process to avoid cancellation phenomenon. The detection algorithm of transient overload signal segment adopts a multi-condition joint judgment mechanism, which considers not only the energy ratio but also the signal waveform characteristics. The waveform slope change rate is used as an auxiliary judgment feature, and the slope calculation is realized through adjacent sample difference operation. The overload region boundary uses a hysteresis threshold to prevent frequent switching, and the opening threshold is slightly higher than the closing threshold to form a stable band. The detection result is subjected to morphological filtering to eliminate isolated noise points, and the morphological filtering uses opening operation and closing operation to smooth the region boundary.

[0050] The psychoacoustic model part of the distortion suppressor based on the hearing model calculates the masking threshold of each critical band in real time, and the critical band division adopts Bark scale. The input of the psychoacoustic model includes signal spectrum envelope and sound pressure level information, and the output is the maximum acceptable distortion degree of each frequency band. The dynamic amplitude limiting threshold calculation introduces a safety margin, which is set to 3-6 decibels to prevent marginal effects. The amplitude limiting threshold curve is smoothed to avoid jumping, and the smoothing uses a median filter to remove outliers.

[0051] The implementation of the soft limiting algorithm uses a lookup table method to improve the calculation efficiency, and the lookup table stores the corresponding relationship between the input and the output. The lookup table index is calculated by normalizing the signal amplitude, and the normalization reference is adaptively adjusted according to the long-term statistics of the signal. The soft limiting process monitors the distortion index in real time, and the distortion index is obtained by calculating the total harmonic distortion rate of the output signal. When the distortion exceeds the preset threshold, the amplitude limiting parameters are automatically adjusted, and the parameter adjustment uses the gradient descent method to optimize the auditory quality.

[0052] The analysis filter bank of the harmonic compensation processing is consistent with the front-end decomposition filter bank, ensuring frequency band alignment. The harmonic component extraction uses spectral subtraction, which subtracts the ideal amplitude limiting spectrum from the amplitude limiting signal spectrum to obtain the distortion component. The selective compensation mechanism sets a minimum compensation threshold to avoid introducing noise caused by excessive processing of small distortions. The compensated spectrum is inverse transformed to the time domain using the overlap-add method to ensure continuity.

[0053] The whole process of nonlinear distortion suppression is integrated in the digital signal processing chain, and the processing delay is controlled within 2 milliseconds. The system adopts a parallel processing architecture, and multiple subbands are processed simultaneously to improve throughput. The processing parameters are automatically configured according to the type of input signal, and different parameter sets are used for speech signals and music signals. The processing effect is monitored by objective indicators, including peak factor, harmonic distortion rate and other parameters. The system supports manual adjustment of parameters to meet the individual needs of different application scenarios. The nonlinear distortion suppression module is connected to the front and rear modules through a standard interface, and the data format uses floating-point numbers to ensure accuracy. The module is designed in a modular manner, making it easy to update algorithms and expand functions. Real-time processing performance is guaranteed through instruction-level optimization, and the key loop is implemented using assembly code. The module test uses standard test signals and actual recording materials to verify the correctness of the function and auditory transparency.

[0054] In example 4, the determination process of the dynamic clipping threshold refers to an energy distribution characteristic database, which stores energy mode characteristics of different audio signal types. The construction of the energy distribution characteristic database is based on a large number of sample analyses, covering speech, music and various environmental sounds. The dynamic clipping threshold calculation uses an energy statistical feature vector, which includes three components: instantaneous peak, short-term average and long-term average. The dynamic clipping threshold is obtained through a lookup table, and the index of the lookup table is generated by the normalized value of the feature vector. The dynamic clipping threshold is updated using a sliding window mechanism, and the window size matches the signal sampling rate. For a 48kHz sampling rate, a 1024-sample window is used. The dynamic clipping threshold is smoothed using a median filter, and the filter length is set to 3 frames to prevent threshold mutations.

[0055] The implementation of the soft clipping algorithm uses a polynomial approximation method, and the polynomial degree is selected as 5 to balance the calculation complexity and accuracy. The input-output characteristics of the soft clipping algorithm are divided into three regions: linear region, transition region and clipping region. The linear region maintains unit gain, the transition region uses cubic spline interpolation, and the clipping region uses asymptotic line approximation. The parameters of the soft clipping algorithm are adjusted adaptively according to the signal type. For speech signals, a harder transition characteristic is used to preserve clarity, and for music signals, a softer transition characteristic is used to maintain naturalness. The real-time processing of the soft clipping algorithm uses a parallel computing architecture, and each subband is processed independently to improve efficiency.

[0056] The output of the soft clipping algorithm is aligned in time to ensure the synchronization of multiple subband signals. The extraction process of the harmonic distortion component uses an adaptive filter bank, and the center frequency of the adaptive filter bank is aligned with the signal fundamental harmonic. The coefficients of the adaptive filter bank are updated through the LMS algorithm, and the step factor is set according to the convergence speed requirement. The harmonic distortion component is represented using a harmonic amplitude spectrum, which contains the first 10 harmonic components. The phase information of the harmonic distortion component is extracted through Hilbert transform, which preserves the phase relationship between harmonics.

[0057] The storage of harmonic distortion components uses a ring buffer, and the buffer depth meets the processing delay requirement. The core of the selective compensation processing is an auditory perceptual weighting matrix, and the dimension of the auditory perceptual weighting matrix is the same as the number of critical bands. The coefficients of the auditory perceptual weighting matrix are calculated by a masking threshold distribution, and the masking threshold distribution is from a front-end auditory model. The selective compensation processing adopts a frequency domain multiplication operation, and the compensation gain coefficient is determined by the ratio of the harmonic distortion amplitude to the masking threshold. The implementation of the selective compensation processing adopts an overlap-save method to avoid the boundary effect introduced by block processing. The parameters of the selective compensation processing are optimized through perceptual experiments, and the experimental samples contain various audio materials. Referring to Table 1, the corresponding relationship of the feature vector for dynamic clipping threshold calculation is shown.

[0058] Table 1: Dynamic clipping threshold lookup table Feature vector type Normalization range Clipping threshold (dB) Transition zone width (dB) Speech signal 0.2-0.5 -3.0 2.0 Music signal 0.1-0.8 -6.0 4.0 Impulse signal 0.5-1.0 -1.0 1.0 Sustained signal 0.05-0.2 -9.0 6.0 The real-time calculation process of the dynamic clipping threshold includes four steps of feature extraction, normalization, table lookup and application. The energy statistics of the current frame are calculated in the feature extraction stage, and the statistics include the peak value, the mean value and the variance. The normalization process uses historical extreme value scaling, and the historical extreme value is maintained through a sliding window. The table lookup operation uses bilinear interpolation to improve the accuracy, and the transition is smooth when non-integer index is processed. In the application stage, the threshold is converted into a linear gain, and the gain coefficient is calculated through the dB conversion formula.

[0059] The polynomial coefficients of the soft clipping algorithm are obtained by least squares fitting of the ideal curve, and the fitting error is controlled within 0.1%. The implementation of the soft clipping algorithm adopts fixed-point arithmetic optimization, and the fixed-point digit width is selected as 16 bits to balance the precision and efficiency. The anti-saturation processing of the soft clipping algorithm includes output clamping and integrator reset mechanism to prevent abnormal signals from causing system instability. The frequency response of the soft clipping algorithm remains flat, and the spectral distortion introduced by nonlinearity is corrected through pre-distortion compensation.

[0060] The analysis frame length of the harmonic distortion component is set to 1024 points, and the frame shift of 256 points meets the time resolution requirement. The tracking of the harmonic distortion component uses a peak detection algorithm, and the detection sensitivity is adaptively adjusted according to the signal-to-noise ratio. The verification of the harmonic distortion component is performed through harmonic verification to exclude non-harmonic noise interference. The interpolation processing of the harmonic distortion component is performed between frames to maintain the time continuity.

[0061] The gain calculation of the selective compensation processing uses a perceptual weighting function, and the shape of the weighting function conforms to the equal-loudness curve. The gain limit of the selective compensation processing sets the maximum compensation amount to prevent the introduction of new distortion by overcompensation. The smoothing processing of the selective compensation processing adopts a first-order IIR filter, and the filter coefficients are matched with the signal frame rate. The bypass mechanism of the selective compensation processing is activated when invalid harmonics are detected to ensure the robustness of the system.

[0062] The system integration test uses standard test signals, including sine sweep, pulse sequence and multi-tone signals. The system performance evaluation adopts objective indicators, including harmonic distortion rate, noise masking ratio and spectral distortion degree. The system parameter calibration is completed through listening experiments, and the experimenters are professionally trained. The system real-time verification is verified through delay measurement, and the overall processing delay is controlled within 3 milliseconds.

[0063] The parameter configuration interface of the dynamic limiting and harmonic compensation processing provides a professional mode and a simplified mode. The professional mode opens all parameter adjustments, and the simplified mode only provides preset options. The log recording function of the dynamic limiting and harmonic compensation processing saves processing statistical information, including limiting times, compensation amount and distortion degree. The version management of the dynamic limiting and harmonic compensation processing supports parameter preset storage and calling, which facilitates quick switching in different scenarios. The help system of the dynamic limiting and harmonic compensation processing provides detailed parameter explanations, including theoretical basis and practical guidance.

[0064] The dynamic limiting and harmonic compensation processing fault detection mechanism monitors the processing state, and the abnormal state triggers the protection strategy. The resource usage of the dynamic limiting and harmonic compensation processing is displayed in real time, including CPU occupancy and memory usage. The software implementation of the dynamic limiting and harmonic compensation processing adopts modular design, and the core algorithm is packaged as an independent library function. The hardware acceleration scheme of the dynamic limiting and harmonic compensation processing uses SIMD instruction parallel computing to improve processing efficiency. The cross-platform support of the dynamic limiting and harmonic compensation processing covers Windows, Linux and embedded systems, ensuring deployment flexibility.

[0065] In embodiment 5, the phase consistency correction processing is performed on the distortion-suppressed sub-band signal set. The phase consistency correction processing is based on the group delay equalization principle to eliminate the phase deviation between sub-bands. The phase consistency correction processing uses an all-pass filter network to adjust the phase response of each sub-band signal, and the coefficients of the all-pass filter network are calculated by an iterative optimization algorithm. The goal of the phase consistency correction processing is to make all sub-band signals have consistent phase characteristics when synthesized, and the accuracy requirement of the phase consistency correction processing reaches sample-level alignment. The group delay change is monitored during the phase consistency correction processing, and the group delay change is obtained by calculating the phase derivative of the analytic signal through Hilbert transform. The phase consistency correction processing adopts a relatively loose correction tolerance for low-frequency sub-bands and a strict correction tolerance for high-frequency sub-bands, matching the frequency perception characteristics of the human auditory system.

[0066] The specific implementation of the phase consistency correction process includes the following steps: establishing a reference phase reference, the reference phase reference selects a center subband or a subband with maximum energy as a reference point for time alignment. Calculate the phase offset of each subband relative to the reference phase reference, the phase offset is determined by peak detection of cross-correlation function. Design a phase compensation filter, the phase compensation filter uses a finite impulse response filter structure to achieve accurate phase adjustment. Apply the phase compensation filter to the subband signal for phase correction, the phase correction process keeps the amplitude response unchanged. Verify the effect of phase consistency, confirm the correction accuracy by measuring the phase difference between the subband signals.

[0067] The corrected subband signals are synthesized into complete time-domain audio signals by using the overlap-add method, and the implementation of the overlap-add method is based on the frame processing architecture. The overlap-add method divides each subband signal into consecutive time frames, and the time frame length is consistent with the frame length used in the analysis stage. The overlap-add method applies a synthesis filter bank to each time frame, and the synthesis filter bank and the analysis filter bank form a perfect reconstruction pair. The overlap-add method sets an overlap region between frames, and the length of the overlap region is usually half or three quarters of the frame length. The overlap-add method uses a windowing function in the overlap region to achieve smooth transition, and the selection of the windowing function considers the main lobe width and the sidelobe attenuation characteristics.

[0068] The detailed process of the overlap-add method is as follows: frame division is performed on the phase-corrected subband signals, and the frame division uses a sliding window method to continuously cut signal segments. A synthesis filter is applied to each signal frame of the subband, and the impulse response of the synthesis filter satisfies the biorthogonal condition with the analysis filter. The filtered subband signal frames are arranged in time sequence, and the sample points in the overlap region are weighted and added. The weighting function used by the overlap-add method is usually a cosine window or a raised cosine window, and the shape of the weighting function affects the smoothness of the synthesized signal. The output of the overlap-add method is subjected to gain normalization to ensure that the energy level of the synthesized signal is consistent with that of the original signal.

[0069] The special considerations of the phase consistency correction process include the adaptive mechanism when processing non-stationary signals, which refer to signal segments with rapidly changing instantaneous frequencies. For non-stationary signals, the phase consistency correction process uses a dynamic adjustment strategy, which modifies the phase compensation parameters according to the instantaneous frequency change rate. The instantaneous frequency change rate is calculated by phase difference, which uses the phase difference between the previous and next frames divided by the time interval. When a rapidly changing instantaneous frequency is detected, the phase consistency correction process temporarily relaxes the correction requirement to avoid distortion caused by excessive correction. For quasi-stationary signal segments, the phase consistency correction process uses strict phase alignment standards to ensure synthesis quality.

[0070] The implementation optimization of overlap-add method involves the improvement of computational efficiency, which is achieved by polyphase filtering and fast convolution techniques. Polyphase filtering decomposes the synthesis filter into multiple sub-filters, each of which processes a sub-sequence of down-sampled signal. Fast convolution replaces time-domain convolution with frequency-domain multiplication, significantly reducing computational complexity. The real-time performance of overlap-add method is guaranteed by a frame pipelining architecture, which enables the parallel execution of the previous frame synthesis operation and the current frame analysis operation. The memory management of overlap-add method adopts a ring buffer structure, which efficiently handles continuous data streams.

[0071] The quality assessment of phase coherence correction processing uses objective indicators, including phase root mean square error and group delay fluctuation coefficient. The phase root mean square error measures the degree of consistency between the corrected sub-band signal and the reference phase, and the group delay fluctuation coefficient reflects the linearity of the phase. The perceptual verification of phase coherence correction processing is carried out through subjective listening tests, which use a double-blind AB comparison method. The parameter optimization of phase coherence correction processing is based on auditory experimental data, which comes from the evaluation results of professional listeners.

[0072] The special processing of overlap-add method targets the boundary effect, which refers to the possible signal discontinuity at the frame boundary. Measures to alleviate the boundary effect include increasing the length of the overlap region, using a window function with smoother transition characteristics, and using a boundary interpolation algorithm. The processing of transient signals by overlap-add method uses an adaptive frame length strategy, which automatically switches to a shorter frame length when a transient is detected. The processing of steady-state signals by overlap-add method uses long frames to improve frequency resolution, while the processing of transient signals uses short frames to maintain time resolution.

[0073] The integration test of band synthesis model includes normal signal and edge case verification, including full-band sine wave, impulse signal and silent section. The performance benchmark test of band synthesis model uses standard audio data set, which contains various music types and speech materials. The real-time performance analysis of band synthesis model measures the calculation delay and CPU occupancy, and the time delay from input to output is required to be controlled within 10 milliseconds. The hardware acceleration of band synthesis model utilizes the SIMD instruction set of modern processors, which realizes parallel processing of multiple sub-band data.

[0074] The phase coherence correction process cooperates with the overlap-add method through a global scheduler that coordinates the timing relationship of the data processing pipeline. Precondition checks of the phase coherence correction process ensure that the input signals meet the processing requirements, including sample rate consistency and subband quantity verification. The post-processing stage of the overlap-add method includes DC offset removal and quantization noise shaping, with a high-pass filter used for DC offset removal to remove low-frequency distortion, and bit allocation optimized by a noise transfer function for quantization noise shaping. Parameter configurability of the frequency band synthesis model supports different application scenarios, which is implemented through a graphical interface or a configuration file. Debugging tools of the frequency band synthesis model provide intermediate signal observation points that allow monitoring of the signal quality of each processing stage. Version management of the frequency band synthesis model records the algorithm modification history, which facilitates tracking of performance improvements and problem fixes. Documentation of the frequency band synthesis model specifies the interface specifications and processing flow in detail, including code comments and design specifications.

[0075] Numerical stability of the phase coherence correction process is ensured by fixed-point arithmetic, which uses saturation processing and rounding modes. Anti-aliasing measures of the overlap-add method include upsampling filtering and mirror frequency suppression, with an interpolation filter used for upsampling filtering to eliminate high-order harmonics. Power consumption optimization of the frequency band synthesis model is designed for mobile devices, which is implemented through dynamic voltage frequency adjustment and clock gating. Multi-channel extension of the frequency band synthesis model supports stereo and surround sound formats, while maintaining inter-channel phase coherence.

[0076] An exception handling mechanism of the frequency band synthesis model detects illegal inputs and system errors, which uses a safe recovery strategy to ensure system reliability. A calibration program of the frequency band synthesis model uses standard test signals, which is run regularly to maintain processing accuracy. A life test of the frequency band synthesis model accelerates aging experiments, which verifies the long-term running stability of the system. Compatibility tests of the frequency band synthesis model cover different operating systems and hardware platforms, which ensure software portability.

[0077] It should be noted that, in the present text, relational terms such as first and second are used merely to distinguish one entity or action from another, without necessarily requiring or implying any such actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variant thereof are intended to cover non-exclusive inclusions, so that a process, method, article, or apparatus that comprises a list of elements does not only include those elements, but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus.

[0078] While embodiments of the application have been shown and described, it is to be understood that the embodiments described are merely exemplary of the principles and application of the present application. Numerous modifications and adaptions can be effected without departing from the spirit and scope of the present application, which is not limited to the exact construction and arrangement described. It is intended, therefore, to cover all modifications and adaptions that fall within the scope of the claims and their equivalents.

Claims

1. An audio processing method based on auditory model optimization, characterized in that, include: The input audio signal is subjected to multi-scale frequency band decomposition to generate a set of sub-band signals containing different frequency ranges; The sub-band signal set is input into the auditory masking effect calculation model, and the masking threshold distribution of each sub-band signal is calculated based on the frequency sensitivity characteristics of the human auditory system. Based on the masking threshold distribution, the sub-band signal set is dynamically gain-adjusted to generate a gain-optimized sub-band signal set. The gain-optimized subband signal set is subjected to nonlinear distortion suppression processing to generate a distortion-suppressed subband signal set; The set of sub-band signals after distortion suppression is input into the frequency band synthesis model and reconstructed into a complete time-domain audio signal.

2. The audio processing method based on auditory model optimization according to claim 1, characterized in that, The process of performing multi-scale frequency band decomposition on the input audio signal to generate a set of sub-band signals containing different frequency ranges includes: The audio signal is divided into frequency bands using a non-uniform filter bank to generate an initial set of sub-band signals. The initial sub-band signal set is phase aligned to eliminate phase distortion introduced by the filter bank; Based on a preset frequency band energy distribution threshold, the initial sub-band signal set is subjected to frequency band energy normalization processing to generate the sub-band signal set.

3. The audio processing method based on auditory model optimization according to claim 2, characterized in that, The step of inputting the sub-band signal set into the auditory masking effect calculation model, and calculating the masking threshold distribution of each sub-band signal based on the frequency sensitivity characteristics of the human auditory system, includes: Extract the instantaneous energy distribution and spectral envelope features of each sub-band signal in the sub-band signal set; The pre-trained auditory critical frequency band analysis model is invoked to calculate the inter-band masking effect of the instantaneous energy distribution and spectral envelope characteristics; Based on the calculation results of the inter-band masking effect, the masking threshold distribution of each sub-band signal in different time windows is generated.

4. The audio processing method based on auditory model optimization according to claim 3, characterized in that, The step of dynamically adjusting the gain of the sub-band signal set based on the masking threshold distribution to generate a gain-optimized sub-band signal set includes: The dynamic gain adjustment range for each sub-band signal is determined based on the masking threshold distribution; An adaptive gain control algorithm is used to adjust the gain of the sub-band signal set in real time within the dynamic gain adjustment range; The subband signal after gain adjustment is subjected to energy smoothing to eliminate transient distortion introduced by gain abrupt changes, thereby generating the set of subband signals after gain optimization.

5. The audio processing method based on auditory model optimization according to claim 4, characterized in that, The step of performing nonlinear distortion suppression processing on the gain-optimized subband signal set to generate a distortion-suppressed subband signal set includes: Detect transient overload signal segments in the gain-optimized subband signal set; A distortion suppressor based on an auditory model is invoked to perform dynamic amplitude limiting on the transient overload signal segment; Harmonic compensation processing is performed on the sub-band signals after amplitude limiting to generate the set of sub-band signals after distortion suppression.

6. The audio processing method based on auditory model optimization according to claim 5, characterized in that, The detection of transient overload signal segments in the gain-optimized sub-band signal set includes: Calculate the ratio of the instantaneous peak energy to the long-term average energy of each sub-band signal; Based on a preset overload detection threshold, signal segments exceeding the overload detection threshold are selected as transient overload signal segments.

7. The audio processing method based on auditory model optimization according to claim 6, characterized in that, The invocation of an auditory model-based distortion suppressor to dynamically limit the transient overload signal segment includes: Based on the energy distribution characteristics of the transient overload signal segment, determine the dynamic limiting threshold; A soft limiting algorithm is used to smoothly attenuate the signal amplitude that exceeds the dynamic limiting threshold.

8. The audio processing method based on auditory model optimization according to claim 7, characterized in that, The harmonic compensation processing of the sub-band signal after amplitude limiting includes: Extract the harmonic distortion components of the sub-band signal after amplitude limiting; The harmonic distortion components are selectively compensated based on the masking threshold distribution.

9. The audio processing method based on auditory model optimization according to claim 8, characterized in that, The step of inputting the distortion-suppressed sub-band signal set into the frequency band synthesis model to reconstruct a complete time-domain audio signal includes: The distortion-suppressed subband signal set is subjected to phase consistency correction processing; The corrected sub-band signals are synthesized into the complete time-domain audio signal using the overlapping addition method.

10. An audio processing system optimized based on an auditory model, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the audio processing method based on auditory model optimization as described in any one of claims 1 to 9.