Intelligent voice reduction audio processing method and system
Through intelligent audio reduction method, combined with time-frequency analysis and multi-stage signal processing, the problem of poor echo cancellation in complex acoustic environments is solved, and more efficient audio signal processing and better user experience is achieved.
Patent Information
- Application Number
- CN202510295337.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-13
AI Technical Summary
Existing echo cancellation techniques are not effective in complex acoustic environments, especially under nonlinear distortion and dynamic environmental changes, which lead to impact on audio system performance and user experience.
The intelligent down-return sound audio processing method is adopted to obtain the sound field signal and speaker reference signal, time-frequency analysis and multi-stage signal processing are performed, including linear echo cancellation, nonlinear residual echo cancellation, multimodal residual suppression and acoustic enhancement.
It improves the initial effect of echo cancellation, enhances the clarity of the audio signal and overall processing efficiency, adapts to different acoustic environments and application scenarios, and reduces communication interference.
Smart Images

Figure CN119811412B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio processing, and particularly to an intelligent echo reduction audio processing method and system. Background Art
[0002] As an important medium for modern information exchange, audio communication systems have been widely used in remote conferences, intelligent assistants, and voice control devices. With the rapid development of audio interaction technology and the continuous improvement of user experience requirements, how to effectively eliminate echo interference and improve audio quality and communication clarity has become one of the research focuses. Existing echo cancellation technologies often only adopt a single processing method, such as simple linear filtering or spectrum suppression with fixed parameters, while ignoring the complexity of the actual acoustic environment and the comprehensive impact of multi-path propagation on echo formation. Such a simplified method may lead to poor echo cancellation effects, especially under non-linear distortion and dynamic environmental change conditions, thus significantly affecting the overall performance of the audio system and user experience. Summary of the Invention
[0003] The main object of the present invention is to provide an intelligent echo reduction audio processing method and system, which can more accurately identify and filter basic echo components, thereby improving the initial effect of echo cancellation.
[0004] To achieve the above object, the present invention provides an intelligent echo reduction audio processing method, including:
[0005] Obtain a sound field signal and a speaker reference signal, perform signal framing and short-time Fourier transform to obtain time-frequency spectrum audio data;
[0006] Perform linear echo cancellation processing on the time-frequency spectrum audio data and dynamically adjust the convergence speed to obtain a preliminary filtered signal;
[0007] Perform non-linear residual echo cancellation on the preliminary filtered signal and generate multi-reflection paths to obtain a non-linear cancellation signal;
[0008] Perform multi-modal residual suppression on the non-linear cancellation signal, perform spectrum processing through sub-band division and joint decision to obtain an echo suppression signal;
[0009] Obtain real-time ambient sound field change data, perform acoustic enhancement on the echo suppression signal to obtain a target echo reduction audio output.
[0010] Further, the obtaining a sound field signal and a speaker reference signal, performing signal framing and short-time Fourier transform to obtain time-frequency spectrum audio data includes:
[0011] Synchronously digitize sample the sound field signal and the speaker reference signal according to a preset sampling rate and bit depth to obtain an original signal in the digital domain;
[0012] Perform time frame segmentation on the original signal in the digital domain according to preset frame length and frame shift parameters to obtain a framed sequence;
[0013] Optimize spectral leakage of the framed sequence according to a preset Hanning window function to obtain a windowed signal frame;
[0014] Perform a fast Fourier transform on the windowed signal frame to obtain complex frequency domain coefficients;
[0015] Calculate an amplitude spectrum and a phase spectrum based on the complex frequency domain coefficients, and perform band energy analysis to obtain a band energy distribution matrix;
[0016] Perform a Mel frequency scale mapping on the band energy distribution matrix to obtain Mel filter bank coefficients;
[0017] Perform three-dimensional feature construction on the Mel filter bank coefficients and the complex frequency domain coefficients to obtain the time-frequency spectrum audio data.
[0018] Further, the linear echo cancellation processing is performed on the time-frequency spectrum audio data, and the convergence speed is dynamically adjusted to obtain a preliminary filtered signal, including:
[0019] Perform multi-subband decomposition on the time-frequency spectrum audio data to obtain a plurality of frequency subband signals;
[0020] Perform echo path analysis on the plurality of frequency subband signals to obtain a linear echo path signal;
[0021] Calculate an instantaneous mismatch metric between the sound field signal and the speaker reference signal to obtain a convergence state parameter;
[0022] Adjust a step parameter according to the convergence state parameter to obtain an optimized step value;
[0023] Perform upper and lower threshold setting mapping on the optimized step value to obtain a stable step parameter;
[0024] Perform frequency domain filtering on the time-frequency spectrum audio data according to the stable step parameter to obtain a linear echo estimation signal;
[0025] Optimize the time-frequency spectrum audio data according to the linear echo estimation signal, and perform residual correlation compensation to obtain the preliminary filtered signal.
[0026] Further, the non-linear residual echo cancellation is performed on the preliminary filtered signal, and multi-reflection paths are generated to obtain a non-linear cancellation signal, including:
[0027] Construct a non - linear residual basis for the preliminary filtered signal to obtain a coupled feature matrix;
[0028] Construct a Volterra kernel expansion space based on the coupled feature matrix and perform real - time update on the preliminary filtered signal to obtain optimized filter parameters;
[0029] Perform multi - reflection path analysis based on the optimized filter parameters to obtain a set of adversarial reflection paths;
[0030] Perform time - frequency domain superposition of the set of adversarial reflection paths and the preliminary filtered signal to obtain a multi - path suppression signal;
[0031] Analyze the sub - band energy distribution of the multi - path suppression signal and generate a dynamic time - frequency masking matrix according to a preset threshold value;
[0032] Perform point - by - point multiplication operation on the multi - path suppression signal according to the dynamic time - frequency masking matrix to obtain the non - linear cancellation signal.
[0033] Further, the constructing a Volterra kernel expansion space based on the coupled feature matrix and performing real - time update on the preliminary filtered signal to obtain optimized filter parameters includes:
[0034] Perform singular value decomposition on the coupled feature matrix to obtain a dimensionality - reduced feature space;
[0035] Construct separated components for the dimensionality - reduced feature space to obtain a set of separated vectors;
[0036] Construct the second - order expansion term of the Volterra kernel function according to the set of separated vectors and perform weighted summation of the expansion terms to obtain a second - order kernel expansion matrix;
[0037] Construct third - order cross - terms according to the set of separated vectors and the second - order kernel expansion matrix and perform tensor product operation to obtain a third - order kernel expansion tensor;
[0038] Construct an inner - product operation of the second - order kernel expansion matrix and the third - order kernel expansion tensor to obtain a space mapping function;
[0039] Perform non - linear projection on the preliminary filtered signal according to the space mapping function to obtain a projection coefficient vector;
[0040] Perform least - squares iterative optimization on the projection coefficient vector to obtain the optimized filter parameters.
[0041] Further, performing multi-modal residual suppression on the non-linear cancellation signal, and performing spectrum processing through sub-band division and joint decision to obtain an echo suppression signal, including:
[0042] Performing sub-band decomposition on the non-linear cancellation signal to obtain a non-uniform sub-band representation;
[0043] Performing multi-modal feature extraction and integration on the non-uniform sub-band representation to obtain a multi-modal feature set;
[0044] Calculating cross-entropy and mutual information according to the multi-modal feature set to obtain a sub-band grouping structure;
[0045] Performing adaptive grouping on the non-uniform sub-band representation according to the sub-band grouping structure to obtain multiple groups of sub-band sets;
[0046] Evaluating decision rules for each of the multiple groups of sub-band sets to obtain multiple suppression decision values;
[0047] Performing joint decision processing on the multiple suppression decision values to obtain a target suppression strategy;
[0048] Performing differential processing on the multiple groups of sub-band sets according to the target suppression strategy to obtain a modal collaborative suppression signal;
[0049] Performing variational decomposition phase compensation on the modal collaborative suppression signal to obtain a phase coordination signal;
[0050] Performing asymmetric singular value decomposition and reconstruction on the phase coordination signal to obtain the echo suppression signal.
[0051] Further, obtaining real-time environmental sound field change data, and performing acoustic enhancement on the echo suppression signal to obtain an output of the target reduced reverberation audio, including:
[0052] Obtaining environmental sensor data, and performing spatial spectrum decomposition and sub-band analysis on the environmental sensor data according to the sound field signal to obtain environmental sound field spectrum characteristics;
[0053] Calculating acoustic parameter changes between the environmental sound field spectrum characteristics and the environmental sensor data to obtain the real-time environmental sound field change data;
[0054] Performing multi-stage cascade filtering processing on the echo suppression signal according to the real-time environmental sound field change data to obtain a preliminary enhancement signal;
[0055] Performing phase correction and spectrum smoothing processing on the preliminary enhancement signal, and adjusting the gain factor according to the real-time environmental sound field change data to obtain a frequency-domain enhancement signal;
[0056] Perform residual noise selective suppression on the frequency-domain enhanced signal and the real-time ambient sound field change data to obtain a noise suppression signal;
[0057] Perform auditory frequency band energy adjustment on the noise suppression signal to obtain a time-domain enhanced signal;
[0058] Perform harmonic reconstruction and formant enhancement processing on the time-domain enhanced signal to obtain the target degraded voice frequency output.
[0059] Further, the performing residual noise selective suppression on the frequency-domain enhanced signal and the real-time ambient sound field change data to obtain a noise suppression signal includes:
[0060] Perform frequency band region division on the frequency-domain enhanced signal to obtain a frequency band division signal;
[0061] Perform cross-correlation analysis on the frequency band division signal and the real-time ambient sound field change data to obtain the residual noise distribution characteristics;
[0062] Perform frequency band selective masking analysis on the residual noise distribution characteristics to obtain a frequency band noise marking signal;
[0063] Perform Bayesian probability estimation on the frequency band noise marking signal to obtain a noise probability distribution map;
[0064] Perform non-linear suppression analysis on the frequency-domain enhanced signal according to the noise probability distribution map to obtain a preliminary noise suppression signal;
[0065] Perform spectral subtraction balance on the preliminary noise suppression signal and the real-time ambient sound field change data to obtain a balanced suppression signal;
[0066] Perform artifact threshold clearing on the balanced suppression signal to obtain an artifact removal coefficient;
[0067] Perform wavelet reconstruction and phase preservation on the artifact removal coefficient to obtain a noise suppression signal.
[0068] The present invention also provides an intelligent degraded voice frequency processing system, which is applied to the intelligent degraded voice frequency processing method described in any one of the above, and includes:
[0069] An acquisition module, which is used to acquire a sound field signal and a speaker reference signal, and perform signal framing and short-time Fourier transform to obtain time-frequency spectrum audio data;
[0070] An analysis module, which is used to perform linear echo cancellation processing on the time-frequency spectrum audio data and dynamically adjust the convergence speed to obtain a preliminary filtered signal;
[0071] An association module, which is used to perform non - linear residual echo cancellation on the preliminary filtered signal and generate multiple reflection paths to obtain a non - linear cancellation signal;
[0072] A processing module, which is used to perform multi - modal residual suppression on the non - linear cancellation signal, and perform spectrum processing through sub - band division and joint decision to obtain an echo suppression signal;
[0073] A control module, which is used to obtain real - time ambient sound field change data and perform acoustic enhancement on the echo suppression signal to obtain a target reduced - echo audio output.
[0074] An intelligent reduced - echo audio processing method and system provided by the present invention have the following beneficial effects:
[0075] Through time - frequency analysis and processing of the sound field signal and the speaker reference signal, combined with the dynamically adjusted linear echo cancellation technology, it can more accurately identify and filter the basic echo components, thereby improving the initial effect of echo cancellation and providing a more reliable basis for subsequent processing. By introducing non - linear residual echo cancellation and multiple reflection path generation mechanisms, it realizes the refined processing of complex echo paths, helps to solve the processing defects of traditional methods under non - linear conditions, and avoids sound quality distortion. Based on the joint decision mechanism of multi - modal residual suppression and sub - band division, it can ensure that the system can operate efficiently under different acoustic environments and echo characteristics, reduce the impact of residual echo on communication quality, and improve the overall processing efficiency of the system. By obtaining real - time ambient sound field change data and performing acoustic enhancement, a more intelligent adaptive processing strategy is formulated to realize the dynamic optimization of the entire system, thereby effectively improving audio clarity and reducing communication interference. And by comprehensively considering the complexity of the actual acoustic environment, it can flexibly adjust the echo cancellation strategy according to the characteristics and demand changes of different application scenarios, making the system more adaptable to diverse audio interaction scenarios. Description of the Drawings
[0076] Figure 1 is a flowchart of an intelligent reduced - echo audio processing method provided by the present invention;
[0077] Figure 2 is a structural diagram of an intelligent reduced - echo audio processing system provided by the present invention.
[0078] The realization, functional characteristics and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments
[0079] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0080] Next, the present invention will be further described in conjunction with the accompanying drawings and specific embodiments.
[0081] Referring to Figure 1 as shown, the present invention provides 1. An intelligent echo reduction audio processing method, comprising:
[0082] Step S1: Obtain a sound field signal and a speaker reference signal, and perform signal framing and short-time Fourier transform to obtain time-frequency spectrum audio data;
[0083] Step S2: Perform linear echo cancellation processing on the time-frequency spectrum audio data and dynamically adjust the convergence speed to obtain a preliminary filtered signal;
[0084] Step S3: Perform non-linear residual echo cancellation on the preliminary filtered signal and generate multiple reflection paths to obtain a non-linear cancellation signal;
[0085] Step S4: Perform multi-modal residual suppression on the non-linear cancellation signal, and perform spectrum processing through sub-band division and joint decision to obtain an echo suppression signal;
[0086] Step S5: Obtain real-time ambient sound field change data, and perform acoustic enhancement on the echo suppression signal to obtain the target echo reduction audio output.
[0087] Based on the steps shown above, the detailed step process is as follows:
[0088] Step S1: The microphone array captures the sound field signal containing the target speech and echo components, and at the same time the system obtains the reference signal played by the speaker. These two signals are framed after preprocessing. Usually, a frame length of 20 - 30 ms is used, and there is an overlap of about 50% between frames to ensure the continuity of signal processing. The framed signal is multiplied by a window function such as a Hanning window or a Hamming window to reduce spectral leakage. Subsequently, the fast Fourier transform (FFT) is performed on each windowed signal frame to map the time-domain signal to the complex frequency domain, forming a time-frequency spectrum representation. The time-frequency spectrum data contains amplitude spectrum and phase spectrum information. The amplitude spectrum reflects the energy distribution of each frequency component, and the phase spectrum retains the timing characteristics of the signal. This time-frequency representation enables the echo cancellation algorithm to process the signal characteristics separately in different frequency bands, improving the accuracy and efficiency of echo cancellation. Time-frequency domain processing also facilitates the subsequent analysis of the acoustic environment characteristics, such as the room impulse response and reverberation characteristics, providing the basic data for frequency domain analysis in the subsequent steps.
[0089] Step S2: Based on the obtained time-frequency spectrum data, linear echo cancellation processing uses frequency-domain based adaptive algorithms such as the NLMS (Normalized Least Mean Square) or RLS (Recursive Least Squares) algorithm to estimate the linear transfer function from the speaker signal to the microphone. The adaptive filter continuously updates the filter coefficients to minimize the mean square error between the estimated echo and the actual echo. Dynamically adjusting the convergence speed is the key optimization point of this process. The system intelligently adjusts the step size parameter by evaluating the signal-to-noise ratio, signal stationarity, and double-talk detection results. In the single-talk scenario (only the far-end signal), a larger step size is used to accelerate convergence; in the double-talk scenario (both the far-end and near-end signals are present), the step size is reduced to avoid incorrect cancellation of the near-end speech. The filter length is dynamically set according to the room reverberation time, usually covering the echo path of 50 - 300 ms. The preliminary filtered signal output by the linear processing significantly reduces the main echo components, but may still contain non-linear residual echoes caused by factors such as speaker non-linear distortion, microphone non-linear response, and sampling clock drift. This signal is used as the input for the next stage of non-linear processing to achieve progressive optimization of echo cancellation.
[0090] Step S3: The obtained preliminary filtered signal still contains residual echoes introduced by speaker distortion, microphone non-linearity, and various non-linear factors in the signal processing chain. Non-linear residual echo cancellation uses deep learning-based models such as deep neural networks (DNN), long short-term memory networks (LSTM), or convolutional neural networks (CNN) to construct non-linear mapping functions. These models are trained with a large amount of echo data and can identify the spectral characteristics of residual echoes. The multi-reflection path generation technology simulates the multi-path propagation phenomenon of sound waves in a complex acoustic environment. The system constructs a reflection model based on the room geometry and material properties to generate multiple possible echo paths and their corresponding spectral characteristics. The non-linear cancellation processing also considers the high-order harmonic characteristics of the speaker reference signal and captures the energy migration caused by non-linear distortion through octave analysis. The residual echo estimator combines the state information of linear echo cancellation to accurately distinguish between residual echoes and near-end speech, especially protecting the near-end speech quality in the double-talk state. The non-linearly cancelled signal obtained after processing significantly reduces various complex echo interferences, laying a foundation for subsequent multi-modal residual suppression.
[0091] Step S4: The multi-modal residual suppression technology integrates the features of three different representation modalities in the time domain, frequency domain, and perceptual domain to form a more comprehensive description of the residual echo. The sub-band division process decomposes the full-band signal into multiple critical-bandwidth sub-bands, and each sub-band is processed independently to adapt to the echo characteristics in different frequency regions. The joint decision-making mechanism fuses multiple features, including sub-band energy ratio, phase correlation, spectral smoothness, and modulation spectrum, to accurately distinguish the residual echo from the proximal speech. Based on the decision result, the system applies spectral weighting technology to impose greater attenuation on the time-frequency points with stronger residual echo, while keeping the minimum intervention in the region dominated by the proximal speech. The spectral processing adopts a Bayesian statistical framework, combines the prior probability model to estimate the probability distributions of the echo and speech, and achieves optimal suppression in the sense of minimum mean square error. To reduce musical noise and speech distortion, the system introduces perceptual masking threshold control to ensure that the spectral modification does not exceed the range perceptible to the human ear. During the processing, the spectral continuity and phase consistency are maintained to avoid artificial traces. The output echo suppression signal has removed most of the echo components, while maintaining the naturalness and clarity of the proximal speech.
[0092] Step S5: The real-time environmental sound field change data is obtained through a variety of sensors and signal analysis, including room acoustic parameter estimation, background noise characteristic analysis, and user position tracking. The system monitors the reverberation time change, room mode resonance, and acoustic impedance characteristics to capture environmental changes in a timely manner. Based on these data, the acoustic enhancement module applies dynamic equalization adjustment, adaptive gain control, and spectral shaping technologies to compensate for the speech distortion that may be introduced during the echo cancellation process. Special attention is paid to the restoration of the naturalness of the speech perception, and the timbre characteristics of the speech are maintained through harmonic reconstruction and formant enhancement. For the low signal-to-noise ratio situation, the system integrates a noise suppression algorithm to further improve the speech clarity. The end processing includes dereverberation, dynamic range compression, and psychoacoustic enhancement to enable the output audio to obtain the best listening experience on various playback devices. To reduce the algorithm delay, the system adopts an inter-frame progressive processing strategy to achieve low-latency processing. The final de-echoed audio output has extremely low echo residue, natural speech quality, and stable audio characteristics, and can maintain excellent performance even in the case of significant changes in the acoustic environment.
[0093] Step S6: The adaptive quality assessment module calculates a series of objective metrics, including echo return loss enhancement (ERLE), perceptual evaluation of speech quality (PESQ), and speech intelligibility index (SII), to comprehensively evaluate the echo cancellation effect, speech quality retention, and overall listening experience of the system. Without user intervention, the system automatically fine-tunes the key parameters of each processing stage according to the evaluation results. Long-term operation data analysis identifies different usage scenario patterns, such as conference room mode, home environment mode, and mobile usage mode, and constructs an optimized configuration file for each scenario. The parameter optimization adopts a hybrid strategy, combining a rule-based expert system and machine learning methods, to achieve a balance between achieving the best echo reduction effect and computational resource consumption. The system also has self-learning ability, and gradually improves the internal decision-making model by recording the parameter combinations of successful processing cases. The robustness guarantee mechanism prevents the system from crashing due to extreme conditions, and sets upper limits for parameter changes and a fallback mechanism. This step enables the entire echo reduction system to form a closed-loop control, continuously evolving and adapting to ensure an excellent audio communication experience in the face of changing technical environments and evolving user needs.
[0094] An intelligent echo reduction audio processing method provided by the present invention can more accurately identify and filter basic echo components through time-frequency analysis processing of the sound field signal and the speaker reference signal, combined with dynamically adjusted linear echo cancellation technology, thereby improving the initial effect of echo cancellation and providing a more reliable basis for subsequent processing. By introducing non-linear residual echo cancellation and multi-reflection path generation mechanisms, refined processing of complex echo paths is achieved, which helps to solve the processing defects of traditional methods under non-linear conditions and avoid sound quality distortion. Based on the joint decision-making mechanism of multi-modal residual suppression and sub-band division, the system can operate efficiently under different acoustic environments and echo characteristics, reducing the impact of residual echo on communication quality and improving the overall processing efficiency of the system. By obtaining real-time environmental sound field change data and performing acoustic enhancement, a more intelligent adaptive processing strategy is developed to achieve dynamic optimization of the entire system, thereby effectively improving audio clarity and reducing communication interference. By comprehensively considering the complexity of the actual acoustic environment, the echo cancellation strategy can be flexibly adjusted according to the characteristics and requirements changes of different application scenarios, making the system more adaptable to diverse audio interaction scenarios.
[0095] In one embodiment, the sound field signal and the speaker reference signal are acquired, signal framing and short-time Fourier transform are performed to obtain time-frequency spectrum audio data, including:
[0096] In the initial stage of signal processing, it is necessary to convert analog signals (acoustic field signals and speaker reference signals) into digital signals. This process depends on two parameters: the sampling rate and the bit depth. The sampling rate specifies the number of times the signal is sampled per second, and the bit depth defines the precision of each sampled value. Through these preset conditions, the acoustic field signal and the speaker reference signal are divided into equally spaced digital sampling points in time, thus being transformed into the original digital-domain signal. After sampling, the signal is no longer a continuous analog waveform but a time series composed of a series of discrete digital values, forming the basis of the digital signal.
[0097] The digitized signal still needs to be further segmented for more detailed spectral information in time-frequency analysis. Time-frame segmentation is carried out by setting preset frame length and frame shift parameters. The frame length determines the duration of each segment of the signal, usually represented by a certain number of sampling points. The frame shift parameter sets the degree of overlap between consecutive frames, controlling the sliding of the signal frames on the time axis. The purpose of frame segmentation is to perform local analysis on the overall signal through small segments of the signal, enabling the analysis to capture the details that change over time in the signal. These frame sequences become the basis for subsequent frequency-domain processing.
[0098] Since directly performing a Fourier transform on each frame of the signal will introduce the problem of spectral leakage, that is, the spectral components of the signal will spread to adjacent frequency bands, affecting the analysis accuracy, using a window function can reduce this effect. The Hanning window function is a commonly used windowing technique. It optimizes spectral leakage by applying lower weights to both ends of the signal frame and keeping a higher weight in the middle part of the signal frame. The result of windowing the signal frame is that the window function effectively reduces the spectral aliasing caused by signal segmentation, making the frequency-domain analysis more accurate.
[0099] The fast Fourier transform is a key step in converting a time-domain signal into a frequency-domain signal. Each windowed signal frame undergoes FFT processing, and the result is complex frequency-domain coefficients. The complex frequency-domain coefficients contain the amplitude and phase information of the signal, where the real part represents the amplitude information of the signal and the imaginary part represents the phase information of the signal. The FFT has high computational efficiency and can convert the time-domain signal into the frequency-domain signal in a short time, which provides the basis for subsequent spectral analysis.
[0100] Through the complex frequency-domain coefficients, the amplitude spectrum and phase spectrum of the signal can be extracted respectively. The amplitude spectrum reflects the energy magnitude of the signal at each frequency component and is a description of the signal's distribution in the frequency dimension; the phase spectrum describes the phase information of the signal at different frequency components. The calculation of the amplitude spectrum and phase spectrum is based on the modulus length and phase angle of the complex frequency-domain coefficients, thereby obtaining the frequency components and phase information of the signal.
[0101] Band energy analysis refers to dividing the frequency-domain signal into certain frequency bands and calculating the energy distribution within each frequency band. The band energy distribution matrix records the corresponding energy values for each frequency band. The frequency band division is based on the spectral characteristics of the audio signal. Usually, the frequency domain is divided into several frequency bands with equal widths or set according to actual needs to facilitate the analysis of the energy distribution of different frequency components. The purpose of band energy analysis is to capture the energy information of different frequency bands, which helps with spectral optimization and feature extraction.
[0102] The Mel frequency scale is a non-linear frequency scale designed to mimic the auditory characteristics of the human ear. Through the mapping of the Mel frequency scale, the band energy distribution matrix can be converted into a spectral representation that conforms to the auditory characteristics of the human ear. The Mel frequency scale compresses the low-frequency part while stretching the high-frequency part to improve the frequency resolution in the frequency bands where the human ear is more sensitive. The Mel filter bank coefficients are obtained by mapping the band energy distribution matrix through the Mel frequency scale, which reflects the spectral distribution of the signal on the Mel scale.
[0103] Combining the Mel filter bank coefficients with the complex frequency-domain coefficients constructs a three-dimensional feature. The three-dimensional feature is a data set that combines time, frequency, and feature information and is usually used for subsequent signal denoising, classification, or other processing tasks. By combining the Mel filter bank coefficients and the frequency-domain coefficients, the characteristics of the signal can be comprehensively described in the time-frequency domain, obtaining a comprehensive and information-rich time-frequency spectrum audio data. This three-dimensional feature not only retains the time structure of the audio signal but also captures the energy changes of the signal at different frequencies, thus providing strong support for further audio analysis.
[0104] In this embodiment, through the intelligent downmix audio processing of the sound field signal and the speaker reference signal, it is possible to effectively optimize spectral leakage, improve the accuracy of band energy analysis, and achieve a spectral representation that conforms to the auditory characteristics of the human ear through the Mel frequency scale mapping, thereby improving the processing quality of the audio signal. By adopting synchronous digital sampling and frame processing, it can better adapt to signal changes on different time scales, ensuring that the signal is fully analyzed in the time-frequency domain and providing accurate data support for subsequent denoising and signal enhancement. The use of the Mel filter bank coefficients can convert the spectral characteristics of the audio signal into a format that is more in line with the auditory characteristics of the human ear, effectively improving the natural perception of audio analysis. Through this processing method, the frequency characteristics of the audio signal are more accurately characterized, reducing the errors caused by spectral leakage and time-frequency resolution problems, making the denoising process more accurate and effectively improving the audio quality. The application of this method in intelligent audio systems can improve the intelligent level of audio processing, especially in complex noise environments, significantly improving the clarity and restoration degree of audio signals.
[0105] In one embodiment, linear echo cancellation processing is performed on the time-frequency spectrum audio data, and the convergence speed is dynamically adjusted to obtain a preliminary filtered signal, including:
[0106] The time-frequency spectrum audio data is divided in the frequency domain through a polyphase filter bank. The multi-subband decomposition adopts the critical bandwidth division criterion, and the full frequency band is divided into 32 non-uniformly distributed frequency subband signals. The bandwidth of each frequency subband signal follows the equivalent rectangular bandwidth (ERB) auditory model to ensure that the subband division matches the human ear auditory characteristics. The prototype filter of the polyphase filter bank uses a cosine-modulated filter with 60 dB stopband attenuation, and a 15% overlap region is set between adjacent subbands to avoid frequency domain leakage. After the processing is completed, the original time-frequency spectrum audio data is converted into multiple frequency subband signals, and each subband signal contains independent time-frequency energy distribution characteristics.
[0107] Parallel processing is performed on the multiple frequency subband signals, and a frequency domain block least mean square (FDAF) algorithm is used to construct a linear echo path model. Within each subband, the cross-correlation operation is performed between the speaker reference signal and the sound field signal to calculate the filter impulse response with a length of 256 points. The filter order is dynamically adjusted according to the subband center frequency. A 128th-order filter is used for the low-frequency subbands to capture long-delay reflections, and a 64th-order filter is used for the high-frequency subbands to reduce the computational complexity. Through the block update method, the filter coefficient matrix is updated once every 8 time-frequency frames. A forgetting factor of 0.95 is introduced during the update process to balance the tracking speed and stability. The linear echo path signal is generated through the processing, and this signal characterizes the transmission characteristics of the direct sound and early reflections from the speaker to the microphone.
[0108] Based on the short-time coherence function of the sound field signal and the speaker reference signal, the instantaneous degree of mismatch metric is calculated. In the time-frequency domain, for each frequency subband signal with a frame length of 20 ms as a unit, the amplitude coherence coefficient between the two signals is calculated. The sliding window method is used for the calculation of the amplitude coherence coefficient, and the window length is 5 frames. By comparing the coherence change rate between the current frame and the historical frame, the convergence state parameter is generated. The value range of the convergence state parameter is [0, 1]. When the parameter value is lower than 0.3, it is determined as the under-convergence state, and when it is higher than 0.7, it is determined as the over-convergence state. This parameter reflects the tracking performance of the adaptive filter in real time and provides a quantitative basis for the step size adjustment.
[0109] According to the non-linear mapping relationship of the convergence state parameter, a piecewise proportional-integral (PI) controller is used to generate an optimized step size value. In the under-convergence region (parameter < 0.3), the integral gain is set to 0.8 and the proportional gain is set to 1.2 to accelerate the convergence speed; in the stable convergence region (0.3 ≤ parameter ≤ 0.7), the integral gain is reduced to 0.2 and the proportional gain is set to 0.5 to maintain the convergence accuracy; in the over-convergence region (parameter > 0.7), the integral term is turned off and a proportional gain of 0.1 is enabled to prevent the filter from diverging. A saturation limiting mechanism is introduced during the adjustment process to ensure that the step size value is always within the effective range of [0.001, 0.1]. The output frequency of the optimized step size value is synchronized with the filter update period and is updated every 8 time-frequency frames.
[0110] Dynamic range compression is performed on the optimized step size value, and a hyperbolic tangent function is used to achieve non-linear mapping. The upper threshold of the step size is set to 0.05, and the lower threshold is set to 0.005. When the input step size value exceeds the upper limit, the output value is fixed at 1.2 times the upper limit; when it is lower than the lower limit, the output value is fixed at 0.8 times the lower limit. Within the threshold interval, the step size value is scaled according to a logarithmic relationship, and the scaling formula is: stable step size parameter = lower threshold + (upper threshold - lower threshold) × log10(optimized step size value / lower threshold); this processing eliminates the sudden change interference of the step size parameter and generates a stable step size parameter with smooth transition characteristics.
[0111] The stable step size parameter is used to control the update process of the frequency-domain block LMS algorithm, and filtering operations are performed on the time-frequency spectrum audio data. Within each frequency sub-band, the linear echo path signal is convolved with the speaker reference signal to generate a linear echo estimation signal. The overlapping save method is used for the convolution operation, with a 50% frame overlap rate combined with a Hann window function to avoid time-domain aliasing. During the filtering process, the time delay change of the echo path signal is monitored in real time. When the time delay deviation exceeds 2 ms, the filter coefficient re-initialization mechanism is triggered to prevent performance degradation caused by environmental mutations.
[0112] After subtracting the linear echo estimation signal from the original time-frequency spectrum audio data, residual component processing is performed. Residual correlation compensation includes two parallel operations: amplitude compensation uses the sub-band energy matching method, and the energy ratio of the residual signal to the original signal is calculated within each frequency sub-band. When the energy ratio exceeds -15 dB, the amplitude of the residual signal is increased proportionally; phase compensation uses complex domain least squares fitting to correct the phase offset of the residual signal. The compensated signal is reconstructed into a time-domain waveform through a synthesis filter bank, and a preliminary filtered signal is output. This signal retains the target speech component and suppresses the linear echo component by at least 25 dB.
[0113] In this embodiment, by introducing the multi - sub - band decomposition technology into the intelligent echo - reduction audio processing method, the full - band audio signal can be effectively divided into 32 frequency sub - bands. The bandwidth of each sub - band signal matches the auditory characteristics of the human ear, thereby improving the accuracy of echo cancellation and the naturalness of the human ear's listening perception. The use of the polyphase filter bank ensures the accuracy of frequency - domain division and reduces frequency - domain leakage through a 15% sub - band overlapping region, enhancing the frequency resolution ability. A linear echo - path model is constructed by the frequency - domain block least - mean - square (FDAF) algorithm, and the filter order is dynamically adjusted within each sub - band, enabling the low - frequency sub - bands to capture long - delay reflections while the high - frequency sub - bands reduce the computational complexity, thus achieving efficient and accurate echo - path analysis. The method of updating the filter coefficient matrix in blocks combined with the use of the forgetting factor ensures the stability and real - time performance of the algorithm.
[0114] In one embodiment, non - linear residual echo cancellation is performed on the preliminarily filtered signal, and multi - reflection path generation is carried out to obtain a non - linear cancellation signal, including:
[0115] Non - linear residual echo cancellation is performed on the preliminarily filtered signal. The preliminarily filtered signal is usually the signal processed by the traditional echo - suppression algorithm. Although it has been preliminarily processed, it still contains certain residual echo components. To further eliminate these residual echoes, non - linear residual echo cancellation technology is adopted. This process analyzes the non - linear characteristics of the signal, identifies and extracts those echo components, and continuously adjusts the parameters of the filter during the optimization process to enable it to more precisely eliminate these echoes. In this way, the obtained signal is the echo signal after preliminary elimination, usually called the non - linear cancellation signal.
[0116] On the basis of eliminating the echo signal, the coupled - feature matrix is then constructed. The coupled - feature matrix is obtained by analyzing the relationship between the characteristics of the preliminarily filtered signal and the echo components, and extracting the key features of the signal. These features can help judge the performance of the echo components in the signal, and thus guide the subsequent echo - cancellation process. The method of constructing the coupled - feature matrix usually analyzes the time - domain and frequency - domain characteristics of the signal, combines its time - delay and attenuation characteristics with the echo, and obtains the relevant feature information of the signal.
[0117] After obtaining the coupled - feature matrix, the Volterra - kernel extended space is further constructed. The Volterra - kernel extended space is a mathematical tool for modeling signals through multi - order convolution operations, which can describe the non - linear changes of signals during multiple reflection processes. In this space, the non - linear characteristics of the signal in a complex propagation environment can be captured, and it provides support for the subsequent optimized filter design. By further processing the coupled - feature matrix, a multi - dimensional extended space is constructed, which can describe the non - linear changes of the signal, thus providing a basis for optimizing the filter parameters.
[0118] After constructing the Volterra kernel extended space, the next step is to update the initially filtered signal in real time to obtain more accurate filter parameters. This process adjusts the filter parameters in real time based on the features in the extended space to eliminate finer echo components. At each moment, the filter is optimized according to the current signal state and the constructed extended space, so that the filter response better meets the actual requirements and gradually approaches the optimal solution. The key to this step lies in the ability to adjust the filter in real time to ensure that the echo cancellation effect is continuously optimized over time.
[0119] After being processed by the optimized filter, the echo components in the signal have been effectively suppressed. However, in some complex environments, there may be multipath propagation phenomena. Multipath propagation refers to the situation where sound signals reach the receiving end through multiple different paths due to reflection, refraction, etc., resulting in the superposition of multiple reflected signals, which further affects the clarity of speech. To effectively identify and suppress these reflected signals, it is necessary to analyze the multi-reflection paths.
[0120] By analyzing the multi-reflection paths of the optimized signal, different reflection paths in the signal can be identified. The characteristics of these paths are mainly reflected in the signal delay, amplitude, and attenuation mode. After identifying these paths, an "adversarial reflection path set" can be accurately constructed. This set contains the reflection paths that have the greatest impact on the received signal, and they have a relatively high degree of interference to the signal, so they need to be focused on for processing.
[0121] After obtaining the adversarial reflection path set, the next step is to perform time-frequency domain superposition of these paths with the initially filtered signal. Time-frequency domain superposition combines the characteristics of the signal in the time domain and the frequency domain, and superposes these reflected path signals so as to comprehensively suppress these interference components both in frequency and time. Through time-frequency domain superposition, the influence of these multipath signals can be effectively reduced, thereby further improving the signal quality.
[0122] The obtained signal is called the multipath suppression signal, which has removed most of the reflections and multipath interferences through time-frequency domain superposition. At this time, the remaining noise components in the signal will mainly be of low intensity and more in line with the characteristics of natural speech.
[0123] To further improve the speech quality and suppress the low-intensity noise that is difficult to eliminate by traditional methods, the next step is to perform sub-band energy distribution analysis on the multipath suppression signal. The purpose of sub-band energy distribution analysis is to decompose the signal into several sub-band signals, and by analyzing the energy distribution of each sub-band, identify which frequency bands in the signal have more prominent noise components.
[0124] Based on the analysis results of the sub-band energy distribution, a dynamic time-frequency masking matrix is generated. This matrix can dynamically adjust the suppression degree of each sub-band according to the energy distribution at different times and frequencies. For those frequency bands with strong noise, the masking matrix will give stronger suppression; for those with weak noise, it will give lighter suppression. This dynamic adjustment mechanism can maximize the removal of noise while ensuring speech clarity.
[0125] According to the generated dynamic time-frequency masking matrix, a point-by-point multiplication operation is performed on the multipath suppression signal. The point-by-point multiplication operation multiplies each element in the time-frequency masking matrix by the corresponding frequency component of the multipath suppression signal, thereby precisely suppressing the frequency of the signal. This operation can effectively suppress the remaining noise at all levels in the time domain and frequency domain, and optimize the signal quality. Through the point-by-point multiplication operation, the finally obtained signal is the non-linear cancellation signal. This signal has undergone multi-level processing, including echo cancellation, reflection path analysis, time-frequency domain suppression, etc., and has greatly improved the quality and clarity of the speech.
[0126] In this embodiment, by combining non-linear residual echo cancellation and multi-reflection path generation technology, the quality of the audio signal is effectively improved. By performing non-linear residual echo cancellation on the preliminarily filtered signal, the echo components can be significantly removed, avoiding the problem of echo residues in traditional echo cancellation methods, thereby improving the clarity of the speech. The construction of the non-linear residual basis and the application of the Volterra kernel expansion space enable this method to accurately capture the complex non-linear characteristics in the signal, further optimizing the filter parameters, enabling it to dynamically adapt to echo changes in different environments, and providing a more efficient echo cancellation effect. Through multi-reflection path analysis and the generation of adversarial reflection path sets, the interference of multi-path propagation on the audio signal can be effectively identified and eliminated. The time-frequency domain superposition technology further reduces multi-path interference and improves the overall quality of the signal. The generation of the dynamic time-frequency masking matrix precisely suppresses the noise in different frequency bands, thereby ensuring the naturalness and clarity of the speech.
[0127] In one embodiment, a Volterra kernel expansion space is constructed according to the coupling feature matrix, and the preliminarily filtered signal is updated in real time to obtain optimized filter parameters, including:
[0128] The coupled feature matrix is obtained by extracting features from the audio signal during the preprocessing stage. Each moment of the audio signal is analyzed to extract multi-dimensional feature data, which reflects various information such as the frequency, amplitude, and phase of the signal. The coupled feature matrix is composed of these feature information, where each row represents the feature information within a time window, and each column represents the variation of the feature in different time windows. The purpose of constructing the coupled feature matrix is to comprehensively describe the complex characteristics of the signal and provide a complete data basis for subsequent processing.
[0129] By performing singular value decomposition (SVD) on the coupled feature matrix, the matrix is decomposed into the product of three matrices: a left singular matrix, a diagonal matrix, and a right singular matrix. The goal of singular value decomposition is to find the principal components of the matrix, thereby achieving dimensionality reduction. In signal processing, dimensionality reduction can not only reduce the computational complexity but also help remove noise and unnecessary redundant information, retaining the main effective features. After dimensionality reduction, a low-dimensional feature space is obtained, in which the main features of the signal are better represented.
[0130] In the dimensionality-reduced feature space, separated component construction is performed to obtain a separated vector group. The main task of this step is to further process the dimensionality-reduced feature space and extract the most representative components of the audio signal. Through separation algorithms (such as independent component analysis ICA), the dimensionality-reduced data set is decomposed into several mutually independent components. Each component represents an independent feature in the signal, and the separated vector group is the set of these components. The separated vector group will provide the basis for subsequent kernel function construction.
[0131] Using the separated vector group, the second-order expansion term of the Volterra kernel function is further constructed. The Volterra series is a polynomial model that contains the various-order nonlinear relationships of the signal. When constructing the second-order expansion term, a weighted summation method of the separated vector group is adopted to obtain a second-order kernel expansion matrix. This matrix represents the interaction of the second-order nonlinear features in the signal. In this way, the second-order nonlinear effects in the signal can be captured more accurately, providing a more effective mathematical model for echo cancellation.
[0132] Among them, the formula for constructing the second-order kernel expansion matrix includes ;
[0133] represents the second-order kernel expansion matrix, is the weighting coefficient, and are two components respectively extracted from the separated vector group, represents the tensor product operation. is the total number of separated vector groups, that is, the number of independent components extracted from the audio signal. i is an index variable used to traverse all separated vectors.
[0134] This formula is used to construct the second-order non-linear feature matrix in the signal by weighted summation.
[0135] When continuing to construct the third-order kernel extension tensor, the separated vector group and the second-order kernel extension matrix are used to construct the third-order cross terms. At this time, using the tensor product operation, the separated vector group and the second-order kernel extension matrix are combined to generate a third-order kernel extension tensor. The tensor product operation is a method of operation between multi-dimensional arrays, which can combine features of different orders and further improve the expression ability of the model. Through the third-order kernel extension tensor, more complex high-order non-linear relationships in the signal can be captured, thus making the echo reduction process more accurate.
[0136] After obtaining the third-order kernel extension tensor, an inner product operation is performed to construct the spatial mapping function. The role of the spatial mapping function is to map the audio signal to a new space, in which the non-linear features of each order of the signal are fully expressed. Through the inner product operation, the relationship between the second-order kernel extension matrix and the third-order kernel extension tensor is transformed into the spatial mapping function, which can effectively represent the high-order non-linear features of the signal. The mapped signal has a more distinct structure in the new space, which helps to further optimize the filter parameters.
[0137] The spatial mapping function is used to perform non-linear projection on the preliminarily filtered signal. The projection operation maps the signal from the original space to the new space, in which the non-linear features of each order of the signal are enhanced. The result of this process is to obtain the projection coefficient vector, which contains the main information of the signal in the new space. These coefficient vectors will be used as key inputs in the subsequent optimization steps.
[0138] The projection coefficient vector will undergo least squares iterative optimization. The least squares method is a common optimization algorithm that continuously adjusts the filter parameters by minimizing the error between the predicted result and the actual result. Here, the goal of the least squares iterative optimization is to find the optimal filter parameters to achieve the best echo cancellation effect. This process requires continuous iteration until it converges to the minimum error value. Finally, the optimized filter parameters will be used for the echo reduction processing of the signal.
[0139] In this embodiment, by introducing a coupling feature matrix into the audio processing method, it is possible to comprehensively and accurately describe the characteristics of the audio signal such as frequency, amplitude, and phase, providing a good data basis for subsequent processing, thereby improving the processing accuracy. By performing singular value decomposition (SVD) on the coupling feature matrix, the computational complexity can be effectively reduced, and noise and redundant information can be removed, ensuring the retention and optimization of the main features. Using independent component analysis (ICA) to separate the independent feature components of the audio signal further enhances the accuracy of signal processing. Constructing second-order and third-order kernel extension matrices and tensors, and capturing the complex second-order and higher-order nonlinear relationships in the signal through tensor product operations can significantly improve the effect of echo cancellation. The spatial mapping function constructed through the inner product operation maps the signal into a new feature space, enabling the more complete expression of nonlinear features, thereby optimizing the adjustment of filter parameters. In the final least squares iterative optimization process, the filter parameters are dynamically adjusted to ensure the optimization of the echo cancellation effect and improve the clarity of the audio signal. Through the organic combination of multiple steps, the entire method achieves efficient and accurate audio echo cancellation, significantly improving the audio signal quality and user experience.
[0140] In one embodiment, multi-modal residual suppression is performed on the non-linear cancellation signal, and spectrum processing is carried out through sub-band division and joint decision-making to obtain an echo suppression signal, including:
[0141] The first step of processing the signal is to perform sub-band decomposition. Sub-band decomposition divides the original audio signal into multiple sub-bands according to the frequency range, obtaining a non-uniform sub-band representation. Through sub-band decomposition, the different frequency components of the signal are separated, making the processing of each frequency band more accurate. The non-uniform sub-band representation means that the band density in different frequency ranges is uneven, and more refined adjustment is made according to the characteristics of the frequency distribution, thereby improving the echo suppression effect.
[0142] The non-uniform sub-band representation will perform multi-modal feature extraction and integration. The goal of this step is to extract the feature information of the signal from multiple dimensions, such as time domain, frequency domain, and phase information, etc., to form a multi-modal feature set. These features can reflect various aspects of the signal and contribute to subsequent echo suppression decisions. The multi-modal feature set refers to a set that comprehensively combines multiple feature information, and these information are different perception methods of the signal state, which helps to achieve more accurate echo suppression processing.
[0143] Based on the multi-modal feature set, cross-entropy and mutual information are calculated to obtain the sub-band grouping structure. Cross-entropy and mutual information are important mathematical tools for measuring the similarity and dependence relationship between signal features. Cross-entropy reflects the information difference between different sub-bands, while mutual information is used to measure the mutual dependence degree between different features. Through calculation, the correlation between each sub-band can be determined, thereby obtaining an optimized sub-band grouping structure. The sub-band grouping structure refers to the reasonable grouping of the sub-bands after spectrum decomposition according to their similarity or information dependence relationship.
[0144] After obtaining the sub-band grouping structure, the non-uniform sub-band representation is adaptively grouped according to this structure to form multiple groups of sub-band sets. The process of adaptive grouping will automatically adjust the grouping rules according to the specific characteristics of the signal, so that each sub-band set can better represent the actual situation of the signal. These adaptively grouped sub-band sets have a more effective suppression effect in the subsequent processing stage.
[0145] The decision rule evaluation is performed on each of the multiple groups of sub-band sets to obtain multiple suppression decision values. At this stage, each group of sub-band sets will be evaluated according to the set decision rules to generate corresponding suppression decision values. The decision rules are based on the understanding of the signal characteristics and processing objectives, including the intensity of the signal, noise characteristics, etc., to determine whether further echo suppression is required. The decision value is the result obtained after processing the signal according to the preset rules, which determines whether and how to perform echo suppression.
[0146] All the multiple suppression decision values will undergo joint decision processing to obtain the final target suppression strategy. The process of joint decision processing is to combine the suppression decisions of different sub-bands to form a unified suppression strategy. In this process, the suppression decision values of each sub-band are involved, and the overall echo suppression effect is comprehensively evaluated to finally obtain the optimal suppression strategy. The target suppression strategy refers to the optimal echo suppression scheme provided by the system for the entire audio signal, which can minimize the echo and improve the sound quality to the greatest extent.
[0147] According to the target suppression strategy, the system performs differential processing on the multiple groups of sub-band sets to obtain the modal collaborative suppression signal. In this process, each sub-band set is adjusted specifically according to different suppression strategies. For example, for the frequency band with more obvious echo, a stronger suppression strategy is adopted, while for the frequency band with lighter echo, a relatively mild processing is adopted. The purpose of this differential processing is to flexibly adjust according to the different degrees of echo in order to maximize the elimination of echo while maintaining the sound quality.
[0148] The modal collaborative suppression signal is the result after differential processing. It integrates the characteristics of each sub-band and can effectively eliminate the echo component. Next, variational decomposition phase compensation is performed on the modal collaborative suppression signal. Variational decomposition is a technique in signal processing that aims to decompose a signal into multiple components for independent processing of each component. During this process, the system compensates the phase of the signal to ensure that the phases of all frequency bands of the signal are coordinated, thereby avoiding sound quality loss caused by phase misalignment.
[0149] Asymmetric singular value decomposition reconstruction is performed on the phase-coordinated signal to obtain the echo suppression signal. Asymmetric singular value decomposition is an efficient signal reconstruction method that can denoise and remove echoes from complex signals while maintaining the original characteristics of the signal. During the reconstruction process, the system analyzes and adjusts the singular values of the signal to ensure that echoes are effectively suppressed while maximizing the preservation of the naturalness and clarity of the audio. The echo suppression signal is an audio signal obtained after a series of processes, with its echo component greatly reduced and the sound quality significantly improved.
[0150] In this embodiment, by adopting the multi-modal residual suppression technology, fine sub-band decomposition and spectrum processing are performed on the non-linear cancellation signal, effectively improving the accuracy and effect of echo suppression. The non-uniform representation and multi-modal feature extraction of each sub-band can capture the multi-dimensional information of the signal more comprehensively, thereby implementing a more accurate echo suppression strategy. The introduction of cross-entropy and mutual information calculations optimizes the sub-band grouping structure, making the adaptive grouping more intelligent and further enhancing the system's ability to identify and suppress echoes. Through the evaluation of the decision rules of multiple sub-band sets, multiple suppression decision values are obtained, and through joint decision processing, the optimality and consistency of the final suppression strategy are ensured. Such a processing method not only improves the accuracy of echo suppression but also performs differential processing according to the characteristics of different frequency bands, avoiding sound quality loss caused by over-suppression. At the same time, the combination of variational decomposition phase compensation and asymmetric singular value decomposition reconstruction ensures the phase coordination of the signal and the restoration of the overall sound quality.
[0151] In one embodiment, real-time environmental sound field change data is acquired, and acoustic enhancement is performed on the echo suppression signal to obtain the target reduced echo audio output, including:
[0152] Obtain environmental sensor data. The environmental sensor data includes audio signals collected by microphones, sound pressure level data, directivity data, etc. Through multiple sensors, sound information at different positions and directions in the environment can be obtained. These data are preprocessed to remove interference signals and unnecessary noise, and then the short-time Fourier transform (STFT) is used to transform the time-domain signal into the frequency domain to obtain spectral information. According to the sensor positions and the sound propagation model, spatial spectral decomposition is performed on the spectral information to obtain the spectral components in each direction, and sub-band analysis is carried out. The spectrum is divided into several sub-bands, and the characteristics of each sub-band are analyzed to obtain detailed spectral features.
[0153] Calculate the acoustic parameter changes for the environmental sound field spectral features and the environmental sensor data to obtain real-time environmental sound field change data. Extract instantaneous parameters such as sound pressure level, frequency, and phase from the environmental sensor data, and compare these instantaneous parameters with the environmental sound field spectral features to analyze their change trends over time. Through time series analysis methods, predict the sound field changes in the future for a period of time, and compare the current and predicted sound field data to calculate the real-time environmental sound field change data.
[0154] Perform multi-stage cascaded filtering on the echo suppression signal according to the real-time environmental sound field change data to obtain a preliminary enhanced signal. Use the real-time environmental sound field change data to set the parameters of each filter. Use the primary filter to perform preliminary processing on the signal to filter out obvious echo components. Perform cascaded processing through multiple filters in sequence. Each stage of the filter optimizes the signal quality in turn, and finally outputs a preliminary enhanced signal, which has significantly reduced echoes.
[0155] Perform phase correction and spectral smoothing on the preliminary enhanced signal, and adjust the gain factor according to the real-time environmental sound field change data to obtain a frequency-domain enhanced signal. Use a phase correction algorithm to adjust the phase of the preliminary enhanced signal to make the phase of the signal more consistent; apply spectral smoothing technology to smooth the signal spectrum and reduce the sharp changes in the spectrum; according to the real-time environmental sound field change data, dynamically adjust the gain factor to optimize the intensity and equalization of the signal, and output a frequency-domain enhanced signal to make the spectrum of the signal smoother and the quality significantly improved.
[0156] Perform residual noise selective suppression on the frequency-domain enhanced signal and the real-time environmental sound field change data to obtain a noise suppression signal. Identify the residual noise components in the frequency-domain enhanced signal, use the real-time environmental sound field change data to analyze the characteristics and change trends of the noise, apply a selective suppression algorithm, and select appropriate suppression strategies for different types of noise, and output a noise suppression signal, which significantly reduces the residual noise.
[0157] Perform auditory band energy adjustment on the noise suppression signal to obtain a time-domain enhanced signal. Analyze the spectral characteristics of the noise suppression signal, determine the energy distribution of each frequency band, adjust the energy of each frequency band according to the human ear's auditory characteristics to make the signal more in line with auditory habits, apply the auditory masking effect, optimize the frequency band energy distribution of the signal, and output the time-domain enhanced signal with a more natural sound quality.
[0158] Perform harmonic reconstruction and formant enhancement processing on the time-domain enhanced signal to obtain the target downsampled audio output. Analyze the fundamental frequency and harmonic components of the time-domain enhanced signal, use harmonic reconstruction technology to reconstruct the harmonic components of the signal, enhance the harmonic structure of the signal, identify the formant positions and characteristics in the signal, apply the formant enhancement algorithm to strengthen the formant components of the signal, make the sound quality of the signal more full and layered, and finally output the target downsampled audio signal to make the signal quality reach the best state.
[0159] In this embodiment, through multi-level processing of the environmental sensor data, various types of sound information in the environment can be accurately captured, and the echo suppression signal can be adjusted in real time, effectively reducing the impact of environmental noise and echo, thereby improving the clarity and quality of the audio signal. Through spatial spectrum decomposition and sub-band analysis, the sound characteristics in different directions and frequency bands can be analyzed in detail, providing accurate spectral feature support for subsequent audio processing, and effectively improving the system's adaptability to complex environmental sound fields. The use of multi-stage cascaded filtering processing makes the suppression of echo and noise more refined, further enhancing the signal quality. In addition, through processing steps such as phase correction, spectral smoothing, and gain adjustment, the smoothness and balance of the signal are significantly improved, which helps to improve the naturalness and auditory comfort of the sound output. Through selective suppression of residual noise and auditory band energy adjustment, the interference of background noise is effectively reduced, making the finally output audio signal more in line with the auditory characteristics of the human ear and improving the fineness and layering of the sound quality.
[0160] In one embodiment, perform selective suppression of residual noise on the frequency-domain enhanced signal and real-time environmental sound field change data to obtain a noise suppression signal, including:
[0161] The frequency-domain enhanced signal is extracted from the audio source signal and processed to enhance the main audio components in the signal, especially in a noisy environment, to ensure the clarity of important information (such as speech). After the frequency-domain enhanced signal undergoes frequency band region division, the frequency components of the signal are divided into multiple independent frequency bands, and each frequency band contains signals within a specific frequency range. Through this division method, the processing of different frequency components can be more refined, specifically reducing the impact of noise on specific frequency bands. The goal of this step is to convert the signal into a frequency-domain representation that is convenient for subsequent processing, allowing independent analysis and processing of each frequency band.
[0162] The cross-correlation analysis is performed on the signal after frequency band division and the real-time ambient sound field change data. By calculating the correlation between two signals (the frequency band divided signal and the ambient sound field change data), the cross-correlation analysis can reveal the influence of the ambient sound field change on the noise components in the signal. This analysis helps to identify which frequency bands may contain more complex noise components and the temporal variation patterns of the noise components. Through the cross-correlation analysis, the distribution characteristics of the residual noise can be obtained, which will provide key information for subsequent noise suppression.
[0163] Based on the obtained distribution characteristics of the residual noise, the frequency band selective masking analysis is carried out. This step determines which frequency bands have obvious noise by analyzing the intensity and distribution of the residual noise in different frequency bands, and performs masking processing on these frequency bands. Masking means reducing or completely suppressing the noise components in the frequency band to improve the clarity of the signal. The result of the frequency band selective masking analysis is the frequency band noise marked signal, which indicates the noise regions that need to be processed in the frequency band and helps with subsequent noise reduction operations.
[0164] The frequency band noise marked signal is used for Bayesian probability estimation. Bayesian probability estimation combines the ambient sound field change data and the noise marking information, and calculates the probability of whether each frequency band region belongs to noise according to Bayes' formula. This analysis process obtains a noise probability distribution map, which shows the probability of noise existence in each frequency band. Based on this map, it can be further determined which frequency bands need to be focused on for noise reduction.
[0165] Based on the noise probability distribution map, the non-linear suppression analysis is performed on the frequency domain enhanced signal. Non-linear suppression is a common method in signal processing. It suppresses the more prominent noise components in the frequency band according to the noise probability of each frequency band in the noise probability distribution map. During the non-linear suppression analysis process, the suppression intensity can be automatically adjusted according to the noise intensity, with stronger suppression for frequency bands with stronger noise and appropriate retention for frequency bands with weaker noise. Through this process, a preliminary noise suppression signal is obtained, that is, the noise has been partially suppressed, but certain audio information is still retained.
[0166] The preliminary noise suppression signal is further combined with the real-time ambient sound field change data for spectral subtraction balance processing. Spectral subtraction balance performs subtraction operations on the signal through spectral analysis to reduce the abnormal components caused by noise in the signal. At the same time, it combines with the real-time ambient sound field change data to adjust the balance parameters in the spectral subtraction process, making the noise suppression more accurate. The signal after spectral subtraction balance is the balance suppression signal, which has a stronger noise suppression effect than the preliminary noise suppression signal and retains more original audio content.
[0167] After obtaining the balanced suppression signal, the next step is to perform artifact threshold clearing. Artifact clearing is a processing method for removing the untrue noise components in the signal. By setting the artifact clearing threshold, those parts that seem to be noise but actually belong to signal errors are removed. The artifact removal coefficient obtained after artifact threshold clearing represents the signal strength after removing artifacts. At this stage, the invalid noise components in the signal have been effectively removed.
[0168] At this stage, through a series of spectral analyses of the signal (such as Fourier transform or wavelet transform), the components with sudden, discrete, extreme or irregular changes in the signal are identified. These mutated components are usually artifacts. By comparing with the waveform and spectral characteristics of the real signal, the characteristics of the artifacts are extracted. The artifact threshold is usually set according to the statistical characteristics of the signal (such as mean, variance, peak, etc.). The purpose of setting the threshold is to distinguish artifacts from real noise or valid signals. This threshold can be a fixed value or can be dynamically adjusted according to the real-time changes of the signal. For example, if the spectral amplitude of a signal changes very violently and is far from its normal waveform, it may be judged as an artifact, and the part exceeding a certain amplitude threshold will be marked as an artifact.
[0169] After the threshold is set, the artifacts are removed through a clearing algorithm. In the frequency domain, each frequency component of the signal has an amplitude and a phase information. When clearing artifacts, first the signal amplitude is detected and compared with the set threshold, and the components with too large amplitude or too fast change (i.e., possible artifacts) are removed. For the frequency components outside the threshold range, usually their amplitudes are set to zero, or their values are reduced through smoothing to avoid excessive influence on the natural components of the signal.
[0170] After artifact clearing, the remaining signal is the artifact removal coefficient. At this time, the noise components of the signal (especially the artifact components) have been effectively removed. The artifact removal coefficient usually includes the adjusted amplitude of each frequency component (the smoother amplitude after removing artifacts), and the phase information of the original signal. Through wavelet reconstruction or other signal reconstruction methods, these artifact-removed amplitudes are combined with the original phase information to finally restore the time-domain signal.
[0171] During the process of reconstructing the artifact-removed signal, detailed adjustments are also required to ensure that the artifact removal process does not introduce other distortions or unnatural changes. For example, further smoothing of the phase is performed to avoid unnatural phase jumps during reconstruction.
[0172] After wavelet reconstruction and phase preservation processing of the artifact removal coefficient, the noise suppression signal is finally obtained. Wavelet reconstruction is to reconstruct the processed signal coefficients into a time-domain signal through wavelet transform, and phase preservation ensures that the phase information of the signal is not lost during the reconstruction process.
[0173] The combination of the frequency-domain enhanced signal and the real-time environmental sound field change data in this embodiment can achieve efficient noise suppression and speech enhancement. Through the frequency band division and cross-correlation analysis of the audio signal, the noise components in different frequency bands are accurately identified, enabling the noise reduction process to be targeted and avoiding the global suppression of the entire signal in traditional methods, thereby improving the audio quality and the retention of details. By combining Bayesian probability estimation and nonlinear suppression techniques, the intensity of noise suppression can be dynamically adjusted to ensure the balance between noise and signal, avoiding the situation of audio distortion or information loss caused by over-suppression. Further spectral subtraction balance processing and artifact threshold clearing techniques effectively remove artifacts and invalid noise, ensuring the clarity and naturalness of the audio signal. Combining wavelet reconstruction and phase preservation processing ensures that important phase information is not lost during the signal reconstruction process, making the finally output noise suppression signal retain both the fineness of the original audio and achieve an efficient noise suppression effect.
[0174] Referring to Figure 2 As shown in the figure, the present invention also provides an intelligent reverberation reduction audio processing system, which is applied to the intelligent reverberation reduction audio processing method of any one of the above, and includes:
[0175] An acquisition module, which is used to acquire the sound field signal and the speaker reference signal, and perform signal framing and short-time Fourier transform to obtain time-frequency spectrum audio data;
[0176] An analysis module, which is used to perform linear echo cancellation processing on the time-frequency spectrum audio data and dynamically adjust the convergence speed to obtain a preliminary filtered signal;
[0177] An association module, which is used to perform nonlinear residual echo cancellation on the preliminary filtered signal and generate multiple reflection paths to obtain a nonlinearly cancelled signal;
[0178] A processing module, which is used to perform multi-modal residual suppression on the nonlinearly cancelled signal, and perform spectrum processing through sub-band division and joint decision to obtain an echo suppression signal;
[0179] A control module, which is used to acquire the real-time environmental sound field change data and perform acoustic enhancement on the echo suppression signal to obtain the target reverberation reduction audio output.
[0180] An intelligent voice reduction audio processing system provided by the present invention can more accurately identify and filter out basic echo components through time-frequency analysis and processing of sound field signals and speaker reference signals, combined with dynamically adjusted linear echo cancellation technology, thereby improving the initial effect of echo cancellation and providing a more reliable basis for subsequent processing. By introducing non-linear residual echo cancellation and multi-reflection path generation mechanisms, refined processing of complex echo paths is achieved, which helps to solve the processing defects of traditional methods under non-linear conditions and avoid sound quality distortion. Based on the joint decision-making mechanism of multi-modal residual suppression and sub-band division, the system can ensure efficient operation under different acoustic environments and echo characteristics, reduce the impact of residual echo on communication quality, and improve the overall processing efficiency of the system. By obtaining real-time environmental sound field change data and performing acoustic enhancement, a more intelligent adaptive processing strategy is formulated to achieve dynamic optimization of the entire system, thereby effectively improving audio clarity and reducing communication interference. By comprehensively considering the complexity of the actual acoustic environment, the echo cancellation strategy can be flexibly adjusted according to the characteristics and requirements changes of different application scenarios, making the system more adaptable to diverse audio interaction scenarios.
[0181] It should be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described system and each module can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0182] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. An intelligent echo reduction audio processing method, characterized in that: include: Acquire the sound field signal and the speaker reference signal, perform signal framing and short-time Fourier transform, and obtain time-frequency spectrum audio data; Performing linear echo cancellation processing on the time-frequency spectrum audio data, and dynamically adjusting the convergence speed to obtain a preliminary filtered signal; performing nonlinear residual echo elimination on the preliminary filtered signal and performing multi-reflection path generation to obtain a nonlinear elimination signal; Performing multimodal residual suppression on the nonlinear elimination signal, performing spectrum processing through sub-band division and joint decision, and obtaining an echo suppression signal; Acquire real-time environmental sound field change data, acoustically enhance the echo suppression signal, and obtain target echo-reduced audio output; The performing multi-modal residual suppression on the nonlinear elimination signal, performing spectrum processing by sub-band division and joint decision, and obtaining an echo suppression signal, comprises: performing sub-band decomposition on the nonlinear elimination signal to obtain a non-uniform sub-band representation; Performing multimodal feature extraction and integration on the non-uniform sub-band representation to obtain a multimodal feature set; Calculating cross entropy and mutual information according to the multimodal feature set to obtain a subband grouping structure; Adaptively grouping the non-uniform subband representations according to the subband grouping structure to obtain multiple subband sets; Performing decision rule evaluation on the multiple groups of subband sets respectively to obtain multiple suppression decision values; Performing joint decision processing on the multiple suppression decision values to obtain a target suppression strategy; Performing differential processing on the multiple groups of subband sets according to the target suppression strategy to obtain a modal cooperative suppression signal; performing variational decomposition phase compensation on the modal cooperative suppression signal to obtain a phase coordination signal; Asymmetric singular value decomposition and reconstruction are performed on the phase coordination signal to obtain the echo suppression signal.
2. The intelligent echo reduction audio processing method according to claim 1, characterized in that: The step of acquiring the sound field signal and the loudspeaker reference signal, performing signal framing and short-time Fourier transform, and obtaining time-frequency spectrum audio data includes: Performing synchronous digital sampling on the sound field signal and the loudspeaker reference signal according to a preset sampling rate and bit depth to obtain a digital domain original signal; Divide the original digital domain signal into time frames according to preset frame length and frame shift parameters to obtain a frame sequence; Optimizing spectrum leakage of the frame sequence according to a preset Hanning window function to obtain a windowed signal frame; Performing a fast Fourier transform on the windowed signal frame to obtain complex frequency domain coefficients; Calculating the amplitude spectrum and the phase spectrum according to the complex frequency domain coefficients, and performing frequency band energy analysis to obtain a frequency band energy distribution matrix; Performing Mel frequency scale mapping on the frequency band energy distribution matrix to obtain Mel filter bank coefficients; The Mel filter bank coefficients and the complex frequency domain coefficients are used to construct three-dimensional features to obtain the time-frequency spectrum audio data.
3. The intelligent echo reduction audio processing method according to claim 1, characterized in that: The performing of linear echo cancellation processing on the time-frequency spectrum audio data and dynamically adjusting the convergence speed to obtain a preliminary filtered signal includes: Performing multi-sub-band decomposition on the time-frequency spectrum audio data to obtain multiple frequency sub-band signals; Performing echo path analysis on the multiple frequency sub-band signals to obtain a linear echo path signal; Performing instantaneous imbalance measurement calculation on the sound field signal and the loudspeaker reference signal to obtain a convergence state parameter; Adjust the step length parameter according to the convergence state parameter to obtain an optimized step length value; Performing upper and lower threshold setting mapping on the optimized step value to obtain a stable step parameter; Perform frequency domain filtering on the time-frequency spectrum audio data according to the stable step size parameter to obtain a linear echo estimation signal; Signal optimization is performed on the time-frequency spectrum audio data according to the linear echo estimation signal, and residual correlation compensation is performed to obtain the preliminary filtered signal.
4. The intelligent echo reduction audio processing method according to claim 1, characterized in that: The performing nonlinear residual echo elimination on the preliminary filtered signal and generating multiple reflection paths to obtain a nonlinear elimination signal includes: Performing nonlinear residual basis construction on the preliminary filtered signal to obtain a coupling characteristic matrix; Constructing a Volterra kernel expansion space according to the coupling characteristic matrix, and updating the preliminary filtering signal in real time to obtain optimized filter parameters; Perform multi-reflection path analysis based on the optimized filter parameters to obtain a set of adversarial reflection paths; Superimposing the countermeasure reflection path set and the preliminary filtered signal in time and frequency domains to obtain a multipath suppression signal; Performing sub-band energy distribution analysis on the multipath suppression signal, and generating a dynamic time-frequency masking matrix according to a preset threshold value; The multipath suppression signal is subjected to a point-by-point multiplication operation according to the dynamic time-frequency masking matrix to obtain the nonlinear elimination signal.
5. The intelligent echo reduction audio processing method according to claim 4, characterized in that: The step of constructing a Volterra kernel expansion space according to the coupling characteristic matrix and updating the preliminary filtering signal in real time to obtain optimized filter parameters includes: Performing singular value decomposition on the coupling feature matrix to obtain a reduced-dimensional feature space; Constructing separation components for the dimension-reduced feature space to obtain a separation vector group; Constructing the second-order expansion term of the Volterra kernel function according to the separation vector group, and performing weighted summation of the expansion term to obtain a second-order kernel expansion matrix; Constructing a third-order cross term based on the separation vector group and the second-order kernel expansion matrix, and performing a tensor product operation to obtain a third-order kernel expansion tensor; Performing an inner product operation on the second-order kernel expansion matrix and the third-order kernel expansion tensor to construct a spatial mapping function; Performing nonlinear projection on the preliminary filtered signal according to the spatial mapping function to obtain a projection coefficient vector; The projection coefficient vector is optimized by least squares iteration to obtain the optimized filter parameters.
6. The intelligent echo reduction audio processing method according to claim 1, characterized in that: The step of acquiring real-time environmental sound field change data, acoustically enhancing the echo suppression signal, and obtaining a target echo-reduced audio output includes: Acquire environmental sensor data, and perform spatial spectrum decomposition and sub-band analysis on the environmental sensor data according to the sound field signal to obtain spectrum characteristics of the environmental sound field; Calculating acoustic parameter changes on the ambient sound field spectrum characteristics and the ambient sensor data to obtain the real-time ambient sound field change data; Performing multi-stage cascade filtering on the echo suppression signal according to the real-time environmental sound field change data to obtain a preliminary enhanced signal; Performing phase correction and spectrum smoothing processing on the preliminary enhanced signal, adjusting the gain factor according to the real-time environmental sound field change data, and obtaining a frequency domain enhanced signal; Selectively suppressing residual noise on the frequency domain enhanced signal and the real-time ambient sound field change data to obtain a noise suppressed signal; Performing auditory frequency band energy adjustment on the noise suppression signal to obtain a time domain enhanced signal; The time domain enhanced signal is subjected to harmonic reconstruction and formant enhancement processing to obtain a target echo-reduced audio output.
7. The intelligent echo reduction audio processing method according to claim 6, characterized in that: The selectively suppressing residual noise of the frequency domain enhanced signal and the real-time environmental sound field change data to obtain a noise suppressed signal includes: Performing frequency band region division on the frequency domain enhanced signal to obtain a frequency band divided signal; Performing cross-correlation analysis on the frequency band division signal and the real-time environmental sound field change data to obtain residual noise distribution characteristics; Performing a band-selective masking analysis on the residual noise distribution characteristics to obtain a band noise marker signal; Performing Bayesian probability estimation on the frequency band noise marker signal to obtain a noise probability distribution diagram; Performing nonlinear suppression analysis on the frequency domain enhanced signal according to the noise probability distribution graph to obtain a preliminary noise suppression signal; Performing spectral subtraction balance on the preliminary noise suppression signal and the real-time environmental sound field change data to obtain a balanced suppression signal; Performing artifact threshold removal on the balance suppression signal to obtain an artifact removal coefficient; The artifact removal coefficients are subjected to wavelet reconstruction and phase preservation to obtain a noise suppression signal.
8. An intelligent echo reduction audio processing system, characterized in that: The intelligent echo reduction audio processing method applied to any one of claims 1 to 7 above comprises: An acquisition module, which is used to acquire a sound field signal and a loudspeaker reference signal, and perform signal framing and short-time Fourier transform to obtain time-frequency spectrum audio data; An analysis module, the analysis module is used to perform linear echo cancellation processing on the time-frequency spectrum audio data and dynamically adjust the convergence speed to obtain a preliminary filtered signal; An association module, the association module is used to perform nonlinear residual echo elimination on the preliminary filtered signal and generate multiple reflection paths to obtain a nonlinear elimination signal; A processing module, the processing module is used to perform multi-modal residual suppression on the nonlinear elimination signal, perform spectrum processing through sub-band division and joint decision, and obtain an echo suppression signal; A control module, the control module is used to obtain real-time environmental sound field change data, acoustically enhance the echo suppression signal, and obtain a target echo-reduced audio output; The performing multi-modal residual suppression on the nonlinear elimination signal, performing spectrum processing by sub-band division and joint decision, and obtaining an echo suppression signal, comprises: performing sub-band decomposition on the nonlinear elimination signal to obtain a non-uniform sub-band representation; Performing multimodal feature extraction and integration on the non-uniform sub-band representation to obtain a multimodal feature set; Calculating cross entropy and mutual information according to the multimodal feature set to obtain a subband grouping structure; Adaptively grouping the non-uniform subband representations according to the subband grouping structure to obtain multiple subband sets; Performing decision rule evaluation on the multiple groups of subband sets respectively to obtain multiple suppression decision values; Performing joint decision processing on the multiple suppression decision values to obtain a target suppression strategy; Performing differential processing on the multiple groups of subband sets according to the target suppression strategy to obtain a modal cooperative suppression signal; performing variational decomposition phase compensation on the modal cooperative suppression signal to obtain a phase coordination signal; Asymmetric singular value decomposition and reconstruction are performed on the phase coordination signal to obtain the echo suppression signal.
Citation Information
Patent Citations
Nonlinear narrowband active noise control method based on Volterra filter
CN104575512A
Echo cancellation processing method and processing system
CN110838300A
Intelligent sound effect optimization system based on dynamic frequency gain adjustment
CN119541430A