Voice processing method and device with adjustable noise reduction and reverberation removal degree

By introducing a post-processing mechanism with adjustable parameters after speech denoising and dereverberation, and using convex functions to control the amplitude spectrum of speech and noise signals, the problem of the coupling and unadjustable nature of the denoising and dereverberation processes in existing technologies is solved. This enables flexible adjustment of denoising intensity and reverberation suppression level, and improves the adaptability and configurability of the speech processing system.

CN121583276APending Publication Date: 2026-02-27YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511673764.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

In existing technologies, the processes of speech noise reduction and dereverberation are highly coupled, making it impossible to flexibly adjust the noise reduction intensity and reverberation suppression level according to actual needs, resulting in a lack of flexibility in complex and ever-changing application scenarios.

Method used

By introducing a post-processing mechanism based on control parameters after noise reduction and dereverberation, the amplitude of the speech signal and the noise signal are controlled separately. The amplitude spectrum of the speech signal and the noise signal is controlled by a convex function, so as to achieve adjustable control of the noise reduction intensity and the degree of reverberation suppression.

Benefits of technology

It enables flexible adjustment of noise reduction intensity and reverberation suppression without relying on retraining or structural modification of the front-end AI model, improving the configurability and adaptability of the speech processing system and meeting application scenarios with different noise environments and speech quality requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121583276A_ABST
    Figure CN121583276A_ABST
Patent Text Reader

Abstract

The invention discloses a voice processing method and device with adjustable noise reduction and reverberation removal degree. The method comprises the following steps: obtaining an original voice signal; noise reduction and dereverberation processing are carried out on the original voice signal to obtain a signal processing result, and the signal processing result comprises the first voice signal; performing noise estimation based on the original voice signal and the signal processing result to obtain a first noise signal; post-processing the first voice signal and the first noise signal according to preset regulation and control parameters, and outputting a target voice signal; wherein the post-processing comprises the following steps: respectively carrying out amplitude regulation and control on the first voice signal and the first noise signal according to a preset first regulation and control parameter to obtain a corresponding second voice signal and a corresponding second noise signal. According to the method, the noise suppression and reverberation removal degree can be flexibly adjusted, the requirements in different scenes are met, and an AI voice module does not need to be retrained and deployed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of speech signals, in particular to a speech processing method and device with adjustable noise reduction and dereverberation degree. BACKGROUND

[0002] Traditional artificial intelligence-based speech noise reduction technology usually adopts an end-to-end model, which directly suppresses noise and reverberation in the frequency spectrum after collecting the original signal through a speech transmitting and receiving device. Reverberation refers to the phenomenon of delayed sound waves superimposed after sound is reflected multiple times in space, resulting in the spread of speech energy in time and the decline of clarity. Such methods rely on a single AI model to complete the noise reduction and dereverberation tasks simultaneously, and the processing process is highly coupled. The output speech signal is a fixed noise reduction effect preset by the model. Although it can achieve basic optimization in general scenarios, the model parameters and network structure are fixed, and the noise reduction strength and dereverberation degree cannot be dynamically adjusted according to actual needs, resulting in a lack of flexibility in complex and variable application scenarios. Therefore, how to dynamically adjust the noise reduction strength and dereverberation degree more flexibly according to actual needs in complex and variable application scenarios is a technical problem that needs to be solved. SUMMARY

[0003] The main purpose of the present application is to overcome the problem of adjustable noise suppression and dereverberation degree in complex and variable application scenarios in the prior art, and to provide a speech processing method and device with adjustable noise reduction and dereverberation degree, which can flexibly adjust the noise suppression and dereverberation degree to meet the needs of different scenarios.

[0004] In one aspect, the present application provides a speech processing method with adjustable noise reduction and dereverberation degree, comprising

[0005] obtaining an original speech signal;

[0006] performing noise reduction and dereverberation processing on the original speech signal to obtain a signal processing result, wherein the signal processing result includes a first speech signal;

[0007] performing noise estimation based on the original speech signal and the signal processing result to obtain a first noise signal;

[0008] performing post-processing on the first speech signal and the first noise signal according to a preset control parameter to output a target speech signal; wherein the post-processing includes:

[0009] performing amplitude control on the first speech signal and the first noise signal according to a preset first control parameter to obtain corresponding second speech signal and second noise signal.

[0010] Optionally, the amplitude regulation is a convex function regulation, and the preset first regulation parameter is a convex function coefficient.

[0011] Optionally, the convex function adopted in the amplitude regulation is a power function y n = x n α ,

[0012] x n is an amplitude spectrum of the input first speech signal, y n is an amplitude spectrum of the output second speech signal, and a is the first regulation parameter and has a value range of (0, +∞). n is an amplitude spectrum of the input first noise signal, y n is an amplitude spectrum of the output second noise signal, and a is the first regulation parameter and has a value range of (0, +∞).

[0013] Optionally, the post-processing of the first speech signal and the first noise signal according to the preset regulation parameter further includes:

[0014] reverberation regulation of the second speech signal and the second noise signal according to a preset second regulation parameter, to output a target speech signal.

[0015] Optionally, the reverberation regulation of the second speech signal and the second noise signal according to the preset second regulation parameter, to output the target speech signal, includes:

[0016] reverberation regulation of an amplitude spectrum of the second speech signal and an amplitude spectrum of the second noise signal according to the preset second regulation parameter, to obtain an amplitude spectrum of the target speech signal; and combination of the amplitude spectrum of the target speech signal and a phase spectrum of the second speech signal to generate a complex spectrum, and inverse Fourier transform to obtain the target speech signal.

[0017] Optionally, the reverberation regulation of the amplitude spectrum of the second speech signal and the amplitude spectrum of the second noise signal according to the preset second regulation parameter, to obtain the amplitude spectrum of the target speech signal, includes:

[0018] ;

[0019] wherein Ym is the amplitude spectrum of the target speech signal, Y2 is the amplitude spectrum of the second noise signal, Y1 is the amplitude spectrum of the second speech signal, β is the second regulation parameter, and 0≤β≤1.

[0020] The combination of the amplitude spectrum of the target speech signal and the phase spectrum of the second speech signal to generate the complex spectrum, and the inverse Fourier transform to obtain the target speech signal, includes:

[0021] ;

[0022] wherein, output is a complex spectrum of the target speech signal, is a phase spectrum of the second speech signal, represents a complex unit vector.

[0023] Optionally, the signal processing result further comprises a noise signal, the first noise signal is obtained by performing noise estimation on the original speech signal and the signal processing result, comprising:

[0024] The first noise signal is obtained by performing noise estimation on the original speech signal and the noise signal.

[0025] Optionally, the first noise signal is obtained by performing noise estimation on the original speech signal and the signal processing result, comprising:

[0026] The noise signal is obtained by performing noise estimation on the original speech signal and the first speech signal, and the first noise signal is obtained by using spectral subtraction based on the original speech signal and the noise signal.

[0027] In another aspect, the present application also provides a speech processing device with adjustable dereverberation and noise reduction, comprising:

[0028] A speech signal acquisition module acquires an original speech signal;

[0029] A noise reduction and dereverberation processing module performs noise reduction and dereverberation processing on the original speech signal to obtain a signal processing result, and performs noise estimation on the original speech signal and the signal processing result to obtain a first noise signal, wherein the signal processing result at least comprises a first speech signal;

[0030] A post-processing module performs post-processing on the first speech signal and the first noise signal according to a preset control parameter, and outputs a target speech signal, wherein the post-processing comprises:

[0031] According to a preset first control parameter, the first speech signal and the first noise signal are respectively amplitude-controlled to obtain corresponding second speech signal and second noise signal.

[0032] Optionally, the noise reduction and dereverberation processing module uses an AI speech processing model.

[0033] From the above description of the present application, compared with the prior art, the present application has the following beneficial effects: the present application provides a noise reduction and dereverberation degree adjustable speech processing method and device, by introducing a post-processing mechanism based on a control parameter after noise reduction and dereverberation processing, amplitude control is performed on the first speech signal and the first noise signal respectively, and adjustable control of the noise reduction intensity is realized; the mechanism does not depend on retraining or structural modification of the front-end AI model, and only by adjusting the first control parameter can the nonlinear mapping relationship of the speech and noise amplitude spectrum be changed, effectively improving the configurability and adaptability of the speech processing system, and being able to meet the application scenarios of different noise environments and speech quality requirements. BRIEF DESCRIPTION OF DRAWINGS

[0034] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings described in the following are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0035] Figure 1 A noise reduction and dereverberation degree adjustable speech processing method main flow chart is provided for the embodiments of the present application.

[0036] Figure 2 A noise reduction and dereverberation degree adjustable speech processing device composition diagram is provided for the embodiments of the present application.

[0037] Figure 3 A noise reduction degree adjustment schematic diagram is provided for the embodiments of the present application. DETAILED DESCRIPTION

[0038] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely in the following combined with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0039] The present application is further described through specific embodiments.

[0040] Traditional artificial intelligence speech noise reduction technology packages noise reduction and dereverberation into an end-to-end network, and the model outputs a "unique cleanness" after being trained: noise and reverberation are simultaneously suppressed to a fixed level, and users cannot individually reduce noise while retaining spatial sense, nor can they instantly switch between different modes in different scenarios such as conferences, live broadcasts, recording studios, etc. If you want to change the effect, you have to re-collect data, fine-tune and redeploy the entire model, which is long cycle, high cost, and edge device computing power and storage resources cannot withstand frequent iterations.

[0041] In this context, the inventors of the present application have found that the core defect of the prior art is that the strength of noise suppression and reverberation removal is bound as a joint processing result that cannot be separated, and users cannot independently adjust it according to environmental noise level, reverberation duration or speech clarity requirements.

[0042] To solve the problem of how to dynamically adjust the noise reduction strength and reverberation suppression degree according to actual needs in complex and variable application scenarios in the prior art, the embodiments of the present application provide a speech processing method with adjustable noise reduction and dereverberation degree, which comprises: obtaining an original speech signal; performing noise reduction and dereverberation processing on the original speech signal to obtain a signal processing result, wherein the signal processing result comprises a first speech signal; performing noise estimation based on the original speech signal and the signal processing result to obtain a first noise signal; performing post-processing on the first speech signal and the first noise signal according to a preset control parameter to output a target speech signal; wherein the post-processing comprises: performing amplitude control on the first speech signal and the first noise signal according to a preset first control parameter to obtain corresponding second speech signal and second noise signal.

[0043] Embodiment one

[0044] Referring to Figure 1 , the embodiments of the present application propose a speech processing method with adjustable noise reduction and dereverberation degree, comprising the following steps:

[0045] S1 obtains an original speech signal, which is a speech signal with noise.

[0046] S2 performs noise reduction and dereverberation processing on the original speech signal and obtains a signal processing result, which at least includes a first speech signal.

[0047] In this step, the original speech signal is denoised and dereverberated using an AI speech processing model to obtain a signal processing result. Different AI speech processing models can yield different signal processing results. There are mainly two kinds: the first AI speech processing model, whose signal processing result contains a first speech signal; the second AI speech processing model, whose signal processing result can contain a first speech signal and a noise signal. The first speech signal is the speech signal after denoising and dereverberation, and the noise signal is an estimation result of the noise of the original speech signal.

[0048] For the first AI speech processing model, a speech enhancement model or a dereverberation model can be used. Examples of the speech enhancement model include the following: DCCRN (Deep Complex Convolutional Recurrent Network), which is an end-to-end model based on complex spectrum, jointly optimizes denoising and dereverberation, and outputs a pure speech signal (suppresses non-speech components); DPCRN (Dual-Path Convolutional Recurrent Network) can also be used to separate speech and interference through time-frequency domain dual-path separation. SEGAN (Speech Enhancement Generative Adversarial Network) is a time-domain model based on GAN, which directly generates denoised speech waveform without relying on spectral features. Examples of the dereverberation model include the following: WavLM (Waveform-based Dereverberation Model) uses self-supervised pre-training to extract direct sound components from reverberation signals; DereverbNet is an LSTM-based end-to-end network that models room impulse responses (RIR) to suppress reverberation tails.

[0049] For the second AI speech processing model, a speech separation model, a target sound detection model, or a multi-channel processing model, etc. can be adopted. The speech separation model takes the following examples: Conv-TasNet (Convolutional Time-domain Audio Separation Network), which is a time-domain speech separation model, separates multiple sound sources in a mixed signal through an encoder-decoder structure, and outputs a specified channel that retains target speech and certain noise; DPRNN (Dual-Path RNN), which is a speech separation model that processes long-term dependencies in stages. The target sound detection model takes the following examples: TasNet (Time-domain Audio Separation Network) separates target speech through mask estimation, and can adjust the threshold to retain part of the background noise; VoiceFilter, based on speaker feature embedding, extracts target human voice and retains non-speech environmental sound. The multi-channel processing model takes the following examples: BeamformIt (deep learning enhanced beamforming), which combines traditional beamforming and DNN to suppress non-target direction noise, but retains sound sources in the specified direction. The AI speech processing model of the present application is not limited to this, and can also include other similar functional AI speech processing models.

[0050] S3 performs noise estimation based on the original speech signal and the signal processing result to obtain a first noise signal.

[0051] In this step, different noise estimation methods are used for the processing results output by the above two different AI speech processing models: for the signal processing result including the first speech signal, the original speech signal and the first speech signal are first estimated to obtain a noise signal, and then the original speech signal and the noise signal are used to obtain the first noise signal by using the spectral subtraction method. Specifically, for the i-th frame, noise_est i =noisy i -(denoise+derev) i , thus obtaining the noise i of the i-th frame. i-1 =α×noise_est i +(1-α)×noise_est i , noise_est i-1 represents the noise estimated according to the AI processing result of the i-th frame, noise_est i represents the noise estimated according to the AI processing result of the i-1-th frame, noise i represents the noise estimated according to noise_est i-1The i-th frame of actual noise obtained after smoothing. In actual applications, a can be set according to requirements, for example, a under pure noise is preferably 0.5, and a under human voice is preferably 0.95.

[0052] In some embodiments, the following method can also be used to implement:

[0053] Noise estimation based on frequency domain statistical modeling: first, a noise estimation model is constructed using the spectral difference between the original speech signal and the first speech signal. By analyzing the residual energy distribution of the two in the frequency domain, combined with the statistical characteristics of the non-speech segment (such as the mean variance model), the noise component of each frequency point is dynamically calculated; then an improved spectral subtraction method (such as multi-window spectral estimation or adaptive gain control) is used to weight and smooth the noise component in the original signal, generating a high-fidelity first noise signal.

[0054] Deep learning assisted iterative optimization method: the original signal and the first speech signal are compared by a pre-trained convolutional neural network (CNN), and the initial noise features are extracted by using the residual learning mechanism. The features are input into a mixture density network (MDN) to predict the time-frequency mask of the noise signal, and then the spectral subtraction method is used for phase correction. The energy matching degree of the noise mask and the original signal is optimized through iteration, and finally the first noise signal that meets the acoustic characteristics is output.

[0055] Further, for the signal processing results including the noise signal and the first speech signal, the noise estimation of the original speech signal and the noise signal can be used to obtain the first noise signal, which can be implemented by using the spectral subtraction method.

[0056] In some embodiments, the following method can also be used to implement:

[0057] Frequency domain residual energy analysis method: the residual energy spectrum is calculated and the noise features are extracted by comparing the frequency energy difference between the original speech signal and the first speech signal after noise reduction. The time-frequency components of the signal are separated by using the short-time Fourier transform (STFT), the residual spectrum is amplitude corrected and phase aligned by using the spectral subtraction method, and finally the high-precision first noise signal is reconstructed by inverse transform.

[0058] Adaptive filtering noise separation method: the original speech signal is used as a reference input, and a filter is constructed by using the least mean square (LMS) or recursive least square (RLS) adaptive algorithm with the first speech signal as the expected signal. The noise component is dynamically separated by iteratively updating the filter coefficients, and the time-frequency continuity of the noise spectrum is optimized by combining the spectral subtraction method, to realize the extraction of the first noise signal with low distortion.

[0059] S4 performs post-processing on the first speech signal and the first noise signal according to the preset control parameters, and outputs a target speech signal.

[0060] The post-processing includes amplitude regulation of the first speech signal and the first noise signal according to the first regulation parameter to obtain the second speech signal and the second noise signal. The step realizes the regulation of the noise suppression degree and the dereverberation degree by adjusting the amplitude spectrum of the output.

[0061] The amplitude spectrum of the second speech signal is obtained by regulating the amplitude spectrum of the first speech signal by using a convex function, and the coefficient of the convex function is used as the first regulation parameter to regulate the noise suppression degree of the second speech signal.

[0062] The amplitude spectrum of the first speech signal is nonlinearly mapped by introducing a convex function to construct a weighted regulation mechanism of the amplitude spectrum. Specifically, the form of the convex function is defined based on the noise suppression requirement (such as an exponential function or a piecewise linear function), the amplitude spectrum of the first speech signal is taken as the input, and the compression or enhancement degree of the amplitude spectrum of different frequency bands is dynamically controlled by adjusting the first regulation parameter (such as a curvature factor or a weight coefficient) in the convex function. The greater the parameter is, the more significant the suppression of the high-noise frequency band by the convex function is, but at the same time, the trade-off relationship between speech distortion and noise residue needs to be balanced, and finally the optimized amplitude spectrum of the second speech signal is generated to realize flexible and controllable noise suppression intensity.

[0063] Specifically, when a power function is used as the convex function, y n =x n α , x n is the amplitude spectrum of the input first speech signal, y n is the amplitude spectrum of the output second speech signal, and α is the first regulation parameter and its value range is (0, +∞) for controlling the intensity of nonlinear mapping:

[0064] When α = 1, the power function degenerates into a linear mapping y n =x n , at this time, the mask contrast of the power spectrum remains the original proportion, and the output result is consistent with the unregulated module, in the balanced state of noise suppression and speech preservation;

[0065] When α > 1, the power function has an upper convexity, stronger nonlinear compression is applied to the frequency points with higher energy in the amplitude spectrum (usually the noise dominant area), resulting in a significant increase in the mask contrast, and as α increases, the suppression degree of the high-noise frequency band is further increased.

[0066] When 0 < α < 1, the power function becomes a lower convexity, the low-amplitude frequency points (such as weak speech components) are enhanced, and the suppression intensity of the high-noise frequency band is reduced, thereby reducing the mask contrast. A smaller α value can improve the speech smoothness.

[0067] By adjusting the value of a, the noise suppression strength can be dynamically adjusted in the frequency domain: increasing a gives priority to suppressing high-noise frequency bands, and decreasing a focuses on preserving the integrity of the voice, which is suitable for adaptive noise reduction requirements in different signal-to-noise ratio scenarios. For example: for a regular environment, a = 1 can be used to meet most usage requirements. For scenarios where noise is noisy and human voice is expected to be cleaner, a > 1 can be used for regulation. For scenarios where noise is weak and higher voice restoration is expected, 0 < a < 1 can be used for regulation. See Figure 3 , the abscissa is the AI directly output noise suppression coefficient, and the ordinate is the adjusted noise suppression coefficient, which can non-linearly adjust the noise reduction degree.

[0068] In this embodiment, the amplitude spectrum of the first noise signal is adjusted by a convex function to obtain the amplitude spectrum of the second noise signal, and the coefficient of the convex function is used as the first control parameter to control the noise suppression degree of the second noise signal. The original noise amplitude spectrum is used as the input, a preset convex function form (such as a power function or an exponential function) is used, and the function coefficient (i.e. the first control parameter) is adjusted to control the reconstruction strength of the noise spectrum. When the control parameter is large, the non-linear characteristic of the convex function is more pronounced in the high-energy frequency band, forcing the amplitude spectrum of the noise dominant region to be compressed, thereby reducing the noise energy ratio; when the control parameter is small, the function curve tends to be linear, retaining more noise details but reducing the suppression ability.

[0069] Specifically, when a power function is used as the convex function, y n =x n α , x n is the amplitude spectrum of the input first noise signal, y n is the amplitude spectrum of the output second noise signal, and a is the first control parameter with a value range of (0, +∞). The specific form is as follows:

[0070] When a = 1, the power function degenerates into an identity mapping, the output amplitude spectrum is identical to the input, and it is in an uncontrolled state;

[0071] When a > 1, the convexity of the power function causes the high-amplitude frequency points (usually corresponding to the noise dominant region) to be significantly compressed, reducing the noise energy ratio, and the suppression degree increases with a;

[0072] When 0 < a < 1, the convexity of the power function has a more obvious enhancement effect on low-amplitude frequency points, and the noise suppression strength is weakened, but it can alleviate the loss of voice details caused by excessive suppression.

[0073] By adjusting a, adaptive balance between noise suppression depth and signal fidelity can be achieved, providing flexible frequency domain regulation strategies for different noise scenarios.

[0074] The step further comprises outputting a target speech signal after reverberation regulation of the second speech signal and the second noise signal according to a second regulation parameter.

[0075] Specifically, the amplitude spectrum of the second speech signal and the amplitude spectrum of the second noise signal are regulated according to the second regulation parameter to obtain an amplitude spectrum of the target speech signal:

[0076] ;

[0077] Wherein Ym is the amplitude spectrum of the target speech signal, Y2 is the amplitude spectrum of the second noise signal, Y1 is the amplitude spectrum of the second speech signal, and β is the second regulation parameter, 0≤β≤1.

[0078] Based on the second regulation parameter β, the amplitude spectrum of the target speech signal is constructed by linearly superimposing the amplitude spectrum Y1 of the second speech signal and the amplitude spectrum Y2 of the second noise signal, so as to realize continuous adjustment of the residual reverberation.

[0079] When β=0, Ym=Y2, at this time, only noise suppression is completed but the original reverberation characteristics of the speech are retained, which is suitable for scenes that need to maintain the naturalness of the sound field

[0080] When β=1, the amplitude spectrum of the target speech signal is equivalent to Y1, indicating that noise and reverberation are simultaneously eliminated, which is suitable for high-definition speech requirements.

[0081] When 0<β<1, a controllable balance between reverberation suppression intensity and speech naturalness is established through linear interpolation, and while removing noise, the reverberation of the speech is also removed to some extent, and the degree of removal is related to the size of β. The increase of the β value corresponds to the increase of the reverberation suppression weight, and the residual reverberation energy and are positively correlated.

[0082] In the embodiment, the second regulation parameter β is preferably 0.8, so that when the reverberation suppression intensity reaches 80%, the interference of reverberation on speech intelligibility can be effectively eliminated, and the distortion of acoustic characteristics caused by complete elimination of reverberation can be avoided, so that an optimal compromise between noise reduction performance and acoustic scene adaptability is achieved.

[0083] Further, after obtaining the amplitude spectrum of the target speech signal, the phase spectrum of the second speech signal is extracted, the amplitude spectrum and the phase spectrum of the target speech signal are combined to generate a complex spectrum, and the target speech signal is obtained through inverse Fourier transform:

[0084] ;

[0085] Wherein, is the complex spectrum of the target speech signal, is the phase spectrum of the second speech signal, represents a complex unit vector.

[0086] In the present application, the output of the AI speech processing module is regulated in the form of a convex function, combined with reverberation regulation, to achieve adjustment of the effect and degree of the AI speech processing module; by outputting pure noise suppression results or noise suppression & reverberation removal results, the degree of noise reduction and reverberation removal is arbitrarily adjusted to meet the needs of different application scenarios.

[0087] The present application outputs a target speech signal by regulating the reverberation of the second speech signal and the second noise signal after amplitude regulation through the second regulation parameter, realizing the adjustment of the reverberation residual degree in the processing result. This mechanism flexibly controls the dereverberation intensity without relying on the front-end model update, and adapts to application scenarios with different requirements for speech naturalness and environmental feeling.

[0088] The present application constructs an amplitude spectrum weighting regulation mechanism based on a convex function, dynamically adjusts the compression or enhancement amplitude of each frequency band amplitude spectrum by implementing nonlinear mapping on the first speech signal or first noise signal amplitude spectrum. The power function is preferably used as a convex function, and by adjusting its exponential coefficient, the noise suppression depth and speech signal fidelity can be adaptively balanced to form a frequency domain regulation strategy suitable for different noise environments.

[0089] The present application nonlinearly maps the first noise signal amplitude spectrum based on a convex function to generate a second noise signal amplitude spectrum, and uses the convex function coefficient as the first regulation parameter to dynamically adjust the noise suppression intensity. Specifically, a convex function such as a power function is used, and after the original noise amplitude spectrum is input, the noise spectrum reconstruction intensity is controlled by adjusting its coefficient (i.e. the first regulation parameter): when the parameter is large, the nonlinear characteristics of the high-energy frequency band are enhanced, significantly compressing the amplitude spectrum of the noise dominant region; when the parameter is small, the function tends to be linear, the suppression ability is weakened but more noise details are retained.

[0090] The present application linearly superimposes the second speech signal amplitude spectrum Y1 and the second noise signal amplitude spectrum Y2 according to the second regulation parameter to construct a target speech signal amplitude spectrum, and realizes continuous adjustment of the reverberation residual amount.

[0091] Embodiment two

[0092] Based on this, the present application also proposes a speech processing device with adjustable noise reduction and dereverberation degree, which uses the above-mentioned speech processing method with adjustable noise reduction and dereverberation degree, as shown in Figure 2 , comprising:

[0093] The speech signal acquisition module 101 acquires the original speech signal.

[0094] The noise reduction and dereverberation processing module 102 performs noise reduction and dereverberation processing on the original speech signal to obtain a signal processing result, and performs noise estimation on the original speech signal and the signal processing result to obtain a first noise signal, wherein the signal processing result at least includes a first speech signal.

[0095] The post-processing module 103 performs post-processing on the first speech signal and the first noise signal according to preset control parameters to output a target speech signal, wherein the post-processing includes: performing amplitude control on the first speech signal and the first noise signal according to preset first control parameters to obtain corresponding second speech signal and second noise signal.

[0096] The speech signal acquisition module 101 is configured to perform steps S1 of the speech processing method with adjustable noise reduction and dereverberation degree in Embodiment I.

[0097] Specifically, the original speech signal is obtained, and the original speech signal is a speech signal with noise.

[0098] The noise reduction and dereverberation processing module 102 is configured to perform steps S2 and S3 of the speech processing method with adjustable noise reduction and dereverberation degree in Embodiment I, and the noise reduction and dereverberation processing module can be implemented by an AI speech processing model.

[0099] Specifically, the noise reduction and dereverberation processing module 102 is configured to perform step S2 to perform noise reduction and dereverberation processing on the original speech signal and obtain a signal processing result, and the signal processing result at least includes a first speech signal, and the noise estimation is performed on the original speech signal and the signal processing result to obtain a first noise signal.

[0100] In this step, the AI speech processing model is used to perform noise reduction and dereverberation processing on the original speech signal to obtain a signal processing result. Different AI speech processing models can obtain different signal processing results. There are mainly two kinds: the first kind of AI speech processing model, whose signal processing result contains a first speech signal; the second kind of AI speech processing model, whose signal processing result can contain a first speech signal and a noise signal. The first speech signal is the speech signal after noise reduction and dereverberation, and the noise signal is the estimation result of the noise of the original speech signal.

[0101] For the first AI speech processing model, a speech enhancement model or a dereverberation model can be used. For example, a DCCRN (Deep Complex Convolutional Recurrent Network) is used, which is an end-to-end model based on complex spectrum, jointly optimizes noise reduction and dereverberation, and outputs a pure speech signal (suppresses non-speech components); a DPCRN (Dual-Path Convolutional Recurrent Network) can also be used, which separates speech and interference through time-frequency domain double-path separation. SEGAN (Speech Enhancement Generative Adversarial Network) is a time-domain model based on GAN, which directly generates a denoising speech waveform without relying on spectral features. For example, a WavLM (Waveform-based Dereverberation Model) is used, which uses self-supervised pre-training to extract direct sound components from reverberation signals; DereverbNet is an LSTM-based end-to-end network that models room impulse responses (RIR) and suppresses reverberation tails.

[0102] For the second AI speech processing model, a speech separation model, a target sound detection model, or a multi-channel processing model can be used. For example, a Conv-TasNet (Convolutional Time-domain Audio Separation Network) is used, which is a time-domain speech separation model that separates multiple sound sources in a mixed signal through an encoder-decoder structure, and outputs a specified target speech and specific noise channel; DPRNN (Dual-Path RNN) is a speech separation model that handles long-term dependencies in stages. For example, a TasNet (Time-domain Audio Separation Network) is used to separate target speech by mask estimation, and the threshold can be adjusted to retain part of the background noise; VoiceFilter is based on speaker feature embedding to extract target human voice and retain non-speech environment sound. For example, BeamformIt (deep learning enhanced beamforming) combines traditional beamforming with DNN to suppress non-target direction noise, but retains sound sources in the specified direction. The AI speech processing model of the present application is not limited to this, and can also include other similar functional AI speech processing models.

[0103] Specifically, the noise reduction and dereverberation processing module 102 is further configured to perform step S3 of the method of embodiment I, and obtain a first noise signal based on noise estimation of the original speech signal and the signal processing result.

[0104] In this step, different noise estimation methods are used for the processing results output by the two different AI voice processing models: for the signal processing result including the first voice signal, the noise signal is estimated from the original voice signal and the first voice signal, and then the original voice signal and the noise signal are processed by the spectral subtraction method to obtain the first noise signal. Specifically, for the i-th frame, noise_esti = noisyi - (denoise + derevi), and thus noisei = a x noise_esti-1 + (1-a) x noise_esti, where noise_esti represents the noise estimated from the AI processing result of the i-th frame, noise_esti-1 represents the noise estimated from the AI processing result of the (i-1)-th frame, and noisei represents the actual noise of the i-th frame obtained by smoothing noise_esti and noise_esti-1. In actual application, a can be set according to requirements, for example, a is preferably 0.5 under pure noise, and a is preferably 0.95 under human voice.

[0105] In some embodiments, the following methods can also be used:

[0106] Noise estimation based on frequency domain statistical modeling: first, a noise estimation model is constructed using the spectral difference between the original voice signal and the first voice signal. By analyzing the residual energy distribution of the two in the frequency domain, combined with the statistical characteristics of non-speech segments (such as mean variance model), the noise component of each frequency point is dynamically calculated; then, an improved spectral subtraction method (such as multi-window spectral estimation or adaptive gain control) is used to smooth and suppress the noise component in the original signal, generating a high-fidelity first noise signal.

[0107] Deep learning assisted iterative optimization method: the original signal and the first voice signal are compared by a pre-trained convolutional neural network (CNN), and the initial noise features are extracted by using the residual learning mechanism. The features are input into a mixture density network (MDN) to predict the time-frequency mask of the noise signal, and then the spectral subtraction method is used for phase correction. Through iterative optimization of the energy matching degree of the noise mask and the original signal, a first noise signal that meets the acoustic characteristics is finally output.

[0108] Further, for the signal processing result including the noise signal and the first voice signal, the noise estimation method can be used to obtain the first noise signal by performing noise estimation on the original voice signal and the noise signal.

[0109] In some embodiments, the following methods can also be used:

[0110] Frequency domain residual energy analysis method: by comparing the frequency energy difference between the original speech signal and the first speech signal after noise reduction, the residual energy spectrum is calculated and the noise feature is extracted. The short-time Fourier transform (STFT) is used to separate the time-frequency components of the signal, and the residual spectrum is amplitude corrected and phase aligned by using the spectral subtraction method. Finally, a high-precision first noise signal is reconstructed through inverse transform.

[0111] Adaptive filter noise separation method: the original speech signal is used as the reference input, and the least mean square (LMS) or recursive least square (RLS) adaptive algorithm is used to construct the filter with the first speech signal as the expected signal. The filter coefficients are updated iteratively to separate the noise component dynamically, and the spectral subtraction method is used to optimize the time-frequency continuity of the noise spectrum to realize the extraction of the first noise signal with low distortion.

[0112] The post-processing module 103 is used to execute step S4 of the speech processing method with adjustable de-reverberation degree of noise reduction in embodiment one.

[0113] Specifically, the post-processing module 103 is used to execute step S4, and the first speech signal and the first noise signal are post-processed according to the preset control parameter to output a target speech signal.

[0114] Among them, the post-processing includes respectively adjusting the amplitude spectrum of the first speech signal and the first noise signal according to the first control parameter to obtain the second speech signal and the second noise signal. This step adjusts the amplitude spectrum of the output to realize the control of the noise suppression degree and the de-reverberation degree.

[0115] The amplitude spectrum of the second speech signal is obtained by adjusting the amplitude spectrum of the first speech signal with a convex function, and the coefficient of the convex function is used as the first control parameter to control the noise suppression degree of the second speech signal.

[0116] The amplitude spectrum of the first speech signal is nonlinearly mapped by introducing a convex function to construct a weighted control mechanism for the amplitude spectrum. Specifically, the form of the convex function is defined based on the noise suppression requirement (such as an exponential function or a piecewise linear function), the amplitude spectrum of the first speech signal is taken as the input, and the compression or enhancement degree of the amplitude spectrum of different frequency bands is dynamically controlled by adjusting the first control parameter (such as the curvature factor or the weight coefficient) in the convex function. The greater the parameter, the more significant the suppression of the convex function on the high noise frequency band, but at the same time, the trade-off relationship between speech distortion and noise residue needs to be balanced, and finally the optimized amplitude spectrum of the second speech signal is generated to realize the flexible control of the noise suppression intensity.

[0117] Specifically, when the power function is used as the convex function, y n =x n α , x n is the amplitude spectrum of the input first speech signal, and y nis the amplitude spectrum of the output second speech signal, and a is a first control parameter and has a value range of (0, +∞) for controlling the intensity of the nonlinear mapping:

[0118] When a = 1, the power function degenerates into a linear mapping y n =x n , at this time, the mask contrast of the power spectrum remains the original proportion, and the output result is consistent with the uncontrolled module, in a balanced state of noise suppression and speech fidelity;

[0119] When a > 1, the power function has an upward convexity, and stronger nonlinear compression is applied to the frequency points with higher energy in the amplitude spectrum (usually the noise dominant area), resulting in a significant increase in the mask contrast, and as a increases, the suppression degree of the high noise frequency band is further increased.

[0120] When 0 < a < 1, the power function becomes a downward convexity, enhances the low-amplitude frequency points (such as weak speech components), and at the same time reduces the suppression intensity of the high noise frequency band, thereby reducing the mask contrast. A smaller a value can improve the speech smoothness.

[0121] By adjusting the value of a, the noise suppression intensity can be dynamically adjusted in the frequency domain: increasing a preferentially suppresses the high noise frequency band, and reducing a focuses on preserving the integrity of the speech, which is suitable for adaptive noise reduction requirements in different signal-to-noise ratio scenarios. For example: for a regular environment, a = 1 can be used to meet most use requirements. For a noisy environment, a > 1 can be used to control the voice to be cleaner. For a weak noise environment, 0 < a < 1 can be used to control the voice to be restored more accurately. Referring to Figure 3 , the horizontal coordinate is the AI directly output noise suppression coefficient, and the vertical coordinate is the adjusted noise suppression coefficient, which can nonlinearly adjust the noise reduction degree.

[0122] In this embodiment, the amplitude spectrum of the second noise signal is obtained by using a convex function to control the amplitude spectrum of the first noise signal, and the coefficient of the convex function is used as a first control parameter to control the noise suppression degree of the second noise signal. The original noise amplitude spectrum is used as input, a preset convex function form (such as a power function or an exponential function) is used, and the noise spectrum reconstruction intensity is controlled by adjusting the function coefficient (i.e. the first control parameter). When the control parameter is large, the nonlinear characteristics of the convex function are more significant in the high energy frequency band, forcing the amplitude spectrum of the noise dominant area to be compressed, thereby reducing the noise energy proportion; when the control parameter is small, the function curve tends to be linear, and more noise details are retained but the suppression ability is weakened.

[0123] Specifically, when a power function is used as the convex function, y n =x n α , x n is the amplitude spectrum of the input first noise signal, and yn For the amplitude spectrum of the output second noise signal, α is a first control parameter and its value range is (0, +∞), and the specific is as follows:

[0124] When α = 1, the power function degenerates into an identity mapping, the output amplitude spectrum is completely consistent with the input, and it is in a non-control state;

[0125] When α > 1, the convexity of the power function makes the high-amplitude frequency points (usually corresponding to the noise dominant region) be significantly compressed, the noise energy proportion is reduced, and the suppression degree is enhanced with the increase of α;

[0126] When 0 < α < 1, the convexity of the power function has more obvious enhancement effect on the low-amplitude frequency points, the noise suppression strength is weakened, but the loss of voice details caused by excessive suppression can be alleviated.

[0127] By adjusting α, adaptive balance between noise suppression depth and signal fidelity can be achieved, and flexible frequency domain control strategy is provided for different noise scenes.

[0128] The step further comprises outputting a target voice signal after reverberation control of the second voice signal and the second noise signal according to a second control parameter.

[0129] Specifically, the amplitude spectrum of the target voice signal is obtained by performing reverberation control on the amplitude spectrum of the second voice signal and the amplitude spectrum of the second noise signal according to a second control parameter:

[0130]

[0131] Wherein, Ym is the amplitude spectrum of the target voice signal, Y2 is the amplitude spectrum of the second noise signal, Y1 is the amplitude spectrum of the second voice signal, and β is the second control parameter, 0 ≤ β ≤ 1.

[0132] Based on the second control parameter β, the target voice signal amplitude spectrum is constructed by linearly superimposing the second voice signal amplitude spectrum Y1 and the second noise signal amplitude spectrum Y2, so as to realize continuous adjustment of the reverberation residual amount.

[0133] When β = 0, Ym = Y2, at this time, only noise suppression is completed but the original reverberation characteristics of the voice are retained, which is suitable for scenes that need to maintain the naturalness of the sound field

[0134] When β = 1, the amplitude spectrum of the target voice signal is equivalent to Y1, which indicates that noise and reverberation are eliminated at the same time, which is suitable for high-definition voice demand;

[0135] When 0 < β < 1, a controllable balance between reverberation suppression intensity and voice naturalness is established through linear interpolation, the voice reverberation is also removed part of the time when the noise is removed, and the degree of removal is related to the size of β. The increase of β value corresponds to the increase of reverberation suppression weight, and the residual reverberation energy is​ positively correlated.

[0136] In the embodiment, the second control parameter β is preferably 0.8, so that when the reverb suppression intensity reaches 80%, the disturbance of the reverb to the speech intelligibility can be effectively eliminated, and the acoustic characteristic distortion caused by completely eliminating the reverb can be avoided, and an engineering optimal compromise between the noise reduction performance and the acoustic scene adaptability is achieved.

[0137] Further, after obtaining the target speech signal amplitude spectrum, the phase spectrum of the second speech signal is extracted, the amplitude spectrum and the phase spectrum of the target speech signal are combined to generate a complex spectrum, and the target speech signal is obtained through inverse Fourier transform:

[0138]

[0139] wherein, the complex spectrum of the target speech signal, the phase spectrum of the second speech signal, represents a complex unit vector.

[0140] The speech processing device with adjustable noise reduction and dereverberation degree provided by the embodiment has the following beneficial effects compared with the prior art:

[0141] The AI speech processing module output is regulated in the form of a convex function, combined with reverb regulation, to realize adjustment of the AI speech processing module effect and degree; by outputting pure noise suppression results or noise suppression & reverb removal results, the noise reduction and dereverberation degree is realized, and the needs of different application scenarios are met.

[0142] The second control parameter is used to regulate the reverb of the second speech signal and the second noise signal after amplitude regulation, and the target speech signal is output, so that the reverb residual degree in the processing result is adjusted, and the reverb removal intensity is flexibly controlled without relying on front-end model updating, and application scenarios with different requirements for speech naturalness and environmental feeling are adapted.

[0143] The amplitude spectrum weighting regulation mechanism is constructed based on a convex function, the nonlinear mapping of the first speech signal or the first noise signal amplitude spectrum is implemented, and the compression or enhancement amplitude of each frequency band amplitude spectrum is dynamically adjusted. The power function is preferably used as the convex function, and by adjusting the exponential coefficient, the noise suppression depth and the speech signal fidelity can be adaptively balanced, and a frequency domain regulation strategy suitable for different noise environments is formed.

[0144] ​The application is based on nonlinear mapping of the first noise signal amplitude spectrum by a convex function, generates a second noise signal amplitude spectrum, and uses the convex function coefficient as the first control parameter to dynamically adjust the noise suppression strength. Specifically, a convex function such as a power function is used to input the original noise amplitude spectrum, and the noise spectrum reconstruction strength is controlled by adjusting the coefficient (i.e., the first control parameter) of the function: when the parameter is large, the nonlinear characteristics of the high-energy frequency band are enhanced, and the amplitude spectrum of the noise dominant region is significantly compressed; when the parameter is small, the function tends to be linear, and the suppression ability is weakened but more noise details are retained.

[0145] The application linearly superimposes the second speech signal amplitude spectrum Y1 and the second noise signal amplitude spectrum Y2 according to the second control parameter to construct a target speech signal amplitude spectrum, and realizes continuous adjustment of the amount of residual reverberation.

[0146] The method and device for adjusting the degree of noise reduction and dereverberation provided by the embodiment of the application implement adjustable control of the noise reduction strength by introducing a post-processing mechanism based on a control parameter after noise reduction and dereverberation processing, and performing amplitude control on the first speech signal and the first noise signal. The mechanism does not depend on retraining or structural modification of the front-end AI model, and only by adjusting the first control parameter can the nonlinear mapping relationship of the speech and noise amplitude spectrum be changed, effectively improving the configurability and adaptability of the speech processing system, flexibly adjusting the effect of the AI speech processing module output, meeting the demand for AI noise reduction effect in different scenarios, and without the need to retrain and deploy the AI speech processing module.

[0147] It should be noted that although several modules or units of the device for action execution are mentioned in the foregoing detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into a plurality of modules or units.

[0148] From the above description of the embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.) or a network, and includes a plurality of instructions to make a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) execute the method according to the embodiments of the present disclosure.

[0149] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. It is intended that the disclosure be construed as including any patents, patent applications, publications, publications, or other disclosure of the prior art that may be related to the present disclosure in their entirety.

[0150] The above merely is a specific embodiment of the present application, but the design concept of the present application is not limited thereto, and any non-essential change of the present application using the concept should belong to the act of infringing the protection scope of the present application.

Claims

1. A speech processing method with adjustable noise reduction and de-reverberation levels, characterized in that, Includes the following steps: Acquire the raw speech signal; The original speech signal is subjected to noise reduction and dereverberation processing to obtain a signal processing result, which includes a first speech signal. A first noise signal is obtained by noise estimation based on the original speech signal and the signal processing result; The first speech signal and the first noise signal are post-processed according to preset control parameters to output the target speech signal; The post-processing includes: The amplitudes of the first speech signal and the first noise signal are adjusted according to the preset first adjustment parameters to obtain the corresponding second speech signal and second noise signal.

2. The speech processing method with adjustable noise reduction and de-reverberation levels as described in claim 1, characterized in that, The amplitude control is a convex function control, and the preset first control parameter is a convex function coefficient.

3. The speech processing method with adjustable noise reduction and de-reverberation levels as described in claim 2, characterized in that: The amplitude control uses a convex function: a power function y n =x n α , x n When y is the amplitude spectrum of the first input speech signal, n Let x be the amplitude spectrum of the output second speech signal, α be the first modulation parameter with a range of (0, +∞), and x be the amplitude spectrum of the output second speech signal. n When y is the amplitude spectrum of the first input noise signal, n Let α be the amplitude spectrum of the output second noise signal, and let α be the first control parameter with a range of (0,+∞).

4. The speech processing method with adjustable noise reduction and de-reverberation levels as described in claim 1, characterized in that... The post-processing of the first speech signal and the first noise signal according to preset control parameters further includes: The reverberation of the second speech signal and the second noise signal is adjusted according to the preset second control parameters, and the target speech signal is output.

5. The speech processing method with adjustable noise reduction and de-reverberation levels as described in claim 4, characterized in that, The step of performing reverberation control on the second speech signal and the second noise signal according to a preset second control parameter, and outputting a target speech signal, includes: The amplitude spectrum of the second speech signal and the amplitude spectrum of the second noise signal are reverberated according to the preset second control parameters to obtain the amplitude spectrum of the target speech signal; based on the phase spectrum of the second speech signal, the amplitude spectrum of the target speech signal is combined with the phase spectrum to generate a complex spectrum, which is then obtained by inverse Fourier transform.

6. The speech processing method with adjustable noise reduction and de-reverberation level as described in claim 5, characterized in that, The step of performing reverberation modulation on the amplitude spectrum of the second speech signal and the amplitude spectrum of the second noise signal according to a preset second modulation parameter to obtain the amplitude spectrum of the target speech signal includes: ; Wherein, Ym is the amplitude spectrum of the target speech signal, Y2 is the amplitude spectrum of the second noise signal, Y1 is the amplitude spectrum of the second speech signal, β is the second control parameter, and 0≤β≤1; The step of combining the amplitude spectrum of the target speech signal with the phase spectrum based on the phase spectrum of the second speech signal to generate a complex spectrum, and then performing an inverse Fourier transform to obtain the target speech signal, includes: ; Wherein, output is the complex spectrum of the target speech signal. This represents the phase spectrum of the second speech signal. Represents a complex unit vector.

7. The speech processing method with adjustable noise reduction and de-reverberation level as described in claim 1, characterized in that, in, The signal processing result further includes: a noise signal, wherein the first noise signal is obtained by noise estimation based on the original speech signal and the signal processing result, including: The first noise signal is obtained by noise estimation based on the original speech signal and the noise signal.

8. The speech processing method with adjustable noise reduction and de-reverberation levels as described in claim 1, characterized in that, The step of obtaining a first noise signal by noise estimation based on the original speech signal and the signal processing result includes: A noise signal is obtained by performing noise estimation on the original speech signal and the first speech signal, and the first noise signal is obtained by spectral subtraction based on the original speech signal and the noise signal.

9. A speech processing device with adjustable noise reduction and de-reverberation levels, characterized in that: include The speech signal acquisition module acquires the raw speech signal. The noise reduction and dereverberation processing module performs noise reduction and dereverberation processing on the original speech signal to obtain a signal processing result, and performs noise estimation on the original speech signal and the signal processing result to obtain a first noise signal, wherein the signal processing result includes at least the first speech signal; The post-processing module performs post-processing on the first speech signal and the first noise signal according to preset control parameters, and outputs the target speech signal, wherein the post-processing includes: The amplitudes of the first speech signal and the first noise signal are adjusted according to the preset first adjustment parameters to obtain the corresponding second speech signal and second noise signal.

10. The speech processing device with adjustable noise reduction and de-reverberation level as described in claim 9, characterized in that, The noise reduction and dereverberation processing module uses an AI voice processing model.