Voice signal processing method and device, electronic equipment, storage medium and product
By dividing the speech signal into multiple sub-band feature data and using a pre-trained deep learning network to reconstruct the parameters of the speech detection and noise estimation model, the problem of inaccurate noise estimation in non-stationary noise environments by single-channel speech enhancement algorithms is solved, and higher quality speech signal output is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SPREADTRUM COMMUNICATION (SHANGHAI) CO LTD
- Filing Date
- 2026-02-03
- Publication Date
- 2026-05-01
AI Technical Summary
Existing single-channel speech enhancement algorithms based on statistical models are inaccurate in noise estimation in non-stationary noise environments, resulting in poor speech signal quality.
By dividing the speech signal into multiple sub-band feature data and using a pre-trained deep learning network for speech detection, combined with a noise estimation model for parameter reconstruction and gain calculation, more accurate noise estimation can be achieved.
It improves the quality of speech signals and solves the problem of poor speech signal quality caused by inaccurate noise estimation.
Smart Images

Figure CN121963764A_ABST
Abstract
Description
Speech signal processing methods, devices, electronic equipment, storage media and products Technical Field
[0001] This application relates to the field of communication technology, and in particular to a voice signal processing method, apparatus, electronic device, storage medium, and product. Background Technology
[0002] As users place higher demands on the clarity and stability of voice communication, single-channel voice enhancement technology has been applied in communication systems. Single-channel voice enhancement technology plays an important role in communication systems, and statistical model-based single-channel voice enhancement methods are widely used due to their low computational cost.
[0003] However, single-channel speech enhancement algorithms based on statistical models assume that the noise signal is statistically stationary, making them unsuitable for non-stationary noise environments. Furthermore, related techniques indirectly obtain the speech presence probability by using the ratio of the input speech power spectrum to the corresponding minimum tracking result, and then use a soft-decision algorithm based on the speech presence probability to estimate the noise power spectral density. However, the accuracy of the speech presence probability estimation heavily depends on various threshold parameters; speech presence probability estimation with fixed threshold parameters does not yield high-performance results.
[0004] Therefore, the aforementioned related technologies still suffer from inaccurate noise estimation, resulting in poor quality of the final speech signal. Summary of the Invention
[0005] This application provides a speech signal processing method, apparatus, electronic device, storage medium, and product to achieve more accurate noise estimation and thus improve the quality of speech signals.
[0006] In a first aspect, embodiments of this application provide a speech signal processing method, including:
[0007] Acquire pre-trained deep learning networks and speech signals;
[0008] The speech signal is preprocessed to obtain a complex domain signal.
[0009] The complex domain signal is divided into sub-bands to obtain multiple sub-band feature data.
[0010] The sub-band feature data of each sub-band is input into the pre-trained deep learning network to obtain the speech detection result corresponding to each sub-band feature data;
[0011] Based on the speech detection results corresponding to each sub-band feature data, sub-band parameter reconstruction processing is performed to obtain the speech detection parameters for each frequency point of the speech signal;
[0012] Based on the speech detection parameters at each frequency point of the speech signal, the complex domain signal, and the noise estimation model, the noise estimation result is determined;
[0013] Gain calculation is performed based on the noise estimation results to obtain the gain result; the target time-domain output signal is obtained based on each complex domain signal and the gain result.
[0014] In one possible implementation, obtaining the pre-trained deep learning network includes: acquiring multiple types of feature data and an initial neural network model; training the initial neural network model based on the multiple types of feature data to obtain target model parameters; and updating the initial neural network model based on the target model parameters to obtain the pre-trained deep learning network.
[0015] In one possible implementation, the preprocessing of the speech signal to obtain a complex domain signal includes: performing frame-segmentation and windowing processing on the speech signal to obtain a framed signal; and performing time-domain to frequency-domain transformation processing on the framed signal to obtain a complex domain signal.
[0016] In one possible implementation, the step of performing sub-band parameter reconstruction processing based on the voice detection results corresponding to each of the sub-band feature data to obtain noise estimation parameters includes: inputting the voice detection results corresponding to each of the sub-band feature data into a sub-band voice detection system based on a deep learning network model, so as to recover the voice detection parameters of each frequency point of the entire band through an overlapping addition method.
[0017] In one possible implementation, determining the noise estimation result based on the speech detection parameters at each frequency point of the speech signal, the complex domain signal, and the noise estimation model includes: performing frequency domain smoothing on the complex domain signal to obtain a frequency-smoothed noisy speech power spectrum; performing time-domain smoothing on the frequency-smoothed noisy speech power spectrum to obtain a time-smoothed noisy speech power spectrum; obtaining a coarsely estimated noise power spectrum based on the speech detection parameters at each frequency point of the speech signal, the complex domain signal, and the noise estimation model; obtaining the posterior speech non-existence probability based on the coarsely estimated noise power spectrum and the time-smoothed noisy speech power spectrum; calculating the prior signal-to-noise ratio (SNR) using a decision-guided method; obtaining the prior speech presence probability based on the posterior speech non-existence probability, the posterior SNR, and the prior SNR; obtaining a noise estimation power spectrum smoothing parameter based on the prior speech presence probability and a preset smoothing parameter value; and determining the noise estimation result based on the noise estimation power spectrum smoothing parameter and the complex domain signal.
[0018] In one possible implementation, obtaining the target time-domain output signal based on each of the complex domain signals and the gain result includes: determining a first frequency domain signal after noise reduction and improved clarity based on the complex domain signals and the gain result; performing frequency-domain to time-domain transformation on the first frequency domain signal to obtain a first time-domain signal; and performing frame synthesis on the first time-domain signal to obtain the target time-domain output signal.
[0019] Secondly, embodiments of this application provide a speech signal processing apparatus, comprising:
[0020] The acquisition module is used to acquire pre-trained deep learning networks and speech signals;
[0021] The signal preprocessing module is used to preprocess the speech signal to obtain a complex domain signal;
[0022] The sub-band division module is used to perform sub-band division processing based on the complex domain signal to obtain multiple sub-band feature data;
[0023] The deep learning detection module is used to input the feature data of each sub-band into the pre-trained deep learning network to obtain the speech detection results corresponding to each sub-band feature data;
[0024] The parameter reconstruction module is used to perform sub-band parameter reconstruction processing based on the speech detection results corresponding to each sub-band feature data to obtain the speech detection parameters of each frequency point of the speech signal;
[0025] The noise estimation module is used to determine the noise estimation result based on the speech detection parameters at each frequency point of the speech signal, the complex domain signal, and the noise estimation model.
[0026] The gain module is used to perform gain calculation processing based on the noise estimation results to obtain the gain result;
[0027] The output module is used to obtain the target time-domain output signal based on each of the complex domain signals and the gain result.
[0028] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;
[0029] The memory stores computer-executed instructions;
[0030] The processor executes computer execution instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.
[0031] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.
[0032] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.
[0033] The speech signal processing method, electronic device, storage medium, and program product provided in this application embodiment involve dividing a speech signal into multiple sub-band feature data, obtaining speech detection results based on a pre-trained deep learning network, reconstructing sub-band parameters to obtain speech detection parameters for controlling a noise estimation model, and thus obtaining a noise estimation result. Gain calculation is then performed on the noise estimation result to obtain a gain result. Finally, based on the gain result and the complex domain signal, a target time-domain output signal is obtained. The entire processing utilizes a deep learning network and uses speech detection results to control the noise estimation result based on the noise estimation model, achieving more accurate noise estimation and thus obtaining a more accurate gain result to process the complex domain signal, resulting in the final target time-domain output signal and improving the signal quality of the final output signal. Attached Figure Description
[0034] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0035] Figure 1 is a schematic diagram of a scenario for a speech signal processing method provided in this application;
[0036] Figure 2 is a schematic flowchart of a speech signal processing method provided in an embodiment of this application;
[0037] Figure 3 is a schematic diagram of the sub-band feature data speech detection process based on a pre-trained deep learning network provided in an embodiment of this application;
[0038] Figure 4 is a simplified flowchart of a speech signal processing embodiment provided in this application;
[0039] Figure 5 is a schematic diagram of the structure of a speech signal processing device provided in an embodiment of this application;
[0040] Figure 6 is a schematic diagram of the structure of an electronic device provided in this application.
[0041] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0042] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0043] First, let me explain the terms used in this application:
[0044] SPP: Speech Presence Probability;
[0045] RNN: Recurrent Neural Network;
[0046] CNN: Convolutional Neural Network;
[0047] MCRA: Minima Controlled Recursive Averaging;
[0048] MCRA2: Minima Controlled Recursive Averaging 2;
[0049] IMCRA: Improved Minima Controlled Recursive Averaging;
[0050] OLA: Overlap-And-Add;
[0051] VAD: Voice Activity Detection.
[0052] Figure 1 is a schematic diagram of a scenario for a speech signal processing method provided in this application. As shown in Figure 1, the specific application scenario of this application includes: user 101, user terminal 102, and server 103. User 101 can be a person using user terminal 102. User terminal 102 can be a terminal device equipped with speech signal processing functions, such as hardware products like headphones, mobile phones, tablets, televisions, vehicle central control systems, walkie-talkies, and smartwatches. User terminal 102 is also used to acquire and process speech signals, as well as to obtain trained neural network models from server 103 and use the output enhanced speech signals. Server 103 can be a device that performs feature extraction using multiple types of data blocks and trains neural networks. Server 103 can be a backend service or a cloud server.
[0053] Based on the above scenarios, it is clear that while traditional single-channel speech enhancement algorithms can eliminate relatively stable noise, their suppression effect on non-stationary noise remains insufficient. Therefore, single-channel AI noise reduction methods can be used to handle non-stationary noise. However, because the noise reduction effect of AI is affected by the training data and the parameters of the model itself, some small AI noise reduction models are only applicable to specific scenarios. For example, in in-vehicle central control systems, the model parameters and training data required for AI noise reduction models in urban road environments and highway environments are significantly different. While large AI models can achieve better performance, they require more memory and computing resources, and algorithm latency and power consumption are greatly increased, making them unusable on certain devices such as headphones, walkie-talkies, and smartwatches. Therefore, there is an urgent need for a speech signal processing method to solve the technical problem of inaccurate noise estimation in speech signals across various devices, leading to poor speech signal quality, in order to achieve more accurate noise estimation and thus improve speech signal quality.
[0054] To address the aforementioned technical problems, this application provides a speech signal processing method that first uses speech detection parameters obtained from a deep learning-based sub-band speech detection (VAD) algorithm to control a noise estimation model. This adjusted noise estimation model, combined with sub-band feature data, achieves more accurate noise estimation and performs noise reduction, ultimately outputting a higher-quality speech signal. This solves the technical problem of inaccurate noise estimation leading to poor speech signal quality.
[0055] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0056] Figure 2 is a schematic flowchart of a speech signal processing method provided in an embodiment of this application.
[0057] As shown in Figure 2, the entity executing this voice signal processing method can be the user terminal 102 shown in Figure 1, or other hardware devices with the same function. This embodiment does not impose any restrictions on this.
[0058] As shown in Figure 2, the method includes:
[0059] S201: Acquire pre-trained deep learning network and speech signal.
[0060] In this embodiment, the speech signal can be a time-domain signal y(n) containing noise and speech. The pre-trained deep learning network can be a deep learning neural network obtained by training a model using a large amount of training data of various types of noise and speech.
[0061] Specifically, in an optional embodiment of this application, obtaining the pre-trained deep learning network in step S201 includes:
[0062] S201a: Acquire various types of feature data and an initial neural network model.
[0063] In this embodiment, the various types of feature data can be obtained by extracting features from signals of various types of sound, such as noise, speech, different languages, and different groups of people. The initial neural network model can be an open-source neural network such as a recurrent neural network (RNN) or a convolutional neural network (CNN).
[0064] S201b: Train the initial neural network model based on various types of feature data to obtain the target model parameters.
[0065] In this embodiment, the process of training an initial neural network model using multiple types of feature data involves dividing the multiple types of feature data into training sets, test sets, and validation sets, and then sequentially inputting them into the initial neural network model for model training to adjust the model parameters of the neural network model, ultimately obtaining a neural network model with the required accuracy of the data processing results.
[0066] In this embodiment, model training can be offline or online training on a backend server. The user's processing results of the speech signal can also be fed back to the model training stage to continuously adjust the target model parameters through model training.
[0067] S201c: Update the initial neural network model based on the target model parameters to obtain a pre-trained deep learning network.
[0068] In this embodiment, the obtained target model parameters replace the relevant model parameters in the initial neural network model to obtain a pre-trained deep learning network for use in subsequent speech signal processing.
[0069] S202: Preprocess the speech signal to obtain the complex domain signal.
[0070] In this embodiment, preprocessing the speech signal can be performed on the continuously acquired time-domain discrete signal containing noise and speech. After preliminary processing, the complex domain signal is obtained. The process, where n represents the sampling point number. There are a total of 0 to N sampling points, where n∈[0,N]. The index represents the signal frame number, and k represents the signal frequency index. Complex domain signals are frequency domain signals.
[0071] Specifically, in an optional embodiment of this application, step S202 specifically includes:
[0072] S202a: Perform frame-segmentation and windowing processing on the speech signal to obtain the framed signal.
[0073] In this embodiment, frame-by-frame windowing processing can be applied to long-term non-stationary speech signals. This is transformed into a short-time approximately stationary frame signal. The process involves windowing. The window in frame segmentation and windowing can be a window function, such as a Hamming window or a Hanning window. The function of the window function is to suppress spectral leakage at frame edges.
[0074] S202b: Perform time-domain to frequency-domain transformation on the framed signal to obtain a complex domain signal.
[0075] In this embodiment, the time-domain to frequency-domain DFT transform processing can be performed on the framed signal. The complex domain signal obtained after N-point DFT transformation The process. Among them, the complex domain signal... middle, The frame index is looped, where k represents the frequency point index.
[0076] S203: Perform sub-band division processing on the complex domain signal to obtain multiple sub-band feature data.
[0077] In this embodiment, subband partitioning can be performed by dividing the signal into the Mel spectral domain using a Mel filter bank or by using a sliding window to partition the complex domain signal. The process involves dividing the full-band spectrogram into n sub-band spectrograms, and then extracting features to obtain multiple sub-band feature data.
[0078] Based on the above embodiments, in an optional embodiment of this application, taking a sliding window approach as an example, step S203 involves sub-band division processing based on the complex domain signal to obtain multiple sub-band feature data, including:
[0079] S203a: The full-band spectrum is divided into multiple sub-band spectrums by sliding a window along the frequency axis according to the complex domain signal, where the bandwidth of the sub-band spectrum is the size of the sliding window.
[0080] In this embodiment, the overlap rate of adjacent frequency bands is 50%, but other overlap rates are also possible, and the sliding window is a 50% Hanning window.
[0081] S204: Input the feature data of each sub-band into the pre-trained deep learning network to obtain the speech detection results corresponding to the feature data of each sub-band.
[0082] In this embodiment, the pre-trained deep learning network has been described in the above embodiments, so it will not be repeated here. After the feature data of each sub-band is processed by the pre-trained deep learning network, the speech detection results corresponding to the feature data of each sub-band are obtained. Where bandi represents the subband index, and the range is si represents the initial frequency of subband i, and ei represents the ending frequency of subband i. The probability of noisy speech is a number ranging from 0 to 1, and the results for each sub-band are given. After subband parameter reconstruction, the following is obtained Control the noise estimation parameters based on the statistical model.
[0083] S205: Perform sub-band parameter reconstruction processing based on the speech detection results corresponding to the feature data of each sub-band to obtain the speech detection parameters of each frequency point of the speech signal.
[0084] In this embodiment, the voice detection results corresponding to the feature data of each sub-band are processed by sub-band parameter reconstruction. The voice detection parameters of each frequency point in the entire band can be recovered by the overlap-addition OLA method. .
[0085] For example, the subband spectrum is input into a pre-trained deep learning network to obtain the VAD results for each subband. In this embodiment, after inputting the sub-band spectrum, the speech detection result corresponding to each sub-band spectrogram can be obtained.
[0086] Figure 3 is a schematic diagram of the speech detection process based on subband feature data of a pre-trained deep learning network provided in an embodiment of this application.
[0087] As shown in Figure 3, the input complex domain signal (i.e., the complex spectrum of the input signal in Figure 3) is processed by sub-band division to obtain a sub-band spectrum. The overlap rate of adjacent frequency bands is 50%, corresponding to the overlapping area between the red dashed lines and the blue dashed lines. After a sub-band spectrum is input into a pre-trained deep learning network, the corresponding voice detection (VAD) result is obtained. The VAD results corresponding to the spectrograms of each sub-band are fed into the parameter reconstruction module. In this embodiment, the parameter reconstruction uses the overlap-addition OLA method to recover the voice detection VAD parameters for each frequency point. .
[0088] S206: Determine the noise estimation result based on the speech detection parameters at each frequency point of the speech signal, the complex domain signal, and the noise estimation model.
[0089] In this embodiment, the noise estimation model is controlled according to the speech detection parameters of each frequency point of the speech signal. The noise estimation model can be a noise estimation model based on a statistical model. The noise estimation method used by the model can be any one of MCRA, MCRA2, IMCRA and corresponding improved schemes.
[0090] Based on the above embodiments, as an optional embodiment of this application, step S206 specifically includes:
[0091] S206a: Perform frequency domain smoothing on the complex domain signal to obtain the power spectrum of the noisy speech after frequency domain smoothing;
[0092] S206b: Obtain the time-domain smoothed power spectrum of the noisy speech based on the time-domain smoothed power spectrum of the frequency-domain smoothed speech.
[0093] S206c: A rough estimate of the noise power spectrum is obtained based on the speech detection parameters at each frequency point of the speech signal, the complex domain signal, and the noise estimation model;
[0094] S206d: Based on the roughly estimated noise power spectrum and the time-smoothed noisy speech power spectrum, the probability of the posterior speech not existing is obtained;
[0095] S206e: Calculate the prior signal-to-noise ratio using the decision-guided method;
[0096] S206f: Based on the posterior probability of the absence of speech, the posterior signal-to-noise ratio, and the prior signal-to-noise ratio, obtain the prior probability of the presence of speech.
[0097] S206g: Based on the prior speech existence probability and the preset smoothing parameter value, the noise estimation power spectrum smoothing parameter is obtained;
[0098] S206h: Determine the noise estimation result based on the noise estimation power spectrum smoothing parameter and the complex domain signal.
[0099] In this embodiment, firstly based on the complex domain signal The power spectrum of the noisy speech was calculated. Then, frequency and time domain smoothing is performed to obtain the power spectrum of the noisy speech after frequency domain smoothing. The specific calculation method is as follows:
[0100] Where b is the selected normalized window function. It is a convolution.
[0101] In this embodiment, the power spectrum of the noisy speech after time-domain smoothing The power spectrum of noisy speech after frequency domain smoothing Temporal smoothing can be achieved through computation. The calculation method is as follows:
[0102] ;
[0103] In the formula, For time-domain smoothing parameters, The value can range from 0.1 to 0.2.
[0104] In this embodiment, the preset first smoothing parameter set can be the minimum value tracking smoothing parameter when the data is judged as a speech frame. , The preset second smoothing parameter set can be the minimum tracking smoothing parameter when judged as a noisy frame. , .
[0105] Based on the above embodiments, in an optional embodiment of this application, the above-obtained... The method is as follows:
[0106] .
[0107] in , ,in , To determine the minimum tracking smoothing parameter when converting to speech frames, where , To determine the minimum tracking smoothing parameter when a noisy frame is formed, , The suggested value ranges are 0.99~0.998 and 0~0.1, respectively. , The suggested value ranges are 0~0.1 and 0.1~0.2. This indicates the probability of the speech sound existing.
[0108] In this embodiment, according to , It can be calculated Specifically, calculate The formula is:
[0109] .
[0110] Then With preset threshold The probability of the posterior speech not existing is obtained by comparison.
[0111] ,Right now:
[0112] .
[0113] Then, based on the conditional speech existence probability of Bayes, the speech existence probability is obtained.
[0115] ;
[0116] in For the prior signal-to-noise ratio, the prior signal-to-noise ratio The calculation method is as follows:
[0117] ;
[0118] ; ;
[0119] In the above content, G temp Indicates Wiener gain, This indicates the smoothing parameter used in the decision-guided method. It can take the value 0.98. This represents the posterior signal-to-noise ratio. Then, the power spectrum smoothing parameter for noise estimation is calculated. , Therefore, the noise power spectrum is estimated as follows:
[0120] The final noise estimation power spectrum is the noise estimation result. Represents the power spectrum smoothing parameter estimated from noise. .
[0121] S207: Perform gain calculation based on the noise estimation results to obtain the gain result.
[0122] In this embodiment, gain calculation can be performed using Wiener gain calculation, logarithmic spectral amplitude gain LSA, optimal improved logarithmic spectral amplitude gain OMLSA, and other gain calculation methods.
[0123] Based on the above embodiments, in an optional embodiment of this application, taking Wiener gain calculation as an example, the gain calculation method is as follows:
[0124] ;
[0125] ;
[0126] ;
[0127] in, The gain signal represents the noise estimation result after gain calculation. This represents the posterior signal-to-noise ratio.
[0128] S208: Based on the complex domain signals and gain results, obtain the target time domain output signal.
[0129] In this embodiment, the results of each complex domain cyclic sum and gain can be transformed in the frequency domain and then synthesized in the frame to obtain the final target time-domain output signal. This results in an enhanced speech signal, improving speech quality.
[0130] Specifically, in an optional embodiment of this application, step S208 specifically includes:
[0131] S208a: Based on the complex domain signal and the gain result, determine the first frequency domain signal after noise reduction and improved clarity.
[0132] S208b: Perform frequency-domain to time-domain transformation on the first frequency domain signal to obtain the first time domain signal.
[0133] S208c: Perform frame synthesis on the first time-domain signal to obtain the target time-domain output signal.
[0134] In this embodiment, the first frequency domain signal after noise reduction and improved clarity can be obtained by multiplying the complex domain signal and the gain result to obtain the final enhanced frequency signal as the first frequency domain signal.
[0135] In this embodiment, the frequency-to-time domain transformation of the first frequency domain signal can be the inverse process of the time-to-frequency domain transformation, and the final result is the noise-reduced and enhanced speech signal as the first time domain signal.
[0136] In this embodiment, the first time-domain signal is frame-synthesized to obtain the target time-domain output signal. Then, the first time-domain signals corresponding to each frame are sequentially synthesized into a full-band speech signal for output according to the frame number. This achieves the goal of improving the quality of the speech signal.
[0137] To make the speech signal processing method provided in the above embodiments clearer, the entire processing process will be described using a noisy speech signal as the input signal as an example.
[0138] Figure 4 is a simplified flowchart of a speech signal processing method provided in an embodiment of this application.
[0139] As shown in Figure 4, the method includes the following steps: First, after the noisy speech signal is input, it is framed and windowed to obtain framed signals. The framed signals undergo a time-to-frequency domain DFT transformation to obtain complex domain signals. These complex domain signals are then processed by feature extraction to obtain sub-band feature data. The sub-band feature data is input sequentially according to the frame number of each sub-band into a pre-trained deep learning network. The model parameters of this deep learning network are obtained by training the neural network on a dataset containing features extracted from multiple types of data. After the sub-band feature data is processed by the deep learning network with adjusted model parameters, the corresponding Voice Detection (VAD) data for each sub-band feature data is obtained.
[0140] After sub-band parameter reconstruction, the voice detection VAD data yields voice detection parameters. These parameters are then used to control a statistical model-based noise estimation process for the complex domain signal, resulting in a noise estimation result. This process identifies which sounds are noise and which are speech. Gain calculation is performed on the noise estimation result to obtain the gain. This gain is then multiplied by the complex domain signal to obtain the final, noise-reduced and enhanced first frequency domain signal. This first frequency domain signal undergoes a frequency-to-time domain transformation to obtain the first time domain signal. Finally, the first time domain signal is synthesized frame by frame number to produce the target time domain signal output. This process completes the processing of the noisy speech signal. The entire process incorporates low-resource noise reduction methods and collaborates with AI-powered noise reduction through deep learning networks. Controlling the minimum tracking parameter indirectly controls the neural network, thereby achieving more accurate noise estimation and improving speech quality.
[0141] In summary, the speech signal processing method provided in this application divides the speech signal into multiple sub-band feature data, obtains speech detection results based on a pre-trained deep learning network, reconstructs sub-band parameters to obtain speech detection parameters for controlling a noise estimation model, and then obtains noise estimation results. Gain calculation is then performed on the noise estimation results to obtain a gain result. Finally, based on the gain result and the complex domain signal, the target time-domain output signal is obtained. The entire process utilizes a deep learning network and uses speech detection results to control the noise estimation results based on the noise estimation model, achieving more accurate noise estimation and thus obtaining more accurate gain results to process the complex domain signal, resulting in the final target time-domain output signal and improving the signal quality of the final output signal.
[0142] Figure 5 is a schematic diagram of the structure of a speech signal processing device provided in an embodiment of this application.
[0143] As shown in Figure 5, the device includes: an acquisition module 51, a signal preprocessing module 52, a subband division module 53, a deep learning detection module 54, a parameter reconstruction module 55, a noise estimation module 56, a gain module 57, and an output module 58.
[0144] Acquisition module 51 is used to acquire pre-trained deep learning network and speech signals;
[0145] Signal preprocessing module 52 is used to preprocess the speech signal to obtain a complex domain signal;
[0146] Subband division module 53 is used to perform subband division processing on complex domain signals to obtain multiple subband feature data;
[0147] The deep learning detection module 54 is used to input the feature data of each sub-band into the pre-trained deep learning network to obtain the speech detection results corresponding to the feature data of each sub-band.
[0148] The parameter reconstruction module 55 is used to perform sub-band parameter reconstruction processing based on the speech detection results corresponding to the feature data of each sub-band, so as to obtain the speech detection parameters of each frequency point of the speech signal.
[0149] The noise estimation module 56 is used to determine the noise estimation result based on the speech detection parameters at each frequency point of the speech signal, the complex domain signal, and the noise estimation model.
[0150] Gain module 57 is used to perform gain calculation processing based on noise estimation results to obtain the gain result;
[0151] Output module 58 is used to obtain the target time-domain output signal based on each complex domain signal and gain result.
[0152] In an optional embodiment of this application, the acquisition module 51 is specifically used for: acquiring multiple types of feature data and an initial neural network model; training the initial neural network model based on the multiple types of feature data to obtain target model parameters; and updating the initial neural network model based on the target model parameters to obtain a pre-trained deep learning network.
[0153] In an optional embodiment of this application, the signal processing module 52 is specifically used for: performing frame-by-frame windowing processing on the speech signal to obtain a framed signal; and performing time-domain to frequency-domain transformation processing on the framed signal to obtain a complex domain signal.
[0154] In an optional embodiment of this application, the parameter reconstruction module 55 is specifically used to: input the voice detection results corresponding to the feature data of each sub-band into the sub-band voice detection system based on the deep learning network model, so as to recover the voice detection parameters of each frequency point of the whole band through the overlapping addition method.
[0155] In an optional embodiment of this application, the noise estimation module 56 is specifically used for: performing frequency domain smoothing processing on the complex domain signal to obtain the frequency-smoothed noisy speech power spectrum; performing time-domain smoothing processing on the frequency-smoothed noisy speech power spectrum to obtain the time-smoothed noisy speech power spectrum; obtaining a coarsely estimated noise power spectrum based on the speech detection parameters at each frequency point of the speech signal, the complex domain signal, and the noise estimation model; obtaining the posterior speech non-existence probability based on the coarsely estimated noise power spectrum and the time-smoothed noisy speech power spectrum; calculating the prior signal-to-noise ratio using a decision-guided method; obtaining the prior speech existence probability based on the posterior speech non-existence probability, the posterior signal-to-noise ratio, and the prior signal-to-noise ratio; obtaining the noise estimation power spectrum smoothing parameter based on the prior speech existence probability and a preset smoothing parameter value; and determining the noise estimation result based on the noise estimation power spectrum smoothing parameter and the complex domain signal.
[0156] In an optional embodiment of this application, the output module 58 is specifically used to: determine a first frequency domain signal after noise reduction and improved clarity based on the complex domain signal and the gain result; perform frequency domain to time domain transformation processing on the first frequency domain signal to obtain a first time domain signal; and perform frame synthesis on the first time domain signal to obtain a target time domain output signal.
[0157] The speech signal processing device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0158] Figure 6 is a schematic diagram of the structure of an electronic device provided in this application. As shown in Figure 6, the electronic device 60 provided in this embodiment includes at least one processor 601 and a memory 602. Optionally, the device 60 further includes a communication component 603. The processor 601, memory 602, and communication component 603 are connected via a bus 604.
[0159] In a specific implementation, at least one processor 601 executes computer execution instructions stored in memory 602, causing at least one processor 601 to perform the above-described method.
[0160] The specific implementation process of processor 601 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0161] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0162] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0163] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0164] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0165] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0166] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0167] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0168] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0169] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0170] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0171] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0172] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0173] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A speech signal processing method, characterized in that, include: Acquire pre-trained deep learning networks and speech signals; The speech signal is preprocessed to obtain a complex domain signal; The complex domain signal is divided into subbands to obtain multiple subband feature data. Each subband feature data is input into the pre-trained deep learning network to obtain the corresponding speech detection result. Subband parameter reconstruction is performed based on the speech detection results to obtain speech detection parameters for each frequency point of the speech signal. The noise estimation result is determined based on the speech detection parameters for each frequency point of the speech signal, the complex domain signal, and the noise estimation model. The gain is calculated based on the noise estimation results to obtain the gain result; the target time-domain output signal is obtained based on each complex domain signal and the gain result.
2. The method according to claim 1, characterized in that, The process of obtaining a pre-trained deep learning network includes: acquiring multiple types of feature data and an initial neural network model; training the initial neural network model based on the multiple types of feature data to obtain target model parameters; and updating the initial neural network model based on the target model parameters to obtain a pre-trained deep learning network.
3. The method according to claim 1, characterized in that, The step of preprocessing the speech signal to obtain a complex domain signal includes: performing frame-segmentation and windowing processing on the speech signal to obtain a framed signal; and performing time-domain to frequency-domain transformation processing on the framed signal to obtain a complex domain signal.
4. The method according to claim 1, characterized in that, The step of performing sub-band parameter reconstruction processing based on the voice detection results corresponding to each sub-band feature data to obtain noise estimation parameters includes: inputting the voice detection results corresponding to each sub-band feature data into a sub-band voice detection system based on a deep learning network model, so as to recover the voice detection parameters of each frequency point of the entire band through an overlapping addition method.
5. The method according to claim 1, characterized in that, The step of determining the noise estimation result based on the speech detection parameters at each frequency point of the speech signal, the complex domain signal, and the noise estimation model includes: performing frequency domain smoothing on the complex domain signal to obtain a frequency-smoothed noisy speech power spectrum; performing time-domain smoothing on the frequency-smoothed noisy speech power spectrum to obtain a time-smoothed noisy speech power spectrum; obtaining a coarsely estimated noise power spectrum based on the speech detection parameters at each frequency point of the speech signal, the complex domain signal, and the noise estimation model; obtaining the posterior speech non-existence probability based on the coarsely estimated noise power spectrum and the time-smoothed noisy speech power spectrum; calculating the prior signal-to-noise ratio (SNR) using a decision-guided method; obtaining the prior speech presence probability based on the posterior speech non-existence probability, the posterior SNR, and the prior SNR; obtaining a noise estimation power spectrum smoothing parameter based on the prior speech presence probability and a preset smoothing parameter value; and determining the noise estimation result based on the noise estimation power spectrum smoothing parameter and the complex domain signal.
6. The method according to any one of claims 1 to 5, characterized in that, The step of obtaining the target time-domain output signal based on each of the complex domain signals and the gain result includes: determining a first frequency domain signal after noise reduction and improved clarity based on the complex domain signals and the gain result; performing frequency-domain to time-domain transformation on the first frequency domain signal to obtain a first time-domain signal; and performing frame synthesis on the first time-domain signal to obtain the target time-domain output signal.
7. A speech signal processing device, characterized in that, include: The acquisition module is used to acquire pre-trained deep learning networks and speech signals; The signal preprocessing module is used to preprocess the speech signal to obtain a complex domain signal; The sub-band division module is used to perform sub-band division processing based on the complex domain signal to obtain multiple sub-band feature data; The deep learning detection module is used to input the sub-band feature data of each sub-band into the pre-trained deep learning network to obtain the speech detection result corresponding to each sub-band feature data; the parameter reconstruction module is used to perform sub-band parameter reconstruction processing based on the speech detection result corresponding to each sub-band feature data to obtain the speech detection parameters of each frequency point of the speech signal. The noise estimation module is used to determine the noise estimation result based on the speech detection parameters at each frequency point of the speech signal, the complex domain signal, and the noise estimation model. The gain module is used to perform gain calculation processing based on the noise estimation results to obtain the gain result; The output module is used to obtain the target time-domain output signal based on each of the complex domain signals and the gain result.
8. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 6.