Voice Enhancement Method, Device, Equipment and Storage Medium
By predicting the glottic parameters and excitation signals of speech frames, and reconstructing the speech signal using neural networks, the problem of noise interference in voice communication is solved, and high-quality speech enhancement effect is achieved.
Patent Information
- Application Number
- CN202110171244.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-08
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-02-08
AI Technical Summary
Voice signals are easily mixed with noise during communication, resulting in poor communication quality and affecting the user's auditory experience.
By predicting the glottal parameters, gain and excitation signals of the speech frame, the neural network is used for speech enhancement processing, and the original speech signal is reconstructed to remove noise.
Effectively remove noise, retain the original structure of the voice signal, improve voice quality, and improve user experience.
Smart Images

Figure CN113571079B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing technologies, and more particularly, to a speech enhancement method, apparatus, device, and storage medium. Background Art
[0002] Due to the convenience and timeliness of voice communication, the application of voice communication is becoming more and more widespread. For example, voice signals are transmitted between participants in a cloud conference. In voice communication, voice signals may be mixed with noise, and the noise mixed in the voice signals will result in poor communication quality, greatly affecting the user's auditory experience. Therefore, how to perform speech enhancement processing to remove the noise part is a technical problem that needs to be solved urgently in the prior art. Summary of the Invention
[0003] Embodiments of the present application provide a speech enhancement method, apparatus, device, and storage medium to achieve speech enhancement.
[0004] Other features and advantages of the present application will become apparent from the following detailed description, or will be partially learned through the practice of the present application.
[0005] According to one aspect of the embodiments of the present application, a speech enhancement method is provided, including:
[0006] Predicting glottal parameters according to the frequency-domain representation of a target speech frame to obtain the glottal parameters corresponding to the target speech frame;
[0007] Predicting the gain of the target speech frame according to the gain corresponding to the historical speech frames of the target speech frame to obtain the gain corresponding to the target speech frame;
[0008] Predicting an excitation signal according to the frequency-domain representation of the target speech frame to obtain the excitation signal corresponding to the target speech frame;
[0009] Performing synthesis processing on the glottal parameters corresponding to the target speech frame, the gain corresponding to the target speech frame, and the excitation signal corresponding to the target speech frame to obtain the enhanced speech signal corresponding to the target speech frame.
[0010] According to one aspect of the embodiments of the present application, a speech enhancement apparatus is provided, including:
[0011] A glottal parameter prediction module, configured to predict glottal parameters according to the frequency-domain representation of a target speech frame to obtain the glottal parameters corresponding to the target speech frame;
[0012] A gain prediction module, configured to predict the gain of the target speech frame according to the gain corresponding to the historical speech frames of the target speech frame to obtain the gain corresponding to the target speech frame;
[0013] An excitation signal prediction module, configured to predict an excitation signal according to the frequency-domain representation of the target speech frame, so as to obtain an excitation signal corresponding to the target speech frame;
[0014] A synthesis module, configured to perform synthesis processing on the glottal parameters corresponding to the target speech frame, the gain corresponding to the target speech frame, and the excitation signal corresponding to the target speech frame, so as to obtain an enhanced speech signal corresponding to the target speech frame.
[0015] In some embodiments of the present application, based on the foregoing solution, the synthesis module includes: a glottal filter construction unit, configured to construct a glottal filter according to the glottal parameters corresponding to the target speech frame. A filtering unit, configured to filter the excitation signal corresponding to the target speech frame through the glottal filter to obtain a first speech signal. An amplification unit, configured to amplify the first speech signal according to the gain corresponding to the target speech frame to obtain an enhanced speech signal corresponding to the target speech frame.
[0016] In some embodiments of the present application, based on the foregoing solution, the target speech frame includes a plurality of sample points; the glottal filter is a K-order filter, where K is a positive integer; the excitation signal includes excitation signal values corresponding to the plurality of sample points in the target speech frame respectively; the filtering unit includes: a convolution unit, configured to convolve the excitation signal values corresponding to the first K sample points of each sample point in the target speech frame with the K-order filter to obtain a target signal value of each sample point in the target speech frame; a combination unit, configured to combine the target signal values corresponding to all sample points in the target speech frame in chronological order to obtain the first speech signal.
[0017] In some embodiments of the present application, based on the foregoing solution, the glottal filter is a K-order filter, and the glottal parameters include K-order line spectrum frequency parameters or K-order linear prediction coefficients.
[0018] In some embodiments of the present application, based on the foregoing solution, the glottal parameter prediction module includes: a first input unit, configured to input the frequency-domain representation of the target speech frame into a first neural network, and the first neural network is trained according to the frequency-domain representation of the sample speech frame and the glottal parameters corresponding to the sample speech frame; a first output unit, configured to output the glottal parameters corresponding to the target speech frame by the first neural network according to the frequency-domain representation of the target speech frame.
[0019] In some embodiments of the present application, based on the foregoing solution, the glottal parameter prediction module 1210 is further configured to: use the glottal parameters corresponding to the historical speech frames of the target speech frame as a reference, and perform glottal parameter prediction according to the frequency-domain representation of the target speech frame to obtain the glottal parameters corresponding to the target speech frame.
[0020] In some embodiments of the present application, based on the foregoing solution, the glottal parameter prediction module includes: a second input unit, configured to input the frequency-domain representation of the target speech frame and the glottal parameters corresponding to the historical speech frames of the target speech frame into a first neural network, where the first neural network is trained by using the frequency-domain representation of the sample speech frames, the glottal parameters corresponding to the sample speech frames, and the glottal parameters corresponding to the historical speech frames of the sample speech frames; a second output unit, configured to predict by the first neural network according to the frequency-domain representation of the target speech frame and the glottal parameters corresponding to the historical speech frames of the target speech frame, and output the glottal parameters corresponding to the target speech frame.
[0021] In some embodiments of the present application, based on the foregoing solution, the gain prediction module includes: a third input unit, configured to input the gain corresponding to the historical speech frames of the target speech frame into a second neural network; the second neural network is trained according to the gain corresponding to the sample speech frames and the gain corresponding to the historical speech frames of the sample speech frames; a third output unit, configured to output the target gain by the second neural network according to the gain corresponding to the historical speech frames of the target speech frame.
[0022] In some embodiments of the present application, based on the foregoing solution, the excitation signal prediction module includes: a fourth input unit, configured to input the frequency-domain representation of the target speech frame into a third neural network; the third neural network is trained according to the frequency-domain representation of the sample speech frames and the frequency-domain representation of the excitation signals corresponding to the sample speech frames; a fourth output unit, configured to output the frequency-domain representation of the excitation signal corresponding to the target speech frame by the third neural network according to the frequency-domain representation of the target speech frame.
[0023] In some embodiments of the present application, based on the foregoing solution, the speech enhancement device further includes: an acquisition module, configured to acquire the time-domain signal of the target speech frame; a time-frequency transformation module, configured to perform time-frequency transformation on the time-domain signal of the target speech frame to obtain the frequency-domain representation of the target speech frame.
[0024] In some embodiments of the present application, based on the foregoing solution, the acquisition module is further configured to: acquire a second speech signal, where the second speech signal is a collected speech signal or a speech signal obtained by decoding an encoded speech; perform frame segmentation on the second speech signal to obtain the time-domain signal of the target speech frame.
[0025] In some embodiments of the present application, the speech enhancement device further includes: a processing module, configured to play or encode and transmit the enhanced speech signal corresponding to the target speech frame.
[0026] According to one aspect of the embodiments of the present application, there is provided an electronic device, including: a processor; a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the voice enhancement method described above is implemented.
[0027] According to one aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the voice enhancement method described above is implemented.
[0028] In the solution of the present application, the glottal parameters and the excitation signal for reconstructing the original voice signal in the target voice frame are predicted based on the frequency-domain representation of the target voice frame, and the gain for reconstructing the original voice signal in the target voice frame is predicted based on the gain of the historical voice frames of the target voice frame. Then, voice synthesis is performed according to the predicted glottal parameters, the corresponding excitation signal, and the corresponding gain corresponding to the target voice frame, which is equivalent to reconstructing the original voice signal in the target voice frame. The signal obtained through the synthesis process is the enhanced voice signal corresponding to the target voice frame, realizing the enhancement of the voice frame.
[0029] Moreover, in this solution, the predicted glottal parameters are strongly correlated with the glottal characteristics of the physical process of voice generation. Therefore, synthesizing voice according to the predicted glottal parameters can maintain the voice structure of the original voice signal and can maximize the avoidance of voice clipping.
[0030] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Brief Description of the Drawings
[0031] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the following-described accompanying drawings are only some embodiments of the present application. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts. In the accompanying drawings:
[0032] Figure 1 is a schematic diagram of a voice communication link in a VoIP system shown according to a specific embodiment.
[0033] Figure 2 shows a schematic diagram of a digital model of voice signal generation.
[0034] Figure 3 shows a schematic diagram of decomposing an original voice signal into an excitation signal and the frequency response of a glottal filter.
[0035] Figure 4It is a flowchart of a voice enhancement method shown according to an embodiment of the present application.
[0036] Figure 5 is Figure 4 A flowchart of step 440 of the corresponding embodiment in one embodiment.
[0037] Figure 6 It is a schematic diagram of performing short-time Fourier transform on a voice frame in a windowed overlapping manner shown according to an embodiment of the present application.
[0038] Figure 7 It is a flowchart of voice enhancement shown according to a specific embodiment of the present application;
[0039] Figure 8 It is a schematic diagram of a first neural network shown according to an embodiment of the present application.
[0040] Figure 9 It is a schematic diagram of the input and output of a first neural network shown according to another embodiment of the present application.
[0041] Figure 10 It is a schematic diagram of a second neural network shown according to an embodiment of the present application.
[0042] Figure 11 It is a schematic diagram of a third neural network shown according to an embodiment of the present application.
[0043] Figure 12 It is a block diagram of a voice enhancement device shown according to an embodiment of the present application.
[0044] Figure 13 It shows a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners
[0045] Now, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0046] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0047] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0048] The flowcharts shown in the drawings are only illustrative descriptions and do not necessarily include all content and operations / steps, nor do they have to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.
[0049] It should be noted that: "a plurality of" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0050] Noise in the voice signal will greatly reduce the voice quality and affect the user's auditory experience. Therefore, in order to improve the quality of the voice signal, it is necessary to perform enhancement processing on the voice signal to remove the noise as much as possible and retain the original voice signal in the signal (i.e., the pure signal without noise). In order to implement the enhancement processing of the voice, the solution of this application is proposed.
[0051] The solution of this application can be applied to the application scenarios of voice calls, such as voice communication through instant messaging applications and voice calls in game applications. Specifically, voice enhancement can be performed according to this solution at the voice sending end, the voice receiving end, or the server providing the voice communication service.
[0052] Cloud conferencing is an important part of online office work. In a cloud conference, after the voice collection device of the participants in the cloud conference collects the voice signal of the speaker, it needs to send the collected voice signal to other conference participants. This process involves the transmission and playback of the voice signal among multiple participants. If the noise signal mixed in the voice signal is not processed, it will greatly affect the auditory experience of the conference participants. In this scenario, the solution of this application can be applied to enhance the voice signal in the cloud conference, so that the voice signal heard by the conference participants is the enhanced voice signal, improving the quality of the voice signal.
[0053] Cloud conferencing is an efficient, convenient, and low-cost conferencing form based on cloud computing technology. Users only need to perform simple and easy operations through an Internet interface to quickly and efficiently synchronously share voice, data files, and videos with teams and customers around the world. The cloud conferencing service provider helps users operate complex technologies such as data transmission and processing during the conference.
[0054] Currently, domestic cloud conferencing mainly focuses on service contents in the form of SaaS (Software as a Service), including service forms such as telephone, network, and video. Video conferencing based on cloud computing is called cloud conferencing. In the era of cloud conferencing, the transmission, processing, and storage of data are all processed by the computer resources of the video conferencing provider. Users no longer need to purchase expensive hardware and install cumbersome software at all. They only need to open the client and enter the corresponding interface to conduct efficient remote conferences.
[0055] The cloud conferencing system supports multi-server dynamic cluster deployment and provides multiple high-performance servers, greatly improving the stability, security, and availability of conferences. In recent years, video conferencing has been welcomed by many users because it can significantly improve communication efficiency, continuously reduce communication costs, and bring about an upgrade in internal management level. It has been widely applied in various fields such as transportation, logistics, finance, operators, education, and enterprises.
[0056] Figure 1 It is a schematic diagram of a voice communication link in a VoIP (Voice over Internet Protocol) system shown according to a specific embodiment. As Figure 1 shown, based on the network connection between the sending end 110 and the receiving end 120, the sending end 110 and the receiving end 120 can perform voice transmission.
[0057] As Figure 1 shown, the sending end 110 includes an acquisition module 111, a pre-enhancement processing module 112, and an encoding module 113. Among them, the acquisition module 111 is used to acquire voice signals, and it can convert the acquired acoustic signals into digital signals; the pre-enhancement processing module 112 is used to enhance the acquired voice signals to remove the noise in the acquired voice signals and improve the quality of the voice signals. The encoding module 113 is used to encode the enhanced voice signals to improve the anti-interference ability of the voice signals during transmission. The pre-enhancement processing module 112 can perform voice enhancement according to the method of this application. After enhancing the voice, it is then encoded, compressed, and transmitted, so as to ensure that the signals received by the receiving end are no longer affected by noise.
[0058] The receiving end 120 includes a decoding module 121, a post-enhancement module 122, and a playback module 123. The decoding module 121 is used to decode the received encoded voice to obtain a decoded signal; the post-enhancement module 122 is used to perform enhancement processing on the decoded voice signal; the playback module 123 is used to play the enhanced voice signal. The post-enhancement module 122 can also perform voice enhancement according to the method of this application. In some embodiments, the receiving end 120 may further include a sound effect adjustment module, and this sound effect adjustment module is used to adjust the sound effect of the enhanced voice signal.
[0059] In a specific embodiment, it may be that only the receiving end 120 or only the sending end 110 performs voice enhancement according to the method of this application. Of course, it may also be that both the sending end 110 and the receiving end 120 perform voice enhancement according to the method of this application.
[0060] In some application scenarios, in addition to supporting VoIP communication, the terminal device in a VoIP system can also support other third-party protocols, such as traditional PSTN (Public Switched Telephone Network) circuit domain telephones. However, traditional PSTN services cannot perform voice enhancement. In this scenario, voice enhancement can be performed according to the method of this application in the terminal acting as the receiving end.
[0061] When specifically describing the solution of this application, it is necessary to introduce the generation of voice signals. Voice signals are generated by the physiological movements of the human vocal organs under the control of the brain, that is: at the trachea, an impact signal similar to noise with a certain energy is generated (equivalent to an excitation signal); the impact signal impacts the human vocal cords (the vocal cords are equivalent to a glottis filter), generating a quasi-periodic opening and closing; after being amplified by the oral cavity, a sound is emitted (the voice signal is output).
[0062] Figure 2 The schematic diagram of the digital model for voice signal generation is shown, and the generation process of voice signals can be described through this digital model. As Figure 2 shown, after the excitation signal impacts the glottis filter, it is then subjected to gain control and the voice signal is output, where the glottis filter is defined by glottis parameters. This process can be expressed by the following formula:
[0063] x(n) = G · r(n) · ar(n); (Formula 1)
[0064] Among them, x(n) represents the input voice signal; G represents the gain, which can also be called the linear prediction gain; r(n) represents the excitation signal; ar(n) represents the glottis filter.
[0065] Figure 3A schematic diagram showing the decomposition of an original speech signal into an excitation signal and the frequency response of a glottal filter is shown. Figure 3 a shows a schematic diagram of the frequency response of the original speech signal. Figure 3 b shows a schematic diagram of the frequency response of the glottal filter decomposed according to the original speech signal. Figure 3 c shows a schematic diagram of the frequency response of the excitation signal decomposed according to the original speech signal. As Figure 3 shown, the undulating part in the frequency response diagram of the original speech signal corresponds to the peak position in the frequency response diagram of the glottal filter. The excitation signal is equivalent to the residual signal after LP (Linear Prediction) analysis of the original speech signal. Therefore, its corresponding frequency response is relatively flat.
[0066] It can be seen from the above that an excitation signal, a glottal filter, and a gain can be decomposed from an original speech signal (i.e., a speech signal without noise). The decomposed excitation signal, glottal filter, and gain can be used to represent the original speech signal. Among them, the glottal filter can be expressed by glottal parameters. Conversely, if the excitation signal corresponding to an original speech signal, the glottal parameters for determining the glottal filter, and the gain are known, then the original speech signal can be reconstructed according to the corresponding excitation signal, glottal filter, and gain.
[0067] The solution of this application is exactly based on this principle. The glottal parameters, excitation signal, and gain corresponding to the original speech signal in a speech signal to be processed are predicted, and then speech synthesis is performed based on the obtained glottal parameters, excitation signal, and gain. The synthesized speech signal is equivalent to the original speech signal in the speech signal to be processed. Therefore, the synthesized signal is equivalent to the signal with noise removed. This process realizes the enhancement of the speech signal to be processed. Therefore, the synthesized signal can also be called the enhanced speech signal corresponding to the speech signal to be processed.
[0068] Figure 4 is a flowchart of a speech enhancement method shown according to an embodiment of this application. This method can be executed by a computer device with processing capabilities, such as a server, a terminal, etc., which is not specifically limited here. Referring to Figure 4 shown, this method at least includes steps 410 to 440, which are introduced in detail as follows:
[0069] Step 410, perform glottal parameter prediction according to the frequency domain representation of the target speech frame to obtain the glottal parameters corresponding to the target speech frame.
[0070] The speech signal varies non-stationarily and randomly over time. However, the characteristics of the speech signal are strongly correlated within a short period, that is, the speech signal has short-term correlation. Therefore, in the solution of this application, speech enhancement is performed in units of speech frames. The target speech frame refers to the speech frame to be enhanced currently.
[0071] The frequency-domain representation of the target speech frame can be obtained by performing time-frequency transformation on the time-domain signal of the target speech frame. The time-frequency transformation is, for example, the Short-term Fourier transform (STFT). The frequency-domain representation can be the amplitude spectrum, complex frequency spectrum, etc., which are not specifically limited herein.
[0072] The glottal parameters refer to the parameters used to construct the glottal filter. When the glottal parameters are determined, the glottal filter is correspondingly determined. The glottal filter is a digital filter. The glottal parameters can be the Linear Prediction Coefficients (LPC) coefficients or the Line Spectral Frequency (LSF) parameters. The number of glottal parameters corresponding to the target speech frame is related to the order of the glottal filter. If the glottal filter is a K-order filter, the glottal parameters include K-order LSF parameters or K-order LPC coefficients, where the LSF parameters and LPC coefficients can be converted into each other.
[0073] A glottal filter of order p can be expressed as:
[0074] A p (z)=1 + a1z -1 + a2z -2 +…+ a p z -p ;(Formula 2)
[0075] where a1, a2, …, a p are the LPC coefficients; p is the order of the glottal filter; z is the input signal of the glottal filter.
[0076] Based on Formula 2, if we let:
[0077] P(z)=A p (z)-z -(p+1) A p (z -1 );(Formula 3)
[0078] Q(z)=A p (z)+z -(p+1) A p (z -1 );(Formula 4)
[0079] We can get:
[0080]
[0081] Physically speaking, P(z) and Q(z) respectively represent the periodic change laws of glottal opening and glottal closing. The roots of the polynomials P(z) and Q(z) appear alternately in the complex plane, and their distribution is a series of angular frequencies on the unit circle of the complex plane. The LSF parameters are the angular frequencies corresponding to the roots of P(z) and Q(z) on the unit circle of the complex plane. The LSF parameter LSF(n) corresponding to the nth speech frame can be expressed as ω n , of course, the LSF parameter LSF(n) corresponding to the nth speech frame can also be directly expressed by the roots of P(z) corresponding to the nth speech frame and the roots of Q(z). Define the roots of P(z) and Q(z) corresponding to the nth speech frame in the complex plane as θ n , then the LSF parameter corresponding to the nth speech frame is expressed as:
[0082]
[0083] where Rel{θ n} represents the real part of the complex number θ n ; Image{θ n} represents the imaginary part of the complex number θ n .
[0084] In step 410, the glottal parameter prediction performed refers to predicting the glottal parameters used to reconstruct the original speech signal in the target speech frame. In one embodiment, the glottal parameters corresponding to the target speech frame can be predicted through a trained neural network model.
[0085] In some embodiments of the present application, step 410 includes: inputting the frequency-domain representation of the target speech frame into a first neural network, where the first neural network is trained according to the frequency-domain representation of the sample speech frame and the glottal parameters corresponding to the sample speech frame; and outputting, by the first neural network, the glottal parameters corresponding to the target speech frame according to the frequency-domain representation of the target speech frame.
[0086] The first neural network refers to a neural network model for performing glottal parameter prediction. Among them, the first neural network can be a model constructed by a long short-term memory neural network, a convolutional neural network, a recurrent neural network, a fully connected neural network, etc., which is not specifically limited here.
[0087] The frequency-domain representation of the sample speech frame is obtained by performing time-frequency transformation on the time-domain signal of the sample speech frame. This frequency-domain representation can be an amplitude spectrum, a complex frequency spectrum, etc., which is not specifically limited here.
[0088] In some embodiments of the present application, the signal indicated by the sample speech frame may be obtained by combining a known original speech signal and a known noise signal. Then, when the original speech signal is known, the glottal parameters corresponding to each sample speech frame can be obtained by performing linear prediction analysis on the original speech signal.
[0089] During the training process, after inputting the frequency-domain representation of the sample speech frame into the first neural network, the first neural network predicts the glottal parameters according to the frequency-domain representation of the sample speech frame and outputs the predicted glottal parameters. Then, the predicted glottal parameters are compared with the glottal parameters corresponding to the original speech signal in this sample speech frame. If the two are inconsistent, the parameters of the first neural network are adjusted until the predicted glottal parameters output by the first neural network according to the frequency-domain representation of the sample speech frame are consistent with the glottal parameters corresponding to the original speech signal in this sample speech frame. After the training is completed, the first neural network learns the ability to accurately predict the glottal parameters corresponding to the original speech signal in the input speech frame according to the frequency-domain representation of the input speech frame.
[0090] In some embodiments of the present application, since there is a correlation between speech frames and the similarity of the frequency-domain features between two adjacent speech frames is relatively high, therefore, the glottal parameters corresponding to the target speech frame can be predicted by combining the glottal parameters corresponding to the historical speech frames before the target speech frame. In this embodiment, step 410 includes: using the glottal parameters corresponding to the historical speech frames of the target speech frame as a reference, and predicting the glottal parameters according to the frequency-domain representation of the target speech frame to obtain the glottal parameters corresponding to the target speech frame.
[0091] Since there is a correlation between the historical speech frame and the target speech frame, and there is a similarity between the glottal parameters corresponding to the historical speech frame of the target speech frame and the glottal parameters corresponding to the target speech frame, therefore, using the glottal parameters corresponding to the original speech signal in the historical speech frame of the target speech frame as a reference to supervise the prediction process of the glottal parameters of the target speech frame can improve the accuracy of glottal parameter prediction.
[0092] In an embodiment of the present application, since the similarity of the glottal parameters of the closer speech frames is higher, therefore, using the glottal parameters corresponding to the historical speech frames closer to the target speech frame as a reference can further ensure the prediction accuracy. For example, the glottal parameters corresponding to the previous speech frame of the target speech frame can be used as a reference. In a specific embodiment, the number of historical speech frames used as a reference can be one frame or multiple frames, which can be specifically selected according to actual needs.
[0093] The glottal parameters corresponding to the historical speech frames of the target speech frame can be the glottal parameters predicted for the historical speech frame. In other words, during the process of glottal parameter prediction, the glottal parameters predicted for the historical speech frame are reused to supervise the glottal parameter prediction process of the current speech frame.
[0094] In some embodiments of the present application, in the scenario of predicting glottal parameters using a first neural network, in addition to using the frequency-domain representation of the target speech frame as an input, the glottal parameters corresponding to the historical speech frames of the target speech frame are also used as an input to the first neural network to perform glottal parameter prediction. In this embodiment, step 410 includes: inputting the frequency-domain representation of the target speech frame and the glottal parameters corresponding to the historical speech frames of the target speech frame into the first neural network, where the first neural network is trained using the frequency-domain representation of the sample speech frames, the glottal parameters corresponding to the sample speech frames, and the glottal parameters corresponding to the historical speech frames of the sample speech frames; and predicting, by the first neural network, based on the frequency-domain representation of the target speech frame and the glottal parameters corresponding to the historical speech frames of the target speech frame, and outputting the glottal parameters corresponding to the target speech frame.
[0095] During the training process of the first neural network in this embodiment, the frequency-domain representation of the sample speech frame and the glottal parameters corresponding to the historical speech frames of the sample speech frame are input into the first neural network, and the first neural network outputs predicted glottal parameters. If the output predicted glottal parameters are inconsistent with the glottal parameters corresponding to the original speech signal in the sample speech frame, the parameters of the first neural network are adjusted until the output predicted glottal parameters are consistent with the glottal parameters corresponding to the original speech signal in the sample speech frame. After the training is completed, the first neural network learns the ability to predict the glottal parameters for reconstructing the original speech signal in the speech frame based on the frequency-domain representation of the speech frame and the glottal parameters corresponding to the historical speech frames of the speech frame.
[0096] Please continue to refer to Figure 4 , step 420, perform gain prediction on the target speech frame according to the gain corresponding to the historical speech frames of the target speech frame to obtain the gain corresponding to the target speech frame.
[0097] The gain corresponding to the historical speech frame refers to the gain used to reconstruct the original speech signal in the historical speech frame. Similarly, the predicted gain corresponding to the target speech frame in step 420 is used to reconstruct the original speech signal in the target speech frame.
[0098] In some embodiments of the present application, deep learning can be used to predict the gain of the target speech frame. That is, the gain prediction is performed through a constructed neural network model. For ease of description, the neural network model used for gain prediction is referred to as the second neural network. The second neural network can be a model constructed by a long short-term memory neural network, a convolutional neural network, a fully connected neural network, etc.
[0099] In an embodiment of the present application, step 420 may include: inputting the gain corresponding to the historical speech frame of the target speech frame into the second neural network; the second neural network is trained according to the gain corresponding to the sample speech frame and the gain corresponding to the historical speech frame of the sample speech frame; and outputting the target gain by the second neural network according to the gain corresponding to the historical speech frame of the target speech frame.
[0100] The signal indicated by the sample speech frame can be obtained by combining a known original speech signal and a known noise signal. Therefore, when the original speech signal is known, linear prediction analysis can be performed on the original speech signal to correspondingly determine the gain corresponding to each sample speech frame, that is, the gain used to reconstruct the original speech signal in the sample speech frame.
[0101] The gain corresponding to the historical speech frame of the target speech frame can be obtained by the second neural network predicting the gain for the historical speech frame. In other words, the gain predicted for the historical speech frame is reused as the input to the second neural network model during the gain prediction process for the target speech frame.
[0102] During the training of the second neural network, the gain corresponding to the historical speech frame of the sample speech frame is input into the second neural network, and then the second neural network performs gain prediction according to the input gain corresponding to the historical speech frame of the sample speech frame and outputs a predicted gain; then, the parameters of the second neural network are adjusted according to the predicted gain and the gain corresponding to the sample speech frame, that is: if the predicted gain is inconsistent with the gain corresponding to the sample speech frame, the parameters of the second neural network are adjusted until the predicted gain output by the second neural network for the sample speech frame is consistent with the gain corresponding to the sample speech frame. Through the above training process, the second neural network can learn the ability to predict the gain corresponding to a speech frame according to the gain corresponding to the historical speech frame of the speech frame, so as to accurately perform gain prediction.
[0103] Step 430, predicting an excitation signal according to the frequency domain representation of the target speech frame to obtain the excitation signal corresponding to the target speech frame.
[0104] The excitation signal prediction performed in step 430 refers to predicting the excitation signal corresponding to the original speech signal for reconstructing the target speech frame. Therefore, the excitation signal corresponding to the target speech frame can be used to reconstruct the original speech signal in the target speech frame.
[0105] In some embodiments of the present application, the excitation signal prediction can be performed in a deep learning manner, that is, the excitation signal prediction is performed through a constructed neural network model. For ease of description, the neural network model used for excitation signal prediction is referred to as the third neural network. The third neural network can be a model constructed by a long short-term memory neural network, a convolutional neural network, a fully connected neural network, etc.
[0106] In some embodiments of the present application, step 430 includes: inputting the frequency domain representation of the target speech frame into the third neural network; the third neural network is trained according to the frequency domain representation of the sample speech frame and the frequency domain representation of the excitation signal corresponding to the sample speech frame; and the third neural network outputs the frequency domain representation of the excitation signal corresponding to the target speech frame according to the frequency domain representation of the target speech frame.
[0107] The excitation signal corresponding to the sample speech frame refers to the excitation signal that can be used to reconstruct the original speech signal in the sample speech frame. The excitation signal corresponding to the sample speech frame can be determined by performing linear prediction analysis on the original speech signal in the sample speech frame. The frequency domain representation of the excitation signal can be the amplitude spectrum or complex spectrum of the excitation signal, which is not specifically limited herein.
[0108] During the training of the third neural network, the frequency domain representation of the sample speech frame is input into the third neural network model, and then the third neural network performs excitation signal prediction according to the input frequency domain representation of the sample speech frame and outputs the frequency domain representation of the predicted excitation signal; then, the parameters of the third neural network are adjusted according to the frequency domain representation of the predicted excitation signal and the frequency domain representation of the excitation signal corresponding to the sample speech frame, that is: if the frequency domain representation of the predicted excitation signal is inconsistent with the frequency domain representation of the excitation signal corresponding to the sample speech frame, the parameters of the third neural network are adjusted until the frequency domain representation of the predicted excitation signal output by the third neural network for the sample speech frame is consistent with the frequency domain representation of the excitation signal corresponding to the sample speech frame. Through the above training process, the third neural network can learn the ability to predict the excitation signal corresponding to a speech frame according to the frequency domain representation of the speech frame, so as to accurately perform excitation signal prediction.
[0109] Step 440, performing synthesis processing on the glottal parameters corresponding to the target speech frame, the gain corresponding to the target speech frame, and the excitation signal corresponding to the target speech frame to obtain the enhanced speech signal corresponding to the target speech frame.
[0110] After obtaining the glottal parameters corresponding to the target speech frame, the gain corresponding to the target speech frame, and the excitation signal corresponding to the target speech frame, linear prediction analysis can be performed based on these three parameters to achieve synthesis processing, and an enhanced signal corresponding to the target speech frame can be obtained. Specifically, a glottal filter can be constructed according to the glottal parameters corresponding to the target speech frame first, and then speech synthesis can be performed in accordance with the above formula (1) by combining the gain corresponding to the target speech frame and the corresponding excitation signal to obtain an enhanced speech signal corresponding to the target speech frame.
[0111] In some embodiments of the present application, as Figure 5 shown, step 440 includes:
[0112] Step 510, constructing a glottal filter according to the glottal parameters corresponding to the target speech frame.
[0113] If the glottal parameters are LPC coefficients, the glottal filter can be directly constructed according to the above formula (2). If the glottal filter is a K-order filter, the glottal parameters corresponding to the target speech frame include K-order LPC coefficients, that is, a1, a2,..., a K , in other embodiments, the constant 1 in the above formula (2) can also be used as an LPC coefficient.
[0114] If the glottal parameters are LSF parameters, the LSF parameters can be converted into LPC coefficients, and then the glottal filter can be constructed correspondingly according to the above formula (2).
[0115] Step 520, filtering the excitation signal corresponding to the target speech frame through the glottal filter to obtain a first speech signal.
[0116] The filtering process is convolution in the time domain. Therefore, the process of filtering the excitation signal through the glottal filter as described above can be transformed into the time domain. Then, based on the prediction of the frequency-domain representation of the excitation signal corresponding to the target speech frame, the frequency-domain representation of the excitation signal is transformed into the time domain to obtain the time-domain signal of the excitation signal corresponding to the target speech frame.
[0117] In the solution of the present application, the target speech frame is a digital signal, which includes multiple sample points. The excitation signal is filtered through a glottal filter, that is, the historical sample points before a sample point are convolved with the glottal filter to obtain the target signal value corresponding to the sample point. In some embodiments of the present application, the target speech frame includes multiple sample points; the glottal filter is a K-order filter, where K is a positive integer; the excitation signal includes the excitation signal values corresponding to the multiple sample points in the target speech frame; according to the above filtering process, step 520 includes: convolving the excitation signal values corresponding to the first K sample points of each sample point in the target speech frame with the K-order filter to obtain the target signal value of each sample point in the target speech frame; combining the target signal values corresponding to all the sample points in the target speech frame in chronological order to obtain the first speech signal. Among them, the expression of the K-order filter can refer to the above formula (1). That is to say, for each sample point in the target speech frame, the excitation signal values corresponding to the first K sample points before it are used to convolve with the K-order filter to obtain the target signal value corresponding to each sample point.
[0118] It can be understood that for the first sample point in the target speech frame, it is necessary to calculate the target signal value corresponding to the first sample point by means of the excitation signal values of the last K sample points in the previous speech frame of the target speech frame. Similarly, for the second sample point in the target speech frame, it is necessary to convolve the excitation signal values of the last (K - 1) sample points in the previous speech frame of the target speech frame and the excitation signal value of the first sample point in the target speech frame with the K-order filter to obtain the target signal value corresponding to the second sample point in the target speech frame.
[0119] In summary, step 520 also requires the participation of the excitation signal values corresponding to the historical speech frames of the target speech frame. The number of sample points in the required historical speech frame is related to the order of the glottal filter, that is, if the glottal filter is of K order, the participation of the excitation signal values of the last K sample points in the previous speech frame of the target speech frame is required.
[0120] Step 530, amplify the first speech signal according to the gain corresponding to the target speech frame to obtain the enhanced speech signal corresponding to the target speech frame.
[0121] Through the above steps 510 - 530, speech synthesis is realized according to the glottal parameters, excitation signal and gain predicted for the target speech frame, and the enhanced speech signal of the target speech frame is obtained.
[0122] In the solution of this application, based on the frequency-domain representation of the target speech frame, the glottal parameters and excitation signal for reconstructing the original speech signal in the target speech frame are predicted, and the gain of the historical speech frames of the target speech frame is predicted for reconstructing the original speech signal in the target speech frame. Then, according to the predicted glottal parameters, the corresponding excitation signal, and the corresponding gain of the target speech frame, speech synthesis is performed, which is equivalent to reconstructing the original speech signal in the target speech frame. The signal obtained through the synthesis process is the enhanced speech signal corresponding to the target speech frame, achieving the enhancement of the speech frame.
[0123] In the related art, speech enhancement is performed through methods such as spectral estimation and spectral regression prediction. The speech enhancement method based on spectral estimation believes that a mixed speech contains a speech part and a noise part. Therefore, the noise can be estimated through a statistical model, etc. The spectrum corresponding to the mixed speech is subtracted from the spectrum corresponding to the noise, and the remaining is the speech spectrum. Thus, a clean speech signal is restored from the spectrum obtained by subtracting the spectrum corresponding to the noise from the spectrum corresponding to the mixed speech. The speech enhancement method based on spectral regression prediction predicts the masking threshold corresponding to the speech frame through a neural network. This masking threshold reflects the proportion of the speech component and the noise component in each frequency point of the speech frame. Then, the gain control of the mixed signal spectrum is performed according to this masking threshold to obtain the enhanced spectrum.
[0124] The above speech enhancement methods based on spectral estimation and spectral regression prediction are based on the estimation of the posterior probability of the noise spectrum, and there may be inaccurate estimation of the noise. For example, transient noises such as keyboard typing occur instantaneously, and the estimated noise spectrum is very inaccurate, resulting in poor noise suppression effect. In the case of inaccurate prediction of the noise spectrum, if the original mixed speech signal is processed according to the estimated noise spectrum, it may lead to speech distortion in the mixed speech signal or poor noise suppression effect. Therefore, in this case, a compromise needs to be made between speech fidelity and noise suppression.
[0125] In the solution of this application, since the glottal parameters are strongly correlated with the glottal characteristics in the physical process of sound generation, the predicted glottal parameters effectively guarantee the speech structure of the original speech signal in the target speech frame. Therefore, synthesizing based on the predicted glottal parameters, excitation signal, and gain to obtain the enhanced speech signal of the target speech frame can effectively avoid the reduction of the original speech signal in the target speech frame and effectively protect the speech structure. At the same time, after predicting the glottal parameters, excitation signal, and gain corresponding to the target speech frame, since the original noisy speech is no longer processed, there is no need to make a compromise between speech fidelity and noise suppression.
[0126] In some embodiments of the present application, before step 410, the method further includes: obtaining the time-domain signal of the target speech frame; performing time-frequency transformation on the time-domain signal of the target speech frame to obtain the frequency-domain representation of the target speech frame.
[0127] The time-frequency transformation can be a short-time Fourier transform (STFT). The frequency-domain representation can be an amplitude spectrum, a complex frequency spectrum, etc., which are not specifically limited herein.
[0128] In the short-time Fourier transform, the operation of windowing and overlapping is adopted to eliminate the non-smoothness between frames. Figure 6 Figure is a schematic diagram showing windowing and overlapping in a specifically illustrated short-time Fourier transform. In Figure 6 a 50% windowing and overlapping operation is adopted. If the short-time Fourier transform is applied to 640 sample points, the number of overlapping samples (hop-size) of the window function is 320. The window function used for windowing can be a Hanning window. Of course, other window functions can also be used, which are not specifically limited herein.
[0129] In other embodiments, a non-50% windowing and overlapping operation can also be adopted. For example, if the short-time Fourier transform is applied to 512 sample points, in this case, if a speech frame includes 320 sample points, only 192 sample points of the previous speech frame need to be overlapped.
[0130] In some embodiments of the present application, the obtaining the time-domain signal of the target speech frame includes: obtaining a second speech signal, where the second speech signal is a collected speech signal or a speech signal obtained by decoding an encoded speech; performing frame segmentation on the second speech signal to obtain the time-domain signal of the target speech frame.
[0131] In some instances, the second speech signal can be frame-segmented according to a set frame length, and the frame length can be set according to actual needs. For example, the frame length can be set to 20 ms.
[0132] As described above, the solution of the present application can be applied to the transmitting end for speech enhancement, or can also be applied to the receiving end for speech enhancement.
[0133] When the solution of the present application is applied to the transmitting end, the second speech signal is the speech signal collected by the transmitting end. Then, the second speech signal is frame-segmented to obtain multiple speech frames. After frame-segmenting to obtain speech frames, each speech frame can be used as the target speech frame and the target speech frame can be enhanced according to the above process of steps 410-440. Further, after obtaining the enhanced speech signal corresponding to the target speech frame, the enhanced speech signal can also be encoded for transmission based on the obtained encoded speech.
[0134] In one embodiment, since the directly collected voice signal is an analog signal, in order to facilitate signal processing, before frame segmentation, the signal further needs to be digitized. The collected voice signal can be sampled at a set sampling rate, and the set sampling rate can be 16000Hz, 8000Hz, 32000Hz, 48000Hz, etc., which can be specifically set according to actual needs.
[0135] When the solution of the present application is applied to the receiving end, the second voice signal is the voice signal obtained by decoding the received encoded voice. After obtaining voice frames by performing frame segmentation on the second voice signal, the voice frames are used as target voice frames and enhanced according to the process of steps 410-440 above to obtain the enhanced voice signal of the target voice frames. Further, the enhanced voice signal corresponding to the target voice frames can also be played. Since the obtained enhanced voice signal has less noise and higher quality compared to the signal before the enhancement of the target voice frames, for users, the auditory experience is better.
[0136] Next, the solution of the present application will be further described in conjunction with specific embodiments:
[0137] Figure 7 is a flowchart of a voice enhancement method shown according to a specific embodiment. Assume that the nth frame of voice frame is used as the target voice frame, and the time-domain signal of the nth frame of voice frame is s(n). As Figure 7 shown, perform time-frequency transformation on the nth frame of voice frame according to step 710 to obtain the frequency-domain representation S(n) of the nth frame of voice frame, where S(n) can be an amplitude spectrum or a complex frequency spectrum, which is not specifically limited herein.
[0138] After obtaining the frequency-domain representation S(n) of the nth frame of voice frame, the glottal parameters corresponding to the nth frame of voice frame can be predicted through step 720, and the excitation signal corresponding to the target voice frame can be obtained through steps 730 and 740.
[0139] In step 720, either only the frequency-domain representation S(n) of the nth frame of voice frame can be used as the input of the first neural network, or the glottal parameters P_pre(n) corresponding to the historical voice frames of the target voice frame and the frequency-domain representation S(n) of the nth frame of voice frame can be used as the input of the first neural network. The first neural network can perform glottal parameter prediction based on the input information to obtain the glottal parameters ar(n) corresponding to the nth frame of voice frame.
[0140] In step 730, the frequency-domain representation S(n) of the nth speech frame is used as the input of the third neural network. This third neural network predicts the excitation signal based on the input information and outputs the frequency-domain representation R(n) of the excitation signal corresponding to the nth speech frame. On this basis, frequency-time transformation can be performed through step 740 to obtain the time-domain signal r(n) of the excitation signal corresponding to the nth speech frame.
[0141] The gain corresponding to the nth speech frame is obtained through step 750. In step 750, the gain G_pre(n) of the historical speech frame of the nth speech frame is input into the input of the second neural network, and the second neural network correspondingly performs gain prediction to obtain the gain G_(n) corresponding to the nth speech frame.
[0142] After obtaining the glottal parameter ar(n), the corresponding excitation signal r(n), and the corresponding gain G_(n) of the nth speech frame, synthetic filtering is performed according to step 760 based on these three parameters to obtain the enhanced speech signal s_e(n) corresponding to the nth speech frame. Specifically, speech synthesis can be performed according to the principle of linear prediction analysis. It is worth mentioning that in the process of speech synthesis according to the principle of linear prediction analysis, the information of historical speech frames needs to be utilized. Specifically, in the filtering process of the excitation signal by the glottal filter, that is, for the sample point at time t, it is obtained by convolving the excitation signal values of the previous p historical sample points with the 16th-order glottal filter. If the glottal filter is a 16th-order digital filter, then in the process of synthesizing the nth speech frame, the information of the last p sample points in the (n - 1)th frame also needs to be utilized.
[0143] The above steps 720, 730, and 750 are further described below in conjunction with specific embodiments. Assume that the sampling frequency Fs of the speech signal to be processed is 16000 Hz and the frame length is 20 ms, then each speech frame includes 320 sample points; assume that the short-time Fourier transform in this method is performed in a manner of 640 sample points and 320 overlapping sample points. And further assume that the glottal parameter is the line spectral frequency coefficient, that is, the glottal parameter ar(n) corresponding to the nth speech frame is LSF(n), and the glottal filter is set as a 16th-order filter.
[0144] Figure 8 is a schematic diagram of the first neural network shown according to a specific embodiment, as Figure 8As shown, the first neural network includes one LSTM (Long-Short Term Memory) layer and three cascaded FC (Full Connected) layers. Among them, the LSTM layer has 1 hidden layer, which includes 256 units. The input of the LSTM layer is the frequency-domain representation S(n) of the nth frame of speech. In this embodiment, the input of the LSTM layer is 321-dimensional STFT coefficients. In the three cascaded FC layers, activation functions σ() are provided in the first two FC layers. The set activation functions are used to increase the non-linear expression ability of the first neural network. No activation function is provided in the last FC layer, and this last FC layer is used as a classifier for classification output. As Figure 8 shown, from bottom to top, the three FC layers include 512, 512, and 16 units respectively. The output of the last FC layer is the 16-dimensional line spectral frequency coefficients LSF(n) corresponding to the nth frame of speech, that is, the 16th-order line spectral frequency coefficients.
[0145] Figure 9 is a schematic diagram of the input and output of the first neural network shown according to another embodiment. Among them, Figure 9 the structure of the first neural network in Figure 8 is the same as that in Figure 8 . Compared with Figure 9 , the input of the first neural network in Figure 9 also includes the line spectral frequency coefficients LSF(n - 1) of the previous speech frame (i.e., the (n - 1)th frame) of the nth frame of speech. As
[0146] Figure 10 is a schematic diagram of the second neural network shown according to a specific embodiment. As Figure 10 shown, the second neural network includes one LSTM layer and one FC layer. Among them, the LSTM layer has 1 hidden layer, which includes 128 units; the input of the FC layer is a 512-dimensional vector, and the output is a 1-dimensional gain. In a specific embodiment, the gain G_pre(n) of the historical speech frame of the nth frame of speech can be defined as the gains corresponding to the first 4 speech frames of the nth frame of speech, that is:
[0147] G_pre(n) = {G(n - 1), G(n - 2), G(n - 3), G(n - 4)};
[0148] Of course, the number of historical speech frames selected for gain prediction is not limited to the examples above and can be specifically selected according to actual needs.
[0149] In the structures of the first neural network and the second neural network shown above, the network presents an M-to-N mapping relationship (N << M), that is, the dimension of the input information of the neural network is M, and the dimension of the output information is N, greatly streamlining the structures of the first neural network and the second neural network and reducing the complexity of the model.
[0150] Figure 11 It is a schematic diagram of a third neural network shown according to a specific embodiment, as Figure 11 shown. The third neural network includes one LSTM layer and three FC layers. Among them, the LSTM layer has one hidden layer, including 256 units. The input of the LSTM is the 321-dimensional STFT coefficient S(n) corresponding to the nth frame of the speech frame. The number of units included in the three FC layers are 512, 512, and 321 respectively. The last FC layer outputs the frequency-domain representation R(n) of the excitation signal corresponding to the nth frame of the speech frame with 321 dimensions. From bottom to top, activation functions are provided in the first two FC layers of the three FC layers to enhance the non-linear expression ability of the model, and no activation function is provided in the last FC layer for classification output.
[0151] It is worth mentioning that Figures 8 - 11 the structures of the first neural network, the second neural network, and the third neural network shown are only exemplary examples. In other embodiments, corresponding network structures can also be set in the open source platform of deep learning and trained accordingly.
[0152] The following introduces the device embodiments of the present application, which can be used to execute the methods in the above embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the above method embodiments of the present application.
[0153] Figure 12 It is a block diagram of a voice enhancement device shown according to an embodiment, as Figure 12 shown. The voice enhancement device includes:
[0154] A glottal parameter prediction module 1210, configured to predict glottal parameters according to the frequency-domain representation of the target speech frame to obtain the glottal parameters corresponding to the target speech frame.
[0155] A gain prediction module 1220, configured to perform gain prediction on the target speech frame according to the gain corresponding to the historical speech frames of the target speech frame to obtain the gain corresponding to the target speech frame.
[0156] An excitation signal prediction module 1230, configured to predict an excitation signal according to the frequency-domain representation of the target speech frame, so as to obtain the excitation signal corresponding to the target speech frame.
[0157] A synthesis module 1240, configured to perform synthesis processing on the glottal parameters corresponding to the target speech frame, the gain corresponding to the target speech frame, and the excitation signal corresponding to the target speech frame, so as to obtain the enhanced speech signal corresponding to the target speech frame.
[0158] In some embodiments of the present application, the synthesis module 1240 includes: a glottal filter construction unit, configured to construct a glottal filter according to the glottal parameters corresponding to the target speech frame; a filtering unit, configured to filter the excitation signal corresponding to the target speech frame through the glottal filter to obtain a first speech signal; an amplification unit, configured to amplify the first speech signal according to the gain corresponding to the target speech frame to obtain the enhanced speech signal corresponding to the target speech frame.
[0159] In some embodiments of the present application, the target speech frame includes a plurality of sample points; the glottal filter is a K-order filter, where K is a positive integer; the excitation signal includes the excitation signal values corresponding to the plurality of sample points in the target speech frame; the filtering unit includes: a convolution unit, configured to convolve the excitation signal values corresponding to the first K sample points of each sample point in the target speech frame with the K-order filter to obtain the target signal value of each sample point in the target speech frame; a combination unit, configured to combine the target signal values corresponding to all the sample points in the target speech frame in chronological order to obtain the first speech signal. In some embodiments of the present application, the glottal filter is a K-order filter, and the glottal parameters include K-order line spectrum frequency parameters or K-order linear prediction coefficients.
[0160] In some embodiments of the present application, the glottal parameter prediction module 1210 includes: a first input unit, configured to input the frequency-domain representation of the target speech frame into a first neural network, and the first neural network is trained according to the frequency-domain representation of the sample speech frame and the glottal parameters corresponding to the sample speech frame; a first output unit, configured to output the glottal parameters corresponding to the target speech frame by the first neural network according to the frequency-domain representation of the target speech frame.
[0161] In some embodiments of the present application, the glottal parameter prediction module 1210 is further configured to: use the glottal parameters corresponding to the historical speech frames of the target speech frame as a reference, and perform glottal parameter prediction according to the frequency-domain representation of the target speech frame to obtain the glottal parameters corresponding to the target speech frame.
[0162] In some embodiments of the present application, the glottal parameter prediction module 1210 includes: a second input unit, configured to input the frequency-domain representation of the target speech frame and the glottal parameters corresponding to the historical speech frames of the target speech frame into a first neural network, where the first neural network is trained using the frequency-domain representation of the sample speech frames, the glottal parameters corresponding to the sample speech frames, and the glottal parameters corresponding to the historical speech frames of the sample speech frames; a second output unit, configured to predict, by the first neural network based on the frequency-domain representation of the target speech frame and the glottal parameters corresponding to the historical speech frames of the target speech frame, and output the glottal parameters corresponding to the target speech frame.
[0163] In some embodiments of the present application, the gain prediction module 1220 includes: a third input unit, configured to input the gain corresponding to the historical speech frames of the target speech frame into a second neural network; the second neural network is trained based on the gain corresponding to the sample speech frames and the gain corresponding to the historical speech frames of the sample speech frames; a third output unit, configured to output the target gain by the second neural network based on the gain corresponding to the historical speech frames of the target speech frame.
[0164] In some embodiments of the present application, the excitation signal prediction module 1230 includes: a fourth input unit, configured to input the frequency-domain representation of the target speech frame into a third neural network; the third neural network is trained based on the frequency-domain representation of the sample speech frames and the frequency-domain representation of the excitation signals corresponding to the sample speech frames; a fourth output unit, configured to output the frequency-domain representation of the excitation signal corresponding to the target speech frame by the third neural network based on the frequency-domain representation of the target speech frame.
[0165] In some embodiments of the present application, the speech enhancement device further includes: an acquisition module, configured to acquire the time-domain signal of the target speech frame; a time-frequency transformation module, configured to perform time-frequency transformation on the time-domain signal of the target speech frame to obtain the frequency-domain representation of the target speech frame.
[0166] In some embodiments of the present application, the acquisition module is further configured to: acquire a second speech signal, where the second speech signal is a collected speech signal or a speech signal obtained by decoding an encoded speech; perform frame segmentation on the second speech signal to obtain the time-domain signal of the target speech frame.
[0167] In some embodiments of the present application, the speech enhancement device further includes: a processing module, configured to play or encode and transmit the enhanced speech signal corresponding to the target speech frame.
[0168] Figure 13 The structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown.
[0169] It should be noted that Figure 13 the computer system 1300 of the illustrated electronic device is only an example and should not impose any limitation on the functions and the scope of use of the embodiments of the present application.
[0170] As Figure 13 shown, the computer system 1300 includes a central processing unit (CPU) 1301, which can perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) 1302 or the programs loaded from the storage section 1308 into the random access memory (RAM) 1303, such as executing the methods in the above embodiments. In the RAM 1303, various programs and data required for system operations are also stored. The CPU 1301, the ROM 1302, and the RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.
[0171] The following components are connected to the I / O interface 1305: an input section 1306 including a keyboard, a mouse, etc.; an output section 1307 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1308 including a hard disk, etc.; and a communication section 1309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1309 performs communication processing via a network such as the Internet. A drive 1310 is also connected to the I / O interface 1305 as required. A removable medium 1311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1310 as required so that the computer program read from it can be installed into the storage section 1308 as required.
[0172] Specifically, according to the embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 1309, and / or installed from the removable medium 1311. When the computer program is executed by the central processing unit (CPU) 1301, various functions defined in the system of the present application are executed.
[0173] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0174] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0175] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation to the unit itself in some cases.
[0176] As another aspect, the present application also provides a computer-readable storage medium, which can be included in the electronic device described in the above embodiments; or can exist alone without being assembled into the electronic device. The above computer-readable storage medium carries computer-readable instructions, and when the computer-readable storage instructions are executed by a processor, the methods in any of the above embodiments are implemented.
[0177] According to one aspect of the present application, an electronic device is also provided, which includes: a processor; a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the methods in any of the above embodiments are implemented.
[0178] According to one aspect of the embodiments of the present application, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods in any of the above embodiments.
[0179] It should be noted that although several modules or units of devices for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0180] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described here can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the methods according to the embodiments of the present application.
[0181] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known knowledge or conventional technical means in the technical field not disclosed in the present application.
[0182] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A voice enhancement method, characterized in that, Including: Predicting glottal parameters based on the frequency-domain representation of the target speech frame to obtain the glottal parameters corresponding to the target speech frame; the target speech frame refers to the speech frame to be enhanced currently, and the glottal parameter prediction is to predict the glottal parameters used to reconstruct the original speech signal in the target speech frame through a neural network; Predicting the gain of the target speech frame based on the gain corresponding to the historical speech frames of the target speech frame to obtain the gain corresponding to the target speech frame; the gain corresponding to the historical speech frames refers to the gain used to reconstruct the original speech signal in the historical speech frames, and the original speech signal is a speech signal without noise; Predicting an excitation signal based on the frequency-domain representation of the target speech frame to obtain the excitation signal corresponding to the target speech frame; Performing synthesis processing on the glottal parameters corresponding to the target speech frame, the gain corresponding to the target speech frame, and the excitation signal corresponding to the target speech frame to obtain the enhanced speech signal corresponding to the target speech frame.
2. The method according to claim 1, wherein The performing synthesis processing on the glottal parameters corresponding to the target speech frame, the gain corresponding to the target speech frame, and the excitation signal corresponding to the target speech frame to obtain the enhanced speech signal corresponding to the target speech frame includes: Constructing a glottal filter according to the glottal parameters corresponding to the target speech frame; Filtering the excitation signal corresponding to the target speech frame through the glottal filter to obtain a first speech signal; Amplifying the first speech signal according to the gain corresponding to the target speech frame to obtain the enhanced speech signal corresponding to the target speech frame.
3. The method according to claim 2, wherein The target speech frame includes multiple sample points; the glottal filter is a K-order filter, where K is a positive integer; the excitation signal includes the excitation signal values corresponding to the multiple sample points in the target speech frame respectively; The filtering the excitation signal corresponding to the target speech frame through the glottal filter to obtain a first speech signal includes: Convolving the excitation signal values corresponding to the first K sample points of each sample point in the target speech frame with the K-order filter to obtain the target signal value of each sample point in the target speech frame; Combining the target signal values corresponding to all the sample points in the target speech frame in chronological order to obtain the first speech signal.
4. The method according to claim 2, wherein The glottal filter is a K-order filter, and the glottal parameters include K-order line spectral frequency parameters or K-order linear prediction coefficients; K is a positive integer.
5. The method according to claim 1, wherein The predicting glottal parameters based on the frequency-domain representation of the target speech frame to obtain the glottal parameters corresponding to the target speech frame includes: Inputting the frequency-domain representation of the target speech frame into a first neural network, and the first neural network is trained according to the frequency-domain representation of the sample speech frames and the glottal parameters corresponding to the sample speech frames; Outputting, by the first neural network, the glottal parameters corresponding to the target speech frame according to the frequency-domain representation of the target speech frame.
6. The method according to claim 1, wherein The predicting glottal parameters based on the frequency-domain representation of the target speech frame to obtain the glottal parameters corresponding to the target speech frame includes: Using the glottal parameters corresponding to the historical speech frames of the target speech frame as a reference, perform glottal parameter prediction based on the frequency-domain representation of the target speech frame to obtain the glottal parameters corresponding to the target speech frame.
7. The method according to claim 6, characterized in that, The step of using the glottal parameters corresponding to the historical speech frames of the target speech frame as a reference, performing glottal parameter prediction based on the frequency-domain representation of the target speech frame, and obtaining the glottal parameters corresponding to the target speech frame includes: Input the frequency-domain representation of the target speech frame and the glottal parameters corresponding to the historical speech frames of the target speech frame into a first neural network, where the first neural network is trained by using the frequency-domain representation of sample speech frames, the glottal parameters corresponding to the sample speech frames, and the glottal parameters corresponding to the historical speech frames of the sample speech frames; The first neural network makes a prediction based on the frequency-domain representation of the target speech frame and the glottal parameters corresponding to the historical speech frames of the target speech frame, and outputs the glottal parameters corresponding to the target speech frame.
8. The method according to claim 1, wherein The step of performing gain prediction on the target speech frame according to the gain corresponding to the historical speech frames of the target speech frame to obtain the gain corresponding to the target speech frame includes: Input the gain corresponding to the historical speech frames of the target speech frame into a second neural network; the second neural network is trained according to the gain corresponding to the sample speech frames and the gain corresponding to the historical speech frames of the sample speech frames; The second neural network outputs the gain corresponding to the target speech frame according to the gain corresponding to the historical speech frames of the target speech frame.
9. The method according to claim 1, characterized in that The step of performing excitation signal prediction based on the frequency-domain representation of the target speech frame to obtain the excitation signal corresponding to the target speech frame includes: Input the frequency-domain representation of the target speech frame into a third neural network; the third neural network is trained according to the frequency-domain representation of sample speech frames and the frequency-domain representation of the excitation signals corresponding to the sample speech frames; The third neural network outputs the frequency-domain representation of the excitation signal corresponding to the target speech frame according to the frequency-domain representation of the target speech frame.
10. The method according to claim 1, wherein Before performing glottal parameter prediction based on the frequency-domain representation of the target speech frame to obtain the glottal parameters corresponding to the target speech frame, the method further includes: Obtain the time-domain signal of the target speech frame; Perform time-frequency transformation on the time-domain signal of the target speech frame to obtain the frequency-domain representation of the target speech frame.
11. The method according to claim 10, wherein The step of obtaining the time-domain signal of the target speech frame includes: Obtain a second speech signal, where the second speech signal is a collected speech signal or a speech signal obtained by decoding an encoded speech; Perform frame segmentation on the second speech signal to obtain the time-domain signal of the target speech frame.
12. The method according to claim 1, wherein After performing synthesis processing on the glottal parameters corresponding to the target speech frame, the gain corresponding to the target speech frame, and the excitation signal corresponding to the target speech frame to obtain the enhanced speech signal corresponding to the target speech frame, the method further includes: Play or encode and transmit the enhanced speech signal corresponding to the target speech frame.
13. A voice enhancement device, characterized in that, including: A glottal parameter prediction module, configured to predict glottal parameters based on the frequency-domain representation of a target speech frame, so as to obtain the glottal parameters corresponding to the target speech frame; the target speech frame refers to the speech frame to be enhanced currently, and the glottal parameter prediction performed refers to predicting, through a neural network, the glottal parameters for reconstructing the original speech signal in the target speech frame; A gain prediction module, configured to predict the gain of the target speech frame based on the gain corresponding to the historical speech frames of the target speech frame, so as to obtain the gain corresponding to the target speech frame; the gain corresponding to the historical speech frames refers to the gain for reconstructing the original speech signal in the historical speech frames, and the original speech signal is a speech signal without noise; An excitation signal prediction module, configured to predict an excitation signal based on the frequency-domain representation of the target speech frame, so as to obtain the excitation signal corresponding to the target speech frame; A synthesis module, configured to perform synthesis processing on the glottal parameters corresponding to the target speech frame, the gain corresponding to the target speech frame, and the excitation signal corresponding to the target speech frame, so as to obtain the enhanced speech signal corresponding to the target speech frame.
14. An electronic device, characterized in that, Comprising: A processor; A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method according to any one of claims 1-12 is implemented.
15. A computer-readable storage medium, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the method according to any one of claims 1-12 is implemented.
16. A computer program product, characterized in that, The computer program product includes a computer program, and the computer program is adapted to be loaded and executed by a processor to perform the method according to any one of claims 1-12.
Citation Information
Patent Citations
Voice processing method, device and equipment and storage medium
CN111554309A
Voice processing method, device and equipment and storage medium
CN111554322A
Voice processing method, device and equipment and storage medium
CN111554323A