Speech Enhancement Method, Apparatus, Device and Storage Medium
By pre-enhancing, decomposing and synthesizing the voice signal, the problem of noise influence in voice communication is solved, and efficient enhancement and quality improvement of voice signals is achieved.
Patent Information
- Application Number
- CN202110182834.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-08
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-02-08
AI Technical Summary
In voice communication, noise is mixed in voice signals, resulting in poor communication quality and affecting the user's auditory experience.
By pre-enhancing the target speech frame, the first amplitude spectrum is obtained, and then the speech decomposition is performed based on the amplitude spectrum to obtain the glottal parameters, gain and excitation signals, and finally synthesis processing is performed to obtain the enhanced speech signal.
Effectively remove noise, improve the quality of voice signals, and ensure enhanced accuracy and fidelity of voice signals.
Smart Images

Figure CN113571081B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech processing technology, and in particular to a speech enhancement method, apparatus, device and storage medium. Background Art
[0002] Due to the convenience and timeliness of voice communication, voice communication is increasingly used, for example, voice signals are transmitted between participants in a cloud conference. In voice communication, the voice signal may be mixed with noise, which will lead to poor communication quality and greatly affect the user's auditory experience. Therefore, how to enhance the voice to remove the noise is a technical problem that needs to be solved in the prior art. Summary of the invention
[0003] The embodiments of the present application provide a speech enhancement method, apparatus, device and storage medium to achieve speech enhancement.
[0004] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.
[0005] According to one aspect of an embodiment of the present application, a speech enhancement method is provided, comprising: performing pre-enhancement processing on a target speech frame according to an amplitude spectrum corresponding to the target speech frame to obtain a first amplitude spectrum; performing speech decomposition on the target speech frame according to the first amplitude spectrum to obtain glottal parameters, gain and excitation signal corresponding to the target speech frame; performing synthesis processing according to the glottal parameters, the gain and the excitation signal to obtain an enhanced speech signal corresponding to the target speech frame.
[0006] According to one aspect of an embodiment of the present application, a speech enhancement device is provided, comprising: a pre-enhancement module, used to perform pre-enhancement processing on a target speech frame according to an amplitude spectrum corresponding to the target speech frame, so as to obtain a first amplitude spectrum; a speech decomposition module, used to perform speech decomposition on the target speech frame according to the first amplitude spectrum, so as to obtain glottal parameters, gain and excitation signal corresponding to the target speech frame; and a synthesis module, used to perform synthesis processing according to the glottal parameters, the gain and the excitation signal, so as to obtain an enhanced speech signal corresponding to the target speech frame.
[0007] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor; a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the speech enhancement method as described above is implemented.
[0008] According to one aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, the above-mentioned voice enhancement method is implemented.
[0009] In the solution of the present application, on the basis of pre-enhancing the target speech frame to obtain the first amplitude spectrum, the target speech frame is decomposed and synthesized based on the first amplitude spectrum, realizing the enhancement of the target speech frame in two stages, which can effectively ensure the voice enhancement effect. Moreover, compared with the amplitude spectrum before the pre-enhancement of the target speech frame, the first amplitude spectrum has less noise information. And in the process of speech decomposition, noise will affect the accuracy of speech decomposition. Therefore, using the first amplitude spectrum as the basis for speech decomposition can reduce the difficulty of speech decomposition, improve the accuracy of the glottal parameters, excitation signal and gain obtained by speech decomposition, and thus ensure the accuracy of the enhanced speech signal obtained subsequently.
[0010] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0012] Figure 1 is a schematic diagram of a voice communication link in a VoIP (Voice over Internet Protocol) system shown according to a specific embodiment.
[0013] Figure 2 shows a schematic diagram of a digital model for generating a voice signal.
[0014] Figure 3 shows a schematic diagram of decomposing an original voice signal into an excitation signal and the frequency response of a glottal filter.
[0015] Figure 4 shows a flowchart of a voice enhancement method shown according to an embodiment of the present application.
[0016] Figure 5 is a flowchart of a voice enhancement method shown according to an embodiment of the present application.
[0017] Figure 6 is a schematic diagram of the structure of a first neural network shown according to a specific embodiment.
[0018] Figure 7 is a flowchart of step 410 shown according to an embodiment of the present application.
[0019] Figure 8 is a schematic structural diagram of a second neural network shown according to a specific embodiment.
[0020] Figure 9 is a flowchart of step 420 shown according to an embodiment of the present application.
[0021] Figure 10 is a flowchart of step 430 shown according to an embodiment of the present application.
[0022] Figure 11 is a schematic diagram of a third neural network shown according to a specific embodiment
[0023] Figure 12 is a schematic diagram of the input and output of a third neural network shown according to another embodiment.
[0024] Figure 13 is a schematic diagram of a fourth neural network shown according to a specific embodiment.
[0025] Figure 14 is a schematic diagram of a fifth neural network shown according to a specific embodiment.
[0026] Figure 15 is a schematic diagram of performing short-time Fourier transform on speech frames in a windowed overlapping manner shown according to an embodiment of the present application.
[0027] Figure 16 is a block diagram of a speech enhancement device shown according to an embodiment.
[0028] Figure 17 shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners
[0029] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0030] In addition, the described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0031] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0032] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the content and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps may be decomposed, while some operations / steps may be combined or partially combined, so the actual execution order may change according to the actual situation.
[0033] It should be noted that: "a plurality" as mentioned herein means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0034] Noise in the voice signal will greatly reduce the voice quality and affect the user's auditory experience. Therefore, in order to improve the quality of the voice signal, it is necessary to perform enhancement processing on the voice signal to remove the noise as much as possible and retain the original voice signal in the signal (i.e., the pure signal without noise). In order to implement voice enhancement processing, the solution of the present application is proposed.
[0035] The solution of the present application can be applied to the application scenarios of voice calls, such as voice communication through instant messaging applications and voice calls in game applications. Specifically, voice enhancement can be performed according to this solution at the voice sending end, the voice receiving end, or the server providing voice communication services.
[0036] Cloud conferencing is an important part of remote work. In cloud conferencing, after the voice collection device of the participants in the cloud conference collects the voice signal of the speaker, it needs to send the collected voice signal to other conference participants. This process involves the transmission and playback of voice signals among multiple participants. If the noise signals mixed in the voice signals are not processed, it will greatly affect the auditory experience of the conference participants. In this scenario, the solution of this application can be applied to enhance the voice signals in the cloud conference, so that the voice signals heard by the conference participants are enhanced voice signals, improving the quality of the voice signals.
[0037] Cloud conferencing is an efficient, convenient, and low-cost conference form based on cloud computing technology. Users only need to perform simple and easy operations through the Internet interface to quickly and efficiently synchronously share voices, data files, and videos with teams and customers around the world. The complex technologies such as data transmission and processing in the conference are operated by the cloud conferencing service provider to assist users.
[0038] Currently, domestic cloud conferencing mainly focuses on service contents in the form of SaaS (Software as a Service). It includes service forms such as telephone, network, and video. The video conference based on cloud computing is called cloud conferencing. In the era of cloud conferencing, the transmission, processing, and storage of data are all processed by the computer resources of the video conference provider. Users no longer need to purchase expensive hardware and install cumbersome software at all. They only need to open the client and enter the corresponding interface to conduct efficient remote conferences.
[0039] The cloud conferencing system supports multi-server dynamic cluster deployment and provides multiple high-performance servers, greatly improving the stability, security, and availability of the conference. In recent years, video conferencing has been welcomed by many users because it can significantly improve communication efficiency, continuously reduce communication costs, and bring about an upgrade in internal management level. It has been widely applied in various fields such as transportation, logistics, finance, operators, education, and enterprises.
[0040] Figure 1 It is a schematic diagram of the voice communication link in a VoIP (Voice over Internet Protocol) system shown according to a specific embodiment. As Figure 1 shown, based on the network connection between the sending end 110 and the receiving end 120, the sending end 110 and the receiving end 120 can perform voice transmission.
[0041] As Figure 1As shown in the figure, the transmitting end 110 includes an acquisition module 111, a pre-enhancement processing module 112, and an encoding module 113. Among them, the acquisition module 111 is used to acquire voice signals, and it can convert the acquired acoustic signals into digital signals; the pre-enhancement processing module 112 is used to enhance the acquired voice signals to remove the noise in the acquired voice signals and improve the quality of the voice signals. The encoding module 113 is used to encode the enhanced voice signals to improve the anti-interference ability of the voice signals during transmission. The pre-enhancement processing module 112 can perform voice enhancement according to the method of this application. After enhancing the voice, encoding compression and transmission are performed, so that the signals received by the receiving end are no longer affected by noise.
[0042] The receiving end 120 includes a decoding module 121, a post-enhancement module 122, and a playback module 123. The decoding module 121 is used to decode the received encoded voice to obtain a decoded signal; the post-enhancement module 122 is used to perform enhancement processing on the decoded voice signal; the playback module 123 is used to play the enhanced voice signal. The post-enhancement module 122 can also perform voice enhancement according to the method of this application. In some embodiments, the receiving end 120 may further include a sound effect adjustment module, and this sound effect adjustment module is used to adjust the sound effect of the enhanced voice signal.
[0043] In a specific embodiment, voice enhancement may be performed only at the receiving end 120 or only at the transmitting end 110 according to the method of this application. Of course, voice enhancement may also be performed at both the transmitting end 110 and the receiving end 120 according to the method of this application.
[0044] In some application scenarios, in addition to supporting VoIP communication, the terminal device in the VoIP system can also support other third-party protocols, such as traditional PSTN (Public Switched Telephone Network) circuit-switched telephones. However, traditional PSTN services cannot perform voice enhancement. In this scenario, voice enhancement can be performed at the terminal acting as the receiving end according to the method of this application.
[0045] In order to specifically describe the solution of this application, it is necessary to introduce the generation of voice signals. Voice signals are generated by the physiological movements of the human vocal organs under the control of the brain, that is: at the trachea, an impact signal similar to noise with a certain energy is generated (equivalent to an excitation signal); the impact signal impacts the human vocal cords (the vocal cords are equivalent to a glottal filter), generating a quasi-periodic opening and closing; after being amplified by the oral cavity, a sound is emitted (the output voice signal).
[0046] Figure 2 The figure shows a schematic diagram of the digital model of voice signal generation, and the generation process of voice signals can be described through this digital model. AsFigure 2 As shown, after the excitation signal impacts the glottal filter and then undergoes gain control, a speech signal is output. Among them, the glottal filter is defined by glottal parameters. This process can be represented by the following formula:
[0047] x(n) = G·r(n)·ar(n); (Formula 1)
[0048] Among them, x(n) represents the input speech signal; G represents the gain, which can also be called the linear prediction gain; r(n) represents the excitation signal; ar(n) represents the glottal filter.
[0049] Figure 3 The figure shows a schematic diagram of decomposing the excitation signal and the frequency response of the glottal filter from an original speech signal. Figure 3 a shows a schematic diagram of the frequency response of the original speech signal. Figure 3 b shows a schematic diagram of the frequency response of the glottal filter decomposed from the original speech signal. Figure 3 c shows a schematic diagram of the frequency response of the excitation signal decomposed from the original speech signal. As Figure 3 shown, the undulating part in the frequency response diagram of the original speech signal corresponds to the peak position in the frequency response diagram of the glottal filter. The excitation signal is equivalent to the residual signal after performing LP (Linear Prediction) analysis on the original speech signal. Therefore, its corresponding frequency response is relatively flat.
[0050] It can be seen from the above that an excitation signal, a glottal filter, and a gain can be decomposed from an original speech signal (i.e., a speech signal without noise). The decomposed excitation signal, glottal filter, and gain can be used to represent the original speech signal. Among them, the glottal filter can be expressed by glottal parameters. Conversely, if the excitation signal corresponding to an original speech signal, the glottal parameters for determining the glottal filter, and the gain are known, then the original speech signal can be reconstructed based on the corresponding excitation signal, glottal filter, and gain.
[0051] The solution of this application is exactly based on this principle. The original speech signal in the speech frame is reconstructed according to the glottal parameters, excitation signal, and gain corresponding to the speech frame, so as to achieve speech enhancement.
[0052] The following elaborates in detail on the implementation details of the technical solution of the embodiments of this application:
[0053] Figure 4 The figure shows a flowchart of a speech enhancement method according to an embodiment of this application. This method can be executed by a computer device with processing capabilities, such as a terminal, a server, etc., which is not specifically limited here. Refer to Figure 4As shown, the method at least includes steps 410 to 440, which are introduced in detail as follows:
[0054] Step 410, perform pre-enhancement processing on the target speech frame according to the amplitude spectrum corresponding to the target speech frame to obtain a first amplitude spectrum.
[0055] The speech signal changes non-stationarily randomly over time, but the characteristics of the speech signal are strongly correlated within a short period of time, that is, the speech signal has short-time correlation. Therefore, in the solution of this application, speech enhancement is performed in units of speech frames. The target speech frame refers to the speech frame to be enhanced currently.
[0056] The amplitude spectrum corresponding to the target speech frame can be obtained by performing time-frequency transformation on the time-domain signal of the target speech frame. The time-frequency transformation is, for example, the Short-term Fourier transform (STFT). By performing time-frequency transformation, the amplitude spectrum and phase spectrum of the target speech frame can be obtained. The phase spectrum of the target speech frame indicates the phase information of the target speech frame.
[0057] The first amplitude spectrum is the amplitude spectrum obtained after pre-enhancing the target speech frame. By pre-enhancing the target speech frame, part of the noise in the target speech frame can be removed. Therefore, compared with the amplitude spectrum obtained by performing time-frequency transformation on the target speech frame, the influence of noise in the first amplitude spectrum obtained by pre-enhancement is less.
[0058] In this solution, the pre-enhancement of the target speech frame aims to obtain the first amplitude spectrum. Therefore, during the pre-enhancement process, the phase information of the target speech frame does not need to be concerned, thereby reducing the computational amount.
[0059] In some embodiments of this application, deep learning can be used to pre-enhance the target speech frame. By training a neural network model to predict the amplitude spectrum of the noise in the speech frame according to the amplitude spectrum of the speech frame, and then subtracting the predicted amplitude spectrum of the noise from the amplitude spectrum of the speech frame, the first amplitude spectrum is obtained. For ease of description, the neural network model used to predict the amplitude spectrum of the noise in the speech frame is called the noise amplitude prediction model. After training, the noise amplitude model can output the predicted amplitude spectrum of the noise according to the input amplitude spectrum of the speech frame, and then subtract the amplitude spectrum of the noise from the amplitude spectrum of the speech frame, that is, the first amplitude spectrum is obtained.
[0060] In some embodiments of the present application, a neural network model can also be trained to predict the magnitude spectrum of an enhanced speech frame based on the magnitude spectrum of the speech frame. For ease of description, the neural network model used to predict the enhanced magnitude spectrum is referred to as the magnitude spectrum prediction model. During the training process, the magnitude spectrum of the sample speech frame is input into the magnitude spectrum prediction model, and the magnitude spectrum prediction model predicts the enhanced magnitude spectrum, and adjusts the parameters of the magnitude spectrum prediction model according to the predicted enhanced magnitude spectrum and the label information of the sample speech frame until the difference between the predicted enhanced magnitude spectrum and the magnitude spectrum indicated by the label information meets the preset requirements. The label information of the sample speech frame is used to indicate the magnitude spectrum of the original speech signal in the sample speech frame. After the training is completed, the magnitude spectrum prediction model can output the first magnitude spectrum according to the magnitude spectrum of the target speech frame.
[0061] Step 420, perform speech decomposition on the target speech frame according to the first magnitude spectrum to obtain the glottal parameters, gain, and excitation signal corresponding to the target speech frame.
[0062] The glottal parameters, corresponding gain, and corresponding excitation signal obtained by the speech decomposition are used to reconstruct the original speech signal in the target speech frame according to the Figure 2 process shown.
[0063] As described above, an original speech signal is obtained by exciting a glottal filter with an excitation signal and then performing gain control. The first magnitude spectrum includes information about the original speech signal of the target speech frame. Therefore, based on the first magnitude spectrum, linear prediction analysis can be performed to reversely determine the glottal parameters, excitation signal, and gain used to reconstruct the original speech signal in the target speech frame.
[0064] The glottal parameters refer to the parameters used to construct the glottal filter. Once the glottal parameters are determined, the glottal filter is correspondingly determined. The glottal filter is a digital filter. The glottal parameters can be linear prediction coding (Linear Prediction Coefficients, LPC) coefficients or line spectral frequency (Line Spectral Frequency, LSF) parameters. The number of glottal parameters corresponding to the target speech frame is related to the order of the glottal filter. If the glottal filter is a K - order filter, the glottal parameters include K - order LSF parameters or K - order LPC coefficients, where the LSF parameters and LPC coefficients can be converted into each other.
[0065] A glottal filter of order p can be expressed as:
[0066] A p (z)=1 + a 1 z -1 + a 2 z -2 +…+ ap z -p ; (Equation 2)
[0067] where a 1 , a 2 ,..., a p are LPC coefficients; p is the order of the glottal filter; z is the input signal of the glottal filter.
[0068] Based on Equation 2, if we let:
[0069] P(z) = A p (z) - z -(p+1) A p (z -1 ); (Equation 3)
[0070] Q(z) = A p (z) + z -(p+1) A p (z -1 ); (Equation 4)
[0071] We can obtain:
[0072]
[0073] Physically speaking, P(z) and Q(z) respectively represent the periodic change laws of glottal opening and glottal closing. The roots of the polynomials P(z) and Q(z) appear alternately in the complex plane, and their distribution is a series of angular frequencies on the unit circle of the complex plane. The LSF parameter is the angular frequency corresponding to the roots of P(z) and Q(z) on the unit circle of the complex plane. The LSF parameter LSF(n) corresponding to the nth speech frame can be expressed as ω n . Of course, the LSF parameter LSF(n) corresponding to the nth speech frame can also be directly expressed by the roots of P(z) and the roots of Q(z) corresponding to this nth speech frame.
[0074] Define the roots of P(z) and Q(z) corresponding to the nth speech frame in the complex plane as θ n , then the LSF parameter corresponding to the nth speech frame is expressed as:
[0075]
[0076] where Rel{θ n} represents the real part of the complex number θ n ; Imag{θ n} represents the imaginary part of the complex number θ n .
[0077] In some embodiments of the present application, voice decomposition can be performed in a deep learning manner. Neural network models for predicting glottal parameters, predicting excitation signals, and predicting gains can be trained first, so that the three neural network models can respectively predict the predicted glottal parameters, excitation signals, and gains corresponding to the target voice frame based on the first magnitude spectrum.
[0078] In some embodiments of the present application, according to the principle of linear prediction analysis, signal processing can also be performed based on the first magnitude spectrum, and the glottal parameters, excitation signals, and gains corresponding to the target voice frame can be calculated. For the specific process, refer to the following description.
[0079] Step 430, perform synthesis processing according to the glottal parameters, the gain, and the excitation signal to obtain the enhanced voice signal corresponding to the target voice frame.
[0080] When the glottal parameters corresponding to the target voice frame are determined, the corresponding glottal filter is determined accordingly. On this basis, according to Figure 2 the generation process of the original voice signal shown, impact the determined glottal filter with the excitation signal corresponding to the target voice frame, and perform gain control on the filtered signal according to the gain corresponding to the target voice frame to realize the reconstruction of the original voice signal.
[0081] In the solution of the present application, based on pre-enhancing the target voice frame according to the magnitude spectrum corresponding to the target voice frame to obtain the first magnitude spectrum, voice decomposition and synthesis are performed on the target voice frame based on the first magnitude spectrum, realizing the enhancement of the target voice frame in two stages, which can effectively ensure the voice enhancement effect. Compared with the magnitude spectrum before the pre-enhancement of the target voice frame, there is less noise information in the first magnitude spectrum. And in the voice decomposition process, noise will affect the accuracy of voice decomposition. Therefore, using the first magnitude spectrum as the basis for voice decomposition can reduce the difficulty of voice decomposition, improve the accuracy of the glottal parameters, excitation signals, and gains obtained by voice decomposition, and further ensure the accuracy of the subsequent obtained enhanced voice signal.
[0082] In some embodiments of the present application, step 410 includes: inputting the magnitude spectrum of the target voice frame into a first neural network, where the first neural network is trained according to the magnitude spectrum corresponding to the sample voice frame and the magnitude spectrum corresponding to the original voice signal in the sample voice frame; and outputting the first magnitude spectrum by the first neural network according to the magnitude spectrum of the target voice frame.
[0083] The first neural network can be a model constructed by a long short-term memory neural network, a convolutional neural network, a recurrent neural network, a fully connected neural network, etc., and no specific limitation is made here.
[0084] In some embodiments of the present application, multiple sample speech frames can be obtained by framing a sample speech signal. Among them, the sample speech signal can be obtained by combining a known original speech signal with a known noise signal. Then, when the original speech signal is known, the original speech signal in the sample speech frames can be subjected to time-frequency transformation correspondingly to obtain the amplitude spectrum corresponding to the original speech signal in the sample speech frames. The amplitude spectrum corresponding to the sample speech frame can be obtained by performing time-frequency transformation on the time-domain signal of the sample speech frame.
[0085] During the training process, the amplitude spectrum corresponding to the sample speech frame is input into the first neural network. The first neural network makes a prediction based on the amplitude spectrum corresponding to the sample speech frame and outputs a predicted first amplitude spectrum. Then, the predicted first amplitude spectrum is compared with the amplitude spectrum corresponding to the original speech signal in the sample speech frame. If the similarity between the two does not meet the preset requirement, the parameters of the first neural network are adjusted until the difference between the predicted first amplitude spectrum output by the first neural network and the amplitude spectrum corresponding to the original speech signal in the sample speech frame meets the preset requirement. Among them, the preset requirement can be that the similarity between the predicted first amplitude spectrum and the amplitude spectrum corresponding to the original speech signal in the sample speech frame is not lower than the similarity threshold, and the similarity threshold can be set as needed, such as 100%, 98%, etc. Through the above training process, the first neural network can learn the ability to predict the first amplitude spectrum based on the input amplitude spectrum.
[0086] Figure 5 is a schematic structural diagram of the first neural network shown according to a specific embodiment. As Figure 5 shown, the first neural network includes two LSTM (Long-Short Term Memory) layers and two FC (Full-Connected) layers. The two LSTM layers are hidden layers. The input information first passes through two cascaded LSTM layers and then passes through two cascaded FC layers to obtain the output information. Along the direction from input to output, the two LSTM layers respectively include 512 units and 256 units, and the two FC layers respectively include 512 units and 256 units. An activation function σ() is provided in the first FC layer to increase the non-linear expression ability of the first neural network; no activation function is provided in the second FC layer, which is used as a classifier for classification output.
[0087] In a specific embodiment of the present application, the input of the first LSTM layer in the first neural network can be a 320-dimensional vector. In other embodiments, considering the DC component in the target speech frame, the DC component can also be input into the first neural network, and then the input of the first LSTM layer is a 321-dimensional vector. Of course, Figure 5This is merely an exemplary illustration of the structure of the first neural network and should not be considered as a limitation on the scope of use of this application.
[0088] In some embodiments of the present application, as Figure 6 shown, step 410 includes:
[0089] Step 610: Input the amplitude spectrum corresponding to the target speech frame into a second neural network, where the second neural network is trained based on the amplitude spectrum corresponding to the sample speech frame and the amplitude envelopes of each sub-band in the amplitude spectrum corresponding to the original speech signal in the sample speech frame.
[0090] Step 620: The second neural network outputs the amplitude envelope corresponding to each sub-band in the target speech frame according to the amplitude spectrum of the target speech frame.
[0091] The second neural network refers to a neural network model for predicting the amplitude envelope. The second neural network can be a model constructed by a long short-term memory neural network, a convolutional neural network, a recurrent neural network, a fully connected neural network, etc., and is not specifically limited here.
[0092] By dividing the amplitude spectrum along the frequency, multiple sub-bands in the amplitude spectrum can be obtained. The band division performed on the amplitude spectrum can be uniform frequency band division (i.e., each sub-band corresponds to the same frequency width) or non-uniform frequency band division, which is not specifically limited here. It can be understood that each sub-band corresponds to a frequency range, which includes multiple frequency points.
[0093] The non-uniform frequency band division can be Bark band division. Bark band division is performed according to the Bark frequency scale. The Bark frequency scale maps the frequency to multiple critical frequency bands in psychoacoustics. The number of frequency bands can be set according to the sampling rate and actual needs. For example, the number of frequency points is set to 24. Bark band division conforms to the characteristics of the auditory system. Generally, the lower the frequency, the fewer the number of coefficients included in the sub-band, or even just a single coefficient, and the higher the frequency, the more the number of coefficients included in the sub-band.
[0094] The amplitude spectrum is a set of many different frequencies, forming a very wide frequency range. The amplitudes of different frequencies may be different. The curve formed by connecting the highest amplitude points of different frequencies is called the amplitude spectrum envelope line. The amplitude envelope of a sub-band refers to the corresponding value of the sub-band on the amplitude spectrum envelope line. The amplitude envelope of a sub-band is the opening of the average energy of adjacent STFT amplitude spectrum coefficients.
[0095] Similarly, when the original speech signal in the sample speech frame is known, the amplitude spectrum of the original speech signal in the sample speech frame can be correspondingly determined. Based on this, the amplitude envelope of each subband can be determined according to the amplitudes of each frequency point in the subband. The amplitude spectrum corresponding to the sample speech frame is obtained by performing time-frequency transformation on the time-domain signal of the sample speech frame.
[0096] During the training process, the amplitude spectrum corresponding to the sample speech frame is input into the second neural network, and the second neural network predicts the amplitude envelope of each subband to obtain the predicted amplitude envelope of each subband. If the predicted amplitude envelope of a subband is inconsistent with the amplitude envelope of the subband in the amplitude spectrum corresponding to the original speech signal in the sample speech frame, the parameters of the second neural network are adjusted. Through this training process, the second neural network can learn the ability to predict the amplitude envelope of a subband based on the amplitude spectrum of the speech frame.
[0097] Step 630: Generate the first amplitude spectrum according to the amplitude envelope corresponding to each subband in the target speech frame and the amplitudes of each frequency point in the amplitude spectrum of the target speech frame.
[0098] In some embodiments of the present application, step 630 further includes: step 631, determining the first gain corresponding to each subband according to the amplitude envelope respectively corresponding to each subband in the target speech frame; step 632, adjusting the amplitude value of each frequency point in the corresponding subband of the amplitude spectrum of the target speech frame according to the first gain corresponding to each subband to obtain the first amplitude of each frequency point in each subband; step 633, combining the first amplitudes of each frequency point in the target speech frame to obtain the first amplitude spectrum.
[0099] In the process of determining the first gain corresponding to a subband according to the amplitude envelope corresponding to the subband, the amplitudes of each frequency point in the subband in the amplitude spectrum of the target speech frame can also be combined. For example, calculate the average amplitude in the subband according to the amplitudes of each frequency point in the subband in the amplitude spectrum of the target speech frame, and use the ratio of the amplitude envelope of the subband to the average amplitude in the subband as the first gain corresponding to the subband.
[0100] After obtaining the first gain corresponding to each subband, the frequency points in the same subband share the first gain corresponding to the subband, and the amplitude value of each frequency point in the subband in the amplitude spectrum of the target speech frame is adjusted according to the first gain corresponding to the subband to obtain the first amplitude of each frequency point in the subband. Based on this, the first amplitudes of each frequency point in the target speech frame are combined to obtain the first amplitude spectrum.
[0101] In the solution of this embodiment, since the frequency points in the same sub-band are adjacent, and there is a correlation between the STFT coefficients of adjacent frequency points, all the frequency points in the same sub-band can share the same first gain. On this basis, compared with the first neural network, the number of parameters to be predicted by the second neural network is smaller, so that the dimension of the output data of the second neural network is smaller than the dimension of the output data of the first neural network. That is to say, in this embodiment, it is not necessary to determine the first gain for each frequency point in the sub-band respectively. Relatively speaking, the computational complexity of the neural network model is reduced. Therefore, the method of this embodiment can use a neural network model with a simpler structure than the first neural network as the second neural network.
[0102] Figure 7 is a schematic structural diagram of a second neural network shown according to a specific embodiment. As Figure 7 shown, in the direction from input to output, the second neural network includes two cascaded LSTM layers and two cascaded FC layers. Among them, the two LSTM layers are hidden layers, each including 256 units and 128 units respectively. The input of the first LSTM layer is 320-dimensional STFT (Short-term Fourier Transform) coefficients, and the STFT coefficients are the coefficients representing amplitude values obtained by performing short-term Fourier transform on the target speech frame; the two FC layers include 256 units and 64 units respectively. Among them, the last FC layer has no activation function and is used as a classifier for classification output, and the output m′(n) is the amplitude envelope corresponding to the target speech frame.
[0103] Compared with Figure 5 the first neural network shown, Figure 7 the last FC layer of the second neural network in Figure 6 includes 64 units, indicating that the dimension of m′(n) output by the last FC layer is 64; Figure 8 the last FC layer of the first neural network in Figure 5 includes 256 units, indicating that the dimension of the data output by the last FC layer is 256; if it is evenly divided into bands, it is equivalent to taking Figure 7 four adjacent coefficients among the 256 STFT coefficients output in Figure 5 as a sub-band. In this sub-band division method, it can be ensured that Figure 7 the output dimension of the second neural network in
[0104] In some embodiments of the present application, as Figure 8 shown, step 420 includes:
[0105] Step 810: Calculate the power spectrum after pre - enhancement corresponding to the target speech frame based on the first magnitude spectrum and the phase spectrum corresponding to the target speech frame.
[0106] In the solution of this application, since only the magnitude spectrum is focused on for enhancement during the pre - enhancement process of the target speech frame, and the phase spectrum is not enhanced synchronously. Therefore, to avoid loss of phase information, the phase spectrum corresponding to the target speech frame is reused, and combined with the first magnitude spectrum obtained by pre - enhancement, the power spectrum calculated based on the phase spectrum of the target speech frame before pre - enhancement and the first magnitude spectrum after pre - enhancement is used as the power spectrum after pre - enhancement corresponding to the target speech frame.
[0107] Combining the first magnitude spectrum and the phase spectrum corresponding to the target speech frame can obtain a complex frequency spectrum, which can express the amplitude and phase information of the signal after pre - enhancement; in other words, using the phase spectrum corresponding to the target speech frame as the real part and the first magnitude spectrum as the imaginary part to form a complex frequency spectrum. Assuming that the complex frequency spectrum obtained by combining the first magnitude spectrum and the phase spectrum corresponding to the target speech frame is S′_c(n), then the power spectrum Pa(n) after pre - enhancement corresponding to the target speech frame obtained in step 810 is:
[0108] Pa(n)=(RealS′_c(n)) 2 +(ImagS′_c(n)) 2 ; (Formula 7)
[0109] where RealS′_c(n) represents the real part of the complex frequency spectrum S′_c(n), and ImagS′_c(n) represents the imaginary part of the complex frequency spectrum S′_c(n).
[0110] Step 820: Calculate the autocorrelation coefficient based on the power spectrum after pre - enhancement.
[0111] According to the Wiener - Khinchin theorem: The power spectrum of a stationary random process and its autocorrelation function are a pair of Fourier transform relationships. In this solution, a frame of speech is regarded as a stationary random signal. Therefore, based on the power spectrum after pre - enhancement corresponding to the target speech frame, the inverse Fourier transform can be performed on the power spectrum after pre - enhancement corresponding to the target speech frame to obtain the autocorrelation coefficient corresponding to this power spectrum after pre - enhancement.
[0112] Specifically, step 820 includes: performing an inverse Fourier transform on the power spectrum to obtain an inverse transform result; extracting the real part of the inverse transform result to obtain the autocorrelation coefficient. That is:
[0113] AC(n)=Real(iFFT(Pa(n))); (Formula 8)
[0114] AC(n) represents the autocorrelation coefficient corresponding to the nth speech frame. iFFT (Inverse Fast Fourier Transform) refers to the inverse transform of FFT (Fast Fourier Transform), and Real represents the real part of the result obtained by the inverse fast Fourier transform. AC(n) includes multiple parameters, and the coefficients in AC(n) can be further expressed as AC j (n), where 0 ≤ j ≤ p and p is the order of the glottal filter.
[0115] Step 830, calculate the glottal parameters according to the autocorrelation coefficient.
[0116] According to the Yule-Walker equation, for the nth speech frame, there is the following relationship between its corresponding autocorrelation coefficient and the corresponding glottal parameters:
[0117] k - KA = 0; (Formula 9)
[0118] Among them, k is the autocorrelation vector; K is the autocorrelation matrix; A is the LPC coefficient matrix. Specifically,
[0119]
[0120] Among them, AC j (n) = E[s(n)s(n - j)], where 0 ≤ j ≤ p; (Formula 10)
[0121] p is the order of the glottal filter; a 1 (n), a 2 (n),..., a p (n) are all the LPC coefficients corresponding to the nth speech frame, which are respectively a 1 and a 2 ,..., a p in Formula 2 above; since a 0 (n) is a constant 1, a 0 (n) can also be regarded as an LPC coefficient corresponding to the nth speech frame.
[0122] Based on the obtained autocorrelation coefficient, the autocorrelation vector and the autocorrelation matrix can be determined correspondingly, and then by solving Formula 9, the LPC coefficients can be obtained. In a specific embodiment, the Levinson-Durbin algorithm can be used to solve Formula 9.
[0123] Since the LSF parameters and the LPC coefficients can be converted into each other, when the LPC coefficients are calculated, the LSF parameters can be determined accordingly. In other words, regardless of whether the glottal parameters are LPC coefficients or LSF parameters, they can be determined through the above process.
[0124] Step 840, calculate the gain according to the autocorrelation coefficient and the glottal parameters.
[0125] The gain corresponding to the nth frame of speech can be calculated according to the following formula.
[0126]
[0127] It is worth mentioning that G(n) calculated according to Formula 11 is the square of the gain corresponding to the target speech frame in the time domain representation.
[0128] Step 850, calculate the power spectrum of the excitation signal according to the gain and the power spectrum of the glottal filter, where the glottal filter is a filter constructed according to the glottal parameters.
[0129] Assume that the amplitude spectrum corresponding to the target speech frame is obtained by performing a Fourier transform on m (m is a positive integer) sample points. To calculate the power spectrum of the glottal filter, first construct a zero array s_AR(n) of dimension m for the nth frame of speech; then, assign the (p + 1)-dimensional a j (n) to the first (p + 1) dimensions of this zero array, where j = 0, 1, 2,... p; by calling the Fast Fourier Transform (FFT) of m sample points, obtain the FFT coefficients:
[0130] S_AR(n) = FFT(s_AR(n)); (Formula 12)
[0131] Based on obtaining the FFT coefficients S_AR(n), the power spectrum of the glottal filter corresponding to the nth frame of speech can be obtained for each sample according to the following Formula 13:
[0132] AR_LPS(n, k) = (Real(S_AR(n, k))) 2 +(Imag(S_AR(n, k))) 2 (Formula 13) where Real(S_AR(n, k)) represents the real part of S_AR(n, k), Imag(S_AR(n, k)) represents the imaginary part of S_AR(n, k), k represents the series number of the FFT coefficients, 0 ≤ k ≤ m, and k is a positive integer.
[0133] After obtaining the power spectrum AR_LPS(n) of the glottal filter corresponding to the nth speech frame, for the convenience of calculation, the power spectrum AR_LPS(n) of the glottal filter is converted from the natural number domain to the logarithmic domain according to Equation 14:
[0134] AR_LPS 1 (n) = log 10 (AR_LPS(n)); (Equation 14)
[0135] Taking the negative of the above AR_LPS 1 (n) according to the following Equation 15, the power spectrum AR_LPS 2 (n) corresponding to the inverse glottal filter is obtained:
[0136] AR_LPS 2 (n) = -1 * AR_LPS 1 (n); (Equation 15)
[0137] Then, the power spectrum R(n) of the excitation signal corresponding to the target speech frame can be calculated according to the following Equation 16:
[0138] R(n) = Pa(n) * (G 1 (n)) 2 * AR_LPS 3 (n); (Equation 16)
[0139] where
[0140]
[0141] Through the above process, the power spectra of the glottal parameters, gain, and excitation signal corresponding to the target speech frame are calculated correspondingly, as well as the power spectrum of the glottal filter defined by the glottal parameters.
[0142] After obtaining the gain corresponding to the target speech frame, the power spectrum of the corresponding excitation signal, and the power spectrum of the glottal filter defined by the glottal parameters, the synthesis process can be carried out according to the Figure 9 process shown. As Figure 9 shown, step 430 includes:
[0143] Step 910, generating a second amplitude spectrum according to the power spectrum of the glottal filter and the power spectrum of the excitation signal.
[0144] The first amplitude spectrum S_filt(n) can be calculated according to the following Equation 19:
[0145]
[0146] where R 1(n) = 10 * log 10 (R(n)); (Equation 20)
[0147] Step 920, enhance the second amplitude spectrum according to the said gain to obtain an enhanced amplitude spectrum.
[0148] The enhanced amplitude spectrum S_e(n) can be obtained according to the following formula:
[0149] S_e(n) = G 2 (n) * S_filt(n); (Equation 21)
[0150] Wherein,
[0151] Step 930, determine the enhanced speech signal corresponding to the target speech frame according to the phase spectrum corresponding to the target speech frame and the enhanced amplitude spectrum.
[0152] In some embodiments of the present application, step 930 further includes: combining the phase spectrum corresponding to the target speech frame and the enhanced amplitude spectrum to obtain a target complex frequency spectrum; transforming the target complex frequency spectrum into the time domain to obtain the time domain signal of the enhanced speech signal corresponding to the target speech frame. Specifically, the real part of the target complex frequency spectrum is the enhanced amplitude spectrum, and the imaginary part of the target complex frequency spectrum is the phase spectrum corresponding to the target speech frame.
[0153] The phase spectrum corresponding to the target speech frame is obtained by performing time-frequency transformation on the time domain signal of the target speech frame. Since the obtained enhanced amplitude spectrum does not carry the phase information of the signal, the phase spectrum corresponding to the target speech frame is reused to provide the phase information.
[0154] The complex frequency spectrum of the signal is in complex form, including a real part and an imaginary part. Among them, the real part reflects the amplitude of the signal, and the imaginary part reflects the phase of the signal. In step 930, the enhanced amplitude spectrum is used as the real part, and the phase spectrum corresponding to the target speech frame is used as the imaginary part. The obtained complex expression is the target complex frequency spectrum. On this basis, the target complex frequency spectrum is transformed into the time domain, and the obtained signal is the time domain signal of the enhanced speech signal corresponding to the target speech frame.
[0155] In Figure 8 and Figure 9Based on the embodiments, if in step 410, the target speech frame is pre-enhanced in the manner of deep learning, this solution realizes the deep combination of traditional signal processing and deep learning, performs secondary enhancement on the target speech frame, and realizes multi-stage enhancement of the target speech frame. That is, in the first stage, deep learning is used to pre-enhance according to the amplitude spectrum of the target speech frame, which can reduce the difficulty of obtaining glottal parameters, excitation signals, and gains in speech decomposition in the second stage. In the second stage, glottal parameters, excitation signals, and gains for reconstructing the original speech signal are obtained through signal processing. Moreover, in the second stage, speech synthesis is performed according to the digital model of speech generation, and the signal of the target speech frame is not directly processed. Therefore, the situation of voice clipping can be avoided in the second stage.
[0156] Figure 10 It is a flowchart of a speech enhancement method shown according to a specific embodiment. Assume that the nth speech frame is used as the target speech frame, and the time-domain signal of the nth speech frame is s(n). As Figure 10 shown, it specifically includes steps 1010 - 1050.
[0157] Step 1010, time-frequency transformation; through step 1010, the time-domain signal s(n) of the nth speech frame is subjected to time-frequency transformation to obtain the amplitude spectrum S(n) corresponding to the nth speech frame and the phase spectrum Ph(n) corresponding to the nth speech frame.
[0158] Step 1020, pre-enhancement; based on the amplitude spectrum S(n) corresponding to the nth speech frame, pre-enhancement processing is performed on the nth speech frame to obtain the first amplitude spectrum S′(n) after pre-enhancement of the nth speech frame.
[0159] Step 1030, speech decomposition; based on the first amplitude spectrum S′(n) of the nth speech frame, speech decomposition is performed to obtain the glottal parameter set P(n) corresponding to the nth speech frame and the frequency-domain representation R(n) of the excitation signal corresponding to the nth speech frame. Among them, the glottal parameter set P(n) includes the glottal parameter ar(n) and the gain G(n). Among them, the obtained frequency-domain representation of the excitation signal can be the power spectrum of the excitation signal in the above Figure 8 and Figure 9 shown embodiments. The glottal parameter ar(n) can be defined by the LPC coefficients calculated above or the power spectrum of the glottal filter obtained based on the LPC coefficients.
[0160] In some embodiments, the phase information required in the speech decomposition process can reuse the phase spectrum Ph(n) corresponding to the nth speech frame.
[0161] Step 1040, speech synthesis. Based on the glottal parameters ar(n), gain G(n) corresponding to the nth speech frame obtained and the frequency-domain representation R(n) of the excitation signal corresponding to the nth speech frame, speech synthesis is performed to obtain the enhanced amplitude spectrum S_e(n) corresponding to the nth speech frame.
[0162] Step 1050, frequency-time transformation. The phase spectrum of the nth speech frame is reused as the phase spectrum of the enhanced speech signal corresponding to the nth speech frame. Therefore, the enhanced complex spectrum corresponding to the nth speech frame is obtained by combining the phase spectrum Ph(n) corresponding to the nth speech frame and the enhanced amplitude spectrum S_e(n) corresponding to the nth speech frame. The obtained enhanced complex spectrum is transformed into the time domain, that is, the time-domain signal s_e(n) of the enhanced speech signal corresponding to the nth speech frame is obtained.
[0163] In addition, in this solution, in the two stages of pre-enhancement and enhancement based on speech decomposition and synthesis, the goal is to obtain the amplitude spectrum. Therefore, in these two stages, the phase information of the target speech frame does not need to be concerned about, and the phase spectrum of the target speech frame can be directly reused, reducing the processing amount in the two speech enhancement stages on the premise of not losing the phase information.
[0164] In some embodiments of the present application, step 410 includes: predicting the glottal parameters of the target speech frame according to the first amplitude spectrum to obtain the glottal parameters corresponding to the target speech frame. Predicting the excitation signal of the target speech frame according to the first amplitude spectrum to obtain the excitation signal corresponding to the target speech frame. Predicting the gain of the target speech frame according to the gain corresponding to the historical speech frame of the target speech frame to obtain the gain corresponding to the target speech frame.
[0165] In some embodiments of the present application, a neural network model for glottal parameter prediction, a neural network model for excitation signal prediction, and a neural network model for gain prediction can be trained respectively. Among them, these three neural network models can be models constructed by long short-term memory neural networks, convolutional neural networks, recurrent neural networks, fully connected neural networks, etc., and specific limitations are not provided here.
[0166] In some embodiments of the present application, the step of predicting the glottal parameters of the target speech frame according to the first amplitude spectrum to obtain the glottal parameters corresponding to the target speech frame further includes: inputting the first amplitude spectrum into a third neural network, and the third neural network is trained according to the glottal parameters corresponding to the sample speech frame and the amplitude spectrum corresponding to the sample speech frame; the third neural network outputs the glottal parameters corresponding to the target speech frame according to the first amplitude spectrum.
[0167] The third neural network refers to a neural network model used for glottal parameter prediction. Among them, the third neural network can be a model constructed by a long short-term memory neural network, a convolutional neural network, a recurrent neural network, a fully connected neural network, etc., and specific limitations are not provided here.
[0168] The amplitude spectrum of the sample speech frame is obtained by performing time-frequency transformation on the time-domain signal of the sample speech frame. In some embodiments of the present application, the sample speech signal can be framed to obtain multiple sample speech frames. Among them, the sample speech signal can be obtained by combining a known original speech signal with a known noise signal. Then, when the original speech signal is known, the glottal parameters corresponding to the sample speech frame can be obtained by performing linear prediction analysis on the original speech signal.
[0169] During the training process, after inputting the amplitude spectrum of the sample speech frame into the third neural network, the third neural network predicts the glottal parameters according to the amplitude spectrum of the sample speech frame and outputs the predicted glottal parameters. Then, the predicted glottal parameters are compared with the glottal parameters corresponding to the sample speech frame. If the two are inconsistent, the parameters of the third neural network are adjusted until the predicted glottal parameters output by the third neural network according to the amplitude spectrum of the sample speech frame are consistent with the glottal parameters corresponding to the sample speech frame. After the training is completed, the third neural network learns the ability to accurately predict the glottal parameters for reconstructing the original speech signal in the speech frame according to the input amplitude spectrum of the speech frame.
[0170] Figure 11 is a schematic diagram of the third neural network shown according to a specific embodiment. As Figure 11 shown, the third neural network includes one layer of LSTM (Long-Short Term Memory) layer and three cascaded FC (Full Connected) layers. Among them, the LSTM layer has 1 hidden layer, which includes 256 units. The input of the LSTM layer is the frequency-domain representation S(n) of the nth frame of the speech frame. In this embodiment, the input of the LSTM layer is 321-dimensional STFT coefficients. In the three cascaded FC layers, activation functions σ() are provided in the first two FC layers. The set activation functions are used to increase the non-linear expression ability of the third neural network. No activation function is provided in the last FC layer, and this last FC layer is used as a classifier for classification output. As Figure 11 shown, from bottom to top, the three FC layers include 512, 512, and 16 units respectively. The output of the last FC layer is the 16-dimensional line spectral frequency coefficients LSF(n) corresponding to the nth frame of the speech frame, that is, the 16th-order line spectral frequency coefficients.
[0171] In some embodiments of the present application, since there is a correlation between speech frames and the similarity of frequency-domain features between adjacent two speech frames is relatively high, therefore, the glottal parameters corresponding to the target speech frame can be predicted by combining the glottal parameters corresponding to the historical speech frames before the target speech frame. In one embodiment, the step of predicting the glottal parameters of the target speech frame according to the first amplitude spectrum to obtain the glottal parameters corresponding to the target speech frame further includes: inputting the first amplitude spectrum and the glottal parameters corresponding to the historical speech frames of the target speech frame into a third neural network, and the third neural network is trained according to the glottal parameters corresponding to the sample speech frames, the glottal parameters corresponding to the historical speech frames of the sample speech frames, and the amplitude spectrum corresponding to the sample speech frames; and outputting the glottal parameters corresponding to the target speech frame by the third neural network according to the first amplitude spectrum and the glottal parameters corresponding to the historical speech frames of the target speech frame.
[0172] Since there is a correlation between the historical speech frames and the target speech frame, and there is a similarity between the glottal parameters corresponding to the historical speech frames of the target speech frame and the glottal parameters corresponding to the target speech frame, therefore, taking the glottal parameters corresponding to the historical speech frames of the target speech frame as a reference to supervise the prediction process of the glottal parameters of the target speech frame can improve the accuracy of glottal parameter prediction.
[0173] In one embodiment of the present application, since the similarity of glottal parameters of speech frames that are closer is higher, therefore, taking the glottal parameters corresponding to the historical speech frames closer to the target speech frame as a reference can further ensure the prediction accuracy. For example, the glottal parameters corresponding to the previous speech frame of the target speech frame can be taken as a reference. In a specific embodiment, the number of historical speech frames used as a reference can be one frame or multiple frames, and can be specifically selected according to actual needs.
[0174] The glottal parameters corresponding to the historical speech frames of the target speech frame can be the glottal parameters predicted for the historical speech frame. In other words, during the process of predicting glottal parameters, the glottal parameters predicted for the historical speech frames are reused as a reference for the process of predicting the glottal parameters of the current speech frame.
[0175] The training process of the third neural network in this embodiment is similar to the training process of the third neural network in the previous embodiment, and the training process will not be elaborated here.
[0176] Figure 12 It is a schematic diagram of the input and output of the third neural network shown according to another embodiment, where Figure 12 the structure of the third neural network in Figure 11 is the same as that in Figure 11 , compared with Figure 12The input of the third neural network also includes the line spectral frequency parameter LSF(n - 1) of the previous speech frame (i.e., the (n - 1)-th frame) of the n-th speech frame. As Figure 12 shown, the line spectral frequency coefficient LSF(n - 1) of the previous speech frame of the n-th speech frame is embedded in the second FC layer as reference information. Since the similarity of LSF parameters between two adjacent speech frames is very high, therefore, if the LSF parameters corresponding to the historical speech frames of the n-th speech frame are used as reference information, the prediction accuracy of LSF parameters can be improved.
[0177] In some embodiments of the present application, the step of predicting an excitation signal for the target speech frame according to the first magnitude spectrum to obtain the excitation signal corresponding to the target speech frame further includes: inputting the first magnitude spectrum into a fourth neural network, where the fourth neural network is trained according to the magnitude spectrum corresponding to the sample speech frame and the magnitude spectrum of the excitation signal corresponding to the sample speech frame; and outputting, by the fourth neural network, the magnitude spectrum of the excitation signal corresponding to the target speech frame according to the first magnitude spectrum.
[0178] During the training of the fourth neural network, the magnitude spectrum of the sample speech frame is input into the fourth neural network model, and then the fourth neural network predicts the excitation signal according to the input magnitude spectrum of the sample speech frame and outputs the magnitude spectrum of the predicted excitation signal; and then the parameters of the fourth neural network are adjusted according to the magnitude spectrum of the predicted excitation signal and the magnitude spectrum of the excitation signal corresponding to the sample speech frame, that is: if the similarity between the magnitude spectrum of the predicted excitation signal and the magnitude spectrum of the excitation signal corresponding to the sample speech frame does not meet the preset requirement, then the parameters of the fourth neural network are adjusted until the similarity between the magnitude spectrum of the predicted excitation signal output by the fourth neural network for the sample speech frame and the magnitude spectrum of the excitation signal corresponding to the sample speech frame meets the preset requirement. Through the above training process, the fourth neural network can learn the ability to predict the magnitude spectrum of the excitation signal corresponding to a speech frame according to the magnitude spectrum of the speech frame, so as to accurately predict the excitation signal.
[0179] Figure 13 is a schematic diagram of the fourth neural network shown according to a specific embodiment. As Figure 13As shown in the figure, the fourth neural network includes one LSTM layer and three FC layers. Among them, the LSTM layer has one hidden layer, including 256 units. The input of the LSTM is the first amplitude spectrum S′(n) obtained by pre-enhancing the nth speech frame, and its dimension can be 321. The number of units included in the three FC layers are 512, 512, and 321 respectively. The last FC layer outputs the amplitude spectrum R(n) of the excitation signal corresponding to the nth speech frame with a dimension of 321. Along the direction from input to output, activation functions are provided in the first two FC layers of the three FC layers to enhance the non-linear expression ability of the model, and there is no activation function in the last FC layer for classification output.
[0180] In some embodiments of the present application, the step of predicting the gain of the target speech frame according to the gain corresponding to the historical speech frames of the target speech frame to obtain the gain corresponding to the target speech frame further includes: inputting the gain corresponding to the historical speech frames of the target speech frame into a fifth neural network, and the fifth neural network is trained according to the gain corresponding to the sample speech frames and the gain corresponding to the historical speech frames of the sample speech frames; outputting the gain corresponding to the target speech frame by the fifth neural network according to the gain corresponding to the historical speech frames of the target speech frame.
[0181] The gain corresponding to the historical speech frames of the target speech frame can be the gain predicted by the fifth neural network for the historical speech frame. In other words, the gain predicted for the historical speech frame is reused as the input of the fifth neural network model in the process of predicting the gain of the target speech frame.
[0182] Figure 14 It is a schematic diagram of the fifth neural network shown according to a specific embodiment. As Figure 14 shown, the fifth neural network includes one LSTM layer and one FC layer. Among them, the LSTM layer has one hidden layer, which includes 128 units; the input of the FC layer is a 512-dimensional vector, and the output is a 1-dimensional gain. In a specific embodiment, the historical speech frame gain G_pre(n) of the nth speech frame can be defined as the gains corresponding to the first 4 speech frames of the nth speech frame, that is:
[0183] G_pre(n) = {G(n - 1), G(n - 2), G(n - 3), G(n - 4)};
[0184] Of course, the number of historical speech frames selected for gain prediction is not limited to the above examples, and can be specifically selected according to actual needs.
[0185] As shown above, the second neural network, the third neural network, and the fifth neural network as a whole present an M-to-N mapping relationship (N << M), that is, the dimension of the input information of the neural network model is M, and the dimension of the output information is N, greatly streamlining the structure of the neural network model and reducing the complexity of the model.
[0186] It is worth mentioning that the structures of the first neural network, the second neural network, the third neural network, the fourth neural network, and the fifth neural network shown above are only exemplary examples. In other embodiments, a neural network model with a corresponding structure can also be set on an open-source platform for deep learning and trained accordingly.
[0187] In some embodiments of the present application, on the basis of predicting the glottal parameters, gain, and amplitude spectrum of the excitation signal, step 430 includes: constructing a glottal filter according to the glottal parameters corresponding to the target speech frame; filtering the excitation signal corresponding to the target speech frame through the glottal filter to obtain a first speech signal; and amplifying the first speech signal according to the gain corresponding to the target speech frame to obtain the enhanced speech signal corresponding to the target speech frame.
[0188] If the glottal parameters are LPC coefficients, the glottal filter can be directly constructed according to the above formula (2). If the glottal filter is a K-order filter, the glottal parameters corresponding to the target speech frame include K-order LPC coefficients, that is, a 1 ,a 2 ,...,a K In other embodiments, the constant 1 in the above formula (2) can also be used as an LPC coefficient.
[0189] If the glottal parameters are LSF parameters, the LSF parameters can be converted into LPC coefficients, and then the glottal filter can be constructed according to the above formula (2).
[0190] The filtering process is convolution in the time domain. Therefore, the process of filtering the excitation signal through the glottal filter as described above can be converted to the time domain. Then, on the basis of predicting the amplitude spectrum of the excitation signal corresponding to the target speech frame, the amplitude spectrum of the excitation signal is transformed into the time domain to obtain the time-domain signal of the excitation signal corresponding to the target speech frame.
[0191] In the solution of the present application, the target speech frame is a digital signal, which includes a plurality of sample points. The excitation signal is filtered by a glottal filter, that is, the historical sample points before a sample point are convolved with the glottal filter to obtain the target signal value corresponding to the sample point. In some embodiments of the present application, the target speech frame includes a plurality of sample points; the glottal filter is a K - order filter, where K is a positive integer; the excitation signal includes the excitation signal values corresponding to the plurality of sample points in the target speech frame; the step of filtering the excitation signal corresponding to the target speech frame by the glottal filter to obtain the first speech signal further includes: convolving the excitation signal values corresponding to the first K sample points before each sample point in the target speech frame with the K - order filter to obtain the target signal value corresponding to each sample point in the target speech frame; combining the target signal values corresponding to all the sample points in the target speech frame in chronological order to obtain the first speech signal. Among them, the expression of the K - order filter can refer to the above formula 1. That is to say, for each sample point in the target speech frame, the excitation signal values corresponding to the first K sample points before it are used to convolve with the K - order filter to obtain the target signal value corresponding to each sample point.
[0192] It can be understood that for the first sample point in the target speech frame, it is necessary to calculate the target signal value corresponding to the first sample point by means of the excitation signal values of the last K sample points in the previous speech frame of the target speech frame. Similarly, for the second sample point in the target speech frame, it is necessary to convolve the excitation signal values of the last (K - 1) sample points in the previous speech frame of the target speech frame and the excitation signal value of the first sample point in the target speech frame with the K - order filter to obtain the target signal value corresponding to the second sample point in the target speech frame.
[0193] In the related art, there are methods of speech enhancement through spectral estimation and spectral regression prediction. The speech enhancement method of spectral estimation believes that a mixed speech contains a speech part and a noise part. Therefore, the noise can be estimated through a statistical model, etc. The spectrum corresponding to the mixed speech is subtracted from the spectrum corresponding to the noise, and the remaining is the speech spectrum. Thus, a clean speech signal is restored through the spectrum obtained by subtracting the spectrum corresponding to the noise from the spectrum corresponding to the mixed speech. The speech enhancement method of spectral regression prediction predicts the masking threshold corresponding to the speech frame through a neural network, and this masking threshold reflects the proportion of the speech component and the noise component in each frequency point of the speech frame; then, the gain control of the mixed signal spectrum is performed according to the masking threshold to obtain the enhanced spectrum.
[0194] The above speech enhancement methods predicted by spectral estimation and spectral regression are based on the estimation of the posterior probability of the noise spectrum, and there may be inaccurate estimation of the noise. For example, transient noises such as keyboard typing occur instantaneously, and the estimated noise spectrum is very inaccurate, resulting in poor noise suppression effect. In the case of inaccurate noise spectrum prediction, if the original mixed speech signal is processed according to the estimated noise spectrum, it may cause speech distortion in the mixed speech signal or poor noise suppression effect; therefore, in this case, a compromise needs to be made between speech fidelity and noise suppression.
[0195] In the solution of this application, since the glottal parameters are strongly correlated with the glottal characteristics in the physical process of sound generation, the predicted glottal parameters effectively guarantee the speech structure of the original speech signal in the target speech frame. Therefore, synthesizing the enhanced speech signal of the target speech frame based on the predicted glottal parameters, excitation signal and gain can effectively avoid the reduction of the original speech and effectively protect the speech structure; at the same time, after predicting the glottal parameters, excitation signal and gain corresponding to the target speech frame, since the original noisy speech will no longer be processed, there is no need to make a compromise between speech fidelity and noise suppression.
[0196] In some embodiments of this application, before step 410, the method further includes: obtaining the time-domain signal of the target speech frame; performing time-frequency transformation on the time-domain signal of the target speech frame to obtain the amplitude spectrum corresponding to the target speech frame and the phase spectrum corresponding to the target speech frame.
[0197] The time-frequency transformation can be the short-term Fourier transform (STFT).
[0198] In the short-term Fourier transform, the operation of windowing and overlapping is adopted to eliminate the non-smoothness between frames. Figure 15 It is a schematic diagram showing the windowing and overlapping in a specific short-term Fourier transform. In Figure 15 it, the operation of 50% windowing and overlapping is adopted. If the short-term Fourier transform is applied to 640 sample points, the number of overlapping samples (hop-size) of the window function is 320. The window function used for windowing can be the Hanning window, Hamming window, etc. Of course, other window functions can also be used, and no specific limitation is made here.
[0199] In other embodiments, the operation of non-50% windowing and overlapping can also be adopted. For example, if the short-term Fourier transform is applied to 512 sample points, in this case, if a speech frame includes 320 sample points, only 192 sample points of the previous speech frame need to be overlapped.
[0200] In some embodiments of the present application, the time-domain signal of the target speech frame can be obtained through the following steps: Obtain the speech signal to be processed, where the speech signal to be processed is the collected speech signal or the speech signal obtained by decoding the encoded speech; Frame the speech signal to be processed to obtain the time-domain signal of the target speech frame.
[0201] In some instances, the speech signal to be processed can be framed according to a set frame length, and the frame length can be set according to actual needs. For example, the frame length can be set to 20 ms.
[0202] As described above, the solution of the present application can be applied to the transmitting end for speech enhancement or to the receiving end for speech enhancement.
[0203] When the solution of the present application is applied to the transmitting end, the speech signal to be processed is the speech signal collected by the transmitting end. Then, the speech signal to be processed is framed to obtain multiple speech frames.
[0204] After framing, the speech signal to be processed is segmented into multiple speech frames. Then, each speech frame can be used as the target speech frame, and the target speech frame can be enhanced according to the process of steps 410-440 above. Further, after obtaining the enhanced speech signal corresponding to the target speech frame, the enhanced speech signal can also be encoded for transmission based on the obtained encoded speech.
[0205] In one embodiment, since the directly collected speech signal is an analog signal, for the convenience of signal processing, before framing, the signal further needs to be digitized to convert the time-continuous speech signal into a time-discrete speech signal. During the digitization process, the collected speech signal can be sampled according to a set sampling rate. The set sampling rate can be 16000 Hz, 8000 Hz, 32000 Hz, 48000 Hz, etc., and can be specifically set according to actual needs.
[0206] When the solution of the present application is applied to the receiving end, the voice signal to be processed is the voice signal obtained by decoding the received encoded voice. In this case, it may be that the transmitting end does not enhance the voice signal to be transmitted. Therefore, in order to improve the signal quality, it is necessary to enhance the voice at the receiving end. After obtaining multiple voice frames by framing the voice signal to be processed, they are used as target voice frames and enhanced according to the process of steps 410-440 above to obtain the enhanced voice signal of the target voice frame. Further, the enhanced voice signal corresponding to the target voice frame can also be played. Since the obtained enhanced voice signal has less noise and higher quality compared to the signal before the target voice frame is enhanced, the auditory experience for the user is better.
[0207] The following describes the device embodiments of the present application, which can be used to execute the methods in the above embodiments of the present application. For the details not disclosed in the device embodiments of the present application, please refer to the above method embodiments of the present application.
[0208] Figure 16 is a block diagram of a voice enhancement device shown according to an embodiment, as Figure 16 shown, the voice enhancement device includes: a pre-enhancement module 1610, configured to perform pre-enhancement processing on the target voice frame according to the amplitude spectrum corresponding to the target voice frame to obtain a first amplitude spectrum. A voice decomposition module 1620, configured to decompose the target voice frame according to the first amplitude spectrum to obtain the glottal parameters, gain, and excitation signal corresponding to the target voice frame; a synthesis module 1630, configured to perform synthesis processing according to the glottal parameters, the gain, and the excitation signal to obtain the enhanced voice signal corresponding to the target voice frame.
[0209] In some embodiments of the present application, the voice decomposition module 1620 includes: a power spectrum calculation unit, configured to calculate the pre-enhanced power spectrum corresponding to the target voice frame according to the first amplitude spectrum and the phase spectrum corresponding to the target voice frame; an autocorrelation coefficient calculation unit, configured to calculate the autocorrelation coefficient according to the pre-enhanced power spectrum; a glottal parameter calculation unit, configured to calculate the glottal parameters according to the autocorrelation coefficient; a gain calculation unit, configured to calculate the gain according to the autocorrelation coefficient and the glottal parameters; an excitation signal determination unit, configured to calculate the power spectrum of the excitation signal according to the gain and the power spectrum of the glottal filter, and the glottal filter is a filter constructed according to the glottal parameters.
[0210] In some embodiments of the present application, the synthesis module 1630 includes: a first frequency response acquisition unit configured to acquire the frequency response of a glottal filter, the glottal filter being a filter constructed according to the glottal parameters; a second frequency domain response acquisition unit configured to acquire the frequency response of the excitation signal; a first amplitude spectrum generation unit configured to generate a first amplitude spectrum based on the power spectrum of the glottal filter and the power spectrum of the excitation signal; an enhancement unit configured to enhance the first amplitude spectrum according to the gain to obtain an enhanced amplitude spectrum; and an enhanced speech signal determination unit configured to determine the enhanced speech signal corresponding to the target speech frame based on the phase spectrum corresponding to the target speech frame and the enhanced amplitude spectrum.
[0211] In some embodiments of the present application, the enhanced speech signal determination unit includes: a combination unit configured to combine the phase spectrum corresponding to the target speech frame and the enhanced amplitude spectrum to obtain a target complex frequency spectrum; and an enhanced speech signal determination unit configured to transform the target complex frequency spectrum into the time domain to obtain the time domain signal of the enhanced speech signal corresponding to the target speech frame.
[0212] In some embodiments of the present application, the pre-enhancement module 1610 includes: a first input unit configured to input the amplitude spectrum of the target speech frame into a first neural network, the first neural network being trained according to the amplitude spectrum corresponding to the sample speech frame and the amplitude spectrum of the original speech signal in the sample speech frame; and a first output unit configured to output the first amplitude spectrum by the first neural network according to the amplitude spectrum of the target speech frame.
[0213] In some embodiments of the present application, the pre-enhancement module 1610 includes: a second input unit configured to input the amplitude spectrum corresponding to the target speech frame into a second neural network, the second neural network being trained according to the amplitude spectrum corresponding to the sample speech frame and the amplitude envelopes of each sub-band in the amplitude spectrum of the original speech signal in the sample speech frame; a second output unit configured to output the amplitude envelope corresponding to each sub-band in the target speech frame by the second neural network according to the amplitude spectrum of the target speech frame; and a first amplitude spectrum generation unit configured to generate the first amplitude spectrum based on the amplitude envelope corresponding to each sub-band in the target speech frame and the amplitudes of each frequency point in the amplitude spectrum of the target speech frame.
[0214] In some embodiments of the present application, the first amplitude spectrum generation unit includes: a first gain determination unit configured to determine a first gain corresponding to each sub-band according to the amplitude envelope corresponding to each sub-band in the target speech frame; a first amplitude value determination unit configured to adjust the amplitude value of each frequency point in the corresponding sub-band of the amplitude spectrum of the target speech frame according to the first gain corresponding to each sub-band to obtain a first amplitude value of each frequency point in each sub-band; and a first amplitude value combination unit configured to combine the first amplitude values of each frequency point in the target speech frame to obtain the first amplitude spectrum.
[0215] In some embodiments of the present application, the speech decomposition module 1620 includes: a glottal parameter prediction unit configured to perform glottal parameter prediction on the target speech frame according to the first amplitude spectrum to obtain the glottal parameters corresponding to the target speech frame; an excitation signal prediction unit configured to perform excitation signal prediction on the target speech frame according to the first amplitude spectrum to obtain the excitation signal corresponding to the target speech frame; and a gain prediction unit configured to perform gain prediction on the target speech frame according to the gain corresponding to the historical speech frame of the target speech frame to obtain the gain corresponding to the target speech frame.
[0216] In some embodiments of the present application, the glottal parameter prediction unit includes: a third input unit configured to input the first amplitude spectrum into a third neural network, where the third neural network is trained according to the glottal parameters corresponding to the sample speech frame and the amplitude spectrum corresponding to the sample speech frame; and a third output unit configured to output, by the third neural network, the glottal parameters corresponding to the target speech frame according to the first amplitude spectrum.
[0217] In some embodiments of the present application, the glottal parameter prediction unit includes: a fourth input unit configured to input the first amplitude spectrum and the glottal parameters corresponding to the historical speech frame of the target speech frame into a third neural network, where the third neural network is trained according to the glottal parameters corresponding to the sample speech frame, the glottal parameters corresponding to the historical speech frame of the sample speech frame, and the amplitude spectrum corresponding to the sample speech frame; and a fourth output unit configured to output, by the third neural network, the glottal parameters corresponding to the target speech frame according to the first amplitude spectrum and the glottal parameters corresponding to the historical speech frame of the target speech frame.
[0218] In some embodiments of the present application, the excitation signal prediction unit includes: a fifth input unit configured to input the first amplitude spectrum into a fourth neural network, where the fourth neural network is trained according to the amplitude spectrum corresponding to the sample speech frame and the amplitude spectrum of the excitation signal corresponding to the sample speech frame; and a fifth output unit configured to output, by the fourth neural network, the amplitude spectrum of the excitation signal corresponding to the target speech frame according to the first amplitude spectrum.
[0219] In some embodiments of the present application, the gain prediction unit includes: a sixth input unit configured to input the gain corresponding to the historical speech frames of the target speech frame into a fifth neural network, where the fifth neural network is trained based on the gain corresponding to the sample speech frames and the gain corresponding to the historical speech frames of the sample speech frames; and a sixth output unit configured to output, by the fifth neural network, the gain corresponding to the target speech frame based on the gain corresponding to the historical speech frames of the target speech frame.
[0220] Figure 17 FIG. shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application.
[0221] It should be noted that Figure 17 the computer system 1700 of the shown electronic device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0222] As Figure 17 shown, the computer system 1700 includes a central processing unit (CPU) 1701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1702 or the program loaded from the storage section 1708 into the random access memory (RAM) 1703, such as executing the methods in the above embodiments. In the RAM 1703, various programs and data required for system operations are also stored. The CPU 1701, ROM 1702, and RAM 1703 are connected to each other via a bus 1704. An input / output (I / O) interface 1705 is also connected to the bus 1704.
[0223] The following components are connected to the I / O interface 1705: an input portion 1706 including a keyboard, a mouse, etc.; an output portion 1707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 1708 including a hard disk, etc.; and a communication portion 1709 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication portion 1709 performs communication processing via a network such as the Internet. A drive 1710 is also connected to the I / O interface 1705 as needed. A removable medium 1711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1710 as needed so that a computer program read therefrom can be installed into the storage portion 1708 as needed.
[0224] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1709 and / or installed from the removable medium 1711. When the computer program is executed by the central processing unit (CPU) 1701, various functions defined in the system of the present application are executed.
[0225] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0226] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutively represented blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0227] The units involved in the embodiments described in the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the unit itself.
[0228] On the other hand, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist separately and not be assembled into the electronic device. The above computer-readable storage medium carries computer-readable instructions, and when the computer-readable storage instructions are executed by a processor, the methods in any of the above embodiments are implemented.
[0229] According to one aspect of the present application, an electronic device is also provided, which includes: a processor; a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the methods in any of the above embodiments are implemented.
[0230] According to one aspect of the embodiments of the present application, a computer program product or a computer program is provided, and the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods in any of the above embodiments.
[0231] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0232] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by the way of software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0233] After considering the specification and practicing the embodiments disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application.
[0234] It should be understood that the present application is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A voice enhancement method, characterized in that, it includes: Performing pre-enhancement processing on the target voice frame according to the amplitude spectrum corresponding to the target voice frame to obtain a first amplitude spectrum; Performing voice decomposition on the target voice frame according to the first amplitude spectrum to obtain the glottal parameters, gain, and excitation signal corresponding to the target voice frame; Performing synthesis processing according to the glottal parameters, the gain, and the excitation signal to obtain an enhanced voice signal corresponding to the target voice frame; Among them, the performing pre-enhancement processing on the target voice frame according to the amplitude spectrum corresponding to the target voice frame to obtain a first amplitude spectrum includes: Inputting the amplitude spectrum corresponding to the target voice frame into a second neural network, and the second neural network is trained according to the amplitude spectrum corresponding to the sample voice frame and the amplitude envelope of each sub-band in the amplitude spectrum corresponding to the original voice signal in the sample voice frame; Outputting, by the second neural network, the amplitude envelope corresponding to each sub-band in the target voice frame according to the amplitude spectrum of the target voice frame; Generating the first amplitude spectrum according to the amplitude envelope corresponding to each sub-band in the target voice frame and the amplitudes of each frequency point in the amplitude spectrum of the target voice frame.
2. The method according to claim 1, characterized in that, the performing voice decomposition on the target voice frame according to the first amplitude spectrum to obtain the glottal parameters, gain, and excitation signal corresponding to the target voice frame includes: Calculating the pre-enhanced power spectrum corresponding to the target voice frame according to the first amplitude spectrum and the phase spectrum corresponding to the target voice frame; Calculating the autocorrelation coefficient according to the pre-enhanced power spectrum; Calculating the glottal parameters according to the autocorrelation coefficient; Calculating the gain according to the autocorrelation coefficient and the glottal parameters; Calculating the power spectrum of the excitation signal according to the gain and the power spectrum of the glottal filter, and the glottal filter is a filter constructed according to the glottal parameters.
3. The method according to claim 2, characterized in that, the performing synthesis processing according to the glottal parameters, the gain, and the excitation signal to obtain an enhanced voice signal corresponding to the target voice frame includes: Generating a second amplitude spectrum according to the power spectrum of the glottal filter and the power spectrum of the excitation signal; Enhancing the second amplitude spectrum according to the gain to obtain an enhanced amplitude spectrum; Determining the enhanced voice signal corresponding to the target voice frame according to the phase spectrum corresponding to the target voice frame and the enhanced amplitude spectrum.
4. The method according to claim 3, characterized in that, the determining the enhanced voice signal corresponding to the target voice frame according to the phase spectrum corresponding to the target voice frame and the enhanced amplitude spectrum includes: Combining the phase spectrum corresponding to the target voice frame and the enhanced amplitude spectrum to obtain a target complex frequency spectrum; Performing a transformation of the target complex frequency spectrum to the time domain to obtain the time domain signal of the enhanced voice signal corresponding to the target voice frame.
5. The method according to claim 1, characterized in that, Generating the first amplitude spectrum according to the amplitude envelope corresponding to each sub-band in the target speech frame and the amplitude values of each frequency point in the amplitude spectrum of the target speech frame includes: Determining a first gain corresponding to each sub-band according to the amplitude envelope respectively corresponding to each sub-band in the target speech frame; adjusting the amplitude value of each frequency point in the corresponding sub-band of the amplitude spectrum of the target speech frame according to the first gain corresponding to each sub-band to obtain the first amplitude value of each frequency point in each sub-band; Combining the first amplitude values of each frequency point in the target speech frame to obtain the first amplitude spectrum.
6. The method according to claim 1, wherein, Performing speech decomposition on the target speech frame according to the first amplitude spectrum to obtain the glottal parameters, gain, and excitation signal corresponding to the target speech frame includes: Performing glottal parameter prediction on the target speech frame according to the first amplitude spectrum to obtain the glottal parameters corresponding to the target speech frame; Performing excitation signal prediction on the target speech frame according to the first amplitude spectrum to obtain the excitation signal corresponding to the target speech frame; Performing gain prediction on the target speech frame according to the gain corresponding to the historical speech frame of the target speech frame to obtain the gain corresponding to the target speech frame.
7. The method according to claim 6, wherein, Performing glottal parameter prediction on the target speech frame according to the first amplitude spectrum to obtain the glottal parameters corresponding to the target speech frame includes: Inputting the first amplitude spectrum into a third neural network, where the third neural network is trained according to the glottal parameters corresponding to the sample speech frame and the amplitude spectrum corresponding to the sample speech frame; Outputting the glottal parameters corresponding to the target speech frame by the third neural network according to the first amplitude spectrum.
8. The method according to claim 6, wherein, Performing glottal parameter prediction on the target speech frame according to the first amplitude spectrum to obtain the glottal parameters corresponding to the target speech frame includes: Inputting the first amplitude spectrum and the glottal parameters corresponding to the historical speech frame of the target speech frame into a third neural network, where the third neural network is trained according to the glottal parameters corresponding to the sample speech frame, the glottal parameters corresponding to the historical speech frame of the sample speech frame, and the amplitude spectrum corresponding to the sample speech frame; Outputting the glottal parameters corresponding to the target speech frame by the third neural network according to the first amplitude spectrum and the glottal parameters corresponding to the historical speech frame of the target speech frame.
9. The method according to claim 6, wherein, Performing excitation signal prediction on the target speech frame according to the first amplitude spectrum to obtain the excitation signal corresponding to the target speech frame includes: Inputting the first amplitude spectrum into a fourth neural network, where the fourth neural network is trained according to the amplitude spectrum corresponding to the sample speech frame and the amplitude spectrum of the excitation signal corresponding to the sample speech frame; Outputting the amplitude spectrum of the excitation signal corresponding to the target speech frame by the fourth neural network according to the first amplitude spectrum.
10. The method according to claim 6, wherein, Performing gain prediction on the target speech frame according to the gain corresponding to the historical speech frames of the target speech frame to obtain the gain corresponding to the target speech frame includes: Inputting the gain corresponding to the historical speech frames of the target speech frame into a fifth neural network, where the fifth neural network is trained according to the gain corresponding to the sample speech frames and the gain corresponding to the historical speech frames of the sample speech frames; Outputting, by the fifth neural network, the gain corresponding to the target speech frame according to the gain corresponding to the historical speech frames of the target speech frame.
11. A speech enhancement device Characterized in that It includes: A pre-enhancement module for performing pre-enhancement processing on the target speech frame according to the amplitude spectrum corresponding to the target speech frame to obtain a first amplitude spectrum; A speech decomposition module for performing speech decomposition on the target speech frame according to the first amplitude spectrum to obtain the glottal parameters, gain, and excitation signal corresponding to the target speech frame; A synthesis module for performing synthesis processing according to the glottal parameters, the gain, and the excitation signal to obtain an enhanced speech signal corresponding to the target speech frame; Wherein the pre-enhancement module is further configured to perform the following steps: Inputting the amplitude spectrum corresponding to the target speech frame into a second neural network, where the second neural network is trained according to the amplitude spectrum corresponding to the sample speech frames and the amplitude envelopes of each sub-band in the amplitude spectrum corresponding to the original speech signal in the sample speech frames; Outputting, by the second neural network, the amplitude envelope corresponding to each sub-band in the target speech frame according to the amplitude spectrum of the target speech frame; Generating the first amplitude spectrum according to the amplitude envelope corresponding to each sub-band in the target speech frame and the amplitudes of each frequency point in the amplitude spectrum of the target speech frame.
12. An electronic device Characterized in that It includes: A processor; A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method described in any one of claims 1-10 is implemented.
13. A computer-readable storage medium, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the method described in any one of claims 1-10 is implemented.
14. A computer program product Characterized in that It includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method described in any one of claims 1-10.
Citation Information
Patent Citations
Voice processing method, device and equipment and storage medium
CN111554322A
Voice gain control method and computer storage medium
CN112242147A