A speech processing method, apparatus, electronic device, and readable medium
By calculating the glottal gain and excitation gain of the speech frame, and combining energy compensation to process the spectrum coefficients of the speech frame, the dependence on the completeness of the training data in the prior art is solved, and effective processing of the non-contained noise and improving the speech quality are achieved.
Patent Information
- Application Number
- CN202111238478.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-22
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-10-22
AI Technical Summary
When processing noise-containing speech, the prior art needs to collect training data for various types of noise, resulting in the model processing effect being affected by the completeness of the training data, and it is impossible to effectively process the noise types and noise environments not included in the training data.
By obtaining the spectral coefficients of the to-process speech frame, calculating the glottal gain and excitation gain, and performing gain control based on these gains, combined with energy compensation prediction, the target speech frame is obtained. This method does not require training for various types of noises, and can effectively deal with unincluded noise types and noise environments.
Improves the effect of speech noise reduction, reduces dependence on the completeness of training data, and maintains good performance in the face of untrained noise.
Smart Images

Figure CN114333893B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular, to a voice processing method, apparatus, electronic device, and readable medium. Background Art
[0002] With the development of computer technology, various voice communication or voice control technologies have emerged. Through such technologies, users are allowed to communicate over long distances or the efficiency of human-computer interaction can be improved. In a real environment, when a user is in the surrounding environment, various environmental noises will be collected by devices such as microphones, and the quality of voice communication will be affected to varying degrees. Therefore, voice enhancement has become an important topic.
[0003] In related technologies, a deep learning method is used to learn the signal features of a noisy voice audio, so as to predict the proportion of the voice component and the noise component, and then the noisy voice is enhanced according to the prediction result to achieve the noise reduction effect.
[0004] However, in the above solution, it is necessary to collect training data for various noises to train the model, so that the trained model can process the noise types covered in the training data. Therefore, the processing effect of the model is affected by the completeness of the training data, and the noise reduction effect is poor when facing situations not included in the training data. Summary of the Invention
[0005] Based on the above technical problems, the present application provides a voice processing method, apparatus, electronic device, and readable medium, thereby reducing the influence of the completeness of training data, and being able to effectively process noise types and noise environments not included in the training data, and improving the noise reduction effect.
[0006] Other features and advantages of the present application will become apparent through the following detailed description, or will be partially learned through the practice of the present application.
[0007] According to one aspect of the embodiments of the present application, a voice processing method is provided, including:
[0008] Obtain the spectral coefficients of the voice frame to be processed;
[0009] Calculate the glottal gain according to the spectral coefficients of the voice frame to be processed to obtain a first gain, where the first gain corresponds to the glottal characteristics of the voice frame to be processed;
[0010] Calculate the excitation gain according to the spectral coefficients of the voice frame to be processed to obtain a second gain, where the second gain corresponds to the excitation signal of the voice frame to be processed;
[0011] Compensate and predict based on the spectral coefficients of the to-be-processed speech frame to obtain a control coefficient, where the control coefficient is determined according to the energy of the spectral coefficients of the to-be-processed audio frame;
[0012] Perform gain control on the to-be-processed speech frame according to the first gain, the second gain, and the control coefficient to obtain a target speech frame.
[0013] According to one aspect of the embodiments of the present application, there is provided a speech processing device, including:
[0014] A spectral coefficient acquisition module, configured to acquire the spectral coefficients of a to-be-processed speech frame;
[0015] A glottal gain module, configured to perform glottal gain calculation according to the spectral coefficients of the to-be-processed speech frame to obtain a first gain, where the first gain corresponds to the glottal characteristics of the to-be-processed speech frame;
[0016] An excitation gain module, configured to perform excitation gain calculation according to the spectral coefficients of the to-be-processed speech frame to obtain a second gain, where the second gain corresponds to the excitation signal of the to-be-processed speech frame;
[0017] A compensation prediction module, configured to perform compensation prediction according to the spectral coefficients of the to-be-processed speech frame to obtain a control coefficient, where the control coefficient is determined according to the energy of the spectral coefficients of the to-be-processed audio frame;
[0018] A gain control module, configured to perform gain control on the to-be-processed speech frame according to the first gain, the second gain, and the control coefficient to obtain a target speech frame.
[0019] In some embodiments of the present application, based on the above technical solution, the glottal gain module includes:
[0020] A first input sub-unit, configured to input the spectral coefficients of the to-be-processed speech frame into a first neural network, where the first neural network is trained according to the glottal characteristics corresponding to the noisy speech frame and the glottal characteristics corresponding to the denoised speech frame corresponding to the noisy speech frame;
[0021] A gain prediction sub-module, configured to perform gain prediction according to the spectral coefficients of the to-be-processed speech frame through the first neural network to obtain the first gain.
[0022] In some embodiments of the present application, based on the above technical solution, the speech processing device further includes:
[0023] A historical frame spectral coefficient acquisition module, configured to acquire the spectral coefficients of the historical speech frames of the to-be-processed speech frame;
[0024] The first input sub-module includes: a historical frame input unit, configured to input the spectral coefficients of the to-be-processed speech frame and the spectral coefficients of the historical speech frame into a first neural network.
[0025] In some embodiments of the present application, based on the above technical solution, the glottal gain module includes:
[0026] A first gain calculation sub-module, configured to perform gain calculation on the spectral coefficients of the to-be-processed speech frame through the first neural network to obtain first glottal gains corresponding to each sub-band in the to-be-processed speech frame, where the sub-band corresponds to at least one frequency band in the spectral coefficients of the to-be-processed speech frame;
[0027] A sub-band gain merging sub-module, configured to merge the first glottal gains corresponding to each sub-band as the first gain.
[0028] In some embodiments of the present application, based on the above technical solution, the speech processing device further includes:
[0029] A second gain calculation sub-module, configured to perform predictive analysis on the spectral coefficients of the to-be-processed speech frame and the pitch period of the to-be-processed speech frame through the first neural network to determine a second glottal gain, where the second glottal gain corresponds to the long-term correlation feature of the spectral coefficients of the to-be-processed speech frame;
[0030] The sub-band gain merging sub-module includes:
[0031] A glottal gain merging unit, configured to merge the first glottal gains corresponding to each sub-band and the second glottal gain as the first gain.
[0032] In some embodiments of the present application, based on the above technical solution, the first gain calculation sub-module includes:
[0033] A first glottal parameter prediction unit, configured to perform parameter prediction on the spectral coefficients of the to-be-processed speech frame through the first neural network to obtain first glottal parameters, where the first glottal parameters are used to represent the short-term correlation feature of the spectral coefficients of the to-be-processed speech frame;
[0034] A first prediction unit, configured to perform gain prediction on the basis of the first glottal parameters through the first neural network to obtain the first prediction result;
[0035] A result determination unit, configured to determine the first gain according to the first prediction result.
[0036] In some embodiments of the present application, based on the above technical solution, the speech processing device further includes:
[0037] A second glottal parameter prediction unit, configured to perform parameter prediction through the first neural network according to the spectral coefficients of the to-be-processed speech frame and the pitch period of the to-be-processed speech frame, so as to obtain a second glottal parameter, where the first glottal parameter is used to represent the long-term correlation characteristics of the spectral coefficients of the to-be-processed speech frame;
[0038] A second prediction unit, configured to perform gain prediction through the first neural network according to the second glottal parameter, so as to obtain a second prediction result;
[0039] The result determination unit includes:
[0040] A result merging subunit, configured to merge and determine the first prediction result and the second prediction result as the first gain.
[0041] In some embodiments of the present application, based on the above technical solutions, the excitation gain module includes:
[0042] A second input sub-module, configured to input the spectral coefficients of the to-be-processed speech frame into a second neural network, where the second neural network is trained according to the excitation signal of the noise speech frame and the excitation signal of the denoised speech frame corresponding to the noise speech frame;
[0043] A speech decomposition sub-module, configured to perform speech decomposition on the spectral coefficients of the to-be-processed speech frame through the second neural network, so as to obtain an excitation signal;
[0044] A gain prediction sub-module, configured to perform gain prediction through the second neural network according to the excitation signal, so as to obtain the second gain.
[0045] In some embodiments of the present application, based on the above technical solutions, the gain control module includes:
[0046] A first enhancement sub-module, configured to enhance the to-be-processed speech frame according to the second gain, so as to obtain a first enhancement result;
[0047] A second enhancement sub-module, configured to perform gain operation on each sub-band in the first enhancement result according to the first gain, so as to obtain a second enhancement result;
[0048] An energy compensation sub-module, configured to perform energy compensation on the second enhancement result according to the control coefficient, so as to obtain a third enhancement result;
[0049] An inverse time-frequency conversion sub-module, configured to perform inverse time-frequency conversion according to the third enhancement result, so as to obtain an enhanced speech frame as the target speech frame.
[0050] In some embodiments of the present application, based on the above technical solutions, the compensation prediction module includes:
[0051] A gain control sub-module, configured to perform gain control on the to-be-processed speech frame according to the first gain and the second gain, so as to obtain a gain control result;
[0052] A control coefficient prediction sub-module, configured to perform compensation prediction according to the gain control result and the spectral coefficients of the to-be-processed speech frame, so as to obtain the control coefficient.
[0053] In some embodiments of the present application, based on the above technical solutions, the speech processing device further includes:
[0054] An amplitude spectrum calculation module, configured to calculate the amplitude spectrum and phase spectrum corresponding to the to-be-processed speech frame according to the to-be-processed speech frame;
[0055] The gain control module includes:
[0056] An amplitude spectrum gain control sub-module, configured to perform gain control on the amplitude spectrum corresponding to the to-be-processed speech frame according to the first gain and the second gain, so as to obtain an enhanced amplitude spectrum;
[0057] An amplitude spectrum energy compensation sub-module, configured to perform energy compensation according to the enhanced amplitude spectrum and the control coefficient, so as to obtain a compensated amplitude spectrum;
[0058] An amplitude spectrum inverse time-frequency conversion sub-module, configured to perform inverse time-frequency conversion according to the compensated amplitude spectrum and the phase spectrum corresponding to the to-be-processed speech frame, so as to obtain a target speech frame.
[0059] In some embodiments of the present application, based on the above technical solutions, the compensation prediction module includes:
[0060] A historical spectral coefficient acquisition sub-module, configured to acquire the spectral coefficients of the historical speech frames of the to-be-processed speech frame;
[0061] A third input sub-module, configured to input the spectral coefficients of the to-be-processed speech frame and the spectral coefficients of the historical speech frames into a third neural network, where the third neural network is trained according to the energy of the spectral coefficients corresponding to the noisy speech frames and the energy of the spectral coefficients corresponding to the denoised speech frames corresponding to the noisy speech frames;
[0062] A compensation prediction sub-module, configured to perform compensation prediction according to the spectral coefficients of the to-be-processed speech frame through the third neural network, so as to obtain the control coefficient.
[0063] According to one aspect of the embodiments of the present application, there is provided an electronic device, including: a processor; and a memory, configured to store executable instructions of the processor; wherein, the processor is configured to execute the speech processing method in the above technical solutions by executing the executable instructions.
[0064] According to one aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the voice processing method in the above technical solution is implemented.
[0065] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the voice processing method provided in the above various optional implementation manners.
[0066] In the embodiments of the present application, a first gain and a second gain are respectively calculated for the glottal features and the excitation signal of the noisy voice signal, and then gain control is performed according to the first gain and the second gain, so as to denoise the noisy voice signal and perform energy compensation on the denoised voice signal. Performing noise reduction processing according to the glottal features can specifically identify the human voice part in the voice signal. Therefore, the human voice part is also processed during the noise reduction process, and it is no longer necessary to train for various types of noise. Therefore, the influence on the completeness of the training data is reduced, and noise types and noise environments not included in the training data can be effectively processed, improving the noise reduction effect.
[0067] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] The accompanying drawings here are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0069] Figure 1 It is a schematic diagram of an exemplary system architecture in an application scenario of the technical solution of the present application;
[0070] Figure 2 It shows a schematic diagram of a digital model for generating a voice signal;
[0071] Figure 3 It is a schematic diagram of an exemplary implementation of a glottal filter in the embodiments of the present application;
[0072] Figure 4 It is a schematic diagram of another exemplary implementation of a glottal filter in the embodiments of the present application;
[0073] Figure 5 Schematic diagram showing the frequency responses of the excitation signal and the glottal filter decomposed from the original speech signal at different signal-to-noise ratios;
[0074] Figure 6 Flowchart showing the speech processing method according to an embodiment of the present application;
[0075] Figure 7 Schematic structural diagram of the first neural network according to a specific embodiment;
[0076] Figure 8 Schematic structural diagram of the second neural network according to a specific embodiment;
[0077] Figure 9 Schematic diagram of the overall process in the embodiment of the present application;
[0078] Figure 10 Schematic diagram of the overall process of another solution in the embodiment of the present application;
[0079] Figure 11 Schematic structural diagram of the third neural network according to a specific embodiment;
[0080] Figure 12 Schematic block diagram showing the composition of the speech processing device in the embodiment of the present application;
[0081] Figure 13 Schematic structural diagram of the computer system of the electronic device suitable for implementing the embodiment of the present application. Detailed implementation manners
[0082] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0083] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0084] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0085] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the content and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0086] Noise in the speech signal will greatly reduce the speech quality and affect the user's auditory experience. Therefore, in order to improve the quality of the speech signal, it is necessary to perform enhancement processing on the speech signal to remove the noise as much as possible and retain the original speech information in the speech signal, that is, to obtain a clean signal after denoising.
[0087] The solution of this application can be applied to the scenario of voice calls, such as voice calls through instant messaging software, multi-person calls in game applications, etc., and can also be applied to various cloud technology-based services, such as cloud games, cloud conferences, cloud calls, and cloud education. Among them, voice enhancement can be performed according to this solution at the voice sending end, the voice receiving end, or the server providing the voice communication service.
[0088] Cloud conference is an important part of online office. In a cloud conference, after the voice collection device of the participants in the cloud conference collects the speech signal of the speaker, it needs to send the collected speech signal to other conference participants. This process involves the transmission and playback of the speech signal among multiple participants. If the noise signal mixed in the speech signal is not processed, it will greatly affect the auditory experience of the conference participants. In this scenario, the solution of this application can be applied to enhance the speech signal in the cloud conference, so that the speech signal heard by the conference participants is the enhanced speech signal, improving the quality of the speech signal.
[0089] Cloud conference is an efficient, convenient, and low-cost conference form based on cloud computing technology. Users only need to perform simple and easy operations through the Internet interface to quickly and efficiently synchronously share voice, data files, and videos with teams and customers around the world, and complex technologies such as data transmission and processing in the conference are helped by the cloud conference service provider for the users to operate.
[0090] Currently, domestic cloud conferences mainly focus on service content based on the SaaS (Software as a Service) model, including service forms such as telephone, network, and video. A video conference based on cloud computing is called a cloud conference. In the era of cloud conferences, the transmission, processing, and storage of data are all handled by the computer resources of the video conference provider. Users no longer need to purchase expensive hardware and install cumbersome software. They only need to open the client and enter the corresponding interface to conduct efficient remote conferences.
[0091] The cloud conference system supports multi-server dynamic cluster deployment and provides multiple high-performance servers, greatly enhancing the stability, security, and availability of conferences. In recent years, video conferences have been welcomed by many users because they can significantly improve communication efficiency, continuously reduce communication costs, and bring about an upgrade in internal management level. They have been widely applied in various fields such as government affairs, transportation, finance, operators, education, and enterprises.
[0092] Next, take Voice over Internet Protocol (VoIP) as an example to introduce the application scenario of the embodiments of this application. Please refer to Figure 1 , Figure 1 which is an exemplary system architecture schematic diagram of the technical solution of this application in an application scenario.
[0093] As Figure 1 shown, this system architecture includes a sending end 110 and a receiving end 120. There is a network connection between the sending end 110 and the receiving end 120, and the sending end 110 and the receiving end 120 can conduct voice communication through the network connection.
[0094] As Figure 1 shown, the sending end 110 includes an acquisition module 111, a pre-enhancement module 112, and an encoding module 113. Among them, the acquisition module 111 is used to acquire voice signals, and it can convert the acquired acoustic signals into digital signals; the pre-enhancement module 112 is used to enhance the acquired voice signals to remove the noise in the acquired voice signals and improve the quality of the voice signals. The encoding module 113 is used to encode the enhanced voice signals to improve the anti-interference ability of the voice signals during transmission. The pre-enhancement module 112 can perform voice enhancement according to the method of this application. After enhancing the voice, it is then encoded, compressed, and transmitted, so as to ensure that the signals received by the receiving end are no longer affected by noise.
[0095] The receiving end 120 includes a decoding module 121, a post-enhancement module 122, and a playback module 123. The decoding module 121 is used to decode the received encoded voice to obtain a decoded signal; the post-enhancement module 122 is used to perform enhancement processing on the decoded voice signal; the playback module 123 is used to play the enhanced voice signal. The post-enhancement module 122 can also perform voice enhancement according to the method of this application. In some embodiments, the receiving end 120 may further include a sound effect adjustment module, and this sound effect adjustment module is used to adjust the sound effect of the enhanced voice signal.
[0096] In a specific embodiment, it may be that only the receiving end 120 or only the sending end 110 performs voice enhancement according to the method of this application. Of course, it may also be that both the sending end 110 and the receiving end 120 perform voice enhancement according to the method of this application.
[0097] In some application scenarios, the terminal device in the VoIP system can support other third-party protocols in addition to supporting VoIP communication. For example, traditional PSTN (Public Switched Telephone Network) circuit domain telephones. However, traditional PSTN services cannot perform voice enhancement. In this scenario, voice enhancement can be performed in the terminal acting as the receiving end according to the method of this application.
[0098] Before specifically describing this solution, a voice generation method based on an excitation signal is first introduced. The human voice generation method is that air flow drives the vocal cords to vibrate and generate sound when passing through the vocal cords. The voice generation process of the voice generation method based on an excitation signal includes: at the trachea, an impact signal of pseudo-noise with a certain energy is generated, that is, an excitation signal, which is equivalent to the air flow; the impact signal impacts the glottis filter (equivalent to the human vocal cords), generating pseudo-periodic opening and closing, and thus making a sound. It can be seen that this process simulates the human voice generation process.
[0099] Figure 2 The schematic diagram of the digital model for voice signal generation is shown. Through this digital model, the generation process of the voice signal can be described. As Figure 2 shown, the excitation signal impacts the glottis filter to output a voice signal. Among them, the glottis filter is usually configured according to glottis parameters. The glottis filter can adopt the filter in various schemes for generating voice using the source-filter model. Specifically, please refer to Figure 3 , Figure 3 which is the schematic diagram of the exemplary implementation of the glottis filter in the embodiment of this application. Considering the short-term correlation of the voice signal, the glottis filter can be implemented by a linear predictive coding (LPC) filter. The excitation signal impacts the LPC filter to generate a voice signal.
[0100] On the other hand, according to the classical speech signal processing theory, the LPC filter only reflects the short-term correlation in vocalization. However, for voiced sounds (such as vowels), there is a long-term correlation (Long-Term Prediction, LTP) (or quasi-periodicity); the glottal filter can also be implemented using multiple filters. Specifically, please refer to Figure 4 , Figure 4 which is a schematic diagram of another exemplary implementation of the glottal filter in the embodiments of the present application. As Figure 4 shown, the glottal filter consists of two parts: an LPC filter and an LTP filter. Among them, the LTP filter also receives the pitch period as an input. The pitch period indicates that when calculating the nth sample, the (n - p)th sample point is required, where p is the pitch period.
[0101] Figure 5 shows a schematic diagram of the frequency responses of the excitation signal and the glottal filter decomposed from the original speech signal under different signal-to-noise ratios. Figure 5 a shows a schematic diagram of the frequency response of the original speech signal. Figure 5 b shows a schematic diagram of the frequency response of the glottal filter decomposed from the original speech signal. Figure 5 c shows a schematic diagram of the frequency response of the excitation signal decomposed from the original speech signal. Figure 5 shows two original speech signals and their corresponding decomposition results, represented by solid lines and dashed lines respectively. One of the original signals is a 30 dB signal, and the other is a 0 dB signal. The undulating part in the frequency response diagram of the original speech signal corresponds to the peak position in the frequency response diagram of the glottal filter. The excitation signal is equivalent to the residual signal (i.e., the excitation signal) after linear prediction analysis of the original speech signal. Therefore, its corresponding frequency response is relatively flat. In Figure 5 a, although there are certain differences between the two original speech signals of 30 dB and 0 dB, the differences are relatively insignificant, and there are many overlapping parts. After decomposition, Figure 5 in the frequency response of the glottal filter in Figure 5 b, the differences between the two are relatively obvious, and the overlapping parts are significantly reduced. In
[0102] As can be seen from the above, an excitation signal and a glottal filter can be decomposed from an original speech signal (i.e., a speech signal without noise), and the decomposed excitation signal and glottal filter can be used to represent the original speech signal. Among them, the glottal filter can be represented by glottal parameters. Conversely, if the excitation signal corresponding to an original speech signal and the glottal parameters for determining the glottal filter are known, the original speech signal can be reconstructed according to the corresponding excitation signal and glottal filter.
[0103] Based on this principle, the solution of this application calculates the gain corresponding to the glottal filter and the gain corresponding to the excitation signal respectively to perform gain control on the original speech signal, thereby realizing speech enhancement.
[0104] The implementation details of the technical solution of the embodiments of this application are elaborated in detail below. For the convenience of introduction, please refer to Figure 6 , Figure 6 FIG. shows a flowchart of a speech processing method according to an embodiment of this application. This method can be executed by a computer device with processing capabilities, such as a terminal, a server, etc., which is not specifically limited here. As Figure 6 shown, this method at least includes the following steps S610 to S650:
[0105] Step S610, obtain the spectral coefficients of the speech frame to be processed.
[0106] The speech signal changes randomly over time and is not stationary, but the characteristics of the speech signal are strongly correlated within a short period of time, that is, the speech signal has short-term correlation. Therefore, in the solution of this application, speech processing is performed in units of speech frames. The speech frame to be processed is the currently to-be-processed speech frame, which is any frame in the original noisy audio to be processed.
[0107] When obtaining the spectral coefficients of the speech frame to be processed, it can be obtained by performing a time-frequency transform on the time-domain signal of the speech frame to be processed. The time-frequency transform is, for example, a short-time Fourier transform (STFT). The dimension of the spectral coefficients usually depends on the number of sample points of the speech frame to be processed and the overlapping ratio of the windowing used in the STFT transform. For example, for the frequency-domain representation of 257 sample points, the dimension of the spectral coefficients is [2, 257].
[0108] Step S620, perform glottal gain calculation according to the spectral coefficients of the speech frame to be processed to obtain a first gain, and the first gain corresponds to the glottal characteristics of the speech frame to be processed.
[0109] Glottal gain calculation is a process of calculating the gain for the glottal filter part corresponding to the speech frame to be processed. The calculated first gain is associated with the glottal characteristics of the speech frame to be processed. Depending on the glottal model adopted for the speech frame to be processed, the first gain specifically includes multiple sub-gains. For example, for the LPC+LTP glottal model, the first gain can include the sub-gain corresponding to LPC and the auto-gain corresponding to LTP.
[0110] Glottal gain calculation can be performed in the way of a neural network. The trained neural network is used to directly output the corresponding first gain according to the spectral coefficients of the speech frame to be processed. The neural network is trained in a supervised training manner. The training data includes the noisy speech and the data annotation corresponding to each speech frame calculated for the noisy speech, that is, the denoised speech. The neural network is trained based on the noisy speech and the denoised speech to output the first gain.
[0111] Glottal gain calculation can also be performed in other ways. For example, first, the speech frame to be processed is decomposed according to the glottal model to obtain the glottal parameters of the corresponding glottal filter. Then, both the glottal parameters and the spectral coefficients of the speech frame to be processed are input into the neural network for processing. The neural network simulates the denoised speech based on the glottal parameters and the spectral coefficients of the speech frame to be processed, and then determines the first gain through the simulated denoised speech and the noisy speech.
[0112] Step S630, perform excitation gain calculation according to the spectral coefficients of the speech frame to be processed to obtain the second gain, where the second gain corresponds to the excitation signal of the speech frame to be processed.
[0113] Excitation gain calculation is a process of calculating the gain for the excitation signal part corresponding to the speech frame to be processed. The calculated second gain is associated with the excitation signal of the speech frame to be processed. Specifically, the dimension of the second gain usually corresponds to the spectral coefficients of the speech frame to be processed.
[0114] Excitation gain calculation can be performed in the way of a neural network. The trained neural network is used to directly output the corresponding second gain according to the spectral coefficients of the speech frame to be processed. The neural network is trained in a supervised training manner. The training data includes the noisy speech and the excitation signal obtained after decomposing the corresponding denoised speech of the noisy speech. The neural network is trained based on the noisy speech and the excitation signal of the denoised speech to output the second gain.
[0115] The excitation gain calculation can also be performed in other ways. For example, first, the voice frame to be processed is decomposed according to the glottis model to obtain the corresponding excitation signal. Then, both the excitation signal and the spectral coefficients of the voice frame to be processed are used as inputs to the neural network for processing. The neural network uses the glottis parameters obtained during the decomposition of the voice frame to be processed to simulate the denoised speech based on the excitation signal and the spectral coefficients of the voice frame to be processed, and then determines the second gain through the simulated denoised speech and the noisy speech.
[0116] Step S640: Perform compensation prediction based on the spectral coefficients of the voice frame to be processed to obtain a control coefficient, where the control coefficient is determined based on the energy of the spectral coefficients of the audio frame to be processed.
[0117] The control system is a scheme for energy compensation. During the enhancement of the voice frame to be processed, a part of the energy will be lost during denoising, which will affect the auditory effect of the obtained denoised result. Compensation prediction is performed based on the spectral coefficients of the voice frame to be processed to obtain a control coefficient. The control coefficient is usually a two-dimensional vector, representing the real part and the imaginary part of the spectral coefficients of the voice frame to be processed respectively. The control coefficient can directly act on the enhancement results of the first gain and the second gain to perform energy compensation.
[0118] The process of compensation prediction can be estimated based on the enhancement result and the voice frame to be processed. For example, by calculating the energy difference between the two, the control coefficient required for compensation is estimated. The process of compensation prediction can also be in the form of deep learning to learn the control coefficient required for the voice frame to be processed to restore to its original energy level after denoising. For example, the noisy voice frame is used as a training sample, and the control coefficient required after denoising is manually calculated as the training target, so as to obtain the corresponding model to predict the control coefficient.
[0119] Step S650: Perform gain control on the voice frame to be processed according to the first gain, the second gain, and the control coefficient to obtain the target voice frame.
[0120] Specifically, first, the frequency-domain representation of the voice frame to be processed can be enhanced according to the second gain, and then the obtained result is further gained according to the first gain to obtain the enhanced frequency-domain representation. Subsequently, energy compensation calculation is performed on the enhanced frequency representation according to the control coefficient to obtain the compensated frequency-domain representation. Then, inverse STFT is performed on the compensated frequency-domain representation to obtain the enhanced voice frame to be processed.
[0121] In an embodiment of the present application, a first gain and a second gain are respectively calculated for the glottal features and the excitation signal of the noisy speech signal, and then gain control is performed according to the first gain and the second gain, so as to denoise the noisy speech signal and perform energy compensation on the denoised speech signal. Performing noise reduction processing according to the glottal features can specifically identify the human voice part in the speech signal. Therefore, in the noise reduction process, the human voice part is also processed, and there is no need to train for various types of noise. Thus, the influence of the completeness of the training data is reduced, and noise types and noise environments not included in the training data can be effectively processed, improving the noise reduction effect.
[0122] In some embodiments of the present application, based on the above technical solution, for step S620 of calculating the glottal gain according to the spectral coefficients of the speech frame to be processed to obtain the first gain, the following steps may be included:
[0123] Input the spectral coefficients of the speech frame to be processed into a first neural network, and the first neural network is trained according to the glottal features corresponding to the noisy speech frame and the glottal features corresponding to the denoised speech frame corresponding to the noisy speech frame;
[0124] The first neural network performs gain prediction according to the spectral coefficients of the speech frame to be processed to obtain the first gain.
[0125] The first neural network may be a model constructed by a long short-term memory neural network, a convolutional neural network, a recurrent neural network, a fully connected neural network, etc., and specific limitations are not provided here.
[0126] During the training process, first, the noisy speech frames in the training data are decomposed to obtain the frequency response corresponding to the glottal filter in the glottal model. Then, training is performed according to the spectral coefficients of the noisy speech frames and the frequency response corresponding to the decomposed glottal filter. By adjusting the model parameters of the first neural network until the difference between the glottal features of the noisy speech frame and the glottal features of the denoised speech frame output by the first neural network meets the preset requirements. Among them, the preset requirements can be calculated by the mean square error. Through training, the mean square error between the glottal features of the noisy speech frame and the glottal features of the denoised speech frame output by the first gain is made to meet the set mean square error threshold, thereby determining that the trained model can achieve the expected purpose. Through this training process, the first gain predicted by the first neural network can make the glottal filter of the speech frame to be processed under the glottal model (i.e., glottal filter + excitation signal) sufficiently similar to the glottal filter of the clean speech under the glottal model, thus having the noise reduction ability.
[0127] The first neural network performs gain prediction according to the spectral coefficients to obtain the first gain.Figure 7 It is a schematic structural diagram of a first neural network shown according to a specific embodiment. As Figure 7 shown, the first neural network includes three fully connected (FC) layers. The input F(n) is a spectral coefficient of [257, 2] dimensions. The output of the first FC layer is a vector of [256, 1] dimensions, the output of the second FC layer is a vector of [128, 1] dimensions, and the output of the third FC layer is a vector of [32, 1] dimensions, that is, the first gain g1(n). Of course, Figure 7 it is merely an exemplary example of the structure of the first neural network and cannot be considered as a limitation on the scope of use of this application.
[0128] In an embodiment of this application, the first gain for the glottal features is obtained through a neural network, and the relationship between the glottal features and the first gain is learned through the neural network, so that a model can be obtained based on limited training data to process various situations in the actual scenario, improving the flexibility of the solution.
[0129] In an embodiment of this application, based on the above technical solution, before inputting the spectral coefficient of the speech frame to be processed into the first neural network, the speech processing method further includes:
[0130] Obtain the spectral coefficients of the historical speech frames of the speech frame to be processed;
[0131] The above step of inputting the spectral coefficient of the speech frame to be processed into the first neural network includes:
[0132] Input the spectral coefficient of the speech frame to be processed and the spectral coefficients of the historical speech frames into the first neural network.
[0133] The process of calculating the glottal gain can also incorporate the historical speech frames of the speech frame to be processed into the calculation process, facilitating the further accurate prediction of the required gain based on the correlation relationship between speech frames. The historical speech frames are other speech frames in the audio where the speech frame to be processed is located. For example, for the nth frame, the historical speech frames can be the (n - 1)th frame, the (n - 2)th frame, etc. Perform STFT transformation on the historical speech frames to obtain the spectral coefficients of the historical speech frames. Subsequently, when predicting the first gain, input the spectral coefficient of the speech frame to be processed and the spectral coefficients of the historical speech frames into the first neural network. Correspondingly, when training the first neural network, it is also necessary to input the spectral coefficients of the historical speech frames corresponding to the noisy speech frames in the training data as training samples into the neural network to be trained for training, so that the neural network can learn the correlation relationship between the speech frame to be processed and the historical speech frames and the corresponding gain time, facilitating the output of an appropriate first gain during the actual application process.
[0134] In an embodiment of the present application, the historical speech frames are used as inputs and processed together with the speech frames to be processed, so that the relationship between adjacent speech frames can be more fully considered during the feature extraction process, thereby improving the accuracy of feature extraction.
[0135] In an embodiment of the present application, based on the above technical solution, the above step of predicting the gain according to the spectral coefficients of the speech frames to be processed by the first neural network to obtain the first gain may include the following steps:
[0136] The first neural network calculates the gain of the speech frames to be processed according to the spectral coefficients of the speech frames to be processed, and obtains the first glottal gain corresponding to each sub-band in the speech frames to be processed, where the sub-band corresponds to at least one frequency band in the spectral coefficients of the speech frames to be processed;
[0137] The first glottal gains corresponding to each sub-band are combined as the first gain.
[0138] Specifically, the spectral response related to the glottal filter is a low-pass-like smoothing effect. Therefore, although the dimension of the frequency-domain representation of the speech frames to be processed is 257 dimensions, when calculating the first gain, it is not necessary to reach a resolution of 257 dimensions. Therefore, several adjacent coefficients can be combined during the calculation of the first gain, and a common first gain is used. Each sub-band includes the features in at least two adjacent dimensions of the spectral coefficients of the speech frames to be processed.
[0139] According to the frequency-domain representation of the speech frames to be processed, band division is performed along the frequency, and multiple sub-bands in the frequency-domain representation can be obtained. The band division performed on the frequency-domain representation can be uniform band division of the frequency (that is, the frequency width corresponding to each sub-band is the same), or non-uniform band division, which is not specifically limited herein. It can be understood that each sub-band corresponds to a frequency range, which includes multiple frequency points.
[0140] The non-uniform band division can be Bark band division. Bark band division is performed according to the Bark frequency scale. The Bark frequency scale maps the frequency to multiple critical frequency bands in psychoacoustics. The number of bands can be set according to the sampling rate and actual needs. For example, the number of frequency points is set to 24. Bark band division conforms to the characteristics of the auditory system. Generally, the lower the frequency, the fewer the number of coefficients included in the sub-band, or even a single coefficient, and the higher the frequency, the more the number of coefficients included in the sub-band.
[0141] In one embodiment, for 257 coefficients, every 8 adjacent coefficients are combined into a sub-band (the first element of the FFT coefficients is the DC component and can be ignored). Therefore, the dimension of the finally output first gain g1(n) is 32-dimensional. Through the first neural network, the first glottal gain corresponding to each sub-band can be output. The first glottal gains of each sub-band are combined to obtain the first gain. That is, 32 sub-bands correspond to the 32 dimensions of the first gain.
[0142] In the embodiments of the present application, the first gain is calculated according to the spectral coefficients of the speech frame to be processed by means of sub-band combination, so that the dimension of the calculation process can be reduced, the overall calculation amount of the scheme can be reduced, and the calculation efficiency can be improved.
[0143] In one embodiment of the present application, based on the above technical solution, the speech processing method may further include the following steps:
[0144] The spectral coefficients of the speech frame to be processed and the pitch period of the speech frame to be processed are predicted and analyzed through the first neural network to determine the second glottal gain, and the second glottal gain corresponds to the long-term correlation feature of the spectral coefficients of the speech frame to be processed;
[0145] Combining the first glottal gains corresponding to each sub-band as the first gain includes:
[0146] Combining the first glottal gains corresponding to each sub-band and the second glottal gain as the first gain.
[0147] In this embodiment, the first gain includes two parts, the first glottal feature corresponding to the short-term correlation feature of the speech frame and the second glottal feature corresponding to the long-term correlation feature of the speech frame. The pitch period of the speech frame to be processed can be obtained by pre-performing speech decomposition and analysis on the speech frame to be processed. The first neural network can directly output the second glottal gain according to the spectral coefficients and pitch period of the speech frame to be processed. The second glottal gain corresponds to the glottal parameters of the LTP filtering in the glottal filter. Therefore, during the training process, the model is trained based on the glottal parameters corresponding to the LTP filter obtained from the speech and the denoised speech decomposition. By adjusting the model parameters, the mean square error similarity between the finally output first gain and the first gain corresponding to the denoising result reaches the mean square error threshold, thereby completing the training. The first neural network can output the first glottal gain and the second glottal gain together. In one embodiment, the first neural network may be composed of two sub-networks, which are respectively used to output the first glottal feature and the second glottal feature.
[0148] In an embodiment of the present application, during the calculation of the first gain, the long-term correlation of the speech frame is further considered, making the solution more refined in recognizing the speech part in the speech frame, thereby avoiding the influence of the gain on the original speech and improving the accuracy of the solution.
[0149] In an embodiment of the present application, based on the above technical solution, the above step of predicting the gain based on the spectral coefficients of the speech frame to be processed by the first neural network to obtain the first gain may include the following steps:
[0150] Predicting, by the first neural network, a first glottal parameter based on the spectral coefficients of the speech frame to be processed, where the first glottal parameter is used to represent the short-term correlation characteristics of the spectral coefficients of the speech frame to be processed;
[0151] Predicting, by the first neural network, a first prediction result based on the first glottal parameter;
[0152] Determining the first gain according to the first prediction result.
[0153] In this embodiment, the first neural network will predict the first glottal parameter corresponding to the speech frame to be processed according to the spectral coefficients of the speech frame to be processed. The first glottal parameter is used to represent the short-term correlation characteristics of the spectral coefficients of the speech frame to be processed. Specifically, the first glottal parameter corresponds to an LPC filter. During the training process of the first neural network, by decomposing the denoised speech corresponding to the noisy speech in the pre-training data, the configuration parameters of the LPC filter of the denoised speech can be determined. According to the spectral coefficients of the noisy speech and the configuration parameters of the LPC filter of the denoised speech, the first neural network can be trained so that the mean square error between the first glottal parameter predicted by the first neural network and the configuration parameters of the LPC filter of the denoised speech meets the set accuracy requirement. Then, according to the obtained first glottal parameter, the first neural network can predict the spectrum of the denoised speech, and determine the first prediction result according to the predicted spectrum and the spectrum of the speech frame to be processed. The first prediction result is the first gain.
[0154] In this embodiment, by predicting the first glottal parameter corresponding to the speech frame to be processed through the first neural network and then predicting the first gain according to the first glottal parameter, the glottal parameter prediction target corresponding to the glottal filter simplifies the complexity of the training data compared with the annotation result of annotating each speech frame of the speech to be processed, thereby improving the calculation efficiency.
[0155] In an embodiment of the present application, based on the above technical solution, the speech processing method further includes:
[0156] The first neural network performs parameter prediction based on the spectral coefficients of the speech frame to be processed and the pitch period of the speech frame to be processed, and obtains second glottal parameters, where the second glottal parameters are used to represent the long-term correlation characteristics of the spectral coefficients of the speech frame to be processed;
[0157] The first neural network performs gain prediction based on the second glottal parameters to obtain a second prediction result;
[0158] Determining a first gain according to the first prediction result includes:
[0159] Combining and determining the first prediction result and the second prediction result as the first gain.
[0160] Specifically, the first neural network predicts second glottal parameters according to the spectral coefficients of the speech frame to be processed and the pitch period of the speech frame to be processed. The first glottal parameters are used to represent the long-term correlation characteristics of the spectral coefficients of the speech frame to be processed. Specifically, the second glottal parameters correspond to the LTP filter. In this embodiment, the glottal model of the speech frame further includes an LTP filter. The speech signal generated by the LPC filter configured according to the first glottal parameters is further processed by the LTP filter configured according to the second glottal parameters to simulate the speech in the speech frame to be processed. During the training process, by decomposing the denoised speech corresponding to the noisy speech in the pre-training data, the configuration parameters of the LTP filter of the denoised speech can be determined. According to the spectral coefficients of the noisy speech and the configuration parameters of the LTP filter of the denoised speech, the first neural network can be trained so that the mean square error between the second glottal parameters predicted by the first neural network and the configuration parameters of the LTP filter of the denoised speech meets the set accuracy requirements. Then, according to the obtained second glottal parameters, the first neural network can combine the first prediction result to predict the spectrum of the denoised speech, thereby obtaining a second prediction result. The second prediction result is also a part of the first gain, which makes the long-term correlation characteristics in the denoising result obtained based on the second prediction result similar to the long-term correlation characteristics in the denoised speech. Combining the first prediction result and the second prediction result can obtain the first gain. In the subsequent application process, the speech frame to be processed is enhanced according to the first prediction result and the second prediction result in sequence, so as to achieve the effect of noise reduction.
[0161] In the embodiment of the present application, by predicting the second glottal parameters, the long-term correlation of the speech frame is further considered in the prediction process of the first gain, so that the prediction of the voiced part in the speech frame is more accurate, thereby improving the accuracy of the solution.
[0162] In an embodiment of the present application, based on the above technical solution, in step S630, the excitation gain is calculated according to the spectral coefficients of the speech frame to be processed to obtain a second gain, including:
[0163] Input the spectral coefficients of the processed speech frames into a second neural network, which is trained based on the excitation signal of the noisy speech frames and the excitation signal of the corresponding denoised speech frames of the noisy speech frames;
[0164] The second neural network predicts a second gain based on the excitation signal of the speech frames to be processed, obtaining the second gain.
[0165] The second neural network refers to a neural network model used to predict the second gain corresponding to the excitation signal. The second neural network can be a model constructed by a long short-term memory neural network, a convolutional neural network, a recurrent neural network, a fully connected neural network, etc., and is not specifically limited herein.
[0166] During the training process, first decompose the speech frames with noise in the training data to obtain the frequency response corresponding to the excitation signal in the glottal model, and then perform training based on the spectral coefficients of the speech frames with noise and the frequency response of the obtained excitation signal. By adjusting the model parameters of the second neural network until the second gain output by the second neural network can make the difference between the excitation signal of the noisy speech frame and the excitation signal of the denoised speech frame meet the preset requirements. Among them, the preset requirement can be that the similarity between the excitation signal of the noisy speech frame and the excitation signal of the denoised speech frame is not lower than the similarity threshold. Through this training process, the second gain predicted by the second neural network can make the excitation signal of the speech frames to be processed under the glottal model be sufficiently similar to the excitation signal of the clean speech under the glottal model, thus having the ability to reduce noise.
[0167] The second neural network predicts a second gain based on the spectral coefficients of the speech frames to be processed, obtaining the second gain. Figure 8 is a schematic structural diagram of the second neural network shown according to a specific embodiment, as Figure 8 shown, the first neural network includes three fully connected (FC) layers. The input F(n) is a spectral coefficient of [257, 2] dimensions. The output of the first FC layer is a vector of [1024, 1] dimensions, the output of the second FC layer is a vector of [512, 1] dimensions, and the output of the third FC layer is a vector of [257, 1] dimensions, that is, the second gain g12(n). Of course, Figure 8 This is only an exemplary example of the structure of the second neural network and should not be considered as a limitation on the scope of use of this application.
[0168] In the embodiments of the present application, the second gain for the excitation signal is obtained through a neural network, and the relationship between the excitation signal and the second gain is learned through the neural network, so that the noisy speech can be denoised according to the glottal model without performing speech decomposition on the noisy speech, thereby saving computing resources.
[0169] In one embodiment of the present application, based on the above technical solution, step S650 of performing gain control on the speech frame to be processed according to the first gain, the second gain, and the control coefficient to obtain the target speech frame includes:
[0170] Enhancing the speech frame to be processed according to the second gain to obtain a first enhancement result;
[0171] Performing gain operation on each subband in the first enhancement result according to the first gain to obtain a second enhancement result;
[0172] Performing energy compensation on the second enhancement result according to the control coefficient to obtain a third enhancement result;
[0173] Performing inverse time-frequency conversion according to the third enhancement result to obtain the enhanced speech frame as the target speech frame.
[0174] Specifically, for the frequency-domain representation of the speech frame to be processed, first perform a multiplication operation on a sample-by-sample basis according to the corresponding parameters in the second gain to obtain the first enhancement result. As described above, the dimension of the second gain corresponds to the frequency-domain representation of the speech frame to be processed. That is, if the frequency-domain representation of the speech frame to be processed is 257-dimensional, then the second gain is also 257-dimensional. Therefore, when enhancing according to the second gain, the multiplication operation can be directly performed according to the corresponding relationship of the dimensions to obtain the first enhancement result. Based on the first enhancement result, perform a gain operation according to the first gain. Specifically, when calculating the first gain, the first gain is merged according to the subband division. Therefore, when calculating according to the first gain, the multiplication is also performed according to the corresponding relationship of the subband merging. For example, in the first gain result, every 8 dimensions correspond to one subband, and the first gain is a 32-dimensional variable. Then, when calculating according to the first gain, every 8 dimensions in the first gain result correspond to one dimension in the first gain for calculation, so as to obtain the second gain result. Performing energy compensation on the second enhancement result according to the control coefficient to obtain the third enhancement result. Specifically, the process of energy compensation is to directly sum the control coefficient and the second gain result to obtain the third enhancement result. The calculation formula is as follows:
[0175] S_e2(n) = S_e1(n) + g2(n)
[0176] where S_e2(n) is the third enhancement result, g2(n) is the control coefficient, and S_e1(n) is the second enhancement result. Performing an inverse STFT transformation according to the third enhancement result, that is, the frequency-domain representation can be transformed into a time-domain signal, so as to obtain the enhanced speech frame, that is, the target speech frame.
[0177] In an embodiment of the present application, a specific method for gain control is provided, which improves the feasibility of the solution.
[0178] The overall process of the voice processing method of the present application will be introduced below. For ease of introduction, please refer to Figure 9 , Figure 9 which is a schematic diagram of the overall process in the embodiment of the present application. As Figure 9 shown, the input of the solution is the voice frame s(n) to be processed. The voice frame s(n) is subjected to STFT time-frequency transformation to obtain the spectral coefficients S(n). Based on the spectral coefficients S(n), a first neural network is called to obtain a first gain g11(n), and a second neural network is called to obtain a second gain g12(n). The first gain g11(n) and the second gain g12(n) are jointly used to perform a first spectral control (i.e., gain control) on the spectral coefficients S(n), so as to output a first enhanced spectrum S_e1(n). The first spectral control is mainly used to suppress noise in the voice frame. In particular, the inputs of the first neural network and the second neural network may also include the spectral coefficients S_pre(n) of the historical frames (such as the n-1, n-2 frames, etc., and so on) of the voice frame s(n). The processing process of the third neural network can be executed in parallel with the processes of the first neural network and the second neural network. Based on the spectral coefficients S(n), a third neural network is called to obtain a control coefficient g2(n). Applying the control coefficient g2(n) to the first enhanced spectrum S_e1(n) can obtain a second enhanced spectrum S_e2(n). The second spectral control is mainly used for energy compensation. In particular, the input of the third neural network may also include the spectral coefficients S_pre(n) of the historical frames of the voice frame s(n). Finally, inverse time-frequency transformation is performed according to the second enhanced spectrum S_e2(n) to obtain the enhanced and supplemented signal s_e(n) corresponding to the voice frame.
[0179] In an embodiment of the present application, based on the above technical solution, Figure 9 the parallel process shown in
[0180] can also be executed serially. In this regard, step S640 above, performing compensation prediction according to the spectral coefficients of the voice frame to be processed to obtain a control coefficient, includes:
[0181] Performing gain control on the voice frame to be processed according to the first gain and the second gain to obtain a gain control result;
[0182] Performing compensation prediction according to the gain control result and the spectral coefficients of the voice frame to be processed to obtain a control coefficient. Figure 9The calculation processes of the first neural network and the second neural network shown in Figure 10 are executed serially with the calculation process of the third neural network. Specifically, first, according to the first gain and the second gain, gain control is performed on the speech frame to be processed to obtain a gain control result, and then, based on the gain control result and the spectral coefficients of the speech frame to be processed, compensation prediction is performed to obtain a control coefficient. Specifically, for the convenience of introduction, please refer to Figure 10 and Figure 10 which is a schematic diagram of the overall process of another solution in the embodiments of the present application. As Figure 10 shown, the input of the solution is the speech frame s(n) to be processed. The speech frame s(n) is subjected to STFT time-frequency transformation to obtain spectral coefficients S(n). Based on the spectral coefficients S(n), the first neural network is called to obtain the first gain g11(n), and the second neural network is called to obtain the second gain g12(n). The first gain g11(n) and the second gain g12(n) are jointly used to perform the first spectral control (i.e., gain control) on the spectral coefficients S(n), so as to output the first enhanced spectrum S_e1(n). Then, the calculation of the third neural network is performed based on the first enhanced spectrum S_e1(n). Specifically, based on S(n) and S_e1(n), the third neural network is called to obtain the control coefficient, g2(n), of the speech frame. The control coefficient g2(n) is applied to the first enhanced spectrum S_e1(n) to obtain the second enhanced spectrum, S_e2(n). The second spectral control is mainly used for energy compensation. Subsequently, inverse time-frequency conversion is performed on the second enhanced spectrum S_e2(n) to obtain the enhanced and supplemented signal second enhanced spectrum s_e(n) corresponding to the speech frame.
[0183] It can be understood that in the solution of the present application, the input of the third neural network includes the gain result of gain control according to the output results of the first neural network and the second neural network. Therefore, the training process of the third neural network can be jointly executed with the training processes of the first neural network and the second neural network, and the model parameters of the three neural networks can be adjusted according to the final result. The training process of the third neural network can also be executed separately. The training data used can be generated by the trained first neural network and second neural network according to the noisy speech frames in the training data, or can be directly generated through an additional data preparation process, such as obtained by manual calculation.
[0184] In the embodiments of the present application, the compensation prediction process is based on the result of gain control, so that the compensation process can perform corresponding compensation estimation according to the result of gain control, so that the obtained control coefficient is more in line with the actually required compensation, and the accuracy and compensation effect of the control coefficient are improved.
[0185] In an embodiment of the present application, based on the above technical solution, the gain control and compensation control of the present application can also be performed based on the amplitude spectrum of the speech frame to be processed. In this regard, the speech processing method further includes:
[0186] Calculate the amplitude spectrum and phase spectrum corresponding to the speech frame to be processed according to the speech frame to be processed;
[0187] The above step S650, performing gain control on the speech frame to be processed according to the first gain, the second gain, and the control coefficient to obtain the target speech frame, includes:
[0188] Perform gain control on the amplitude spectrum corresponding to the speech frame to be processed according to the first gain and the second gain to obtain the enhanced amplitude spectrum;
[0189] Perform energy compensation according to the enhanced amplitude spectrum and the control coefficient to obtain the compensated amplitude spectrum;
[0190] Perform inverse time-frequency conversion according to the compensated amplitude spectrum and the phase spectrum corresponding to the speech frame to be processed to obtain the target speech frame.
[0191] Specifically, performing time-frequency transformation on the speech frame to be processed can obtain the amplitude spectrum and phase spectrum corresponding to the speech frame to be processed. Subsequently, perform gain control on the amplitude spectrum corresponding to the speech frame to be processed according to the first gain and the second gain to obtain the enhanced amplitude spectrum. In one embodiment, the calculation processes of the first gain and the second gain can also be calculated according to the amplitude spectrum. Specifically, the calculation and processing based on the spectral coefficients of the speech frame to be processed in the above embodiment can be replaced with those based on the amplitude spectrum. It can be understood that in the scheme of performing gain control according to the first neural network and the second neural network, these two neural networks will also use the amplitude spectrum for training during the training process. Then, perform energy compensation according to the enhanced amplitude spectrum and the control coefficient to obtain the compensated amplitude spectrum. Reuse the phase spectrum of the speech frame to be processed, generate the second enhanced spectrum together with the compensated amplitude spectrum, and then perform inverse time-frequency conversion according to the second enhanced spectrum to obtain the target speech frame.
[0192] In one embodiment, the energy compensation can be performed based on spectral coefficients, that is, calculate the first enhanced spectrum by calculating the enhanced amplitude spectrum and the phase spectrum corresponding to the speech frame to be processed. Subsequently, perform subsequent energy compensation according to the first enhanced spectrum.
[0193] In the solution of the present application, the calculation processes of performing gain control and energy compensation according to the amplitude spectrum provide another specific implementation manner for the solution of the present application, improving the diversity of the solution.
[0194] In an embodiment of the present application, based on the above technical solution, step S640 of obtaining a control coefficient by performing compensation prediction according to the spectral coefficients of the speech frame to be processed may include the following steps:
[0195] Obtain the spectral coefficients of the historical speech frames of the speech frame to be processed;
[0196] Input the spectral coefficients of the speech frame to be processed and the spectral coefficients of the historical speech frames into a third neural network, where the third neural network is trained according to the energy of the spectral coefficients corresponding to the noisy speech frames and the energy of the spectral coefficients corresponding to the denoised speech frames corresponding to the noisy speech frames;
[0197] Perform compensation prediction according to the spectral coefficients of the speech frame to be processed through the third neural network to obtain a control coefficient.
[0198] The third neural network refers to a neural network model used for compensation prediction. The third neural network can be a model constructed by a long short-term memory neural network, a convolutional neural network, a recurrent neural network, a fully connected neural network, etc., and is not specifically limited herein.
[0199] The input of the third neural network is the spectral coefficients of the speech frame to be processed, and the output is the control coefficient corresponding to the speech frame to be processed. The control coefficient is a group of two-dimensional vectors, respectively representing the real part and the imaginary part of the spectral coefficients; among them, the control coefficient will act on the enhanced result according to the first gain and the second gain to obtain a second enhanced spectrum for energy compensation.
[0200] The training process of the third neural network can be carried out independently or together with the processes of the first neural network and the second neural network. During training, input the spectral coefficients of the noisy speech frames and the spectral coefficients of the denoised speech frames corresponding to the noisy speech frames in the training data into the third neural network for prediction to obtain the predicted control coefficient of the output. By adjusting the model parameters of the third neural network, the difference between the energy of the compensated target speech frame and the energy of the unprocessed noisy speech frame can meet the preset requirements. In this embodiment, input the spectral coefficients of the speech frame to be processed and the spectral coefficients of the historical speech frames into the third neural network. Therefore, during the training process, the spectral coefficients of the historical speech frames are also input into the third neural network for training. Perform compensation prediction according to the spectral coefficients of the speech frame to be processed through the third neural network to obtain a control coefficient.
[0201] Figure 11 It is a schematic structural diagram of the third neural network shown according to a specific embodiment, as Figure 11As shown, the preprocessing neural network includes six convolutional layers and one Long Short-Term Memory (LSTM) layer. The input S(n) is represented by spectral coefficients and is thus a [2, 257]-dimensional spectral coefficient. Figure 11 For each convolutional layer and LSTM layer in Figure 11 , the dimensions of the variables output by that layer are marked. The first convolutional layer outputs a [16, 127]-dimensional variable, the second convolutional layer outputs a [32, 62]-dimensional variable, the third convolutional layer outputs a [64, 29]-dimensional variable, the fourth convolutional layer outputs a [128, 13]-dimensional variable, the fifth convolutional layer outputs a [128, 5]-dimensional variable, the sixth convolutional layer outputs a [128, 1]-dimensional variable, and the LSTM layer outputs a [2, 257]-dimensional variable. The variable output by the LSTM increase is the control coefficient g2(n), which is a [2, 257]-dimensional vector. It should be understood that Figure 11 This is merely an exemplary example of the structure of the preprocessing neural network and should not be considered as a limitation on the scope of use of this application. In another embodiment, the input of the third neural network can be S(n) and the second enhanced result S_e1(n), and other structures are the same as those above and will not be elaborated here.
[0202] In the embodiments of this application, historical speech frames are used as inputs and processed together with the speech frames to be processed, so that the relationship between adjacent speech frames can be more fully considered during feature extraction, thereby improving the accuracy of feature extraction.
[0203] It should be noted that although the steps of the method in this application are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution, etc.
[0204] The apparatus embodiments of this application are introduced below and can be used to execute the speech processing method in the above embodiments of this application. Figure 12 The block diagram of the speech processing apparatus in the embodiments of this application is schematically shown. As Figure 12 shown, the speech processing apparatus 1200 mainly may include:
[0205] A spectral coefficient acquisition module 1210, configured to acquire the spectral coefficients of the speech frames to be processed;
[0206] A glottal gain module 1220, configured to calculate the glottal gain based on the spectral coefficients of the speech frames to be processed to obtain a first gain, where the first gain corresponds to the glottal characteristics of the speech frames to be processed;
[0207] An excitation gain module 1230, configured to calculate an excitation gain based on the spectral coefficients of the to-be-processed speech frame to obtain a second gain, where the second gain corresponds to the excitation signal of the to-be-processed speech frame;
[0208] A compensation prediction module 1240, configured to perform compensation prediction based on the spectral coefficients of the to-be-processed speech frame to obtain a control coefficient, where the control coefficient is determined according to the energy of the spectral coefficients of the to-be-processed audio frame;
[0209] A gain control module 1250, configured to perform gain control on the to-be-processed speech frame according to the first gain, the second gain, and the control coefficient to obtain a target speech frame.
[0210] In some embodiments of the present application, based on the above technical solutions, the glottal gain module 1220 includes:
[0211] A first input sub-unit, configured to input the spectral coefficients of the to-be-processed speech frame into a first neural network, where the first neural network is trained according to the glottal features corresponding to the noisy speech frame and the glottal features corresponding to the denoised speech frame corresponding to the noisy speech frame;
[0212] A gain prediction sub-module, configured to perform gain prediction based on the spectral coefficients of the to-be-processed speech frame through the first neural network to obtain the first gain.
[0213] In some embodiments of the present application, based on the above technical solutions, the speech processing device 1200 further includes:
[0214] A historical frame spectral coefficient acquisition module, configured to acquire the spectral coefficients of the historical speech frames of the to-be-processed speech frame;
[0215] The first input sub-module includes: a historical frame input unit, configured to input the spectral coefficients of the to-be-processed speech frame and the spectral coefficients of the historical speech frames into the first neural network.
[0216] In some embodiments of the present application, based on the above technical solutions, the glottal gain module 1220 includes:
[0217] A first gain calculation sub-module, configured to perform gain calculation on the spectral coefficients of the to-be-processed speech frame through the first neural network to obtain the first glottal gain corresponding to each sub-band in the to-be-processed speech frame, where the sub-band corresponds to at least one frequency band in the spectral coefficients of the to-be-processed speech frame;
[0218] A sub-band gain merging sub-module, configured to merge the first glottal gains corresponding to the respective sub-bands as the first gain.
[0219] In some embodiments of the present application, based on the above technical solutions, the voice processing device 1200 further includes:
[0220] A second gain calculation sub-module, configured to perform predictive analysis on the spectral coefficients and the pitch period of the to-be-processed speech frame through the first neural network to determine a second glottal gain, where the second glottal gain corresponds to the long-term correlation feature of the spectral coefficients of the to-be-processed speech frame;
[0221] The sub-band gain merging sub-module includes:
[0222] A glottal gain merging unit, configured to merge the first glottal gain corresponding to each sub-band and the second glottal gain as the first gain.
[0223] In some embodiments of the present application, based on the above technical solutions, the first gain calculation sub-module includes:
[0224] A first glottal parameter prediction unit, configured to perform parameter prediction on the spectral coefficients of the to-be-processed speech frame through the first neural network to obtain a first glottal parameter, where the first glottal parameter is used to represent the short-term correlation feature of the spectral coefficients of the to-be-processed speech frame;
[0225] A first prediction unit, configured to perform gain prediction based on the first glottal parameter through the first neural network to obtain the first prediction result;
[0226] A result determination unit, configured to determine the first gain according to the first prediction result.
[0227] In some embodiments of the present application, based on the above technical solutions, the voice processing device 1200 further includes:
[0228] A second glottal parameter prediction unit, configured to perform parameter prediction on the spectral coefficients and the pitch period of the to-be-processed speech frame through the first neural network to obtain a second glottal parameter, where the first glottal parameter is used to represent the long-term correlation feature of the spectral coefficients of the to-be-processed speech frame;
[0229] A second prediction unit, configured to perform gain prediction based on the second glottal parameter through the first neural network to obtain a second prediction result;
[0230] The result determination unit includes:
[0231] A result merging sub-unit, configured to merge and determine the first prediction result and the second prediction result as the first gain.
[0232] In some embodiments of the present application, based on the above technical solutions, the excitation gain module 1230 includes:
[0233] A second input sub-module, configured to input the spectral coefficients of the to-be-processed speech frame into a second neural network, where the second neural network is trained according to the excitation signal of the noise speech frame and the excitation signal of the denoised speech frame corresponding to the noise speech frame;
[0234] A speech decomposition sub-module, configured to perform speech decomposition on the spectral coefficients of the to-be-processed speech frame through the second neural network to obtain an excitation signal;
[0235] A gain prediction sub-module, configured to perform gain prediction according to the excitation signal through the second neural network to obtain the second gain.
[0236] In some embodiments of the present application, based on the above technical solution, the gain control module 1250 includes:
[0237] A first enhancement sub-module, configured to enhance the to-be-processed speech frame according to the second gain to obtain a first enhancement result;
[0238] A second enhancement sub-module, configured to perform gain calculation on each sub-band in the first enhancement result according to the first gain to obtain a second enhancement result;
[0239] An energy compensation sub-module, configured to perform energy compensation on the second enhancement result according to the control coefficient to obtain a third enhancement result;
[0240] An inverse time-frequency conversion sub-module, configured to perform inverse time-frequency conversion according to the third enhancement result to obtain an enhanced speech frame as the target speech frame.
[0241] In some embodiments of the present application, based on the above technical solution, the compensation prediction module 1240 includes:
[0242] A gain control sub-module, configured to perform gain control on the to-be-processed speech frame according to the first gain and the second gain to obtain a gain control result;
[0243] A control coefficient prediction sub-module, configured to perform compensation prediction according to the gain control result and the spectral coefficients of the to-be-processed speech frame to obtain the control coefficient.
[0244] In some embodiments of the present application, based on the above technical solution, the speech processing device 1200 further includes:
[0245] An amplitude spectrum calculation module, configured to calculate the amplitude spectrum and phase spectrum corresponding to the to-be-processed speech frame according to the to-be-processed speech frame;
[0246] The gain control module 1250 includes:
[0247] The amplitude spectrum gain control sub-module is used to perform gain control on the amplitude spectrum corresponding to the to-be-processed speech frame according to the first gain and the second gain, so as to obtain an enhanced amplitude spectrum;
[0248] The amplitude spectrum energy compensation sub-module is used to perform energy compensation according to the enhanced amplitude spectrum and the control coefficient, so as to obtain a compensated amplitude spectrum;
[0249] The amplitude spectrum inverse time-frequency conversion sub-module is used to perform inverse time-frequency conversion according to the compensated amplitude spectrum and the phase spectrum corresponding to the to-be-processed speech frame, so as to obtain a target speech frame.
[0250] In some embodiments of the present application, based on the above technical solutions, the compensation prediction module 1240 includes:
[0251] The historical spectrum coefficient acquisition sub-module is used to acquire the spectrum coefficients of the historical speech frames of the to-be-processed speech frame;
[0252] The third input sub-module is used to input the spectrum coefficients of the to-be-processed speech frame and the spectrum coefficients of the historical speech frames into a third neural network, and the third neural network is trained according to the energy of the spectrum coefficients corresponding to the noise speech frames and the energy of the spectrum coefficients corresponding to the denoised speech frames corresponding to the noise speech frames;
[0253] The compensation prediction sub-module is used to perform compensation prediction according to the spectrum coefficients of the to-be-processed speech frame through the third neural network to obtain the control coefficient
[0254] It should be noted that the device provided in the above embodiment and the method provided in the above embodiment belong to the same concept. The specific manners in which each module performs operations have been described in detail in the method embodiment, and will not be repeated here.
[0255] Figure 13 The structural schematic diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown.
[0256] It should be noted that Figure 13 The computer system 1300 of the electronic device shown is only an example, and should not bring any limitation to the functions and usage scopes of the embodiments of the present application.
[0257] Such as Figure 13As shown, the computer system 1300 includes a Central Processing Unit (CPU) 1301, which can perform various appropriate actions and processes according to the program stored in the Read-Only Memory (ROM) 1302 or the program loaded from the storage section 1308 into the Random Access Memory (RAM) 1303. In the RAM 1303, various programs and data required for system operation are also stored. The CPU 1301, ROM 1302, and RAM 1303 are connected to each other via a bus 1304. An Input / Output (I / O) interface 1305 is also connected to the bus 1304.
[0258] The following components are connected to the I / O interface 1305: an input section 1306 including a keyboard, a mouse, etc.; an output section 1307 including, for example, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc. and a speaker, etc.; a storage section 1308 including a hard disk, etc.; and a communication section 1309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1309 performs communication processing via a network such as the Internet. A drive 1310 is also connected to the I / O interface 1305 as needed. A removable medium 1311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1310 as needed so that a computer program read from it can be installed into the storage section 1308 as needed.
[0259] Specifically, according to the embodiments of the present application, the processes described in each method flowchart can be implemented as computer software programs. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 1309, and / or installed from the removable medium 1311. When the computer program is executed by the Central Processing Unit (CPU) 1301, various functions defined in the system of the present application are executed.
[0260] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0261] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0262] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0263] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on the network, and includes several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0264] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.
[0265] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A voice processing method, characterized in that, Including: Obtain the spectral coefficients of the speech frame to be processed; Input the spectral coefficients of the speech frame to be processed into a first neural network, where the first neural network is trained based on the glottal features corresponding to the noisy speech frame and the glottal features corresponding to the denoised speech frame corresponding to the noisy speech frame; Perform gain calculation on the spectral coefficients of the speech frame to be processed through the first neural network to obtain the first glottal gain corresponding to each sub-band in the speech frame to be processed, where the sub-band corresponds to at least one frequency band in the spectral coefficients of the speech frame to be processed; Merge the first glottal gains corresponding to the respective sub-bands as a first gain, where the first gain corresponds to the glottal features of the speech frame to be processed; Perform excitation gain calculation based on the spectral coefficients of the speech frame to be processed to obtain a second gain, where the second gain corresponds to the excitation signal of the speech frame to be processed; Perform compensation prediction based on the spectral coefficients of the speech frame to be processed to obtain a control coefficient, where the control coefficient is determined based on the energy of the spectral coefficients of the speech frame to be processed; Perform gain control on the speech frame to be processed based on the first gain, the second gain, and the control coefficient to obtain a target speech frame.
2. The method according to claim 1, characterized in that, Before inputting the spectral coefficients of the speech frame to be processed into the first neural network, the method further includes: Obtain the spectral coefficients of the historical speech frames of the speech frame to be processed; The inputting the spectral coefficients of the speech frame to be processed into the first neural network includes: Input the spectral coefficients of the speech frame to be processed and the spectral coefficients of the historical speech frames into the first neural network.
3. The method according to claim 1, characterized in that, The method further includes: Perform prediction analysis on the spectral coefficients of the speech frame to be processed and the pitch period of the speech frame to be processed through the first neural network to determine a second glottal gain, where the second glottal gain corresponds to the long-term correlation feature of the spectral coefficients of the speech frame to be processed; The merging the first glottal gains corresponding to the respective sub-bands as the first gain includes: Merge the first glottal gains corresponding to the respective sub-bands and the second glottal gain as the first gain.
4. The method according to claim 1, wherein The obtaining the first gain by performing gain prediction on the spectral coefficients of the speech frame to be processed through the first neural network includes: Perform parameter prediction on the spectral coefficients of the speech frame to be processed through the first neural network to obtain a first glottal parameter, where the first glottal parameter is used to represent the short-term correlation feature of the spectral coefficients of the speech frame to be processed; Perform gain prediction based on the first glottal parameter through the first neural network to obtain a first prediction result; Determine the first gain according to the first prediction result.
5. The method according to claim 4, wherein The method further includes: Perform parameter prediction on the spectral coefficients of the speech frame to be processed and the pitch period of the speech frame to be processed through the first neural network to obtain a second glottal parameter, where the first glottal parameter is used to represent the long-term correlation feature of the spectral coefficients of the speech frame to be processed; Perform gain prediction based on the second glottal parameter through the first neural network to obtain a second prediction result; Determining the first gain according to the first prediction result includes: Combining the first prediction result and the second prediction result to determine the first gain.
6. The method according to claim 1, characterized in that Calculating an excitation gain according to the spectral coefficients of the to-be-processed speech frame to obtain a second gain, including: Inputting the spectral coefficients of the to-be-processed speech frame into a second neural network, where the second neural network is trained according to the excitation signal of a noisy speech frame and the excitation signal of the denoised speech frame corresponding to the noisy speech frame; Performing speech decomposition on the spectral coefficients of the to-be-processed speech frame through the second neural network to obtain an excitation signal; Performing gain prediction according to the excitation signal through the second neural network to obtain the second gain.
7. The method according to claim 1, wherein Performing gain control on the to-be-processed speech frame according to the first gain, the second gain, and the control coefficient to obtain a target speech frame, including: Enhancing the to-be-processed speech frame according to the second gain to obtain a first enhancement result; Performing a gain operation on each subband in the first enhancement result according to the first gain to obtain a second enhancement result; Performing energy compensation on the second enhancement result according to the control coefficient to obtain a third enhancement result; Performing inverse time-frequency conversion according to the third enhancement result to obtain the enhanced speech frame as the target speech frame.
8. The method according to claim 1, characterized in that, Performing compensation prediction according to the spectral coefficients of the to-be-processed speech frame to obtain a control coefficient, including: Performing gain control on the to-be-processed speech frame according to the first gain and the second gain to obtain a gain control result; Performing compensation prediction according to the gain control result and the spectral coefficients of the to-be-processed speech frame to obtain the control coefficient.
9. The method according to claim 1, wherein The method further includes: Calculating the magnitude spectrum and phase spectrum corresponding to the to-be-processed speech frame according to the to-be-processed speech frame; Performing gain control on the to-be-processed speech frame according to the first gain, the second gain, and the control coefficient to obtain a target speech frame, including: Performing gain control on the magnitude spectrum corresponding to the to-be-processed speech frame according to the first gain and the second gain to obtain an enhanced magnitude spectrum; Performing energy compensation according to the enhanced magnitude spectrum and the control coefficient to obtain a compensated magnitude spectrum; Performing inverse time-frequency conversion according to the compensated magnitude spectrum and the phase spectrum corresponding to the to-be-processed speech frame to obtain the target speech frame.
10. The method according to claim 1, wherein Performing compensation prediction according to the spectral coefficients of the to-be-processed speech frame to obtain a control coefficient, including: Obtaining the spectral coefficients of the historical speech frame of the to-be-processed speech frame; Inputting the spectral coefficients of the to-be-processed speech frame and the spectral coefficients of the historical speech frame into a third neural network, where the third neural network is trained according to the energy of the spectral coefficients corresponding to the noisy speech frame and the energy of the spectral coefficients corresponding to the denoised speech frame corresponding to the noisy speech frame; Performing compensation prediction according to the spectral coefficients of the to-be-processed speech frame through the third neural network to obtain the control coefficient.
11. A voice processing device, characterized in that, Includes: A spectral coefficient acquisition module for acquiring the spectral coefficients of the to-be-processed speech frame; A first input sub-unit, configured to input spectral coefficients of the to-be-processed speech frame into a first neural network, where the first neural network is trained according to glottal features corresponding to a noise speech frame and glottal features corresponding to a denoised speech frame corresponding to the noise speech frame; A first gain calculation sub-module, configured to perform gain calculation on the spectral coefficients of the to-be-processed speech frame through the first neural network to obtain first glottal gains corresponding to each sub-band in the to-be-processed speech frame, where the sub-band corresponds to at least one frequency band in the spectral coefficients of the to-be-processed speech frame; A sub-band gain merging sub-module, configured to merge the first glottal gains corresponding to the respective sub-bands as a first gain; An excitation gain module, configured to perform excitation gain calculation according to the spectral coefficients of the to-be-processed speech frame to obtain a second gain, where the second gain corresponds to an excitation signal of the to-be-processed speech frame; A compensation prediction module, configured to perform compensation prediction according to the spectral coefficients of the to-be-processed speech frame to obtain a control coefficient, where the control coefficient is determined according to the energy of the spectral coefficients of the to-be-processed speech frame; A gain control module, configured to perform gain control on the to-be-processed speech frame according to the first gain, the second gain, and the control coefficient to obtain a target speech frame.
12. An electronic device, characterized in that, Comprising: A processor; A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the speech processing method according to any one of claims 1 to 10 by executing the executable instructions.
13. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the speech processing method according to any one of claims 1 to 10.
14. A computer program product, characterized in that, The computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the speech processing method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Voice gain control method and computer storage medium
CN112242147A