A voice processing method, device, electronic device and readable medium

By calculating the audio feature vector of the voice signal and performing gain control, the problem of poor noise type and environment processing in the prior art is solved, and a more efficient voice signal denoising effect is achieved.

CN114333892BActive Publication Date: 2025-06-24TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111237543.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-22
Publication Date
2025-06-24
Estimated Expiration
2041-10-22

AI Technical Summary

Technical Problem

When processing noise-containing voice signals, the prior art needs to collect training data for various types of noise, resulting in the model processing effect being affected by the completeness of the training data, especially when facing the noise type and environment that does not include noise, the noise reduction effect is poor.

Method used

By obtaining the spectral coefficients of the to-process speech frame, calculating the audio feature vector, and calculating the glottal gain and excitation gain based on this, and gain control is performed in combination with the control coefficient to achieve denoising of the noisy voice signal.

Benefits of technology

This method can effectively reduce the impact of training data completeness, improve processing capabilities for unincluded noise types and environments, and improve noise reduction effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114333892B_ABST
    Figure CN114333892B_ABST
Patent Text Reader

Abstract

The present application provides a voice processing method, apparatus, electronic device and readable medium. The method includes: obtaining spectral coefficients of a voice frame to be processed; performing feature calculation according to the spectral coefficients of the voice frame to be processed to obtain an audio feature vector of the voice frame to be processed; performing glottal gain calculation according to the audio feature vector to obtain a first gain, the first gain corresponding to the glottal feature of the voice frame to be processed; performing excitation gain calculation according to the audio feature vector to obtain a second gain, the second gain corresponding to the excitation signal of the voice frame to be processed; performing compensation prediction according to the audio feature vector to obtain a control coefficient, the control coefficient being determined according to the energy of the spectral coefficients of the voice frame to be processed; and performing gain control on the voice frame to be processed according to the first gain, the second gain and the control coefficient to obtain a target voice frame. This method can effectively process noise types and noise environments not included in the training data and improve the noise reduction effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular, to a voice processing method, apparatus, electronic device, and readable medium. Background Art

[0002] With the development of computer technology, various voice communication or voice control technologies have emerged. Through such technologies, users are allowed to communicate over long distances or the efficiency of human-computer interaction can be improved. In a real environment, when a user is in the surrounding environment, various environmental noises will be collected by devices such as microphones, and the quality of voice communication will be affected to varying degrees. Therefore, voice enhancement has become an important topic.

[0003] In related technologies, a deep learning method is used to learn the signal features of noisy speech audio, so as to predict the proportion of the voice component and the noise component, and then enhance the noisy speech according to the prediction result to achieve the noise reduction effect.

[0004] However, in the above solution, it is necessary to collect training data for various noises to train the model, so that the trained model can process the noise types covered in the training data. Therefore, the processing effect of the model is affected by the completeness of the training data, and the noise reduction effect is poor when facing situations not covered in the training data. Summary of the Invention

[0005] Based on the above technical problems, the present application provides a voice processing method, apparatus, electronic device, and readable medium to reduce the influence of the completeness of training data, and can effectively process noise types and noise environments not included in the training data, and improve the noise reduction effect.

[0006] Other features and advantages of the present application will become apparent through the following detailed description, or be learned in part through the practice of the present application.

[0007] According to one aspect of an embodiment of the present application, a voice processing method is provided, including:

[0008] Obtain the spectral coefficients of the voice frame to be processed;

[0009] Perform feature calculation according to the spectral coefficients of the voice frame to be processed to obtain the audio feature vector of the voice frame to be processed;

[0010] Perform glottal gain calculation according to the audio feature vector to obtain a first gain, where the first gain corresponds to the glottal feature of the voice frame to be processed;

[0011] Perform excitation gain calculation according to the audio feature vector to obtain a second gain, where the second gain corresponds to the excitation signal of the voice frame to be processed;

[0012] Perform compensation prediction based on the audio feature vector to obtain a control coefficient, where the control coefficient is determined according to the energy of the spectral coefficients of the audio frame to be processed;

[0013] Perform gain control on the speech frame to be processed according to the first gain, the second gain, and the control coefficient to obtain a target speech frame.

[0014] According to one aspect of the embodiments of the present application, a speech processing device is provided, including:

[0015] A spectral coefficient acquisition module, configured to acquire the spectral coefficients of a speech frame to be processed;

[0016] A vector acquisition module, configured to perform feature calculation according to the spectral coefficients of the speech frame to be processed to obtain an audio feature vector of the speech frame to be processed.

[0017] A glottal gain module, configured to perform glottal gain calculation according to the audio feature vector to obtain a first gain, where the first gain corresponds to the glottal feature of the speech frame to be processed;

[0018] An excitation gain module, configured to perform excitation gain calculation according to the audio feature vector to obtain a second gain, where the second gain corresponds to the excitation signal of the speech frame to be processed;

[0019] A compensation prediction module, configured to perform compensation prediction according to the audio feature vector to obtain a control coefficient, where the control coefficient is determined according to the energy of the spectral coefficients of the audio frame to be processed;

[0020] A gain control module, configured to perform gain control on the speech frame to be processed according to the first gain, the second gain, and the control coefficient to obtain a target speech frame.

[0021] In some embodiments of the present application, based on the above technical solution, the glottal gain module includes:

[0022] A first neural network sub-module, configured to input the audio feature vector into a first neural network, where the first neural network is trained according to the glottal features corresponding to the noisy speech frames and the glottal features corresponding to the denoised speech frames corresponding to the noisy speech frames;

[0023] A glottal gain prediction sub-module, configured to perform gain prediction according to the audio feature vector through the first neural network to obtain the first gain.

[0024] In some embodiments of the present application, based on the above technical solution, the glottal gain prediction sub-module includes:

[0025] A gain calculation unit, configured to calculate a gain for the audio feature vector through the first neural network to obtain a first glottal gain corresponding to each sub-band in the to-be-processed speech frame, where the sub-band corresponds to at least one frequency band in the to-be-processed speech frame;

[0026] A gain generation unit, configured to combine the first glottal gains corresponding to the respective sub-bands as the first gain.

[0027] In some embodiments of the present application, based on the above technical solutions, the speech processing device includes:

[0028] A gain analysis unit, configured to perform predictive analysis on the audio feature vector and the fundamental period of the to-be-processed speech frame through the first neural network to determine a second glottal gain, where the second glottal gain corresponds to the long-term correlation feature of the audio feature vector;

[0029] The gain generation unit includes:

[0030] A gain merging subunit, configured to combine the first glottal gains corresponding to the respective sub-bands and the second glottal gain as the first gain.

[0031] In some embodiments of the present application, based on the above technical solutions, the glottal gain prediction sub-module includes:

[0032] A first parameter prediction unit, configured to perform parameter prediction according to the audio feature vector through the first neural network to obtain a first glottal parameter, where the first glottal parameter is used to represent the short-term correlation feature of the audio feature vector;

[0033] A first gain prediction unit, configured to perform gain prediction according to the first glottal parameter through the first neural network to obtain the first prediction result;

[0034] A gain determination unit, configured to determine the first gain according to the first prediction result.

[0035] In some embodiments of the present application, based on the above technical solutions, the glottal gain prediction sub-module further includes:

[0036] A second parameter prediction unit, configured to perform parameter prediction according to the audio feature vector and the fundamental period of the to-be-processed speech frame through the first neural network to obtain a second glottal parameter, where the second glottal parameter is used to represent the long-term correlation feature of the audio feature vector;

[0037] A second gain prediction unit, configured to perform gain prediction according to the second glottal parameter through the first neural network to obtain a second prediction result;

[0038] The gain determination unit includes:

[0039] A prediction result merging sub-unit, configured to merge the first prediction result and the second prediction result to determine the first gain.

[0040] In some embodiments of the present application, based on the above technical solution, the excitation gain module includes:

[0041] A second neural network sub-module, configured to input the audio feature vector into a second neural network, where the second neural network is trained according to the excitation signal of the noisy speech frame and the excitation signal of the denoised speech frame corresponding to the noisy speech frame;

[0042] An excitation gain prediction sub-module, configured to perform gain prediction on the excitation signal corresponding to the audio feature vector through the second neural network to obtain the second gain.

[0043] In some embodiments of the present application, based on the above technical solution, the vector acquisition module includes:

[0044] A feature calculation sub-module, configured to input the spectral coefficients of the to-be-processed speech frame into a preprocessing neural network for feature calculation to obtain the audio feature vector of the to-be-processed speech frame, where the preprocessing neural network is trained according to the spectral coefficients of the noisy speech frame and the spectral coefficients of the denoised speech frame corresponding to the noisy speech frame.

[0045] In some embodiments of the present application, based on the above technical solution, the speech processing device further includes:

[0046] A historical spectral coefficient acquisition module, configured to acquire the spectral coefficients of the historical speech frames of the to-be-processed speech frame;

[0047] The feature calculation sub-module includes:

[0048] A feature vector calculation unit, configured to input the spectral coefficients of the to-be-processed speech frame and the spectral coefficients of the historical speech frame into the preprocessing neural network for feature calculation to obtain the audio feature vector of the to-be-processed speech frame.

[0049] In some embodiments of the present application, based on the above technical solution, the gain control module includes:

[0050] A first enhancement sub-module, configured to enhance the to-be-processed speech frame according to the second gain to obtain a first enhancement result;

[0051] A second enhancement sub-module, configured to perform gain operation on each sub-band in the first enhancement result according to the first gain to obtain a second enhancement result;

[0052] An energy compensation sub-module, configured to perform energy compensation on the second enhancement result according to the control coefficient to obtain a third enhancement result;

[0053] An inverse time-frequency conversion sub-module, configured to perform inverse time-frequency conversion on the third enhancement result to obtain an enhanced speech frame as a target speech frame.

[0054] In some embodiments of the present application, based on the above technical solutions, the speech processing device further includes:

[0055] A first amplitude spectrum calculation module, configured to calculate the amplitude spectrum and phase spectrum corresponding to the to-be-processed speech frame according to the to-be-processed speech frame;

[0056] The gain control module includes:

[0057] A first amplitude spectrum gain control sub-module, configured to perform gain control on the amplitude spectrum corresponding to the to-be-processed speech frame according to the first gain and the second gain to obtain an enhanced amplitude spectrum;

[0058] A first phase spectrum merging sub-module, configured to merge the enhanced amplitude spectrum and the phase spectrum corresponding to the to-be-processed speech frame to obtain a fourth enhancement result;

[0059] A first amplitude spectrum energy compensation sub-module, configured to perform energy compensation on the fourth enhancement result according to the control coefficient to obtain a compensated enhancement result;

[0060] A first amplitude spectrum inverse time-frequency conversion sub-module, configured to perform inverse time-frequency conversion on the compensated enhancement result to obtain an enhanced speech frame as a target speech frame.

[0061] In some embodiments of the present application, based on the above technical solutions, the speech processing device further includes:

[0062] A second amplitude spectrum calculation module, configured to calculate the amplitude spectrum and phase spectrum corresponding to the to-be-processed speech frame according to the to-be-processed speech frame;

[0063] A second amplitude spectrum gain control sub-module, configured to perform gain control on the amplitude spectrum corresponding to the to-be-processed speech frame according to the first gain and the second gain to obtain an enhanced amplitude spectrum;

[0064] A second amplitude spectrum energy compensation sub-module, configured to perform energy compensation on the enhanced amplitude spectrum according to the control coefficient to obtain a compensated amplitude spectrum;

[0065] A second phase spectrum merging sub-module, configured to merge the compensated amplitude spectrum and the phase spectrum corresponding to the to-be-processed speech frame to obtain a compensated enhancement result;

[0066] The second amplitude spectrum inverse time-frequency conversion sub-module is configured to perform inverse time-frequency conversion according to the compensated enhancement result to obtain an enhanced speech frame as a target speech frame.

[0067] According to one aspect of the embodiments of the present application, there is provided an electronic device, which includes: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the speech processing method in the above technical solution by executing the executable instructions.

[0068] According to one aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the speech processing method in the above technical solution is implemented.

[0069] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the speech processing method provided in the above various optional implementation manners.

[0070] In the embodiments of the present application, the first gain and the second gain are respectively calculated for the glottal features and the excitation signals of the noisy speech signal, and then gain control is performed according to the first gain and the second gain, so as to denoise the noisy speech signal. Denoising according to the glottal features can specifically identify the human voice part in the speech signal. Therefore, the human voice part is also processed during the denoising process, and it is no longer necessary to train for various types of noises. Therefore, the influence of the completeness of the training data is reduced, and it is possible to effectively process noise types and noise environments not included in the training data, improving the denoising effect.

[0071] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings according to these drawings without creative efforts. In the drawings:

[0073] Figure 1 It is a schematic diagram of an exemplary system architecture in an application scenario of the technical solution of the present application;

[0074] Figure 2 Schematic diagram showing a digital model for generating speech signals;

[0075] Figure 3 Schematic diagram of an example implementation of the glottal filter in an embodiment of the present application;

[0076] Figure 4 Schematic diagram of another example implementation of the glottal filter in an embodiment of the present application;

[0077] Figure 5 Schematic diagram showing the frequency responses of the excitation signal and the glottal filter decomposed from the original speech signal at different signal-to-noise ratios;

[0078] Figure 6 Flowchart showing a speech processing method according to an embodiment of the present application;

[0079] Figure 7 Schematic diagram of the structure of a first neural network according to a specific embodiment;

[0080] Figure 8 Schematic diagram of the structure of a second neural network according to a specific embodiment;

[0081] Figure 9 Schematic diagram of the structure of a second neural network according to a specific embodiment;

[0082] Figure 10 Schematic diagram of the structure of a third neural network according to a specific embodiment;

[0083] Figure 11 Schematic diagram of the overall process in an embodiment of the present application;

[0084] Figure 12 Schematic diagram of the overall process of another solution in an embodiment of the present application;

[0085] Figure 13 Schematic diagram of the overall process of another solution in an embodiment of the present application;

[0086] Figure 14 Another schematic diagram of the structure of a third neural network according to a specific embodiment;

[0087] Figure 15 Schematically shows the block diagram of the composition of a speech processing device in an embodiment of the present application;

[0088] Figure 16 Schematic diagram showing the structure of a computer system of an electronic device suitable for implementing an embodiment of the present application. Detailed implementation manners

[0089] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art.

[0090] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be employed. In other instances, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of this application.

[0091] The block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0092] The flowcharts shown in the drawings are merely illustrative and do not necessarily include all of the content and operations / steps, nor do they necessarily have to be executed in the order described. For example, some operations / steps may be decomposed, while some operations / steps may be combined or partially combined, so the actual execution order may change according to the actual situation.

[0093] Noise in the voice signal will greatly reduce the voice quality and affect the user's auditory experience. Therefore, in order to improve the quality of the voice signal, it is necessary to perform enhancement processing on the voice signal to remove the noise as much as possible and retain the original voice information in the voice signal, that is, to obtain a clean signal after denoising.

[0094] The solution of this application can be applied to scenarios of voice calls, such as voice calls through instant messaging software, multi-person calls in game applications, etc., and can also be applied to various cloud technology-based services, such as cloud games, cloud conferences, cloud calls, and cloud education. Among them, voice enhancement can be performed according to this solution at the voice sending end, the voice receiving end, or the server providing the voice communication service.

[0095] Cloud conferencing is an important part of remote work. In cloud conferencing, after the voice collection device of the participants in the cloud conference collects the voice signal of the speaker, it needs to send the collected voice signal to other conference participants. This process involves the transmission and playback of voice signals among multiple participants. If the noise signals mixed in the voice signals are not processed, it will greatly affect the auditory experience of the conference participants. In this scenario, the solution of this application can be applied to enhance the voice signals in the cloud conference, so that the voice signals heard by the conference participants are enhanced voice signals, improving the quality of the voice signals.

[0096] Cloud conferencing is an efficient, convenient, and low-cost conferencing form based on cloud computing technology. Users only need to perform simple and easy operations through the Internet interface to quickly and efficiently synchronously share voice, data files, and videos with teams and customers around the world. The complex technologies such as data transmission and processing in the conference are operated by the cloud conferencing service provider to assist users.

[0097] Currently, domestic cloud conferencing mainly focuses on service contents based on the SaaS (Software as a Service) model, including service forms such as telephone, network, and video. The video conferencing based on cloud computing is called cloud conferencing. In the era of cloud conferencing, the transmission, processing, and storage of data are all processed by the computer resources of the video conferencing provider. Users no longer need to purchase expensive hardware and install cumbersome software at all. They only need to open the client and enter the corresponding interface to conduct efficient remote conferences.

[0098] The cloud conferencing system supports multi-server dynamic cluster deployment and provides multiple high-performance servers, greatly improving the stability, security, and availability of the conference. In recent years, video conferencing has been welcomed by many users because it can significantly improve communication efficiency, continuously reduce communication costs, and bring about an upgrade in internal management level, and has been widely applied in various fields such as government affairs, transportation, finance, operators, education, and enterprises.

[0099] Next, take Voice over Internet Protocol (VoIP) as an example to introduce the application scenario of the embodiments of this application. Please refer to Figure 1 , Figure 1 which is the schematic diagram of the exemplary system architecture of the technical solution of this application in an application scenario.

[0100] As Figure 1 shown, the system architecture includes a sending end 110 and a receiving end 120. There is a network connection between the sending end 110 and the receiving end 120, and the sending end 110 and the receiving end 120 can conduct voice communication through the network connection.

[0101] AsFigure 1 As shown in Figure 1 , the sending end 110 includes an acquisition module 111, a pre-enhancement module 112, and an encoding module 113. Among them, the acquisition module 111 is used to acquire voice signals, and it can convert the acquired acoustic signals into digital signals; the pre-enhancement module 112 is used to enhance the acquired voice signals to remove the noise in the acquired voice signals and improve the quality of the voice signals. The encoding module 113 is used to encode the enhanced voice signals to improve the anti-interference ability of the voice signals during transmission. The pre-enhancement module 112 can perform voice enhancement according to the method of this application. After enhancing the voice, it is then encoded, compressed, and transmitted, so that the signals received by the receiving end are no longer affected by noise.

[0102] The receiving end 120 includes a decoding module 121, a post-enhancement module 122, and a playback module 123. The decoding module 121 is used to decode the received encoded voice to obtain a decoded signal; the post-enhancement module 122 is used to perform enhancement processing on the decoded voice signal; the playback module 123 is used to play the enhanced voice signal. The post-enhancement module 122 can also perform voice enhancement according to the method of this application. In some embodiments, the receiving end 120 may further include a sound effect adjustment module, and this sound effect adjustment module is used to adjust the sound effect of the enhanced voice signal.

[0103] In a specific embodiment, it may be that only the receiving end 120 or only the sending end 110 performs voice enhancement according to the method of this application. Of course, it may also be that both the sending end 110 and the receiving end 120 perform voice enhancement according to the method of this application.

[0104] In some application scenarios, the terminal device in the VoIP system can support other third-party protocols in addition to supporting VoIP communication. For example, traditional PSTN (Public Switched Telephone Network) circuit domain telephones. However, traditional PSTN services cannot perform voice enhancement. In this scenario, voice enhancement can be performed according to the method of this application in the terminal acting as the receiving end.

[0105] Before specifically describing this solution, first introduce the voice generation method based on an excitation signal. The human voice generation method is that when air flows through the vocal cords, it drives the vocal cords to vibrate and produce sound. And the voice generation process of the voice generation method based on an excitation signal includes: at the trachea, a noise-like impact signal with a certain energy is generated, that is, the excitation signal, which is equivalent to the air flow; the impact signal impacts the glottis filter (equivalent to the human vocal cords), generating a quasi-periodic opening and closing, thereby making a sound. It can be seen that this process simulates the human voice generation process.

[0106] Figure 2A schematic diagram of a digital model for generating a speech signal is shown. Through this digital model, the generation process of the speech signal can be described. As Figure 2 shown, the excitation signal impinges on the glottis filter to output the speech signal. Among them, the glottis filter is usually configured according to glottis parameters. The glottis filter can adopt the filter in various schemes for generating speech using the source-filter model. Specifically, please refer to Figure 3 , Figure 3 which is a schematic diagram of an example implementation of the glottis filter in the embodiments of the present application. Considering the short-term correlation of the speech signal, the glottis filter can be implemented by a Linear Predictive Coding (LPC) filter. The excitation signal impinges on the LPC filter to generate the speech signal.

[0107] On the other hand, according to the classical speech signal processing theory, the LPC filter only reflects the short-term correlation in voicing, but for voiced sounds (such as vowels), there is long-term correlation (Long-Term Prediction, LTP) (or quasi-periodicity); the glottis filter can also be implemented by multiple filters. Specifically, please refer to Figure 4 , Figure 4 which is a schematic diagram of another example implementation of the glottis filter in the embodiments of the present application. As Figure 4 shown, the glottis filter consists of two parts: an LPC filter and an LTP filter. Among them, the LTP filter also receives the pitch period as an input. The pitch period indicates that when calculating the nth sample, the (n - p)th sample point is required, where p is the pitch period.

[0108] Figure 5 A schematic diagram showing the frequency responses of the excitation signal and the glottis filter decomposed from the original speech signal under different signal-to-noise ratios is shown. Figure 5 a shows a schematic diagram of the frequency response of the original speech signal. Figure 5 b shows a schematic diagram of the frequency response of the glottis filter decomposed from the original speech signal. Figure 5 c shows a schematic diagram of the frequency response of the excitation signal decomposed from the original speech signal. Figure 5 shows two original speech signals and their corresponding decomposition results, represented by solid lines and dashed lines respectively. One of the original signals is a 30 dB signal, and the other signal is a 0 dB signal. The undulating part in the frequency response diagram of the original speech signal corresponds to the peak position in the frequency response diagram of the glottis filter. The excitation signal is equivalent to the residual signal (i.e., the excitation signal) after performing linear prediction analysis on the original speech signal. Therefore, its corresponding frequency response is relatively flat. In Figure 5In a, although there are certain differences between the two original voice signals of 30 dB and 0 dB, the differences are relatively insignificant and there are many overlapping parts. After decomposition, Figure 5 in the frequency response of the glottal filter in b, the differences between the two are relatively obvious and the overlapping parts are significantly reduced. While in Figure 5 the excitation signal in c, the differences between the two signals are significantly amplified and the two excitation signals can be clearly distinguished. It can be seen that signal decomposition can more fully reflect the differences between the original voice signals, and gain control based on the decomposition results can also make the gain results accurate.

[0109] As can be seen from the above, an excitation signal and a glottal filter can be decomposed from an original voice signal (i.e., a voice signal without noise), and the decomposed excitation signal and glottal filter can be used to express the original voice signal. Among them, the glottal filter can be expressed by glottal parameters. Conversely, if the excitation signal corresponding to an original voice signal and the glottal parameters for determining the glottal filter are known, the original voice signal can be reconstructed according to the corresponding excitation signal and glottal filter.

[0110] The solution of this application is based on this principle, and calculates the gain corresponding to the glottal filter and the gain corresponding to the excitation signal respectively to perform gain control on the original voice signal, so as to achieve voice enhancement.

[0111] The implementation details of the technical solution of the embodiments of this application are elaborated in detail below. For the convenience of introduction, please refer to Figure 6 , Figure 6 shows a flowchart of a voice processing method according to an embodiment of this application. This method can be executed by a computer device with processing capabilities, such as a terminal, a server, etc., which is not specifically limited here. As Figure 6 shown, this method at least includes the following steps S610 to S660:

[0112] Step S610, obtain the spectral coefficients of the voice frame to be processed.

[0113] Voice signals change non-stationarily and randomly over time, but the characteristics of voice signals are strongly correlated in a short period of time, that is, voice signals have short-time correlation. Therefore, in the solution of this application, voice processing is performed in units of voice frames. The voice frame to be processed is the current voice frame to be processed, which is any frame in the original noisy audio to be processed.

[0114] To obtain the spectral coefficients of the speech frame to be processed, the time-frequency transformation can be performed on the time-domain signal of the speech frame to be processed. The time-frequency transformation can be, for example, the Short-term Fourier transform (STFT). The dimension of the spectral coefficients usually depends on the number of sample points of the speech frame to be processed and the overlapping ratio of the windowing used in the STFT transformation. For example, for the frequency-domain representation of 257 sample points, the dimension of the spectral coefficients is [2, 257].

[0115] Step S620: Calculate the features based on the spectral coefficients of the speech frame to be processed to obtain the audio feature vector of the speech frame to be processed.

[0116] Based on the frequency-domain representation of the speech frame to be processed obtained by the STFT transformation, audio feature extraction can be performed to obtain the audio feature vector. The way of feature extraction can be executed according to a preset algorithm or through a trained neural network. The dimension of the audio feature vector usually depends on the number of sample points of the speech frame to be processed and the overlapping ratio of the windowing used in the STFT transformation. For example, for the frequency-domain representation of 257 sample points, the audio feature vector can be 128-dimensional.

[0117] Step S630: Calculate the glottal gain based on the audio feature vector to obtain the first gain, where the first gain corresponds to the glottal feature of the speech frame to be processed.

[0118] The glottal gain calculation is a process of calculating the gain for the glottal filter part corresponding to the speech frame to be processed. The calculated first gain is associated with the glottal feature of the speech frame to be processed. Depending on the glottal model used for the speech frame to be processed, the first gain specifically includes multiple sub-gains. For example, for the LPC+LTP glottal model, the first gain can include the sub-gain corresponding to LPC and the self-gain corresponding to LTP.

[0119] The glottal gain calculation can be performed in the way of a neural network. The trained neural network is used to directly output the corresponding first gain according to the spectral coefficients of the speech frame to be processed. The neural network is trained in a supervised training manner. The training data includes the noisy speech and the data annotation corresponding to each speech frame calculated for the noisy speech, that is, the denoised speech. The neural network is trained based on the noisy speech and the denoised speech to output the first gain.

[0120] The glottal gain calculation can also be performed in other ways. For example, first, the speech frame to be processed is decomposed according to the glottal model to obtain the glottal parameters of the corresponding glottal filter. Then, both the glottal parameters and the spectral coefficients of the speech frame to be processed are input into a neural network for processing. The neural network simulates the denoised speech based on the glottal parameters and the spectral coefficients of the speech frame to be processed, and then determines the first gain through the simulated denoised speech and the noisy speech.

[0121] Step S640: Calculate the excitation gain based on the audio feature vector to obtain a second gain, where the second gain corresponds to the excitation signal of the speech frame to be processed.

[0122] The excitation gain calculation is a process of calculating the gain for the excitation signal part corresponding to the speech frame to be processed. The calculated second gain is associated with the excitation signal of the speech frame to be processed. Specifically, the dimension of the second gain usually corresponds to the spectral coefficients of the speech frame to be processed.

[0123] The excitation gain calculation can be performed using a neural network. The trained neural network directly outputs the corresponding second gain based on the spectral coefficients of the speech frame to be processed. The neural network is trained in a supervised training manner, and the training data includes the noisy speech and the excitation signal obtained after decomposing the corresponding denoised speech of the noisy speech. The neural network is trained based on the noisy speech and the excitation signal of the denoised speech to output the second gain.

[0124] The excitation gain calculation can also be performed in other ways. For example, first, the speech frame to be processed is decomposed according to the glottal model to obtain the corresponding excitation signal. Then, both the excitation signal and the spectral coefficients of the speech frame to be processed are input into a neural network for processing. The neural network uses the glottal parameters obtained when decomposing the speech frame to be processed to simulate the denoised speech based on the excitation signal and the spectral coefficients of the speech frame to be processed, and then determines the second gain through the simulated denoised speech and the noisy speech.

[0125] Step S650: Perform compensation prediction based on the audio feature vector to obtain a control coefficient, where the control coefficient is determined according to the energy of the spectral coefficients of the audio frame to be processed.

[0126] The control system is a scheme for performing energy compensation. During the process of enhancing the speech frame to be processed, a part of the energy is also lost during denoising, which affects the auditory effect of the obtained denoised result. Compensation prediction is performed based on the spectral coefficients of the speech frame to be processed to obtain a control coefficient. The control coefficient is usually a group of two-dimensional vectors, which respectively represent the real part and the imaginary part of the spectral coefficients of the speech frame to be processed. The control coefficient can directly act on the enhancement results of the first gain and the second gain to perform energy compensation.

[0127] The process of compensation prediction can be estimated based on the enhanced result and the speech frame to be processed. For example, by calculating the energy difference between the two, the control coefficient required for compensation is estimated. The process of compensation prediction can also use deep learning to learn the control coefficient required for the speech frame to be processed to restore to its original energy level after denoising. For example, the noisy speech frame is used as a training sample, and the control coefficient required after denoising is calculated manually as the training target, so as to obtain the corresponding model to predict the control coefficient.

[0128] Step S660: Perform gain control on the speech frame to be processed according to the first gain, the second gain, and the control coefficient to obtain the target speech frame.

[0129] Specifically, first, the frequency-domain representation of the speech frame to be processed can be enhanced according to the second gain, and then the obtained result can be further gained according to the first gain to obtain the enhanced frequency-domain representation. Subsequently, energy compensation calculation is performed on the enhanced frequency representation according to the control coefficient to obtain the compensated frequency-domain representation. Then, inverse STFT is performed on the compensated frequency-domain representation to obtain the enhanced speech frame to be processed.

[0130] In the embodiments of the present application, the first gain and the second gain are respectively calculated for the glottal characteristics and the excitation signal of the noisy speech signal, and then gain control is performed according to the first gain and the second gain, so as to denoise the noisy speech signal and perform energy compensation on the denoised speech signal. Denoising according to the glottal characteristics can specifically identify the human voice part in the speech signal. Therefore, the human voice part is also processed during the denoising process, and there is no need to train for various noises. Therefore, the influence on the completeness of the training data is reduced, and noise types and noise environments not included in the training data can be effectively processed, improving the denoising effect.

[0131] In some embodiments of the present application, based on the above technical solution, the above step S630 of calculating the glottal gain according to the audio feature vector to obtain the first gain may include the following steps:

[0132] Input the audio feature vector into the first neural network, and the first neural network is trained according to the glottal characteristics corresponding to the noisy speech frame and the glottal characteristics corresponding to the denoised speech frame corresponding to the noisy speech frame;

[0133] The first neural network predicts the gain according to the audio feature vector to obtain the first gain.

[0134] The first neural network can be a model constructed by a long short-term memory neural network, a convolutional neural network, a recurrent neural network, a fully connected neural network, etc., and no specific limitation is made here.

[0135] During the training process, first, the noisy speech frames in the training data are decomposed to obtain the frequency response corresponding to the glottal filter in the glottal model. Then, based on the audio feature vector of the noisy speech frame and the frequency response of the decomposed glottal filter, training is performed. By adjusting the model parameters of the first neural network until the first gain output by the first neural network enables the difference between the glottal features of the noisy speech frame and the glottal features of the denoised speech frame to meet the preset requirements. Among them, the preset requirements can be calculated by means of the mean square error. Through training, the mean square error between the glottal features of the noisy speech frame and the glottal features of the denoised speech frame satisfied the set mean square error threshold, thereby determining that the trained model can achieve the expected purpose. Through this training process, the first gain predicted by the first neural network can make the glottal filter of the speech frame to be processed under the glottal model (i.e., glottal filter + excitation signal) sufficiently similar to the glottal filter of the clean speech under the glottal model, thus having the ability to reduce noise.

[0136] The first neural network predicts the gain based on the audio feature vector to obtain the first gain. Figure 7 The structure diagram of the first neural network shown according to a specific embodiment is as Figure 7 shown. The first neural network includes three fully connected (FC) layers. The input F(n) is a [128,1]-dimensional audio feature vector. The output of the first FC layer is a [256,1]-dimensional vector, the output of the second FC layer is a [128,1]-dimensional vector, and the output of the third FC layer is a [32,1]-dimensional vector, that is, the first gain g11(n). Of course, Figure 7 This is only an exemplary example of the structure of the first neural network and should not be considered as a limitation on the scope of use of this application.

[0137] In the embodiments of the present application, the first gain for the glottal feature is obtained through a neural network, and the relationship between the glottal feature and the first gain is learned through the neural network, so that a model can be obtained based on limited training data to handle various situations in the actual scenario, improving the flexibility of the solution.

[0138] In an embodiment of the present application, based on the above technical solution, the above step of predicting the gain by the first neural network according to the audio feature vector to obtain the first gain may include the following steps:

[0139] The first neural network calculates the gain for the subbands in the audio feature vector to obtain the first glottal gain corresponding to each subband, where the subband corresponds to at least one frequency band in the speech frame to be processed;

[0140] Merge the first glottal gains corresponding to each sub-band as the first gain.

[0141] Specifically, the spectral response related to the glottal filter is a low-pass-like smoothing effect. Therefore, although the dimension of the frequency-domain representation of the speech frame to be processed is 257 dimensions, when calculating the first gain, it is not necessary to reach a resolution of 257 dimensions. Therefore, during the calculation of the first gain, several adjacent coefficients can be merged and share one first gain. Each sub-band includes features in at least two adjacent dimensions in the audio feature vector.

[0142] By dividing the frequency-domain representation of the speech frame to be processed along the frequency, multiple sub-bands in the frequency-domain representation can be obtained. The frequency division performed on the frequency-domain representation can be uniform frequency division (i.e., each sub-band corresponds to the same frequency width), or non-uniform frequency division, which is not specifically limited here. It can be understood that each sub-band corresponds to a frequency range, which includes multiple frequency points.

[0143] The non-uniform frequency division can be Bark division. Bark division is performed according to the Bark frequency scale. The Bark frequency scale maps the frequency to multiple critical frequency bands in psychoacoustics. The number of frequency bands can be set according to the sampling rate and actual needs. For example, the number of frequency points is set to 24. Bark division conforms to the characteristics of the auditory system. Generally, the lower the frequency, the fewer the number of coefficients included in the sub-band, or even just a single coefficient. The higher the frequency, the more the number of coefficients included in the sub-band.

[0144] In one embodiment, for 257 coefficients, 8 adjacent coefficients are merged into one sub-band (the first element of the FFT coefficient is the DC component and can be ignored). Therefore, the dimension of the finally output first gain g1(n) is 32 dimensions. Through the first neural network, the first glottal gains corresponding to each sub-band can be output. Merge the first glottal gains of each sub-band to obtain the first gain. That is, 32 sub-bands correspond to the 32 dimensions of the first gain.

[0145] In the embodiments of the present application, the first gain is calculated according to the audio feature vector by means of sub-band merging, so as to be able to reduce the dimension of the calculation process, thereby being able to reduce the overall calculation amount of the solution and improve the calculation efficiency.

[0146] In one embodiment of the present application, based on the above technical solution, the speech processing method may further include the following steps:

[0147] Perform predictive analysis on the audio feature vector and the fundamental period of the speech frame to be processed by the first neural network to determine the second glottal gain, and the second glottal gain corresponds to the long-term correlation feature of the audio feature vector;

[0148] Combining the first glottal gains corresponding to each sub-band as the first gain includes:

[0149] Combining the first glottal gain and the second glottal gain corresponding to each sub-band as the first gain.

[0150] In this embodiment, the first gain includes two parts, the first glottal feature corresponding to the short-term correlation feature of the speech frame and the second glottal feature corresponding to the long-term correlation feature of the speech frame. The fundamental period of the speech frame to be processed can be obtained by performing speech decomposition and analysis on the speech frame to be processed in advance. The first neural network can directly output the second glottal gain according to the audio feature vector and the fundamental period. The second glottal gain corresponds to the glottal parameter of the LTP filtering in the glottal filter. Therefore, during the training process, the model is trained based on the glottal parameters corresponding to the LTP filter obtained from the decomposed speech and the denoised speech. By adjusting the model parameters, the mean square error similarity between the finally output first gain and the first gain corresponding to the denoising result reaches the mean square error threshold, thereby completing the training. The first neural network can output the first glottal gain and the second glottal gain together. In one embodiment, the first neural network can be composed of two sub-networks, which are respectively used to output the first glottal feature and the second glottal feature.

[0151] In the embodiment of the present application, the long-term correlation of the speech frame is further considered in the calculation process of the first gain, making the recognition of the speech part in the speech frame more refined, thereby avoiding the influence of the gain on the original speech and improving the accuracy of the solution.

[0152] In an embodiment of the present application, based on the above technical solution, the above step of predicting the gain according to the audio feature vector by the first neural network to obtain the first gain may include the following steps:

[0153] Predicting the first glottal parameter according to the audio feature vector by the first neural network, where the first glottal parameter is used to represent the short-term correlation feature of the audio feature vector;

[0154] Predicting the first prediction result according to the first glottal parameter by the first neural network;

[0155] Determining the first gain according to the first prediction result.

[0156] In this embodiment, the first neural network predicts the first glottal parameter corresponding to the speech frame to be processed according to the audio feature vector. The first glottal parameter is used to represent the short-term correlation feature of the audio feature vector. Specifically, the first glottal parameter corresponds to the LPC filter. During the training process of the first neural network, by decomposing the denoised speech corresponding to the noisy speech in the pre-trained data, the configuration parameters of the LPC filter of the denoised speech can be determined. According to the audio feature vector of the noisy speech and the configuration parameters of the LPC filter of the denoised speech, the first neural network can be trained so that the mean square error between the first glottal parameter predicted by the first neural network and the configuration parameters of the LPC filter of the denoised speech meets the set accuracy requirement. Then, based on the obtained first glottal parameter, the first neural network can predict the spectrum of the denoised speech, and determine the first prediction result according to the predicted spectrum and the spectrum of the speech frame to be processed. The first prediction result is the first gain.

[0157] In this embodiment, the first glottal parameter corresponding to the speech frame to be processed is predicted by the first neural network, and then the first gain is predicted according to the first glottal parameter. The glottal parameter prediction target corresponding to the glottal filter simplifies the complexity of the training data compared with the annotation result of annotating each speech frame of the speech to be processed, thereby improving the calculation efficiency.

[0158] In an embodiment of the present application, based on the above technical solution, the speech processing method further includes:

[0159] The first neural network performs parameter prediction according to the audio feature vector and the fundamental period of the speech frame to be processed of the speech frame to be processed, and obtains a second glottal parameter, where the second glottal parameter is used to represent the long-term correlation feature of the audio feature vector;

[0160] The first neural network performs gain prediction according to the second glottal parameter to obtain a second prediction result;

[0161] Determining the first gain according to the first prediction result includes:

[0162] Combining the first prediction result and the second prediction result to determine the first gain.

[0163] Specifically, the first neural network predicts the second glottal parameter based on the audio feature vector and the fundamental period of the speech frame to be processed. The first glottal parameter is used to represent the long-term correlation feature of the audio feature vector. Specifically, the second glottal parameter corresponds to the LTP filter. In this embodiment, the glottal model of the speech frame further includes an LTP filter. The speech signal generated by the LPC filter configured according to the first glottal parameter is further processed by the LTP filter configured according to the second glottal parameter to simulate the speech in the speech frame to be processed. During the training process, by decomposing the denoised speech corresponding to the noisy speech in the pre-trained data, the configuration parameters of the LTP filter of the denoised speech can be determined. According to the audio feature vector of the noisy speech and the configuration parameters of the LTP filter of the denoised speech, the first neural network can be trained so that the mean square error between the second glottal parameter predicted by the first neural network and the configuration parameters of the LTP filter of the denoised speech meets the set accuracy requirement. Then, according to the obtained second glottal parameter, the first neural network can combine the first prediction result to predict the spectrum of the denoised speech, thereby obtaining the second prediction result. The second prediction result is also a part of the first gain, which makes the long-term correlation feature in the denoising result obtained based on the second prediction result similar to the long-term correlation feature in the denoised speech. By combining the first prediction result and the second prediction result, the first gain can be obtained. In the subsequent application process, the speech frame to be processed is enhanced according to the first prediction result and the second prediction result in sequence, so as to achieve the effect of noise reduction.

[0164] In the embodiment of the present application, by predicting the second glottal parameter, the long-term correlation of the speech frame is further considered in the prediction process of the first gain, making the prediction of the voiced part in the speech frame more accurate, thereby improving the accuracy of the solution.

[0165] In an embodiment of the present application, based on the above technical solution, the above step S640 of calculating the excitation gain according to the audio feature vector to obtain the second gain includes:

[0166] Input the audio feature vector into the second neural network, and the second neural network is trained according to the excitation signal of the noisy speech frame and the excitation signal of the denoised speech frame corresponding to the noisy speech frame;

[0167] The second neural network predicts the gain according to the excitation signal corresponding to the audio feature vector to obtain the second gain.

[0168] The second neural network refers to a neural network model used to predict the second gain corresponding to the excitation signal. The second neural network can be a model constructed by a long short-term memory neural network, a convolutional neural network, a recurrent neural network, a fully connected neural network, etc., and is not specifically limited herein.

[0169] During the training process, first, the noisy speech frames in the training data are decomposed to obtain the frequency response corresponding to the excitation signal in the glottal model. Then, training is performed based on the audio feature vectors of the noisy speech frames and the frequency response of the decomposed excitation signal. By adjusting the model parameters of the second neural network until the second gain output by the second neural network enables the difference between the excitation signal of the noisy speech frame and the excitation signal of the denoised speech frame to meet a preset requirement. Here, the preset requirement may be such that the similarity between the excitation signal of the noisy speech frame and the excitation signal of the denoised speech frame is not lower than the similarity threshold. Through this training process, the second gain predicted by the second neural network can make the excitation signal of the speech frame to be processed under the glottal model sufficiently similar to the excitation signal of the clean speech under the glottal model, thus having the ability to reduce noise.

[0170] The second neural network predicts the gain according to the audio feature vector to obtain the second gain. Figure 8 FIG. shows a schematic structural diagram of the second neural network according to a specific embodiment, as Figure 8 shown, the first neural network includes three fully connected (FC) layers. The input F(n) is a 128-dimensional audio feature vector. The output of the first FC layer is a [1024,1]-dimensional vector, the output of the second FC layer is a [512,1]-dimensional vector, and the output of the third FC layer is a [257,1]-dimensional vector, that is, the second gain g12(n). Of course, Figure 8 This is only an exemplary example of the structure of the second neural network and should not be considered as a limitation on the scope of use of this application.

[0171] In the embodiments of the present application, the second gain for the excitation signal is obtained through a neural network, and the relationship between the excitation signal and the second gain is learned through the neural network, so that the noisy speech can be denoised according to the glottal model without performing speech decomposition on the noisy speech, thereby saving computing resources.

[0172] In an embodiment of the present application, based on the above technical solution, the above step S620 of calculating the audio feature vector of the speech frame to be processed according to the spectral coefficients of the speech frame to be processed may include the following steps:

[0173] The spectral coefficients of the speech frame to be processed are input into a preprocessing neural network for feature calculation to obtain the audio feature vector of the speech frame to be processed. The preprocessing neural network is trained according to the spectral coefficients of the noisy speech frames and the spectral coefficients of the corresponding denoised speech frames of the noisy speech frames.

[0174] Specifically, by performing STFT transformation on the speech frame to be processed, the frequency-domain representation of the speech frame to be processed can be obtained. By decomposing the frequency-domain representation into real and imaginary parts, the spectral coefficients of the speech frame to be processed can be obtained.

[0175] The preprocessing neural network refers to a neural network model used to convert the spectral coefficients of the speech frame to be processed into audio feature vectors. The preprocessing neural network can be a model constructed by long short-term memory neural network, convolutional neural network, recurrent neural network, fully connected neural network, etc., and no specific limitation is made here.

[0176] The input of the preprocessing neural network is the spectral coefficients of the speech frame to be processed, and the output is the audio feature vector corresponding to the speech frame to be processed. The training process of the preprocessing neural network is usually carried out together with the processes of the first neural network and the second neural network. Therefore, during the training process, the adjustment of the network parameters of this neural network is carried out together with the adjustment processes of the first neural network and the second neural network. During training, the spectral coefficients of the noisy speech frames and the spectral coefficients of the denoised speech frames corresponding to the noisy speech frames in the training data are input into the preprocessing neural network for prediction, and then the first neural network and the second neural network are trained according to the predicted audio feature vectors. The model parameters of the three models are adjusted according to the results output by the first neural network and the second neural network. By jointly adjusting the model parameters of the preprocessing neural network with the model parameters of the first neural network and the second neural network, the first gain and the second gain can make the difference between the enhanced denoising result and the denoised speech meet the requirements.

[0177] The preprocessing neural network performs feature calculation according to the audio feature vector to obtain the audio feature vector of the speech frame to be processed. Figure 9 The structural schematic diagram of the second neural network shown according to a specific embodiment is as Figure 9 shown. The preprocessing neural network includes 6 convolutional layers and one long short-term memory (LSTM) layer. The input S(n) is represented by spectral coefficients, so it is a [2,257]-dimensional spectral coefficient. Figure 9 The dimension of the variable output by each convolutional layer and LSTM layer in is marked. The first convolutional layer outputs a [16,127]-dimensional variable, the second convolutional layer outputs a [32,62]-dimensional variable, the third convolutional layer outputs a [64,29]-dimensional variable, the fourth convolutional layer outputs a [128,13]-dimensional variable, the fifth convolutional layer outputs a [128,5]-dimensional variable, the sixth convolutional layer outputs a [128,1]-dimensional variable, and the LSTM layer outputs a [128,1]-dimensional variable. The variable output by the LSTM layer is the audio feature vector F(n), which is a [128,1]-dimensional vector. It should be understood that Figure 9This is merely an exemplary example of the structure of the preprocessing neural network and should not be considered as a limitation on the scope of use of this application.

[0178] In an embodiment of this application, the neural network is used to extract features from the speech frame to be processed, thereby reducing the influence of noise on the audio feature vector during the feature extraction process, making the audio feature vector better reflect the speech features in the speech frame to be processed, and improving the accuracy of the solution.

[0179] In an embodiment of this application, based on the above technical solution, the speech processing method further includes:

[0180] Obtain the spectral coefficients of the historical speech frames of the speech frame to be processed;

[0181] The above step of inputting the spectral coefficients of the speech frame to be processed into the preprocessing neural network for feature calculation to obtain the audio feature vector of the speech frame to be processed includes:

[0182] Input the spectral coefficients of the speech frame to be processed and the spectral coefficients of the historical speech frames into the preprocessing neural network for feature calculation to obtain the audio feature vector of the speech frame to be processed.

[0183] Specifically, in this embodiment, when extracting the features of the audio feature vector, the historical speech frames of the speech frame to be processed can also be used as inputs. Specifically, first, obtain the spectral coefficients of the historical speech frames of the speech frame to be processed. The historical speech frames are other speech frames in the audio where the speech frame to be processed is located. For example, for the nth frame, the historical speech frames can be the (n - 1)th frame, the (n - 2)th frame, etc. When performing feature calculation, input the spectral coefficients of the speech frame to be processed and the spectral coefficients of the historical speech frames into the preprocessing neural network for feature calculation. In this embodiment, the structure of the preprocessing neural network is similar to the structure Figure 9 described previously, and it also uses 6 convolutional layers and 1 LSTM layer. The difference is that depending on the number of input historical speech frames, the dimension of the input variable increases. For example, if m historical speech frames are input, the dimension of the input variable is [2, 257 * m].

[0184] Correspondingly, during the training process of the preprocessing neural network, the spectral coefficients of the historical speech frames of the noisy speech frames, together with the spectral coefficients of the noisy speech frames and the spectral coefficients of the denoised speech frames, are also used as inputs for training. The specific training principle is the same as that in the above embodiment and will not be elaborated here.

[0185] In an embodiment of this application, the historical speech frames are used as inputs and processed together with the speech frame to be processed, so that the relationship between adjacent speech frames can be more fully considered during the feature extraction process, thereby improving the accuracy of feature extraction.

[0186] In one embodiment of the present application, based on the above technical solution, step S650 of performing gain control on the speech frame to be processed according to the first gain, the second gain, and the control coefficient to obtain the target speech frame includes:

[0187] Enhancing the speech frame to be processed according to the second gain to obtain a first enhancement result;

[0188] Performing a gain operation on each subband in the first enhancement result according to the first gain to obtain a second enhancement result;

[0189] Performing energy compensation on the second enhancement result according to the control coefficient to obtain a third enhancement result;

[0190] Performing an inverse time-frequency conversion according to the third enhancement result to obtain the enhanced speech frame as the target speech frame.

[0191] Specifically, for the frequency-domain representation of the speech frame to be processed, first perform a multiplication operation according to the corresponding parameters in the second gain for each sample point to obtain the first enhancement result. As described above, the dimension of the second gain corresponds to the frequency-domain representation of the speech frame to be processed. That is, if the frequency-domain representation of the speech frame to be processed is 257-dimensional, then the second gain is also 257-dimensional. Therefore, when enhancing according to the second gain, a multiplication operation can be directly performed according to the corresponding relationship of the dimensions to obtain the first enhancement result. Based on the first enhancement result, perform a gain operation according to the first gain. Specifically, when calculating the first gain, the first gain is merged according to the subband division. Therefore, when calculating according to the first gain, multiplication is also performed according to the corresponding relationship of the subband merging. For example, in the first gain result, every 8 dimensions correspond to one subband, and the first gain is a 32-dimensional variable. Then, when calculating according to the first gain, every 8 dimensions in the first gain result correspond to one dimension in the first gain for calculation, thereby obtaining the second gain result. Performing energy compensation on the second enhancement result according to the control coefficient to obtain the third enhancement result. Specifically, the process of energy compensation is to directly sum the control coefficient and the second gain result to obtain the third enhancement result, and the calculation formula is as follows:

[0192] S_e2(n) = S_e1(n) + g2(n)

[0193] where S_e2(n) is the third enhancement result, g2(n) is the control coefficient, and S_e1(n) is the second enhancement result. Performing an inverse STFT on the third enhancement result, that is, the frequency-domain representation can be transformed into a time-domain signal, thereby obtaining the enhanced speech frame, that is, the target speech frame.

[0194] In the embodiment of the present application, a specific method for performing gain control is provided, which improves the feasibility of the solution.

[0195] In one embodiment, compensating and predicting based on the spectral coefficients of the to-be-processed speech frame to obtain a control coefficient includes:

[0196] Inputting the spectral coefficients of the to-be-processed speech frame into a third neural network, where the third neural network is trained based on the energy of the spectral coefficients corresponding to the noisy speech frame and the energy of the spectral coefficients corresponding to the denoised speech frame corresponding to the noisy speech frame;

[0197] Compensating and predicting based on the spectral coefficients of the to-be-processed speech frame through the third neural network to obtain the control coefficient.

[0198] The third neural network refers to a neural network model for compensating and predicting. The third neural network can be a model constructed by a long short-term memory neural network, a convolutional neural network, a recurrent neural network, a fully connected neural network, etc., and is not specifically limited herein.

[0199] The input of the third neural network is the spectral coefficients of the to-be-processed speech frame, and the output is the control coefficient corresponding to the to-be-processed speech frame. The control coefficient is a set of two-dimensional vectors, respectively representing the real part and the imaginary part of the spectral coefficients; among them, the control coefficient will act on the enhancement result according to the first gain and the second gain to obtain a second enhanced spectrum for energy compensation.

[0200] The training process of the third neural network can be carried out independently or together with the processes of the first neural network and the second neural network. During training, the spectral coefficients of the noisy speech frame in the training data and the spectral coefficients of the denoised speech frame corresponding to the noisy speech frame are input into the third neural network for prediction to obtain the predicted control coefficient of the output. By adjusting the model parameters of the third neural network, the difference between the energy of the compensated target speech frame and the energy of the unprocessed noisy speech frame can meet the preset requirements. In this embodiment, the spectral coefficients of the to-be-processed speech frame and the spectral coefficients of the historical speech frame are input into the third neural network. Therefore, during the training process, the spectral coefficients of the historical speech frame are also input into the third neural network for training. Compensating and predicting based on the spectral coefficients of the to-be-processed speech frame through the third neural network to obtain the control coefficient.

[0201] Specifically, please refer to Figure 10 , Figure 10 which is a schematic structural diagram of the third neural network shown according to a specific embodiment. As Figure 10 shown, the preprocessing neural network includes 5 FC layers. The input F(n) is represented by spectral coefficients of [128,1]. Figure 10The dimensions of the variables output by each convolutional layer are marked. The first FC layer outputs a variable of dimension [2048, 1], the second FC layer outputs a variable of dimension [1024, 1], the third FC layer outputs a variable of dimension [1024, 1], the fourth FC layer outputs a variable of dimension [512, 1], and g2(n) output by the fifth FC layer is of dimension [257, 2]. The variable output by the fifth FC layer is the audio feature vector F(n). It should be understood that Figure 10 This is only an exemplary example of the structure of the preprocessing neural network and should not be considered as a limitation on the scope of use of this application.

[0202] The overall process of the speech processing method of this application will be introduced below. For the convenience of introduction, please refer to Figure 11 , Figure 11 which is a schematic diagram of the overall process in the embodiments of this application. As Figure 11 shown, the input of the scheme is the speech frame s(n) to be processed. The speech frame s(n) is subjected to STFT time-frequency transformation to obtain the spectral coefficients S(n). Based on the spectral coefficients S(n), the preprocessing neural network is called to obtain the audio feature vector F(n). Based on the audio feature vector F(n), the first neural network is called to obtain the first gain g11(n), and the second neural network is called to obtain the second gain g12(n). The first gain g11(n) and the second gain g12(n) are jointly used to perform the first spectral control (i.e., gain control) on the spectral coefficients S(n), so as to output the first enhanced spectrum S_e1(n). The first spectral control is mainly used to suppress the noise in the speech frame. In particular, the input of the preprocessing neural network may also include the spectral coefficients S_pre(n) of the historical frames (such as the n-1th, n-2th frames, etc., and so on) of the speech frame s(n). The processing process of the third neural network can be executed in parallel with the processes of the first neural network and the second neural network. Based on the audio feature vector F(n), the third neural network is called to obtain the control coefficient g2(n). Applying the control coefficient g2(n) to the first enhanced spectrum S_e1(n) can obtain the second enhanced spectrum S_e2(n). The second spectral control is mainly used for energy compensation. Finally, inverse time-frequency transformation is performed according to the second enhanced spectrum S_e2(n) to obtain the enhanced and supplemented signal s_e(n) corresponding to the speech frame.

[0203] In an embodiment of this application, based on the above technical solution, the method further includes:

[0204] Calculating the amplitude spectrum and phase spectrum corresponding to the speech frame to be processed according to the speech frame to be processed;

[0205] The above step S650 of performing gain control on the speech frame to be processed according to the first gain, the second gain, and the control coefficient to obtain the target speech frame includes:

[0206] Perform gain control on the amplitude spectrum corresponding to the speech frame to be processed according to the first gain and the second gain, and obtain an enhanced amplitude spectrum;

[0207] Merge the enhanced amplitude spectrum and the phase spectrum corresponding to the speech frame to be processed to obtain a fourth enhanced result;

[0208] Perform energy compensation on the fourth enhanced result according to the control coefficient to obtain a compensated enhanced result;

[0209] Perform inverse time-frequency conversion on the compensated enhanced result to obtain an enhanced speech frame as the target speech frame.

[0210] In this embodiment, the gain control process is performed based on the amplitude spectrum of the speech frame to be processed. Specifically, the device for executing this method, in addition to obtaining the spectral coefficients of the speech frame to be processed, also calculates the amplitude spectrum and the phase spectrum of the speech frame to be processed. After obtaining the first gain and the second gain, perform gain control on the amplitude spectrum corresponding to the speech frame to be processed according to the first gain and the second gain, and obtain an enhanced amplitude spectrum. Subsequently, reuse the phase spectrum of the speech frame to be processed and calculate it together with the enhanced amplitude spectrum to obtain the fourth enhanced result corresponding to the speech frame to be processed. Perform energy compensation according to the fourth enhanced result and the control coefficient to obtain a compensated enhanced result, and perform inverse time-frequency conversion on the compensated enhanced result to obtain an enhanced speech frame as the target speech frame.

[0211] Specifically, for the convenience of introduction, please refer to Figure 12 , Figure 12 which is a schematic diagram of the overall process of another solution in the embodiments of the present application. As Figure 12As shown, the input of the solution is the speech frame s(n) to be processed. The speech frame s(n) is subjected to STFT time-frequency transformation to obtain the spectral coefficients S(n). When calculating S(n), the amplitude spectrum M(n) and the phase spectrum Ph(n) of the speech frame s(n) to be processed are also calculated. Based on the spectral coefficients S(n), a preprocessing neural network is called to obtain the audio feature vector F(n). Based on the audio feature vector F(n), a first neural network is called to obtain the first gain g11(n), and a second neural network is called to obtain the second gain g12(n). When performing gain control, the first gain g11(n) and the second gain g12(n) are jointly used to perform a first spectral control (i.e., gain control) on the amplitude spectrum M(n), so as to obtain the output first enhanced amplitude spectrum M_e1(n). Combining the first enhanced amplitude spectrum M_e1(n) with the phase spectrum Ph(n) can obtain the first enhanced spectrum S_e1(n). The first spectral control is mainly used to suppress the noise in the speech frame. In particular, the input of the preprocessing neural network may further include the spectral coefficients S_pre(n) of the historical frames (such as the n-1th, n-2th frames, etc., and so on) of the speech frame s(n). The processing process of the third neural network can be executed in parallel with the processes of the first neural network and the second neural network. Based on the audio feature vector F(n), a third neural network is called to obtain the control coefficient g2(n). Applying the control coefficient g2(n) to the first enhanced spectrum S_e1(n) can obtain the second enhanced spectrum S_e2(n). The second spectral control is mainly used for energy compensation. Finally, inverse time-frequency conversion is performed according to the second enhanced spectrum S_e2(n) to obtain the enhanced and supplemented signal s_e(n) corresponding to the speech frame.

[0212] In the solution of the present application, the process of performing gain control according to the amplitude spectrum provides another specific implementation manner for the solution of the present application, improving the diversity of the solution.

[0213] In an embodiment of the present application, based on the above technical solution, the method further includes:

[0214] Calculating the amplitude spectrum and phase spectrum corresponding to the speech frame to be processed according to the speech frame to be processed;

[0215] The above step S650, performing gain control on the speech frame to be processed according to the first gain, the second gain and the control coefficient to obtain the target speech frame, includes:

[0216] Performing gain control on the amplitude spectrum corresponding to the speech frame to be processed according to the first gain and the second gain to obtain the enhanced amplitude spectrum;

[0217] Performing energy compensation on the enhanced amplitude spectrum according to the control coefficient to obtain the compensated amplitude spectrum;

[0218] Merge the compensated magnitude spectrum with the phase spectrum corresponding to the speech frame to be processed to obtain a compensated enhancement result;

[0219] Perform inverse time-frequency conversion based on the compensated enhancement result to obtain the enhanced speech frame as the target speech frame.

[0220] In this embodiment, gain control and energy compensation processes are performed based on the magnitude spectrum of the speech frame to be processed. Specifically, the device for executing this method, in addition to obtaining the spectral coefficients of the speech frame to be processed, also calculates the magnitude spectrum and phase spectrum of the speech frame to be processed. And after obtaining the first gain and the second gain, according to the first gain and the second gain, perform gain control on the magnitude spectrum corresponding to the speech frame to be processed to obtain the enhanced magnitude spectrum. Subsequently, during the energy compensation process, perform energy compensation on the enhanced magnitude spectrum according to the control coefficient to obtain the compensated magnitude spectrum. Then reuse the phase spectrum of the speech frame to be processed and merge it with the compensated magnitude spectrum to obtain the compensated enhancement result. Finally, perform inverse time-frequency conversion based on the compensated enhancement result to obtain the enhanced speech frame as the target speech frame.

[0221] Specifically, for the convenience of introduction, please refer to Figure 13 , Figure 13 which is a schematic diagram of the overall process of another solution in the embodiment of the present application. As Figure 13As shown, the input of the solution is the speech frame s(n) to be processed. The STFT time-frequency transformation is used for the speech frame s(n) to obtain the spectral coefficients S(n). When calculating S(n), the amplitude spectrum M(n) and the phase spectrum Ph(n) of the speech frame s(n) to be processed are also calculated. Based on the spectral coefficients S(n), a preprocessing neural network is called to obtain the audio feature vector F(n). Based on the audio feature vector F(n), a first neural network is called to obtain the first gain g11(n), and a second neural network is called to obtain the second gain g12(n). When performing gain control, the first gain g11(n) and the second gain g12(n) are jointly used for the first spectral control (i.e., gain control) of the amplitude spectrum M(n), so as to obtain the output first enhanced amplitude spectrum M_e1(n). The first spectral control is mainly used to suppress the noise in the speech frame. In particular, the input of the preprocessing neural network may also include the spectral coefficients S_pre(n) of the historical frames (such as the n-1, n-2 frames, etc., and so on) of the speech frame s(n). The processing process of the third neural network can be executed in parallel with the processes of the first neural network and the second neural network. Based on the audio feature vector F(n), a third neural network is called to obtain the control coefficient g2(n). Applying the control coefficient g2(n) to the first enhanced amplitude spectrum M_e1(n) can obtain the second enhanced amplitude spectrum M_e2(n). The second spectral control is mainly used for energy compensation. Finally, according to the combination of the second enhanced amplitude spectrum M_e2(n) and the phase spectrum Ph(n), an inverse time-frequency transformation is performed to obtain the enhanced and supplemented signal s_e(n) corresponding to the speech frame.

[0222] It can be understood that in this embodiment, the control coefficient output by the third neural network only controls the amplitude spectrum. Therefore, the dimension of the control coefficient output by it can be different. Specifically, please refer to Figure 14 , Figure 14 which is another structural schematic diagram of the third neural network shown according to a specific embodiment. As Figure 14 shown, the preprocessing neural network includes 5 FC layers. The input F(n) is represented by spectral coefficients of [128,1]. Figure 14 The dimension of the variable output by each convolutional layer in Figure 14 is marked. The first FC layer outputs a variable of [2048,1] dimension, the second FC layer outputs a variable of [1024,1] dimension, the third FC layer outputs a variable of [1024,1] dimension, the fourth FC layer outputs a variable of [512,1] dimension, and the g2(n) output by the fifth FC layer is of [257,1] dimension. The variable output by the LSTM addition is the audio feature vector F(n). It should be understood that

[0223] It should be noted that although the steps of the method in this application are described in a specific order in the accompanying drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.

[0224] The following introduces the device implementation of this application, which can be used to execute the voice processing method in the above embodiments of this application. Figure 15 The block diagram showing the composition of the voice processing device in the embodiments of this application is schematically shown. As Figure 15 shown, the voice processing device 1500 mainly may include:

[0225] A spectral coefficient acquisition module 1510, configured to acquire the spectral coefficients of the voice frame to be processed;

[0226] A vector acquisition module 1520, configured to perform feature calculation based on the spectral coefficients of the voice frame to be processed to obtain the audio feature vector of the voice frame to be processed.

[0227] A glottal gain module 1530, configured to perform glottal gain calculation based on the audio feature vector to obtain a first gain, where the first gain corresponds to the glottal feature of the voice frame to be processed;

[0228] An excitation gain module 1540, configured to perform excitation gain calculation based on the audio feature vector to obtain a second gain, where the second gain corresponds to the excitation signal of the voice frame to be processed;

[0229] A compensation prediction module 1550, configured to perform compensation prediction based on the audio feature vector to obtain a control coefficient, where the control coefficient is determined according to the energy of the spectral coefficients of the audio frame to be processed;

[0230] A gain control module 1560, configured to perform gain control on the voice frame to be processed according to the first gain, the second gain, and the control coefficient to obtain the target voice frame.

[0231] In some embodiments of this application, based on the above technical solution, the glottal gain module 1530 includes:

[0232] A first neural network sub-module, configured to input the audio feature vector into a first neural network, where the first neural network is trained according to the glottal features corresponding to the noisy voice frames and the glottal features corresponding to the denoised voice frames corresponding to the noisy voice frames;

[0233] A glottal gain prediction sub-module, configured to perform gain prediction based on the audio feature vector through the first neural network to obtain the first gain.

[0234] In some embodiments of the present application, based on the above technical solutions, the glottal gain prediction sub-module includes:

[0235] A gain calculation unit, configured to calculate the gain of the audio feature vector through a first neural network to obtain the first glottal gain corresponding to each sub-band in the to-be-processed speech frame, where the sub-band corresponds to at least one frequency band in the to-be-processed speech frame;

[0236] A gain generation unit, configured to merge the first glottal gains corresponding to each sub-band as the first gain.

[0237] In some embodiments of the present application, based on the above technical solutions, the speech processing device 1500 includes:

[0238] A gain analysis unit, configured to perform prediction and analysis on the audio feature vector and the pitch period of the to-be-processed speech frame through a first neural network to determine the second glottal gain, where the second glottal gain corresponds to the long-term correlation feature of the audio feature vector;

[0239] The gain generation unit includes:

[0240] A gain merging sub-unit, configured to merge the first glottal gains corresponding to each sub-band and the second glottal gain as the first gain.

[0241] In some embodiments of the present application, based on the above technical solutions, the glottal gain prediction sub-module includes:

[0242] A first parameter prediction unit, configured to perform parameter prediction according to the audio feature vector through a first neural network to obtain the first glottal parameter, where the first glottal parameter is used to represent the short-term correlation feature of the audio feature vector;

[0243] A first gain prediction unit, configured to perform gain prediction according to the first glottal parameter through a first neural network to obtain a first prediction result;

[0244] A gain determination unit, configured to determine the first gain according to the first prediction result.

[0245] In some embodiments of the present application, based on the above technical solutions, the glottal gain prediction sub-module further includes:

[0246] A second parameter prediction unit, configured to perform parameter prediction according to the audio feature vector and the pitch period of the to-be-processed speech frame through a first neural network to obtain the second glottal parameter, where the second glottal parameter is used to represent the long-term correlation feature of the audio feature vector;

[0247] A second gain prediction unit, configured to perform gain prediction according to the second glottal parameter through a first neural network to obtain a second prediction result;

[0248] The gain determination unit includes:

[0249] A prediction result merging sub-unit, configured to merge the first prediction result and the second prediction result to determine a first gain.

[0250] In some embodiments of the present application, based on the above technical solution, the excitation gain module 1540 includes:

[0251] A second neural network sub-module, configured to input an audio feature vector into a second neural network, where the second neural network is trained according to the excitation signal of a noise speech frame and the excitation signal of the denoised speech frame corresponding to the noise speech frame;

[0252] An excitation gain prediction sub-module, configured to perform gain prediction on the excitation signal corresponding to the audio feature vector through the second neural network to obtain a second gain.

[0253] In some embodiments of the present application, based on the above technical solution, the vector acquisition module 1520 includes:

[0254] A feature calculation sub-module, configured to input the spectral coefficients of a speech frame to be processed into a preprocessing neural network for feature calculation to obtain an audio feature vector of the speech frame to be processed, where the preprocessing neural network is trained according to the spectral coefficients of a noise speech frame and the spectral coefficients of the denoised speech frame corresponding to the noise speech frame.

[0255] In some embodiments of the present application, based on the above technical solution, the speech processing device 1500 further includes:

[0256] A historical spectral coefficient acquisition module, configured to acquire the spectral coefficients of historical speech frames of the speech frame to be processed;

[0257] The feature calculation sub-module includes:

[0258] A feature vector calculation unit, configured to input the spectral coefficients of the speech frame to be processed and the spectral coefficients of the historical speech frames into a preprocessing neural network for feature calculation to obtain an audio feature vector of the speech frame to be processed.

[0259] In some embodiments of the present application, based on the above technical solution, the gain control module 1560 includes:

[0260] A first enhancement sub-module, configured to enhance the speech frame to be processed according to the second gain to obtain a first enhancement result;

[0261] A second enhancement sub-module, configured to perform gain calculation on each sub-band in the first enhancement result according to the first gain to obtain a second enhancement result;

[0262] An energy compensation sub-module, configured to perform energy compensation on the second enhancement result according to a control coefficient to obtain a third enhancement result;

[0263] An inverse time-frequency conversion sub-module, configured to perform inverse time-frequency conversion on the third enhancement result to obtain an enhanced speech frame as a target speech frame.

[0264] In some embodiments of the present application, based on the above technical solutions, the speech processing device 1500 further includes:

[0265] A first amplitude spectrum calculation module, configured to calculate the amplitude spectrum and phase spectrum corresponding to the speech frame to be processed according to the speech frame to be processed;

[0266] The gain control module 1560 includes:

[0267] A first amplitude spectrum gain control sub-module, configured to perform gain control on the amplitude spectrum corresponding to the speech frame to be processed according to a first gain and a second gain to obtain an enhanced amplitude spectrum;

[0268] A first phase spectrum merging sub-module, configured to merge according to the enhanced amplitude spectrum and the phase spectrum corresponding to the speech frame to be processed to obtain a fourth enhancement result;

[0269] A first amplitude spectrum energy compensation sub-module, configured to perform energy compensation on the fourth enhancement result according to the control coefficient to obtain a compensated enhancement result;

[0270] A first amplitude spectrum inverse time-frequency conversion sub-module, configured to perform inverse time-frequency conversion on the compensated enhancement result to obtain an enhanced speech frame as a target speech frame.

[0271] In some embodiments of the present application, based on the above technical solutions, the speech processing device 1500 further includes:

[0272] A second amplitude spectrum calculation module, configured to calculate the amplitude spectrum and phase spectrum corresponding to the speech frame to be processed according to the speech frame to be processed;

[0273] A second amplitude spectrum gain control sub-module, configured to perform gain control on the amplitude spectrum corresponding to the speech frame to be processed according to a first gain and a second gain to obtain an enhanced amplitude spectrum;

[0274] A second amplitude spectrum energy compensation sub-module, configured to perform energy compensation on the enhanced amplitude spectrum according to the control coefficient to obtain a compensated amplitude spectrum;

[0275] A second phase spectrum merging sub-module, configured to merge the compensated amplitude spectrum with the phase spectrum corresponding to the speech frame to be processed to obtain a compensated enhancement result;

[0276] The second amplitude spectrum inverse time-frequency conversion sub-module is used to perform inverse time-frequency conversion according to the compensation enhancement result to obtain an enhanced speech frame as the target speech frame.

[0277] It should be noted that the device provided in the above embodiment and the method provided in the above embodiment belong to the same concept. The specific manners in which each module performs operations have been described in detail in the method embodiment, and will not be elaborated here.

[0278] Figure 16 The structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown.

[0279] It should be noted that Figure 16 The computer system 1600 of the electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0280] As Figure 16 shown, the computer system 1600 includes a central processing unit (CPU) 1601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1602 or the program loaded from the storage section 1608 into the random access memory (RAM) 1603. In the RAM 1603, various programs and data required for system operations are also stored. The CPU 1601, ROM 1602, and RAM 1603 are connected to each other via a bus 1604. The input / output (I / O) interface 1605 is also connected to the bus 1604.

[0281] The following components are connected to the I / O interface 1605: an input section 1606 including a keyboard, a mouse, etc.; an output section 1607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1608 including a hard disk, etc.; and a communication section 1609 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1609 performs communication processing via a network such as the Internet. A drive 1610 is also connected to the I / O interface 1605 as needed. A removable medium 1611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1610 as needed so that the computer program read from it can be installed into the storage section 1608 as needed.

[0282] In particular, according to an embodiment of the present application, the processes described in each method flowchart can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 1609 and / or installed from the removable medium 1611. When the computer program is executed by the central processing unit (CPU) 1601, various functions defined in the system of the present application are performed.

[0283] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0284] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, and the above-mentioned module, segment of a program, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0285] It should be noted that although several modules or units of devices for performing actions are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0286] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0287] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.

[0288] It should be understood that the present application is not limited to the exact structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A voice processing method, characterized in that, Including: Obtain the spectral coefficients of the speech frame to be processed; Perform feature calculation based on the spectral coefficients of the speech frame to be processed to obtain the audio feature vector of the speech frame to be processed; Input the audio feature vector into a first neural network, where the first neural network is trained based on the glottal features corresponding to the noisy speech frame and the glottal features corresponding to the denoised speech frame corresponding to the noisy speech frame; Perform gain prediction based on the audio feature vector through the first neural network to obtain a first gain, where the first gain corresponds to the glottal features of the speech frame to be processed; Perform excitation gain calculation based on the audio feature vector to obtain a second gain, where the second gain corresponds to the excitation signal of the speech frame to be processed; Perform compensation prediction based on the audio feature vector to obtain a control coefficient, where the control coefficient is determined according to the energy of the spectral coefficients of the speech frame to be processed; Perform gain control on the speech frame to be processed according to the first gain, the second gain, and the control coefficient to obtain a target speech frame; Wherein, the performing gain prediction based on the audio feature vector through the first neural network to obtain a first gain includes: Perform gain calculation on the audio feature vector through the first neural network to obtain the first glottal gain corresponding to each sub-band in the speech frame to be processed, where the sub-band corresponds to at least one frequency band in the speech frame to be processed; Merge the first glottal gains corresponding to the respective sub-bands as the first gain.

2. The method according to claim 1, characterized in that, The method further includes: Perform prediction analysis on the audio feature vector and the fundamental period of the speech frame to be processed through the first neural network to determine a second glottal gain, where the second glottal gain corresponds to the long-term correlation feature of the audio feature vector; The merging the first glottal gains corresponding to the respective sub-bands as the first gain includes: Merge the first glottal gains corresponding to the respective sub-bands and the second glottal gain as the first gain.

3. The method according to claim 1, characterized in that, The performing gain prediction based on the audio feature vector through the first neural network to obtain a first gain further includes: Perform parameter prediction based on the audio feature vector through the first neural network to obtain a first glottal parameter, where the first glottal parameter is used to represent the short-term correlation feature of the audio feature vector; Perform gain prediction based on the first glottal parameter through the first neural network to obtain a first prediction result; Determine the first gain according to the first prediction result.

4. The method according to claim 3, wherein The method further includes: Perform parameter prediction based on the audio feature vector and the fundamental period of the speech frame to be processed through the first neural network to obtain a second glottal parameter, where the second glottal parameter is used to represent the long-term correlation feature of the audio feature vector; Perform gain prediction based on the second glottal parameter through the first neural network to obtain a second prediction result; The determining the first gain according to the first prediction result includes: Merge the first prediction result and the second prediction result to determine the first gain.

5. The method according to claim 1, wherein Performing excitation gain calculation based on the audio feature vector to obtain a second gain, including: Inputting the audio feature vector into a second neural network, where the second neural network is trained based on the excitation signal of a noisy speech frame and the excitation signal of the denoised speech frame corresponding to the noisy speech frame; Predicting a gain based on the excitation signal corresponding to the audio feature vector through the second neural network to obtain the second gain.

6. The method according to claim 1, wherein Performing feature calculation based on the spectral coefficients of the speech frame to be processed to obtain the audio feature vector of the speech frame to be processed, including: Inputting the spectral coefficients of the speech frame to be processed into a preprocessing neural network for feature calculation to obtain the audio feature vector of the speech frame to be processed, where the preprocessing neural network is trained based on the spectral coefficients of a noisy speech frame and the spectral coefficients of the denoised speech frame corresponding to the noisy speech frame.

7. The method according to claim 6, characterized in that, The method further includes: Obtaining the spectral coefficients of the historical speech frames of the speech frame to be processed; The step of inputting the spectral coefficients of the speech frame to be processed into a preprocessing neural network for feature calculation to obtain the audio feature vector of the speech frame to be processed includes: Inputting the spectral coefficients of the speech frame to be processed and the spectral coefficients of the historical speech frames into the preprocessing neural network for feature calculation to obtain the audio feature vector of the speech frame to be processed.

8. The method according to claim 1, characterized in that Performing gain control on the speech frame to be processed based on the first gain, the second gain, and the control coefficient to obtain a target speech frame, including: Enhancing the speech frame to be processed according to the second gain to obtain a first enhancement result; Performing gain operation on each subband in the first enhancement result according to the first gain to obtain a second enhancement result; Performing energy compensation on the second enhancement result according to the control coefficient to obtain a third enhancement result; Performing inverse time-frequency conversion on the third enhancement result to obtain the enhanced speech frame as the target speech frame.

9. The method according to claim 1, wherein The method further includes: Calculating the magnitude spectrum and phase spectrum corresponding to the speech frame to be processed according to the speech frame to be processed; Performing gain control on the speech frame to be processed based on the first gain, the second gain, and the control coefficient to obtain a target speech frame, including: Performing gain control on the magnitude spectrum corresponding to the speech frame to be processed according to the first gain and the second gain to obtain an enhanced magnitude spectrum; Merging the enhanced magnitude spectrum and the phase spectrum corresponding to the speech frame to be processed to obtain a fourth enhancement result; Performing energy compensation on the fourth enhancement result according to the control coefficient to obtain a compensated enhancement result; Performing inverse time-frequency conversion on the compensated enhancement result to obtain the enhanced speech frame as the target speech frame.

10. The method according to claim 1, characterized in that, The method further includes: Calculating the magnitude spectrum and phase spectrum corresponding to the speech frame to be processed according to the speech frame to be processed; Performing gain control on the speech frame to be processed based on the first gain, the second gain, and the control coefficient to obtain a target speech frame, including: Perform gain control on the amplitude spectrum corresponding to the speech frame to be processed according to the first gain and the second gain to obtain an enhanced amplitude spectrum; Perform energy compensation on the enhanced amplitude spectrum according to the control coefficient to obtain a compensated amplitude spectrum; Merge the compensated amplitude spectrum with the phase spectrum corresponding to the speech frame to be processed to obtain a compensated enhancement result; Perform inverse time-frequency conversion according to the compensated enhancement result to obtain an enhanced speech frame as the target speech frame.

11. A voice processing device, characterized in that, Comprising: A spectrum coefficient acquisition module for acquiring the spectrum coefficients of the speech frame to be processed; A vector acquisition module for calculating features according to the spectrum coefficients of the speech frame to be processed to obtain an audio feature vector of the speech frame to be processed; A first neural network sub-module that inputs the audio feature vector into a first neural network, and the first neural network is trained according to the glottal features corresponding to the noisy speech frame and the glottal features corresponding to the denoised speech frame corresponding to the noisy speech frame; A glottal gain prediction sub-module for predicting a gain through the first neural network according to the audio feature vector to obtain a first gain, and the first gain corresponds to the glottal features of the speech frame to be processed; An excitation gain module for calculating an excitation gain according to the audio feature vector to obtain a second gain, and the second gain corresponds to the excitation signal of the speech frame to be processed; A compensation prediction module for predicting a compensation according to the audio feature vector to obtain a control coefficient, and the control coefficient is determined according to the energy of the spectrum coefficients of the speech frame to be processed; A gain control module for performing gain control on the speech frame to be processed according to the first gain, the second gain and the control coefficient to obtain a target speech frame; Wherein, the predicting a gain through the first neural network according to the audio feature vector to obtain a first gain includes: Performing gain calculation on the audio feature vector through the first neural network to obtain a first glottal gain corresponding to each sub-band in the speech frame to be processed, wherein the sub-band corresponds to at least one frequency band in the speech frame to be processed; Merging the first glottal gains corresponding to the respective sub-bands as the first gain.

12. An electronic device, characterized in that, Comprising: A processor; A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the method for speech processing according to any one of claims 1 to 10 by executing the executable instructions.

13. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the speech processing method according to any one of claims 1 to 10.

14. A computer program product, characterized in that, The computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the speech processing method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Voice processing method, device and equipment and storage medium

    CN111554322A

  • Voice gain control method and computer storage medium

    CN112242147A