A clean speech reconstruction method, device, equipment and medium

By extracting deep feature information, determining the fundamental frequency signal and building a filter, the problem that the voice signal after denoising in the prior art cannot match the speaker's voice characteristics, and a clean speech reconstruction that is more suitable for natural speech is achieved.

CN114495965BActive Publication Date: 2025-05-13COMMUNICATION UNIVERSITY OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210111339.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-29
Publication Date
2025-05-13
Estimated Expiration
2042-01-29

AI Technical Summary

Technical Problem

The voice signal after denoising in the prior art may not fit the voice characteristics of the speaker when making a voice.

Method used

By extracting deep feature information from the noisy voice to be enhanced, determining the fundamental frequency information of the voice content, and constructing a filter matching the noisy voice, simulating the modulation of the sound wave signal by the sound channel, and finally reconstructing the clean voice.

Benefits of technology

The reconstructed clean voice is more in line with the voice characteristics of the speaker when he speaks, improving the quality and reliability of voice enhancement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114495965B_ABST
    Figure CN114495965B_ABST
Patent Text Reader

Abstract

The present application provides a method, device, equipment and medium for reconstructing clean speech, the method comprising: extracting deep feature information from the noisy speech to be enhanced, determining the fundamental frequency information of the speech content in the noisy speech according to the deep feature information, constructing a filter matching the noisy speech based on the deep feature information, and reconstructing the clean speech according to the filter matching the fundamental frequency signal and the noisy speech. The present application takes into account the vocalization principle of the human vocal system, extracts deep feature information that can characterize the speech characteristics of the speaker when emitting the clean speech corresponding to the noisy speech, determines the fundamental frequency information of the speech content in the noisy speech according to the deep feature information, so as to simulate the sound wave signal generated by the vibration of the vocal cords, constructs a filter matching the noisy speech according to the deep feature information, so as to simulate the vocal channel that modulates the sound wave signal, and reconstructs the clean speech that is more in line with the speech characteristics of the speaker when speaking based on the filter matching the fundamental frequency signal and the noisy speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of speech signal processing, and in particular to a clean speech reconstruction method, device, equipment and medium. Background Art

[0002] Speech enhancement refers to the technology of extracting target speech information from a noisy background and improving its intelligibility when the speech signal is interfered by various noises. The practical application of speech enhancement technology is extensive and in-depth. It is needed in communications in noisy environments in daily life and military communications in harsh environments. In addition, speech coding and speech recognition often need to be carried out under laboratory conditions, that is, in an environment without noise or with a very high signal-to-noise ratio. Therefore, in order to promote speech coding and speech recognition, speech enhancement processing is also needed.

[0003] Current speech enhancement methods usually denoise the phase spectrum of noisy speech, that is, in the rough processing stage, the time domain signal or spectrum amplitude information is denoised, the noisy phase information is retained, and the result of the rough processing stage is used as input in the tuning stage. At the same time, the amplitude information and phase information of the frequency spectrum are jointly denoised to refine the enhancement result, thereby improving the subjective auditory experience of the speech enhancement result. However, this method of denoising the phase spectrum of noisy speech only treats the noisy speech as a segment of electronic signal for processing. The processing process ignores the physical movement laws of the vocal cords and the modulation laws of the vocal tract on the sound waves during human vocalization, resulting in the denoised speech signal may not be well matched to the speech characteristics of the speaker when speaking. Summary of the invention

[0004] In view of this, the present application provides a clean speech reconstruction method, device, equipment and medium for solving the problem in the prior art that the denoised speech signal may not be well matched to the speech characteristics of the speaker when speaking. The technical solution is as follows:

[0005] A clean speech reconstruction method, characterized by comprising:

[0006] Extracting deep feature information from the noisy speech to be enhanced, wherein the deep feature information represents the speech characteristics of the speaker when uttering the clean speech corresponding to the noisy speech, and the clean speech refers to the speech information in a noise-free environment composed of the speech content of the noisy speech;

[0007] Determine the fundamental frequency information of the speech content in the noisy speech according to the deep feature information, wherein the fundamental frequency signal is used to characterize the sound wave signal generated by the vibration of the vocal cords;

[0008] Constructing a filter matched to the noisy speech according to the deep feature information, wherein the filter matched to the noisy speech is used to characterize a vocal channel that modulates the sound wave signal;

[0009] Based on the baseband signal and the filter matching the noisy speech, a clean speech corresponding to the noisy speech is reconstructed.

[0010] Optionally, determining the fundamental frequency signal of the speech content in the noisy speech according to the deep feature information includes:

[0011] The deep feature information is input into an ordinary differential equation filter solving network to obtain a group of ordinary differential equations corresponding to the speech content as the fundamental frequency signal, wherein the ordinary differential equation filter solving network is trained based on the deep feature information extracted from the training speech.

[0012] Optionally, constructing a filter matching the noisy speech according to the deep feature information includes:

[0013] Determining the optimal order and optimal coefficient for matching the noisy speech according to the deep feature information;

[0014] A filter matching the noisy speech is constructed by using the optimal order and optimal coefficient.

[0015] Optionally, the filter matched to the noisy speech is a linear prediction filter.

[0016] Optionally, the noisy speech is a frequency domain signal;

[0017] The reconstructing a clean speech corresponding to the noisy speech based on a filter matching the baseband signal and the noisy speech includes:

[0018] Multiplying the baseband signal with the filter matched to the noisy speech to obtain a product result;

[0019] The product result is subjected to inverse Fourier transform to obtain the reconstructed clean speech.

[0020] A clean speech reconstruction device comprises: a feature extraction module, a fundamental frequency signal determination module, a filter construction module and a speech reconstruction module;

[0021] The feature extraction module is used to extract deep feature information from the noisy speech to be enhanced, wherein the deep feature information represents the speech characteristics of the speaker when uttering the clean speech corresponding to the noisy speech, and the clean speech refers to the speech information in a noise-free environment composed of the speech content of the noisy speech;

[0022] The fundamental frequency signal determination module is used to determine the fundamental frequency information of the speech content in the noisy speech according to the deep feature information, wherein the fundamental frequency signal is used to characterize the sound wave signal generated by the vibration of the vocal cords;

[0023] The filter construction module is used to construct a filter matching the noisy speech according to the deep feature information, wherein the filter matching the noisy speech is used to characterize the vocal channel that modulates the sound wave signal;

[0024] The speech reconstruction module is used to reconstruct the clean speech corresponding to the noisy speech based on the baseband signal and the filter matched with the noisy speech.

[0025] Optionally, the fundamental frequency signal determination module is specifically used to input the deep feature information into an ordinary differential equation filter solving network to obtain a group of ordinary differential equations corresponding to the speech content as the fundamental frequency signal, wherein the ordinary differential equation filter solving network is trained based on the deep feature information extracted from the training speech.

[0026] Optionally, the filter construction module includes: an order coefficient prediction submodule and a filter construction submodule;

[0027] The order coefficient prediction submodule is used to determine the optimal order and optimal coefficient matching the noisy speech according to the deep feature information;

[0028] The filter construction submodule is used to construct a filter matching the noisy speech through the optimal order and optimal coefficient.

[0029] A clean speech reconstruction device, comprising a memory and a processor;

[0030] The memory is used to store programs;

[0031] The processor is used to execute the program to implement each step of the clean speech reconstruction method as described in any one of the above items.

[0032] A readable storage medium stores a computer program, and when the computer program is executed by a processor, the computer program implements the steps of any of the above-mentioned clean speech reconstruction methods.

[0033] Through the above technical solutions, it can be known that the clean speech reconstruction method provided by the present application takes into account the vocalization principle of the human vocal system, first extracts deep feature information from the noisy speech to be enhanced, and then determines the fundamental frequency information of the speech content in the noisy speech based on the deep feature information, thereby simulating the sound wave signal generated by the vibration of the vocal cords in the human vocal system, and constructs a filter matching the noisy speech based on the deep feature information, thereby simulating the vocal channel that modulates the sound wave signal in the human vocal system, and finally reconstructs the clean speech corresponding to the noisy speech based on the filter matching the fundamental frequency signal and the noisy speech. Because the present application first extracts the deep feature information that can characterize the speech characteristics of the speaker when emitting the clean speech corresponding to the noisy speech at the beginning of the processing, and then reconstructs the clean speech corresponding to the noisy speech based on the deep feature information, and the reconstruction process simulates the human vocal system, so that the clean speech reconstructed by the present application is more in line with the speech characteristics of the speaker when speaking. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0035] Figure 1 A flow chart of a clean speech reconstruction method provided in an embodiment of the present application;

[0036] Figure 2 A schematic diagram of the structure of the vocal cord channel simulation speech enhancement system provided in an embodiment of the present application;

[0037] Figure 3 A schematic diagram of the structure of a clean speech reconstruction device provided in an embodiment of the present application;

[0038] Figure 4 This is a hardware structure block diagram of the clean speech reconstruction device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0039] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0040] In view of the problems existing in the prior art, the inventor of this case conducted in-depth research and came up with the idea that the powerful fitting ability of neural networks can be used to simulate the vocalization process of the human vocal system, that is, to establish an acoustic enhancement or recognition system based on neural networks that is more in line with human vocalization habits and rules, so that the enhanced noisy speech (also called noisy speech) is closer to the sound produced by humans. Based on this, this application is proposed.

[0041] Before introducing the present application in detail, a brief introduction to the human vocal system is first given to help those skilled in the art better understand the present application.

[0042] The human vocal system includes the glottis (vocal cords) that produce vibration excitation and the vocal tract that modulates sound waves. The vocal cords produce different sounds depending on their state. When the vocal cords open and close periodically, the air flows through the tight vocal cords, generating a quasi-periodic pulse of air flow. The sound waves radiated through the lips are voiced sounds (voiced sounds); when the vocal cords are fully stretched, the air passes through the glottis without hindrance and is only modulated by the oral cavity. The sound produced is called friction sound or unvoiced sound (voiceless sound). Different types of sounds may only be reflected in the different energy levels on the spectrum, but they are very different in the human ear. If the overall vocal cord model can be effectively fitted, different types of sounds can be effectively distinguished and targeted processing can be made.

[0043] Next, the clean speech reconstruction method provided by the present application is described in detail through the following embodiments. Figure 1 , shows a flow chart of a clean speech reconstruction method provided by an embodiment of the present application, and the clean speech reconstruction method may include:

[0044] Step S101: extract deep feature information from the noisy speech to be enhanced.

[0045] The deep feature information represents the speech characteristics of the speaker when uttering the clean speech corresponding to the noisy speech, and the clean speech refers to the speech information in a noise-free environment composed of the speech content of the noisy speech.

[0046] In this step, different noisy speech corresponds to different speech characteristics, so the deep feature information extracted is different.

[0047] Optional, see Figure 2 The structural diagram of the vocal cord channel simulation speech enhancement system shown in FIG. 1 , in this step, deep feature information can be extracted from noisy speech through a deep network. After the noisy speech is input into the deep network, the deep network can output comprehensive data containing feature information of different dimensions of the speech signal, which is also the deep feature information in this step. Here, the deep network is trained based on the training speech (i.e., the training data).

[0048] The above-mentioned "feature information of different dimensions" refers to the deep network's deep representation of data features that is representative after multiple extractions and classifications. Optionally, if the noisy speech input to the deep network is a frequency domain signal, since the deep network processes a complex spectrum, the "feature information of different dimensions" includes the amplitude and phase information in the spectrum obtained after the noisy speech is subjected to Fourier transform (stft).

[0049] Step S102: Determine fundamental frequency information of speech content in the noisy speech according to the deep feature information.

[0050] The baseband signal is used to characterize the sound wave signal generated by the vibration of the vocal cords. The baseband signal is equivalent to the excitation signal emitted by the speaker's vocal cords, and the speaker's voice signal can be obtained after the vocal tract modulates the excitation signal.

[0051] Optionally, this step can determine the fundamental frequency signal of the speech content in the noisy speech by solving the ordinary differential equation filter network, see Figure 2 As shown in the "ordinary differential equation solving" branch, this step inputs the deep feature information into the ordinary differential equation filter solving network to obtain the ordinary differential equation group corresponding to the speech content. The ordinary differential equation group refers to the combination of equations for the vibration of the vocal cords, that is, the combination of equations for the vibration of the vocal cords when the speaker makes a noisy speech.

[0052] Here, in the ordinary differential equation filter solving network, a deep neural network and a small learning rate are used to solve the ordinary differential equation, and the ordinary differential equation filter solving network is trained based on the deep feature information extracted from the training speech. Specifically, when training the network, an adjoint sensitivity method can be used for training, wherein an ordinary differential equation solver (ODE Solver, which can be regarded as a black box and can be used to solve the integral value of the neural network from the previous moment to the next moment) provides a true value as a reference, and the gradient is calculated by reversely solving the second enhanced ODE (ordinary differential equation) in time. Based on the calculated gradient, the optimal parameters of the ordinary differential equation filter solving network can be trained, thereby realizing the simulation of the OED solver by the neural network.

[0053] This embodiment determines the baseband signal by filtering and solving the network through ordinary differential equations, thereby ensuring the operation efficiency, reducing the calculation burden of the computer, and improving the efficiency and accuracy.

[0054] Step S103: constructing a filter matching the noisy speech based on the deep feature information.

[0055] The filter matched to the noisy speech is used to characterize the vocal channel that modulates the sound wave signal.

[0056] In this step, a filter matching the noisy speech can be constructed based on the deep feature information, so as to simulate the vocal tract (such as throat, mouth, etc.) that modulates the sound wave signal through the filter, that is, to simulate the deformation of the throat, mouth, etc. made by the speaker when making noisy speech.

[0057] Optionally, the filter matched to the noisy speech can be a linear prediction (LPC) filter, which is an analog filter for the sound source and the sound channel. Optionally, the LPC filter is an all-pole filter, and the baseband signal obtained in the previous step can generate clean speech by passing through the all-pole filter.

[0058] The coefficients of the above all-pole filter depend on the shape of the vocal tract of the specific sound being produced. See the following three formulas. In the LPC filter, the LPC algorithm minimizes the difference between the real signal and the predicted signal by converging the mean squared error (MSE) function, and then calculates the partial derivative of each filter coefficient, sets it equal to 0 and solves the linear equations about the filter coefficients to obtain the coefficients of each filter.

[0059]

[0060]

[0061]

[0062] In an optional embodiment, this step can simulate the modulation function of the throat, oral cavity, etc. by using a deep neural network plus LPC coding. Then the process of this step of "constructing a filter matching the noisy speech based on the deep feature information" may include:

[0063] Step S1031: Determine the optimal order and optimal coefficient for matching the noisy speech according to the deep feature information.

[0064] It should be understood that when the noisy speech in the above step S101 is different, the optimal order and optimal coefficient determined in this step may be different.

[0065] Optionally, in the scenario where the filter matching the noisy speech can be a linear prediction filter, this step corresponds to Figure 2 The "LPC order prediction" module shown in the figure inputs the deep feature information into the LPC order prediction module to obtain the optimal order and optimal coefficients matching the noisy speech. Here, the number of optimal coefficients is equal to the optimal order.

[0066] Optionally, this step can determine the optimal order and optimal coefficient that matches the noisy speech through the LPC order prediction network. Specifically, this step can input the deep feature information into the LPC order prediction network, and the LPC order prediction network can feedback the optimal order and optimal coefficient of the LPC filter that best matches the current noisy speech through learning data.

[0067] The above-mentioned LPC order prediction network is a neural network, which can be obtained by training deep feature information extracted from training speech.

[0068] Step S1032: construct a filter matching the noisy speech by using the optimal order and optimal coefficient.

[0069] Optionally, in the scenario where the filter matching the noisy speech can be a linear prediction filter, this step corresponds to Figure 2 The "LPC" module shown in the figure sends the optimal order and optimal coefficient obtained in the previous step to the LPC filter module, so as to construct an LPC filter specially designed for the current noisy speech.

[0070] As mentioned above for the introduction of the "all-pole filter", the coefficients of the all-pole filter depend on the shape of the vocal tract of the specific sound produced. Therefore, based on the optimal order and optimal coefficients determined based on the deep feature information, a filter that matches the noisy speech is determined so that the filter can well simulate the modulation function of the vocal tract.

[0071] Step S104: reconstructing a clean speech corresponding to the noisy speech according to the baseband signal and the filter matched with the noisy speech.

[0072] As explained above, the baseband signal can be used to simulate the sound wave signal generated by the vibration of the vocal cords, and the filter matching the noisy speech can be used to simulate the vocal channel that modulates the sound wave signal. Therefore, the process of this step of "reconstructing the clean speech corresponding to the noisy speech according to the baseband signal and the filter matching the noisy speech" is equivalent to the process of the speaker's vocal channel modulating the sound wave signal emitted by the speaker's vocal cords.

[0073] It can be seen that this step simulates the vocal principle of the human vocal system, that is, this step takes into account the physical movement laws of the vocal cords and the modulation laws of the vocal tract on sound waves during the human vocalization process, so that the clean speech corresponding to the reconstructed noisy speech is more in line with the speech characteristics of the speaker when speaking.

[0074] In an optional embodiment, the noisy speech to be enhanced may be a time domain signal, and the process of the step of "reconstructing the clean speech corresponding to the noisy speech according to the baseband signal and the filter matching the noisy speech" may include: convolving the baseband signal with the filter matching the noisy speech, and using the convolution result as the reconstructed clean speech.

[0075] Considering that noisy speech is a time domain signal, the process of the filter modulating the baseband signal only considers the waveform information of the baseband signal. Compared with considering the phase and amplitude information of the baseband signal, the effect of reconstructing the clean speech is relatively poor.

[0076] Based on this, in a preferred case, the noisy speech to be enhanced may be a frequency domain signal, and the process of the step of "reconstructing the clean speech corresponding to the noisy speech based on the baseband signal and the filter matching the noisy speech" may include: multiplying the baseband signal with the filter matching the noisy speech to obtain a product result; and performing an inverse Fourier transform on the product result to obtain a reconstructed clean speech.

[0077] Since the phase and amplitude of the baseband signal are taken into account when reconstructing clean speech in the frequency domain, the effect of clean speech is better.

[0078] The clean speech reconstruction method provided by the present application takes into account the vocalization principle of the human vocal system, first extracts deep feature information from the noisy speech to be enhanced, and then determines the fundamental frequency information of the speech content in the noisy speech based on the deep feature information, thereby simulating the sound wave signal generated by the vibration of the vocal cords in the human vocal system, and constructs a filter matching the noisy speech based on the deep feature information, thereby simulating the vocal channel that modulates the sound wave signal in the human vocal system, and finally reconstructs the clean speech corresponding to the noisy speech based on the fundamental frequency signal and the filter matching the noisy speech. Because the present application first extracts the deep feature information that can characterize the speech characteristics of the speaker when uttering the clean speech corresponding to the noisy speech at the beginning of the processing, and then reconstructs the clean speech corresponding to the noisy speech based on the deep feature information, and the reconstruction process simulates the human vocal system, the clean speech reconstructed by the present application is more in line with the speech characteristics of the speaker when speaking.

[0079] In addition, it should be noted that the architectural concept provided in this application can also be applied to the research of various speech generation problems such as speech generation, speech conversion, and speaker style conversion.

[0080] An embodiment of the present application also provides a clean speech reconstruction device. The clean speech reconstruction device provided in the embodiment of the present application is described below. The clean speech reconstruction device described below and the clean speech reconstruction method described above can be referenced to each other.

[0081] See also Figure 3 , which shows a schematic diagram of the structure of a clean speech reconstruction device provided in an embodiment of the present application, such as Figure 3 As shown, the clean speech reconstruction device may include: a feature extraction module 301, a fundamental frequency signal determination module 302, a filter construction module 303 and a speech reconstruction module 304.

[0082] The feature extraction module 301 is used to extract deep feature information from the noisy speech to be enhanced, wherein the deep feature information represents the speech characteristics of the speaker when uttering the clean speech corresponding to the noisy speech, and the clean speech refers to the speech information in a noise-free environment composed of the speech content of the noisy speech.

[0083] The fundamental frequency signal determination module 302 is used to determine the fundamental frequency information of the speech content in the noisy speech according to the deep feature information, wherein the fundamental frequency signal is used to represent the sound wave signal generated by the vibration of the vocal cords.

[0084] The filter construction module 303 is used to construct a filter that matches the noisy speech according to the deep feature information, wherein the filter that matches the noisy speech is used to characterize the vocal channel that modulates the sound wave signal.

[0085] The speech reconstruction module 304 is used to reconstruct the clean speech corresponding to the noisy speech based on the baseband signal and the filter matched with the noisy speech.

[0086] The clean speech reconstruction device provided by the present application takes into account the vocalization principle of the human vocal system, first extracts deep feature information from the noisy speech to be enhanced, and then determines the fundamental frequency information of the speech content in the noisy speech based on the deep feature information, thereby simulating the sound wave signal generated by the vibration of the vocal cords in the human vocal system, and constructs a filter matching the noisy speech based on the deep feature information, thereby simulating the vocal channel that modulates the sound wave signal in the human vocal system, and finally reconstructs the clean speech corresponding to the noisy speech based on the fundamental frequency signal and the filter matching the noisy speech. Because the present application first extracts the deep feature information that can characterize the speech characteristics of the speaker when emitting the clean speech corresponding to the noisy speech at the beginning of the processing, and then reconstructs the clean speech corresponding to the noisy speech based on the deep feature information, and the reconstruction process simulates the human vocal system, the clean speech reconstructed by the present application is more in line with the speech characteristics of the speaker when speaking.

[0087] In a possible implementation, the fundamental frequency signal determination module 302 can be specifically used to input the deep feature information into an ordinary differential equation filter solving network to obtain a set of ordinary differential equations corresponding to the speech content as the fundamental frequency signal, wherein the ordinary differential equation filter solving network is trained based on the deep feature information extracted from the training speech.

[0088] In a possible implementation, the filter construction module 303 may include: an order coefficient prediction submodule and a filter construction submodule.

[0089] The order coefficient prediction submodule is used to determine the optimal order and optimal coefficient matching the noisy speech according to the deep feature information.

[0090] The filter construction submodule is used to construct a filter matching the noisy speech through the optimal order and optimal coefficient.

[0091] In a possible implementation, the filter matched to the noisy speech is a linear prediction filter.

[0092] In a possible implementation manner, the noisy speech is a frequency domain signal.

[0093] Then the above-mentioned speech reconstruction module 304 may include: a multiplication module and an inverse Fourier transform module;

[0094] The multiplication module is used to multiply the baseband signal with the filter matched with the noisy speech to obtain a product result;

[0095] The inverse Fourier transform module is used to perform inverse Fourier transform on the product result to obtain the reconstructed clean speech.

[0096] The present application embodiment also provides a clean speech reconstruction device. Optionally, Figure 4 The hardware structure diagram of the clean speech reconstruction device is shown in FIG. Figure 4 , the hardware structure of the clean speech reconstruction device may include: at least one processor 401, at least one communication interface 402, at least one memory 403 and at least one communication bus 404;

[0097] In the embodiment of the present application, the number of the processor 401, the communication interface 402, the memory 403, and the communication bus 404 is at least one, and the processor 401, the communication interface 402, and the memory 403 communicate with each other through the communication bus 404;

[0098] The processor 401 may be a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention;

[0099] The memory 403 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory;

[0100] The memory 403 stores a program, and the processor 401 can call the program stored in the memory 403, and the program is used to:

[0101] Extracting deep feature information from the noisy speech to be enhanced, wherein the deep feature information represents the speech characteristics of the speaker when uttering the clean speech corresponding to the noisy speech, and the clean speech refers to the speech information in a noise-free environment composed of the speech content of the noisy speech;

[0102] Determine the fundamental frequency information of the speech content in the noisy speech according to the deep feature information, wherein the fundamental frequency signal is used to characterize the sound wave signal generated by the vibration of the vocal cords;

[0103] Constructing a filter matched to the noisy speech according to the deep feature information, wherein the filter matched to the noisy speech is used to characterize a vocal channel that modulates the sound wave signal;

[0104] Based on the baseband signal and the filter matching the noisy speech, a clean speech corresponding to the noisy speech is reconstructed.

[0105] Optionally, the detailed functions and extended functions of the program may refer to the above description.

[0106] The embodiment of the present application also provides a readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the clean speech reconstruction method as described above is implemented.

[0107] Optionally, the detailed functions and extended functions of the program may refer to the above description.

[0108] Finally, it should be noted that, in this article, relational terms such as and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprises a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0109] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0110] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A clean speech reconstruction method, characterized in that: include: Extracting deep feature information from the noisy speech to be enhanced, wherein the deep feature information represents the speech characteristics of the speaker when uttering the clean speech corresponding to the noisy speech, and the clean speech refers to the speech information in a noise-free environment composed of the speech content of the noisy speech; Determine the fundamental frequency signal of the speech content in the noisy speech according to the deep feature information, wherein the fundamental frequency signal is used to characterize the sound wave signal generated by the vibration of the vocal cords; Constructing a filter matched to the noisy speech based on the deep feature information, wherein the filter matched to the noisy speech is used to characterize a vocal channel that modulates the sound wave signal; A clean speech corresponding to the noisy speech is reconstructed according to the filter matching the baseband signal and the noisy speech.

2. The clean speech reconstruction method according to claim 1, characterized in that: The determining the fundamental frequency signal of the speech content in the noisy speech according to the deep feature information includes: The deep feature information is input into an ordinary differential equation filter solving network to obtain a group of ordinary differential equations corresponding to the speech content as the fundamental frequency signal, wherein the ordinary differential equation filter solving network is trained based on the deep feature information extracted from the training speech.

3. The clean speech reconstruction method according to claim 1, characterized in that: The step of constructing a filter matching the noisy speech based on the deep feature information includes: Determining the optimal order and optimal coefficient for matching the noisy speech according to the deep feature information; A filter matching the noisy speech is constructed by using the optimal order and optimal coefficient.

4. The clean speech reconstruction method according to claim 3, characterized in that: The filter matched to the noisy speech is a linear prediction filter.

5. The clean speech reconstruction method according to claim 1, characterized in that: The noisy speech is a frequency domain signal; The step of reconstructing a clean speech corresponding to the noisy speech according to a filter matching the baseband signal and the noisy speech includes: Multiplying the baseband signal with the filter matched to the noisy speech to obtain a product result; The product result is subjected to inverse Fourier transform to obtain the reconstructed clean speech.

6. A clean speech reconstruction device, characterized in that: include: Feature extraction module, fundamental frequency signal determination module, filter construction module and speech reconstruction module; The feature extraction module is used to extract deep feature information from the noisy speech to be enhanced, wherein the deep feature information represents the speech characteristics of the speaker when uttering the clean speech corresponding to the noisy speech, and the clean speech refers to the speech information in a noise-free environment composed of the speech content of the noisy speech; The fundamental frequency signal determination module is used to determine the fundamental frequency signal of the speech content in the noisy speech according to the deep feature information, wherein the fundamental frequency signal is used to characterize the sound wave signal generated by the vibration of the vocal cords; The filter construction module is used to construct a filter matching the noisy speech based on the deep feature information, wherein the filter matching the noisy speech is used to characterize the vocal channel that modulates the sound wave signal; The speech reconstruction module is used to reconstruct the clean speech corresponding to the noisy speech according to the baseband signal and the filter matched with the noisy speech.

7. The clean speech reconstruction device according to claim 6, characterized in that: The fundamental frequency signal determination module is specifically used to input the deep feature information into an ordinary differential equation filter solving network to obtain a group of ordinary differential equations corresponding to the speech content as the fundamental frequency signal, wherein the ordinary differential equation filter solving network is trained based on the deep feature information extracted from the training speech.

8. The clean speech reconstruction device according to claim 6, characterized in that: The filter construction module comprises: an order coefficient prediction submodule and a filter construction submodule; The order coefficient prediction submodule is used to determine the optimal order and optimal coefficient matching the noisy speech according to the deep feature information; The filter construction submodule is used to construct a filter matching the noisy speech through the optimal order and optimal coefficient.

9. A clean speech reconstruction device, characterized in that: including memory and processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the clean speech reconstruction method as described in any one of claims 1 to 5.

10. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the clean speech reconstruction method as claimed in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Speech quality improvement under heavy noise conditions in hands-free communication

    CN105938714A

  • Voice enhancement method and device and electronic equipment

    CN109427340A