Speech enhancement method and system based on generative adversarial network
By adding multi-head attention layer and multi-generator multi-stage enhancement to the speech enhancement method of generating adversarial networks, the problem of failure to fully utilize speech timing characteristics in the prior art is solved, and a higher quality and intelligible speech enhancement effect is achieved.
Patent Information
- Application Number
- CN202210301250.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-03-25
AI Technical Summary
When dealing with noisy speech, existing speech enhancement methods fail to fully consider the timing characteristics of speech, resulting in poor noise removal effect and insufficient speech quality and intelligibility.
The speech enhancement method based on the generative adversarial network is adopted. By adding a multi-head attention layer to the generator, utilizing the timing characteristics of the speech, and combining the multi-generator multi-stage enhancement and attention mechanism, the generator's ability to approach a clean speech signal is improved.
The quality and intelligibility of the voice signal are significantly improved, and the enhanced voice has higher speech quality and short-term intelligibility.
Smart Images

Figure CN114664318B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech signal processing, and in particular to a speech enhancement method and system based on a generative adversarial network. Background Art
[0002] The statements in this section merely mention background art related to the present invention and do not necessarily constitute prior art.
[0003] Voice is the most direct way to transmit information, but there will be a lot of noise interference in various scenarios of our lives, which will affect the quality of voice. Noise will interfere with human-to-human communication and human-computer interaction. The quality of noisy voice will greatly affect the operating efficiency of the voice system. In the voice signal, there are various interfering noises. The purpose of speech enhancement is to remove the unnecessary noise contained in the signal as much as possible, improve the quality of noisy speech, and increase the intelligibility of speech.
[0004] The speech enhancement methods based on digital signal processing mainly include spectral subtraction, Wiener filtering, and subspace-based algorithms. However, these algorithms have certain limitations and introduce some idealized assumptions, such as stable and additive noise. Only when the noise is stable can better results be achieved.
[0005] At present, in speech enhancement methods based on generative adversarial networks, the generator design is mostly a single generator, and the generator and discriminator are mostly fully convolutional neural networks. The fully convolutional neural networks of the generator and discriminator do not take into account the temporal characteristics of speech well. Summary of the invention
[0006] In order to solve the shortcomings of the prior art, the present invention provides a speech enhancement method and system based on a generative adversarial network; the Speech Enhancement by Generative Adversarial Network (SEGAN) network is improved to remove noise from noisy speech as much as possible, and improve the intelligibility and speech quality of noisy speech. The improvement of adding a multi-head attention layer can better utilize the temporal characteristics of speech.
[0007] In a first aspect, the present invention provides a speech enhancement method based on a generative adversarial network;
[0008] Speech enhancement methods based on generative adversarial networks include:
[0009] Obtain a noisy speech signal; input the noisy speech signal into the trained generative adversarial network, and output an enhanced speech signal;
[0010] Wherein, the generative adversarial network includes two generators and two discriminators;
[0011] The generative adversarial network improves the ability of the generator to approach the target signal through the mutual game between two generators and two discriminators during the training process.
[0012] In a second aspect, the present invention provides a speech enhancement system based on a generative adversarial network;
[0013] The speech enhancement system based on generative adversarial network includes:
[0014] An acquisition module is configured to: acquire a noisy speech signal;
[0015] A speech enhancement module is configured to: input a noisy speech signal into a trained generative adversarial network and output an enhanced speech signal;
[0016] Wherein, the generative adversarial network includes two generators and two discriminators;
[0017] The generative adversarial network improves the ability of the generator to approximate the clean speech target signal through the mutual game between two generators and two discriminators during the training process.
[0018] In a third aspect, the present invention further provides an electronic device, comprising:
[0019] a memory for non-transitory storage of computer-readable instructions; and
[0020] a processor for executing the computer readable instructions,
[0021] When the computer-readable instructions are executed by the processor, the method described in the first aspect is executed.
[0022] In a fourth aspect, the present invention further provides a storage medium that non-temporarily stores computer-readable instructions, wherein when the non-temporary computer-readable instructions are executed by a computer, the instructions of the method described in the first aspect are executed.
[0023] In a fifth aspect, the present invention further provides a computer program product, comprising a computer program, wherein the computer program is used to implement the method described in the first aspect when running on one or more processors.
[0024] Compared with the prior art, the present invention has the following beneficial effects:
[0025] The present invention mainly utilizes a generative adversarial network speech enhancement (Speech EnhancementGenerativeAdversarial Network) network for improvement, and the generated enhanced speech has the purpose of higher speech quality and short-term intelligibility.
[0026] The present invention fully considers the temporal relationship of speech signals, improves the previous full convolution design of the generator and discriminator, adds a multi-head attention mechanism to the generator, and combines multi-generator multi-stage enhancement with the attention mechanism, making full use of the multi-head attention mechanism and the idea of generative adversarial network game. This method can make the enhanced speech have higher quality and intelligibility. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0028] Figure 1 This is a flowchart of a GAN-based speech enhancement method according to Embodiment 1 of the present application;
[0029] Figure 2 This is a structural diagram of a generator in a GAN-based speech enhancement method according to Embodiment 1 of the present application;
[0030] Figure 3 This is a structural diagram of the discriminator in the GAN-based speech enhancement method of Example 1 of the present application. DETAILED DESCRIPTION
[0031] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0032] It should be noted that the terms used herein are only for describing specific embodiments, and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "include" and "have" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0033] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.
[0034] In this embodiment, all data is obtained in compliance with laws and regulations and based on the user's consent, and is used legally.
[0035] With the development of deep learning, many speech enhancement algorithms based on neural networks have emerged, overcoming many existing assumptions and inaccurate noise estimation problems. Compared with other speech enhancement algorithms based on neural networks, speech enhancement algorithms based on generative adversarial networks have the advantages of good generalization performance under different noise types.
[0036] Embodiment 1
[0037] This embodiment provides a speech enhancement method based on a generative adversarial network;
[0038] Speech enhancement methods based on generative adversarial networks include:
[0039] S101: Acquire a noisy speech signal;
[0040] S102: Inputting the noisy speech signal into the trained generative adversarial network, and outputting an enhanced speech signal;
[0041] Wherein, the generative adversarial network includes two generators and two discriminators;
[0042] The generative adversarial network improves the ability of the generator to approach the target signal through the mutual game between two generators and two discriminators during the training process.
[0043] The training process of the two generators is to minimize the following loss function:
[0044]
[0045] The training process of the two discriminators is to minimize the following loss function:
[0046]
[0047] During the training process, the training input of the generator is a noisy speech signal. Z is the hidden layer random noise, n is 2, G 1 represents the first generator; G 2 represents the second generator; D 1 represents the first discriminator; D 2 represents the second discriminator; λ is the hyperparameter of L1 loss, which is set to 100.
[0048] Furthermore, if Figure 1 As shown, the generative adversarial network includes: a first generator, a second generator, a first discriminator and a second discriminator;
[0049] The input end of the first generator is used to input a noisy speech signal;
[0050] The output terminal of the first generator outputs a first enhanced speech signal;
[0051] The input end of the second generator is used to input the first enhanced speech signal;
[0052] The output terminal of the second generator is used to output a second enhanced speech signal;
[0053] The input end of the first discriminator is used to input the second enhanced speech signal and the noise-free speech signal; the first discriminator outputs the recognition result of the noise-free speech signal or the noisy speech signal;
[0054] The input end of the second discriminator is used to input the second enhanced speech signal and the noise-free speech signal; the second discriminator outputs the recognition result of the noise-free speech signal or the noisy speech signal.
[0055] Furthermore, the internal structures of the first generator and the second generator are consistent.
[0056] like Figure 2 As shown, the first generator includes an encoder and a decoder connected to each other;
[0057] The encoder comprises: a convolutional layer c1, a convolutional layer c2, a convolutional layer c3, a convolutional layer c4, a convolutional layer c5, a convolutional layer c6, a multi-head attention mechanism layer, a convolutional layer c7, a convolutional layer c8, a convolutional layer c9, a convolutional layer c10 and a convolutional layer c11 connected in sequence;
[0058] The decoder comprises: a deconvolution layer d11, a deconvolution layer d10, a deconvolution layer d9, a deconvolution layer d8, a deconvolution layer d7, a multi-head attention mechanism layer, a deconvolution layer d6, a deconvolution layer d5, a deconvolution layer d4, a deconvolution layer d3, a deconvolution layer d2 and a deconvolution layer d1 connected in sequence;
[0059] Among them, the convolution layer of the encoder and the deconvolution layer of the decoder are added with residual connections.
[0060] Furthermore, the working principle of the encoder is: analyzing the input speech signal sequence; using a multi-head attention mechanism layer to learn speech features from different aspects, especially for noise processing, to improve the quality of generated speech.
[0061] Furthermore, the working principle of the decoder is: generating an output speech signal sequence, wherein a multi-head attention mechanism layer is used to learn speech features from different aspects, especially for noise processing, to improve the quality of generated speech.
[0062] Furthermore, the convolutional layer of the encoder is residually connected to the convolutional layer of the decoder; specifically comprising:
[0063] The convolution layer c1 of the encoder is connected to the deconvolution layer d1 of the decoder;
[0064] The convolutional layer c2 of the encoder is connected to the deconvolutional layer d2 of the decoder;
[0065] The convolutional layer c3 of the encoder is connected to the deconvolutional layer d3 of the decoder;
[0066] And so on; the convolutional layer c11 of the encoder is connected to the deconvolutional layer d11 of the decoder.
[0067] Furthermore, the internal structures of the first discriminator and the second discriminator are the same.
[0068] like Figure 3 As shown, the first discriminator includes: a convolutional layer e1, a convolutional layer e2, a convolutional layer e3, a convolutional layer e4, a convolutional layer e5, a convolutional layer e6, a convolutional layer e7, a convolutional layer e8, a GRU layer, a multi-head attention mechanism layer and a softmax activation function layer connected in sequence.
[0069] Furthermore, the working principle of the first discriminator is: judging the signal generated by the second generator as false, and judging the real speech signal as true; using the GRU layer network with small parameters to reduce the risk of overfitting; using multi-head attention to learn speech features from different aspects to judge whether it is real speech or generated speech.
[0070] Furthermore, the working principle of the second discriminator is the same as that of the first discriminator.
[0071] Furthermore, the trained generative adversarial network; the training process includes:
[0072] Construct a training set; the data set is a database provided by Voice Bank of the University of Edinburgh, UK. The clean speech and noisy speech of the database are composed of 28 speakers, each with about 400 speech;
[0073] Inputting the noisy speech into the first generator, the first generator generates a first enhanced speech signal;
[0074] The first enhanced speech signal is input into the second generator, and the second generator generates a second enhanced speech signal;
[0075] Inputting the second enhanced speech signal and the noise-free signal into the first discriminator for discrimination, and outputting a first discrimination result;
[0076] Inputting the second enhanced speech signal and the noise-free signal into a second discriminator for discrimination, and outputting a second discrimination result;
[0077] When the accuracy of the first and second discrimination results reaches 50%, the training is stopped to obtain the trained generative adversarial network.
[0078] For the selection of network structure, two generators and two discriminators are used in GAN. A convolutional network is selected and a multi-head attention mechanism layer is added to it. In particular, a GRU network design is used in the two discriminators. Two generators are used to enhance the speech signal in two independent stages until the two discriminators cannot distinguish. Among them, the enhanced speech generated by the second generator is the final enhanced speech. The noisy speech is input into the trained first generator, and the speech signal is generated through the second generator. Gaussian noise is used as random noise input, and a clean speech signal is used as the target signal. A convolutional neural network is used and an attention layer is added as the network structure of the generator and discriminator.
[0079] Furthermore, the trained generative adversarial network, prior to the training process, also includes: an initialization phase;
[0080] Furthermore, the initialization stage includes: a step of processing a data set, a step of initializing the first generator and the second generator, a step of initializing the first discriminator and the second discriminator, and a stage of optimizing weights.
[0081] Furthermore, the steps of processing the data set include:
[0082] (1.1) The data in the dataset are integrated into tfrecords files, clean speech data (noise-free speech signals) are classified into the wav class, and random noise is classified into the noisy class.
[0083] Embodiment: In this step, the data type in the tfrecords file is int type, the data size range is -32767 to 32767, and the input data set sampling rate is 16KHZ, so each data size is set to 16384, but each data size is not limited to this and can be adjusted according to the data sampling rate.
[0084] (1.2) Determine the optimizer of the entire GAN and read out the random noise and clean speech of the tfrecords file.
[0085] Example: Determine the optimizer as RMSProp.
[0086] (1.3) Change the size of random noise and clean speech, and apply pre-emphasis in the range of 0.9 to 1.
[0087] Example: Change the range of random noise and clean speech to -1 to 1 to prevent problems such as gradient explosion, and implement 0.95 pre-emphasis to make its high-frequency characteristics have better performance
[0088] (1.4) Put the random noise and clean speech into the queue, and take out the required enhanced speech and clean speech batches each time.
[0089] Example: Batch size is 50, 16384 frames long;
[0090] Considering that multiple generators collaborate in multiple stages to generate speech, a two-generator training method is adopted to reconstruct and generate clean speech.
[0091] Furthermore, the first generator and the second generator initialization step specifically include:
[0092] (2.1) Take out the random noise separately and adjust the dimension.
[0093] Example: Adjust the random noise dimension to 4 dimensions, and the dimension size is [150, 16384, 1, 1].
[0094] (2.2) Determine the convolution kernel size of the two-dimensional convolution to be 32 and the stride to be 2. Adjust the dimension after the two-dimensional convolution and use the activation function. Concatenate the two-dimensional convolution result with Gaussian noise of the same size. Perform two-dimensional deconvolution and perform skip residual connection with the vector of the same size in the two-dimensional convolution process. Use the activation function PReLU in each deconvolution layer. In this example, only the batch size is 50.
[0095] (2.3) Add a multi-head attention layer at the end to get the output of the last layer, apply the activation function to it, and get the generated enhanced speech.
[0096] Example: Use the PReLU activation function, whose formula is
[0097] Furthermore, the first discriminator and the second discriminator initialization step specifically includes:
[0098] (3.1) The clean speech obtained in the data processing stage is set as the w sequence.
[0099] (3.2) Create a Gaussian noise sequence with the same dimension and size as the w sequence, and add it to w to obtain a new w.
[0100] Embodiment: The mean value of Gaussian noise is set to 0 and the variance is set to 0.5.
[0101] (3.3) Adjust the dimension of the w sequence. Determine the size, step size, padding method, etc. of the two-dimensional convolution filter. Perform virtual batch normalization on w after the two-dimensional convolution and use the activation function to obtain a new w.
[0102] Embodiment: The parameter selection is the same as the configuration of the initialization stage of the first generator and the second generator, wherein the purpose of virtual batch normalization is to accelerate the convergence speed of the model.
[0103] (3.4) The two-dimensional convolution result is subjected to one-dimensional convolution and then sent to the GRU layer. The output of the GRU layer is sent to the multi-head attention layer, and finally the probability of the true data with an output probability value close to 1 is obtained.
[0104] Furthermore, the weight optimization stage specifically includes:
[0105] (4.1) The first discriminator and the second discriminator use clean speech as real data, and the probability of outputting close to 1 during the initialization phase of the first discriminator and the second discriminator is represented as true data. The first discriminator and the second discriminator input the enhanced speech generated by the generator as false data, and the first discriminator and the second discriminator will perform the operation of the initialization phase and the probability of outputting close to 0 is represented as false data. Calculate the loss value of the first discriminator and the second discriminator.
[0106] (4.2) Update the filter values of the convolution and deconvolution in the initialization of the first generator, the second generator, the first discriminator and the second discriminator, and the gama and beta values in the virtual batch normalization according to the loss values of the first generator, the second generator, the first discriminator and the second discriminator.
[0107] Furthermore, after training, the generative adversarial network, training phase:
[0108] (5.1) Repeat the three steps of initializing the first generator and the second generator, initializing the first discriminator and the second discriminator, and optimizing the weights;
[0109] (5.2) Determine whether the current number of training data is greater than the number of data in the tfrecords file, and repeat the training until the number of training data is reached.
[0110] The random noise z is input into the trained second generator, and the enhanced speech signal is generated by the generator. The process is as follows:
[0111] (6.1) Read the random noise file and determine whether the sampling rate is 16KHz.
[0112] (6.2) Configure the weights of the trained model.
[0113] (6.3) Convert the read data size to -1~1.
[0114] (6.4) Determine the data size.
[0115] (6.5) Send data to the generator at intervals of 16384 and save the generated results.
[0116] (6.6) Write the saved data to a wav file.
[0117] The speech enhancement method based on generative adversarial network has the following innovations: through generative adversarial network technology, the input noisy speech is enhanced in multiple stages through multiple generators, and a multi-head attention layer is added to the multi-layer convolutional neural network of the generator. The noisy speech is output after passing through multiple generators, and the input of the discriminator is the enhanced speech and the real clean speech generated by multiple generators. The discriminator judges the probability that the enhanced speech is a real clean speech through a multi-layer convolutional neural network. The ability of the generator to approximate the clean speech signal can be improved through the mutual game between the generator and the discriminator. It should be noted that the generator design in the generative adversarial network described in the present invention not only includes the two generators in the example, but also includes multiple generators; and the multi-head attention-based combination model in the generator and the discriminator, and the combination of the two in the present invention.
[0118] The present invention provides a speech enhancement method based on a generative adversarial network. Based on the generative adversarial network, an input signal of noisy speech passes through multiple generator multi-layer convolutional neural networks, an attention layer is added to the multi-layer convolutional neural network, and is converted into an enhanced speech output. The input of a discriminator is the enhanced speech and a clean signal generated by the generator. The discriminator determines the probability that the input is a target signal through the multi-layer convolutional neural network. The mutual game between the generator and the discriminator can improve the ability of the enhanced speech generated by the generator to approach the clean signal. The enhanced speech obtained by this method has higher speech quality and intelligibility.
[0119] Embodiment 2
[0120] This embodiment provides a speech enhancement system based on a generative adversarial network;
[0121] The speech enhancement system based on generative adversarial network includes:
[0122] An acquisition module is configured to: acquire a noisy speech signal;
[0123] A speech enhancement module is configured to: input a noisy speech signal into a trained generative adversarial network and output an enhanced speech signal;
[0124] Wherein, the generative adversarial network includes two generators and two discriminators;
[0125] The generative adversarial network improves the ability of the generator to approach the target signal through the mutual game between two generators and two discriminators during the training process.
[0126] It should be noted that the acquisition module and the speech enhancement module correspond to steps S101 to S102 in Embodiment 1, and the examples and application scenarios implemented by the modules and the corresponding steps are the same, but are not limited to the contents disclosed in Embodiment 1. It should be noted that the modules as part of the system can be executed in a computer system such as a set of computer executable instructions.
[0127] The description of each embodiment in the above embodiments has different emphases. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0128] The proposed system can be implemented in other ways. For example, the system embodiment described above is only illustrative, and the division of the modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0129] Embodiment 3
[0130] This embodiment also provides an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device is running, the processor executes the one or more computer programs stored in the memory so that the electronic device executes the method described in the above embodiment one.
[0131] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0132] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0133] In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in a processor or an instruction in the form of software.
[0134] The method in the first embodiment can be directly embodied as a hardware processor, or a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here.
[0135] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0136] Embodiment 4
[0137] This embodiment further provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first embodiment is completed.
[0138] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A speech enhancement method based on a generative adversarial network, characterized in that: include: Acquire a noisy speech signal; Input the noisy speech signal into the trained generative adversarial network and output the enhanced speech signal; Wherein, the generative adversarial network includes two generators and two discriminators; The generative adversarial network improves the ability of the generator to approach the target signal by playing a game between two generators and two discriminators during training. The generative adversarial network comprises: a first generator, a second generator, a first discriminator and a second discriminator; The input end of the first generator is used to input a noisy speech signal; The output terminal of the first generator outputs a first enhanced speech signal; The input end of the second generator is used to input the first enhanced speech signal; The output terminal of the second generator is used to output a second enhanced speech signal; The input end of the first discriminator is used to input the second enhanced speech signal and the noise-free speech signal; the first discriminator outputs the recognition result of the noise-free speech signal or the noisy speech signal; The input end of the second discriminator is used to input the second enhanced speech signal and the noise-free speech signal; the second discriminator outputs the recognition result of the noise-free speech signal or the noisy speech signal; The first generator includes an encoder and a decoder connected to each other; The encoder comprises: a plurality of convolutional layers and an attention mechanism layer; The decoder comprises: a plurality of deconvolution layers and an attention mechanism layer; Among them, the convolution layer of the encoder and the deconvolution layer of the decoder add residual connections; The working principle of the encoder is: analyzing the input speech signal sequence; using a multi-head attention mechanism layer to learn speech features from different aspects.
2. The speech enhancement method based on generative adversarial network as claimed in claim 1, characterized in that: The first discriminator includes: multiple convolutional layers, GRU layers, multi-head attention mechanism layers and softmax activation function layers.
3. The speech enhancement method based on generative adversarial network as claimed in claim 1, characterized in that: The trained GAN. The training process includes: Construct a training set; Inputting the noisy speech into the first generator, the first generator generates a first enhanced speech signal; The first enhanced speech signal is input into the second generator, and the second generator generates a second enhanced speech signal; Inputting the second enhanced speech signal and the noise-free signal into a first discriminator for discrimination, and outputting a first discrimination result; Inputting the second enhanced speech signal and the noise-free signal into a second discriminator for discrimination, and outputting a second discrimination result; When the accuracy of the first and second discrimination results reaches 50%, the training is stopped to obtain the trained generative adversarial network.
4. The speech enhancement method based on generative adversarial network as claimed in claim 1, characterized in that: The training process of the two generators is to minimize the following loss function: The training process of the two discriminators is to minimize the following loss function: During the training process, the training input of the generator is a noisy speech signal. Z is the hidden layer random noise, n is 2, G1 represents the first generator; G2 represents the second generator; D1 represents the first discriminator; D2 represents the second discriminator; λ is the hyperparameter of L1 loss.
5. A speech enhancement system based on generative adversarial networks, characterized by: include: An acquisition module is configured to: acquire a noisy speech signal; A speech enhancement module is configured to: input a noisy speech signal into a trained generative adversarial network and output an enhanced speech signal; Wherein, the generative adversarial network includes two generators and two discriminators; The generative adversarial network improves the ability of the generator to approximate the clean speech target signal by playing a game between two generators and two discriminators during training. The generative adversarial network comprises: a first generator, a second generator, a first discriminator and a second discriminator; The input end of the first generator is used to input a noisy speech signal; The output terminal of the first generator outputs a first enhanced speech signal; The input end of the second generator is used to input the first enhanced speech signal; The output terminal of the second generator is used to output a second enhanced speech signal; The input end of the first discriminator is used to input the second enhanced speech signal and the noise-free speech signal; the first discriminator outputs the recognition result of the noise-free speech signal or the noisy speech signal; The input end of the second discriminator is used to input the second enhanced speech signal and the noise-free speech signal; the second discriminator outputs the recognition result of the noise-free speech signal or the noisy speech signal; The first generator includes an encoder and a decoder connected to each other; The encoder comprises: a plurality of convolutional layers and an attention mechanism layer; The decoder comprises: a plurality of deconvolution layers and an attention mechanism layer; Among them, the convolution layer of the encoder and the deconvolution layer of the decoder add residual connections; The working principle of the encoder is: analyzing the input speech signal sequence; using a multi-head attention mechanism layer to learn speech features from different aspects.
6. An electronic device, comprising: a memory for non-transitory storage of computer readable instructions; as well as a processor for executing the computer readable instructions, When the computer-readable instructions are executed by the processor, the method described in any one of claims 1 to 4 is executed.
7. A storage medium, characterized in that it non-temporarily stores computer-readable instructions, wherein: When the non-transitory computer-readable instructions are executed by a computer, the instructions of the method according to any one of claims 1 to 4 are executed.
Citation Information
Patent Citations
Voice data amplification method and system
CN108922518A
Method, apparatus and equipment for establishing voice enhancement network and computer storage medium
CN109147810A