A Speech Enhancement Method Based on Deep Neural Network Technology

Through the speech enhancement method based on deep neural networks, the pre-trained neural network model is used to encode, decode and remove voice data redundantly, solving the problem that traditional methods cannot quickly deal with transient noise, and achieving efficient noise suppression and voice signal quality improvement.

CN114360561BActive Publication Date: 2025-06-24GUANGDONG ELECTRIC POWER COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111484420.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-07
Publication Date
2025-06-24
Estimated Expiration
2041-12-07

AI Technical Summary

Technical Problem

Traditional speech signal processing methods cannot quickly process transient noise, resulting in poor results when processing voice signals containing transient noise.

Method used

The speech enhancement method based on deep neural network is adopted to process the speech data to be processed through the pre-trained neural network model, the speech data is encoded and decoded using the encoding structure and the decoding structure, and the redundancy in the decoding output information is removed through the gated structure to achieve rapid noise suppression.

Benefits of technology

This method can quickly and effectively suppress noise data, improve the quality of voice signals, and significantly improve the noise reduction effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114360561B_ABST
    Figure CN114360561B_ABST
Patent Text Reader

Abstract

The present application discloses a speech enhancement method based on deep neural network technology, which includes obtaining speech data to be processed; inputting the speech data to be processed into a pre-trained neural network model to obtain enhanced speech data output by the neural network model; the enhanced speech data is obtained by filtering noise data from the speech data to be processed; wherein, the pre-trained neural network model includes an encoding structure and a decoding structure, and is trained by encoding the training speech data and transmitting it to the decoding structure, and performing redundancy removal and transmission processing on the decoding output information between adjacent decoding layers. Thus, by processing the speech data to be processed through the pre-trained neural network model, noise data can be quickly and effectively suppressed, and the pre-trained neural network model focuses more on effective information by performing redundancy removal processing on the decoding output information, significantly improving the noise reduction effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, and more particularly, to a speech enhancement method based on deep neural network technology. Background Art

[0002] In related technologies, the acquired speech signal often includes various interfering noises, and speech enhancement technology is a technology that suppresses the interfering noises therein and enhances the effective speech.

[0003] In traditional speech signal processing, speech noise reduction algorithms are often used to achieve speech enhancement. In this way, it is necessary to estimate the noise spectrum, and the tracking estimation of the noise spectrum takes a certain amount of time to be accurately completed. Therefore, traditional signal processing methods cannot quickly process transient noises. Summary of the Invention

[0004] In view of the above problems, the present invention proposes a speech enhancement method based on deep neural network technology to improve the above problems.

[0005] An embodiment of this application provides a speech enhancement method based on deep neural network technology. The method includes: acquiring speech data to be processed; inputting the speech data to be processed into a pre-trained neural network model to obtain enhanced speech data output by the neural network model; the enhanced speech data is obtained by filtering noise data from the speech data to be processed; wherein, the pre-trained neural network model includes an encoding structure and a decoding structure, and is trained by encoding the training speech data and passing it to the decoding structure, and performing redundancy removal and transfer processing on the decoding output information between adjacent decoding layers.

[0006] In some embodiments of this application, based on the foregoing solution, the pre-trained neural network model is also trained by performing speech recognition processing on the training speech data.

[0007] In some embodiments of this application, based on the foregoing solution, the pre-trained neural network model is trained on a gated connection network based on the training speech data, and the gated connection network includes an encoding structure, a gating structure, and a decoding structure.

[0008] In some embodiments of this application, based on the foregoing solution, the gated connection network further includes a timing structure, and the timing structure is arranged between the encoding structure and the decoding structure, and the timing structure is used to obtain the timing information in the encoding output information.

[0009] In some embodiments of this application, based on the foregoing solution, the gated connection network further includes a speech recognition structure, and the speech recognition structure is connected to the encoding structure, and the speech recognition structure is used to perform speech recognition processing on the training speech data.

[0010] In some embodiments of the present application, based on the foregoing solution, the encoding structure includes multiple encoding layers, and the decoding structure includes multiple decoding layers; the encoded output information of each encoding layer is input to the corresponding decoding layer; a gating structure is provided between adjacent decoding layers.

[0011] In some embodiments of the present application, based on the foregoing solution, the pre-trained neural network model is obtained through the following steps: obtaining training speech data, where the training speech data includes mixed speech data; obtaining a gating connection network, which includes an encoding structure, a gating structure, and a decoding structure; training the gating connection network with the training speech data until the gating connection network meets a preset condition, and obtaining the trained gating connection network as the pre-trained neural network model; wherein, the encoding structure is used to encode the training speech data to obtain encoded output information, the gating structure is used to remove redundancy and perform transfer processing on the decoded output information to obtain key decoded output information, the decoding structure is used to obtain enhanced speech data corresponding to the training speech data according to the encoded output information and the key decoded output information; the timing structure is used to obtain timing information according to the encoded output information and transfer it to the decoding structure; the speech recognition structure is used to determine whether the training speech data includes target speech data according to the encoded output information.

[0012] In some embodiments of the present application, based on the foregoing solution, training the gating connection network with the training speech data until the gating connection network meets a preset condition includes: inputting the training speech data into the encoding structure, and generating encoded output information corresponding to the training speech data through the encoding structure; inputting the encoded output information into the decoding structure, and outputting enhanced speech data corresponding to the training speech data through the decoding structure and the gating structure; inputting the encoded output information into the speech recognition structure, and generating a speech recognition result corresponding to the speech data through the speech recognition structure; obtaining the total loss value of the gating connection network according to the enhanced speech data and the speech recognition result; performing iterative training on the gating connection network according to the total loss value until the gating connection network meets a preset condition.

[0013] In some embodiments of the present application, based on the foregoing solution, obtaining the total loss value of the gating connection network according to the enhanced speech data and the speech recognition result includes: obtaining the first loss value of the gating connection network according to the enhanced speech data; obtaining the second loss value of the gating connection network according to the speech recognition result; determining the total loss value of the gating connection network based on the first loss value and the second loss value.

[0014] In some embodiments of the present application, based on the foregoing solution, the mixed speech data is obtained by mixing noise data into clean speech data.

[0015] The technical solution provided by the present invention includes obtaining voice data to be processed; inputting the voice data to be processed into a pre-trained neural network model to obtain enhanced voice data output by the neural network model; the enhanced voice data is obtained by filtering noise data from the voice data to be processed; wherein, the pre-trained neural network model includes an encoding structure and a decoding structure, and is trained by encoding training voice data and passing it to the decoding structure, and performing redundancy removal and passing processing on the decoding output information between adjacent decoding layers. In this way, by processing the voice data to be processed with the pre-trained neural network model, noise data can be quickly and effectively suppressed, and the pre-trained neural network model focuses more on effective information by performing redundancy removal processing on the decoding output information, significantly improving the noise reduction effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments and drawings obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0017] Figure 1 FIG. shows a flowchart of a voice enhancement method based on deep neural network technology proposed in an embodiment of the present application.

[0018] Figure 2 FIG. shows a schematic structural diagram of a gated connection network proposed in an embodiment of the present application.

[0019] Figure 3 shows Figure 2 the connection schematic diagram of the gated structure in

[0020] Figure 4 FIG. shows a schematic structural diagram of a gated connection network proposed in another embodiment of the present application.

[0021] Figure 5 FIG. shows a flowchart of the training process of the gated connection network in an embodiment of the present application.

[0022] Figure 6 FIG. shows a block diagram of the structure of a voice enhancement device based on deep neural network technology proposed in an embodiment of the present application.

[0023] Figure 7 FIG. shows a block diagram of the structure of an electronic device proposed in an embodiment of the present application.

[0024] Figure 8 FIG. shows a block diagram of the structure of a computer-readable storage medium proposed in an embodiment of the present application. Detailed implementation manners

[0025] In order to enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application.

[0026] In related technologies, the acquired voice signals often include various interference noises, and voice enhancement technology is a technology for suppressing the interference noises therein and enhancing the effective voice.

[0027] In traditional voice signal processing, voice noise reduction algorithms are often used to achieve voice enhancement. For example, methods such as Wiener filter, adaptive filter (LMS, Least Mean Square), and minimum mean square error (MMSE-LSA, Minimum Mean-Square Error Log-Spectral Amplitude) are used to achieve voice noise reduction.

[0028] However, when using a voice noise reduction algorithm to achieve voice enhancement, it is necessary to estimate the noise spectrum, and the tracking and estimation of the noise spectrum takes a certain amount of time to be accurately completed. Therefore, traditional signal processing methods cannot quickly process transient noises.

[0029] In order to improve the above technical problems, the inventor proposes a voice enhancement method based on deep neural network technology provided by this application. By acquiring the voice data to be processed; inputting the voice data to be processed into a pre-trained neural network model to obtain the enhanced voice data output by the neural network model; the enhanced voice data is obtained after filtering out the noise data from the voice data to be processed; wherein, the pre-trained neural network model includes an encoding structure and a decoding structure, and is trained by encoding the training voice data and passing it to the decoding structure, and performing redundancy removal and passing processing on the decoding output information between adjacent decoding layers. In this way, by using the pre-trained neural network model to process the voice data to be processed, the noise data can be quickly and effectively suppressed, and the pre-trained neural network model focuses more on the effective information by performing redundancy removal processing on the decoding output information, significantly improving the noise reduction effect.

[0030] The embodiments of this application will be specifically described below in conjunction with the accompanying drawings.

[0031] Please refer to Figure 1 , Figure 1 which shows a schematic flowchart of a voice enhancement method based on deep neural network technology provided by an embodiment of this application. In a specific embodiment, the voice enhancement method is applied to a voice enhancement device 400 as shown in Figure 6 and as shown in Figure 7An electronic device 600 configured with the voice enhancement device 400 as shown. Taking the electronic device as an example, the specific process of this embodiment will be described below. Next, the following will be directed at Figure 1 The process shown will be elaborated in detail. The voice enhancement method may specifically include: Step 110 to Step 130.

[0032] Step 110: Obtain the voice data to be processed.

[0033] In the embodiments of the present application, the electronic device may obtain the voice data to be processed that needs to be enhanced.

[0034] Among them, the voice data to be processed may include noise data, that is, the data that needs to be filtered out. It can be understood that during the voice recording process, due to the influence of background noise in the recording environment, the recorded voice data often contains noise. The noise is, for example, environmental sounds or some sounds that we do not need. The noise is, for example, wind sounds, rain sounds, machine operation sounds, etc. During a voice call, usually the caller is close to the microphone. At this time, the noise can also be the voices of people farther away from the microphone. In the voice data containing noise, the effective voice data may be masked by the noise data. For example, when the noise volume is greater than the volume of the effective audio and the noise frequency is greater than the frequency of the effective audio, the effective audio is easily masked by the noise data, resulting in the inability to obtain effective information in application scenarios such as voice playback, voice recognition, and voice calls.

[0035] In some embodiments, the electronic device may be a smart phone, a tablet computer, a smart watch, smart glasses, a notebook computer, etc., which is not limited herein. The electronic device may be provided with an audio acquisition device, and the audio acquisition device is, for example, a microphone, a microphone array, etc. Thus, the voice data can be collected through the audio acquisition device as the voice data to be processed.

[0036] In other embodiments, the electronic device may obtain the voice data to be processed from a locally stored file. For example, when the electronic device is a mobile terminal, it may obtain the voice data to be processed from a local folder, that is, the electronic device pre-collects the voice data through the audio acquisition device and stores it in the local folder. Another example is that the electronic device may also pre-download the voice data from the network and store it in the local folder, and then when it is necessary to perform voice enhancement processing on the voice data, read the voice data to be processed from the local folder.

[0037] In still other embodiments, the electronic device may also download the voice data to be processed from the network. For example, the electronic device may download the required voice data to be processed from the corresponding server through a wireless network, etc.

[0038] In some other embodiments, the electronic device can also obtain the speech data to be processed based on the voice data received by the running application program. For example, the voice chat data received in an instant messaging software, etc.

[0039] Of course, the specific manner in which the electronic device obtains the speech data to be processed may not be limited.

[0040] Step 120: Input the speech data to be processed into a pre-trained neural network model to obtain the enhanced speech data output by the neural network model. The enhanced speech data is obtained by filtering out the noise data from the speech data to be processed. Among them, the pre-trained neural network model includes an encoding structure and a decoding structure, and is trained by encoding the training speech data and passing it to the decoding structure, and performing redundancy removal and passing processing on the decoding output information between adjacent decoding layers.

[0041] After obtaining the speech data to be processed, the electronic device can input the speech data to be processed into a pre-trained neural network model to obtain the enhanced speech data corresponding to the speech data to be processed. Among them, the enhanced speech data is obtained by filtering out the noise data from the speech data to be processed.

[0042] In the embodiments of the present application, the above-mentioned pre-trained neural network model can adopt the structure of an autoencoder. An autoencoder (AE) is a type of artificial neural network (ANNs) used in semi-supervised learning and unsupervised learning, and its function is to perform representation learning on the input information by taking the input information as the learning target.

[0043] The pre-trained neural network model can include an encoding structure and a decoding structure. The encoding structure can encode the input training speech data, and the encoding structure passes the encoding result to the decoding structure, and the decoding structure outputs the enhanced speech data with noise filtered out corresponding to the training speech data. The neural network is trained through the training gap between the enhanced speech data and the clean speech data, as well as the difference between it and the preset gap, to obtain a neural network model with a noise filtering degree meeting the preset requirements.

[0044] Using the pre-trained neural network model to perform speech enhancement processing on the speech data to be processed can quickly and effectively suppress the noise data in the speech data to be processed and obtain the enhanced speech data corresponding to the speech data to be processed.

[0045] Generally, the encoding structure includes multiple encoding layers, and the decoding structure includes multiple decoding layers. The number of decoding layers in the decoding structure corresponds to the number of encoding layers in the encoding structure to obtain enhanced speech data with the same dimension as the data to be processed. Each encoding layer in the encoding structure corresponds to one of the decoding layers in the decoding structure and has the same dimension. For example, the first encoding layer corresponds to the last decoding layer and has the same dimension; the last encoding layer corresponds to the first decoding layer and has the same dimension.

[0046] In the encoding structure, the first encoding layer receives the speech data to be processed, performs downsampling processing, and then passes the encoded output information obtained after the downsampling processing to the next encoding layer. The encoding result of the encoding structure is passed to the decoding structure. The first decoding layer receives the encoding result, performs upsampling processing, and then passes the decoded output information obtained after the upsampling processing to the next decoding layer.

[0047] To prevent the network from experiencing gradient disappearance during training and to reduce the information loss caused by the downsampling processing in the encoding structure, the decoded input information of the decoding layer includes not only the decoded output information of the previous decoding layer but also the encoded output information of its corresponding encoding layer. However, the method of directly concatenating the encoded output information and the decoded output information as the decoded input information can reduce information loss but also cause redundancy in the decoded input information of the decoding layer, that is, adding unnecessary information back into the network's results. Unnecessary information can be, for example, duplicate features. For example, for the same feature A, if the decoded output information contains feature A and the encoded output information also includes feature A, then feature A is a duplicate feature. Duplicate features can lead to insufficient fitting ability of the network.

[0048] Furthermore, in the embodiments of the present application, the pre-trained neural network model is trained by encoding the training speech data and passing it to the decoding structure, and performing redundancy removal and transfer processing on the decoded output information between adjacent decoding layers. Thus, the redundant information in the decoded output information is removed before being passed to the next decoding layer. The network trained in this way can focus on the effective information, avoid the loss of effective information, and has stronger fitting ability.

[0049] In some embodiments, the time-frequency masking (the part to be masked in the speech data to be processed) of the speech data to be processed is estimated by a pre-trained neural network, and then the time-frequency masking is multiplied item by item with the speech data to be processed (the speech data with noise) to obtain the enhanced speech data (the speech data obtained by removing the part to be masked from the speech data to be processed).

[0050] In some embodiments, according to the different degrees of noise filtering of the pre-trained neural network model, the noise filtering effect of the enhanced speech data obtained is different. It can be understood that the higher the degree of noise filtering, the better the noise filtering effect of the obtained enhanced speech data.

[0051] In some embodiments, the degree of noise filtering of the pre-trained neural network model for noise is a fixed value, that is, for different speech data to be processed, the degree of noise filtering is the same.

[0052] In other embodiments, the degree of noise filtering of the pre-trained neural network model for noise is a floating value, that is, in different situations, the degree of noise filtering is different. For example, the speech data to be processed may include multiple frames of audio. When a target data is included in one of the frames of audio, the degree of noise filtering can be reduced to reduce the loss of effective information. When a target data is not included in one of the frames of audio, the degree of noise filtering can be enhanced. Specifically, the magnitude of the degree of noise filtering for each frame of audio can be set according to actual usage needs, and the present application does not limit this.

[0053] In some embodiments, the pre-trained neural network model is used to perform noise filtering on the speech data to be processed to obtain enhanced speech data. Since the time-frequency masking is directly estimated by the pre-trained neural network model, in actual use, if the masking intensity is too strong (the degree of noise filtering is too strong), some effective information in the speech data to be processed will also be masked, and if the masking intensity is insufficient, there will be more noise data in the obtained enhanced speech data, which cannot meet the usage needs. To solve this technical problem, in the speech enhancement method of the embodiments of the present application, the pre-trained neural network model is also obtained by performing speech recognition processing on the training speech data. The speech recognition processing can identify whether the training speech data includes target speech data. The target speech data can be, for example, human voices, musical instrument sounds, etc.

[0054] Furthermore, by performing speech recognition processing on the training speech data, the degree of noise filtering for the training speech data is adjusted according to the speech recognition result. For example, when the training speech data includes target speech data, the degree of noise filtering can be reduced to reduce the loss of effective information. When the training speech data does not include target speech data, the degree of noise filtering can be unchanged or enhanced to ensure the effective filtering of noise.

[0055] In some embodiments, the pre-trained neural network model of the embodiments of the present application is obtained by training a gated connection network based on the training speech data.

[0056] Please refer to Figure 2 as Figure 2As shown, in the embodiments of the present application, the gated connection network 200 may include an encoding structure 210, a decoding structure 220, and a gating structure 230.

[0057] Among them, the encoding structure 210 includes a plurality of encoding layers 211. The encoding layer 211 may adopt a convolutional layer to extract the features of the training speech data. The number of encoding layers 211 in the encoding structure 210 may be 4 to 5 layers. Preferably, the encoding structure 210 may adopt 5 encoding layers. Optionally, the encoding layer 210 may adopt a 3×3 convolutional layer. As Figure 2 shown, the encoding structure 210 includes 4 layers of 3×3 convolutional layers. It can be understood that, according to actual usage needs, the number of encoding layers 211 can be adjusted, and the present application does not limit this.

[0058] The decoding structure 220 includes multiple decoding layers 221. The decoding layer 221 may adopt a transposed convolutional layer. The number of decoding layers 221 corresponds to the number of encoding layers to obtain enhanced speech data with the same dimension as the training speech data. The number of decoding layers 221 in the decoding structure 220 may be 4 to 5 layers. Preferably, the decoding structure 210 may adopt 5 decoding layers.. Optionally, the decoding layer 221 may adopt a 3×3 transposed convolutional layer. As Figure 2 shown, the decoding structure 220 includes 4 layers of 3×3 transposed convolutional layers. It can be understood that, according to actual usage needs, the number of decoding layers 221 can be adjusted, and the present application does not limit this.

[0059] Each encoding layer 211 in the encoding structure 210 corresponds to one of the decoding layers 221 in the decoding structure 220 and has the same dimension. As Figure 2 shown, the first encoding layer corresponds to the last decoding layer and has the same dimension; the second encoding layer corresponds to the penultimate decoding layer and has the same dimension; the third encoding layer corresponds to the second decoding layer and has the same dimension; the last encoding layer corresponds to the first decoding layer and has the same dimension.

[0060] The gating structure 230 is disposed between adjacent decoding layers 211. The gating structure 230 can be used to perform redundancy removal and transfer processing on the decoded output information. As Figure 2 、 Figure 3 shown, after the gating structure 230 performs redundancy removal processing on the decoded output information of the previous decoding layer 221, it is transferred to the next decoding layer 221. As Figure 2As shown, the gated connection network 200 is provided with three gated structures 230, which are respectively arranged between adjacent decoding layers 221. Optionally, the gated connection network 230 can be implemented by a gating function, which can identify the decoded output information, prevent the redundant information in the decoded output information from passing through, and allow the non-redundant information in the decoded output information to pass through and enter the next decoding layer 221. Optionally, the gating function (gateController) can adopt the Sigmoid function. The score range of Sigmoid is between [0, 1]. For example, gateController = Sigmoid(ux), where x can be the decoded output information and u is the control parameter of the gating system. When the decoded output information passes through the gated structure, the value of the control parameter u of the gating system can be obtained. If u = 1, the decoded output information can all pass through the gated structure. If u = 0, the decoded output information cannot pass through the gated structure.

[0061] In some embodiments, as Figure 2 shown, the gated connection network 200 further includes a timing structure 240. The timing structure 240 is arranged between the encoding structure 210 and the decoding structure 220. Voice data itself is a kind of timing data, and the timing structure 240 is used to obtain the timing information in the encoded output information. Optionally, the timing structure 240 can include at least one layer of long short-term memory layer (LSTM, Long Short Term). As Figure 2 shown, the timing structure 240 is provided with one layer of long short-term memory layer, which is arranged between the encoding structure 210 and the decoding structure 230. Thus, the timing structure 240 extracts the timing features from the encoded output information output by the encoding structure 210 and transmits the timing features to the decoding structure 220.

[0062] As Figure 4 shown, in some embodiments, the gated connection network 200 further includes a speech recognition structure 250. Among them, the speech recognition structure 250 is connected to the encoding structure 210, and the speech recognition structure 250 is used to perform speech recognition processing on the training speech data to determine whether the training speech data contains the target speech data. Optionally, the speech recognition structure 250 can include at least one layer of long short-term memory layer. Preferably, the speech recognition structure 250 can include two layers of long short-term memory layer. As Figure 4 shown, the speech recognition structure 250 is provided with one layer of long short-term memory layer. The speech recognition structure 250 is connected to the encoding structure 210 and is used to determine whether the training speech data includes the target speech data according to the encoded output information of the encoding structure 210.

[0063] In some embodiments, before step 110, the speech enhancement method of the embodiment of the present application can also perform the training steps of the pre-trained neural network, that is, steps 310 to 330.

[0064] Step 310: Obtain training speech data, where the training speech data includes mixed speech data.

[0065] In an embodiment of the present application, clean speech data with a preset sampling rate and a preset sampling precision can be used to generate a clean speech data set. Optionally, the preset sampling rate can be 16 kHz, and the preset sampling precision can be 16 bit. It can be understood that the present application is not limited thereto, and the preset sampling rate and the preset sampling precision can be adjusted according to actual training needs.

[0066] In an embodiment of the present application, the noise data set can use publicly available noise data and self-recorded noise data. Among them, the publicly available noise database can include, but is not limited to, publicly available noise databases such as the DEMEND noise library, the NoiseX-92 noise library, and the TUT acoustic scene database. In some embodiments, self-recorded noise data can also be used. It can be understood that the more types of noise, the better the training effect.

[0067] In an embodiment of the present application, the reverberation data is in two ways, the simulated RIR (Room Impulse Response) method and the recorded RIR reverberation data. The simulated RIR uses the Image Source Method. According to the size of the room and the set RT60 (Reverberation Time 60dB, which refers to the time used for the sound field to decay by 60dB, with the unit of second), the RIR is generated. The recorded data is the RIR reverberation data actually recorded in the first room, the second room, and the third room, where the size of the first room is larger than that of the second room, and the size of the second room is larger than that of the third room.

[0068] Further, the mixed speech data can be obtained by mixing noise data into the clean speech data. For example, through an aliasing tool, reverberation with different RT60s can be randomly added with a probability of 0.7, and noise data with a preset signal-to-noise ratio is randomly generated to form corresponding clean speech-noise speech pairs, which are used as training speech data to train the network. The preset signal-to-noise ratio can be [-10, 25] dB. It can be understood that the range of the preset signal-to-noise ratio can be set according to actual training needs, and the present application is not limited thereto.

[0069] Further, set target labels for the training data. Optionally, a pre-trained or prepared VAD (Voice Activity Detection) model can be used to set VAD labels for the training data. The VAD labels can be, for example, human voice labels and non-human voice labels. It can be understood that the present application is not limited thereto, and VAD labels can also be set according to actual training needs.

[0070] Further, a preset number of data for a preset time can be randomly selected from the dataset each time as the input of the gated connection network. The preset time can be, for example, 4s, and the preset number can be, for example, 16kHz. Thus, the number of sample points contained in 16kHz data for 4s is 4 × 16000 sampling points, and the speech spectrum is extracted with a preset frame length, a preset frame shift, and a preset number of Fourier points as the input of the gated connection network. Among them, the preset frame length can be 25ms (milliseconds), that is, 400 sampling points per frame. The preset frame shift can be 12.5ms, and the preset number of Fourier points can be 512 points. It can be understood that this application is not limited thereto, and the preset frame length, the preset frame shift, and the preset number of Fourier points can be set according to actual training needs, and this application does not limit this.

[0071] Step 320: Obtain the gated connection network.

[0072] In an embodiment of this application, the gated connection network can be, for example, Figure 4 the gated connection network shown, that is, the gated connection network includes an encoding structure, a decoding structure, a gating structure, a timing structure, and a speech recognition structure.

[0073] Step 330: Train the gated connection network with training speech data until the gated connection network meets the preset conditions, and obtain the trained gated connection network as the pre-trained neural network model.

[0074] Among them, the encoding structure is used to encode the training speech data to obtain encoded output information, the gating structure is used to remove redundancy and perform transfer processing on the decoded output information to obtain key decoded output information, the decoding structure is used to obtain enhanced speech data corresponding to the training speech data according to the encoded output information and the key decoded output information. The timing structure is used to obtain timing information according to the encoded output information and transfer it to the decoding structure. The speech recognition structure is used to determine whether the training speech data includes target speech data according to the encoded output information.

[0075] In some embodiments, step 330 may further include the following steps.

[0076] (1) Input the training speech data into the encoding structure, and generate encoded output information corresponding to the training speech data through the encoding structure.

[0077] (2) Input the encoded output information into the decoding structure, and output enhanced speech data corresponding to the training speech data through the decoding structure and the gating structure.

[0078] (3) Input the encoded output information into the speech recognition structure, and generate a speech recognition result corresponding to the speech data through the speech recognition structure.

[0079] (4) Obtain the total loss value of the gated connection network according to the enhanced speech data and the speech recognition result.

[0080] Specifically, obtaining the total loss value of the gated connection network according to the enhanced speech data and the speech recognition result may further include the following steps.

[0081] (4.1) Obtain the first loss value of the gated connection network according to the enhanced speech data.

[0082] Among them, the first loss value is used to characterize the difference between the spectrum of the enhanced speech data and the spectrum of the clean speech. The MSE (Mean Square Error) loss function can be used.

[0083] (4.2) Obtain the second loss value of the gated connection network according to the speech recognition result.

[0084] Among them, the second loss value can use the CrossEntropy (cross entropy) loss function.

[0085] (4.3) Based on the first loss value and the second loss value, determine the total loss value of the gated connection network.

[0086] (5) Perform iterative training on the gated connection network according to the total loss value until the gated connection network meets the preset conditions.

[0087] The following introduces the device embodiments of the present application, which can be used to execute the methods in the above embodiments of the present application. For the details not disclosed in the device embodiments of the present application, please refer to the above method embodiments of the present application.

[0088] Please refer to Figure 6 , which shows a speech enhancement device based on deep neural network technology provided by an embodiment of the present invention. The speech enhancement device 400 includes: a speech acquisition module 410 and a speech enhancement module 420.

[0089] Specifically, the speech acquisition module 410 is used to acquire the speech data to be processed. The speech enhancement module 420 is used to input the speech data to be processed into a pre-trained neural network model to obtain the enhanced speech data output by the neural network model. The enhanced speech data is obtained by filtering the noise data from the speech data to be processed. Among them, the pre-trained neural network model includes an encoding structure and a decoding structure, and is trained by encoding the training speech data and passing it to the decoding structure, and performing redundancy removal and transfer processing on the decoding output information between adjacent decoding layers.

[0090] In some embodiments of the present application, the pre-trained neural network model is also trained by performing speech recognition processing on the training speech data.

[0091] In some embodiments of the present application, the pre-trained neural network model is obtained by training a gated training network based on training speech data.

[0092] In some embodiments of the present application, the speech enhancement module 410 further includes: a training data acquisition unit, a network acquisition unit, and a network training unit. Among them, the training data acquisition unit is used to acquire training speech data, and the training speech data includes mixed speech data. The network acquisition unit is used to acquire a gated connection network, and the gated connection network includes an encoding structure, a gating structure, and a decoding structure. The network training unit is used to train the gated connection network with the training speech data until the gated connection network meets a preset condition, and obtain the trained gated connection network as the pre-trained neural network model. Among them, the encoding structure is used to encode the training speech data to obtain encoded output information, the gating structure is used to perform redundancy removal and transfer processing on the decoded output information to obtain key decoded output information, and the decoding structure is used to obtain enhanced speech data corresponding to the training speech data according to the encoded output information and the key decoded output information.

[0093] In some embodiments of the present application, the encoding structure includes a plurality of encoding layers, and the decoding structure includes a plurality of decoding layers. The encoded output information of each encoding layer is input to the corresponding decoding layer; a gating structure is arranged between adjacent decoding layers.

[0094] In some embodiments of the present application, the gated connection network further includes a timing structure, and the timing structure is arranged between the encoding structure and the decoding structure, and the timing structure is used to obtain the timing information in the encoded output information.

[0095] In some embodiments of the present application, the gated connection network further includes a speech recognition structure, and the speech recognition structure is connected to the encoding structure, and the speech recognition structure is used to perform speech recognition processing on the training speech data.

[0096] In some embodiments of the present application, the gating training unit includes: an encoding output subunit, an enhanced speech subunit, a speech recognition subunit, a total loss subunit, and an iteration subunit. Among them, the encoding output subunit is configured to input training speech data into an encoding structure and generate encoding output information corresponding to the training speech data through the encoding structure. The speech recognition subunit is configured to input the encoding output information into a decoding structure and output enhanced speech data corresponding to the training speech data through the decoding structure and a gating structure. The total loss subunit is configured to input the encoding output information into a speech recognition structure and generate a speech recognition result corresponding to the speech data through the speech recognition structure; the total loss subunit is configured to obtain the total loss value of the gating connection network according to the enhanced speech data and the speech recognition result. The iteration subunit is configured to perform iterative training on the gating connection network according to the total loss value until the gating connection network meets a preset condition.

[0097] In some embodiments of the present application, the total loss subunit includes a first loss sub-subunit, a second loss sub-subunit, and a total loss sub-subunit. Among them, the first loss sub-subunit is configured to obtain the first loss value of the gating connection network according to the enhanced speech data. The second loss sub-subunit is configured to obtain the second loss value of the gating connection network according to the speech recognition result. The total loss sub-subunit is configured to determine the total loss value of the gating connection network based on the first loss value and the second loss value.

[0098] In some embodiments of the present application, the mixed speech data is obtained by mixing noise data into clean speech data.

[0099] It should be noted that the various embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For device embodiments, since they are basically similar to method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments. For any processing method described in the method embodiments, it can be implemented by a corresponding processing module in the device embodiments, and will not be elaborated one by one in the device embodiments.

[0100] In addition, in each embodiment of the present application, the various functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.

[0101] Please refer to Figure 7, based on the above speech enhancement method based on deep network learning, another electronic device 500 including a processor that can execute the foregoing speech enhancement method is further provided in an embodiment of the present application. The electronic device 500 further includes one or more processors 510, a memory 520, and one or more applications. Among them, a program that can execute the content in the foregoing embodiments is stored in the memory 520, and the processor 510 can execute the program stored in the memory.

[0102] Among them, the processor 510 may include one or more cores for processing data and a message matrix unit. The processor 510 connects various parts within the entire electronic device through various interfaces and lines, and executes various functions of the electronic device 500 and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory, and calling data stored in the memory. Optionally, the processor 10 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor may integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, and application programs, etc.; the GPU is responsible for rendering and drawing display content; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor and may be implemented separately by a communication chip.

[0103] The memory 520 may include random access memory (RAM) and may also include read-only memory. The memory 520 can be used to store instructions, programs, code, code sets, or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function, instructions for implementing the following various method embodiments, etc. The data storage area may also store data created during the use of the terminal.

[0104] Please refer to Figure 7, which shows a structural block diagram of a computer-readable storage medium provided by an embodiment of the present application. Program code 610 is stored in the computer-readable medium, and the program code 610 can be called by a processor to execute the image recognition method described in the above method embodiment.

[0105] The computer-readable storage medium 600 can be an electronic memory such as a flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, a hard disk, or a ROM. Optionally, the computer-readable storage medium 600 includes a non-transitory computer-readable storage medium. The computer-readable storage medium has a storage space for program code that executes any method step in the above method. These program codes can be read out from or written into one or more computer program products. The program code can be compressed in an appropriate form, for example.

[0106] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method in any of the above embodiments.

[0107] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-mentioned modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0108] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0109] In summary, a speech enhancement method based on deep neural network technology provided by the present invention includes obtaining speech data to be processed; inputting the speech data to be processed into a pre-trained neural network model to obtain enhanced speech data output by the neural network model; the enhanced speech data is obtained by filtering noise data from the speech data to be processed; wherein, the pre-trained neural network model includes an encoding structure and a decoding structure, and is trained by encoding the training speech data and passing it to the decoding structure, and performing redundancy removal and transfer processing on the decoding output information between adjacent decoding layers. Thus, by processing the speech data to be processed through the pre-trained neural network model, noise data can be quickly and effectively suppressed, and the pre-trained neural network model focuses more on effective information by performing redundancy removal processing on the decoding output information, significantly improving the noise reduction effect.

[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A speech enhancement method based on deep neural network technology, characterized in that, The method includes: Obtaining the speech data to be processed; Inputting the speech data to be processed into a pre-trained neural network model to obtain enhanced speech data output by the neural network model; the enhanced speech data is obtained by filtering out noise data from the speech data to be processed; Wherein, the pre-trained neural network model includes an encoding structure and a decoding structure, and is trained by encoding the training speech data and passing it to the decoding structure, and performing redundancy removal and passing processing on the decoding output information between adjacent decoding layers; The pre-trained neural network model is obtained by training a gated connection network based on the training speech data, and the gated connection network includes an encoding structure, a gating structure and a decoding structure; The encoding structure includes a plurality of encoding layers, and the decoding structure includes a plurality of decoding layers; The number of the encoding layers corresponds to the number of the decoding layers; The gating structure is arranged between adjacent decoding layers, and the gating structure is used to perform redundancy removal processing on the decoding output information of the previous decoding layer and then pass it to the next decoding layer.

2. The method according to claim 1, wherein The pre-trained neural network model is trained by performing speech recognition processing on the training speech data.

3. The method according to claim 2, characterized in that, The gated connection network further includes a timing structure, and the timing structure is arranged between the encoding structure and the decoding structure, and the timing structure is used to obtain the timing information in the encoding output information output by the encoding structure.

4. The method according to claim 3, characterized in that, The gated connection network further includes a speech recognition structure, and the speech recognition structure is connected to the encoding structure, and the speech recognition structure is used to perform speech recognition processing on the training speech data.

5. The method according to claim 1, wherein The encoding output information of each encoding layer is input to the corresponding decoding layer.

6. The method according to claim 4, wherein The pre-trained neural network model is trained through the following steps: Obtaining training speech data, and the training speech data includes mixed speech data; Obtaining a gated connection network; Training the gated connection network with the training speech data until the gated connection network meets a preset condition, and obtaining the trained gated connection network as the pre-trained neural network model; Wherein, the encoding structure is used to encode the training speech data to obtain encoding output information, the gating structure is used to perform redundancy removal and passing processing on the decoding output information to obtain key decoding output information, and the decoding structure is used to obtain the enhanced speech data corresponding to the training speech data according to the encoding output information and the key decoding output information; the timing structure is used to obtain timing information according to the encoding output information and pass it to the decoding structure; The speech recognition structure is used to determine whether the training speech data includes target speech data according to the encoding output information.

7. The method according to claim 6, wherein The training of the gated connection network with the training speech data until the gated connection network meets a preset condition includes: Inputting the training speech data into the encoding structure, and generating encoding output information corresponding to the training speech data through the encoding structure; Input the encoded output information into the decoding structure, and output enhanced speech data corresponding to the training speech data through the decoding structure and the gating structure; Input the encoded output information into the speech recognition structure, and generate a speech recognition result corresponding to the speech data through the speech recognition structure; Obtain the total loss value of the gating connection network according to the enhanced speech data and the speech recognition result; Iteratively train the gating connection network according to the total loss value until the gating connection network meets the preset conditions.

8. The method according to claim 7, characterized in that The obtaining the total loss value of the gating connection network according to the enhanced speech data and the speech recognition result includes: Obtain the first loss value of the gating connection network according to the enhanced speech data; Obtain the second loss value of the gating connection network according to the speech recognition result; Based on the first loss value and the second loss value, determine the total loss value of the gating connection network.

9. The method according to any one of claims 6 - 8, characterized in that, The mixed speech data is obtained by mixing noise data into clean speech data.

Citation Information

Patent Citations

  • Audio-visual speech enhancement

    US20210134312A1

  • Methods and apparatus of residual and coefficient coding

    WO2021062014A1