Speech enhancement method, device and computer-readable storage medium

Through the bone conduction sensor combined with the decoding prediction model, the high-frequency part of the speech data is predicted and spliced, which solves the problems of noise suppression and non-stable noise in the speech signal collected by the microphone, and achieves high-quality enhancement of speech data in the full-band.

CN116403591BActive Publication Date: 2025-08-26GOERTEK INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310494248.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-26
Publication Date
2025-08-26
Estimated Expiration
2043-04-26

AI Technical Summary

Technical Problem

In the prior art, when processing voice signals, especially voice signals collected through microphones, it is difficult to effectively suppress non-steady state noise such as wind noise, and the voice enhancement effect is poor when the signal-to-noise ratio is extremely low.

Method used

The bone conduction sensor is used to collect voice data, and the decoding prediction model predicts and splices the high-frequency part data based on the correlation between the low-frequency part of the human voice spectrum to achieve the enhancement of voice data in the full-band.

Benefits of technology

Effectively reduce noise, improve voice quality, and obtain natural listening effects, make up for the problem of missing parts of bone conduction sensor acquisition, and improve voice enhancement effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116403591B_ABST
    Figure CN116403591B_ABST
Patent Text Reader

Abstract

The present invention discloses a speech enhancement method, device, and computer-readable storage medium. The method comprises: obtaining first speech data collected by a bone conduction sensor; inputting speech data within a first preset frequency band in the first speech data into a preset decoding prediction model for prediction, thereby obtaining decoded speech data within a second preset frequency band, wherein the first preset frequency band is smaller than the second preset frequency band; and concatenating the decoded speech data with the speech data within the first preset frequency band in the first speech data to obtain enhanced speech data. The present invention predicts the high-frequency portion of the audio data collected by the bone conduction sensor, obtaining speech data across the entire frequency band, and improving the speech enhancement effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of audio processing technology, and in particular to a speech enhancement method, device, and computer-readable storage medium. Background Art

[0002] In real life, voice signals are common signals in our lives. When a microphone picks up a voice signal, it will inevitably be interfered with by the surrounding environment noise, transmission medium noise, internal electrical noise of the communication equipment, room reverberation, and the voices of other speakers. Therefore, the quality of the picked up voice is affected.

[0003] In conventional technology, noise reduction or speech enhancement is usually performed on the data collected by the microphone to eliminate interference noise with minimal loss of the speech signal. However, this method only has a good suppressing effect on steady-state noise, and is less effective for non-steady-state noise (such as wind noise). At the same time, the speech enhancement effect will be poor when the signal-to-noise ratio is extremely low. Summary of the Invention

[0004] The main purpose of the present invention is to provide a speech enhancement method, device and computer-readable storage medium, aiming to provide a solution for predicting the high-frequency portion of audio data collected by a bone conduction sensor, thereby obtaining speech data of the entire frequency band and improving the speech enhancement effect.

[0005] To achieve the above object, the present invention provides a speech enhancement method, which comprises the following steps:

[0006] Acquiring first voice data collected by the bone conduction sensor;

[0007] Inputting speech data in a first preset frequency band of the first speech data into a preset decoding prediction model for prediction to obtain decoded speech data in a second preset frequency band, wherein the first preset frequency band is smaller than the second preset frequency band;

[0008] The decoded voice data is concatenated with voice data in the first preset frequency band of the first voice data to obtain enhanced voice data.

[0009] Optionally, before the step of inputting the speech data in the first preset frequency band in the first speech data into a preset decoding prediction model for prediction to obtain decoded speech data in the second preset frequency band, the method further includes:

[0010] Acquire initial training voice data collected by a microphone in a quiet environment, and extract first training voice data within the first preset frequency band from the initial training voice data;

[0011] Obtaining a decoding label of the first training speech data;

[0012] Taking the first training speech data and the decoding label as one piece of training data, and obtaining a first training data set according to each piece of acquired training data;

[0013] The first training data set is used to train a preset decoding prediction model to be trained to obtain the decoding prediction model.

[0014] Optionally, the step of obtaining a decoding label of the first training speech data includes:

[0015] Extracting second training voice data in the second preset frequency band from the initial training voice data;

[0016] The second training speech data is used as a decoding label of the first training speech data.

[0017] Optionally, after the step of obtaining the enhanced voice data, the method further includes:

[0018] The enhanced speech data is input into a preset restoration prediction model for prediction to obtain restoration speech data.

[0019] Optionally, before the step of inputting the enhanced speech data into a preset restoration prediction model for prediction to obtain the restored speech data, the method further includes:

[0020] Acquire repair training data obtained based on the decoding prediction model, and acquire repair labels for the repair training data;

[0021] Taking the repair training data and the repair label as one piece of training data, and obtaining a second training data set according to each piece of acquired training data;

[0022] The preset repair prediction model to be trained is trained using the second training data set to obtain the repair prediction model.

[0023] Optionally, the step of obtaining repair training data based on the decoding prediction model includes:

[0024] Inputting the first training speech data into the decoding prediction model for prediction to obtain decoding training data within the second preset frequency band;

[0025] The decoded training data is concatenated with the first training speech data to obtain the repaired training data.

[0026] Optionally, the step of obtaining the repair label of the repair training data includes:

[0027] The initial training speech data is used as the repair label of the repair training data.

[0028] Optionally, the step of concatenating the decoded voice data with voice data in the first preset frequency band in the first voice data to obtain enhanced voice data includes:

[0029] Performing a time-domain to frequency-domain conversion on the voice data in the first preset frequency band in the first voice data to obtain a first frequency spectrum;

[0030] Inputting the first spectrum into the decoding prediction model for prediction to obtain a second spectrum;

[0031] The second spectrum is converted from the frequency domain to the time domain to obtain the decoded speech data.

[0032] To achieve the above-mentioned objectives, the present invention also provides a speech enhancement device, which includes: a memory, a processor, and a speech enhancement program stored in the memory and executable on the processor. When the speech enhancement program is executed by the processor, the steps of the speech enhancement method described above are implemented.

[0033] In addition, to achieve the above-mentioned purpose, the present invention also proposes a computer-readable storage medium, on which a speech enhancement program is stored. When the speech enhancement program is executed by a processor, the steps of the speech enhancement method described above are implemented.

[0034] In the present invention, by acquiring first voice data collected by a bone conduction sensor, a large amount of noise present in voice data collected using a microphone is greatly reduced. Then, voice data in the first voice data within a first preset frequency band is input into a preset decoding prediction model for prediction to obtain decoded voice data in a second preset frequency band, wherein the preset first preset frequency band is smaller than the preset second preset frequency band. Based on the correlation between the low-frequency and high-frequency parts of the spectrum of the human voice, the decoding prediction model is used to predict the high-frequency part data corresponding to the low-frequency voice data, thereby decoding the voice data from low frequency to high frequency. Then, the decoded voice data is spliced ​​with the voice data in the first preset frequency band of the first voice data to obtain enhanced voice data. By splicing the low-frequency and high-frequency parts of the voice data, full-band voice data is obtained, which makes up for the lack of high-frequency parts in the voice data collected by the bone conduction sensor, resulting in a poor listening experience of the voice data. The obtained enhanced voice data has a good noise reduction effect while exhibiting a more natural listening effect, further improving the voice enhancement effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1A schematic diagram of the hardware operating environment involved in an embodiment of the present invention;

[0036] Figure 2 This is a flow chart of a first embodiment of a speech enhancement method according to the present invention;

[0037] Figure 3 2 is a flow chart of a second embodiment of a speech enhancement method according to the present invention;

[0038] Figure 4 1 is a flow chart of a third embodiment of a speech enhancement method according to the present invention;

[0039] Figure 5 The figure is a flow chart of a relatively complete embodiment of the speech enhancement method of the present invention.

[0040] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0041] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0042] like Figure 1 As shown, Figure 1 It is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment of the present invention.

[0043] It should be noted that the voice enhancement device in the embodiment of the present invention can be a headset, a virtual reality device, a smart phone, a personal computer, a server and other devices, and no specific limitation is made here.

[0044] like Figure 1 As shown, the speech enhancement device may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. The communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0045] Those skilled in the art will understand that Figure 1The device structure shown in the figure does not constitute a limitation on the speech enhancement device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0046] like Figure 1 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a voice enhancement program. The operating system is a program that manages and controls the hardware and software resources of the device and supports the operation of the voice enhancement program and other software or programs. Figure 1 In the device shown, the user interface 1003 is mainly used to communicate data with the client; the network interface 1004 is mainly used to establish a communication connection with the server; and the processor 1001 can be used to call the voice enhancement program stored in the memory 1005 and perform the following operations:

[0047] Acquiring first voice data collected by the bone conduction sensor;

[0048] Inputting speech data in a first preset frequency band of the first speech data into a preset decoding prediction model for prediction to obtain decoded speech data in a second preset frequency band, wherein the first preset frequency band is smaller than the second preset frequency band;

[0049] The decoded voice data is concatenated with voice data in the first preset frequency band of the first voice data to obtain enhanced voice data.

[0050] Furthermore, before inputting the speech data in the first preset frequency band of the first speech data into a preset decoding prediction model for prediction to obtain decoded speech data in the second preset frequency band, the processor 1001 may also be configured to call a speech enhancement program stored in the memory 1005 and perform the following operations:

[0051] Acquire initial training voice data collected by a microphone in a quiet environment, and extract first training voice data within the first preset frequency band from the initial training voice data;

[0052] Obtaining a decoding label of the first training speech data;

[0053] Taking the first training speech data and the decoding label as one piece of training data, and obtaining a first training data set according to each piece of acquired training data;

[0054] The first training data set is used to train a preset decoding prediction model to be trained to obtain the decoding prediction model.

[0055] Furthermore, the operation of obtaining the decoding label of the first training speech data includes:

[0056] Extracting second training voice data in the second preset frequency band from the initial training voice data;

[0057] The second training speech data is used as a decoding label of the first training speech data.

[0058] Furthermore, after the operation of obtaining the enhanced voice data, the processor 1001 may also be configured to call a voice enhancement program stored in the memory 1005 and perform the following operations:

[0059] The enhanced speech data is input into a preset restoration prediction model for prediction to obtain restoration speech data.

[0060] Furthermore, before inputting the enhanced speech data into a preset repair prediction model for prediction to obtain the repaired speech data, the processor 1001 may also be configured to call a speech enhancement program stored in the memory 1005 to perform the following operations:

[0061] Acquire repair training data obtained based on the decoding prediction model, and acquire repair labels for the repair training data;

[0062] Taking the repair training data and the repair label as one piece of training data, and obtaining a second training data set according to each piece of acquired training data;

[0063] The preset repair prediction model to be trained is trained using the second training data set to obtain the repair prediction model.

[0064] Furthermore, the operation of obtaining the repair training data based on the decoding prediction model includes:

[0065] Inputting the first training speech data into the decoding prediction model for prediction to obtain decoding training data within the second preset frequency band;

[0066] The decoded training data is concatenated with the first training speech data to obtain the repaired training data.

[0067] Furthermore, the operation of obtaining the repair label of the repair training data includes:

[0068] The initial training speech data is used as the repair label of the repair training data.

[0069] Furthermore, the operation of splicing the decoded voice data with voice data in the first preset frequency band in the first voice data to obtain enhanced voice data includes:

[0070] Performing a time-domain to frequency-domain conversion on the voice data in the first preset frequency band in the first voice data to obtain a first frequency spectrum;

[0071] Inputting the first spectrum into the decoding prediction model for prediction to obtain a second spectrum;

[0072] The second spectrum is converted from the frequency domain to the time domain to obtain the decoded speech data.

[0073] Based on the above structure, various embodiments of the speech enhancement method are proposed.

[0074] Reference Figure 2 , Figure 2 FIG. 4 is a flow chart of a first embodiment of a speech enhancement method according to the present invention.

[0075] The embodiments of the present invention provide embodiments of a method for speech enhancement. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown. In this embodiment, the execution subject of the speech enhancement method may be a headset, a virtual reality device, a personal computer, a smart phone, or other device, which is not limited in this embodiment. For ease of description, the following description of each embodiment will be omitted. In this embodiment, the speech enhancement method includes:

[0076] Step S10, acquiring first voice data collected by a bone conduction sensor;

[0077] A signal (hereinafter referred to as a bone conduction signal for distinction) may be collected by a bone conduction sensor, and the bone conduction signal may be converted and processed to obtain voice data (hereinafter referred to as the first voice data for distinction).

[0078] In one feasible embodiment, a bone conduction sensor can be set in a product to collect voice data; for example, it is set in headphones or virtual reality equipment; the specific setting location is designed according to needs, for example, the bone conduction sensor is usually set in a position where it contacts the human skull.

[0079] In another feasible implementation, the first voice data can be voice data collected in real time, or it can be non-real-time recorded voice data. Different implementation methods can be selected according to different real-time requirements for voice enhancement in application scenarios. This embodiment does not limit this.

[0080] Step S20: Inputting the speech data in the first preset frequency band into a preset decoding prediction model for prediction to obtain decoded speech data in a second preset frequency band, wherein the first preset frequency band is smaller than the second preset frequency band;

[0081] The first voice data is split into different frequency bands, and then the voice data obtained by the splitting and located in the first preset frequency band is input into a pre-trained decoding prediction model for prediction. Based on the correlation between the low frequency (first preset frequency band) and high frequency (second preset frequency band) of the human voice, the decoded voice data (hereinafter referred to as decoded voice data for distinction) in the second preset frequency band is obtained.

[0082] Exemplarily, if the first voice data all belongs to the first preset frequency band, the first voice data is input into the decoding prediction model for prediction.

[0083] Exemplarily, the decoding prediction model can be a pre-trained neural network model; based on the correlation between the low-frequency part of the speech (first preset frequency band) and the high-frequency part (second preset frequency band), the low-frequency speech data and the associated high-frequency speech data are predicted, thereby performing speech enhancement on the low-frequency speech data collected by the bone conduction sensor to improve the speech quality; in this embodiment, there is no restriction on the specific structure of the decoding prediction model, for example, it can be implemented using network structures such as convolutional neural networks or recurrent neural networks.

[0084] A frequency band refers to a frequency range, which includes multiple frequency points. A first preset frequency band being smaller than a second preset frequency band means that the maximum frequency point of the first preset frequency band is smaller than the minimum frequency point of the second preset frequency band. The boundary frequency point between the first preset frequency band and the second preset frequency band can be set as needed and is not limited in this embodiment. For example, the first preset frequency band is within 1 kHz, and the second preset frequency band is above 1 kHz, including 1 kHz.

[0085] In this embodiment, since the noise reduction effect of voice data collected by a microphone is poor, and the signal collected by the bone conduction sensor is the signal generated by the vibration of the skull when a person speaks, the interference of external noise on the bone conduction signal is minimal. Therefore, using a bone conduction sensor instead of a conventional microphone to collect voice data can greatly reduce the noise in the voice data; however, the upper limit frequency of the bone conduction sensor is generally 1kHz, that is, the bone conduction sensor can only collect voice data within 1kHz, making the voice data relatively low and unpleasant to listen to; therefore, the correlation between the low-frequency and high-frequency parts of the human voice spectrum is used to train a decoding prediction model, so that the decoding prediction model can predict the high-frequency voice data corresponding to the low-frequency voice data, realize low-frequency to high-frequency decoding of the data, and thus achieve voice enhancement.

[0086] In one feasible implementation, step S20, inputting speech data in the first preset frequency band of the first speech data into a preset decoding prediction model for prediction to obtain decoded speech data in the second preset frequency band includes:

[0087] Step S201, performing a time domain to frequency domain conversion on the first voice data in a first preset frequency band to obtain a first spectrum;

[0088] The voice data in the first voice data that is within the first preset frequency band is converted from the time domain to the frequency domain, thereby obtaining a first spectrum associated with the first preset frequency band. The spectrum may further include amplitude, phase angle value, wavelength, etc., which is not limited in this embodiment; wherein, the conversion from the time domain to the frequency domain may be achieved through Fourier transform (FFT); illustratively, the complex numbers of each frame or each frequency point are first converted, and then the spectrum amplitude and phase angle value are calculated based on the complex numbers.

[0089] Step S202: input the first spectrum into a decoding prediction model for prediction to obtain a second spectrum;

[0090] Inputting a first spectrum associated with a first preset frequency band into a pre-trained decoding prediction model for prediction to obtain a second spectrum associated with a second preset frequency band;

[0091] Exemplarily, the decoding prediction model can be a pre-trained neural network model; based on the correlation between the low-frequency part (first preset frequency band) and the high-frequency part (second preset frequency band) of the speech, the low-frequency speech data and the associated high-frequency speech data are predicted, thereby performing speech enhancement on the low-frequency speech data collected by the bone conduction sensor to improve the speech quality.

[0092] In one feasible implementation, the decoding prediction model training process can be performed by using training speech data collected by a microphone in a quiet environment (hereinafter referred to as initial training speech data for distinction), splitting the initial training speech data into a first preset frequency band and a second preset frequency band, and then using the first training speech data in the first preset frequency band as input data, and the second training speech data in the associated second preset frequency band as decoding labels, and using a supervised training method for training. That is, the decoding labels are used to supervise the results predicted by the decoding prediction model, and then the model parameters of the decoding prediction model are continuously updated, so that the enhanced speech predicted by the updated decoding prediction model becomes closer and closer to the speech data collected by the microphone in a quiet environment. The type of the decoding prediction model can be a deep neural network, a recurrent neural network, a convolutional neural network, or a combination of multiple neural networks, etc., which is not limited in this embodiment.

[0093] In another feasible implementation, the first amplitude and first phase angle value of each frame are input into a decoding prediction model for prediction to obtain the second amplitude and second phase angle value of each frame; accurate voice data is predicted by the amplitude, and voice data that is more natural to the user is predicted based on the phase angle value, thereby further improving the voice enhancement effect.

[0094] Step S203: convert the second spectrum from the frequency domain to the time domain to obtain decoded speech data.

[0095] The obtained second spectrum is converted from the frequency domain to the time domain to obtain decoded speech data; wherein the conversion from the frequency domain to the time domain can be achieved by inverse Fourier transform.

[0096] In this embodiment, a first spectrum is obtained by converting the voice data in the first voice data that is within the first preset frequency band from the time domain to the frequency domain; and then the first spectrum is input into the decoding prediction model for prediction to obtain a second spectrum, so that the decoding prediction model can obtain the second spectrum in the associated second preset frequency band based on the first spectrum corresponding to the first preset frequency band, thereby realizing decoding of the voice data from low frequency to high frequency, improving the voice effect of the voice data collected by the bone conduction sensor, making the user's listening experience of the voice data more natural, and further improving the voice enhancement effect.

[0097] Step S30 : Concatenate the decoded voice data with the voice data in the first voice data that is within the first preset frequency band to obtain enhanced voice data.

[0098] The frequency domain of the voice data collected by the microphone is relatively complete and can include low-frequency and high-frequency parts, but its noise resistance is poor; while the voice data collected by the bone conduction sensor is mainly concentrated in the low-frequency part and has excellent noise resistance, but the voice data collected by bone conduction will lose the high-frequency information of the data, making the voice listening experience poor. Therefore, the decoded voice data in the second preset frequency band (high frequency) obtained by the decoding prediction model is spliced ​​with the voice data in the first preset frequency band (low frequency) to obtain enhanced voice data of the entire frequency band, which overcomes the defect of losing high-frequency information in the voice data collected by bone conduction and further improves the voice enhancement effect.

[0099] In a specific embodiment, voice data is collected through a bone conduction sensor provided in the headphone device, i.e., recording is performed, and then the voice data is processed through a pre-selected and trained decoding prediction model to obtain enhanced voice data with better listening experience. When the user uses or plays this enhanced voice data next time, a high-quality recording can be obtained.

[0100] In this embodiment, by acquiring first voice data collected by a bone conduction sensor, a large amount of noise present in voice data collected using a microphone is greatly reduced. Voice data within a first preset frequency band in the first voice data is then input into a preset decoding prediction model for prediction, thereby obtaining decoded voice data within a second preset frequency band, wherein the preset first preset frequency band is smaller than the preset second preset frequency band. Based on the correlation between the low-frequency and high-frequency portions of the human voice spectrum, the decoding prediction model is used to predict high-frequency data corresponding to the low-frequency voice data, thereby decoding the voice data from low frequency to high frequency. The decoded voice data is then concatenated with the voice data within the first preset frequency band in the first voice data to obtain enhanced voice data. By concatenating the low-frequency and high-frequency portions of the voice data, full-band voice data is obtained, compensating for the poor audibility of the voice data due to the lack of high-frequency portions in the voice data collected by the bone conduction sensor. This allows the obtained enhanced voice data to exhibit a more natural listening experience while achieving good noise reduction, further enhancing the voice enhancement effect. It makes up for the shortcomings of high noise in voice data collected by microphones and the poor listening experience of voice data collected by conventional bone conduction sensors; based on the correlation between the low-frequency and high-frequency parts of the human voice spectrum, the decoding prediction model is used to predict voice data with a complete spectrum, thereby improving the voice enhancement effect.

[0101] Further, based on the above first embodiment, a second embodiment of the speech enhancement method of the present invention is proposed, referring to Figure 3 In this embodiment, before step S20, the method further includes:

[0102] Step A10: acquiring initial training voice data collected by a microphone in a quiet environment, and extracting first training voice data within a first preset frequency band from the initial training voice data;

[0103] In order to train the decoding prediction model so that the decoding prediction model can clearly obtain the relationship between the first preset frequency band and the second preset frequency band of the speech data, clean speech data (hereinafter referred to as initial training speech data) is collected through a microphone in a quiet environment to avoid the influence of noise on training; and then the first training speech data in the initial training speech data that is within the first preset frequency band is extracted to realize the splitting of different frequency band data of the initial training speech data.

[0104] In one feasible implementation, the initial training voice data collected by the microphone is converted from the time domain to the frequency domain, and then the first training voice data in the first preset frequency band is extracted from the initial training voice data.

[0105] Step A20, obtaining a decoding label of the first training speech data;

[0106] Obtain a decoded label of the first training speech data, wherein the decoded label is a correct output or category of each single frame of the first training speech data instance.

[0107] In one feasible implementation, step A20, the step of obtaining a decoding label of the first training speech data includes:

[0108] Step A201, extracting second training voice data in a second preset frequency band from the initial training voice data;

[0109] Since the initial training voice data collected by the microphone in a quiet environment includes data of the entire frequency band, the second training voice data in the second preset frequency band can be extracted from the initial training voice data collected by the microphone in a quiet environment.

[0110] Step A202: Use the second training speech data as the decoding label of the first training speech data.

[0111] The extracted second training voice data is used as the decoding label of the first training voice data; since the initial voice data collected by the microphone contains data of the entire frequency band, that is, it contains voice data of the first preset frequency band and the second preset frequency band at the same time, the first training voice data of the first preset frequency band is used as input data, and the second training voice data of the associated second frequency band is used as the decoding label to judge the prediction accuracy of the decoding prediction model, and continuously train and optimize the decoding prediction model.

[0112] In this embodiment, since the microphone can collect voice data in a wider frequency band, the decoding prediction model can be trained by analyzing the relationship between the first preset frequency band and the second preset frequency band in the voice data collected by the microphone, and then the second training voice data in the second preset frequency band in the initial training voice data is extracted, and the second training voice data is used as the decoding label of the first training voice data; by obtaining the decoding label, the decoding prediction model is trained, and at the same time, the decoding label can be directly extracted from the voice data collected by the microphone without manual operation. When the data volume of the training data set is large, the efficiency of constructing the training data set is greatly improved, thereby improving the training efficiency of the decoding prediction model.

[0113] Step A30: taking the first training speech data and the decoding label as one piece of training data, and obtaining a first training data set based on each piece of acquired training data;

[0114] The first training speech data and the decoding label of the first training speech data are taken as a separate piece of training data, and the pieces of training data are combined to obtain a first training data set, that is, a training data set for the decoding prediction model.

[0115] In one feasible implementation, training speech data collected by a bone conduction sensor in a quiet environment (hereinafter referred to as bone conduction training speech data) is obtained, and first bone conduction training data within a first preset frequency band is extracted from the bone conduction training speech data; wherein the bone conduction training speech data and the initial training speech data collected by a microphone are collected synchronously in the same environment; the first bone conduction training data, the first training speech data, and the decoding label are used as one piece of training data, and a first training data set is obtained based on each piece of the obtained training data; the first training data set is used to train a preset decoding prediction model to be trained to obtain a decoding prediction model; and the decoding prediction model to be trained is trained using the bone conduction speech data and microphone speech data collected synchronously in the same environment, thereby enhancing the credibility of the training results of the decoding prediction model and thereby improving the speech enhancement effect of the decoding prediction model.

[0116] In a specific embodiment, the initial training speech data and bone conduction training speech data used in training can be obtained by playing the same speech in an experimental environment and then collecting it using a microphone and bone conduction sensor. The number of samples used for training can be set as needed and is not limited in this embodiment.

[0117] Step A40: Use the first training data set to train the preset decoding prediction model to obtain a decoding prediction model.

[0118] The first training data set is used to train a preset decoding prediction model to be trained, thereby setting and updating parameters of the decoding prediction model. After multiple rounds of training, a decoding prediction model is obtained.

[0119] In one feasible implementation, the updated decoding model is used as the basis for the next round of training; after multiple rounds of training, the updated decoding model is used as the trained decoding model.

[0120] In another feasible implementation manner, step A40, using the first training data set to train the preset decoding prediction model to obtain the decoding prediction model, includes:

[0121] Step A401: inputting first training speech data into a decoding prediction model to be trained for prediction to obtain predicted decoding data within a second preset frequency band;

[0122] The first training speech data is input into the decoding prediction model to be trained for prediction to obtain predicted decoding data within the second preset frequency band. The predicted decoding data is the data predicted by the decoding prediction model to be trained or the decoding prediction model that has not completed training.

[0123] Step A402 : determining a decoding loss value between the predicted decoding data and the decoding label, and updating the decoding prediction model to be trained according to the decoding loss value to obtain a trained decoding prediction model.

[0124] The predicted decoding data is compared with the corresponding decoding label to determine the decoding loss value between the predicted decoding data and the corresponding decoding label, and then the decoding prediction model to be trained is updated according to the obtained decoding loss value. After multiple rounds of training, the trained decoding prediction model is obtained.

[0125] In a feasible implementation manner, the predicted decoding data and the corresponding decoding label may be compared by subtracting the two and taking the absolute value, or taking the square root, etc., which is not limited in this embodiment.

[0126] In another feasible implementation, the decoding prediction model may be subjected to multiple rounds of iterative training. In the first round of training, the initialized decoding prediction model is updated, and in subsequent rounds of training, the decoding prediction model updated in the previous round of training is used as the basis for updating.

[0127] In another feasible implementation, the decoding loss value can be expressed as a model error value, and the model error value is compared with a preset model error threshold. If the model error value is greater than the preset model error threshold, the decoding prediction model is updated based on the model error value; if the model error value is less than or equal to the preset model error threshold, the training of the decoding prediction model is completed.

[0128] In this embodiment, since the microphone can collect voice data in a wider frequency band, the decoding prediction model can be trained by analyzing the relationship between the first preset frequency band and the second preset frequency band of each frame in the voice data collected by the microphone; then, the initial training voice data collected by the microphone in a quiet environment is obtained, and the first training voice data in the initial training voice data that is within the first preset frequency band is extracted, and then the decoding label of the first training voice data is obtained; the first training voice data and the decoding label are used as a training data, and a first training data set is obtained based on the obtained training data, and the preset decoding prediction model to be trained is trained using the first training data set to obtain a decoding prediction model; by training the decoding prediction model, the credibility of the decoding prediction model training result is enhanced, thereby improving the voice enhancement effect of the decoding prediction model.

[0129] Furthermore, based on the above-mentioned first and / or second embodiments, a third embodiment of the speech enhancement method of the present invention is proposed, referring to Figure 4 In this embodiment, after step S30, the method further includes:

[0130] Step S40: input the enhanced speech data into a preset restoration prediction model for prediction to obtain restoration speech data.

[0131] The enhanced speech data is the full-band speech data obtained by prediction through the decoding prediction model. However, since the data of the second preset frequency band is based on the prediction result of the first preset frequency band, there may be some loss. In order to obtain speech data with better listening experience, the enhanced speech data is input into the preset repair prediction model for prediction to obtain the repair result of the enhanced speech data, that is, the repaired speech data.

[0132] Exemplarily, the repair prediction model can be a pre-trained neural network model; the missing parts between the enhanced speech data obtained based on the decoding prediction model and the speech data collected by the real microphone are repaired, which can be the repair of the imaginary part and the real part of the speech data, wherein the imaginary part can be the amplitude and the real part can be the phase angle value; so as to improve the intelligibility and restoration of the speech data, thereby further realizing speech enhancement of the speech data collected by the bone conduction sensor to improve the speech quality.

[0133] In this embodiment, the enhanced speech data obtained based on the decoding prediction model is input into the repair prediction model for prediction to achieve repair of the possible missing parts of the enhanced speech data, thereby further achieving speech enhancement on the speech data collected by the bone conduction sensor to improve the speech quality.

[0134] In one feasible implementation, before the step S40 of inputting the enhanced speech data into a preset restoration prediction model for prediction to obtain the restored speech data, the method further includes:

[0135] Step B10: obtaining repair training data obtained based on the decoding prediction model, and obtaining repair labels for the repair training data;

[0136] Acquire repair training data for training a repair prediction model, wherein the repair training data is speech data of the full frequency band obtained by splicing the decoded speech data in the second preset frequency band obtained by the decoding prediction model with the pre-decoded speech data in the first preset frequency band (hereinafter referred to as repair training data); and simultaneously obtain repair labels for the repair training data, wherein the repair labels are the correct outputs or categories of the repair training data instances of each single frame.

[0137] In one feasible implementation, step B10, the step of obtaining repair training data based on the decoding prediction model, includes:

[0138] Step B101: inputting the first training speech data into a decoding prediction model for prediction to obtain decoding training data within a second preset frequency band;

[0139] The first training speech data in the first preset frequency band among the initial training speech data collected by a microphone in a quiet environment is input into a pre-trained decoding prediction model for prediction to obtain decoding prediction training data in a second frequency band.

[0140] In one feasible implementation, the speech data in the first preset frequency band in the bone conduction training speech data collected by the bone conduction sensor in a quiet environment is input into a pre-trained decoding prediction model for prediction to obtain decoding prediction training data in the second frequency band.

[0141] Step B102: concatenate the decoded training data with the first training speech data to obtain repaired training data.

[0142] The obtained decoding prediction training data and the first training speech data of the person in the first frequency band are spliced ​​to obtain speech data of the full frequency band, that is, the repaired training data.

[0143] In this embodiment, the repair prediction model further repairs the enhanced speech data obtained based on the decoding prediction model. Therefore, the decoding training data obtained by decoding the preset model and the first training speech data are spliced ​​to directly obtain the repair training data required for training the repair prediction model. No manual operation is required. When the data volume of the training data set is large, the construction efficiency of the training data set is greatly improved, thereby improving the training efficiency of the repair prediction model.

[0144] In another feasible implementation, step B10, the step of obtaining the repair label of the repair training data includes:

[0145] Step B103: Using the initial training speech data as repair labels for repairing training data.

[0146] In this embodiment, since the repair prediction model repairs the imaginary part and real part of the enhanced speech data obtained based on the decoding prediction model to improve the intelligibility and restoration of the speech data, where the imaginary part can be the amplitude and the real part can be the phase angle value; therefore, the initial training speech data collected by the microphone in a quiet environment can be directly used as the repair label; the repair label can be directly obtained from the initial training speech data collected by the microphone without manual operation. When the data volume of the training data set is large, the efficiency of constructing the training data set is greatly improved, thereby improving the training efficiency of the repair prediction model.

[0147] Step B20: taking the repair training data and the repair label as one piece of training data, and obtaining a second training data set based on each piece of acquired training data;

[0148] The repair training data and the repair label are taken as one piece of training data, and the training data are combined to obtain a second training data set, namely, a training set of the repair prediction model.

[0149] Step B30: Use the second training data set to train the preset repair prediction model to obtain a repair prediction model.

[0150] The training data in the second training data set is used to train the preset repair prediction model to be trained to obtain a repair prediction model.

[0151] In one feasible implementation, the repair training data is input into the repair prediction model to be trained for prediction to obtain predicted repair data; the repair loss value between the predicted repair data and the repair label is determined, and the repair prediction model to be trained is updated according to the repair loss value to obtain a trained repair prediction model; wherein, the predicted repair data and the corresponding repair label can be compared by subtracting the absolute value, or taking the square root, etc., which is not limited in this embodiment.

[0152] In another feasible implementation, the repair prediction model may be subjected to multiple rounds of iterative training. In the first round of training, the initialized repair prediction model is updated, and in subsequent rounds of training, the repair prediction model updated in the previous round of training is used as the basis for updating.

[0153] In another feasible implementation, the repair loss value can be expressed as a model error value, and the model error value is compared with a preset model error threshold. If the model error value is greater than the preset model error threshold, the repair prediction model is updated based on the model error value; if the model error value is less than or equal to the preset model error threshold, the training of the repair prediction model is completed.

[0154] In this embodiment, since the decoding prediction model is the second preset frequency band speech data obtained by predicting the first preset frequency band speech data, there may be some missing compared to the speech data collected by the microphone. Therefore, the repair training data and the repair label are obtained based on the decoding prediction model respectively, and then the repair training data and the repair label are used as a training data, and a second training data set is obtained based on each piece of training data obtained; and then the second training data set is used to train the preset repair prediction model to be trained to obtain a repair prediction model; through the training of the repair prediction model, the credibility of the training results of the repair prediction model is enhanced, and then on the basis of obtaining enhanced speech data based on the decoding prediction model, the voice quality is further improved.

[0155] Furthermore, based on the above first, second and / or third embodiments, a more complete embodiment of the speech enhancement method of the present invention is proposed, referring to Figure 5, first voice data is collected through a bone conduction sensor, voice data in a first preset frequency band from the first voice data is extracted, and input into a pre-trained decoding prediction model for prediction to obtain decoded voice data in a second preset frequency band, the decoded voice data is spliced ​​with the voice data in the first preset frequency band to obtain enhanced voice data of the full frequency band, and in order to further improve the voice data instruction, the enhanced voice data is input into a pre-trained repair prediction model for prediction to improve the intelligibility and restoration of the voice data, and obtain repaired voice data.

[0156] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which a speech enhancement program is stored. When the speech enhancement program is executed by a processor, the steps of the speech enhancement method described below are implemented.

[0157] The various embodiments of the speech enhancement device and the computer-readable storage medium of the present invention may refer to the various embodiments of the speech enhancement method of the present invention, and will not be described in detail here.

[0158] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0159] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0160] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including several qualities for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0161] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A speech enhancement method, characterized in that: The speech enhancement method comprises the following steps: Acquire initial training voice data collected by a microphone in a quiet environment, and extract first training voice data within a first preset frequency band from the initial training voice data; Obtaining a decoding label of the first training speech data; Taking the first training speech data and the decoding label as one piece of training data, and obtaining a first training data set according to each piece of acquired training data; Using the first training data set to train a preset decoding prediction model to be trained to obtain a decoding prediction model; Acquiring first voice data collected by the bone conduction sensor; Inputting speech data in a first preset frequency band of the first speech data into the decoding prediction model for prediction to obtain decoded speech data in a second preset frequency band, wherein the first preset frequency band is smaller than the second preset frequency band; The decoded voice data is concatenated with voice data in the first preset frequency band of the first voice data to obtain enhanced voice data.

2. The speech enhancement method according to claim 1, wherein: The step of obtaining the decoding label of the first training speech data includes: Extracting second training voice data in the second preset frequency band from the initial training voice data; The second training speech data is used as a decoding label of the first training speech data.

3. The speech enhancement method according to any one of claims 1 or 2, characterized in that: After the step of obtaining the enhanced voice data, the method further includes: The enhanced speech data is input into a preset restoration prediction model for prediction to obtain restoration speech data.

4. The speech enhancement method according to claim 3, wherein: Before the step of inputting the enhanced speech data into a preset restoration prediction model for prediction to obtain the restoration speech data, the method further includes: Acquire repair training data obtained based on the decoding prediction model, and acquire repair labels for the repair training data; Taking the repair training data and the repair label as one piece of training data, and obtaining a second training data set according to each piece of acquired training data; The preset repair prediction model to be trained is trained using the second training data set to obtain the repair prediction model.

5. The speech enhancement method according to claim 4, wherein: The step of obtaining repair training data based on the decoding prediction model includes: Inputting the first training speech data into the decoding prediction model for prediction to obtain decoding training data within the second preset frequency band; The decoded training data is concatenated with the first training speech data to obtain the repaired training data.

6. The speech enhancement method according to claim 4, wherein: The step of obtaining the repair label of the repair training data includes: The initial training speech data is used as the repair label of the repair training data.

7. The speech enhancement method according to claim 1 or 2, wherein: The step of splicing the decoded voice data with the voice data in the first preset frequency band in the first voice data to obtain enhanced voice data includes: Performing a time-domain to frequency-domain conversion on the voice data in the first preset frequency band in the first voice data to obtain a first frequency spectrum; Inputting the first spectrum into the decoding prediction model for prediction to obtain a second spectrum; The second spectrum is converted from the frequency domain to the time domain to obtain the decoded speech data.

8. A speech enhancement device, characterized in that: The speech enhancement device includes: a memory, a processor, and a speech enhancement program stored in the memory and executable on the processor. When the speech enhancement program is executed by the processor, the steps of the speech enhancement method according to any one of claims 1 to 7 are implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a speech enhancement program, which, when executed by a processor, implements the steps of the speech enhancement method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Voice noise reduction method, device and equipment and computer readable storage medium

    CN115171713A

  • Voice noise reduction method, device and equipment and computer readable storage medium

    CN115631760A