Voice noise reduction method, device and computer readable storage medium

By combining bone conduction sensors and microphones to collect signals, and utilizing a pre-trained speech denoising model, the problem of weak noise resistance of microphone-collected signals is solved, achieving deep noise reduction and sound quality improvement.

CN116403593BActive Publication Date: 2025-11-07GOERTEK INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310491956.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-26
Publication Date
2025-11-07
Estimated Expiration
2043-04-26

AI Technical Summary

Technical Problem

In existing technologies, the noise resistance of the voice signal collected by the microphone is weak, resulting in poor voice noise reduction effect, which affects the quality of the voice signal and the success rate of recognition.

Method used

By combining signals collected by bone conduction sensors and microphones, noise signals in the microphone signals are determined through bone conduction signals, and noise reduction is performed using a pre-trained speech denoising model. The noise reduction gain is then calculated to achieve deep noise reduction.

Benefits of technology

It improves the noise reduction effect and sound quality of voice signals, thus enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116403593B_ABST
    Figure CN116403593B_ABST
Patent Text Reader

Abstract

The application discloses a voice noise reduction method and device and a computer readable storage medium. The method comprises the following steps: acquiring a microphone signal collected by a microphone and acquiring a bone conduction signal collected synchronously by a bone conduction sensor; determining a microphone noise signal in the microphone signal according to the bone conduction signal, wherein the microphone noise signal is a noise signal in the microphone signal; inputting the microphone signal and the microphone noise signal into a preset voice noise reduction model for prediction to obtain a noise reduction gain of the microphone signal; and performing noise reduction processing on the microphone signal according to the noise reduction gain to obtain a noise reduction voice signal. The application realizes a voice noise reduction scheme combining the signal collected by the bone conduction sensor and the signal collected by the microphone, and improves the voice noise reduction effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of audio processing, in particular to a speech noise reduction method, device and computer readable storage medium. BACKGROUND

[0002] With the continuous development of the communication industry, the speech function of mobile terminals such as mobile phones is widely used in voice call, speech recognition and other scenarios. There will inevitably be background noise in the speech signal input by the user to the mobile terminal, and the existence of background noise seriously affects the quality of the speech signal and reduces the success rate and accuracy of speech recognition on the speech signal.

[0003] In conventional technology, the signal collected by the microphone is usually subjected to noise reduction processing, and the useful and clean speech signal is extracted from the noisy speech signal as much as possible, and the noise interference is suppressed or reduced; but the speech signal collected by the microphone has weak noise resistance, and the conventional speech noise reduction method has poor noise reduction effect on the speech signal collected by the microphone.

[0004] The above content is only used to assist in understanding the technical solutions of the present application and does not represent the acknowledgement of the above content as prior art. SUMMARY

[0005] The main purpose of the present application is to provide a speech noise reduction method, device and computer readable storage medium, aiming to provide a scheme of combining the signals collected by the bone conduction sensor and the signals collected by the microphone for speech noise reduction, so as to improve the speech noise reduction effect.

[0006] To achieve the above purpose, the present application provides a speech noise reduction method, which comprises the following steps:

[0007] Obtaining a microphone signal collected by a microphone, and obtaining a bone conduction signal synchronously collected by a bone conduction sensor;

[0008] According to the bone conduction signal, determining a microphone noise signal in the microphone signal, the microphone noise signal being a noise signal in the microphone signal;

[0009] Inputting the microphone signal and the microphone noise signal into a preset speech noise reduction model for prediction to obtain a noise reduction gain of the microphone signal;

[0010] According to the noise reduction gain, performing noise reduction processing on the microphone signal to obtain a noise reduction speech signal.

[0011] Optionally, the step of determining the microphone noise signal in the microphone signal according to the bone conduction signal comprises:

[0012] obtaining a first speech presence probability of the bone conduction signal, and taking the first speech presence probability as a second speech presence probability of the microphone signal;

[0013] determining a microphone noise signal in the microphone signal according to the second speech presence probability.

[0014] Optionally, the step of obtaining the first speech presence probability of the bone conduction signal comprises:

[0015] judging whether an energy of the bone conduction signal is greater than a preset speech energy threshold;

[0016] if yes, determining that the first speech presence probability is a first preset probability;

[0017] if no, determining that the first speech presence probability is a second preset probability, wherein the first preset probability is greater than the second preset probability.

[0018] Optionally, the step of determining the microphone noise signal in the microphone signal according to the second speech presence probability comprises:

[0019] obtaining a gradient of a first noisy signal in the microphone signal, wherein the first noisy signal is a signal in the microphone signal with the second speech presence probability being the first preset probability;

[0020] judging whether the gradient is less than a preset noise gradient threshold;

[0021] if yes, determining that the first noisy signal is the microphone noise signal.

[0022] Optionally, the step of inputting the microphone signal and the microphone noise signal into a preset speech noise reduction model for prediction comprises:

[0023] replacing a second noisy signal in the microphone signal with a preset background sound signal to generate a microphone pre-noise reduction signal, wherein the second noisy signal is a signal in the microphone signal with the second speech presence probability being the second preset probability;

[0024] inputting the microphone pre-noise reduction signal and the microphone noise signal into the preset speech noise reduction model for prediction.

[0025] Optionally, before the step of inputting the microphone signal and the microphone noise signal into the preset speech noise reduction model for prediction, the method further comprises:

[0026] obtaining a speech training signal collected by a microphone in a quiet environment, and obtaining a preset noise training signal;

[0027] add the noise training signal to the speech training signal according to a preset signal-to-noise ratio to obtain a mixed training signal;

[0028] obtain a noise reduction label of the mixed training signal;

[0029] use the mixed training signal, the noise training signal and the noise reduction label as a training signal, and obtain a noise reduction training signal set according to the obtained training signals;

[0030] train a preset to-be-trained speech noise reduction model using the noise reduction training signal set to obtain the speech noise reduction model.

[0031] Optionally, the step of obtaining the noise reduction label of the mixed training signal comprises:

[0032] determine an energy relationship between the speech training signal and the mixed training signal;

[0033] use the energy relationship as the noise reduction label.

[0034] Optionally, the microphone is a microphone array comprising a plurality of array elements, and the step of obtaining the microphone signal collected by the microphone comprises:

[0035] obtain a plurality of signals collected by the microphone array, and extract a signal feature of the plurality of signals;

[0036] perform beamforming processing on the plurality of signals according to the signal feature to obtain a single-channel noise reduction signal, and use the single-channel noise reduction signal as the microphone signal.

[0037] To achieve the above object, the present application further provides a speech noise reduction device, which comprises a memory, a processor and a speech noise reduction program stored in the memory and executable on the processor, and the speech noise reduction program implements the steps of the speech noise reduction method when executed by the processor.

[0038] In addition, to achieve the above object, the present application further provides a computer readable storage medium, which stores a speech noise reduction program, and the speech noise reduction program implements the steps of the speech noise reduction method when executed by a processor.

[0039] In the present application, the microphone signal collected by the microphone and the bone conduction signal collected by the bone conduction sensor are acquired; then the microphone noise signal in the microphone signal is determined according to the bone conduction signal; the noise existing in the synchronously collected microphone signal is accurately analyzed by using the excellent noise resistance of the bone conduction signal collected by the bone conduction sensor; then the microphone signal and the microphone noise signal are input into the preset voice noise reduction model for prediction to obtain the noise reduction gain of the microphone signal; the part of the microphone signal with noise is accurately located on the basis of the microphone noise signal by using the pre-trained voice noise reduction model, so that the noise reduction gain of the microphone signal is calculated to obtain the ideal noise reduction gain of the microphone signal; then the microphone signal is processed by noise reduction according to the noise reduction gain to obtain a noise reduction voice signal; on the basis of retaining the complete frequency domain signal collected by the microphone, the microphone signal is deeply processed to enhance the noise reduction effect and sound quality of the voice signal, and then the user experience is improved. BRIEF DESCRIPTION OF DRAWINGS

[0040] Figure 1 The present application relates to a hardware running environment structure diagram of an embodiment scheme.

[0041] Figure 2 The present application relates to a hardware running environment structure diagram of an embodiment scheme.

[0042] Figure 3 The present application relates to a hardware running environment structure diagram of an embodiment scheme.

[0043] Figure 4 The present application relates to a hardware running environment structure diagram of an embodiment scheme.

[0044] Figure 5 The present application relates to a hardware running environment structure diagram of an embodiment scheme.

[0045] The present application relates to a hardware running environment structure diagram of an embodiment scheme. DETAILED DESCRIPTION

[0046] It should be understood that the specific embodiments described herein are merely intended to explain the present application and are not intended to limit the present application.

[0047] As shown in Figure 1 , the present application relates to a hardware running environment structure diagram of an embodiment scheme. Figure 1 It should be understood that the specific embodiments described herein are merely intended to explain the present application and are not intended to limit the present application.

[0048] It should be understood that the specific embodiments described herein are merely intended to explain the present application and are not intended to limit the present application. It should be understood that the specific embodiments described herein are merely intended to explain the present application and are not intended to limit the present application.

[0049] As Figure 1 shown, the voice noise reduction device can include a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. The communication bus 1002 is used to realize the connection communication between the components. The user interface 1003 can include a display screen (Display), an input unit such as a keyboard (Keyboard), and the optional user interface 1003 can also include a standard wired interface, a wireless interface. The network interface 1004 can optionally include a standard wired interface, a wireless interface (such as a WI-FI interface). The memory 1005 can be a high-speed RAM memory, or a stable memory (non-volatile memory) such as a disk memory. The memory 1005 can also be an independent storage device from the aforementioned processor 1001.

[0050] Those skilled in the art can understand that Figure 1 the device structure shown in the foregoing embodiments does not constitute a limitation on the voice noise reduction device, and can include more or fewer components than those shown, or combine certain components, or different component arrangements.

[0051] As Figure 1 shown, the memory 1005 as a computer storage medium can include an operating system, a network communication module, a user interface module, and a voice noise reduction program. The operating system is a program that manages and controls device hardware and software resources, supports the running of the voice noise reduction program and other software or programs. In Figure 1 the device shown, the user interface 1003 is mainly used for signal communication with the client; the network interface 1004 is mainly used for establishing a communication connection with the server; and the processor 1001 can be used to call the voice noise reduction program stored in the memory 1005 and perform the following operations:

[0052] obtain a microphone signal collected through a microphone, and obtain a bone conduction signal collected synchronously through a bone conduction sensor;

[0053] According to the bone conduction signal, determine a microphone noise signal in the microphone signal, the microphone noise signal being a noise signal in the microphone signal;

[0054] input the microphone signal and the microphone noise signal into a preset voice noise reduction model for prediction to obtain a noise reduction gain of the microphone signal;

[0055] According to the noise reduction gain, perform noise reduction processing on the microphone signal to obtain a noise reduction voice signal.

[0056] Further, the operation of determining the microphone noise signal in the microphone signal according to the bone conduction signal comprises:

[0057] obtaining a first speech presence probability of the bone conduction signal, and taking the first speech presence probability as a second speech presence probability of the microphone signal;

[0058] determining the microphone noise signal in the microphone signal according to the second speech presence probability.

[0059] Further, the operation of obtaining the first speech presence probability of the bone conduction signal comprises:

[0060] judging whether the energy of the bone conduction signal is greater than a preset speech energy threshold;

[0061] if yes, determining that the first speech presence probability is a first preset probability;

[0062] if no, determining that the first speech presence probability is a second preset probability, wherein the first preset probability is greater than the second preset probability.

[0063] Further, the operation of determining the microphone noise signal in the microphone signal according to the second speech presence probability comprises:

[0064] obtaining a gradient of a first noise signal in the microphone signal, wherein the first noise signal is a signal in the microphone signal with the second speech presence probability being the first preset probability;

[0065] judging whether the gradient is less than a preset noise gradient threshold;

[0066] if yes, determining that the first noise signal is the microphone noise signal.

[0067] Further, the operation of inputting the microphone signal and the microphone noise signal into a preset speech noise reduction model for prediction comprises:

[0068] replacing a second noise signal in the microphone signal with a preset background sound signal to generate a microphone pre-noise reduction signal, wherein the second noise signal is a signal in the microphone signal with the second speech presence probability being the second preset probability;

[0069] inputting the microphone pre-noise reduction signal and the microphone noise signal into a preset speech noise reduction model for prediction.

[0070] Further, before the operation of inputting the microphone signal and the microphone noise signal into a preset voice noise reduction model for prediction, the processor 1001 can also be configured to invoke a voice noise reduction program stored in the memory 1005 to perform the following operations:

[0071] obtain a voice training signal collected by a microphone in a quiet environment, and obtain a preset noise training signal;

[0072] add the noise training signal to the voice training signal according to a preset signal-to-noise ratio to obtain a mixed training signal;

[0073] obtain a noise reduction label of the mixed training signal;

[0074] use the mixed training signal, the noise training signal, and the noise reduction label as a training signal, and obtain a noise reduction training signal set according to each training signal obtained;

[0075] train a preset voice noise reduction model to be trained using the noise reduction training signal set to obtain the voice noise reduction model.

[0076] Further, the operation of obtaining the noise reduction label of the mixed training signal comprises:

[0077] determine an energy relationship between the voice training signal and the mixed training signal;

[0078] use the energy relationship as the noise reduction label.

[0079] Further, the microphone is a microphone array comprising a plurality of array elements, and the operation of obtaining the microphone signal collected by the microphone comprises:

[0080] obtain a plurality of signals collected by the microphone array, and extract a signal feature of the plurality of signals;

[0081] perform beamforming processing on the plurality of signals according to the signal feature to obtain a single-channel noise reduction signal, and use the single-channel noise reduction signal as the microphone signal.

[0082] Based on the above structure, various embodiments of the voice noise reduction method are proposed.

[0083] Reference Figure 2 , Figure 2 is a flowchart of the first embodiment of the voice noise reduction method of the present application.

[0084] Embodiments of the present application provide a voice noise reduction method. It should be noted that although a logical sequence is shown in the flowchart, in some cases, the steps shown or described can be performed in an order different from that shown. In the present embodiment, the subject performing the voice noise reduction method can be a headset, a virtual display device, a personal computer, a smartphone, or the like. In the present embodiment, the subject performing the voice noise reduction method is not limited. The following is a description of the embodiments for the convenience of description. In the present embodiment, the voice noise reduction method comprises:

[0085] In step S10, a microphone signal collected by a microphone is obtained, and a bone conduction signal collected synchronously by a bone conduction sensor is obtained.

[0086] The microphone signal can be collected by the microphone, and the bone conduction signal can be collected synchronously by the bone conduction sensor.

[0087] In an embodiment, the microphone and the bone conduction sensor can be arranged in a product for collecting a voice signal. For example, the bone conduction sensor can be arranged in a headset or a virtual reality device. The specific arrangement position can be designed as required, for example, the bone conduction sensor is usually arranged at a position in contact with the human skull.

[0088] In another embodiment, the microphone signal and the bone conduction signal can be real-time collected voice signals, or can be non-real-time recorded voice signals. The specific embodiment can be selected according to the real-time requirement of the voice noise reduction in the application scenario. The present embodiment does not limit this.

[0089] In yet another embodiment, the step S10 of obtaining the microphone signal collected by the microphone comprises:

[0090] In step S11, a plurality of signals collected by a microphone array are obtained, and signal characteristics of the plurality of signals are extracted.

[0091] The microphone can be a microphone array comprising a plurality of elements. The microphone array collects a plurality of channels of voice signals (hereinafter referred to as a plurality of signals for the sake of distinction), and then extracts signal characteristics in the plurality of signals, wherein the signal characteristics can be one or more of phase difference, beam, sound area, power spectrum, etc.

[0092] In step S12, the plurality of signals are subjected to beamforming processing according to the signal characteristics, a single-channel noise reduction signal is obtained, and the single-channel noise reduction signal is taken as the microphone signal.

[0093] According to the signal characteristics of the acquired multi-path signals, beamforming processing is performed on the multi-path signals to eliminate interference signals, such as steady-state noise signals, in the collected multi-path signals, thereby combining the multi-path signals to generate a single-path noise-reduced signal (hereinafter referred to as a single-path noise-reduced signal for distinction).

[0094] For example, beamforming refers to a technique of eliminating interference signals by weighting and combining the signals collected by a multi-element array (microphone array) to form a desired single-path ideal signal, i.e., a single-path noise-reduced signal.

[0095] In an implementation, after obtaining the multi-path signals, time-frequency transformation, such as Fourier transform or short-time Fourier transform, is performed to obtain the multi-path signals in the frequency domain, and then a plurality of fixed beams in different directions are generated by using the amplitude and phase differences between the signals of the channels (paths) in the multi-path signals, and the correlation (commonality and difference) between the beams is determined, the correlation of each beam is accumulated as the importance of the beam, and then the beams are weighted and summed by using the importance as the weight, and then the target beam in the beams is determined, and the interference signals of the other beams are eliminated based on the target beam to obtain a single-path noise-reduced signal with interference eliminated, and the single-path noise-reduced signal is used as the microphone signal.

[0096] In the embodiment, the multi-path signals collected by the microphone array are used to perform preliminary noise reduction processing on the microphone signals to eliminate interference signals in the microphone signals and improve the noise reduction effect and sound quality of the voice signals.

[0097] In step S20, the microphone noise signal in the microphone signal is determined based on the bone conduction signal, and the microphone noise signal is the noise signal in the microphone signal.

[0098] Since the bone conduction signal collected by the bone conduction sensor is the signal generated by the skull vibration when a person speaks, the interference of external noise on the bone conduction signal is minimal, and therefore, by comparing the bone conduction signal collected by the bone conduction sensor with the microphone signal synchronously collected by the microphone, the noise signal (hereinafter referred to as a microphone noise signal for distinction) existing in the microphone signal can be quickly and accurately determined. For example, the frequency band of the noise signal in the microphone signal is determined based on the bone conduction signal, the spectrum of the frequency band is extracted, and a noisy spectrum is generated, and the noisy spectrum is used as the microphone noise signal.

[0099] For example, the microphone noise signal is a steady-state noise signal, a non-steady-state noise signal, or a signal containing steady-state and non-steady-state noise.

[0100] The frequency band refers to a frequency range, and a frequency range includes a plurality of frequency points.

[0101] Step S30, input the microphone signal and the microphone noise signal to a preset voice noise reduction model for prediction to obtain a noise reduction gain of the microphone signal;

[0102] The microphone signal and the microphone noise signal are input to the pre-trained voice noise reduction model for prediction. Based on the correlation between the energy of the mixed signal containing noise and the corresponding noise-free voice signal collected synchronously, and the powerful adaptive computing capability of the voice noise reduction model, the noise reduction gain of the noise part in the microphone signal is accurately calculated. For example, the noise reduction gain is the gain value of the noise-containing voice signal converted to the noise-free voice signal with the same voice content.

[0103] For example, the voice noise reduction model can be a pre-trained neural network model. The noise part in the corresponding microphone signal is marked by the microphone noise signal, and then the noise reduction gain of the noise part in the mixed signal is predicted based on the correlation between the energy of the mixed voice signal containing noise and the corresponding noise-free clean voice signal, so as to perform voice noise reduction on the microphone signal collected by the microphone, for example, non-steady noise reduction and / or steady noise reduction, to improve the voice quality. The specific structure of the decoding prediction model is not limited in this embodiment, for example, a convolutional neural network or a recurrent neural network can be used to realize the network structure.

[0104] For example, the voice noise reduction model can be a pre-trained neural network model; the microphone signal and the microphone noise signal are input, and the noise reduction gain of the microphone signal is output. The frequency band containing steady noise in the corresponding microphone signal containing noise is marked by the microphone noise signal, and the remaining frequency band is the frequency band that may contain non-steady noise. Then, based on the correlation between the energy of the microphone signal containing noise and the corresponding noise-free voice signal, the noise reduction gain of the frequency band containing non-steady noise and steady noise in the microphone signal is predicted, so as to perform voice noise reduction on the microphone signal collected by the microphone.

[0105] In the specific embodiment, the voice signal is synchronously collected by the microphone and the bone conduction sensor arranged in the earphone device, that is, recording is performed to obtain the microphone signal and the bone conduction signal, and then the noise signal existing in the microphone signal, that is, the microphone noise signal, is determined by using the excellent anti-noise capability of the bone conduction signal; and then the microphone signal and the microphone noise signal are input into the preselected and trained voice noise reduction model, the frequency band of the signal with noise existing in the microphone signal is determined by using the microphone noise signal, and the corresponding noise reduction gain value is obtained, and then the microphone signal is subjected to noise reduction processing according to the obtained noise reduction gain value, and a high-quality voice noise reduction signal is obtained, and when the user uses or plays the signal corresponding to the voice noise reduction signal next time, a high-quality recording can be obtained.

[0106] In step S40, the microphone signal is subjected to noise reduction processing according to the noise reduction gain, and a noise reduction voice signal is obtained.

[0107] According to the obtained noise reduction gain, the microphone signal is subjected to noise reduction processing, and a noise reduction voice signal is obtained.

[0108] For example, the noise reduction gain is the noise reduction gain value corresponding to each frequency band or each frequency point of the microphone signal.

[0109] In an available embodiment, the noise reduction gain obtained according to the voice noise reduction model is the noise reduction gain of each frequency band, and the noise reduction gain of each frequency point is calculated by interpolation method, the noise reduction gain of each frequency point is multiplied by the frequency point, the amplitude spectrum of the microphone signal after noise reduction is obtained, and the amplitude spectrum is converted from the frequency domain to the time domain to obtain the noise reduction voice signal.

[0110] In the embodiment, the microphone signal collected by the microphone and the bone conduction signal synchronously collected by the bone conduction sensor are obtained; then the microphone noise signal in the microphone signal is determined according to the bone conduction signal; the signal with noise existing in the synchronously collected microphone signal is accurately analyzed by using the excellent anti-noise capability of the bone conduction signal collected by the bone conduction sensor; then the microphone signal and the microphone noise signal are input into the preselected voice noise reduction model for prediction to obtain the noise reduction gain of the microphone signal; the energy correlation between the microphone signal with noise and the corresponding noise-free voice signal is used to accurately locate the part with noise in the microphone signal and calculate the noise reduction gain of the microphone signal by using the powerful adaptive calculation capability of the pre-trained voice noise reduction model on the basis of the microphone noise signal, so that the ideal noise reduction gain of the microphone signal is obtained, and then the microphone signal is subjected to noise reduction processing according to the noise reduction gain to obtain a noise reduction voice signal; on the basis of retaining the complete frequency domain signal collected by the microphone, deep noise reduction of the microphone signal is realized, the noise reduction effect and the sound quality of the voice signal are enhanced, and then the user experience is improved.

[0111] Further, based on the first embodiment, the second embodiment of the voice noise reduction method of the present application is proposed, referring to Figure 3 In the present embodiment, step S20 comprises:

[0112] Step S21, obtaining the first voice presence probability of the bone conduction signal, and taking the first voice presence probability as the second voice presence probability of the microphone signal;

[0113] Since the voice signal collected by the bone conduction sensor has excellent anti-noise capability, the signal almost does not contain noise, and thus can accurately reflect whether the user is currently speaking; therefore, the voice presence probability of the bone conduction signal collected synchronously with the microphone signal (the first voice presence probability) is obtained, and then the first voice presence probability is taken as the voice presence probability of the microphone signal (the second voice presence probability); wherein the voice presence probability can be the voice presence probability of each frequency band, frequency point, time, or time period.

[0114] For example, the voice presence probability can contain multiple preset probability values, such as a first preset probability of 20%, a second preset probability of 40%, a third preset probability of 60%, etc.; or contain two preset probability values, one corresponding to the presence of voice signal, i.e. 1, and the other corresponding to the absence of voice signal, i.e. 0, which is not limited in the present embodiment.

[0115] In the specific embodiment, when it is determined that the bone conduction signal of the current time (time domain) and / or the current frequency band (frequency domain) contains voice content, since the microphone signal and the bone conduction signal are synchronously collected signals, the microphone signal of the corresponding time and / or frequency band also contains voice content, and thus the voice presence probability of the microphone signal is determined.

[0116] In a feasible embodiment, referring to Figure 4 Step S21, the step of obtaining the first voice presence probability of the bone conduction signal comprises:

[0117] Step S211, determining whether the energy of the bone conduction signal is greater than a preset voice energy threshold;

[0118] The energy of the bone conduction signal is calculated, and it is determined whether the energy of the bone conduction signal is greater than a preset voice energy threshold; since the signal collected by the bone conduction sensor has excellent anti-noise capability, the signal almost does not contain noise, and thus according to the energy of the voice signal collected by the bone conduction sensor, by setting a preset voice energy threshold, the voice presence probability in the bone conduction signal can be accurately determined, and thus the voice presence probability in the synchronously collected microphone signal is determined.

[0119] Step S212, if yes, determining that the first voice presence probability is a first preset probability;

[0120] If the energy of the bone conduction signal is greater than the preset voice energy threshold, it indicates that the bone conduction signal contains voice signals, and the voice existing probability of the bone conduction signal (first voice existing probability) is determined as a first preset probability; for example, the first preset probability is 1.

[0121] If not, the first voice existing probability is determined as a second preset probability in step S213, where the first preset probability is greater than the second preset probability.

[0122] If the energy of the bone conduction signal is less than or equal to the preset voice energy threshold, it indicates that the bone conduction signal does not contain voice signals, and the first voice existing probability of the bone conduction signal is determined as the second preset probability, where the first preset probability is greater than the second preset probability; for example, the second preset probability is 0.

[0123] In this embodiment, since the voice signal collected by the bone conduction sensor has excellent anti-noise capability and almost no noise in the signal, the voice existing probability in the bone conduction signal can be accurately determined by setting a preset voice energy threshold according to the energy of the voice signal collected by the bone conduction sensor, and the voice existing probability in the microphone signal is determined, and the noise part in the microphone signal is accurately determined.

[0124] In step S22, the microphone noise signal in the microphone signal is determined according to the second voice existing probability.

[0125] According to the second voice existing probability of the microphone signal, the noise part in the microphone signal, i.e. the microphone noise signal, is estimated, which can be a steady-state noise signal, a non-steady-state noise signal or a signal containing steady-state and non-steady-state noise, where the microphone noise signal can be the spectrum of any frequency band or frequency point in the microphone signal, which is not limited in this embodiment.

[0126] In this embodiment, since the voice signal collected by the bone conduction sensor has excellent anti-noise capability and almost no noise in the signal, the voice existing probability in the bone conduction signal can be accurately determined by setting a preset voice energy threshold according to the energy of the voice signal collected by the bone conduction sensor, and the voice existing probability in the microphone signal is determined, and the noise part in the microphone signal is accurately determined.

[0127] In a feasible implementation, step S20 includes:

[0128] In step S23, a gradient of the first noisy signal in the microphone signal is obtained, wherein the first noisy signal is a signal in the microphone signal with a second speech existence probability being a first preset probability.

[0129] A gradient of a signal in the microphone signal with a speech existence probability being a first preset probability is obtained. For example, a signal corresponding to a first frequency band in the microphone signal has a speech existence probability being the first preset probability, i.e., the signal corresponding to the first frequency band is the first noisy signal, and the first frequency band has a speech signal. Then, the gradient of the first noisy signal is obtained. The gradient can be a rate of change of energy or amplitude of a signal in any frequency band.

[0130] In step S24, it is determined whether the gradient is less than a preset noise gradient threshold.

[0131] It is determined whether the gradient of the first noisy signal is less than the preset noise gradient threshold. Since the energy change rate of the stationary noise is relatively stable, if the gradient is less than the preset noise gradient threshold, it is indicated that the noise existing in the first noisy signal is the stationary noise. If the gradient is greater than or equal to the preset noise gradient threshold, it is indicated that the first noisy signal has no noise or has non-stationary noise.

[0132] In step S25, if yes, it is determined that the first noisy signal is the microphone noise signal.

[0133] If the gradient is less than the preset noise gradient threshold, it is indicated that the energy change rate of the first noisy signal is relatively stable, i.e., the noise existing in the first noisy signal is the stationary noise. Then, it is determined that the first noisy signal is the microphone noise signal, wherein the noise existing in the microphone noise signal is the stationary noise.

[0134] In the embodiment, the gradient of the first noisy signal in the microphone signal is obtained, wherein the first noisy signal is a signal in the microphone signal with a second speech existence probability being a first preset probability. Then, it is determined whether the gradient is less than a preset noise gradient threshold. Since the energy change rate of the stationary noise is relatively stable, if the gradient is less than the preset noise gradient threshold, it is indicated that the energy change of the first noisy signal is relatively stable. Then, it is determined that the noise existing in the first noisy signal is the stationary noise, and the first noisy signal is regarded as the microphone noise signal. Thus, the microphone noise signal existing in the microphone signal is quickly and accurately determined.

[0135] In an implementation, in step S30, the microphone signal and the microphone noise signal are input into a preset speech noise reduction model for prediction, and the step of generating the microphone pre-noise reduction signal by replacing the second noisy signal in the microphone signal with the preset background sound signal includes:

[0136] In step S31, the second noisy signal in the microphone signal is replaced with the preset background sound signal to generate the microphone pre-noise reduction signal, wherein the second noisy signal is a signal in the microphone signal with a second speech existence probability being a second preset probability.

[0137] The second noisy signal, which has a second preset probability of presence in the microphone signal, does not contain any speech signal, meaning the user is not speaking; therefore, the second noisy signal is determined to be pure noise. If the second noisy signal is directly deleted, the sudden appearance of speech when someone speaks again will feel abrupt, resulting in poor listening quality of the speech signal. Therefore, the second noisy signal is replaced with a preset background sound signal to generate a microphone pre-noise reduction signal, thereby reducing the impact of noise on the speech quality of the speech signal. The background color noise signal can be a weak, comfortable noise.

[0138] Step S32: Input the microphone pre-denoising signal and the microphone noise signal into the preset speech denoising model for prediction.

[0139] The microphone signal is replaced by the microphone pre-denoising signal to achieve preliminary noise reduction of the pure noise part of the microphone signal. Then, the microphone pre-denoising signal and the microphone noise signal are input into the preset speech noise reduction model for prediction.

[0140] In this embodiment, if the probability of the second speech is the second preset probability, it indicates that there is no speech signal in the corresponding signal, that is, the corresponding signal is a pure noise signal. If the second noisy signal is directly deleted, when someone speaks again, the sudden appearance of the speech will feel abrupt, resulting in a poor listening experience of the speech signal. Therefore, the second noisy signal is replaced with a preset background sound signal to generate a microphone pre-denoising signal, and the microphone pre-denoising signal is used as the input of the speech denoising model for prediction. This reduces the impact of the pure noise part on the speech quality of the speech signal and improves the speech denoising efficiency.

[0141] Furthermore, based on the first and / or second embodiments described above, a third embodiment of the speech noise reduction method of the present invention is proposed, referring to... Figure 5 In this embodiment, step S30 includes:

[0142] Step A10: Obtain the voice training signal collected by the microphone in a quiet environment, and obtain the preset noise training signal;

[0143] In order to train the speech denoising model so that it can clearly obtain the energy relationship between the mixed signal containing noise and the corresponding quiet signal without noise, a quiet speech signal (hereinafter referred to as the speech training signal) and a preset noise signal (hereinafter referred to as the noise training signal) are collected through a microphone in a quiet environment.

[0144] Step A20: Add the noise training signal to the speech training signal according to the preset signal-to-noise ratio to obtain the mixed training signal;

[0145] According to a preset signal-to-noise ratio, the noise training signal is added to the speech training signal to obtain a plurality of mixed training signals, wherein the mixed training signal is a noisy signal under different signal-to-noise ratios.

[0146] In step A30, a noise reduction label of the mixed training signal is obtained.

[0147] The noise reduction label of the mixed training signal is obtained, wherein the noise reduction label is the correct output or category of each mixed training data instance.

[0148] In an implementation, the step of obtaining the noise reduction label of the mixed training signal in step A30 includes:

[0149] In step A31, an energy relationship between the speech training signal and the mixed training signal is determined.

[0150] Since the speech training signal is the noise-free part in the mixed training signal, by determining the energy relationship between the speech training signal and the corresponding mixed training signal, the energy proportion of the noise part in the mixed training signal can be known, and the ideal noise reduction gain when performing noise reduction can be obtained.

[0151] For example, the energy relationship can be an energy ratio between the speech training signal and the mixed training signal, a square root of the energy ratio, etc., which is not limited in the embodiment.

[0152] In step A32, the energy relationship is used as the noise reduction label.

[0153] The obtained energy relationship is used as the noise reduction label of the mixed training signal; since the mixed training signal is obtained by adding noise training signals with different signal-to-noise ratios to the speech training signal collected under quiet conditions, by taking the mixed training signal and the corresponding noise training signal as the input signal, the frequency band containing noise in the mixed training signal can be determined through the frequency band of the noise training signal; and the energy ratio relationship between the mixed training signal and the corresponding speech training signal is used as the noise reduction label to determine the ideal predicted noise reduction gain of the speech noise reduction model, judge the prediction accuracy of the speech noise reduction model, and continuously train and optimize the speech noise reduction model.

[0154] In the embodiment, since the mixed training signal is obtained by adding noise training signals with different signal-to-noise ratios to the speech training signal collected in a quiet condition, the energy relationship between the speech training signal and the mixed training signal is determined, and then the energy relationship is used as the noise reduction label. Through the acquisition of the noise reduction label, the training of the speech noise reduction model is realized, and the noise reduction label can be directly obtained by comparing the energy relationship between the speech training signal and the mixed training signal, without manual operation. When the data amount of the training data set is large, the construction efficiency of the training data set is greatly improved, and then the training efficiency of the speech noise reduction model is improved.

[0155] In step A40, the mixed training signal, the noise training signal and the noise reduction label are used as a training signal, and the noise reduction training signal set is obtained according to each training signal.

[0156] The mixed training signal and the corresponding added noise training signal and the noise reduction label are used as a single training signal, and each training signal is combined to obtain the noise reduction training signal set, that is, the training signal set of the speech noise reduction model.

[0157] In the specific embodiment, the number of samples used for training can be set as needed, which is not limited in the embodiment.

[0158] In step A50, the noise reduction training signal set is used to train the preset to-be-trained speech noise reduction model to obtain the speech noise reduction model.

[0159] The noise reduction training signal set is used to train the preset to-be-trained speech noise reduction model, so that the parameters of the speech noise reduction model are set and updated. After multiple rounds of training, the speech noise reduction model is obtained.

[0160] In an available embodiment, the updated speech noise reduction model is used as the basis for the next round of training. After multiple rounds of training, the updated speech noise reduction model is used as the trained speech noise reduction model.

[0161] In the embodiment, since there is a certain correlation between the energy of the quiet voice signal, the pure noise signal and the mixed signal with noise, the voice training signal collected by the microphone in a quiet environment can be obtained, the preset noise training signal is obtained, then the noise training signal is added to the voice training signal according to the preset signal-to-noise ratio, the mixed training signal is obtained, and the noise reduction label of the mixed training signal is obtained, so that the mixed training signal, the noise training signal and the noise reduction label are used as a training signal, the noise reduction training signal set is obtained according to the obtained each training signal, the preset voice noise reduction model to be trained is trained by using the noise reduction training signal set, and the voice noise reduction model is obtained; through the training of the voice noise reduction model, the credibility of the voice noise reduction model training result is enhanced, and then the voice noise reduction effect of the voice noise reduction model is improved.

[0162] In addition, the embodiment of the present application also provides a computer readable storage medium, wherein the storage medium stores a voice noise reduction program, and the voice noise reduction program is executed by a processor to realize the steps of the voice noise reduction method.

[0163] The embodiments of the voice noise reduction device and the computer readable storage medium can refer to the embodiments of the voice noise reduction method, and details are not described herein.

[0164] It should be noted that in this document, the terms "comprising", "containing", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.

[0165] The above-mentioned embodiment numbers of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0166] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be realized by means of software and necessary general hardware platform, of course, they can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a plurality of instructions for making a terminal device (which can be a mobile phone, computer, server, air conditioner or network device) execute the methods described in various embodiments of the present application.

[0167] The above merely describes the preferred embodiments of the present application, and is not intended to limit the patent scope of the present application, and any equivalent structure or equivalent process conversion, or direct or indirect application in other related technical fields, which are made by using the content of the present application specification and drawings, are also included in the patent protection scope of the present application.

Claims

1. A voice noise reduction method, characterized by, The voice noise reduction method comprises the following steps: obtaining a microphone signal collected by a microphone, and obtaining a bone conduction signal collected synchronously by a bone conduction sensor; obtaining a first voice existence probability of the bone conduction signal, and taking the first voice existence probability as a second voice existence probability of the microphone signal; determining a microphone noise signal in the microphone signal according to the second voice existence probability, the microphone noise signal being a noise signal in the microphone signal; inputting the microphone signal and the microphone noise signal into a preset voice noise reduction model for prediction to obtain a noise reduction gain of the microphone signal; performing noise reduction processing on the microphone signal according to the noise reduction gain to obtain a noise reduction voice signal.

2. The voice noise reduction method of claim 1, wherein, The step of obtaining the first voice existence probability of the bone conduction signal comprises: determining whether the energy of the bone conduction signal is greater than a preset voice energy threshold value; if yes, determining that the first voice existence probability is a first preset probability; if no, determining that the first voice existence probability is a second preset probability, wherein the first preset probability is greater than the second preset probability.

3. The voice noise reduction method of claim 2, wherein, The step of determining the microphone noise signal in the microphone signal according to the second voice existence probability comprises: obtaining a gradient of a first noise signal in the microphone signal, wherein the first noise signal is a signal in the microphone signal with the second voice existence probability being the first preset probability; determining whether the gradient is less than a preset noise gradient threshold value; if yes, determining that the first noise signal is the microphone noise signal.

4. The voice noise reduction method of claim 2, wherein, The step of inputting the microphone signal and the microphone noise signal into a preset voice noise reduction model for prediction comprises: replacing a second noise signal in the microphone signal with a preset background sound signal to generate a microphone pre-noise reduction signal, wherein the second noise signal is a signal in the microphone signal with the second voice existence probability being the second preset probability; inputting the microphone pre-noise reduction signal and the microphone noise signal into a preset voice noise reduction model for prediction.

5. The voice noise reduction method according to any one of claims 1 to 4, wherein, Before the step of inputting the microphone signal and the microphone noise signal into a preset voice noise reduction model for prediction, the method further comprises: obtaining a voice training signal collected by a microphone in a quiet environment, and obtaining a preset noise training signal; adding the noise training signal to the voice training signal according to a preset signal-to-noise ratio to obtain a mixed training signal; obtaining a noise reduction label of the mixed training signal; taking the mixed training signal, the noise training signal and the noise reduction label as a training signal, and obtaining a noise reduction training signal set according to the obtained training signals; training a preset to-be-trained voice noise reduction model using the noise reduction training signal set to obtain the voice noise reduction model.

6. The voice noise reduction method of claim 5, wherein, The step of obtaining the noise reduction label of the mixed training signal comprises: determining an energy relationship between the voice training signal and the mixed training signal; taking the energy relationship as the noise reduction label.

7. The voice noise reduction method according to any one of claims 1 to 4, wherein, The microphone is a microphone array comprising a plurality of array elements, and the step of acquiring the microphone signal collected by the microphone comprises: acquiring a plurality of signals collected by the microphone array, and extracting signal features of the plurality of signals; performing beamforming processing on the plurality of signals according to the signal features to obtain a single-channel noise-reduced signal, and taking the single-channel noise-reduced signal as the microphone signal.

8. A speech noise reduction device, characterized by The voice noise reduction device comprises a memory, a processor, and a voice noise reduction program stored on the memory and executable on the processor, and the voice noise reduction program, when executed by the processor, implements the steps of the voice noise reduction method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a voice noise reduction program, and the voice noise reduction program, when executed by the processor, implements the steps of the voice noise reduction method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Speech enhancement method and device, earphone equipment and computer readable storage medium

    CN114822573A