Speech noise reduction method, device, equipment and computer-readable storage medium

By correcting the frequency of the voice data collected by the microphone using a speech model based on a Gaussian mixture model and bone conduction sensor data, the problem of poor noise reduction in high-noise scenarios in the existing technology is solved, achieving a better noise reduction effect.

CN115602185BActive Publication Date: 2025-09-19GEER TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211212384.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-28
Publication Date
2025-09-19
Estimated Expiration
2042-09-28

AI Technical Summary

Technical Problem

Existing speech noise reduction algorithms have poor noise reduction effects in high-noise scenarios (above 80dB) and are unable to effectively remove non-stationary noise.

Method used

A speech model based on a Gaussian mixture model is used to determine the correction standard data. The frequency data of the preliminary noise reduction data collected by the microphone is corrected. Combined with the bone conduction sensor data, the frequency data is adjusted to fall within the correction range, further optimizing the noise reduction effect.

Benefits of technology

The speech noise reduction effect is improved, especially the noise reduction performance in high noise and non-stationary noise scenarios, and the ability to remove non-stationary noise is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115602185B_ABST
    Figure CN115602185B_ABST
Patent Text Reader

Abstract

The present invention discloses a speech noise reduction method, apparatus, device, and computer-readable storage medium. The method comprises: obtaining preliminary noise reduction data obtained by performing preliminary noise reduction on original speech data collected by a microphone; determining correction standard data corresponding to each first frequency point based on a speech model, wherein the speech model is spectral data determined based on a Gaussian mixture model; comparing the frequency point data of each first frequency point in the preliminary noise reduction data with the correction range defined by the correction standard data of the corresponding frequency point, correcting the frequency point data that exceeds the corresponding correction range to fall within the corresponding correction range, obtaining corrected noise reduction data, and using the corrected noise reduction data as the noise reduction result of the original speech data. The present invention provides a speech noise reduction scheme that further improves the speech noise reduction effect by further correcting the noise reduction speech data after preliminary noise reduction using a speech noise reduction algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech processing technology, and in particular to a speech noise reduction method, apparatus, device, and computer-readable storage medium. Background Art

[0002] Speech noise reduction removes noise from voice data captured by a microphone. It is commonly used to reduce the noise in conversations. Currently, one speech noise reduction algorithm estimates the noise carried in the speech data, extracts this noise estimate from the speech data, and then generates the noise-reduced speech data. This speech noise reduction algorithm is effective for stationary noise, but less effective for non-stationary noise and high-noise scenarios (above 80dB). Summary of the Invention

[0003] The main purpose of the present invention is to provide a speech noise reduction method, device, equipment and computer-readable storage medium, aiming to provide a speech noise reduction solution, by further correcting the noise-reduced speech data after preliminary noise reduction using a speech noise reduction algorithm, so as to further improve the noise reduction effect.

[0004] To achieve the above object, the present invention provides a method for reducing speech noise, which comprises the following steps:

[0005] Obtaining preliminary noise reduction data obtained by performing preliminary noise reduction on original voice data collected by a microphone;

[0006] Determining corrected standard data corresponding to each first frequency point based on a speech model, wherein the speech model is spectral data determined based on a Gaussian mixture model, and the Gaussian mixture model is constructed based on a preset number of sample data, and the sample data is frequency point data of each first frequency point extracted from spectral data obtained by converting speech data collected by a microphone for each user sample;

[0007] The frequency point data of each of the first frequency points in the preliminary noise reduction data is compared with the correction range defined by the correction standard data of the corresponding frequency point, and the frequency point data exceeding the corresponding correction range is corrected to fall within the corresponding correction range to obtain corrected noise reduction data, and the corrected noise reduction data is used as the noise reduction result of the original speech data.

[0008] Optionally, the step of determining the correction standard data corresponding to each frequency point according to the speech model includes:

[0009] Acquire bone conduction voice data collected by a bone conduction sensor, and extract first frequency point data of each second frequency point in spectrum data obtained by converting the bone conduction voice data, wherein the bone conduction voice data is collected synchronously with the original voice data, and each second frequency point is a frequency point of each first frequency point that is below a preset frequency;

[0010] Converting each first frequency point data into second frequency point data according to a preset mapping relationship between each second frequency point bone conduction frequency point data and microphone frequency point data;

[0011] The maximum frequency point data in the second frequency point data corresponding to each second frequency point is multiplied by the normalized frequency point data of each first frequency point in the speech model to obtain the corrected standard data corresponding to each first frequency point, wherein the speech model is to normalize the frequency point data of each first frequency point in the original spectrum data determined based on the Gaussian mixture model to obtain normalized spectrum data, and the normalized spectrum data includes the normalized frequency point data of each first frequency point.

[0012] Optionally, when the correction standard data corresponding to the first frequency point is standard frequency point data corresponding to the first frequency point, the correction range represented by the correction standard data is a range obtained by floating the standard frequency point data up and down by a preset error value;

[0013] The step of comparing the frequency point data of each first frequency point in the preliminary noise reduction data with the correction range defined by the correction standard data of the corresponding frequency point, and correcting the frequency point data that exceeds the corresponding correction range to fall within the corresponding correction range includes:

[0014] Comparing frequency data of a target frequency point in the preliminary noise reduction data with a correction range defined by correction standard data corresponding to the target frequency point, wherein the target frequency point is any one of the first frequency points;

[0015] When the target frequency point data is greater than the upper limit of the correction range or less than the lower limit of the correction range, the target frequency point data is subtracted from the standard frequency point data corresponding to the target frequency point and then multiplied by a first preset correction coefficient. The calculation result is then added to the standard frequency point data corresponding to the target frequency point to obtain the corrected frequency point data corresponding to the target frequency point, wherein the target frequency point data is the frequency point data of the target frequency point in the preliminary noise reduction data.

[0016] Optionally, after the step of comparing the frequency point data of each first frequency point in the preliminary noise reduction data with a correction range defined by the correction standard data of the corresponding frequency point, and correcting the frequency point data that exceeds the corresponding correction range to fall within the corresponding correction range, to obtain the corrected noise reduction data, the method further includes:

[0017] Obtaining a first noise estimate used when performing preliminary noise reduction on the original speech data, the first noise estimate including frequency point noise estimates corresponding to each of the first frequency points;

[0018] When the target frequency data is greater than an upper limit of the correction range, increasing the frequency noise estimate corresponding to the target frequency in the first noise estimation to correct the frequency noise estimate corresponding to the target frequency in the first noise estimation;

[0019] When the target frequency data is less than a lower limit of the correction range, reducing the frequency noise estimate corresponding to the target frequency in the first noise estimation, so as to correct the frequency noise estimate corresponding to the target frequency in the first noise estimation;

[0020] The second noise estimate obtained by correcting the first noise estimate is used as the noise estimate used when performing noise reduction on the next frame of original speech data collected by the microphone.

[0021] Optionally, the step of increasing the frequency noise estimate corresponding to the target frequency in the first noise estimate includes:

[0022] The target frequency data minus the corrected frequency data corresponding to the target frequency is multiplied by a second preset correction coefficient, and the calculated result is added to the frequency noise estimate corresponding to the target frequency to obtain the frequency noise estimate after the target frequency is increased.

[0023] Optionally, before the step of determining the correction standard data corresponding to each first frequency point according to the speech model, the method further includes:

[0024] Acquiring a preset amount of sample data;

[0025] Constructing a preset number of Gaussian mixture models using the sample data, wherein each of the Gaussian mixture models includes a target number of Gaussian mixture components, where the target number is the number of the first frequency points;

[0026] The means of the Gaussian mixture components corresponding to the same frequency point in various Gaussian mixture models are weighted averaged to obtain the weighted average values ​​corresponding to each first frequency point, and the original spectrum data is composed of the weighted average values ​​corresponding to each first frequency point to obtain a speech model.

[0027] Optionally, the step of constructing a Gaussian mixture model of a preset number using the sample data includes:

[0028] Dividing the sample data into a preset number of sample data groups, wherein the preset number of groups is the same as the preset number of species;

[0029] A Gaussian mixture model is constructed using each of the sample data groups to obtain the preset number of Gaussian mixture models.

[0030] To achieve the above object, the present invention further provides a speech noise reduction device, comprising:

[0031] An acquisition module is used to obtain preliminary noise reduction data obtained by performing preliminary noise reduction on the original voice data collected by the microphone;

[0032] a determination module for determining, based on a speech model, corrected standard data corresponding to each first frequency point, wherein the speech model is spectral data determined based on a Gaussian mixture model, the Gaussian mixture model being constructed based on a preset number of sample data, the sample data being frequency data of each first frequency point extracted from spectral data obtained by converting speech data collected by a microphone for each user sample;

[0033] a correction module, configured to compare the frequency point data of each of the first frequency points in the preliminary noise reduction data with a correction range defined by the correction standard data of the corresponding frequency point, correct the frequency point data that exceeds the corresponding correction range to fall within the corresponding correction range, thereby obtaining corrected noise reduction data, and using the corrected noise reduction data as a noise reduction result for the original speech data.

[0034] To achieve the above-mentioned objectives, the present invention further provides a speech noise reduction device, comprising: a memory, a processor, and a speech noise reduction program stored in the memory and executable on the processor, wherein the speech noise reduction program, when executed by the processor, implements the steps of the speech noise reduction method described above.

[0035] In addition, to achieve the above objectives, the present invention also proposes a computer-readable storage medium, on which a speech noise reduction program is stored. When the speech noise reduction program is executed by a processor, the steps of the speech noise reduction method described above are implemented.

[0036] In the present invention, preliminary noise reduction data is obtained by performing preliminary noise reduction on original speech data collected by a microphone; correction standard data corresponding to each first frequency point is determined based on a speech model, wherein the speech model is spectral data determined based on a Gaussian mixture model, and the Gaussian mixture model is constructed based on a preset number of sample data, and the sample data is the frequency data of each first frequency point extracted from the speech data collected by the microphone for each user sample after converting the spectral data; the frequency data of each first frequency point in the preliminary noise reduction data is compared with the correction range defined by the correction standard data of the corresponding frequency point, and the frequency data exceeding the corresponding correction range is corrected to fall within the corresponding correction range to obtain corrected noise reduction data, and the corrected noise reduction data is used as the noise reduction result of the original speech data. The present invention provides a speech noise reduction scheme, which further improves the speech noise reduction effect by further correcting the noise reduction speech data after preliminary noise reduction using a speech noise reduction algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 A schematic diagram of the hardware operating environment involved in an embodiment of the present invention;

[0038] Figure 2 This is a flow chart of a first embodiment of a method for reducing speech noise according to the present invention;

[0039] Figure 3 Schematic diagram of the functional modules of a preferred embodiment of the speech noise reduction device of the present invention.

[0040] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0041] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0042] like Figure 1 As shown, Figure 1 It is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment of the present invention.

[0043] It should be noted that the speech noise reduction device in the embodiment of the present invention can be a headset, a smart phone, a personal computer, a server and other devices, and is not specifically limited here.

[0044] like Figure 1As shown, the speech noise reduction device may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. The communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0045] Those skilled in the art will understand that Figure 1 The device structure shown in the figure does not constitute a limitation on the speech noise reduction device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0046] like Figure 1 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a voice noise reduction program. The operating system is a program that manages and controls the hardware and software resources of the device and supports the operation of the voice noise reduction program and other software or programs. Figure 1 In the device shown, the user interface 1003 is mainly used to communicate data with the client; the network interface 1004 is mainly used to establish a communication connection with the server; and the processor 1001 can be used to call the speech noise reduction program stored in the memory 1005 and perform the following operations:

[0047] Obtaining preliminary noise reduction data obtained by performing preliminary noise reduction on original voice data collected by a microphone;

[0048] Determining corrected standard data corresponding to each first frequency point based on a speech model, wherein the speech model is spectral data determined based on a Gaussian mixture model, and the Gaussian mixture model is constructed based on a preset number of sample data, and the sample data is frequency point data of each first frequency point extracted from spectral data obtained by converting speech data collected by a microphone for each user sample;

[0049] The frequency point data of each of the first frequency points in the preliminary noise reduction data is compared with the correction range defined by the correction standard data of the corresponding frequency point, and the frequency point data exceeding the corresponding correction range is corrected to fall within the corresponding correction range to obtain corrected noise reduction data, and the corrected noise reduction data is used as the noise reduction result of the original speech data.

[0050] Furthermore, the operation of determining the correction standard data corresponding to each frequency point according to the speech model includes:

[0051] Acquire bone conduction voice data collected by a bone conduction sensor, and extract first frequency point data of each second frequency point in spectrum data obtained by converting the bone conduction voice data, wherein the bone conduction voice data is collected synchronously with the original voice data, and each second frequency point is a frequency point of each first frequency point that is below a preset frequency;

[0052] Converting each first frequency point data into second frequency point data according to a preset mapping relationship between each second frequency point bone conduction frequency point data and microphone frequency point data;

[0053] The maximum frequency point data in the second frequency point data corresponding to each second frequency point is multiplied by the normalized frequency point data of each first frequency point in the speech model to obtain the corrected standard data corresponding to each first frequency point, wherein the speech model is to normalize the frequency point data of each first frequency point in the original spectrum data determined based on the Gaussian mixture model to obtain normalized spectrum data, and the normalized spectrum data includes the normalized frequency point data of each first frequency point.

[0054] Furthermore, when the correction standard data corresponding to the first frequency point is the standard frequency point data corresponding to the first frequency point, the correction range represented by the correction standard data is a range obtained by floating the standard frequency point data up and down by a preset error value;

[0055] The operation of comparing the frequency point data of each first frequency point in the preliminary noise reduction data with the correction range defined by the correction standard data of the corresponding frequency point, and correcting the frequency point data that exceeds the corresponding correction range to fall within the corresponding correction range includes:

[0056] Comparing frequency data of a target frequency point in the preliminary noise reduction data with a correction range defined by correction standard data corresponding to the target frequency point, wherein the target frequency point is any one of the first frequency points;

[0057] When the target frequency point data is greater than the upper limit of the correction range or less than the lower limit of the correction range, the target frequency point data is subtracted from the standard frequency point data corresponding to the target frequency point and then multiplied by a first preset correction coefficient. The calculation result is then added to the standard frequency point data corresponding to the target frequency point to obtain the corrected frequency point data corresponding to the target frequency point, wherein the target frequency point data is the frequency point data of the target frequency point in the preliminary noise reduction data.

[0058] Furthermore, after comparing the frequency data of each of the first frequency points in the preliminary noise reduction data with the correction range defined by the correction standard data of the corresponding frequency point, correcting the frequency data that exceeds the corresponding correction range to fall within the corresponding correction range, and obtaining the corrected noise reduction data, the processor 1001 may also be used to call the speech noise reduction program stored in the memory 1005 and perform the following operations:

[0059] Obtaining a first noise estimate used when performing preliminary noise reduction on the original speech data, the first noise estimate including frequency point noise estimates corresponding to each of the first frequency points;

[0060] When the target frequency data is greater than an upper limit of the correction range, increasing the frequency noise estimate corresponding to the target frequency in the first noise estimation to correct the frequency noise estimate corresponding to the target frequency in the first noise estimation;

[0061] When the target frequency data is less than a lower limit of the correction range, reducing the frequency noise estimate corresponding to the target frequency in the first noise estimation, so as to correct the frequency noise estimate corresponding to the target frequency in the first noise estimation;

[0062] The second noise estimate obtained by correcting the first noise estimate is used as the noise estimate used when performing noise reduction on the next frame of original speech data collected by the microphone.

[0063] Furthermore, the operation of increasing the frequency noise estimate corresponding to the target frequency in the first noise estimate includes:

[0064] The target frequency data minus the corrected frequency data corresponding to the target frequency is multiplied by a second preset correction coefficient, and the calculated result is added to the frequency noise estimate corresponding to the target frequency to obtain the frequency noise estimate after the target frequency is increased.

[0065] Furthermore, before determining the corrected standard data corresponding to each first frequency point according to the speech model, the processor 1001 may also be configured to call a speech noise reduction program stored in the memory 1005 and perform the following operations:

[0066] Acquiring a preset amount of sample data;

[0067] Constructing a preset number of Gaussian mixture models using the sample data, wherein each of the Gaussian mixture models includes a target number of Gaussian mixture components, where the target number is the number of the first frequency points;

[0068] The means of the Gaussian mixture components corresponding to the same frequency point in various Gaussian mixture models are weighted averaged to obtain the weighted average values ​​corresponding to each first frequency point, and the original spectrum data is composed of the weighted average values ​​corresponding to each first frequency point to obtain a speech model.

[0069] Furthermore, the operation of constructing a Gaussian mixture model of a preset number using the sample data includes:

[0070] Dividing the sample data into a preset number of sample data groups, wherein the preset number of groups is the same as the preset number of species;

[0071] A Gaussian mixture model is constructed using each of the sample data groups to obtain the preset number of Gaussian mixture models.

[0072] Based on the above structure, various embodiments of a speech noise reduction method are proposed.

[0073] Reference Figure 2 , Figure 2 FIG. 4 is a flow chart of a first embodiment of a speech noise reduction method according to the present invention.

[0074] The embodiment of the present invention provides an embodiment of a method for speech noise reduction. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than here. In this embodiment, the execution subject of the speech noise reduction method can be a device such as headphones, a personal computer, a smart phone, etc., which is not limited in this embodiment. For the sake of ease of description, the following description of each embodiment is omitted. In this embodiment, the speech noise reduction method includes:

[0075] Step S10, obtaining preliminary noise reduction data obtained by performing preliminary noise reduction on the original voice data collected by the microphone;

[0076] In this embodiment, speech noise reduction is performed on speech data collected by a microphone. This speech data is hereinafter referred to as original speech data for distinction. Initial noise reduction is first performed on the original speech data. The speech data after initial noise reduction is hereinafter referred to as initial noise reduction data for distinction. The speech noise reduction algorithm used for initial noise reduction is not limited in this embodiment and can be any speech noise reduction algorithm that can be used independently. That is, the speech noise reduction method proposed in this embodiment further modifies the results of initial noise reduction using other speech noise reduction algorithms to further improve the effect of speech noise reduction.

[0077] The original voice data can be data collected over a period of time at a certain sampling frequency. The duration of a frame of voice data can be pre-set, for example, 10ms. The original voice data can include a single frame of voice data, meaning that voice noise reduction can be performed once per collected frame. The original voice data can also include multiple frames of voice data, which can be divided into frames and then subjected to voice noise reduction separately, meaning that multiple frames of data can be continuously collected and then subjected to voice noise reduction separately. In this embodiment, the original recording data is described as including only one frame of voice data.

[0078] Step S20: determining correction standard data corresponding to each first frequency point based on a speech model, wherein the speech model is spectral data determined based on a Gaussian mixture model, and the Gaussian mixture model is constructed based on a preset number of sample data, and the sample data is frequency point data of each first frequency point extracted from spectral data obtained by converting speech data collected by a microphone for each user sample;

[0079] A Gaussian mixture model can be obtained in advance through sample data training, and a speech model can be obtained based on the Gaussian mixture model. Multiple user samples can be selected in advance, and each user sample can be a user of different genders and age groups; a microphone can be used to collect multiple voice data in a quiet environment for multiple user samples, and at least one voice data can be collected for each user sample; for each voice data, the voice data is converted into spectrum data, and the spectrum data includes frequency data of each frequency point (hereinafter referred to as the first frequency point), and the frequency data is the amplitude or energy of the corresponding frequency point; the frequency data of each first frequency point in the spectrum data converted from a frame of voice data is respectively used as a sample data; after constructing a preset number of sample data, each sample data can be used to construct a Gaussian mixture model, and then a speech model is constructed based on the Gaussian mixture model. The first frequency point can be all the frequency points contained in the spectrum data, or it can be a frequency point within a certain frequency range. It can be set according to specific needs. During subsequent corrections, the object of correction is also the frequency data of the first frequency point in the preliminary noise reduction data. For example, in one embodiment, assuming that the sampling rate is set to 16000, one frame is 15 ms, one frame of speech data includes 240 sampling points, and the spectrum data is converted with a window length of 256, the first 128 frequency points in the spectrum data are used as the first frequency point, that is, the frequency point below 8000 Hz is used as the first frequency point.

[0080] The training method of the Gaussian mixture model is not limited in this embodiment. For example, the EM algorithm can be used for training. The Gaussian mixture model includes multiple Gaussian mixture components. Training the Gaussian mixture model means training to obtain the parameters corresponding to each Gaussian mixture component, such as the mean, variance, and mixing coefficient; generating a spectrum data according to the parameters corresponding to each Gaussian mixture component, and the spectrum data includes frequency data of multiple frequency points; the spectrum data can be directly used as a speech model, or the spectrum data can be further processed as needed to obtain a speech model, such as normalizing each frequency point data. Among them, this embodiment does not limit the method of generating spectrum data according to the parameters corresponding to each Gaussian mixture component. For example, the mean of each Gaussian mixture component can be multiplied by its mixing coefficient as a frequency point data, and the frequency point data corresponding to each Gaussian mixture component can be combined to obtain a spectrum data.

[0081] In a specific implementation, the Gaussian mixture model may be constructed locally on the device performing speech noise reduction, or the Gaussian mixture model may be trained in other devices and the speech model may be determined, and then the speech model may be configured in the device requiring speech noise reduction.

[0082] The correction standard data corresponding to each first frequency point can be determined based on the speech model. The correction standard data can be used to limit a correction range. In a specific embodiment, the correction standard data can be the correction range itself, or it can be a standard frequency point data. The correction range can be obtained by floating the standard frequency point data up and down by a certain error value. There are many ways to determine the correction standard data based on the speech model, which are not limited in this embodiment. For example, the frequency point data of each first frequency point in the speech model can be directly used as the standard frequency point data, and the correction standard data corresponding to the first frequency point is the standard frequency point data corresponding to the first frequency point. The error value can be set as needed.

[0083] Step S30: Compare the frequency data of each of the first frequency points in the preliminary noise reduction data with the correction range defined by the correction standard data of the corresponding frequency point, correct the frequency data that exceeds the corresponding correction range to fall within the corresponding correction range, and obtain corrected noise reduction data. The corrected noise reduction data is used as the noise reduction result of the original speech data.

[0084] If the preliminary noise reduction data is speech signal data in the time domain, the preliminary noise reduction data can be converted into time domain spectrum data before subsequent correction processing. If the preliminary noise reduction data is frequency domain spectrum data, subsequent correction processing can be performed directly. The method of converting time domain data into frequency domain spectrum data is not limited in this embodiment, and for example, it can be obtained through framing, windowing, and fast Fourier transform processing.

[0085] For any first frequency point, the frequency data of the first frequency point in the preliminary noise reduction data is compared with the correction range defined by the correction standard data corresponding to the first frequency point. If the frequency data corresponding to the first frequency point in the preliminary noise reduction data exceeds the correction standard data corresponding to the first frequency point, the frequency data corresponding to the first frequency point in the preliminary noise reduction data is corrected so that the corrected frequency data falls within the corresponding correction range. After correcting the frequency data of the first frequency point that needs to be corrected in the preliminary noise reduction data, the corrected preliminary noise reduction data (hereinafter referred to as the corrected noise reduction data for distinction) can be obtained. The corrected noise reduction data can be used as the final noise reduction result of the original voice data. For example, in a specific application scenario, the corrected noise reduction data can be converted into a time-domain voice signal and played through a speaker or sent to the caller.

[0086] It should be noted that since the Gaussian mixture model is trained using sample data extracted from voice data based on user samples, the voice model is based on spectral data determined by the Gaussian mixture model. The corrected standard data corresponding to each first frequency point determined by the voice model reflects the range of the standard frequency point data of human voice at each first frequency point. When the frequency point data of the first frequency point in the preliminary noise reduction data exceeds the corresponding correction range, it indicates that it may be non-stationary noise or noise in high-noise scenes that remains after preliminary noise reduction. By correcting the frequency point data in the preliminary noise reduction data that exceeds the corresponding correction range to fall within the correction range, the noise caused by non-stationarity or high-noise scenes can be further removed, thereby further improving the noise reduction effect.

[0087] In this embodiment, preliminary noise reduction data is obtained by performing preliminary noise reduction on original speech data collected by a microphone; correction standard data corresponding to each first frequency point is determined according to a speech model, wherein the speech model is spectral data determined based on a Gaussian mixture model, and the Gaussian mixture model is constructed based on a preset number of sample data, and the sample data is the frequency data of each first frequency point extracted from the spectral data obtained after converting the speech data collected by a microphone for each user sample; the frequency data of each first frequency point in the preliminary noise reduction data is compared with the correction range defined by the correction standard data of the corresponding frequency point, and the frequency data exceeding the corresponding correction range is corrected to fall within the corresponding correction range to obtain corrected noise reduction data, and the corrected noise reduction data is used as the noise reduction result of the original speech data.

[0088] Furthermore, based on the above-mentioned first embodiment, a second embodiment of the speech noise reduction method of the present invention is proposed. In this embodiment, step S20 includes:

[0089] Step S201, acquiring bone conduction voice data collected by a bone conduction sensor, and extracting first frequency point data of each second frequency point in spectrum data obtained by converting the bone conduction voice data, wherein the bone conduction voice data is collected synchronously with the original voice data, and each second frequency point is a frequency point of each first frequency point that is below a preset frequency;

[0090] Taking into account the different voice characteristics of each user and the different volume of each user when speaking, in this embodiment, the correction standard data corresponding to each first frequency point is determined based on the user's personal microphone voice data and voice model.

[0091] Since the original voice data collected by the microphone contains noise, which will affect the judgment of the user's personal voice characteristics and speaking volume, in this embodiment, the microphone pure voice data can be obtained through a bone conduction sensor.

[0092] Specifically, the bone conduction sensor can simultaneously capture the user's bone-conducted voice data while the microphone captures the original voice data. It's important to note that in scenarios with high background noise, the voice picked up by traditional microphones is heavily contaminated by noise, while the bone conduction sensor effectively blocks out background noise interference. In other words, bone-conducted voice data can be considered pure voice data without background noise.

[0093] The bone conduction voice data is converted into spectrum data, and then the frequency data of the second frequency point is extracted (hereinafter referred to as the first frequency point data for distinction). Among them, the second frequency point is the frequency point below the preset frequency among the first frequency points. The voice picked up by the bone conduction sensor is generally below 500Hz to 1000Hz, so the preset frequency can be selected from 500Hz-1000Hz according to the requirements of the specific application scenario. For example, assuming that the sampling rate is set to 16000, one frame is 15ms, and one frame of voice data includes 240 sampling points, the spectrum data is converted with a window length of 256, and the first 128 frequency points in the spectrum data are used as the first frequency point. When the preset frequency is set to 500Hz, the second frequency point is the first 8 frequency points among the 128 first frequency points.

[0094] Step S202: converting each first frequency point data into second frequency point data according to a preset mapping relationship between each second frequency point bone conduction frequency point data and microphone frequency point data;

[0095] To obtain microphone-only voice data, a mapping relationship between the bone conduction frequency data and the microphone frequency data for each second frequency point in a quiet scene can be pre-established. Specifically, multiple sets of bone conduction voice data and microphone voice data can be collected synchronously in a quiet scene, and the mapping relationship between the bone conduction frequency data and the microphone frequency data for each second frequency point can be obtained through statistics of multiple sets of data. For example, the gain correspondence between the 500Hz spectrum data of the bone conduction sensor path and the 500Hz spectrum data of the microphone path in a quiet scene can be established:

[0096]

[0097] Among them, [ω1, ω2, ..., ω8] are empirical weights obtained through statistics of a large amount of data. Represents dot product. [f1 mic ,f1 mic ,…,f1 mic ] is the microphone frequency data of the 8 second frequency points, [f1 vpu ,f1 vpu ,…,f1 vpu ] is the bone conduction frequency data of 8 second frequency points.

[0098] According to the mapping relationship, the first frequency point data corresponding to each second frequency point is converted into the second frequency point data corresponding to each second frequency point. The second frequency point data corresponding to each second frequency point can be regarded as the pure voice data of the user at a preset frequency or below.

[0099] Step S203, multiply the maximum frequency point data in the second frequency point data corresponding to each second frequency point by the normalized frequency point data of each first frequency point in the speech model to obtain the corrected standard data corresponding to each first frequency point, wherein the speech model is to normalize the frequency point data of each first frequency point in the original spectrum data determined based on the Gaussian mixture model to obtain normalized spectrum data, and the normalized spectrum data includes the normalized frequency point data of each first frequency point.

[0100] After determining the original spectrum data based on the Gaussian mixture model, the frequency data of each first frequency point in the original spectrum data can be normalized to obtain normalized spectrum data containing the normalized frequency data of each first frequency point, and the normalized spectrum data can be used as the speech model. It can be understood that the normalized spectrum data represents the characteristics of human speech (relative to noise) that are unrelated to the individual voice characteristics and volume of each first frequency point. The normalization process can be performed by dividing the frequency data of each first frequency point in the original spectrum data by the maximum value of the frequency data of each second frequency point in the original spectrum data.

[0101] After obtaining the second frequency point data corresponding to each second frequency point, multiply the maximum frequency point data in each second frequency point data by the normalized frequency point data of each first frequency point in the speech model, and the obtained results are used as the corrected standard data corresponding to each first frequency point, or the obtained results can be floated up and down by a preset error value to obtain a range as the corrected standard data.

[0102] Furthermore, in one embodiment, step S30 includes:

[0103] Step S301, comparing the frequency data of a target frequency in the preliminary noise reduction data with a correction range defined by correction standard data corresponding to the target frequency, wherein the target frequency is any one of the first frequency points;

[0104] In this embodiment, when the corrected standard data corresponding to the first frequency point is the standard frequency point data corresponding to the first frequency point, the correction range represented by the corrected standard data is the range obtained by floating the standard frequency point data above and below a preset error value. The preset error value can be set as needed and is not limited in this embodiment. For any one of the first frequency points, hereinafter referred to as the target frequency point for distinction, the frequency data of the target frequency point in the preliminary noise reduction data is compared with the correction range defined by the corrected standard data corresponding to the target frequency point.

[0105] Step S302: When the target frequency point data is greater than the upper limit of the correction range or less than the lower limit of the correction range, the target frequency point data is subtracted from the standard frequency point data corresponding to the target frequency point and then multiplied by a first preset correction coefficient. The calculation result is then added to the standard frequency point data corresponding to the target frequency point to obtain the corrected frequency point data corresponding to the target frequency point, wherein the target frequency point data is the frequency point data of the target frequency point in the preliminary noise reduction data.

[0106] The frequency data of the target frequency point in the preliminary noise reduction data is referred to as the target frequency point data for distinction. When the target frequency point data is greater than the upper limit of the correction range or less than the lower limit of the correction range, the target frequency point data needs to be corrected so that the corrected target frequency point data falls within the correction range. In this embodiment, the target frequency point data is subtracted from the standard frequency point data corresponding to the target frequency point and then multiplied by the first preset correction coefficient, and then the calculated result is added to the standard frequency point data corresponding to the target frequency point, and the result obtained is used as the corrected frequency point data corresponding to the target frequency point. For example, assuming that Y k 2 represents the corrected frequency data of the kth first frequency point, ρ k Indicates the standard frequency data corresponding to the kth first frequency point, Y k 1Indicates the frequency data corresponding to the kth first frequency point in the preliminary noise reduction data, thd k Indicates the preset error value corresponding to the kth first frequency point data, △ k represents the first preset correction coefficient corresponding to the kth first frequency point data, then the preliminary noise reduction data can be corrected according to the following formula:

[0107]

[0108] Furthermore, in one embodiment, after step S30, the following steps are further included:

[0109] Step S40, obtaining a first noise estimate used when performing preliminary noise reduction on the original speech data, the first noise estimate including frequency noise estimates corresponding to each of the first frequency points;

[0110] In this embodiment, it is proposed that, on the basis of correcting the preliminary noise reduction data, the noise estimate used in the preliminary noise reduction (hereinafter referred to as the first noise estimate for distinction) is also corrected, so that the corrected noise estimate is used as the noise estimate used for preliminary noise reduction of the next frame of original speech data, thereby further improving the noise reduction effect.

[0111] Obtain a first noise estimate for use in performing preliminary noise reduction on the original speech data. The first noise estimate includes frequency noise estimates corresponding to each first frequency point. When performing noise reduction on the original speech data, preliminary noise reduction data is obtained by subtracting the corresponding frequency noise estimate from the frequency data of each first frequency point in the original speech data.

[0112] Step S50: When the target frequency data is greater than the upper limit of the correction range, increasing the frequency noise estimate corresponding to the target frequency in the first noise estimation to correct the frequency noise estimate corresponding to the target frequency in the first noise estimation;

[0113] When the target frequency data is greater than the upper limit of the correction range, it indicates that the frequency noise estimate of the target frequency in the original speech data estimated during the initial noise reduction is too small. In this case, the frequency noise estimate corresponding to the target frequency in the first noise estimate can be increased, and the increased frequency noise estimate corresponding to the target frequency can be used as the corrected frequency noise estimate for the target frequency. In specific embodiments, there are many ways to increase the frequency noise estimate corresponding to the target frequency in the first noise estimate, such as adding a preset threshold.

[0114] Step S60: When the target frequency data is less than the lower limit of the correction range, reducing the frequency noise estimate corresponding to the target frequency in the first noise estimation to correct the frequency noise estimate corresponding to the target frequency in the first noise estimation;

[0115] If the target frequency data is less than the upper limit of the correction range, it indicates that the frequency noise estimate of the target frequency in the original speech data estimated during the initial noise reduction is too high. In this case, the frequency noise estimate corresponding to the target frequency in the first noise estimate can be reduced, and the reduced frequency noise estimate corresponding to the target frequency can be used as the corrected frequency noise estimate for the target frequency. In specific embodiments, there are many ways to reduce the frequency noise estimate corresponding to the target frequency in the first noise estimate, such as subtracting a preset threshold.

[0116] Step S70: Using the second noise estimate obtained by correcting the first noise estimate as the noise estimate used for performing noise reduction on the next frame of original speech data collected by the microphone.

[0117] After the first noise estimate is corrected, the corrected first noise estimate is referred to as a second noise estimate for distinction. The second noise estimate can be used as the noise estimate for performing noise reduction on the next frame of original speech data collected by the microphone, thereby improving the noise reduction effect on the next frame of original speech data.

[0118] Furthermore, in one embodiment, step S50 includes:

[0119] Step S501 : subtract the corrected frequency data corresponding to the target frequency from the target frequency data and multiply the result by a second preset correction coefficient. Then, the calculation result is added to the frequency noise estimate corresponding to the target frequency to obtain the frequency noise estimate after the target frequency is increased.

[0120] In this embodiment, a method is proposed to increase the frequency noise estimate corresponding to the target frequency in the first noise estimate, that is, the target frequency data minus the corrected frequency data corresponding to the target frequency is multiplied by a second preset correction coefficient, and the calculated result is added to the frequency noise estimate corresponding to the target frequency to obtain the frequency noise estimate after the target frequency is increased. For example, assuming that Y k 2 represents the corrected frequency data of the kth first frequency point, σ k represents the frequency noise estimate corresponding to the kth first frequency point in the first noise estimate, Indicates the corrected frequency noise estimate corresponding to the kth first frequency point, Y k 1 Indicates the frequency data corresponding to the kth first frequency point in the preliminary noise reduction data, thd k represents the preset error value corresponding to the kth first frequency point data, λ k represents the second preset correction coefficient corresponding to the kth first frequency point data, then the preliminary noise reduction data can be corrected according to the following formula:

[0121]

[0122] Furthermore, based on the first and / or second embodiments above, a third embodiment of the speech noise reduction method of the present invention is proposed. In this embodiment, before step S20, the method further includes:

[0123] Step A10, obtaining a preset amount of sample data;

[0124] In order to improve the ability of the speech model to summarize human speech features, in this embodiment, it is proposed to communicate multiple Gaussian mixture models and obtain a speech model based on the multiple Gaussian mixture models.

[0125] Specifically, a preset number of sample data may be obtained, wherein the preset number may be set as needed and is not limited in this embodiment.

[0126] Step A20, constructing a preset number of Gaussian mixture models using the sample data, wherein each of the Gaussian mixture models includes a target number of Gaussian mixture components, and the target number is the number of the first frequency points;

[0127] A Gaussian mixture model of a preset number is constructed using sample data. In a specific embodiment, different types of Gaussian mixture models can be constructed using the same batch of sample data by setting different hyperparameters (such as learning rate, convergence threshold, etc.), or a Gaussian mixture model can be constructed separately using different batches of sample data to obtain a variety of different Gaussian mixture models. In this embodiment, there is no restriction on the manner of constructing a variety of different Gaussian mixture models. The preset number of species can be set as needed, for example, it can be determined based on the amount of sample data. The larger the amount of sample data, the more preset number of species can be set.

[0128] The number of Gaussian mixture components in the Gaussian mixture model can be pre-set to be consistent with the number of first frequency points. That is, since the sample data comes from the first frequency points in the spectral data converted from the speech data, the Gaussian mixture model can be regarded as a model obtained by combining the distribution of the frequency point data of various speech data at each first frequency point, with each first frequency point corresponding to a distribution (Gaussian mixture component).

[0129] Step A30, perform weighted averaging on the means of the Gaussian mixture components corresponding to the same frequency points in the various Gaussian mixture models to obtain the weighted average values ​​corresponding to the first frequency points, and obtain the speech model by composing the original spectrum data based on the weighted average values ​​corresponding to the first frequency points.

[0130] Each Gaussian mixture model includes Gaussian mixture components corresponding to each first frequency point. The weighted average of the means of the Gaussian mixture components of the same first frequency point in various Gaussian mixture models can be taken to obtain the weighted average values ​​corresponding to each first frequency point. The weighted average values ​​corresponding to each first frequency point constitute the original spectrum data. In a specific embodiment, the original spectrum data can be directly used as a speech model, or the normalized spectrum data obtained after the original spectrum data is normalized can be used as a speech model. The weight of the weighted average can be set as needed. For example, in one embodiment, the weights corresponding to various Gaussian mixture models can be calculated based on the amount of sample data used to construct various Gaussian mixture models. That is, when the amount of sample data used to construct the Gaussian mixture model is larger, the weight corresponding to the Gaussian mixture model is larger.

[0131] By constructing multiple Gaussian mixture models based on sample data and then obtaining a speech model based on the multiple Gaussian mixture models, the speech model can be made to more accurately summarize the characteristics of human speech, thereby determining the correction standard data based on the speech model, and correcting the preliminary noise reduction data based on the correction standard data, thereby obtaining a better noise reduction effect.

[0132] Furthermore, in one embodiment, step A20 includes:

[0133] Step A201, dividing the sample data into a preset number of sample data groups, wherein the preset number of groups is the same as the preset number of species;

[0134] In step A202 , a Gaussian mixture model is constructed using each of the sample data groups to obtain the preset number of Gaussian mixture models.

[0135] In this embodiment, a specific implementation method for constructing multiple Gaussian mixture models is proposed. Specifically, the sample data is divided to obtain a preset number of sample data groups. The preset number of groups is the number of Gaussian mixture models to be constructed. The division method is not limited in this embodiment. For example, the division can be random, and the number of sample data in each sample data group is not limited. In one embodiment, when dividing, it can be ensured that each sample data group includes sample data constructed based on the second frequency point data of each frequency point under the preset frequency. By using each sample data group to construct a Gaussian mixture model, multiple Gaussian mixture models can be obtained.

[0136] In addition, the embodiment of the present invention also proposes a speech noise reduction device, referring to Figure 3 , the speech noise reduction device includes:

[0137] An acquisition module 10 is configured to obtain preliminary noise reduction data obtained by performing preliminary noise reduction on the original speech data collected by the microphone;

[0138] a determination module 20 for determining, based on a speech model, corrected standard data corresponding to each first frequency point, wherein the speech model is spectral data determined based on a Gaussian mixture model, the Gaussian mixture model being constructed based on a preset number of sample data, the sample data being frequency data of each first frequency point extracted from spectral data obtained by converting speech data collected by a microphone for each user sample;

[0139] The correction module 30 is used to compare the frequency point data of each first frequency point in the preliminary noise reduction data with the correction range defined by the correction standard data of the corresponding frequency point, correct the frequency point data that exceeds the corresponding correction range to fall within the corresponding correction range, obtain corrected noise reduction data, and use the corrected noise reduction data as the noise reduction result of the original speech data.

[0140] Furthermore, the determining module 20 is further configured to:

[0141] Acquire bone conduction voice data collected by a bone conduction sensor, and extract first frequency point data of each second frequency point in spectrum data obtained by converting the bone conduction voice data, wherein the bone conduction voice data is collected synchronously with the original voice data, and each second frequency point is a frequency point of each first frequency point that is below a preset frequency;

[0142] Converting each first frequency point data into second frequency point data according to a preset mapping relationship between each second frequency point bone conduction frequency point data and microphone frequency point data;

[0143] The maximum frequency point data in the second frequency point data corresponding to each second frequency point is multiplied by the normalized frequency point data of each first frequency point in the speech model to obtain the corrected standard data corresponding to each first frequency point, wherein the speech model is to normalize the frequency point data of each first frequency point in the original spectrum data determined based on the Gaussian mixture model to obtain normalized spectrum data, and the normalized spectrum data includes the normalized frequency point data of each first frequency point.

[0144] Furthermore, when the correction standard data corresponding to the first frequency point is the standard frequency point data corresponding to the first frequency point, the correction range represented by the correction standard data is a range obtained by floating the standard frequency point data up and down by a preset error value;

[0145] The correction module 30 is further configured to:

[0146] Comparing frequency data of a target frequency point in the preliminary noise reduction data with a correction range defined by correction standard data corresponding to the target frequency point, wherein the target frequency point is any one of the first frequency points;

[0147] When the target frequency point data is greater than the upper limit of the correction range or less than the lower limit of the correction range, the target frequency point data is subtracted from the standard frequency point data corresponding to the target frequency point and then multiplied by a first preset correction coefficient. The calculation result is then added to the standard frequency point data corresponding to the target frequency point to obtain the corrected frequency point data corresponding to the target frequency point, wherein the target frequency point data is the frequency point data of the target frequency point in the preliminary noise reduction data.

[0148] Furthermore, the acquisition module 10 is further configured to:

[0149] Obtaining a first noise estimate used when performing preliminary noise reduction on the original speech data, the first noise estimate including frequency point noise estimates corresponding to each of the first frequency points;

[0150] The correction module 30 is further configured to:

[0151] When the target frequency data is greater than an upper limit of the correction range, increasing the frequency noise estimate corresponding to the target frequency in the first noise estimation to correct the frequency noise estimate corresponding to the target frequency in the first noise estimation;

[0152] When the target frequency data is less than a lower limit of the correction range, reducing the frequency noise estimate corresponding to the target frequency in the first noise estimation, so as to correct the frequency noise estimate corresponding to the target frequency in the first noise estimation;

[0153] The second noise estimate obtained by correcting the first noise estimate is used as the noise estimate used when performing noise reduction on the next frame of original speech data collected by the microphone.

[0154] Furthermore, the correction module 30 is further configured to:

[0155] The target frequency data minus the corrected frequency data corresponding to the target frequency is multiplied by a second preset correction coefficient, and the calculated result is added to the frequency noise estimate corresponding to the target frequency to obtain the frequency noise estimate after the target frequency is increased.

[0156] Furthermore, the acquisition module 10 is further configured to:

[0157] Acquiring a preset amount of sample data;

[0158] The speech noise reduction device further includes:

[0159] a construction module, configured to construct a preset number of Gaussian mixture models using the sample data, wherein each of the Gaussian mixture models includes a target number of Gaussian mixture components, where the target number is the number of the first frequency points;

[0160] The means of the Gaussian mixture components corresponding to the same frequency point in various Gaussian mixture models are weighted averaged to obtain the weighted average values ​​corresponding to each first frequency point, and the original spectrum data is composed of the weighted average values ​​corresponding to each first frequency point to obtain a speech model.

[0161] Furthermore, the building block is also used to:

[0162] Dividing the sample data into a preset number of sample data groups, wherein the preset number of groups is the same as the preset number of species;

[0163] A Gaussian mixture model is constructed using each of the sample data groups to obtain the preset number of Gaussian mixture models.

[0164] The various embodiments of the speech noise reduction device of the present invention can refer to the various embodiments of the speech noise reduction method of the present invention, and will not be repeated here.

[0165] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which a speech noise reduction program is stored. When the speech noise reduction program is executed by a processor, the steps of the speech noise reduction method described below are implemented.

[0166] The various embodiments of the speech noise reduction device and the computer-readable storage medium of the present invention can refer to the various embodiments of the speech noise reduction method of the present invention, and will not be repeated here.

[0167] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0168] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0169] Through the description of the above embodiments, those skilled in the art will clearly understand that the above-mentioned embodiments and methods can be implemented by software plus the necessary general-purpose hardware platform. Of course, hardware can also be used, but in many cases the former is a more preferred embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, or optical disk) and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0170] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A speech noise reduction method, characterized in that: The speech noise reduction method comprises the following steps: Obtaining preliminary noise reduction data obtained by performing preliminary noise reduction on original voice data collected by a microphone; Determining corrected standard data corresponding to each first frequency point based on a speech model, wherein the speech model is spectral data determined based on a Gaussian mixture model, and the Gaussian mixture model is constructed based on a preset number of sample data, and the sample data is frequency point data of each first frequency point extracted from spectral data obtained by converting speech data collected by a microphone for each user sample; comparing the frequency data of each of the first frequency points in the preliminary noise reduction data with a correction range defined by the correction standard data of the corresponding frequency point, correcting the frequency data that exceeds the corresponding correction range to fall within the corresponding correction range, thereby obtaining corrected noise reduction data, and using the corrected noise reduction data as a noise reduction result for the original speech data; The step of determining the correction standard data corresponding to each first frequency point according to the speech model includes: Acquire bone conduction voice data collected by a bone conduction sensor, and extract first frequency point data of each second frequency point in spectrum data obtained by converting the bone conduction voice data, wherein the bone conduction voice data is collected synchronously with the original voice data, and each second frequency point is a frequency point of each first frequency point that is below a preset frequency; Converting each first frequency point data into second frequency point data according to a preset mapping relationship between each second frequency point bone conduction frequency point data and microphone frequency point data; The maximum frequency point data in the second frequency point data corresponding to each second frequency point is multiplied by the normalized frequency point data of each first frequency point in the speech model to obtain the corrected standard data corresponding to each first frequency point, wherein the speech model is to normalize the frequency point data of each first frequency point in the original spectrum data determined based on the Gaussian mixture model to obtain normalized spectrum data, and the normalized spectrum data includes the normalized frequency point data of each first frequency point.

2. The speech noise reduction method according to claim 1, wherein: When the correction standard data corresponding to the first frequency point is the standard frequency point data corresponding to the first frequency point, the correction range represented by the correction standard data is the range obtained by floating the standard frequency point data up and down by a preset error value; The step of comparing the frequency point data of each first frequency point in the preliminary noise reduction data with the correction range defined by the correction standard data of the corresponding frequency point, and correcting the frequency point data that exceeds the corresponding correction range to fall within the corresponding correction range includes: Comparing frequency data of a target frequency point in the preliminary noise reduction data with a correction range defined by correction standard data corresponding to the target frequency point, wherein the target frequency point is any one of the first frequency points; When the target frequency point data is greater than the upper limit of the correction range or less than the lower limit of the correction range, the target frequency point data is subtracted from the standard frequency point data corresponding to the target frequency point and then multiplied by a first preset correction coefficient. The calculation result is then added to the standard frequency point data corresponding to the target frequency point to obtain the corrected frequency point data corresponding to the target frequency point, wherein the target frequency point data is the frequency point data of the target frequency point in the preliminary noise reduction data.

3. The speech noise reduction method according to claim 2, wherein: After the step of comparing the frequency point data of each first frequency point in the preliminary noise reduction data with the correction range defined by the correction standard data of the corresponding frequency point, and correcting the frequency point data that exceeds the corresponding correction range to fall within the corresponding correction range to obtain the corrected noise reduction data, the method further includes: Obtaining a first noise estimate used when performing preliminary noise reduction on the original speech data, the first noise estimate including frequency point noise estimates corresponding to each of the first frequency points; When the target frequency data is greater than an upper limit of the correction range, increasing the frequency noise estimate corresponding to the target frequency in the first noise estimation to correct the frequency noise estimate corresponding to the target frequency in the first noise estimation; When the target frequency data is less than a lower limit of the correction range, reducing the frequency noise estimate corresponding to the target frequency in the first noise estimation, so as to correct the frequency noise estimate corresponding to the target frequency in the first noise estimation; The second noise estimate obtained by correcting the first noise estimate is used as the noise estimate used when performing noise reduction on the next frame of original speech data collected by the microphone.

4. The speech noise reduction method according to claim 3, wherein: The step of increasing the frequency noise estimate corresponding to the target frequency in the first noise estimate includes: The target frequency data minus the corrected frequency data corresponding to the target frequency is multiplied by a second preset correction coefficient, and the calculated result is added to the frequency noise estimate corresponding to the target frequency to obtain the frequency noise estimate after the target frequency is increased.

5. The speech noise reduction method according to any one of claims 1 to 4, characterized in that: Before the step of determining the correction standard data corresponding to each first frequency point according to the speech model, the method further includes: Acquiring a preset amount of sample data; Constructing a preset number of Gaussian mixture models using the sample data, wherein each of the Gaussian mixture models includes a target number of Gaussian mixture components, where the target number is the number of the first frequency points; The means of the Gaussian mixture components corresponding to the same frequency point in various Gaussian mixture models are weighted averaged to obtain the weighted average values ​​corresponding to each first frequency point, and the original spectrum data is composed of the weighted average values ​​corresponding to each first frequency point to obtain a speech model.

6. The method for reducing speech noise according to claim 5, wherein: The step of constructing a Gaussian mixture model with a preset number of species using the sample data includes: Dividing the sample data into a preset number of sample data groups, wherein the preset number of groups is the same as the preset number of species; A Gaussian mixture model is constructed using each of the sample data groups to obtain the preset number of Gaussian mixture models.

7. A speech noise reduction device, characterized in that: The speech noise reduction device comprises: An acquisition module is used to obtain preliminary noise reduction data obtained by performing preliminary noise reduction on the original voice data collected by the microphone; a determination module for determining, based on a speech model, corrected standard data corresponding to each first frequency point, wherein the speech model is spectral data determined based on a Gaussian mixture model, the Gaussian mixture model being constructed based on a preset number of sample data, the sample data being frequency data of each first frequency point extracted from spectral data obtained by converting speech data collected by a microphone for each user sample; a correction module, configured to compare the frequency point data of each first frequency point in the preliminary noise reduction data with a correction range defined by the correction standard data of the corresponding frequency point, correct the frequency point data that exceeds the corresponding correction range to fall within the corresponding correction range, thereby obtaining corrected noise reduction data, and using the corrected noise reduction data as a noise reduction result for the original speech data; The determining module is further configured to: Acquire bone conduction voice data collected by a bone conduction sensor, and extract first frequency point data of each second frequency point in spectrum data obtained by converting the bone conduction voice data, wherein the bone conduction voice data is collected synchronously with the original voice data, and each second frequency point is a frequency point of each first frequency point that is below a preset frequency; Converting each first frequency point data into second frequency point data according to a preset mapping relationship between each second frequency point bone conduction frequency point data and microphone frequency point data; The maximum frequency point data in the second frequency point data corresponding to each second frequency point is multiplied by the normalized frequency point data of each first frequency point in the speech model to obtain the corrected standard data corresponding to each first frequency point, wherein the speech model is to normalize the frequency point data of each first frequency point in the original spectrum data determined based on the Gaussian mixture model to obtain normalized spectrum data, and the normalized spectrum data includes the normalized frequency point data of each first frequency point.

8. A speech noise reduction device, characterized in that: The speech noise reduction device includes: a memory, a processor, and a speech noise reduction program stored in the memory and executable on the processor. When the speech noise reduction program is executed by the processor, the steps of the speech noise reduction method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a speech noise reduction program, which, when executed by a processor, implements the steps of the speech noise reduction method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Noise elimination method, device thereof and computer-readable storage medium

    CN107833579A

  • Audio signal noise reducing method, electronic device and storage medium

    CN113539285A