Voice noise reduction method and device

CN115862654BActive Publication Date: 2026-09-11BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211515812.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-29
Publication Date
2026-09-11
Estimated Expiration
2042-11-29

AI Technical Summary

Technical Problem

[0002]目前,在进行语音通话时,由于噪声的存在,语音的质量会下降

Benefits of technology

[0003] The present invention aims to at least partially solve one of the technical problems in the related art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115862654B_ABST
    Figure CN115862654B_ABST
Patent Text Reader

Abstract

The application provides a speech noise reduction method and device, the method comprises the following steps: obtaining air conduction sound signals and bone conduction sound signals, and bone conduction target event signal segments, bone conduction non-target event signal segments, air conduction target event signal segments and air conduction non-target event signal segments; fusing signals in at least part of frequency bands in the bone conduction target event signal segments and the air conduction target event signal segments to generate a fused signal; and performing noise reduction processing on the fused signal and the non-target event signal segments and then outputting. Thus, the bone conduction sound signals and the air conduction sound signals are effectively fused, so that the speech signal is reduced in noise, and the speech quality of the electronic device is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech signal processing technology, and particularly to a speech noise reduction method and apparatus. Background Technology

[0002] Currently, during voice calls, the quality of the voice degrades due to noise. Among related technologies, air conduction for sound signals suffers from the problem that in environments with strong external noise, the human voice signal is difficult to separate from the noise; bone conduction for sound signals is susceptible to interference from other high-frequency signals. Summary of the Invention

[0003] The present invention aims to at least partially solve one of the technical problems in the related art.

[0004] Therefore, the first objective of this invention is to propose a speech noise reduction method that effectively fuses bone conduction sound signals and air conduction sound signals to reduce speech noise and improve the speech quality of electronic devices.

[0005] The second objective of this invention is to provide a speech noise reduction device.

[0006] The third objective of this invention is to provide a speech noise reduction device.

[0007] A fourth objective of this invention is to provide a non-transitory computer-readable storage medium.

[0008] The fifth objective of this invention is to provide a computer program product.

[0009] To achieve the above objectives, a first aspect of the present invention provides a speech denoising method applied to an electronic device, comprising: acquiring an air-conducted sound signal collected by an air-conducting microphone; acquiring a bone-conducted sound signal collected by a bone-conducting microphone; detecting the bone-conducted sound signal using a preset acoustic event monitoring algorithm and a preset target event to obtain the start time and end time of the target event, wherein the target event is a voice emitted by a user holding the electronic device; segmenting the bone-conducted sound signal and the air-conducted sound signal according to the start time and the end time to obtain a bone-conducted target event signal segment, a bone-conducted targetless event signal segment, an air-conducted target event signal segment, and an air-conducted targetless event signal segment; performing fusion calculation on signals in at least a portion of the frequency bands of the bone-conducted target event signal segment and the air-conducted target event signal segment to obtain a fused signal; denoising the fused signal and outputting it; and denoising the bone-conducted targetless event signal segment and the air-conducted targetless event signal segment respectively and outputting them.

[0010] The speech denoising method of this invention acquires air conduction sound signals and bone conduction sound signals, as well as bone conduction target event signal segments, bone conduction non-target event signal segments, air conduction target event signal segments, and air conduction non-target event signal segments. It then fuses signals within at least a portion of the frequency bands of the bone conduction target event signal segments and air conduction target event signal segments to generate a fused signal. The fused signal and the non-target event signal segments are then subjected to denoising processing before being output. Thus, by effectively fusing bone conduction sound signals and air conduction sound signals, speech signal denoising is achieved, improving the speech quality of electronic devices.

[0011] To achieve the above objectives, a second aspect of the present invention provides a speech noise reduction device applied to an electronic device, comprising: a first acquisition module for acquiring air-conducted sound signals collected by an air-conducted microphone; a second acquisition module for acquiring bone-conducted sound signals collected by a bone-conducted microphone; a detection module for detecting the bone-conducted sound signals using a preset acoustic event monitoring algorithm and a preset target event to obtain the start time and end time of the target event, wherein the target event is a voice emitted by a user holding the electronic device; a processing module for segmenting the bone-conducted sound signals and the air-conducted sound signals according to the start time and the end time to obtain a bone-conducted target event signal segment, a bone-conducted targetless event signal segment, an air-conducted target event signal segment, and an air-conducted targetless event signal segment; a fusion module for performing fusion calculations on signals in at least a portion of the frequency bands of the bone-conducted target event signal segment and the air-conducted target event signal segment to obtain a fused signal, and outputting the fused signal after noise reduction; and a noise reduction module for outputting the bone-conducted targetless event signal segment and the air-conducted targetless event signal segment after noise reduction.

[0012] The speech noise reduction device of this invention acquires air-conducted sound signals and bone-conducted sound signals, as well as bone-conducted target event signal segments, bone-conducted targetless event signal segments, air-conducted target event signal segments, and air-conducted targetless event signal segments. It then fuses signals within at least a portion of the frequency bands of the bone-conducted target event signal segments and the air-conducted target event signal segments to generate a fused signal. The fused signal and the targetless event signal segments are then subjected to noise reduction processing before being output. Thus, by effectively fusing bone-conducted sound signals and air-conducted sound signals, speech signal noise reduction is achieved, improving the speech quality of electronic devices.

[0013] To achieve the above objectives, a third aspect of the present invention provides a speech noise reduction device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect of the present invention.

[0014] To achieve the above objectives, a fourth aspect of the present invention provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the method described in the first aspect of the present invention.

[0015] To achieve the above objectives, a fifth aspect of the present invention provides a computer program product comprising a computer program that, when executed by a processor, implements the method described in the first aspect of the present invention.

[0016] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0017] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0018] Figure 1 This is a schematic flowchart of a speech noise reduction method provided in an embodiment of the present invention;

[0019] Figure 2 This is a schematic flowchart of another speech noise reduction method provided in an embodiment of the present invention;

[0020] Figure 3 A schematic diagram of another speech noise reduction method provided in an embodiment of the present invention;

[0021] Figure 4 This is a schematic diagram of the structure of a signal fusion module provided in an embodiment of the present invention;

[0022] Figure 5 This is a schematic diagram of the structure of a speech noise reduction device provided in an embodiment of the present invention;

[0023] Figure 6 This is a block diagram of a voice noise reduction device for implementing voice noise reduction function, provided in an embodiment of the present invention. Detailed Implementation

[0024] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0025] The speech noise reduction method and apparatus of the present invention are described below with reference to the accompanying drawings.

[0026] Figure 1 This is a schematic flowchart of a speech noise reduction method provided in an embodiment of the present invention.

[0027] Currently, during voice calls, the quality of the voice degrades due to noise. Among related technologies, air conduction for sound signals suffers from the problem that in environments with strong external noise, the human voice signal is difficult to separate from the noise; bone conduction for sound signals is susceptible to interference from other high-frequency signals.

[0028] To address this problem, embodiments of the present invention provide a speech noise reduction method and apparatus, which effectively fuses bone conduction sound signals and air conduction sound signals to reduce speech noise and improve the speech quality of electronic devices. Figure 1 As shown, it should be noted that the speech denoising method of the present invention is executed by a speech denoising device. The speech denoising method of the present invention can be executed by the speech denoising device of the present invention. The speech denoising device of the present invention can be configured in any speech denoising device, or in any software on any speech denoising device, to execute the speech denoising method of the present invention.

[0029] The speech denoising device can be a signal fusion module. The following embodiment uses a signal fusion module as an example to illustrate the speech denoising device. The speech denoising method includes the following steps:

[0030] Step 101: Acquire the air-conducted sound signal collected by the air-conducted microphone.

[0031] In this embodiment, the speech noise reduction method can be applied to electronic devices, such as mobile phones or wearable devices with the ability to collect sound wave signals, such as headphones.

[0032] Step 102: Acquire the bone conduction sound signal collected by the bone conduction microphone.

[0033] Among them, air-conducted sound signals and bone-conducted sound signals are both sound wave signals (vibration signals). Air-conducted sound signals are sound wave signals that travel through the air to electronic devices, and the speech noise reduction device collects air-conducted sound signals; bone-conducted sound signals are sound wave signals that travel through the bones to electronic devices, and the speech noise reduction device collects bone-conducted sound signals.

[0034] Step 103: Detect bone conduction sound signals using a preset acoustic event monitoring algorithm and a preset target event to obtain the start and end times of the target event, where the target event is the voice emitted by the user holding the electronic device.

[0035] As one possible implementation, the speech noise reduction device can perform step 103 by preprocessing the bone conduction sound signal, detecting the target event in the bone conduction sound signal, and determining the start and end times of the target event in the bone conduction sound signal. The start time is the moment when the human voice begins to appear in the bone conduction sound signal, and the end time is the moment when the human voice ends in the bone conduction sound signal. There can be multiple start and end times.

[0036] Step 104: Based on the start time and end time, the bone conduction sound signal and the air conduction sound signal are segmented to obtain the bone conduction target event signal segment, the bone conduction targetless event signal segment, the air conduction target event signal segment, and the air conduction targetless event signal segment.

[0037] As one possible implementation, since bone conduction sound signals and air conduction sound signals have multiple start times and multiple end times, the multiple start times and multiple end times of bone conduction sound signals and air conduction sound signals are located respectively. Based on the location of the multiple start times and multiple end times, the bone conduction sound signals and air conduction sound signals are segmented to obtain multiple bone conduction target event signal segments, multiple bone conduction non-target event signal segments, multiple air conduction target event signal segments, and multiple air conduction non-target event signal segments.

[0038] Step 105: Perform fusion calculation on signals in at least a portion of the frequency bands of the bone conduction target event signal segment and the air conduction target event signal segment to obtain a fused signal, and output the fused signal after noise reduction.

[0039] As one possible implementation, after obtaining the bone conduction target event signal segment and the air conduction target event signal segment, the bone conduction target event signal segment and the air conduction target event signal segment are segmented to obtain multiple frequency bands. For at least one of the multiple frequency bands, the bone conduction target event signal segment and the air conduction target event signal segment are fused and calculated. The fused signal of at least one frequency band is spliced ​​and processed to obtain a fused signal.

[0040] Step 106: Denoise the bone conduction targetless event signal segment and the air conduction targetless event signal segment respectively before outputting.

[0041] In this embodiment of the invention, after obtaining the fused signal, noise reduction can be applied to the fused signal, the bone conduction targetless event signal segment, and the air conduction targetless event signal segment. Optionally, a noise reduction method based on a neural network model can be used, employing a trained neural network model to perform noise reduction processing on the fused signal, the bone conduction targetless event signal segment, and the air conduction targetless event signal segment, thereby further improving the quality of the speech.

[0042] The speech noise reduction method of this invention acquires air-conducted sound signals from an air-conducted microphone and bone-conducted sound signals from a bone-conducted microphone. It then detects the bone-conducted sound signals using a preset acoustic event monitoring algorithm and a preset target event to obtain the start and end times of the target event, where the target event is a voice message emitted by a user holding an electronic device. Based on the start and end times, the bone-conducted and air-conducted sound signals are segmented to obtain bone-conducted target event signal segments, bone-conducted target event signal segments, air-conducted target event signal segments, and air-conducted target event signal segments. The signals in at least a portion of the frequency bands of the bone-conducted and air-conducted target event signal segments are fused to obtain a fused signal, which is then denoised and output. The bone-conducted target event signal segments and air-conducted target event signal segments are also denoised and output. Thus, by effectively fusing the bone-conducted and air-conducted sound signals, speech signal noise reduction is achieved, improving the speech quality of electronic devices.

[0043] To clearly illustrate the speech noise reduction method provided by this invention, this embodiment also provides another speech noise reduction method. Figure 2 This is a schematic flowchart of another speech noise reduction method provided in an embodiment of the present invention.

[0044] like Figure 2 As shown, the speech noise reduction method may include the following steps:

[0045] Step 201: Acquire the air-conducted sound signal collected by the air-conducted microphone.

[0046] Step 202: Acquire the bone conduction sound signal collected by the bone conduction microphone.

[0047] Step 203: Detect bone conduction sound signals using a preset acoustic event monitoring algorithm and a preset target event to obtain the start and end times of the target event, where the target event is the voice emitted by the user holding the electronic device.

[0048] Step 204: Based on the start time and end time, the bone conduction sound signal and the air conduction sound signal are segmented to obtain the bone conduction target event signal segment, the bone conduction targetless event signal segment, the air conduction target event signal segment, and the air conduction targetless event signal segment.

[0049] Step 205: Based on the preset window function, the bone conduction target event signal segment and the air conduction target event signal segment are segmented within the frequency band to obtain multiple frequency bands.

[0050] In this embodiment of the invention, since air conduction sound signals and bone conduction sound signals are not frequency-stationary signals, segmenting and smoothing cross-correlation of the bone conduction target event signal segment and the air conduction target event signal of the electronic device in different frequency bands can more effectively determine the correlation of each frequency band, thereby achieving better fusion of air conduction sound signals and bone conduction sound signals.

[0051] The window function can be preset according to the nature and segmentation requirements of the bone conduction target event signal segment and the air conduction target event signal segment.

[0052] Step 206: For each frequency band, determine the corresponding smooth cross-correlation coefficient, where the smooth cross-correlation coefficient is used to represent the similarity between the bone conduction target event signal segment and the air conduction target event signal segment in that frequency band.

[0053] As one possible implementation, the smooth cross-correlation coefficients between the air conduction target event signal segment and the bone conduction target event signal segment in each frequency band are calculated using the following formula (1):

[0054] seg_coff=F -1 (F(mic_air)×F * (mic_bone)×wind)(1), where seg_coff represents the smooth cross-correlation coefficient, F -1 Let F denote the inverse Fourier transform function, and let F denote the Fourier transform function. * The conjugate function of the Fourier transform is represented, wind represents the preset window function, mic_air is the air conduction target event signal segment, and mic_bone is the bone conduction target event signal segment.

[0055] Step 207: Determine the fusion frequency band and the non-fusion frequency band based on the smooth cross-correlation coefficient. The fusion frequency band refers to the frequency band that needs to be fused, and the non-fusion frequency band refers to the frequency band that does not need to be fused.

[0056] As one possible implementation, if the smooth cross-correlation coefficient of any frequency band among the multiple frequency bands is greater than or equal to a preset threshold, then the frequency band is determined to be a fused frequency band; if the smooth cross-correlation coefficient of any frequency band among the multiple frequency bands is less than the preset threshold, then the frequency band is determined to be a non-fused frequency band.

[0057] Optionally, If the smoothed cross-correlation coefficient seg_coff is greater than or equal to the preset threshold β, and the correlation function f_seg is 1, it indicates that the current frequency band is valid and can be fused; if the smoothed cross-correlation coefficient seg_coff is less than the preset threshold β, and the correlation function f_seg is 0, it indicates that the current frequency band is invalid and cannot be fused. The preset threshold β can be modified according to the actual usage of the electronic equipment.

[0058] As one possible approach, for frequency bands that do not require fusion, the signal with the higher signal-to-noise ratio in the current frequency band is selected from the air conduction sound signal and the bone conduction sound signal as the enhanced speech signal.

[0059] Step 208: Perform fusion calculations on the fusion frequency bands according to the preset fusion algorithm to obtain the fused signal.

[0060] As one possible implementation, noise estimation is performed on the bone conduction event signal segment and the air conduction event signal segment respectively to obtain the bone conduction signal-to-noise ratio and the air conduction signal-to-noise ratio; based on the bone conduction signal-to-noise ratio and the air conduction signal-to-noise ratio, the fusion coefficient is determined; based on the fusion coefficient and the preset fusion algorithm, fusion calculation is performed to obtain the fused signal.

[0061] Optionally, the process of determining the fusion coefficient based on the bone conduction signal-to-noise ratio and the air conduction signal-to-noise ratio can be as follows: calculate the fusion coefficients of the air conduction sound signal and the bone conduction sound signal respectively based on the signal-to-noise ratios of the air conduction sound signal and the bone conduction sound signal in multiple frequency bands and the function for determining the fusion coefficient.

[0062] As one possible implementation, noise estimation is performed on the air conduction target event signal segment and the bone conduction target event signal segment at each frequency band to generate the signal-to-noise ratio of the air conduction target event signal segment and the bone conduction target event signal segment at each frequency band.

[0063] Optionally, noise estimation is performed on the gas-conducted target event signal segment, and the power of the gas-conducted target event signal segment is estimated to be P. air Noise estimation is performed on the bone conduction target event signal segment separated by VAD (Voice Activity Detection), and the power of the bone conduction target event signal segment is estimated to be P. bone .

[0064] Assume the estimated noise power is Pnoise Then P noise (t,f)=|X(t,f)| 2 (2), where X represents the Fourier transform of the signal, |X(t,f)| 2 Let X(t,f) represent the power spectrum estimated by the periodogram method, where X(t,f) represents the value at the f-th frequency point in the t-th frame, and P... noise (t,f) is the estimated noise value at the f-th frequency point in the t-th frame.

[0065] So during the voice activity, there is P noise (t,f)=γ(t,f)·P noise (t,f)+[1-γ(t,f)]|X(t,f)| 2 (3), that is Where H1 indicates the presence of speech activity, H0 indicates the absence of speech activity, and γ(t,f) is a smoothing factor for estimating noise power during speech activity. It is a function of time and frequency and can be calculated based on the SNR (signal-to-noise ratio) σ(t,f) of each frequency point of each frame of signal.

[0066] In this embodiment of the invention, γ(t,f) can be calculated by the recursive averaging algorithm of the following formula (5): The sigmoid function (activation function) is used to map the range of γ(t,f) to between 0 and 1. α adjusts the slope of the sigmoid function; typically, α is less than 1, used to reduce the sensitivity of γ(t,f) to the signal-to-noise ratio and prevent P... noise (t,f) becomes a step function, leading to an underestimation of noise.

[0067] When noise is underestimated, P noise The (t,f) curve is relatively smooth with a small variance (e.g., less than 5), in which case adjustment can be made by decreasing the value of α; when the noise is overestimated, P noise If (t,f) changes drastically and has a large variance (e.g., greater than 10), then the variance can be adjusted by increasing the value of α. The value of α can be, for example, 0.5.

[0068] σ(t,f) can be calculated using the following formula (6). Where Q represents the past Q frames, which is generally a value between 10 and 20; q represents the frame number variable, which starts from 1.

[0069] In this embodiment of the invention, according to the above formulas (2) to (6), the noise is replaced by air conduction target event signal segments and bone conduction target event signal segments, and the power of the estimated air conduction target event signal segment is calculated as P. air The power of the target event signal segment in bone conduction is estimated to be P.bone Gas conduction signal-to-noise ratio σ air Bone conduction signal-to-noise ratio σ bone .

[0070] The fusion coefficients of air conduction target event signals and bone conduction target event signals are calculated using the following formulas.

[0071] λ air =g(σ air ,σ bone (7), λ bone =g(σ bone ,σ air (8), where λ air Let λ be the fusion coefficient of the gas-conducted target event signal. bone Let g(·) be the fusion coefficient of the bone conduction target event signal, and σ be the function that determines the fusion coefficient. air For gas conduction signal-to-noise ratio, σ bone Let g be the bone conduction signal-to-noise ratio. The specific formula for the g(·) function can be...

[0072] Step 209: Noise reduction is performed on the bone conduction targetless event signal segment and the air conduction targetless event signal segment before output.

[0073] It should be noted that the specific descriptions of steps 201, 202, 203, 204 and 209 can be found in other embodiments of the present invention, and will not be explained in detail here.

[0074] In summary, the system acquires air-conducted sound signals from an air-conducted microphone and bone-conducted sound signals from a bone-conducted microphone. It then detects the bone-conducted sound signals using a preset acoustic event monitoring algorithm and a preset target event to obtain the start and end times of the target event, where the target event is the voice emitted by a user holding an electronic device. Based on the start and end times, the bone-conducted and air-conducted sound signals are segmented to obtain bone-conducted target event signal segments, bone-conducted no-target event signal segments, air-conducted target event signal segments, and air-conducted no-target event signal segments. Finally, based on a preset window function, the bone-conducted target event is processed... The signal segments and air conduction target event signal segments are segmented within the frequency band to obtain multiple frequency bands. For each frequency band, a corresponding smooth cross-correlation coefficient is determined, which represents the similarity between the bone conduction target event signal segment and the air conduction target event signal segment in that frequency band. Based on the smooth cross-correlation coefficient, fusion frequency bands and non-fusion frequency bands are determined, where fusion frequency bands refer to those that need to be fused, and non-fusion frequency bands refer to those that do not need to be fused. According to a preset fusion algorithm, fusion calculations are performed on the fusion frequency bands to obtain a fused signal. The bone conduction target event signal segment and the air conduction target event signal segment are respectively denoised and output. Thus, by effectively fusing the bone conduction sound signal and the air conduction sound signal, noise reduction of the speech signal is achieved, improving the speech quality of electronic devices.

[0075] For example, Figure 3 This is a schematic diagram of another speech noise reduction method provided in an embodiment of the present invention.

[0076] like Figure 3 As shown, in step S1, the air-conducted sound signal is received.

[0077] Step S2: Receive bone conduction sound signals.

[0078] Step S3, frequency domain segmented smooth cross-correlation (obtain the smooth cross-correlation coefficients between the air conduction target event signal segment and the bone conduction target event signal segment in each frequency band).

[0079] Step S4, environmental noise estimation (noise estimation is performed on each frequency band for the air conduction target event signal segment and the bone conduction target event signal segment).

[0080] Step S5: Perform VAD (Voice Activity Detection) detection based on the bone conduction sound signal.

[0081] Step S6: Estimate the effective frequency bands based on the segment correlation and determine the frequency bands that need to be fused (determine the frequency bands that need to be fused among the bone conduction target event signal segments and air conduction target event signal segments based on the smooth cross-correlation coefficients in each frequency band).

[0082] Step S7: Calculate the signal-to-noise ratio of the current frequency band for the air conduction target event signal segment and the bone conduction target event signal segment, respectively.

[0083] Step S8: Determine the corresponding fusion coefficients based on the signal-to-noise ratios of the air conduction target event signal segment and the bone conduction target event signal segment in each frequency band.

[0084] Step S9: Denoise the fused signal.

[0085] In summary, this method involves receiving air-conduction sound signals; receiving bone-conduction sound signals; obtaining the smooth cross-correlation coefficients between the air-conduction and bone-conduction target event signal segments in each frequency band; performing noise estimation for both air-conduction and bone-conduction target event signal segments in each frequency band; performing VAD (Voice Activity Detection) detection based on the bone-conduction sound signals; determining the frequency bands to be fused within the bone-conduction target event signal segments based on the smooth cross-correlation coefficients in each frequency band; calculating the signal-to-noise ratio (SNR) of the current frequency band for both air-conduction and bone-conduction target event signal segments; determining the corresponding fusion coefficients based on the SNR of the air-conduction and bone-conduction target event signal segments in each frequency band; and denoising the fused signal. Therefore, by determining the smooth cross-correlation coefficients and fusion coefficients of the air-conduction and bone-conduction target event signal segments, the air-conduction and bone-conduction target event signal segments are fused to generate a fused signal, thereby achieving noise reduction of the speech signal and improving the speech quality of electronic devices.

[0086] Figure 4 This is a schematic diagram of the structure of a signal fusion module provided in an embodiment of the present invention.

[0087] like Figure 4 As shown, the signal fusion module includes: a first module 410, a second module 420, a third module 430, and a fourth module 440.

[0088] The first module 410 is used to receive air conduction sound signals and bone conduction sound signals; the second module 420 is used to obtain the smooth cross-correlation coefficients between the air conduction target event signal segment and the bone conduction target event signal segment in each frequency band and to perform noise estimation on the air conduction target event signal segment and the bone conduction target event signal segment in each frequency band; the third module 430 is used to determine the corresponding fusion coefficients based on the signal-to-noise ratio of the air conduction target event signal segment and the bone conduction target event signal segment in each frequency band; and the fourth module 440 is used to denoise the fused signal.

[0089] In summary, this method involves receiving air-conduction sound signals; receiving bone-conduction sound signals; obtaining the smooth cross-correlation coefficients between the air-conduction and bone-conduction target event signal segments in each frequency band; performing noise estimation for both air-conduction and bone-conduction target event signal segments in each frequency band; performing VAD (Voice Activity Detection) detection based on the bone-conduction sound signals; determining the frequency bands to be fused within the bone-conduction target event signal segments based on the smooth cross-correlation coefficients in each frequency band; calculating the signal-to-noise ratio (SNR) of the current frequency band for both air-conduction and bone-conduction target event signal segments; determining the corresponding fusion coefficients based on the SNR of the air-conduction and bone-conduction target event signal segments in each frequency band; and denoising the fused signal. Therefore, by determining the smooth cross-correlation coefficients and fusion coefficients of the air-conduction and bone-conduction target event signal segments, the air-conduction and bone-conduction target event signal segments are fused to generate a fused signal, thereby achieving noise reduction of the speech signal and improving the speech quality of electronic devices.

[0090] To achieve the above embodiments, the present invention also proposes a speech noise reduction device.

[0091] Figure 5 This is a schematic diagram of a speech noise reduction device provided in an embodiment of the present invention.

[0092] like Figure 5 As shown, the voice noise reduction device 500 includes: a first acquisition module 510, a second acquisition module 520, a detection module 530, a processing module 540, a fusion module 550, and a noise reduction module 560.

[0093] The first acquisition module 510 is used to acquire the air-conducted sound signal collected by the air-conducted microphone.

[0094] The second acquisition module 520 is used to acquire bone conduction sound signals collected by the bone conduction microphone;

[0095] The detection module 530 is used to detect the bone conduction sound signal through a preset acoustic event monitoring algorithm and a preset target event, and to obtain the start time and end time of the target event, wherein the target event is the voice emitted by the user holding the electronic device;

[0096] Processing module 540 is used to segment the bone conduction sound signal and the air conduction sound signal according to the start time and the end time, respectively, to obtain bone conduction target event signal segment, bone conduction no-target event signal segment, air conduction target event signal segment and air conduction no-target event signal segment;

[0097] The fusion module 550 is used to perform fusion calculations on signals in at least a portion of the frequency bands of the bone conduction target event signal segment and the air conduction target event signal segment to obtain a fused signal, and output the fused signal after noise reduction.

[0098] The noise reduction module 560 is used to perform noise reduction on the bone conduction targetless event signal segment and the air conduction targetless event signal segment respectively before outputting them.

[0099] Further, in one possible implementation of this invention, the fusion module 550 includes: a first determining unit, a second determining unit, a third determining unit, and a fusion unit; wherein, the first determining unit is used to segment the bone conduction target event signal segment and the air conduction target event signal segment within the frequency band according to a preset window function, thereby obtaining multiple frequency bands; the second determining unit is used to determine a corresponding smooth cross-correlation coefficient for each frequency band, wherein the smooth cross-correlation coefficient is used to represent the similarity between the bone conduction target event signal segment and the air conduction target event signal segment in that frequency band; the third determining unit is used to determine fusion frequency bands and non-fusion frequency bands according to the smooth cross-correlation coefficient, wherein the fusion frequency band refers to the frequency band that needs to be fused, and the non-fusion frequency band refers to the frequency band that does not need to be fused; the fusion unit is used to perform fusion calculations on the fusion frequency bands according to a preset fusion algorithm to obtain the fused signal.

[0100] In one possible implementation of this invention, the fusion unit is specifically used to: estimate the noise of the bone conduction target event signal segment and the air conduction target event signal segment respectively to obtain the bone conduction signal-to-noise ratio and the air conduction signal-to-noise ratio; determine the fusion coefficient based on the bone conduction signal-to-noise ratio and the air conduction signal-to-noise ratio; and perform fusion calculation based on the fusion coefficient and a preset fusion algorithm to obtain the fused signal.

[0101] In one possible implementation of this invention, the second determining unit is specifically configured to: determine the bone conduction target event signal segment and the air conduction target event signal segment corresponding to each frequency band; determine the Fourier transform function, the inverse Fourier transform function, and the conjugate function of the Fourier transform corresponding to each frequency band; and calculate the smooth cross-correlation coefficient between the bone conduction target event signal segment and the air conduction target event signal segment corresponding to each frequency band based on the preset window function, the Fourier transform function, the inverse Fourier transform function, and the conjugate function of the Fourier transform.

[0102] In one possible implementation of this invention, the third determining unit is specifically used to: determine that the frequency band is a fused frequency band if the smooth cross-correlation coefficient of any frequency band among the plurality of frequency bands is greater than or equal to a preset threshold; and determine that the frequency band is a non-fused frequency band if the smooth cross-correlation coefficient of any frequency band among the plurality of frequency bands is less than the preset threshold.

[0103] In one possible implementation of this invention, the fusion unit is specifically used to calculate the fusion coefficients of the air conduction target event signal segment and the bone conduction target event signal segment respectively, based on the signal-to-noise ratios of the air conduction target event signal segment and the bone conduction target event signal segment in the multiple frequency bands and a function for determining the fusion coefficients.

[0104] In one possible implementation of this invention, the fusion coefficients λ of the air conduction target event signal and the bone conduction target event signal segments are calculated using the following formula. air =g(σ air ,σ bone ), λ bone =g(σ bone ,σ air ), where λ air λ is the fusion coefficient of the gas-conducted target event signal segment. bone Let g(·) be the fusion coefficient of the bone conduction target event signal segment, and σ be the function used to determine the fusion coefficient. air Let σ be the gas conduction signal-to-noise ratio. bone The bone conduction signal-to-noise ratio is denoted as .

[0105] In one possible implementation of this invention, the smooth cross-correlation coefficient is calculated using the following formula: seg_coff = F -1 (F(mic_air)×F*(mic_bone)×wind), where seg_coff represents the smooth cross-correlation coefficient, F -1 Let F represent the inverse Fourier transform function, and let F represent the Fourier transform function.* The conjugate function of the Fourier transform is represented by , wind represents the preset window function, mic_air is the air conduction target event signal segment, and mic_bone is the bone conduction target event signal segment.

[0106] It should be noted that the foregoing explanation of the speech noise reduction method embodiment also applies to the speech noise reduction device of this embodiment, and will not be repeated here.

[0107] The speech noise reduction device of this invention acquires air-conducted sound signals collected by an air-conducting microphone and bone-conducted sound signals collected by a bone-conducting microphone. It detects the bone-conducted sound signals using a preset acoustic event monitoring algorithm and a preset target event to obtain the start and end times of the target event, where the target event is a voice message emitted by a user holding an electronic device. Based on the start and end times, the bone-conducted and air-conducted sound signals are segmented to obtain bone-conducted target event signal segments, bone-conducted target event signal segments, air-conducted target event signal segments, and air-conducted target event signal segments. The signals in at least a portion of the frequency bands of the bone-conducted and air-conducted target event signal segments are fused to obtain a fused signal, which is then denoised and output. The bone-conducted target event signal segments and air-conducted target event signal segments are also denoised and output. Thus, by effectively fusing the bone-conducted and air-conducted sound signals, speech signal noise reduction is achieved, improving the speech quality of electronic devices.

[0108] To implement the above embodiments, the present invention also provides a speech noise reduction device, a non-transitory computer-readable storage medium, and a computer program product.

[0109] The speech noise reduction device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the speech noise reduction method proposed in the first aspect embodiment of the present invention as described above.

[0110] As an example, Figure 6 A block diagram of a voice noise reduction device for implementing voice noise reduction function provided in an embodiment of the present invention, such as... Figure 6 As shown, the aforementioned voice noise reduction device 600 may further include:

[0111] The memory 610 and processor 620 are connected by a bus 630, which connects different components (including the memory 610 and the processor 620). The memory 610 stores a computer program, and when the processor 620 executes the program, it implements the speech noise reduction method proposed in the first aspect of the present invention as described above.

[0112] Bus 630 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0113] The speech noise reduction device 600 typically includes a variety of computer-readable media. These media can be any available media that can be accessed by the speech noise reduction device 600, including volatile and non-volatile media, and portable and non-portable media.

[0114] Memory 610 may also include computer system readable media in the form of volatile memory, such as random access memory (RAM) 640 and / or cache memory 650. Voice noise reduction device 600 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 660 can be used to read and write non-removable, non-volatile magnetic media (… Figure 6 Not shown; usually referred to as a "hard drive"). Although Figure 6 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 630 via one or more data media interfaces. Memory 610 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.

[0115] A program / utility 680 having a set (at least one) of program modules 670 may be stored, for example, in memory 610. Such program modules 670 include—but are not limited to—an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 670 typically perform the functions and / or methods described in the embodiments of the present invention.

[0116] The voice noise reduction device 600 can also communicate with one or more external devices 690 (e.g., keyboard, pointing device, display 691, etc.), and with one or more devices that enable a user to interact with the voice noise reduction device 600, and / or with any device that enables the voice noise reduction device 600 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 692. Furthermore, the voice noise reduction device 600 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 693. Figure 6 As shown, network adapter 693 communicates with other modules of voice noise reduction device 600 via bus 630. It should be understood that, although... Figure 6 As not shown in the diagram, other hardware and / or software modules can be used in conjunction with the voice noise reduction device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0117] The processor 620 executes various functional applications and data processing by running programs stored in the memory 610.

[0118] It should be noted that the implementation process and technical principles of the speech noise reduction device in this embodiment are explained in the foregoing description of the speech noise reduction method of this invention, and will not be repeated here.

[0119] The speech noise reduction device provided in this invention acquires air-conducted sound signals from an air-conducted microphone and bone-conducted sound signals from a bone-conducted microphone. It detects the bone-conducted sound signals using a preset acoustic event monitoring algorithm and a preset target event to obtain the start and end times of the target event, where the target event is a voice message emitted by a user holding the electronic device. Based on the start and end times, the bone-conducted and air-conducted sound signals are segmented to obtain bone-conducted target event signal segments, bone-conducted target event signal segments, air-conducted target event signal segments, and air-conducted target event signal segments. The signals in at least a portion of the frequency bands of the bone-conducted and air-conducted target event signal segments are fused to obtain a fused signal, which is then denoised and output. The bone-conducted target event signal segments and air-conducted target event signal segments are also denoised and output. Therefore, by effectively fusing the bone-conducted and air-conducted sound signals, speech signal noise reduction is achieved, improving the speech quality of the electronic device.

[0120] To implement the above embodiments, the present invention also proposes a non-transitory computer-readable storage medium, wherein when the instructions in the computer-readable storage medium are executed by the processor of the speech denoising device, the speech denoising device is able to perform the speech denoising method proposed in the first aspect of the present invention as described above.

[0121] To implement the above embodiments, the present invention also provides a computer program product, which, when executed by the processor of a speech noise reduction device, enables the speech noise reduction device to perform the speech noise reduction method proposed in the first aspect of the present invention as described above.

[0122] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0123] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0124] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of the invention pertain.

[0125] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0126] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any of the following techniques known in the art, or a combination thereof: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0127] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0128] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0129] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A speech noise reduction method, characterized in that, Applied to electronic devices, including: Acquire air-conducted sound signals collected by the air-conducted microphone; Acquire bone conduction sound signals collected by a bone conduction microphone; The bone conduction sound signal is detected by a preset acoustic event monitoring algorithm and a preset target event to obtain the start time and end time of the target event, wherein the target event is the voice emitted by the user holding the electronic device; Based on the start time and the end time, the bone conduction sound signal and the air conduction sound signal are segmented to obtain bone conduction target event signal segment, bone conduction no-target event signal segment, air conduction target event signal segment, and air conduction no-target event signal segment. Based on a preset window function, the bone conduction target event signal segment and the air conduction target event signal segment are segmented within the frequency band to obtain multiple frequency bands; For each frequency band, a corresponding smooth cross-correlation coefficient is determined, wherein the smooth cross-correlation coefficient is used to represent the similarity between the bone conduction target event signal segment and the air conduction target event signal segment in that frequency band; Based on the smooth cross-correlation coefficient, the fusion frequency band and the non-fusion frequency band are determined, wherein the fusion frequency band refers to the frequency band that needs to be fused, and the non-fusion frequency band refers to the frequency band that does not need to be fused; According to the preset fusion algorithm, the fusion frequency band is fused to obtain a fused signal, and the fused signal is output after noise reduction; The bone conduction targetless event signal segment and the air conduction targetless event signal segment are respectively denoised and then output.

2. The method as described in claim 1, characterized in that, The step of performing fusion calculations on the fusion frequency bands according to a preset fusion algorithm to obtain the fused signal includes: Noise estimation is performed on the bone conduction target event signal segment and the air conduction target event signal segment respectively to obtain the bone conduction signal-to-noise ratio and the air conduction signal-to-noise ratio; The fusion coefficient is determined based on the bone conduction signal-to-noise ratio and the air conduction signal-to-noise ratio; The fusion calculation is performed based on the fusion coefficient and the preset fusion algorithm to obtain the fused signal.

3. The method as described in claim 1, characterized in that, The determination of the corresponding smooth cross-correlation coefficient for each frequency band includes: The bone conduction target event signal segment and the air conduction target event signal segment corresponding to each frequency band are determined respectively; Determine the Fourier transform function, inverse Fourier transform function, and conjugate function of the Fourier transform corresponding to each frequency band; Based on the preset window function, the Fourier transform function, the inverse Fourier transform function, and the conjugate function of the Fourier transform, the smooth cross-correlation coefficient between the bone conduction target event signal segment and the air conduction target event signal segment corresponding to each frequency band is calculated.

4. The method as described in claim 1, characterized in that, The step of determining the fused and non-fused frequency bands based on the smooth cross-correlation coefficient includes: If the smooth cross-correlation coefficient of any frequency band among the multiple frequency bands is greater than or equal to a preset threshold, then the frequency band is determined to be a fused frequency band. If the smooth cross-correlation coefficient of any of the multiple frequency bands is less than the preset threshold, then the frequency band is determined to be a non-fusion frequency band.

5. The method as described in claim 2, characterized in that, The determination of the fusion coefficient based on the bone conduction signal-to-noise ratio and the air conduction signal-to-noise ratio includes: Based on the signal-to-noise ratios of the air conduction target event signal segment and the bone conduction target event signal segment in the multiple frequency bands and the function for determining the fusion coefficient, the fusion coefficients of the air conduction target event signal segment and the bone conduction target event signal segment are calculated respectively.

6. The method as described in claim 5, characterized in that, The fusion coefficients of the air conduction target event signal and the bone conduction target event signal segments are calculated using the following formulas. , ,in The fusion coefficient is the signal segment of the gas-conducted target event. The fusion coefficient is the value of the bone conduction target event signal segment. It is the function that determines the fusion coefficient. The gas conduction signal-to-noise ratio is denoted as . The bone conduction signal-to-noise ratio is denoted as .

7. The method as described in claim 3, characterized in that, The smooth cross-correlation coefficient is calculated using the following formula: ,in, This represents the smooth cross-correlation coefficient. This represents the inverse Fourier transform function. This represents the Fourier transform function. Denotes the conjugate function of the Fourier transform. This represents the preset window function. This refers to the gas conduction target event signal segment. This refers to the bone conduction target event signal segment.

8. A voice noise reduction device, characterized in that, Applied to electronic devices, including: The first acquisition module is used to acquire the air-conducted sound signal collected by the air-conducted microphone; The second acquisition module is used to acquire bone conduction sound signals collected by the bone conduction microphone; The detection module is used to detect the bone conduction sound signal through a preset acoustic event monitoring algorithm and a preset target event, and to obtain the start time and end time of the target event, wherein the target event is the voice emitted by the user holding the electronic device; The processing module is used to segment the bone conduction sound signal and the air conduction sound signal according to the start time and the end time, respectively, to obtain the bone conduction target event signal segment, the bone conduction no-target event signal segment, the air conduction target event signal segment, and the air conduction no-target event signal segment; The fusion module is used to perform fusion calculations on signals in at least a portion of the frequency bands of the bone conduction target event signal segment and the air conduction target event signal segment to obtain a fused signal, and output the fused signal after noise reduction; The noise reduction module is used to reduce the noise of the bone conduction targetless event signal segment and the air conduction targetless event signal segment respectively before outputting them; The fusion module includes: a first determining unit, a second determining unit, a third determining unit, and a fusion unit; The first determining unit is used to segment the bone conduction target event signal segment and the air conduction target event signal segment within the frequency band according to a preset window function, thereby obtaining multiple frequency bands. The second determining unit is used to determine the corresponding smooth cross-correlation coefficient for each frequency band, wherein the smooth cross-correlation coefficient is used to represent the similarity between the bone conduction target event signal segment and the air conduction target event signal segment in that frequency band; The third determining unit is used to determine the fusion frequency band and the non-fusion frequency band based on the smooth cross-correlation coefficient, wherein the fusion frequency band refers to the frequency band that needs to be fused, and the non-fusion frequency band refers to the frequency band that does not need to be fused. The fusion unit is used to perform fusion calculations on the fusion frequency band according to a preset fusion algorithm to obtain the fused signal.

9. The apparatus as claimed in claim 8, characterized in that, The fusion unit is specifically used for, Noise estimation is performed on the bone conduction target event signal segment and the air conduction target event signal segment respectively to obtain the bone conduction signal-to-noise ratio and the air conduction signal-to-noise ratio; The fusion coefficient is determined based on the bone conduction signal-to-noise ratio and the air conduction signal-to-noise ratio. The fusion calculation is performed based on the fusion coefficient and the preset fusion algorithm to obtain the fused signal.

10. The apparatus as claimed in claim 8, characterized in that, The second determining unit is specifically used for, The bone conduction target event signal segment and the air conduction target event signal segment corresponding to each frequency band are determined respectively; Determine the Fourier transform function, inverse Fourier transform function, and conjugate function of the Fourier transform corresponding to each frequency band; Based on the preset window function, the Fourier transform function, the inverse Fourier transform function, and the conjugate function of the Fourier transform, the smooth cross-correlation coefficient between the bone conduction target event signal segment and the air conduction target event signal segment corresponding to each frequency band is calculated.

11. The apparatus as claimed in claim 8, characterized in that, The third determining unit is specifically used for, If the smooth cross-correlation coefficient of any frequency band among the multiple frequency bands is greater than or equal to a preset threshold, then the frequency band is determined to be a fused frequency band. If the smooth cross-correlation coefficient of any of the multiple frequency bands is less than the preset threshold, then the frequency band is determined to be a non-fusion frequency band.

12. The apparatus as claimed in claim 9, characterized in that, The fusion unit is specifically used for, Based on the signal-to-noise ratios of the air conduction target event signal segment and the bone conduction target event signal segment in the multiple frequency bands and the function for determining the fusion coefficient, the fusion coefficients of the air conduction target event signal segment and the bone conduction target event signal segment are calculated respectively.

13. The apparatus as claimed in claim 12, characterized in that, The fusion coefficients of the air conduction target event signal and the bone conduction target event signal segments are calculated using the following formulas. , ,in The fusion coefficient is the signal segment of the gas-conducted target event. The fusion coefficient is the value of the bone conduction target event signal segment. It is the function that determines the fusion coefficient. The gas conduction signal-to-noise ratio is denoted as . The bone conduction signal-to-noise ratio is denoted as .

14. The apparatus as claimed in claim 10, characterized in that, The smooth cross-correlation coefficient is calculated using the following formula: ,in, Represents the smooth cross-correlation coefficient. This represents the inverse Fourier transform function. This represents the Fourier transform function. Denotes the conjugate function of the Fourier transform. This represents the preset window function. This refers to the gas conduction target event signal segment. This refers to the bone conduction target event signal segment.

15. A voice noise reduction device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-7.

17. A computer program product comprising a computer program that, when executed by a processor, implements the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Systems and methods for audio signal generation

    US20220150627A1