Speech noise reduction method, device, electronic device, and computer-readable storage medium
By determining and correcting the evaluation gain coefficient of the speech signal, and filtering the speech signal using the cepspectral correlation coefficient, the problem of performance degradation under low signal-to-noise ratio conditions in the prior art is solved, and the effect of retaining the details of the speech signal while decreasing noise is achieved.
Patent Information
- Application Number
- CN202111624173.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-28
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2041-12-28
AI Technical Summary
The existing dual-microphone voice enhancement algorithm has severe performance degraded under low signal-to-noise ratio conditions, and is prone to music noise, making it difficult to retain details of the voice signal while decreasing noise.
By determining the evaluation gain coefficients of the first voice signal and the second voice signal, the evaluation gain coefficient is corrected by using the cepspectral correlation coefficient, the filtering method is used, and the gain K(ω,k) is corrected by combining the cepspectral domain correlation coefficient CCC, the speech and noise frames are filtered and different smoothing schemes are adopted.
While reducing noise, it retains more details of the voice signal and has a certain dereverberation effect, improving the performance of voice enhancement.
Smart Images

Figure CN114495961B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech noise reduction, and in particular to a speech noise reduction method, device, electronic device, and computer-readable storage medium. Background Art
[0002] Speech enhancement algorithms, as a key component of speech signal processing technology, are widely used in fields such as mobile phone calls, simultaneous interpretation, and speech recognition. Speech enhancement technology primarily suppresses noise components in speech to achieve the desired noise reduction goal. Initially, single-microphone-based speech noise reduction algorithms offered good noise suppression for stationary noise, but their performance was significantly limited by their inability to fully utilize the spatial information of the speech signal. This led to the development of dual-microphone speech enhancement algorithms based on microphone arrays. These arrays use multiple microphones to simultaneously receive spatial speech signals and perform positional analysis to accurately extract the target speech signal, thereby improving speech enhancement performance. Furthermore, due to its relatively low implementation cost and simple, easy-to-develop algorithmic foundation, dual-microphone speech enhancement technology has gradually become a major research direction in multichannel speech signal processing.
[0003] Dual-microphone speech enhancement algorithms can spatially localize speech signals, enabling the reception of target speech and filtering of noise. Existing dual-microphone speech enhancement algorithms are mostly based on the amplitude spectrum coherence function. However, these algorithms suffer from severe performance degradation under low signal-to-noise ratio conditions and are prone to musical noise. Summary of the Invention
[0004] The present invention provides a speech noise reduction method, which can retain more details of the speech signal while reducing noise and has a certain dereverberation effect.
[0005] To solve the above technical problems, the first technical solution provided by the present invention is: providing a speech noise reduction method, including: determining an evaluation gain coefficient of a first speech signal and a second speech signal, wherein the first speech signal and the second speech signal are obtained based on dual microphones; determining a cepstral correlation coefficient of the first speech signal and the second speech signal, and using the cepstral correlation coefficient to correct the evaluation gain coefficient to obtain a corrected gain coefficient; and using the corrected gain coefficient to filter the first speech signal and the second speech signal.
[0006] To solve the above technical problems, the second technical solution provided by the present invention is: to provide a speech noise reduction device, including: a determination module, used to determine the evaluation gain coefficient of the first speech signal and the second speech signal, the first speech signal and the second speech signal are obtained based on dual microphones; a correction module, used to determine the cepstral correlation coefficient of the first speech signal and the second speech signal, and use the cepstral correlation coefficient to correct the evaluation gain coefficient to obtain a corrected gain coefficient; a filtering module, used to use the corrected gain coefficient to filter the first speech signal and the second speech signal.
[0007] To solve the above technical problems, the third technical solution provided by the present invention is: to provide an electronic device, including a processor and a memory coupled to each other, wherein the memory is used to store program instructions for implementing any of the above methods; and the processor is used to execute the program instructions stored in the memory.
[0008] In order to solve the above technical problems, the fourth technical solution provided by the present invention is: providing a computer-readable storage medium storing a program file, which can be executed to implement any of the above methods.
[0009] The present invention has the beneficial effects of distinguishing itself from existing technologies by determining evaluation gain coefficients for a first speech signal and a second speech signal, the first and second speech signals being acquired using dual microphones; determining cepstral correlation coefficients between the first and second speech signals, using the cepstral correlation coefficients to modify the evaluation gain coefficients to obtain modified gain coefficients; and filtering the first and second speech signals using the modified gain coefficients. This method can retain more details of the speech signal while reducing noise and has a certain dereverberation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive work, among which:
[0011] Figure 1 A flow chart of an embodiment of a method for reducing speech noise according to the present invention;
[0012] Figure 2 is a schematic diagram of a dual-microphone array;
[0013] Figure 3 for Figure 1 A flow chart of an embodiment of step S11;
[0014] Figure 4 for Figure 3 A flow chart of an embodiment of step S32;
[0015] Figure 5 for Figure 4 A flow chart of an embodiment of step S43;
[0016] Figure 6 for Figure 1 A flow chart of an embodiment of step S12;
[0017] Figure 7 for Figure 1 A flow chart of an embodiment of step S12;
[0018] Figure 8 Schematic diagram of the structure of a speech noise reduction device according to an embodiment of the present invention;
[0019] Figure 9 It is a structural schematic diagram of an embodiment of an electronic device of the present invention;
[0020] Figure 10 A schematic structural diagram of a computer-readable storage medium according to the present invention. Specific implementation methods
[0021] The terms "first," "second," and "third" in this application are used only for descriptive purposes and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of such features. In the description of this application, "multiple" means at least two, for example, two, three, etc., unless otherwise specifically defined. All directional indications in the embodiments of this application (such as up, down, left, right, front, back...) are only used to explain the relative positional relationship, movement, etc. between the components under a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications also change accordingly. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products, or devices.
[0022] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0023] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0024] See Figure 1 , which is a flow chart of an embodiment of a speech noise reduction method of the present invention, specifically comprising:
[0025] Step S11: Determine evaluation gain coefficients of the first speech signal and the second speech signal.
[0026] Specifically, the first speech signal and the second speech signal are obtained by two microphones in a microphone array, such as Figure 2 As shown, Figure 2 This is a schematic diagram of a dual-microphone array. It can be understood that the first voice signal can be a voice signal obtained by microphone 1, the second voice signal can be a voice signal obtained by microphone 2, and the distance between microphone 1 and microphone 2 is d.
[0027] In this embodiment, the evaluation gain coefficients of the first speech signal and the second speech signal are determined. Figure 3 , step S11 includes:
[0028] Step S31: converting the first speech signal into a first frequency domain signal, and converting the second speech signal into a second frequency domain signal.
[0029] Speech signals are time-varying, non-stationary signals. Generally, they are characterized by relative stability within a time range of 10-30ms. Therefore, to obtain a stable speech signal, the speech signal needs to be framed. Specifically, the first and second speech signals are framed separately to produce multiple first and second speech segments. To enhance the coherence between the two frames of speech data, the framing method employed is overlapping segmentation. The overlap between frames is called a frame shift, and the frame shift value is half the frame length. To further enhance the stationarity of the speech data, the framed speech segments need to be windowed. Specifically, the first and second speech segments are windowed separately to produce first and second speech data to be processed. Common window functions include rectangular windows, Hanning windows, Hamming windows, and triangular windows. The Hamming window has a relatively smooth low-pass characteristic and is well suited to reflecting the frequency characteristics of short-duration signals such as speech. The Hamming window is the window type chosen here.
[0030] The microphone collects a real signal, which needs to be converted into a frequency domain signal. Specifically, the first speech data to be processed is subjected to Fourier transform, and then the first speech signal is converted into a first frequency domain signal. The second speech data to be processed is subjected to Fourier transform, and then the second speech signal is converted into a second frequency domain signal. In a specific embodiment, assuming that the number of Fourier transform points is N fft , only N fft / 2 points of data have practical significance, the remaining N fft / 2 points are symmetrical data, and each frequency point can be considered as a narrowband signal. Therefore, the first N frequency points after the first and second voice signals are transformed are selected. fft / 2 point frequency domain data is used as the first frequency domain signal and the second frequency domain signal.
[0031] Step S32: determining the evaluation gain coefficient based on the first frequency domain signal and the second frequency domain signal.
[0032] For details, please combine Figure 4 , step S32 includes:
[0033] Step S41: determining the autopower spectral density of the first speech signal based on the first frequency domain signal, and determining the autopower spectral density of the second speech signal based on the second frequency domain signal.
[0034] The autopower spectral density of the first speech signal is determined based on the first frequency domain signal. The autopower spectral density of the first speech signal is an energy spectrum obtained by square the first frequency domain signal of the first speech signal, and then taking the average. Assume that the first frequency domain signal is represented by Y1(ω, k), and the autopower spectral density of the first speech signal is represented by P y1y1=E[|Y1(ω,k)| 2 ]. Wherein, ω represents the angular frequency, and k represents the number of frames. If the current first frequency domain signal is the third frame, then k=3, that is, the first frequency domain signal of the third frame is represented as Y1(ω,3). In a specific embodiment, the sampling point can be pre-set, and when the speech signal is acquired, or when the pre-set sampling point is met, the currently acquired speech signal is defined as a frame, and the frame speech signal is windowed and Fourier transformed.
[0035] The autopower spectral density of the second speech signal is determined based on the second frequency domain signal. The autopower spectral density of the second speech signal is an energy spectrum obtained by square the second frequency domain signal of the second speech signal, and then taking the average. Assume that the second frequency domain signal is represented by Y2(ω, k), and the autopower spectral density of the second speech signal is represented by P y2y2 =E[|Y2(ω,k)| 2 ]. Wherein, ω represents the angular frequency, and k represents the number of frames. For example, after the second speech signal is framed and processed, 20 frames of second speech segments are obtained. Then, 20 frames of second speech segments are windowed to obtain 20 frames of second speech data to be processed. Fourier transform is performed on these 20 frames of second speech data to be processed, and then 20 frames of second frequency domain signals are obtained. If the current second frequency domain signal is the third frame, then k=3, that is, the third frame of the second frequency domain signal is expressed as Y2(ω,3).
[0036] Step S42: Determine a cross-power spectral density of the first speech signal and the second speech signal based on the first frequency domain signal and the second frequency domain signal.
[0037] The cross power spectrum density is obtained by taking the average of the conjugate product of the first frequency domain signal and the second frequency domain signal. The cross power spectrum density is expressed as:
[0038] P y1y2 =E[Y1(ω,k)Y2 * (ω, k)].
[0039] Step S43: Determine the evaluation gain coefficient based on the cross power spectrum density, the auto power spectrum density of the first speech signal, and the auto power spectrum density of the second speech signal.
[0040] Specifically, in the process of calculating the evaluation gain coefficient in this embodiment, the calculation can be performed by estimating the coherent diffusion power ratio (CDR). Figure 5 , step S43 includes:
[0041] Step S51: determining a coherence function between the first speech signal and the second speech signal based on the cross power spectral density, the auto power spectral density of the first speech signal, and the auto power spectral density of the second speech signal.
[0042] Specifically, the autopower spectral density and cross-power spectral density in the negative domain are obtained for the first speech signal and the second speech signal, and then the coherence function of the first speech signal and the second speech signal is calculated. Specifically, the coherence function of the first speech signal and the second speech signal is:
[0043]
[0044] Step S52: determining a coherent spread power ratio based on the coherence function and a preset coherence function.
[0045] In the scattering scenario, the default coherence function is:
[0046]
[0047] Where ω represents the angular frequency of the speech signal, f s represents the sampling frequency, d represents the distance between the two microphones, and c represents the propagation speed of the speech signal.
[0048] Furthermore, the coherent spread power ratio is determined based on the coherence function, ie, the above formula (1), and the preset coherence function, ie, the above formula (2).
[0049] In one embodiment, the coherent spreading power ratio is expressed as:
[0050]
[0051] Wherein, Δt is the time delay difference between the first voice signal and the second voice signal. When the voice signal source is located in the vertical position direction of the dual microphones, that is, at a 90° direction of the dual microphones, Δt=0.
[0052] Step S53: determining the evaluation gain coefficient based on the coherent spread power ratio.
[0053] The evaluation gain coefficient is determined based on the coherent spread power ratio. It should be noted that the coherent spread power ratio CDR can be considered equivalent to the signal-to-noise ratio in a certain sense. Based on the classical Wiener filtering method, the evaluation gain coefficient can be expressed as:
[0054]
[0055] Step S12: determining the cepstral correlation coefficient between the first speech signal and the second speech signal, and using the cepstral correlation coefficient to correct the evaluation gain coefficient to obtain a corrected gain coefficient.
[0056] For details, please combine Figure 6 The step of determining the cepstral correlation coefficient of the first speech signal and the second speech signal in step S12 includes:
[0057] Step S61: converting the first frequency domain signal into a first cepstrum domain signal, and converting the second frequency domain signal into a second cepstrum domain signal.
[0058] Convert the first frequency domain signal Y1(ω, k) into the first cepstrum domain signal c y1 (q, k); and converting the second frequency domain signal Y2 (ω, k) into a second cepstrum domain signal c y2 (q, k).
[0059] Specifically, in one embodiment, the first frequency domain signal Y1(ω, k) is converted into a first cepstrum domain signal c y1 (q, k) can be expressed as:
[0060] C y1 (q, k) = IDFT(log|Y1(ω, k)| 2 );
[0061] Wherein, IDFT represents inverse Fourier transform, specifically, DFT represents discrete Fourier transform, and IDFT represents inverse discrete Fourier transform, and q represents an index value of the cepstrum domain.
[0062] It can be understood that the second frequency domain signal Y2(ω, k) is converted into the second cepstrum domain signal c y2 (q, k) can be expressed as:
[0063] C y2 (q, k) = IDFT(log|Y2(ω, k)| 2 ).
[0064] Step S62: determining an autocorrelation coefficient of the first speech signal based on the first cepstral domain signal, and determining an autocorrelation coefficient of the second speech signal based on the second cepstral domain signal.
[0065] In one embodiment, it is also necessary to obtain the first cepstrum domain signals c from all y1 (q, k) is searched for the maximum value. Specifically, based on the preset fundamental frequency range, the first cepstral domain signal with the largest index value is searched from the first cepstral domain signal. It can be expressed as:
[0066] q max1 =arg q max{C y1 (q, k), q∈[fs / 300, fs / 70]};
[0067] Among them, q max1 represents the first cepstrum domain signal with the largest index value, [fs / 300, fs / 70] represents a preset sampling frequency, where fs represents the sampling frequency.
[0068] It is understandable that it is also necessary to search for the second cepstral domain signal with the largest index value from the second cepstral domain signal based on the preset fundamental frequency range.
[0069] q max2 =arg q max{C y2 (q, k), q∈[fs / 300, fs / 70]}.
[0070] The first cepstral domain smoothing coefficient of the first cepstral domain signal is determined based on the first cepstral domain signal with the largest index value. Specifically, assuming that the first cepstral domain signal with the largest index value is represented by C y1 (q max1 ,k), compare it with the threshold Cth and VAD (VAD indicates the presence of speech) to determine the first cepstral domain smoothing coefficient. Specifically, the first cepstral domain smoothing coefficient can be expressed as:
[0071]
[0072] Specifically, when there is speech, and the value of the first cepstral domain signal of the current frame is greater than the threshold Cth, and the index value of the first cepstral domain signal of the current frame satisfies q∈{q max1 -1,q max1 ,q max1 +1}, the first cepstral domain smoothing coefficient β cc1 =β p1 , otherwise, the first cepstral domain smoothing coefficient β cc1 =β1.
[0073] The second cepstral domain smoothing coefficient of the second cepstral domain signal is determined based on the second cepstral domain signal with the largest index value. Specifically, assuming that the second cepstral domain signal with the largest index value is represented by C y2 (q max2 ,k), compare it with the threshold Cth and VAD (VAD indicates the presence of speech) to determine the second cepstral domain smoothing coefficient. Specifically, the second cepstral domain smoothing coefficient can be expressed as:
[0074]
[0075] Specifically, when speech exists and the value of the second cepstral domain signal of the current frame is greater than the threshold Cth, and the index value of the second cepstral domain signal of the current frame satisfies q∈{qmax2 -1,q max2 ,q max2 +1}, the second cepstral domain smoothing coefficient β cc2 =β p2 , otherwise, the second cepstral domain smoothing coefficient β cc2 =β2.
[0076] In a specific embodiment, different smoothing coefficients are used for the fundamental tone portion and the non-fundamental tone portion. The fundamental tone portion can use a relatively small value, while the non-fundamental tone portion can use a relatively large value. It is understandable that the fundamental tone portion of the current frame can be understood as the previous frame. That is, the first cepstral domain signal of the previous frame and the first cepstral domain signal of the current frame are processed based on the first cepstral domain smoothing coefficient to obtain a smoothed first cepstral domain signal; and the second cepstral domain signal of the previous frame and the second cepstral domain signal of the current frame are processed based on the second cepstral domain smoothing coefficient to obtain a smoothed second cepstral domain signal.
[0077] The smoothed first cepstrum domain signal can be specifically expressed as:
[0078] Among them, c y1 (q, k-1) is the first cepstrum domain signal of the previous frame, c y1 (q, k) is the first cepstrum domain signal of the current frame.
[0079] The smoothed second cepstrum domain signal can be specifically expressed as:
[0080] Among them, c y2 (q, k-1) is the second cepstrum domain signal of the previous frame, c y2 (q, k) is the cepstrum domain signal of the current second frame.
[0081] The smoothed first cepstral domain signal and the smoothed second cepstral domain signal are obtained in the above manner.
[0082] An autocorrelation coefficient of the first speech signal is determined based on the smoothed first cepstral domain signal, and an autocorrelation coefficient of the second speech signal is determined based on the smoothed second cepstral domain signal.
[0083] Specifically, the autocorrelation coefficient of the first speech signal is expressed as:
[0084]
[0085] The autocorrelation coefficient of the second speech signal is:
[0086]
[0087] Step S63: determining a cross-correlation coefficient between the first speech signal and the second speech signal based on the first cepstral domain signal and the second cepstral domain signal.
[0088] Specifically, the cross-correlation coefficient between the first speech signal and the second speech signal is determined based on the smoothed first cepstral domain signal and the smoothed second cepstral domain signal.
[0089] The mutual correlation coefficient between the first speech signal and the second speech signal is expressed as:
[0090]
[0091] Step S64: Determine the cepstral correlation coefficient based on the cross-correlation coefficient, the autocorrelation coefficient of the first speech signal, and the autocorrelation coefficient of the second speech signal.
[0092] Specifically, the cepstral correlation coefficient can be expressed as:
[0093]
[0094] After obtaining the cepstral domain correlation coefficient, the evaluation gain coefficient is corrected using the cepstral correlation coefficient to obtain a corrected gain coefficient.
[0095] For details, please combine Figure 7 , step S12 further includes:
[0096] Step S71: converting the evaluation gain coefficient into a first cepstrum domain gain coefficient.
[0097] Specifically, the evaluation gain coefficient K(ω, k) is converted into a first cepstrum domain gain coefficient, specifically: Ck(q, k)=IDFT(K(ω, k)).
[0098] Step S72: Process the first cepstral domain gain coefficient using the cepstral correlation coefficient to obtain a second cepstral domain gain coefficient.
[0099] Specifically expressed as: C ks (q,k)=C k (q,k)r c (q,k).
[0100] Step S73: Performing Fourier transform on the second cepstral domain gain coefficient, and selecting the one with the largest cepstral domain value as the modified gain coefficient.
[0101] Specifically expressed as: Kccc(ω,k)=max(DFT(c ks (q,k)),0)
[0102] Step S13: Filtering the first speech signal and the second speech signal using the modified gain coefficient.
[0103] Specifically, after the correction gain coefficient is determined, the first speech signal and the second speech signal are filtered using the correction gain coefficient.
[0104] Specifically, the first speech signal is subjected to weighted filtering using the modified gain coefficient. Specifically expressed as
[0105] The second speech signal is subjected to weighted filtering processing using the modified gain coefficient.
[0106] Specifically expressed as
[0107] In one embodiment, the first speech signal and the second speech signal after filtering are further subjected to inverse Fourier transform. The inverse Fourier transform of the first speech signal after filtering is expressed as: t represents the time domain.
[0108] The inverse Fourier transform of the second speech signal after filtering is expressed as:
[0109] Existing methods are prone to damaging speech while simultaneously reducing noise. This solution fully utilizes useful information in both the dual-channel frequency domain and the cepstral domain, using the cepstral domain correlation coefficient (CCC) to modify the gain K(ω,k), thereby achieving both noise reduction and the preservation of more speech details. 2. In the cepstral domain correlation coefficient estimation, this solution combines harmonic energy with VAD to filter speech and noise frames. Different smoothing schemes are applied to each frame, further highlighting the correlation of speech components in the cepstral domain.
[0110] See Figure 8 , is a structural diagram of an embodiment of a speech noise reduction device of the present invention, which specifically includes a determination module 101, a correction module 102 and a filtering module 103.
[0111] The determination module 101 is configured to determine evaluation gain coefficients for a first speech signal and a second speech signal, where the first speech signal and the second speech signal are acquired based on dual microphones. The correction module 102 is configured to determine cepstral correlation coefficients between the first speech signal and the second speech signal, and to modify the evaluation gain coefficients using the cepstral correlation coefficients to obtain modified gain coefficients. The filtering module 103 is configured to filter the first speech signal and the second speech signal using the modified gain coefficients.
[0112] Existing devices are prone to damaging speech while simultaneously reducing noise. This solution fully utilizes useful information in both the dual-channel frequency domain and the cepstral domain, using the cepstral domain correlation coefficient (CCC) to modify the gain K(ω,k), thereby achieving both noise reduction and the preservation of more speech details. 2. In the cepstral domain correlation coefficient estimation, this solution combines harmonic energy with VAD to filter speech and noise frames. Different smoothing schemes are applied to each frame, further highlighting the correlation of speech components in the cepstral domain.
[0113] See Figure 9 , which is a schematic structural diagram of an electronic device according to an embodiment of the present invention. The electronic device includes a memory 82 and a processor 81 connected to each other.
[0114] The memory 82 is used to store program instructions for implementing any one of the above methods.
[0115] The processor 81 is configured to execute program instructions stored in the memory 82 .
[0116] The processor 81 may also be referred to as a CPU (Central Processing Unit). The processor 81 may be an integrated circuit chip having signal processing capabilities. The processor 81 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The general-purpose processor may be a microprocessor or any conventional processor.
[0117] The memory 82 can be a memory stick, a TF card, etc., which can store all the information in the electronic device, including the input raw data, computer programs, intermediate operation results and final operation results. It is stored in the memory. It stores and retrieves information according to the location specified by the controller. Only with the memory can the electronic device have a memory function and ensure normal operation. The memory of the electronic device can be divided into main memory (internal memory) and auxiliary memory (external memory) according to its purpose. There is also a classification method of dividing it into external memory and internal memory. External memory is usually a magnetic medium or an optical disk, etc., which can store information for a long time. Memory refers to the storage component on the motherboard, which is used to store the data and programs currently being executed, but is only used to temporarily store programs and data. If the power is turned off or the power is cut off, the data will be lost.
[0118] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented by other methods. For example, the device implementation method described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0119] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, i.e., they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the objectives of this embodiment as needed.
[0120] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0121] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, system server, or network device, etc.) or a processor to execute all or part of the steps of each implementation method of the present application.
[0122] See also Figure 10, which is a structural diagram of the computer-readable storage medium of the present invention. The storage medium of the present application stores a program file 91 that can implement all the above methods, wherein the program file 91 can be stored in the above storage medium in the form of a software product, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each implementation method of the present application. The aforementioned storage device includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or terminal devices such as a computer, a server, a mobile phone, and a tablet.
[0123] The above is only an implementation method of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A speech noise reduction method, characterized in that: include: Determining evaluation gain coefficients of a first speech signal and a second speech signal based on a coherent spread power ratio and a Wiener filtering method, wherein the first speech signal and the second speech signal are acquired based on dual microphones; Determine the cepstral correlation coefficient of the first speech signal and the second speech signal, and use the cepstral correlation coefficient to correct the evaluation gain coefficient to obtain a corrected gain coefficient; the step of correcting the evaluation gain coefficient using the cepstral correlation coefficient to obtain the corrected gain coefficient includes: converting the evaluation gain coefficient into a first cepstral domain gain coefficient; processing the first cepstral domain gain coefficient using the cepstral correlation coefficient to obtain a second cepstral domain gain coefficient; performing Fourier transform on the second cepstral domain gain coefficient, and selecting the one with the largest cepstral domain value as the corrected gain coefficient; The first speech signal and the second speech signal are filtered using the modified gain coefficient.
2. The method according to claim 1, characterized in that The step of determining the evaluation gain coefficients of the first speech signal and the second speech signal comprises: Converting the first speech signal into a first frequency domain signal, and converting the second speech signal into a second frequency domain signal; The evaluation gain coefficient is determined based on the first frequency domain signal and the second frequency domain signal.
3. The method according to claim 2, characterized in that The steps of converting the first speech signal into a first frequency domain signal and converting the second speech signal into a second frequency domain signal include: Performing frame processing on the first voice signal and the second voice signal respectively to obtain a plurality of first voice segments and a plurality of second voice segments; Performing windowing processing on the first voice segment and the second voice segment respectively to obtain first voice data to be processed and second voice data to be processed; The first speech data to be processed is subjected to a Fourier transform, thereby converting the first speech signal into a first frequency domain signal; the second speech data to be processed is subjected to a Fourier transform, thereby converting the second speech signal into a second frequency domain signal.
4. The method according to claim 2, characterized in that The step of determining the evaluation gain coefficient based on the first frequency domain signal and the second frequency domain signal includes: Determining an autopower spectral density of the first speech signal based on the first frequency domain signal, and determining an autopower spectral density of the second speech signal based on the second frequency domain signal; Determine a cross power spectral density of the first speech signal and the second speech signal based on the first frequency domain signal and the second frequency domain signal; The evaluation gain coefficient is determined based on the cross power spectrum density, the auto power spectrum density of the first speech signal, and the auto power spectrum density of the second speech signal.
5. The method according to claim 4, characterized in that The step of determining the evaluation gain coefficient based on the cross power spectrum density, the auto power spectrum density of the first speech signal, and the auto power spectrum density of the second speech signal comprises: Determining a coherence function between the first speech signal and the second speech signal based on the cross power spectral density, the auto power spectral density of the first speech signal, and the auto power spectral density of the second speech signal; Determining the coherent spread power ratio based on the coherence function and a preset coherence function; The evaluation gain coefficient is determined based on the coherent spread power ratio.
6. The method according to claim 2, characterized in that The step of determining the cepstral correlation coefficient of the first speech signal and the second speech signal comprises: Converting the first frequency domain signal into a first cepstral domain signal, and converting the second frequency domain signal into a second cepstral domain signal; Determining an autocorrelation coefficient of the first speech signal based on the first cepstral domain signal, and determining an autocorrelation coefficient of the second speech signal based on the second cepstral domain signal; determining a cross-correlation coefficient between the first speech signal and the second speech signal based on the first cepstral domain signal and the second cepstral domain signal; The cepstral correlation coefficient is determined based on the cross-correlation coefficient, the autocorrelation coefficient of the first speech signal, and the autocorrelation coefficient of the second speech signal.
7. The method according to claim 6, characterized in that The steps of determining the autocorrelation coefficient of the first speech signal based on the first cepstral domain signal, and determining the autocorrelation coefficient of the second speech signal based on the second cepstral domain signal, include: Based on a preset pitch frequency range, searching the first cepstral domain signal for the first cepstral domain signal having the largest index value, and based on the preset pitch frequency range, searching the second cepstral domain signal for the second cepstral domain signal having the largest index value; Determine a first cepstral domain smoothing coefficient of a first cepstral domain signal based on the first cepstral domain signal with the largest index value, and determine a second cepstral domain smoothing coefficient of a second cepstral domain signal based on the second cepstral domain signal with the largest index value; processing the first cepstral domain signal of the previous frame and the first cepstral domain signal of the current frame based on the first cepstral domain smoothing coefficient to obtain a smoothed first cepstral domain signal; and processing the second cepstral domain signal of the previous frame and the second cepstral domain signal of the current frame based on the second cepstral domain smoothing coefficient to obtain a smoothed second cepstral domain signal; Determining an autocorrelation coefficient of the first speech signal based on the smoothed first cepstral domain signal, and determining an autocorrelation coefficient of the second speech signal based on the smoothed second cepstral domain signal; The step of determining the cross-correlation coefficient between the first speech signal and the second speech signal based on the first cepstral domain signal and the second cepstral domain signal comprises: A cross-correlation coefficient between the first speech signal and the second speech signal is determined based on the smoothed first cepstral domain signal and the smoothed second cepstral domain signal.
8. The method according to claim 1, characterized in that The step of filtering the first speech signal and the second speech signal by using the modified gain coefficient includes: The first speech signal is subjected to weighted filtering processing using the modified gain coefficient; and the second speech signal is subjected to weighted filtering processing using the modified gain coefficient.
9. The method according to claim 1, characterized in that The method further comprises: Perform inverse Fourier transform on the first speech signal and the second speech signal after filtering.
10. A speech noise reduction device, characterized in that: include: a determination module, configured to determine evaluation gain coefficients of a first speech signal and a second speech signal based on a coherent spread power ratio and a Wiener filtering method, wherein the first speech signal and the second speech signal are acquired based on dual microphones; A correction module is used to determine the cepstral correlation coefficient of the first speech signal and the second speech signal, and use the cepstral correlation coefficient to correct the evaluation gain coefficient to obtain a corrected gain coefficient; the step of correcting the evaluation gain coefficient using the cepstral correlation coefficient to obtain the corrected gain coefficient includes: converting the evaluation gain coefficient into a first cepstral domain gain coefficient; processing the first cepstral domain gain coefficient using the cepstral correlation coefficient to obtain a second cepstral domain gain coefficient; performing Fourier transform on the second cepstral domain gain coefficient, and selecting the one with the largest cepstral domain value as the corrected gain coefficient; A filtering module is used to perform filtering processing on the first speech signal and the second speech signal using the modified gain coefficient.
11. An electronic device, characterized in that: It includes a processor and a memory coupled to each other, wherein: The memory is used to store program instructions for implementing the method according to any one of claims 1 to 9; The processor is configured to execute the program instructions stored in the memory.
12. A computer-readable storage medium, characterized in that A program file is stored, and the program file can be executed to implement the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Multi-channel speech enhancement method and device, terminal and readable storage medium
CN113689870A
Multi-microphone noise reduction method, apparatus and terminal device
WO2019112468A1