A method and device for speech enhancement
By processing far-field speech signals through microphone arrays and multi-level enhancement models, the problems of noise and reverberation in far-field environments are solved, and the accuracy of the speech recognition system is improved.
Patent Information
- Application Number
- CN202111272002.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2041-10-29
AI Technical Summary
In far-field environments, noise and reverberation have a serious impact on speech recognition systems, resulting in a decrease in recognition accuracy and hindering the further promotion of speech recognition technology.
A microphone array is used to collect multi-channel far-field voice signals. Through multi-level enhancement models and filtering technology, sound signals from other sound sources are filtered out to extract pure or relatively pure target sound source voice signals.
It effectively removes noise and reverberation from far-field speech signals, improves the accuracy of the speech recognition system, and enhances the performance of speech recognition technology in far-field environments.
Smart Images

Figure CN116072139B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech enhancement technology, and in particular to a speech enhancement method and device. Background Art
[0002] Language communication is one of the most natural ways of communication for humans. Human research on computer speech covers speech encoding and decoding, speech recognition, speech synthesis, speaker recognition, activation words, speech enhancement, etc.
[0003] In recent years, with the widespread application of deep learning algorithms in speech recognition, speech recognition technology has achieved human-level performance in relatively clean environments. This unprecedented progress has also led to the development of applications in various fields, such as voice assistants and smart speakers. As human-computer interaction scenarios shift from near-field to far-field, people's speech is subject to interference from various external factors, such as ambient noise and far-field reverberation. The combination of noise and speech reverberation in far-field environments can affect hearing and negatively impact speech recognition systems, significantly reducing recognition accuracy and hindering the further promotion of speech recognition technology. Summary of the Invention
[0004] The present invention provides a speech enhancement method, which performs multi-level enhancement on the multi-channel far-field speech signals collected by a microphone array, filters out the sound signals of other sound sources, and obtains a pure or relatively pure speech signal of the target sound source, so as to solve the problems of large reverberation and noise when far-field speech is received.
[0005] In a first aspect, the present invention provides a method for speech enhancement, the method comprising:
[0006] Acquiring a multi-channel far-field speech signal collected by a microphone array, wherein the far-field speech signal includes at least a speech signal of a target sound source and sound signals of other sound sources;
[0007] Extracting the short-time Fourier spectrum of the multi-channel far-field speech signal;
[0008] Inputting the short-time Fourier spectrum of the multi-channel far-field speech signal into the first-stage enhancement model to perform preliminary speech enhancement and determine a first enhanced speech spectrum;
[0009] performing a first linear filtering on the first enhanced speech spectrum to determine a filtered first enhanced speech spectrum;
[0010] Inputting the short-time Fourier spectrum of the multi-channel far-field speech signal and the filtered first enhanced speech spectrum into a second-stage enhancement model for further speech enhancement to determine a second enhanced speech spectrum;
[0011] performing a second linear filtering on the second enhanced speech spectrum to determine a filtered second enhanced speech spectrum;
[0012] Based on the filtered second enhanced speech spectrum, a speech signal of the target sound source is determined.
[0013] In one possible implementation, extracting the short-time Fourier spectrum of the multi-channel far-field speech signal includes:
[0014] Extracting short-time Fourier features of the far-field speech signal of each channel from the far-field speech signals of the multiple channels respectively;
[0015] The short-time Fourier features of the far-field speech signals of each channel in the far-field speech signals of the multiple channels are fused to determine the short-time Fourier spectra of the far-field speech signals of the multiple channels.
[0016] In another possible implementation, the first-stage enhancement model is trained based on a first training set, where the first training set includes multiple first training sample pairs, each of which includes a far-field speech signal and a near-field speech signal corresponding to the far-field speech signal.
[0017] In another possible implementation, performing a first linear filtering on the first enhanced speech spectrum to determine a filtered first enhanced speech spectrum includes:
[0018] The first enhanced speech spectrum is input into a first linear filter to determine the filtered first enhanced speech spectrum, wherein a filter coefficient of the first linear filter is determined based on a short-time Fourier spectrum of the multi-channel speech signal and the first enhanced speech spectrum.
[0019] In another possible implementation, the second-stage enhancement model is trained based on a second training set, and the second training set includes multiple second training sample pairs, one sample of each second training sample pair includes the short-time Fourier spectrum of the multi-channel speech signal corresponding to the far-field speech and the first enhanced speech spectrum, and the other sample includes the near-field speech signal corresponding to the far-field speech.
[0020] In another possible implementation, performing a second linear filtering on the second enhanced speech spectrum to determine a filtered second enhanced speech spectrum includes:
[0021] The first enhanced speech spectrum is input into a second linear filter to determine the filtered second enhanced speech spectrum, wherein the filter coefficient of the second linear filter is determined based on the short-time Fourier spectrum of the multi-channel speech signal and the second enhanced speech spectrum.
[0022] In another possible implementation, determining the target speech signal based on the filtered second enhanced speech spectrum includes:
[0023] Inputting the filtered second enhanced speech spectrum and the short-time Fourier spectrum of the multi-channel speech signal into the second-stage enhanced speech model to determine a third enhanced speech spectrum;
[0024] performing a second linear filtering on the third enhanced speech spectrum to determine a filtered third enhanced speech spectrum;
[0025] Based on the filtered third enhanced speech spectrum, a speech signal of a target sound source is determined.
[0026] In another possible implementation, determining the target speech signal based on the filtered second enhanced speech spectrum includes:
[0027] Performing an inverse Fourier transform on the filtered second enhanced speech spectrum to determine the speech signal of the target sound source.
[0028] In another possible implementation, the first-stage speech enhancement model and the second-stage speech enhancement model both include a bidirectional long short-term memory network layer, a random dropout layer, and a linear layer.
[0029] In a second aspect, the present invention provides a computing device comprising a memory and a processor, wherein the memory stores executable code, and the processor executes the executable code to perform the method of the first aspect.
[0030] In a third aspect, the present invention further provides a storage medium comprising: a readable storage medium and a computer program stored in the readable storage medium, wherein the computer program is used to implement the method of the first aspect.
[0031] The speech enhancement method provided by the present invention performs multi-level enhancement on the multi-channel far-field speech signals collected by the microphone array, filters out the sound signals of other sound sources, and obtains a pure or relatively pure speech signal of the target sound source, so as to solve the problems of large reverberation and noise when receiving far-field speech. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 A flow chart of a speech enhancement method provided by an embodiment of the present invention;
[0033] Figure 2 A schematic structural diagram of a speech enhancement device provided by an embodiment of the present invention;
[0034] Figure 3 A schematic diagram of the structure of a computing device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0035] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments.
[0036] Figure 1 Flowchart of a speech enhancement method provided by an embodiment of the present invention. Figure 1 The method shown includes at least steps S101 to S107.
[0037] In step S101, a multi-channel far-field speech signal collected by a microphone array is obtained.
[0038] Exemplarily, a front-end microphone array is used to pick up voice. For example, the microphone array can be dual-microphone, quad-microphone, six-microphone, etc. The placement of the microphones usually needs to consider multiple factors such as the application environment and the structure of the product. The distribution of the microphone array can be determined according to actual conditions, and this application does not limit it.
[0039] The far-field speech signals collected by the microphone array are sent to the speech processing module for subsequent speech processing, such as speech enhancement and speech recognition.
[0040] The collected far-field voice signals include the voice signals of the target sound source and the sound signals of other sound sources. The voice signals of the target sound source are the voice signals emitted by the user himself, and the sound signals of other sound sources are the sound signals emitted by sound sources other than the user himself, such as other human voice signals or other noises near the user.
[0041] It is easy to understand that the far-field speech mentioned in the present invention means far-field speech in the popular sense. For example, compared with near-field speech, far-field speech is the speech collected in a scenario where the target sound source is 3-10m away from the microphone; the far-field speech signal collected by each microphone in the microphone array is called a far-field speech signal of one channel, and the far-field speech signals collected by multiple microphones in the microphone array are multi-channel far-field speech signals.
[0042] In step S102, the short-time Fourier spectrum of the multi-channel far-field speech signal is extracted.
[0043] The short-time Fourier spectrum features of the far-field speech signals of each channel of the multiple channels are extracted respectively, and then the short-time Fourier features of the far-field speech signals of each channel of the multiple channels are fused to obtain the short-time Fourier spectrum of the far-field speech signals of the multiple channels.
[0044] Exemplarily, the steps for extracting short-time Fourier features from the far-field speech signal of one of the multiple channels are generally as follows: Frame and window the speech of each channel, and calculate the Fourier transform of each frame to obtain the short-time Fourier feature. This feature has a dimension of T×M×F. After concatenating the real and imaginary parts, the feature dimension is T×M×2F, where M is the number of channels, T is the number of frames, which is determined by the window length and window shift, and F is the number of frequency points, which is generally equal to half the Fourier transform length plus 1. After obtaining the short-time Fourier features of each of the multiple channels according to the above method, the short-time Fourier features of each channel are fused to obtain the short-time Fourier spectrum of the multiple channels.
[0045] In step S103, the short-time Fourier spectrum of the multi-channel far-field speech signal is input into the first-stage enhancement model to perform preliminary speech enhancement to determine a first enhanced speech spectrum.
[0046] The first-stage enhancement model is used to initially enhance the far-field speech signal. It is based on a bidirectional long short-term memory (BLSTM) network model and includes a bidirectional long short-term memory layer, a dropout layer, and a linear layer. The first-stage enhancement model is trained based on a first training set, which includes multiple first training sample pairs. Each first training sample pair includes a far-field speech signal and a corresponding near-field speech signal.
[0047] For example, the first-stage speech enhancement network consists of two bidirectional long short-term memory (BLSTM) layers, one dropout layer, and one linear layer. The network input is the short-time Fourier transform spectrum X, whose dimensions are T×(M×2F). The bidirectional long short-term memory network consists of 600 neurons in each direction, with a dropout coefficient of 0.5. The linear layer input is 1200-dimensional, and the output is M×2F-dimensional, representing the real and imaginary parts of the mask respectively. The entire process can be expressed as
[0048] m 1real ,m 1imag =BLSTM1(X)
[0049] Then the first enhanced speech spectrum output after enhancement by the first stage enhancement model is
[0050] Z 1real =m 1real ·X real -m 1imag ·X imag
[0051] Z 1imag =m 1real ·X imag +m 1imag ·X real
[0052] The loss function of the network is
[0053] |Z 1real -S real | 2 +|Z 1imag -S imag | 2
[0054] The neural network updates its parameters based on this loss function. The model is trained until the loss function converges or the training reaches a preset number of times.
[0055] In step S104, a first linear filtering is performed on the first enhanced speech spectrum to determine a filtered first enhanced speech spectrum.
[0056] The first enhanced speech spectrum is input into a first linear filter to determine a filtered first enhanced speech spectrum, wherein a filter coefficient of the first linear filter is determined based on a short-time Fourier spectrum of the multi-channel speech signal and the first enhanced speech spectrum.
[0057] For example, the coefficients of the minimum variance distortionless filter (MVDR) are calculated using the following formula as the beamforming weights:
[0058]
[0059] in is the energy density matrix of the noise,
[0060]
[0061] The steering vector r1 is The eigenvector corresponding to the largest eigenvalue decomposition, X H , are the Hermitian conjugate transposed matrices of X and Z1, respectively.
[0062] Using the beamforming coefficients and the original short-time Fourier spectrum, the signal spectrum pointing to the target sound source is calculated. That is the first enhanced speech spectrum after filtering.
[0063] In step S105, the short-time Fourier spectrum of the multi-channel far-field speech signal and the filtered first enhanced speech spectrum are input into the second-stage enhancement model for further speech enhancement to determine a second enhanced speech spectrum.
[0064] The second-stage enhancement model is used to further enhance the far-field speech signal. It is based on a bidirectional long short-term memory (BLSTM) network model and includes a bidirectional long short-term memory layer, random dropout, and a linear layer. The second-stage enhancement model is trained based on a second training set, which includes multiple second training sample pairs. Each second training sample pair includes one sample that includes the short-time Fourier spectrum of the multi-channel speech signal corresponding to the far-field speech and the first enhanced speech spectrum, and the other sample that includes the near-field speech signal corresponding to the far-field speech.
[0065] Exemplarily, the second-stage speech enhancement network includes two layers of bidirectional long short-term memory (BLSTM) networks, one layer of random dropout, and one layer of linear layer. The network input is the original short-time Fourier spectrum X and the signal output from step S104. The dimensions are all T×(M×2F). The bidirectional long short-term memory network includes 600 neurons in each direction, the random inactivation coefficient is 0.5, the linear layer input is 1200 dimensions, and the output is M×2F dimensions, representing the real and imaginary parts of the mask respectively. The whole process can be expressed as
[0066]
[0067] Then the enhanced speech spectrum is
[0068] Z 2real =m 2real ·X real -m 2imag ·X imag
[0069] Z 2imag =m 2real ·X imag +m 2imag ·X real
[0070] The loss function of the network is
[0071] |Z 2real -S real | 2 +|Z 2imag -S imag | 2
[0072] The neural network updates its parameters based on this loss function. The model is trained until the loss function converges or the training reaches a preset number of times.
[0073] In step S106, a second linear filtering is performed on the second enhanced speech spectrum to determine a filtered second enhanced speech spectrum.
[0074] The first enhanced speech spectrum is input into a second linear filter to determine a filtered second enhanced speech spectrum, wherein a filter coefficient of the second linear filter is determined based on a short-time Fourier spectrum of the multi-channel speech signal and the second enhanced speech spectrum.
[0075] Exemplarily, the coefficients of the minimum variance distortionless filter (MVDR) are calculated using the following formula as the beamforming weights:
[0076]
[0077] in is the energy density matrix of the noise,
[0078]
[0079] The steering vector r2 is The eigenvector corresponding to the largest eigenvalue decomposition. H , are the Hermitian conjugate transposed matrices of X and Z2, respectively.
[0080] Using the beamforming coefficients and the original short-time Fourier spectrum, the signal spectrum pointing to the target sound source is calculated.
[0081]
[0082] That is the second enhanced speech spectrum after filtering.
[0083] In step S107 , the speech signal of the target sound source is determined based on the filtered second enhanced speech spectrum.
[0084] The filtered second enhanced speech spectrum is subjected to inverse Fourier transform to determine the speech signal of the target sound source, that is, the filtered second enhanced speech spectrum output in S107 is restored to a speech signal by inverse Fourier transform.
[0085] In another example, in order to obtain a purer speech signal of the target sound source, the filtered second enhanced speech spectrum and the short-time Fourier spectrum of the multi-channel speech signal can be further input into the second-stage enhanced speech model for speech enhancement. For example, the filtered second enhanced speech spectrum and the short-time Fourier spectrum of the multi-channel speech signal are input into the second-stage enhanced speech model to determine the third enhanced speech spectrum; the third enhanced speech spectrum is then subjected to a second linear filtering to determine the filtered third enhanced speech spectrum; and the target speech signal is determined based on the filtered third enhanced speech spectrum. In other words, steps S106 and S107 are repeated a preset number of times (for example, 1 or 2 times), and then the filtered speech spectrum output from step S107 after the preset number of times is subjected to an inverse Fourier transform to obtain the speech signal of the target sound source.
[0086] Based on the same concept as the above method embodiment, the present invention also provides a speech enhancement device 200, which includes a method for realizing Figure 1 The units or means of each step in the speech enhancement method shown.
[0087] Figure 2 This is a structural diagram of a speech enhancement device provided in an embodiment of the present application. Figure 2 As shown, the speech enhancement device 200 at least includes:
[0088] An acquisition module 201 is configured to acquire a multi-channel far-field speech signal collected by a microphone array, wherein the far-field speech signal includes at least a speech signal of a target sound source and sound signals of other sound sources;
[0089] An extraction module 202 is configured to extract a short-time Fourier spectrum of the multi-channel far-field speech signal;
[0090] The first enhancement module 203 is configured to input the short-time Fourier spectrum of the multi-channel far-field speech signal into the first-stage enhancement model to perform preliminary speech enhancement and determine a first enhanced speech spectrum;
[0091] A first filtering module 204 is configured to perform a first linear filtering on the first enhanced speech spectrum to determine a filtered first enhanced speech spectrum;
[0092] The second enhancement module 205 is configured to input the short-time Fourier spectrum of the multi-channel far-field speech signal and the filtered first enhanced speech spectrum into a second-stage enhancement model for further speech enhancement to determine a second enhanced speech spectrum;
[0093] A second filtering module 206 is configured to perform a second linear filtering on the second enhanced speech spectrum to determine a filtered second enhanced speech spectrum;
[0094] The determination module 207 determines the speech signal of the target sound source based on the filtered second enhanced speech spectrum.
[0095] In one possible implementation, the extraction module 201 is specifically configured to extract the short-time Fourier features of the far-field speech signal of each channel of the far-field speech signals of the multiple channels respectively;
[0096] The short-time Fourier features of the far-field speech signals of each channel in the far-field speech signals of the multiple channels are fused to determine the short-time Fourier spectra of the far-field speech signals of the multiple channels.
[0097] In another possible implementation, the first-stage enhancement model is trained based on a first training set, where the first training set includes multiple first training sample pairs, each of which includes a far-field speech signal and a near-field speech signal corresponding to the far-field speech signal.
[0098] In another possible implementation, the first filtering module 204 is specifically configured to input the first enhanced speech spectrum into a first linear filter to determine the filtered first enhanced speech spectrum, wherein the filtering coefficients of the first linear filter are determined based on the short-time Fourier spectrum of the multi-channel speech signal and the first enhanced speech spectrum.
[0099] In another possible implementation, the second-stage enhancement model is trained based on a second training set, and the second training set includes multiple second training sample pairs, one sample of each second training sample pair includes the short-time Fourier spectrum of the multi-channel speech signal corresponding to the far-field speech and the first enhanced speech spectrum, and the other sample includes the near-field speech signal corresponding to the far-field speech.
[0100] In another possible implementation, the second filtering module 206 is specifically configured to input the first enhanced speech spectrum into a second linear filter to determine the filtered second enhanced speech spectrum, wherein the filtering coefficients of the second linear filter are determined based on the short-time Fourier spectrum of the multi-channel speech signal and the second enhanced speech spectrum.
[0101] In another possible implementation, the determining module 207 is specifically configured to input the filtered second enhanced speech spectrum and the short-time Fourier spectrum of the multi-channel speech signal into the second-stage enhanced speech model to determine a third enhanced speech spectrum;
[0102] performing a second linear filtering on the third enhanced speech spectrum to determine a filtered third enhanced speech spectrum;
[0103] Based on the filtered third enhanced speech spectrum, a speech signal of a target sound source is determined.
[0104] In another possible implementation, the determination module 207 is specifically configured to perform an inverse Fourier transform on the filtered second enhanced speech spectrum to determine the speech signal of the target sound source.
[0105] In another possible implementation, the first-stage speech enhancement model and the second-stage speech enhancement model both include a bidirectional long short-term memory network layer, a random dropout layer, and a linear layer.
[0106] It should be understood that the speech enhancement device 200 according to the embodiment of the present application can be implemented in the embodiment of the present application. Figure 1 The method shown in the figure is described in detail above and will not be repeated here for the sake of brevity.
[0107] Figure 3 A schematic diagram of the structure of a computing device provided in an embodiment of the present application.
[0108] like Figure 3 As shown, the computing device 300 includes at least one processor 301, a memory 302, a communication interface 303, and a microphone array 304. The processor 301 is in communication with the memory 302 and the communication interface 303, and communication can also be achieved through other means such as wireless transmission. The communication interface 303 is used to receive multi-channel far-field speech signals collected by the microphone array 304 (e.g., including microphone 1, microphone 2, ..., microphone N); the memory 302 stores computer instructions, and the processor 301 executes the computer instructions to perform the speech enhancement method in the aforementioned method embodiment.
[0109] It should be understood that in the embodiment of the present application, the processor 301 may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0110] The memory 302 may include a read-only memory and a random access memory, and provides instructions and data to the processor 301. The memory 302 may also include a nonvolatile random access memory.
[0111] The memory 302 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0112] It should be understood that the computing device 300 according to the embodiment of the present application can execute the implementation of the embodiment of the present application. Figure 1 The method shown is described in detail above and will not be repeated here for the sake of brevity.
[0113] The present invention also provides a storage medium, comprising: a readable storage medium and a computer program stored in the readable storage medium, wherein the computer program is used to implement the above-mentioned speech enhancement method.
[0114] Those skilled in the art should further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0115] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0116] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A speech enhancement method, characterized in that: include: Acquiring a multi-channel far-field speech signal collected by a microphone array, wherein the far-field speech signal includes at least a speech signal of a target sound source and sound signals of other sound sources; Extracting the short-time Fourier spectrum of the multi-channel far-field speech signal; Inputting the short-time Fourier spectrum of the multi-channel far-field speech signal into the first-stage enhancement model to perform preliminary speech enhancement and determine a first enhanced speech spectrum; performing a first linear filtering on the first enhanced speech spectrum to determine a filtered first enhanced speech spectrum; Inputting the short-time Fourier spectrum of the multi-channel far-field speech signal and the filtered first enhanced speech spectrum into a second-stage enhancement model for further speech enhancement to determine a second enhanced speech spectrum; performing a second linear filtering on the second enhanced speech spectrum to determine a filtered second enhanced speech spectrum; Based on the filtered second enhanced speech spectrum, a speech signal of the target sound source is determined.
2. The method according to claim 1, characterized in that The extracting of the short-time Fourier spectrum of the multi-channel far-field speech signal comprises: Extracting short-time Fourier features of the far-field speech signal of each channel in the multi-channel far-field speech signal respectively; The short-time Fourier features of the far-field speech signals of each channel in the multi-channel far-field speech signals are fused to determine the short-time Fourier spectrum of the multi-channel far-field speech signals.
3. The method according to claim 1 or 2, characterized in that The first-stage enhancement model is trained based on a first training set, where the first training set includes multiple first training sample pairs, and each first training sample pair includes a far-field speech signal and a near-field speech signal corresponding to the far-field speech signal.
4. The method according to any one of claim 1, characterized in that The performing a first linear filtering on the first enhanced speech spectrum to determine a filtered first enhanced speech spectrum includes: The first enhanced speech spectrum is input into a first linear filter to determine the filtered first enhanced speech spectrum, wherein the filter coefficient of the first linear filter is determined based on the short-time Fourier spectrum of the multi-channel far-field speech signal and the first enhanced speech spectrum.
5. The method according to any one of claim 1, characterized in that The second-stage enhancement model is trained based on a second training set, and the second training set includes multiple second training sample pairs, one sample of each second training sample pair includes the short-time Fourier spectrum of the multi-channel far-field speech signal corresponding to the far-field speech and the first enhanced speech spectrum, and the other sample includes the near-field speech signal corresponding to the far-field speech.
6. The method according to any one of claim 1, characterized in that The performing a second linear filtering on the second enhanced speech spectrum to determine a filtered second enhanced speech spectrum includes: The first enhanced speech spectrum is input into a second linear filter to determine the filtered second enhanced speech spectrum, wherein the filter coefficient of the second linear filter is determined based on the short-time Fourier spectrum of the multi-channel far-field speech signal and the second enhanced speech spectrum.
7. The method according to any one of claim 1, characterized in that The determining the speech signal of the target sound source based on the filtered second enhanced speech spectrum includes: Inputting the filtered second enhanced speech spectrum and the short-time Fourier spectrum of the multi-channel far-field speech signal into the second-stage enhancement model to determine a third enhanced speech spectrum; performing a second linear filtering on the third enhanced speech spectrum to determine a filtered third enhanced speech spectrum; Based on the filtered third enhanced speech spectrum, a speech signal of the target sound source is determined.
8. The method according to any one of claim 1, characterized in that The determining the speech signal of the target sound source based on the filtered second enhanced speech spectrum includes: Performing an inverse Fourier transform on the filtered second enhanced speech spectrum to determine the speech signal of the target sound source.
9. The method according to any one of claims 1, characterized in that The first-stage enhancement model and the second-stage enhancement model both include a bidirectional long short-term memory network layer, a random dropout layer, and a linear layer.
10. A computing device comprising a memory and a processor, characterized in that: The memory stores executable code, and the processor executes the executable code to implement the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Echo cancellation method and device, electronic equipment and readable storage medium
CN112687288A
Audio data processing method and device
CN113096679A