A method of dereverberation for a microphone array

By combining microphone array algorithms and deep learning technology, the problem of voice quality degradation caused by late reflection reverberation in teleconferences has been solved, and clarity and naturalness have been improved in long-distance calls.

CN119626197BActive Publication Date: 2025-12-16广东公信智能会议股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411731485.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-12-16
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

In teleconferences, traditional reverberation suppression techniques struggle to distinguish between early and late reflections, leading to a decline in voice quality. Existing technologies may fail to effectively suppress reverberation or excessively eliminate early reflections, affecting call clarity and naturalness.

Method used

By employing a microphone array algorithm combined with deep learning technology, and through fixed beamforming, noise reduction and dereverberation modules, speech direction estimation, and early reverberation simulation modules, reverberation is finely processed, early and late reflections are distinguished, and speech clarity and fullness are enhanced.

Benefits of technology

It effectively suppresses late-stage reverberation, restores the naturalness and saturation of speech, and provides a better call experience, especially significantly improving voice quality in long-distance calls.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119626197B_ABST
    Figure CN119626197B_ABST
Patent Text Reader

Abstract

The application discloses a late reverberation removing processing method of a microphone array, and relates to the technical field of speech signal processing.The method comprises the following steps: a fixed beam forming module performs spatial filtering processing on an original signal by using a specified beam to obtain a first speech signal; a noise removing and reverberation removing module generates an attenuation gain mask according to the first speech signal, performs noise removing processing on the first speech signal by using the mask, and obtains a second speech signal; a speech direction estimation module performs speech direction estimation by performing weighted calculation on the original signal according to the mask, and obtains a target speech direction; the fixed beam forming module updates the specified beam according to the target speech direction; and an early reverberation simulation module performs reverberation adding processing on the second speech signal, and outputs a final speech signal.Through speech estimation, the fixed beam processing is optimized; after reverberation elimination, the signal is subjected to reverberation adding processing in a spatial reverberation simulation mode, early reflection is simulated, and the sound is clearer and fuller.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech signal processing, and in particular to a late reverberation removal processing method for a microphone array. BACKGROUND

[0002] In the scenario of a conference call, it is very important to obtain clear speech quality, especially in a long-distance situation, which needs to face the challenges of low signal-to-noise ratio and high reverberation, which will seriously affect the clarity of the call. At present, the reverberation suppression technology for the conference scenario generally has the following shortcomings:

[0003] 1. Traditional Wiener filtering or microphone array methods often cause distortion of the main speech signal or poor reverberation suppression effect;

[0004] 2. Many deep learning-based dereverberation technologies can effectively suppress reverberation, but excessive suppression may eliminate the early reverberation that is beneficial to human ear perception, making the target speech sound unnatural and lack of vitality;

[0005] In the process of communication, the generation of speech reverberation is mainly caused by the reflection of sound waves. This reflection can be divided into early reflection and late reflection. The early reflection can enhance the fullness and clarity of the sound, while the late reflection sound will make the speech unclear. Traditional reverberation processing techniques often have difficulty in distinguishing between the two types of reflection sound, and usually treat them the same. Ultimately, either the reverberation is not effectively suppressed and the call quality is still poor, or the reverberation is excessively eliminated, resulting in speech that sounds harsh, unnatural, and even distorted. SUMMARY

[0006] In view of the problems existing in the prior art, the present application provides a late reverberation removal processing method for a microphone array, which solves the problem of speech quality degradation caused by long-distance communication in a conference scenario.

[0007] The present application adopts a microphone array algorithm combined with deep learning technology to finely process reverberation. By simulating and restoring early reverberation, the interference of reverberation on speech is suppressed, and the clarity and fullness of speech are enhanced, thereby providing a better call experience. The technical solution is implemented as follows:

[0008] A late reverberation removal processing method for a microphone array, comprising a fixed beamforming module, a noise reduction and dereverberation module, a speech direction estimation module, and an early reverberation simulation module; further comprising a microphone array composed of multiple microphones; the microphone array is used to collect original signals;

[0009] Further comprising the following steps:

[0010] S1, the fixed beam forming module receives the original signal, and uses the specified beam to perform spatial filtering processing on the original signal, to obtain a first voice signal, that is, to suppress part of the reverberation by using the beam, that is, to perform beam enhancement; the first voice signal is sent to the noise reduction and dereverberation module; the original signal is sent to the voice direction estimation module;

[0011] The beam coefficient is a filter coefficient for an array microphone configuration, which affects the pickup direction, and is a beam designed according to the target direction of pickup, which is used to enhance the sound of the target direction angle;

[0012] S2, the noise reduction and dereverberation module generates an attenuation gain mask according to the first voice signal, and performs noise reduction processing on the first voice signal by using the attenuation gain mask to obtain a second voice signal;

[0013] The attenuation gain mask is sent to the voice direction estimation module to perform step S3;

[0014] The second voice signal is sent to the early reverberation simulation module to perform step S5;

[0015] S3, the voice direction estimation module performs weighted calculation on the original signal according to the attenuation gain mask, and performs voice direction estimation to obtain a target voice direction; the target voice direction is sent to the fixed beam forming module to perform step S4;

[0016] S4, the fixed beam forming module selects and updates the specified beam according to the target voice direction; that is, a new beam of direction is selected, and spatial filtering is performed by using the specified beam;

[0017] S5, the early reverberation simulation module performs reverberation processing on the second voice signal to output a final voice signal.

[0018] The original signal is also an array signal. The fixed beam forming module is fixed beam forming. Through the noise reduction and dereverberation processing of the fixed beam forming module and the noise reduction and dereverberation module, the voice signal, especially the voice signal far away, is distorted, which affects the voice listening feeling; the early reflection recovery processing of the early reverberation simulation module on the voice signal is performed to restore the naturalness and saturation of the voice; the voice direction estimation is performed on the direction of the voice signal to obtain a target voice direction, which can further optimize the beam processing, dynamically adjust the beam direction, that is, update the specified beam, to more accurately process the voice signal.

[0019] As a further optimization of the above scheme, the number of microphones is 8 or 6; the fixed beamforming module fixes the specified beam through a pre-set filter coefficient.

[0020] The number of microphones can also be other numbers, preferably 8 or 6.

[0021] As a further optimization of the above scheme, the weighting calculation is that the second speech signal is weighted by the attenuation gain mask at the time-frequency point, that is:

[0022]

[0023] Where X is the second speech signal, X(l,k) is the frequency point value of the kth frequency point of the first frame;

[0024] mask(l,k) is the attenuation gain value corresponding to the kth frequency point of the first frame;

[0025] X(l,k) and mask(l,k) are multiplied to obtain the time-frequency point

[0026] In speech direction estimation, in order to obtain the de-noised data of each microphone channel, the following formula is used to calculate respectively:

[0027]

[0028] Where X m (l,k) is the frequency point value of the kth frequency point of the first frame of the mth channel; one microphone corresponds to one channel.

[0029] The speech direction estimation module performs speech direction estimation according to the time-frequency point and the doa algorithm.

[0030] The speech signal has corresponding spectral data in each frame, and the frequency point can be used to identify specific phonemes, tones or other features in the speech. For example, the fundamental frequency (F0) is an important frequency point in the speech signal, which is directly related to the pitch of the speaker.

[0031] The Direction of Arrival (DOA) estimation algorithm is a key technology for determining the direction of the signal source, and is widely used in radar, sonar, wireless communication and other fields. The doa estimation algorithm used in this scheme can be music (Multiple Signal Classification), srp-phat (Steered-Response Power PHAT), dsb (Delay and Sum Beamforming), etc., but is not limited to these methods.

[0032] X m (l,k) and The speech signals before and after processing are respectively X and Y. The mask is used to weight the multi-channel spectral data in time-frequency points, so as to remove the non-target speech signal interference.

[0033] As a further optimization of the above scheme, the noise reduction and dereverberation module is a neural network structure based on CRNN.

[0034] The training process of the neural network structure includes inputting clean data and first noise data.

[0035] The clean data and the first noise data are reverberated to obtain second noise data.

[0036] The amplitude spectrum or log spectrum of the second noise data is input as a feature into the neural network structure, and the attenuation gain mask of the amplitude spectrum is output.

[0037] The second noise data is denoised by using the attenuation gain mask to obtain an enhanced speech signal.

[0038] The enhanced speech signal is combined with the clean data to calculate and optimize the neural network structure by using the L1 loss function.

[0039] CRNN, Convolutional Recurrent Neural Network, is a convolutional recurrent neural network that combines convolutional neural network (CNN) and recurrent neural network (RNN); including encoding layer, intermediate layer and decoding layer; the encoding layer uses a cnn convolution operator; the intermediate layer uses a gru recursive operator; and the decoding layer uses a cnn-transpose deconvolution operator.

[0040] In the inference stage, the network parameters of the training process are directly copied, and the network is modified for streaming inference; the data output by beamforming is input for noise reduction processing.

[0041] As a further optimization of the above scheme, the L1 loss function is:

[0042] L1 loss = mean(abs(abs(Y(l,k)) x Mask(l,k) - abs(X(l,k))));

[0043] Wherein, Mask(l, k) represents the attenuation gain mask; X(l, k) represents the clean data; Y(l, k) represents the second noise data, and abs(Y(l, k)) x Mask(l, k) represents the enhanced speech signal, wherein the same weighting calculation method is adopted; the mean() function is used to calculate the average of all differences to obtain the L1 loss; abs() is an absolute value function; (l, k) represents the first frame and the kth frequency point.

[0044] As a further optimization of the above scheme, the early reverberation simulation module comprises a plurality of comb filters connected in parallel and a plurality of all-pass filters connected in series after the comb filters.

[0045] The comb filters filter the second speech signal respectively; the input results of the plurality of comb filters are added to obtain a third speech signal.

[0046] The third speech signal is sent to the all-pass filters for filtering in sequence to obtain a fourth speech signal.

[0047] The fourth speech signal, the second speech signal, and a preset gain signal are superimposed, and the superimposed signal is the final speech signal.

[0048] As a further optimization of the above scheme, the filtering process of the comb filter is represented as:

[0049] y(t) = x(t-τ) + g x y(t-τ);

[0050] The filtering process of the all-pass filter is represented as:

[0051] y(t) = -g x x(t) + x(t-τ) + g x y(t-τ);

[0052] Wherein, x(·) is an input signal, y(·) is an output signal, t is a time, τ is a pre-set delay; g is a pre-set gain coefficient.

[0053] The calculation of y(t) depends on the value of y(t-τ) which has been calculated. g is a reflection attenuation gain, that is, the attenuation amount of the reflected speech signal after a certain time.

[0054] As a further optimization of the above scheme, the processing of the original signal, the first speech signal, the second speech signal, and the final speech signal is in units of frames; a preset estimation period is further included, and the interval of one estimation period is a plurality of frames; in step S2, after each estimation period, the attenuation gain mask is sent to the speech direction estimation module to execute step S3.

[0055] The estimation period is 1 frame, 2 frames or other number of frames as an interval. In speech signal processing, since speech has short-time stationarity, that is, the characteristics of the speech signal do not change much in a very short time, the speech signal is segmented into shorter frames for processing.

[0056] As a further optimization of the above scheme, the fixed beamforming module is preset with a beams, and the angle difference between adjacent two beams is b; a x b = 360; the central angle of one of the beams is 0°, and the central angles of the remaining beams are all integer multiples of b.

[0057] As a further optimization of the above scheme, the target speech direction is an angle; in step S4, the beam with the central angle closest to the target speech direction is selected as the specified beam.

[0058] For example, if it is identified that the current target speaker angle (target speech direction) is 40 degrees, the beam with a central angle of 30 degrees is selected as the specified beam.

[0059] Compared with the prior art, the application has the following beneficial effects:

[0060] (1) Through the noise reduction and dereverberation processing of the fixed beamforming module and the noise reduction and dereverberation module, most of the interference of reverberation is eliminated; for the defects such as speech and distortion caused by the reverberation elimination processing, the signal processed by the reverberation is added to the reverberation processing through the spatial reverberation simulation mode, and the early reflection in the reverberation is simulated to make the sound clearer and fuller.

[0061] (2) The orientation of the speech signal is positioned through the speech direction estimation, and the target speech direction obtained can further optimize the beam processing, dynamically adjust the beam direction, that is, update the specified beam, so as to more accurately process the speech signal;

[0062] (3) The mask posterior estimation is adopted, the multi-channel spectral data is weighted once through the mask, the non-target speech signal interference is removed, and very good speech positioning effect can be obtained in the case that there is very large noise in the outside. BRIEF DESCRIPTION OF DRAWINGS

[0063] Figure 1 is a data flow schematic diagram of a late reflection reverberation processing method of a microphone array provided by an embodiment of the application;

[0064] Figure 2 is a beam pointing schematic diagram provided by an embodiment of the application;

[0065] Figure 3 is a neural network structure schematic diagram provided by an embodiment of the application;

[0066] Figure 4is a data processing flow schematic diagram of an early reverberation simulation module provided by the embodiment of the present application. DETAILED DESCRIPTION

[0067] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the present application.

[0068] As shown in the figure, the embodiment provides a late reverberation removing processing method of a microphone array, which comprises a fixed beam forming module, a noise reducing and reverberation removing module, a voice direction estimation module and an early reverberation simulation module. Figures 1 to 4

[0069] The microphone array comprises a plurality of microphones; in the embodiment, the number of microphones is 8. The microphone array is used for collecting original signals.

[0070] The method further comprises the following steps:

[0071] S1, the fixed beam forming module receives the original signals, and performs spatial filtering processing on the original signals by using a specified beam to obtain a first voice signal, that is, a part of reverberation is suppressed by using a beam; the first voice signal is sent to the noise reducing and reverberation removing module; and the original signals are sent to the voice direction estimation module.

[0072] The beam coefficient is a filtering coefficient for the array microphone configuration, which affects the pickup direction and is a beam designed according to the pickup target direction, and the beam is used to enhance the sound of the target direction angle.

[0073] The fixed beam forming module fixes the specified beam by using a pre-set filtering coefficient.

[0074] In the embodiment, the fixed beam forming module is pre-set with 12 beams, and the angle difference between two adjacent beams is 30°; the central angle of one of the beams is 0°, and the central angles of the remaining beams are all integer multiples of 30°. The central angles of each beam are respectively: 0°, 30°, 60°, 90°, 120°, 150°, 180°, 210°, 240°, 270°, 300° and 330°.

[0075] Figure 2 The provided pointing diagram shows that the direction is at the position of 180°, and the width is 60°.

[0076] ​S2, the noise reduction and dereverberation module generates an attenuation gain mask according to the first voice signal, and performs noise reduction processing on the first voice signal by using the attenuation gain mask to obtain a second voice signal.

[0077] In the embodiment, the noise reduction and dereverberation module is a neural network structure based on CRNN (CRNN-Network).

[0078] The training process of the neural network structure includes inputting clean data clean and first noise data noise.

[0079] The clean data clean and the first noise data noise are reverberated by RIR (Room Impulse Response) to obtain second noise data.

[0080] The amplitude spectrum or log spectrum of the second noise data is input as a feature into the neural network structure, and an attenuation gain mask of the amplitude spectrum is output.

[0081] The second noise data is noise-reduced by using the attenuation gain mask to obtain an enhanced voice signal.

[0082] The enhanced voice signal is combined with the clean data to calculate and optimize the neural network structure by using an L1 loss function.

[0083] In the embodiment, the L1 loss function is as follows:

[0084] L1 loss = mean(abs(abs(Y(l, k)) x Mask(l, k) - abs(X(l, k))));

[0085] Wherein, Mask(l, k) represents the attenuation gain mask; X(l, k) represents the clean data; Y(l, k) represents the second noise data, and abs(Y(l, k)) x Mask(l, k) represents the enhanced voice signal, which adopts the same weighting calculation method as described above; the mean() function is used to calculate the average value of all differences to obtain the L1 loss; abs() is an absolute value function; (l, k) represents the first frame and the kth frequency point.

[0086] CRNN, Convolutional Recurrent Neural Network, is a convolutional recurrent neural network that combines convolutional neural network (CNN) and recurrent neural network (RNN). It includes an encoding layer, an intermediate layer and a decoding layer. The encoding layer adopts a cnn convolution operator; the intermediate layer adopts a gru recursive operator; and the decoding layer adopts a cnn-transpose deconvolution operator.

[0087] In the inference stage, the network parameters of the training process are directly copied, and the network is modified for streaming inference. The data output by beamforming is input for noise reduction processing.

[0088] The attenuation gain mask is sent to the speech direction estimation module, and step S3 is performed.

[0089] The second speech signal is sent to the early reverberation simulation module, and step S5 is performed.

[0090] S3, the speech direction estimation module performs weighted calculation on the original signal according to the attenuation gain mask, and performs speech direction estimation to obtain the target speech direction.

[0091] In this embodiment, the weighted calculation is to perform time-frequency point weighting on the second speech signal through the attenuation gain mask, that is:

[0092]

[0093] Where X is the second speech signal, and X(l,k) is the frequency point value of the kth frequency point of the first frame.

[0094] mask(l,k) is the attenuation gain value corresponding to the kth frequency point of the first frame.

[0095] X(l,k) and mask(l,k) are multiplied to obtain the time-frequency point

[0096] In speech direction estimation, in order to obtain the de-noised data of each microphone channel, the following formula is used to calculate respectively:

[0097]

[0098] Where X m (l,k) is the frequency point value of the kth frequency point of the first frame of the mth channel microphone; one microphone corresponds to one channel.

[0099] The speech direction estimation module performs speech direction estimation according to the time-frequency point and the doa algorithm.

[0100] The speech signal has corresponding spectral data in each frame, and the frequency point can be used to identify specific phonemes, tones or other features in the speech. For example, the fundamental frequency (F0) is an important frequency point in the speech signal, which is directly related to the pitch of the speaker.

[0101] The direction of arrival (DOA) estimation algorithm is a key technology for determining the direction of a signal source, and is widely used in radar, sonar, wireless communication and other fields. The DOA estimation algorithm used in the scheme can be, but is not limited to, music (Multiple Signal Classification), srp-phat (Steered-Response Power PHAT), dsb (Delay and Sum Beamforming) and the like.

[0102] X m (l,k) and The speech signals before and after processing are respectively x(n) and x'(n). The mask is used to weight the multi-channel spectral data in time-frequency points, so as to remove the interference of non-target speech signals. The advantage is that very good speech positioning effect can be obtained even in the case of very large noise in the outside world.

[0103] The target speech direction is sent to the fixed beamforming module, and step S4 is performed. In this embodiment, the target speech direction is an angle.

[0104] S4, the fixed beamforming module selects and updates a specified beam according to the target speech direction, that is, a beam with a new direction is selected, and spatial filtering is performed by using the specified beam. In this embodiment, the beam with the center angle closest to the target speech direction is selected as the specified beam. For example, if the current target speaker angle (target speech direction) is 40 degrees, the beam with a center angle of 30 degrees is selected as the specified beam.

[0105] S5, the early reverberation simulation module performs reverberation processing on the second speech signal, and outputs a final speech signal.

[0106] In this embodiment, the early reverberation simulation module includes 4 comb filters connected in parallel and 2 all-pass filters connected in series after the comb filters.

[0107] The comb filters filter the second speech signal respectively; the input results of the multiple comb filters are added to obtain a third speech signal;

[0108] The third speech signal is sent to the all-pass filters for filtering in sequence to obtain a fourth speech signal;

[0109] The fourth speech signal, the second speech signal and a preset gain signal are superimposed, that is, the final speech signal.

[0110] Figure 4 In this embodiment, x(n) is the second speech signal, g_w is the preset gain signal; x'(n) is the signal obtained by superimposing the fourth speech signal and g_w, and y(n) is the final speech signal.

[0111] In the embodiment, the filtering process of the single comb filter is represented as:

[0112] y(t) = x(t - τ) + g * y(t - τ);

[0113] The filtering process of the single all-pass filter is represented as:

[0114] y(t) = -g * x(t) + x(t - τ) + g * y(t - τ);

[0115] Wherein, x(·) is an input signal, y(·) is an output signal, t is a time, τ is a pre-set delay; g is a pre-set gain coefficient.

[0116] The calculation of y(t) depends on the value of y(t - τ) which has been calculated. g is a reflection attenuation gain, that is, the attenuation amount of the reflected speech signal after a certain time.

[0117] The original signal is also an array signal. The fixed beamforming module is fixed beamforming. Through the noise reduction and dereverberation processing of the fixed beamforming module, the noise reduction and dereverberation module, the obtained speech signal, especially the speech signal far away, is distorted, which affects the speech listening experience; the early reflection simulation module is used to mirror the speech signal for early reflection recovery processing, so as to restore the naturalness and saturation of the speech; the orientation of the speech signal is positioned through the speech direction estimation, and the obtained target speech direction can further optimize the beam processing, dynamically adjust the beam direction, that is, update the specified beam, so as to more accurately process the speech signal.

[0118] In the embodiment, the processing of the original signal, the first speech signal, the second speech signal and the final speech signal is in units of frames; further comprising a pre-set estimation period, the interval of one estimation period is 10 frames; in step S2, the attenuation gain mask is sent to the speech direction estimation module every estimation period, and step S3 is executed. That is, the speech direction estimation is executed once every 10 frames.

[0119] In the speech signal processing, since the speech has short-time stationarity, that is, in a very short time, the characteristics of the speech signal do not change much, therefore, the speech signal is divided into shorter frames for processing.

[0120] According to the disclosure and teaching of the above description, those skilled in the art of the present application can also make changes and modifications to the above embodiments. Therefore, the present application is not limited to the specific embodiments disclosed and described above, and some modifications and changes of the present application should also fall within the protection scope of the claims of the present application. In addition, although some specific terms are used in the specification, these terms are only for convenience of description and do not constitute any limitation on the present application.

Claims

1. A method of dereverberation processing of a microphone array, characterized by, The fixed beam forming module, the noise reduction and dereverberation module, the voice direction estimation module and the early reverberation simulation module are included; a microphone array composed of multiple microphones is further included; the microphone array is used to collect original signals; Further comprising the following steps: S1, the fixed beam forming module receives the original signals, and performs spatial filtering processing on the original signals by using a specified beam to obtain a first voice signal; the first voice signal is sent to the noise reduction and dereverberation module; and the original signal is sent to the voice direction estimation module; S2, the noise reduction and dereverberation module generates an attenuation gain mask according to the first voice signal, and performs noise reduction processing on the first voice signal by using the attenuation gain mask to obtain a second voice signal; The attenuation gain mask is sent to the voice direction estimation module to perform step S3; The second voice signal is sent to the early reverberation simulation module to perform step S5; S3, the voice direction estimation module performs weighted calculation on the original signal according to the attenuation gain mask, and performs voice direction estimation to obtain a target voice direction; the target voice direction is sent to the fixed beam forming module to perform step S4; S4, the fixed beam forming module selects and updates the specified beam according to the target voice direction; S5, the early reverberation simulation module performs reverberation processing on the second voice signal to output a final voice signal.

2. The method of dereverberation of a microphone array according to claim 1, characterized in that, The number of microphones is 8 or 6; the fixed beam forming module fixes the specified beam by using pre-set filter coefficients.

3. The method of claim 1, wherein, The weighted calculation is time-frequency point weighting of the second voice signal by using the attenuation gain mask, that is: Wherein, X is the second voice signal, X(l,k) is the frequency point value of the 1st frame and the kth frequency point; Mask(l,k) is the attenuation gain value corresponding to the 1st frame and the kth frequency point; X(l,k) and mask(l,k) are multiplied to obtain the time-frequency point The voice direction estimation module performs the voice direction estimation according to the time-frequency points and a doa algorithm.

4. The method of dereverberation of a microphone array according to claim 3, characterized in that, The noise reduction and dereverberation module is a neural network structure based on CRNN; The training process of the neural network structure includes inputting clean data and first noise data; The clean data and the first noise data are reverberated to obtain second noise data; The amplitude spectrum or log spectrum of the second noise data is input as a feature into the neural network structure, and the attenuation gain mask of the amplitude spectrum is output; The second noise data is denoised by using the attenuation gain mask to obtain an enhanced voice signal; The enhanced voice signal and the clean data are combined with an L1 loss function to calculate and optimize the neural network structure.

5. The method of dereverberation of a microphone array according to claim 4, characterized in that, The L1 loss function is: L1 loss = mean(abs(abs(Y(l,k))×Mask(l,k)-abs(X(l,k)))) Wherein, Mask(l, k) represents the attenuation gain mask; X(l, k) represents the clean data; Y(l, k) represents the second noise data, abs(Y(l, k)) x Mask(l, k) represents the enhanced speech signal; the mean() function is used to calculate the average of all differences; abs() is an absolute value function; (l, k) represents the first frame k frequency point.

6. The method of dereverberation of a microphone array according to claim 1, wherein, The early reverberation simulation module comprises a plurality of comb filters connected in parallel and a plurality of all-pass filters connected in series after the comb filters. The comb filters filter the second speech signal respectively; the input results of the plurality of comb filters are added to obtain a third speech signal. The third speech signal is sent to the all-pass filters for filtering in sequence to obtain a fourth speech signal. The fourth speech signal, the second speech signal and a preset gain signal are superimposed, and the final speech signal is obtained.

7. The method of dereverberation of a microphone array according to claim 6, characterized in that, The filtering process of the comb filter is represented as: y(t) = x(t-τ) + g x y(t-τ); The filtering process of the all-pass filter is represented as: y(t) = -g x x(t) + x(t-τ) + g x y(t-τ); Wherein, x(·) is an input signal, y(·) is an output signal, t is a time, τ is a pre-set delay; g is a pre-set gain coefficient.

8. The method of claim 1, wherein, The processing of the original signal, the first speech signal, the second speech signal and the final speech signal is in units of frames; a preset estimation period is further included, and the interval of one estimation period is several frames. In step S2, the attenuation gain mask is sent to the speech direction estimation module every estimation period, and step S3 is executed.

9. The method of claim 1, wherein, The fixed beam forming module is pre-set with a beams, and the angle difference between adjacent two beams is b; a x b = 360; the central angle of one of the beams is 0°, and the central angles of the remaining beams are all integer multiples of b.

10. The method of dereverberation of a microphone array according to claim 9, wherein, The target speech direction is an angle; in step S4, the beam closest to the target speech direction in central angle is selected as the specified beam.

Citation Information

Patent Citations

  • Method for eliminating indoor reverberations

    CN103413547A

  • Real-time voice de-reverberation mixing method and system

    CN114255777A