Intelligent voice control system and voice enhancement method in reverberation environment
By analyzing the time-frequency characteristics and interference covariance matrix of speech signals in a reverberant environment, speech enhancement in a reverberant environment is achieved, the acoustic crosstalk problem caused by reverberation is solved, and the accuracy of the speech recognition system and the quality of the speech signal are improved.
Patent Information
- Application Number
- CN202511264320.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-05
AI Technical Summary
In a reverberant environment, the voice signal causes acoustic crosstalk due to the superposition of reflected sound, which affects the accuracy of the voice recognition system. Existing technologies find it difficult to effectively suppress the acoustic crosstalk caused by reverberation.
By obtaining the initial speech signal in a reverberant environment, analyzing its time-frequency spectrum and frequency response attenuation curve, determining the reverberation energy distribution, calculating the reverberation suppression gain, and combining the time-frequency mask and interference covariance matrix, reverberation suppression and harmonic enhancement are performed to weaken the reflected sound energy and preserve the speech harmonic structure.
Significantly reduce the acoustic crosstalk caused by reverberation superposition, improve speech clarity and recognition accuracy, restore the speech harmonic structure, and enhance the naturalness of the speech signal and the spectral details.
Smart Images

Figure CN120808806A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of voice control, more particularly, the present application relates to an intelligent voice control system and a voice enhancement method in a reverberation environment. BACKGROUND
[0002] Voice control is an intelligent interactive way of realizing device control by analyzing human voice instructions. With its convenient and natural characteristics, it has been deeply integrated into multiple scenarios such as smart home, smart wear, and vehicle-mounted systems. The core is to accurately capture and recognize voice information to perform corresponding operations. However, in actual applications, it often faces interference such as noise, echo, and too far distance, which leads to distortion of voice signals and affects recognition effect. Voice enhancement technology, as an important support, can effectively filter environmental interference and strengthen effective voice features through noise reduction, dereverberation, and signal clarity enhancement algorithms, so as to maintain high accuracy and reliability of voice control technology in complex acoustic environments and provide solid protection for natural and smooth human-computer interaction.
[0003] In practical application scenarios such as voice communication and voice recognition, voice signals often form reverberation due to reflection of walls, furniture and other objects in the propagation environment, resulting in superposition of original voice and reflected sound, producing acoustic crosstalk. This acoustic crosstalk can blur the harmonic structure of the voice and mask the key frequency components, not only reducing the clarity and intelligibility of the voice, but also seriously affecting the accuracy of the voice recognition system (for example, spectral distortion caused by reverberation increases the recognition error rate by more than 30%). Although existing reverberation suppression techniques can weaken part of the reflected sound energy through linear filtering, spectral subtraction and other methods, due to the complexity of the environment and the non-stationarity of the signal, some acoustic crosstalk will still remain after processing, especially in the frequency band where the voice harmonics are concentrated. The superposition of crosstalk and useful signals will form an interference area that is difficult to separate, resulting in harmonic energy attenuation and voice naturalness degradation. Therefore, how to suppress the acoustic crosstalk caused by reverberation has become a problem faced by the industry. SUMMARY
[0004] The present application provides an intelligent voice control system and a voice enhancement method in a reverberation environment, which can suppress acoustic crosstalk caused by reverberation.
[0005] In a first aspect, the present application provides a voice enhancement method in a reverberation environment, which is used for voice enhancement in a voice control system. The method comprises the following steps: Obtaining an initial voice signal in a reverberation environment; Determining the reverberation energy distribution in each frequency band according to the time-frequency spectrogram of the initial voice signal and the frequency response attenuation curve in each frequency band, and determining the reverberation suppression gain of the initial voice signal in the time-frequency domain through all the reverberation energy distributions; reverberation suppression gain is determined according to the time-frequency spectrogram of the initial speech signal and the frequency response decay curve in each frequency band, and the reverberation energy distribution in each frequency band is determined according to the time-frequency spectrogram and the reverberation time in each frequency band. The time-frequency mask estimation is performed on the initial speech signal to obtain a time-frequency mask of the initial speech signal, and the residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain is determined according to the time-frequency mask and an interference covariance matrix of a target sound source. The harmonic enhancement is performed on the speech signal after the reverberation suppression through all the residual interference energies to obtain a speech enhancement signal.
[0006] In some embodiments, the determination of the reverberation energy distribution in each frequency band according to the time-frequency spectrogram of the initial speech signal and the frequency response decay curve in each frequency band specifically comprises: The time-frequency transformation is performed on the initial speech signal to obtain a time-frequency spectrogram of the initial speech signal. The reverberation time in each frequency band is determined through the frequency response decay curve in each frequency band. The reverberation energy distribution in each frequency band is determined according to the time-frequency spectrogram and the reverberation time in each frequency band.
[0007] In some embodiments, the determination of the reverberation time in each frequency band through the frequency response decay curve in each frequency band specifically comprises: for each frequency band; the frequency response decay curve in the frequency band is determined; the trend extraction is performed on the frequency response decay curve in the frequency band to filter out an effective section in which the energy monotonically decays with time in each frequency response decay curve; the time interval experienced by the initial peak energy from the initial peak energy to the energy decay reference value in the effective section is calculated, and the obtained time interval is taken as the reverberation time in the frequency band.
[0008] In some embodiments, the determination of the reverberation suppression gain of the initial speech signal in the time-frequency domain through all the reverberation energy distributions specifically comprises: the reverberation energy distributions on the frequency bands are merged to obtain a full-frequency-domain reverberation energy spectrogram; the comparison and analysis are performed on the reverberation energy spectrogram and the power spectrum of the initial speech signal to obtain a reverberation dominant region and a speech dominant region; the reverberation suppression factor of each time-frequency unit of the initial speech signal in the time-frequency domain is determined according to the reverberation dominant region and the speech dominant region; the gain calculation is performed on the reverberation suppression factors of all the time-frequency units to obtain the reverberation suppression gain of the initial speech signal in the time-frequency domain.
[0009] In some embodiments, the reverb suppression on the initial speech signal according to the reverb suppression gain comprises specifically: The reverb threshold value in each frequency band is determined according to the reverb energy distribution in the frequency band. The time-frequency coefficient of each time-frequency unit of the initial speech signal in the time-frequency domain is suppressed by the reverb suppression gain and the reverb threshold value in each frequency band, and then the reverb-suppressed speech signal is obtained.
[0010] In some embodiments, the time-frequency mask estimation on the initial speech signal comprises specifically: The speech presence probability of each time-frequency unit of the initial speech signal in the time-frequency domain is detected by the pre-trained speech activity detection model. The time-frequency mask of the initial speech signal is determined according to all the speech presence probabilities.
[0011] In some embodiments, the residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain is determined according to the time-frequency mask and the interference covariance matrix of the target sound source, which comprises specifically: The interference covariance matrix of the target sound source is determined. The effective area of the initial speech signal in the time-frequency domain is identified according to the time-frequency mask. The interference leakage estimation on the effective area of the initial speech signal in the time-frequency domain is performed by the interference covariance matrix of the target sound source, and the residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain is obtained.
[0012] In some embodiments, the harmonic enhancement on the reverb-suppressed speech signal by all the residual interference energies comprises specifically: The interference area corresponding to the harmonic structure of the reverb-suppressed speech signal in the time-frequency domain is identified by all the residual interference energies. The amplitude enhancement on the interference area corresponding to the harmonic structure in the time-frequency domain is performed, and then the speech enhanced signal is obtained.
[0013] In some embodiments, the reverb environment represents an acoustic environment in which, in addition to direct sound, sound waves propagate in a closed or semi-closed space, multiple reflections and scattering occur on the surfaces of walls, floors, ceilings and objects in the space to form a superimposed sound field.
[0014] In a second aspect, the present application provides a speech control system, which comprises a speech enhancement unit, and the speech enhancement unit comprises: The acquisition module is configured to acquire an initial speech signal in a reverb environment. The processing module is configured to determine reverberation energy distribution in each frequency band according to a time-frequency spectrogram of the initial speech signal and a frequency response attenuation curve in each frequency band, and determine reverberation suppression gain of the initial speech signal in a time-frequency domain through all the reverberation energy distributions. The processing module is further configured to perform reverberation suppression on the initial speech signal according to the reverberation suppression gain to obtain a speech signal after reverberation suppression. The processing module is further configured to perform time-frequency mask estimation on the initial speech signal to obtain a time-frequency mask of the initial speech signal, and determine residual interference energy of each time-frequency unit of the initial speech signal in a time-frequency domain according to the time-frequency mask and an interference covariance matrix of a target sound source. The execution module is configured to perform harmonic enhancement on the speech signal after reverberation suppression through all the residual interference energies to obtain a speech enhancement signal.
[0015] The technical scheme provided by the embodiments disclosed in the present application has the following beneficial effects: In the intelligent speech control system and the speech enhancement method in a reverberation environment provided in the present application, an initial speech signal in a reverberation environment is obtained; reverberation energy distribution in each frequency band is determined according to a time-frequency spectrogram of the initial speech signal and a frequency response attenuation curve in each frequency band, and reverberation suppression gain of the initial speech signal in a time-frequency domain is determined through all the reverberation energy distributions; the initial speech signal is subjected to reverberation suppression according to the reverberation suppression gain to obtain a speech signal after reverberation suppression; time-frequency mask estimation is performed on the initial speech signal to obtain a time-frequency mask of the initial speech signal, and residual interference energy of each time-frequency unit of the initial speech signal in a time-frequency domain is determined according to the time-frequency mask and an interference covariance matrix of a target sound source; harmonic enhancement is performed on the speech signal after reverberation suppression through all the residual interference energies to obtain a speech enhancement signal.
[0016] It can be seen that in the present application, the reverberation suppressed speech signal can be subjected to harmonic enhancement by all residual interference energy to obtain a speech enhancement signal; wherein, firstly, by analyzing the time-frequency spectrogram of the initial speech signal and combining the frequency response attenuation curve of each frequency band, the energy distribution characteristics of reverberation in different frequency bands can be extracted, so as to identify the concentration area of reverberation energy in the time-frequency domain. Since acoustic crosstalk often overlaps with useful speech signals in a specific frequency band to form interference that is difficult to separate, this frequency band level reverberation energy distribution analysis can make the suppression process more targeted, effectively weaken the interference energy and preserve the harmonic structure of the original speech as much as possible; secondly, by integrating the reverberation energy distribution of all frequency bands to determine the reverberation suppression gain of the initial speech signal in the time-frequency domain, the strength of the reverberation can be quantified in the full frequency range, and the suppression intensity can be adaptively adjusted in each time-frequency unit. This gain control based on global reverberation energy characteristics can effectively reduce the reflected sound energy while maximizing the preservation of the amplitude and phase information of the original speech, avoiding excessive attenuation of useful harmonic components, thereby significantly reducing the acoustic crosstalk caused by reverberation overlap; then, the initial speech signal is subjected to reverberation suppression using the calculated reverberation suppression gain, which can specifically weaken the energy components of the reflected sound in the time-frequency domain, thereby significantly reducing the interference intensity of the reverberation in each frequency band; then, time-frequency mask estimation is performed on the initial speech signal, which can identify the energy distribution proportion of useful speech and reverberation interference in different time and frequency units in the time-frequency domain. The introduction of the time-frequency mask can effectively distinguish the target signal from the reverberation component and avoid the mis-suppression of useful harmonics; further, the residual interference energy of the initial speech signal in the time-frequency domain is determined in combination with the time-frequency mask and the interference covariance matrix of the target sound source, which can quantify the residual reverberation and its leakage degree in two dimensions of statistical characteristics and time-frequency distribution. Not only can it identify the weak interference area that still overlaps with useful harmonics after reverberation suppression, but also can avoid the loss of harmonic information caused by simple amplitude reduction, thereby effectively reducing the residual acoustic crosstalk caused by reverberation overlap; finally, the reverberation suppressed speech signal is subjected to harmonic enhancement using the residual interference energy, which can specifically restore the speech harmonic components weakened or covered under the action of reverberation interference, especially the mid-high frequency harmonic region which is highly dependent on speech intelligibility and naturalness. This process not only makes up for the possible loss of useful signal energy in the reverberation suppression stage, but also improves the integrity of the harmonic structure and the spectral detail performance, thereby effectively offsetting the damage to the speech harmonics caused by acoustic crosstalk, and finally obtaining an enhanced speech signal; in summary, the scheme of the present application can realize the suppression of acoustic crosstalk caused by reverberation. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 is an exemplary flowchart of a speech enhancement method in a reverberation environment according to some embodiments of the present application; Figure 2 is a flowchart of determining a reverberation suppression gain according to some embodiments of the present application; Figure 3 is a flowchart of determining residual interference energy according to some embodiments of the present application; Figure 4 is a structural diagram of a speech enhancement unit according to some embodiments of the present application; Figure 5 is a structural diagram of a computer device for implementing a speech enhancement method in a reverberation environment according to some embodiments of the present application. DETAILED DESCRIPTION
[0018] In order to better understand the technical solutions of the present application, the technical solutions of the present application will be described in detail below in combination with the drawings of the specification and specific embodiments.
[0019] Reference Figure 1 The figure is an exemplary flowchart of a speech enhancement method in a reverberation environment according to some embodiments of the present application, which mainly includes the following steps: In step 101, an initial speech signal in a reverberation environment is obtained.
[0020] In specific implementation, a microphone array is used to obtain the initial speech signal in the reverberation environment.
[0021] It should be noted that the reverberation environment in the present application refers to an acoustic environment in which sound waves propagate in a closed or semi-closed space, in addition to direct sound, multiple reflections and scattering of sound waves occur on the surfaces of walls, floors, ceilings and objects in the space to form a superimposed sound field.
[0022] In step 102, the reverberation energy distribution in each frequency band is determined according to the time-frequency spectrogram of the initial speech signal and the frequency response attenuation curve in each frequency band, and the reverberation suppression gain of the initial speech signal in the time-frequency domain is determined by all the reverberation energy distributions.
[0023] In some embodiments, the determination of the reverberation energy distribution in each frequency band according to the time-frequency spectrogram of the initial speech signal and the frequency response attenuation curve in each frequency band can be implemented by the following steps: Time-frequency transform is performed on the initial speech signal to obtain a time-frequency spectrogram of the initial speech signal; The reverberation time in each frequency band is determined by the frequency response attenuation curve in each frequency band; The reverberation energy distribution in each frequency band is determined according to the time-frequency spectrogram and the reverberation time in each frequency band.
[0024] It should be noted that the time-frequency spectrogram in the present application refers to a two-dimensional image of the energy distribution characteristics of the initial speech signal in the time and frequency dimensions.
[0025] In a specific implementation, the time-frequency conversion of the initial voice signal to obtain the time-frequency spectrogram of the initial voice signal can be implemented in the following manner: first, the initial voice signal is divided into overlapping short-time frames with a fixed time length (usually 20-30 milliseconds, referred to as frame length), and the overlapping rate can be set to a value between 50% and 75%; second, a Hanning window function is applied to each short-time frame signal; third, a Fourier transform is performed on each short-time frame signal after the windowing to convert the short-time frame signal from the time domain to the frequency domain, thereby obtaining the frequency components and the energy values of each frequency corresponding to each short-time frame signal; fourth, the frequency components and the energy values of each frequency corresponding to the short-time frame signal are arranged in time sequence to form a two-dimensional image with time as the horizontal axis, frequency as the vertical axis, and pixel brightness (or color) representing the energy intensity, and the obtained two-dimensional image is taken as the time-frequency spectrogram of the initial voice signal.
[0026] In some embodiments, the determination of the reverb time in each frequency band through the frequency response decay curve in each frequency band can be implemented in the following steps: for each frequency band; determining the frequency response decay curve in the frequency band; extracting the trend of the frequency response decay curve in the frequency band, and screening out the effective section in which the energy monotonically decays with time in each frequency response decay curve; calculating the time interval experienced from the initial peak energy to the energy decay reference value in the effective section, and taking the obtained time interval as the reverb time in the frequency band.
[0027] It should be noted that the frequency response decay curve in the present application represents the characteristic curve of the energy of sound waves in the frequency band changing with time in a reverberation environment, and the frequency response decay curve presents the dynamic process of the energy of sound waves gradually decaying from direct sound propagation to multiple reflections.
[0028] In a specific implementation, for each frequency band, first, a standard excitation signal is played to the target reverberation environment, the response signal after the environmental reflection is collected by using a microphone array, the signal energy data of the target frequency band is separated from the collected response signal by using a Butterworth band-pass filter, the energy value at each time point is extracted from the signal energy data of the target frequency band, a curve is further plotted with time as the horizontal axis and energy value as the vertical axis, and the obtained curve is taken as the frequency response decay curve in the target frequency band; second, the moving average filtering algorithm is used to extract the trend of the frequency response decay curve in the frequency band, and effective segment screening is performed based on the curve after the trend extraction: the first-order difference of each point on the curve is calculated to determine the energy change trend, and when the difference result in the continuous time segment is all non-positive values (i.e., the energy does not increase with time), the time segment is determined as an effective segment in which the energy monotonically decays with time; then, the initial peak energy in the effective segment is located, that is, the highest energy value in the effective segment (the highest energy value corresponds to the starting time of the reverberation decay), the energy level that decays by 60 dB relative to the initial peak energy is taken as the energy decay reference value, and the energy decay curve along the effective segment is further tracked to find the time point at which the energy first drops to the energy decay reference value; finally, the difference value between the time point corresponding to the initial peak energy and the time point corresponding to the energy decay reference value is calculated, and the obtained difference value is taken as the reverberation time in the frequency band.
[0029] It should be noted that the reverberation time in the present application represents the duration of the reverberation component in the frequency band.
[0030] In a specific implementation, the determination of the reverberation energy distribution in each frequency band based on the time-frequency spectrogram and the reverberation time in each frequency band can be performed in the following manner: first, the energy values of each frequency band at different time points are obtained based on the time-frequency spectrogram (i.e., the total energy in the vertical axis interval of the time-frequency spectrogram corresponding to each frequency band changes with time); second, for each frequency band, the energy decay mode is determined in combination with the reverberation time of the frequency band: if the energy decay speed at a certain time point matches the reverberation time (i.e., the decay rate is slower than the natural decay of the initial speech signal, which is consistent with the "tail" characteristics of reverberation), the part of the energy value is determined as the reverberation component, the total sum of all energy values determined as the reverberation component in the frequency band is calculated, and the proportion of the total energy value in the frequency band is calculated, and the curve of the proportion value changing with time is further taken as the reverberation energy distribution in the frequency band, thereby obtaining the reverberation energy distribution in each frequency band.
[0031] It should be noted that the reverberation energy distribution in the present application represents the distribution characteristics of the energy of the reverberation component in the time dimension in the frequency band.
[0032] In some embodiments, the reference Figure 2As shown in the figure, the figure is a flowchart for determining the reverb suppression gain in some embodiments of the present application. The reverb suppression gain of the initial speech signal in the time-frequency domain can be determined by the following steps in the present embodiment: The reverb energy distribution of each frequency band is merged to obtain a full-frequency domain reverb energy spectrum; The reverb energy spectrum is compared and analyzed with the power spectrum of the initial speech signal to obtain a reverb dominant area and a speech dominant area; The reverb suppression factor of each time-frequency unit of the initial speech signal in the time-frequency domain is determined according to the reverb dominant area and the speech dominant area; The reverb suppression factors of all time-frequency units are gain calculated to obtain the reverb suppression gain of the initial speech signal in the time-frequency domain.
[0033] It should be noted that the frequency domain reverb energy spectrum in the present application represents the overall distribution characteristics of all reverb energy of the initial speech signal in the time-frequency domain.
[0034] In a specific implementation, the reverb energy distribution of each frequency band is merged to obtain a full-frequency domain reverb energy spectrum, which can be implemented in the following way, that is, first, the reverb energy distribution of each frequency band is sorted according to its corresponding frequency range on the vertical axis (for example, arranged from low frequency to high frequency), second, the time axis of all frequency bands is calibrated, third, the linear interpolation method is used to fill the frequency gaps between adjacent frequency bands, then, the reverb energy proportion data of each frequency band at the same time point is integrated according to the frequency order to form a two-dimensional matrix with time as the horizontal axis, frequency as the vertical axis, and data point value representing the reverb energy proportion of the corresponding time-frequency unit, finally, the two-dimensional matrix is visualized (for example, the color depth represents the proportion), and the spectrum obtained by visualizing is used as the full-frequency domain reverb energy spectrum.
[0035] It should be noted that the speech dominant area in the present application represents a set of time-frequency units in which the speech signal energy dominates the total energy (the sum of speech energy and reverb energy) in the time-frequency domain, the reverb dominant area represents a set of time-frequency units in which the reverb energy dominates the total energy in the time-frequency domain, and the time-frequency unit represents the basic unit for dividing the signal in time-frequency analysis, which is a rectangular area defined by the time interval and the frequency bandwidth in the time-frequency two-dimensional plane.
[0036] In a specific implementation, the reverberation energy map and the power spectrum of the initial speech signal are compared and analyzed to obtain the reverberation dominant region and the speech dominant region, which can be achieved in the following manner: first, the reverberation energy map and the power spectrum of the initial speech signal are mapped to the same time-frequency grid (the time and frequency resolutions remain consistent), so that each time-frequency unit can correspond to compare the reverberation energy value of the reverberation energy map and the total energy value of the power spectrum; second, the proportion of the reverberation energy in each time-frequency unit is calculated, and a threshold judgment method commonly used in the prior art is adopted (set the reverberation proportion greater than 50% as the limit), when the reverberation energy proportion of a certain time-frequency unit exceeds the threshold, it is determined that the time-frequency unit is a reverberation dominant region, otherwise, when the reverberation energy proportion is lower than the threshold, it is determined as a speech dominant region.
[0037] It should be noted that the reverberation suppression factor in the present application represents a parameter for quantifying the reverberation attenuation strength of the time-frequency unit in the time-frequency domain.
[0038] In a specific implementation, the reverberation suppression factor of each time-frequency unit of the initial speech signal in the time-frequency domain can be determined according to the reverberation dominant region and the speech dominant region, which can be achieved in the following manner: first, for the time-frequency unit of the speech dominant region, a reverberation suppression factor close to 1 (for example, 0.8-1.0) is used, second, for the time-frequency unit of the reverberation dominant region, a smaller reverberation suppression factor (for example, 0.2-0.5) is used, and finally, for the transition zone (time-frequency unit with energy proportion close to the threshold) between the two types of regions, a Gaussian smoothing algorithm is used to generate a reverberation suppression factor between the above two.
[0039] It should be noted that the reverberation suppression gain in the present application represents a parameter for quantifying the degree of reverberation energy attenuation of each time-frequency unit of the initial speech signal in the time-frequency domain.
[0040] In a specific implementation, the gain calculation is performed on the reverberation suppression factor of all time-frequency units to obtain the reverberation suppression gain of the initial speech signal in the time-frequency domain, which can be achieved in the following manner: first, the reverberation suppression factor of all time-frequency units is directly used as linear gain, that is, the value of the reverberation suppression factor is the energy proportion that needs to be preserved by the time-frequency unit (for example, the reverberation suppression factor is 0.2, which means only 20% of the energy is preserved, and the reverberation suppression factor is 0.9, which means 90% of the energy is preserved), second, the reverberation suppression factor of each time-frequency unit is multiplied with the time-frequency amplitude of the initial speech signal in the corresponding time-frequency unit, and all the multiplied values are used as the suppression gain value of each time-frequency unit, thereby completing the gain application, and finally, all time-frequency units after multiplication form an overall gain distribution, and the obtained overall gain distribution is used as the reverberation suppression gain of the initial speech signal in the time-frequency domain.
[0041] In step 103, reverberation suppression is performed on the initial speech signal according to the reverberation suppression gain to obtain a reverberation suppressed speech signal.
[0042] In some embodiments, performing reverberation suppression on the initial speech signal according to the reverberation suppression gain to obtain the reverberation suppressed speech signal may be achieved by using the following steps: Determine the reverberation threshold value in each frequency band according to the reverberation energy distribution in each frequency band; The time-frequency coefficient of each time-frequency unit of the initial speech signal in the time-frequency domain is suppressed by the reverberation suppression gain and the reverberation threshold value on each frequency band, thereby obtaining a speech signal after reverberation suppression.
[0043] It should be noted that the reverberation threshold value mentioned in the present application represents the critical energy value of the reverberation energy of the time-frequency unit within the determination frequency band that needs to be suppressed.
[0044] In specific implementation, the reverberation threshold value in each frequency band can be determined based on the reverberation energy distribution in each frequency band. This can be achieved in the following way: for each frequency band, the reverberation energy values at multiple typical moments (such as the energy at different stages such as speech gaps and speech tails) are uniformly sampled from the reverberation energy distribution in the frequency band. Then, the median of these sampled values is calculated, and the median is multiplied by an empirical adjustment coefficient (for example, 1.2 times, which can be flexibly set according to the degree of interference of reverberation on speech in actual scenarios). The multiplied value is used as the reverberation threshold value in the frequency band, thereby obtaining the reverberation threshold value in each frequency band.
[0045] In a specific implementation, the time-frequency coefficient of each time-frequency unit of the initial speech signal in the time-frequency domain is suppressed by the reverberation suppression gain and the reverberation threshold value on each frequency band, and the speech signal after reverberation suppression is obtained can be implemented in the following manner, namely: first, the time-frequency coefficient of each time-frequency unit of the initial speech signal in the time-frequency domain is extracted by short-time Fourier transform; secondly, for each time-frequency unit, the suppression gain value of the time-frequency unit is obtained from the reverberation suppression gain; the suppression gain value of the time-frequency unit is compared with the suppression gain value of the time-frequency unit; The basic attenuation is completed by multiplying the time-frequency coefficients on the element. For the time-frequency coefficient after basic attenuation, it is judged whether the reverberation energy of the frequency band to which it belongs still exceeds the reverberation threshold value in the frequency band. If it exceeds, it means that the reverberation suppression is insufficient. The amplitude of the time-frequency coefficient is multiplied by an enhanced attenuation coefficient (for example, 0.3-0.5). If it does not exceed, the basic attenuation result is maintained. Finally, the amplitudes of all time-frequency units after processing are recombined with the original phases, and converted back to the time domain waveform through inverse short-time Fourier transform to obtain the speech signal after reverberation suppression.
[0046] In step 104, time-frequency mask estimation is performed on the initial speech signal to obtain a time-frequency mask of the initial speech signal, and residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain is determined according to the time-frequency mask and an interference covariance matrix of a target sound source.
[0047] In some embodiments, time-frequency mask estimation performed on the initial speech signal to obtain a time-frequency mask of the initial speech signal can be implemented by the following steps: Detecting, by a pre-trained voice activity detection model, a speech presence probability of each time-frequency unit of the initial speech signal in the time-frequency domain; Determining the time-frequency mask of the initial speech signal according to all speech presence probabilities.
[0048] In a specific implementation, detecting, by a pre-trained voice activity detection model, a speech presence probability of each time-frequency unit of the initial speech signal in the time-frequency domain can be implemented by the following manner, that is: It should be noted that the pre-trained voice activity detection model in the present application is a machine learning model trained by large-scale labeled data (containing speech segments, silence segments, noise segments, and diversified scene samples), and its core function is to determine whether there is speech in the input signal and the start and end time of the speech. In the prior art, a deep learning architecture combining a convolutional neural network and a long short-term memory network is usually used, wherein the convolutional neural network layer can extract local time-frequency features (such as a speech energy pattern of a specific frequency combination) of the signal, and the long short-term memory network layer can capture the context association in the time dimension (for example, the continuity before and after the speech). The voice activity detection model usually inputs features such as a mel spectrum and a short-time energy obtained by preprocessing the speech signal, and outputs a speech presence probability (0-1 range) of each time unit (or time-frequency unit). After training, the voice activity detection model can be generalized to different noise and reverberation environments, so as to accurately distinguish speech and non-speech components.
[0049] It should also be noted that the speech presence probability in the present application represents a numerical index quantifying the possibility of containing valid speech components in each time-frequency unit.
[0050] In a specific implementation, the following method can be used to detect the speech presence probability of each time-frequency unit of the initial speech signal in the time-frequency domain using a pre-trained speech activity detection model: first, the initial speech signal is framed and windowed, and a time-frequency spectrum (including amplitude information of each time frame and frequency point) is obtained through short-time Fourier transform. The time-frequency spectrum is then converted into the input format of the speech activity detection model (for example, Mel spectrum features are extracted, the frequency axis is mapped to a Mel scale that conforms to human ear perception, and 40-80 dimensional features are obtained). Secondly, the processed features are input into the pre-trained speech activity detection model; the speech activity detection model outputs a numerical value in the range of 0-1 for each time-frequency unit, and the obtained value is used as the speech presence probability of each time-frequency unit.
[0051] It should be noted that the time-frequency mask described in this application represents a weight matrix that marks the proportion of speech components in the time-frequency domain of the initial speech signal.
[0052] In specific implementation, the time-frequency mask of the initial speech signal is determined based on all speech presence probabilities. This can be achieved in the following manner: first, the intervals are divided according to the distribution characteristics of the speech presence probability (for example, 0-0.3, 0.3-0.7, 0.7-1.0), and each interval corresponds to a different mask value level. Then, a lower mask value (for example, 0.1-0.2) is assigned to the time-frequency unit in the 0-0.3 interval, a medium mask value (for example, 0.4-0.6) is assigned to the time-frequency unit in the 0.3-0.7 interval, and a higher mask value (for example, 0.8-1.0) is assigned to the time-frequency unit in the 0.7-1.0 interval. The time-frequency mask finally generated is completely aligned with the signal time-frequency structure. The mask values of all time-frequency units are combined according to the corresponding time-frequency positions to form a matrix, and the obtained matrix is used as the time-frequency mask of the initial speech signal.
[0053] In some embodiments, reference Figure 3 As shown in the figure, this figure is a schematic diagram of the process of determining the residual interference energy in some embodiments of the present application. In this embodiment, the residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain is determined according to the time-frequency mask and the interference covariance matrix of the target sound source, which can be achieved by the following steps: First, in step 1041, the interference covariance matrix of the target sound source is determined; Next, in step 1042, a valid region of the initial speech signal in the time-frequency domain is identified according to the time-frequency mask; Then, in step 1043, residual interference analysis is performed on the effective area of the initial speech signal in the time-frequency domain using the interference covariance matrix of the target sound source to obtain the residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain.
[0054] It should be noted that the interference covariance matrix in the present application represents the statistical correlation characteristics of each frequency component of the interference signal in the time-frequency domain in the environment where the target sound source is located.
[0055] In a specific implementation, the interference covariance matrix of the target sound source can be determined in the following manner: first, the initial speech signal is input into a pre-trained speech activity detection model, the time-frequency characteristics of the initial speech signal are analyzed by the speech activity detection model, and the speech activity state of each time frame, "voiced speech" (containing target speech energy) or "unvoiced speech" (without target speech energy), is output. Based on the output result, all "unvoiced speech" time periods are marked, and all "unvoiced speech" time period signals are regarded as non-occurred speech signals (such signals are generated during user silence period, mainly containing reverberation residues, environmental noise, and other pure interference components, without effective speech energy mixed in). The complex time-frequency coefficients of each time frame and each frequency point are extracted from all non-occurred speech signals, all complex time-frequency coefficients are integrated into a two-dimensional data matrix according to the "time-frequency" dimension, the matrix is taken as an interference sample, and then the sample covariance formula is used to calculate the linear correlation degree between different time-frequency units with the interference sample as input, so as to generate a matrix as the interference covariance matrix of the target sound source.
[0056] In a pure interference scene without target speech (for example, only environmental noise, other sound sources, and time periods), the collected interference signals are framed, windowed, and short-time Fourier transformed to obtain complex time-frequency coefficients of the interference signals in the time-frequency domain (arranged as a matrix according to the frequency dimension, each row corresponds to a frequency point, and each column corresponds to a time frame). Then, based on the calculation method of the covariance matrix in the prior art, the time-frequency coefficient matrix is statistically averaged according to the time dimension. Specifically, the cross-correlation values of the time-frequency coefficients of the interference signals of different frequency points are calculated to form a square matrix with the same number of dimensions as the frequency points (the diagonal elements are the energies of the interference signals of the frequency points, and the non-diagonal elements are the correlations of the interference signals of different frequency points). The obtained square matrix is taken as the interference covariance matrix of the target sound source.
[0057] It should be noted that the effective area in the present application represents a set of time-frequency units in the time-frequency domain that are determined to be dominated by target speech components.
[0058] In a specific implementation, the identification of the valid region of the initial speech signal in the time-frequency domain according to the time-frequency mask can be implemented in the following manner: first, based on the statistical distribution characteristics (for example, the probability density function of the mask value) of the time-frequency mask, a plurality of decision thresholds (for example, a low threshold of 0.3 and a high threshold of 0.7) are determined; then, all time-frequency units are traversed, and time-frequency units with a mask value greater than or equal to 0.7 are marked as “strong valid units”, time-frequency units with a mask value greater than 0.3 and less than 0.7 are marked as “weak valid units”, and weak valid units with a mask value less than 0.3 are marked as “invalid units”; subsequently, an expansion operation in morphological filtering is used to expand the “strong valid unit” region so as to include adjacent “weak valid units”; finally, a continuous time-frequency region set is obtained after expansion, and the continuous time-frequency region set is taken as the valid region of the initial speech signal in the time-frequency domain.
[0059] It should be noted that the residual interference energy in the present application represents the interference energy caused by the leakage of reverberation sound energy to the time-frequency units in the time-frequency domain of the initial speech signal to the frequency components of the initial speech signal.
[0060] In a specific implementation, the residual interference analysis of the valid region of the initial speech signal in the time-frequency domain by the interference covariance matrix of the target sound source to obtain the residual interference energy of each time-frequency unit in the time-frequency domain of the initial speech signal can be implemented in the following manner: first, the complex time-frequency coefficients of all time-frequency units in the valid region are extracted; then, the interference energy projection calculation is performed on the complex time-frequency coefficients of each time-frequency unit in the valid region by using the interference covariance matrix of the target sound source, specifically, the quadratic form operation (that is, the conjugate transpose of the complex time-frequency coefficient multiplied by the interference covariance matrix and then multiplied by the complex time-frequency coefficient) is performed on the complex time-frequency coefficient and the interference covariance matrix, and the component obtained after the quadratic form operation is taken as the energy component originating from the interference in the time-frequency unit, and the obtained energy component is taken as the residual interference energy of the time-frequency unit in the time-frequency domain of the initial speech signal, thereby obtaining the residual interference energy of each time-frequency unit in the time-frequency domain of the initial speech signal.
[0061] In step 105, the speech signal after reverberation suppression is subjected to harmonic enhancement by all the residual interference energies, and a speech enhanced signal is obtained.
[0062] In some embodiments, the speech signal after reverberation suppression can be subjected to harmonic enhancement by all the residual interference energies to obtain a speech enhanced signal in the following steps: The interference region corresponding to the harmonic structure of the speech signal after reverberation suppression in the time-frequency domain is identified by all the residual interference energies; The amplitude of the interference region corresponding to the harmonic structure in the time-frequency domain is enhanced, and a speech enhanced signal is obtained.
[0063] It should be noted that the interference region in the present application represents an area in which acoustic crosstalk in a reverberation environment overlaps with the harmonic structure of the speech signal in the time-frequency domain after reverberation suppression.
[0064] In a specific implementation, the interference region corresponding to the harmonic structure of the speech signal in the time-frequency domain after reverberation suppression can be identified by all residual interference energy in the following manner: first, frame the speech signal after reverberation suppression, and then convert each signal frame into a time-frequency matrix in the time-frequency domain by short-time Fourier transform (each element in the time-frequency matrix corresponds to the signal energy of a specific time frame and frequency point, and completely corresponds to the time-frequency unit of the initial speech signal in the time and frequency coordinates); second, match the residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain to the time-frequency matrix according to the same coordinates to form interference energy data of each time-frequency unit; third, generate an interference energy distribution heat map by mapping the interference energy value to a color gradient from low to high; fourth, use the inherent characteristics of the speech signal after reverberation suppression, i.e., the speech is generated by the periodic vibration of the vocal cords, and the frequency components contain integer multiples of the fundamental frequency (i.e., harmonics), which appear as continuous energy bands parallel to the time axis in the time-frequency domain (for example, when the fundamental frequency is 150 Hz, 300 Hz, 450 Hz, etc. will form corresponding harmonic bands); fifth, extract the fundamental frequency from the initial speech signal by autocorrelation method, and then determine the theoretical frequency range of each order harmonic band; and finally, superimpose the interference energy heat map and the theoretical frequency range of each order harmonic band, screen out the time-frequency units that are both within the harmonic band range and satisfy the residual interference energy exceeding a preset threshold (the threshold can be set as the average of all residual interference energy), and form a set of time-frequency units, which are the interference region corresponding to the harmonic structure of the speech signal in the time-frequency domain after reverberation suppression.
[0065] In a specific implementation, the amplitude of the interference region corresponding to the harmonic structure in the time-frequency domain is enhanced, and then the speech enhancement signal is obtained by using the following method, that is, first, the reference features of the harmonic structure are extracted from the initial speech signal, the time-frequency matrix of the non-interference region is obtained by short-time Fourier transform, the amplitude distribution law of each order harmonic (fundamental frequency and integer multiple frequency) in the normal state is counted, including the amplitude change gradient (for example, the amplitude difference of adjacent frames is not more than 15%) of the same harmonic in the continuous time frame and the energy ratio relationship between different harmonics; second, for the identified harmonic structure interference region, the corresponding reference features are matched according to the harmonic order to which the interference region belongs: for the interference region of the continuous time frame, the sliding window average method is used, the amplitude mean value of the same order harmonic of the non-interference region in front and back 3-5 frames is taken as the reference, and the amplitude in the interference region is adjusted to 90%-95% of the reference by linear interpolation, for the isolated interference unit, the amplitude in the interference unit is increased to more than 85% of the average value of the amplitude of the same order harmonic of the adjacent 2-3 non-interference frequency points in the same time frame, and finally, the enhanced interference region time-frequency data and the time-frequency data of the non-interference region are fused, that is, the time-frequency unit at the boundary of the interference region (that is, 1-2 columns of time frames or 1-2 rows of frequency points adjacent to the non-interference region) is processed by weighted smoothing, and the weight is distributed according to the distance (for example, the time-frequency unit closer to the interference region is given higher weight of the enhanced data, and the time-frequency unit closer to the non-interference region is given higher weight of the original non-interference data), for the non-boundary region, the enhanced interference region data and the original non-interference region data are directly retained, and then the inverse short-time Fourier transform is performed on the fused complete time-frequency matrix, and the time domain signal obtained by the inverse short-time Fourier transform is taken as the speech enhancement signal.
[0066] In addition, another aspect of the present application, in some embodiments, the present application provides a speech control system, which comprises a speech enhancement unit, which is used for Figure 4 The figure is a structural schematic diagram of a speech enhancement unit according to some embodiments of the present application, which comprises an acquisition module 401, a processing module 402 and an execution module 403, which are described as follows: The acquisition module 401 is mainly used for acquiring the initial speech signal in the reverberation environment in the present application; The processing module 402 is used for determining the reverberation energy distribution in each frequency band according to the time-frequency spectrum of the initial speech signal and the frequency response attenuation curve in each frequency band, and determining the reverberation suppression gain of the initial speech signal in the time-frequency domain through all the reverberation energy distributions in the present application; It should be noted that the processing module 402 is further configured to perform reverb suppression on the initial speech signal according to the reverb suppression gain to obtain a speech signal after reverb suppression. It should be noted that the processing module 402 is further configured to perform time-frequency mask estimation on the initial speech signal to obtain a time-frequency mask of the initial speech signal, and determine residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain according to the time-frequency mask and an interference covariance matrix of a target sound source. The execution module 403 is mainly configured to perform harmonic enhancement on the speech signal after reverb suppression by all the residual interference energy to obtain a speech enhancement signal.
[0067] In addition, the present application further provides a computer device, which comprises a memory and a processor, the memory stores code, and the processor is configured to acquire the code and execute the speech enhancement method in a reverberation environment.
[0068] In some embodiments, the reference Figure 5 The figure is a structural schematic diagram of a computer device for implementing the speech enhancement method in a reverberation environment according to some embodiments of the present application. The speech enhancement method in a reverberation environment in the above embodiments can be implemented by the computer device shown in the figure, which comprises at least one processor 501, a communication bus 502, a memory 503, and at least one communication interface 504. Figure 5
[0069] The processor 501 can be a general central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more circuits for controlling the execution of the speech enhancement method in a reverberation environment in the present application.
[0070] The communication bus 502 can be used to transmit information between the above components.
[0071] The memory 503 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM), or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magnetic disk storage or other magnetic storage devices, or any other medium capable of storing desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited to this. The memory 503 can exist independently, and is connected to the processor 501 through the communication bus 502. The memory 503 can also be integrated with the processor 501.
[0072] The memory 503 is configured to store program codes for implementing the solutions of the present application, and the processor 501 is configured to control the execution of the program codes. The processor 501 is configured to execute the program codes stored in the memory 503. The program codes can include one or more software modules. The methods described in the above method embodiments can be implemented by the processor 501 and one or more software modules in the program codes in the memory 503.
[0073] The communication interface 504 is configured to communicate with other devices or communication networks, such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc., using any transceiver-like device.
[0074] In specific implementations, as an embodiment, the computer device can include multiple processors, each of which can be a single-CPU processor or a multi-CPU processor. The processor herein can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0075] The computer device described above can be a general-purpose computer device or a special-purpose computer device. In a specific implementation, the computer device can be a desktop computer, a laptop computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. The embodiments of the present application do not limit the type of the computer device.
[0076] In addition, the present application also provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the voice enhancement method in a reverberation environment.
[0077] Although the preferred embodiments of the present application have been described, those skilled in the art who are informed of the basic inventive concept can make additional changes and modifications to the embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0078] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.
Claims
1. A method for speech enhancement in a reverberant environment, used for speech enhancement in a speech control system, characterized in that: The method comprises the following steps: Obtain the initial speech signal in a reverberant environment; Determining the reverberation energy distribution in each frequency band according to the time-frequency spectrum of the initial speech signal and the frequency response attenuation curve in each frequency band, and determining the reverberation suppression gain of the initial speech signal in the time-frequency domain through all the reverberation energy distributions; Performing reverberation suppression on the initial speech signal according to the reverberation suppression gain to obtain a reverberation suppressed speech signal; Performing time-frequency mask estimation on the initial speech signal to obtain a time-frequency mask of the initial speech signal, and determining the residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain according to the time-frequency mask and the interference covariance matrix of the target sound source; The speech signal after reverberation suppression is harmonically enhanced using all the residual interference energy to obtain a speech enhanced signal.
2. The method according to claim 1, wherein Determining the reverberation energy distribution in each frequency band according to the time-spectrogram of the initial speech signal and the frequency response attenuation curve in each frequency band specifically includes: Performing a time-frequency transformation on the initial speech signal to obtain a time-frequency spectrum of the initial speech signal; Determine the reverberation time in each frequency band through the frequency response attenuation curve in each frequency band; The reverberation energy distribution in each frequency band is determined according to the time-frequency spectrum diagram and the reverberation time in each frequency band.
3. The method according to claim 2, wherein Determining the reverberation time in each frequency band through the frequency response attenuation curve in each frequency band specifically includes: For each frequency band; Determine the frequency response attenuation curve within the frequency band; Extract the trend of the frequency response attenuation curve within the frequency band and screen out the effective section where the energy monotonically decays over time in each frequency response attenuation curve; Calculate the time interval from the initial peak energy dropping to the energy attenuation reference value within the effective segment, and use the obtained time interval as the reverberation time within the frequency band.
4. The method according to claim 1, wherein Determining the reverberation suppression gain of the initial speech signal in the time-frequency domain through all reverberation energy distributions specifically includes: The reverberation energy distribution on each frequency band is combined to obtain the full-frequency reverberation energy spectrum; Compare and analyze the reverberation energy spectrum with the power spectrum of the initial speech signal to obtain the reverberation-dominated area and the speech-dominated area; Determining a reverberation suppression factor of each time-frequency unit of the initial speech signal in the time-frequency domain according to the reverberation-dominant area and the speech-dominant area; Gain calculation is performed on the reverberation suppression factors of all time-frequency units to obtain the reverberation suppression gain of the initial speech signal in the time-frequency domain.
5. The method according to claim 1, wherein Performing reverberation suppression on the initial speech signal according to the reverberation suppression gain to obtain a speech signal after reverberation suppression specifically includes: Determine the reverberation threshold value in each frequency band according to the reverberation energy distribution in each frequency band; The time-frequency coefficient of each time-frequency unit of the initial speech signal in the time-frequency domain is suppressed by the reverberation suppression gain and the reverberation threshold value on each frequency band, thereby obtaining a speech signal after reverberation suppression.
6. The method according to claim 1, wherein Performing time-frequency mask estimation on the initial speech signal to obtain the time-frequency mask of the initial speech signal specifically includes: Detecting the speech presence probability of each time-frequency unit of the initial speech signal in the time-frequency domain through a pre-trained speech activity detection model; The time-frequency mask of the initial speech signal is determined according to all speech existence probabilities.
7. The method according to claim 1, wherein Determining the residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain according to the time-frequency mask and the interference covariance matrix of the target sound source specifically includes: Determine the interference covariance matrix of the target sound source; Identifying a valid area of the initial speech signal in the time-frequency domain according to the time-frequency mask; Interference leakage estimation is performed on the effective area of the initial speech signal in the time-frequency domain using the interference covariance matrix of the target sound source to obtain the residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain.
8. The method according to claim 1, wherein The speech signal after reverberation suppression is harmonically enhanced by using all the residual interference energy to obtain the speech enhancement signal, which specifically includes: Identify the interference area corresponding to the harmonic structure of the speech signal after reverberation suppression in the time-frequency domain through all the residual interference energy; The amplitude of the interference area corresponding to the harmonic structure in the time-frequency domain is enhanced to obtain a speech enhancement signal.
9. The method according to claim 1, wherein The reverberation environment refers to the acoustic environment in which, when sound waves propagate in a closed or semi-enclosed space, in addition to the direct sound, multiple reflections and scatterings occur on the walls, floor, ceiling and surfaces of objects in the space to form a superimposed sound field.
10. A voice control system, comprising a voice enhancement unit, characterized in that: The speech enhancement unit comprises: An acquisition module is used to obtain an initial speech signal in a reverberant environment; a processing module, configured to determine the reverberation energy distribution in each frequency band based on the time-frequency spectrum of the initial speech signal and the frequency response attenuation curve in each frequency band, and determine the reverberation suppression gain of the initial speech signal in the time-frequency domain based on all the reverberation energy distributions; The processing module is further configured to perform reverberation suppression on the initial speech signal according to the reverberation suppression gain to obtain a reverberation suppressed speech signal; The processing module is further configured to perform time-frequency mask estimation on the initial speech signal to obtain a time-frequency mask of the initial speech signal, and determine the residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain according to the time-frequency mask and the interference covariance matrix of the target sound source; The execution module is used to perform harmonic enhancement on the speech signal after reverberation suppression by using all the residual interference energy to obtain a speech enhancement signal.
Citation Information
Patent Citations
Speech enhancement method applied to dual-microphone array
CN105788607A
Speech enhancement method and system, computer equipment and storage medium
CN110503972A
Microphone array beam forming method
CN110931036A
Reverberation voice reverberation suppression method and device
CN112687284A
Speech enhancement model training and application method, device and equipment, equipment and storage medium
CN113436643A
Cited By
User input audio low-distortion processing method for multi-source noise environment
CN121617411A