Intelligent voice control system and voice enhancement method in reverberation environment

By analyzing the time-frequency characteristics and interference covariance matrix of speech signals under reverberation conditions, targeted suppression of reverberation and harmonic enhancement are achieved, solving the acoustic crosstalk problem of speech signals under reverberation conditions and improving the accuracy of speech recognition systems.

CN120808806BActive Publication Date: 2025-11-07CHENGDU POLYTECHNIC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511264320.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-11-07
Estimated Expiration
2045-09-05

AI Technical Summary

Technical Problem

In a reverberant environment, acoustic crosstalk occurs due to the superposition of reflected sound in the speech signal, affecting the accuracy of the speech recognition system. Existing technologies are unable to effectively suppress acoustic crosstalk caused by reverberation.

Method used

By acquiring the initial speech signal under reverberation, analyzing its time-frequency spectrum and frequency response attenuation curve, determining the reverberation energy distribution, calculating the reverberation suppression gain, and combining the time-frequency mask and interference covariance matrix, reverberation suppression and harmonic enhancement are performed to obtain the speech enhancement signal.

Benefits of technology

It significantly reduces acoustic crosstalk caused by reverberation superposition, preserves speech harmonic structure, improves speech intelligibility and recognition accuracy, and effectively suppresses interference caused by reverberation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808806B_ABST
    Figure CN120808806B_ABST
Patent Text Reader

Abstract

The application provides a kind of intelligent voice control system and voice enhancement method in reverberation environment, by obtaining initial voice signal in reverberation environment;According to the time-frequency spectrum of initial voice signal and the frequency response attenuation curve in each frequency band, the reverberation energy distribution in each frequency band is determined, and the reverberation suppression gain of initial voice signal in time-frequency domain is determined through all the reverberation energy distribution;According to the reverberation suppression gain, the initial voice signal is subjected to reverberation suppression, and the voice signal after reverberation suppression is obtained;Time-frequency mask estimation is carried out on the initial voice signal, and the time-frequency mask of the initial voice signal is obtained, and the residual interference energy of each time-frequency unit in the initial voice signal in time-frequency domain is determined according to the time-frequency mask and the interference covariance matrix of target sound source;Through all the residual interference energy, the voice signal after reverberation suppression is subjected to harmonic enhancement, and the voice enhancement signal is obtained. By adopting the scheme of the application, the acoustic crosstalk caused by reverberation can be suppressed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of voice control, more particularly, the present application relates to an intelligent voice control system and a voice enhancement method in a reverberation environment. BACKGROUND

[0002] Voice control is an intelligent interactive way of realizing device control by analyzing human voice instructions. With its convenient and natural characteristics, it has been deeply integrated into multiple scenarios such as smart home, smart wear, and vehicle-mounted systems. The core is to accurately capture and recognize voice information to perform corresponding operations. However, in actual applications, it often faces interference such as noise, echo, and too far distance, which leads to distortion of voice signals and affects recognition effect. Voice enhancement technology, as an important support, can effectively filter environmental interference and strengthen effective voice features through noise reduction, dereverberation, and signal clarity enhancement algorithms, so as to maintain high accuracy and reliability of voice control technology in complex acoustic environments and provide solid protection for natural and smooth human-computer interaction.

[0003] In practical application scenarios such as voice communication and voice recognition, voice signals often form reverberation due to reflection of walls, furniture and other objects in the propagation environment, resulting in superposition of original voice and reflected sound, producing acoustic crosstalk. This acoustic crosstalk can blur the harmonic structure of the voice, mask the key frequency components, not only reducing the clarity and intelligibility of the voice, but also seriously affecting the accuracy of the voice recognition system (for example, spectral distortion caused by reverberation increases the recognition error rate by more than 30%). Although existing reverberation suppression techniques can weaken part of the reflected sound energy through linear filtering, spectral subtraction and other methods, due to the complexity of the environment and the non-stationarity of the signal, some acoustic crosstalk will still remain after processing, especially in the frequency band where the voice harmonics are concentrated. The superposition of crosstalk and useful signals will form an interference area that is difficult to separate, resulting in harmonic energy attenuation and voice naturalness degradation. Therefore, how to suppress the acoustic crosstalk caused by reverberation has become a problem faced by the industry. SUMMARY

[0004] The present application provides an intelligent voice control system and a voice enhancement method in a reverberation environment, which can suppress acoustic crosstalk caused by reverberation.

[0005] In a first aspect, the present application provides a voice enhancement method in a reverberation environment, which is used for voice enhancement in a voice control system. The method includes the following steps:

[0006] Obtaining an initial voice signal in a reverberation environment;

[0007] Determining the reverberation energy distribution in each frequency band according to the time-frequency spectrum of the initial voice signal and the frequency response attenuation curve in each frequency band, and determining the reverberation suppression gain of the initial voice signal in the time-frequency domain through all the reverberation energy distributions;

[0008] reverberation suppression gain to the initial speech signal to obtain a speech signal after reverberation suppression;

[0009] performing time-frequency mask estimation on the initial speech signal to obtain a time-frequency mask of the initial speech signal, and determining residual interference energy of each time-frequency unit of the initial speech signal in a time-frequency domain according to the time-frequency mask and an interference covariance matrix of a target sound source;

[0010] performing harmonic enhancement on the speech signal after reverberation suppression through all the residual interference energy to obtain a speech enhancement signal.

[0011] In some embodiments, determining the reverberation energy distribution in each frequency band according to the time-frequency spectrogram of the initial speech signal and the frequency response decay curve in each frequency band specifically comprises:

[0012] performing time-frequency transformation on the initial speech signal to obtain a time-frequency spectrogram of the initial speech signal;

[0013] determining the reverberation time in each frequency band through the frequency response decay curve in each frequency band;

[0014] determining the reverberation energy distribution in each frequency band according to the time-frequency spectrogram and the reverberation time in each frequency band.

[0015] In some embodiments, determining the reverberation time in each frequency band through the frequency response decay curve in each frequency band specifically comprises:

[0016] for each frequency band;

[0017] determining the frequency response decay curve in the frequency band;

[0018] performing trend extraction on the frequency response decay curve in the frequency band to filter out an effective section in which energy monotonically decays with time in each frequency response decay curve;

[0019] calculating a time interval experienced from an initial peak energy to an energy decay reference value in the effective section, and taking the obtained time interval as the reverberation time in the frequency band.

[0020] In some embodiments, determining the reverberation suppression gain of the initial speech signal in the time-frequency domain through all the reverberation energy distributions specifically comprises:

[0021] merging the reverberation energy distributions on each frequency band to obtain a full-frequency-domain reverberation energy spectrogram;

[0022] comparing and analyzing the reverberation energy spectrogram with a power spectrum of the initial speech signal to obtain a reverberation dominant region and a speech dominant region;

[0023] determine a reverb suppression factor of each time-frequency unit of the initial speech signal in the time-frequency domain according to the reverb dominant region and the speech dominant region;

[0024] perform gain calculation on the reverb suppression factors of all time-frequency units to obtain a reverb suppression gain of the initial speech signal in the time-frequency domain.

[0025] In some embodiments, perform reverb suppression on the initial speech signal according to the reverb suppression gain to obtain a reverb-suppressed speech signal, specifically including:

[0026] determine a reverb threshold value in each frequency band according to the reverb energy distribution in the frequency band;

[0027] suppress the time-frequency coefficients of each time-frequency unit of the initial speech signal in the time-frequency domain by the reverb suppression gain and the reverb threshold value in each frequency band, to further obtain a reverb-suppressed speech signal.

[0028] In some embodiments, perform time-frequency mask estimation on the initial speech signal to obtain a time-frequency mask of the initial speech signal, specifically including:

[0029] detect the speech presence probability of each time-frequency unit of the initial speech signal in the time-frequency domain by the pre-trained speech activity detection model;

[0030] determine the time-frequency mask of the initial speech signal according to all speech presence probabilities.

[0031] In some embodiments, determine the residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain according to the time-frequency mask and the interference covariance matrix of the target sound source, specifically including:

[0032] determine the interference covariance matrix of the target sound source;

[0033] identify the effective region of the initial speech signal in the time-frequency domain according to the time-frequency mask;

[0034] perform interference leakage estimation on the effective region of the initial speech signal in the time-frequency domain by the interference covariance matrix of the target sound source to obtain the residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain.

[0035] In some embodiments, perform harmonic enhancement on the reverb-suppressed speech signal by all residual interference energies to obtain a speech-enhanced signal, specifically including:

[0036] identify the interference region corresponding to the harmonic structure of the reverb-suppressed speech signal in the time-frequency domain by all residual interference energies;

[0037] The amplitude of the interference region corresponding to the harmonic structure in the time-frequency domain is enhanced, and then a speech enhancement signal is obtained.

[0038] In some embodiments, the reverberation environment represents an acoustic environment in which, in addition to direct sound, sound waves propagate in a closed or semi-closed space, multiple reflections and scattering occur on the surfaces of walls, floors, ceilings and objects in the space, and a superimposed sound field is formed.

[0039] In a second aspect, the present application provides a speech control system, which comprises a speech enhancement unit, and the speech enhancement unit comprises:

[0040] An acquisition module is configured to acquire an initial speech signal in a reverberation environment.

[0041] A processing module is configured to determine a reverberation energy distribution in each frequency band according to a time-frequency spectrogram of the initial speech signal and a frequency response attenuation curve in each frequency band, and determine a reverberation suppression gain of the initial speech signal in the time-frequency domain through all the reverberation energy distributions.

[0042] The processing module is further configured to perform reverberation suppression on the initial speech signal according to the reverberation suppression gain, and obtain a speech signal after reverberation suppression.

[0043] The processing module is further configured to perform time-frequency mask estimation on the initial speech signal, obtain a time-frequency mask of the initial speech signal, and determine residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain according to the time-frequency mask and an interference covariance matrix of a target sound source.

[0044] An execution module is configured to perform harmonic enhancement on the speech signal after reverberation suppression through all the residual interference energies, and obtain a speech enhancement signal.

[0045] The technical scheme provided by the embodiments of the present application has the following beneficial effects:

[0046] In the intelligent speech control system and the speech enhancement method in a reverberation environment provided by the present application, an initial speech signal in a reverberation environment is acquired, a reverberation energy distribution in each frequency band is determined according to a time-frequency spectrogram of the initial speech signal and a frequency response attenuation curve in each frequency band, a reverberation suppression gain of the initial speech signal in the time-frequency domain is determined through all the reverberation energy distributions, reverberation suppression is performed on the initial speech signal according to the reverberation suppression gain, a speech signal after reverberation suppression is obtained, time-frequency mask estimation is performed on the initial speech signal, a time-frequency mask of the initial speech signal is obtained, residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain is determined according to the time-frequency mask and an interference covariance matrix of a target sound source, and harmonic enhancement is performed on the speech signal after reverberation suppression through all the residual interference energies, and a speech enhancement signal is obtained.

[0047] It can be seen that, in the present application, the reverberation-suppressed speech signal can be subjected to harmonic enhancement by all residual interference energy to obtain an enhanced speech signal. Firstly, the energy distribution characteristics of reverberation in different frequency bands can be extracted by analyzing the time-frequency spectrogram of the initial speech signal and combining the frequency response attenuation curves of each frequency band, so as to identify the concentrated area of reverberation energy in the time-frequency domain. Since acoustic crosstalk often overlaps with useful speech signals in specific frequency bands to form interference that is difficult to separate, this frequency-band-level reverberation energy distribution analysis can make the suppression process more targeted, effectively weaken the interference energy and preserve the harmonic structure of the original speech as much as possible. Secondly, the reverberation suppression gain of the initial speech signal in the time-frequency domain can be determined by integrating the reverberation energy distribution of all frequency bands, which can quantify the strength of reverberation in the full frequency range and adaptively adjust the suppression intensity in each time-frequency unit. This gain control based on global reverberation energy characteristics can effectively reduce the reflected sound energy while maximizing the preservation of the amplitude and phase information of the original speech, avoiding excessive attenuation of useful harmonic components, thereby significantly reducing the acoustic crosstalk caused by reverberation overlap. Then, the initial speech signal can be subjected to reverberation suppression using the calculated reverberation suppression gain, which can specifically weaken the energy components of reflected sound in the time-frequency domain, thereby significantly reducing the interference intensity of reverberation in each frequency band. Then, time-frequency mask estimation can be performed on the initial speech signal to identify the energy distribution proportion of useful speech and reverberation interference in different time and frequency units in the time-frequency domain. The introduction of the time-frequency mask can effectively distinguish the target signal from the reverberation component and avoid the mis-suppression of useful harmonics. Further, the residual interference energy of the initial speech signal in the time-frequency domain can be determined in combination with the time-frequency mask and the interference covariance matrix of the target sound source, which can quantify the residual reverberation and its leakage in two dimensions of statistical characteristics and time-frequency distribution. Not only can it identify the weak interference area that still overlaps with useful harmonics after reverberation suppression, but also can avoid the loss of harmonic information caused by simple amplitude reduction, thereby effectively reducing the residual acoustic crosstalk caused by reverberation overlap. Finally, the reverberation-suppressed speech signal can be subjected to harmonic enhancement using the residual interference energy, which can specifically restore the speech harmonic components weakened or covered under the action of reverberation interference, especially the mid-high frequency harmonic region on which speech intelligibility and naturalness highly depend. This process not only makes up for the possible loss of useful signal energy in the reverberation suppression stage, but also improves the integrity of the harmonic structure and the spectral detail performance, thereby effectively offsetting the damage of acoustic crosstalk to speech harmonics, and finally obtaining an enhanced speech signal. In summary, the scheme of the present application can realize the suppression of acoustic crosstalk caused by reverberation. BRIEF DESCRIPTION OF DRAWINGS

[0048] Figure 1 is an exemplary flowchart of a speech enhancement method in a reverberation environment according to some embodiments of the present application;

[0049] Figure 2 is a flowchart of determining reverberation suppression gain according to some embodiments of the present application;

[0050] Figure 3 is a flowchart of determining residual interference energy according to some embodiments of the present application;

[0051] Figure 4 is a structural diagram of a speech enhancement unit according to some embodiments of the present application;

[0052] Figure 5 is a structural diagram of a computer device for implementing a speech enhancement method in a reverberation environment according to some embodiments of the present application. DETAILED DESCRIPTION

[0053] In order to better understand the technical solutions of the present application, the technical solutions of the present application will be described in detail below in combination with the drawings of the specification and specific implementation manners.

[0054] Reference Figure 1 The figure is an exemplary flowchart of a speech enhancement method in a reverberation environment according to some embodiments of the present application, which mainly includes the following steps:

[0055] In step 101, an initial speech signal in a reverberation environment is obtained.

[0056] In specific implementation, a microphone array is used to obtain the initial speech signal in the reverberation environment.

[0057] It should be noted that the reverberation environment in the present application refers to an acoustic environment in which, when sound waves propagate in a closed or semi-closed space, in addition to direct sound, sound waves are reflected and scattered multiple times on the surfaces of walls, floors, ceilings and objects in the space to form a superimposed sound field.

[0058] In step 102, the reverberation energy distribution in each frequency band is determined according to the time-frequency spectrogram of the initial speech signal and the frequency response attenuation curve in each frequency band, and the reverberation suppression gain of the initial speech signal in the time-frequency domain is determined through all the reverberation energy distributions.

[0059] In some embodiments, the determination of the reverberation energy distribution in each frequency band according to the time-frequency spectrogram of the initial speech signal and the frequency response attenuation curve in each frequency band can be implemented by the following steps:

[0060] Time-frequency transformation is performed on the initial speech signal to obtain a time-frequency spectrogram of the initial speech signal;

[0061] The reverberation time in each frequency band is determined through the frequency response attenuation curve in each frequency band;

[0062] The reverb energy distribution in each frequency band is determined according to the time-frequency spectrogram and the reverb time in each frequency band.

[0063] It should be noted that the time-frequency spectrogram in the present application represents a two-dimensional image of the energy distribution characteristics of the initial speech signal in both time and frequency dimensions.

[0064] In a specific implementation, the time-frequency transformation of the initial speech signal to obtain the time-frequency spectrogram of the initial speech signal can be implemented in the following manner: first, the initial speech signal is segmented into overlapping short-time frames with a fixed time length (usually 20-30 milliseconds, referred to as frame length), and the overlapping rate can be set to a value between 50% and 75%; second, a Hanning window function is applied to each short-time frame signal; third, a Fourier transform is performed on each short-time frame signal after windowing to convert the short-time frame signal from the time domain to the frequency domain, thereby obtaining the frequency components and energy values of each frequency corresponding to the short-time frame signal; fourth, the frequency components and energy values of each frequency corresponding to the short-time frame signal are arranged in time sequence to form a two-dimensional image with time as the horizontal axis, frequency as the vertical axis, and pixel brightness (or color) representing energy intensity; and fifth, the obtained two-dimensional image is taken as the time-frequency spectrogram of the initial speech signal.

[0065] In some embodiments, the reverb time in each frequency band can be determined by the frequency response decay curve in each frequency band in the following steps:

[0066] for each frequency band;

[0067] determining the frequency response decay curve in the frequency band;

[0068] trend extraction is performed on the frequency response decay curve in the frequency band to filter out the effective section in which the energy monotonically decays with time in each frequency response decay curve;

[0069] the time interval experienced by the initial peak energy to drop to the energy decay reference value in the effective section is calculated, and the obtained time interval is taken as the reverb time in the frequency band.

[0070] It should be noted that the frequency response decay curve in the present application represents a characteristic curve of the energy of sound waves in a frequency band changing with time in a reverberation environment, and the frequency response decay curve presents the dynamic process of the energy of sound waves gradually decaying from direct sound propagation to multiple reflections.

[0071] In a specific implementation, for each frequency band, first, a standard excitation signal is played to the target reverberation environment, the response signal after the environmental reflection is collected by using a microphone array, the signal energy data of the target frequency band is separated from the collected response signal by using a Butterworth band-pass filter, the energy value at each time point is extracted from the signal energy data of the target frequency band, a curve is further plotted with time as the horizontal axis and energy value as the vertical axis, and the obtained curve is taken as the frequency response decay curve in the target frequency band; second, the moving average filtering algorithm is used to extract the trend of the frequency response decay curve in the frequency band, and effective segment screening is performed based on the curve after the trend extraction: the first-order difference of each point on the curve is calculated to determine the energy change trend, and when the difference result in the continuous time segment is all non-positive values (i.e., the energy does not increase with time), the time segment is determined as an effective segment in which the energy monotonically decays with time; then, the initial peak energy in the effective segment is located, that is, the highest energy value in the effective segment (the highest energy value corresponds to the starting time of the reverberation decay), the energy level that decays by 60 dB relative to the initial peak energy is taken as the energy decay reference value, and the energy decay curve along the effective segment is further tracked to find the time point at which the energy first drops to the energy decay reference value; finally, the difference value between the time point corresponding to the initial peak energy and the time point corresponding to the energy decay reference value is calculated, and the obtained difference value is taken as the reverberation time in the frequency band.

[0072] It should be noted that the reverberation time in the present application represents the duration of the reverberation component in the frequency band.

[0073] In a specific implementation, the determination of the reverberation energy distribution in each frequency band based on the time-frequency spectrogram and the reverberation time in each frequency band can be performed in the following manner: first, the energy values of each frequency band at different time points are obtained based on the time-frequency spectrogram (i.e., the energy sum in the vertical axis interval of the time-frequency spectrogram corresponding to each frequency band changes with time); second, for each frequency band, the energy decay mode is determined in combination with the reverberation time of the frequency band: if the energy decay speed at a certain time point matches the reverberation time (i.e., the decay rate is slower than the natural decay of the initial speech signal, which is consistent with the "tail" characteristics of reverberation), the part of the energy value is determined as the reverberation component, the total sum of all energy values determined as the reverberation component in the frequency band is calculated, and the proportion of the total energy value in the frequency band is calculated, and the curve of the proportion value changing with time is further taken as the reverberation energy distribution in the frequency band, thereby obtaining the reverberation energy distribution in each frequency band.

[0074] It should be noted that the reverberation energy distribution in the present application represents the distribution characteristics of the energy of the reverberation component in the time dimension in the frequency band.

[0075] In some embodiments, reference is made to Figure 2As shown in the figure, the figure is a flowchart for determining the reverb suppression gain in some embodiments of the present application. The reverb suppression gain of the initial speech signal in the time-frequency domain can be determined by the following steps in the present embodiment:

[0076] The reverb energy distribution of each frequency band is merged to obtain a full frequency domain reverb energy spectrum;

[0077] The reverb energy spectrum and the power spectrum of the initial speech signal are compared and analyzed to obtain a reverb dominant area and a speech dominant area;

[0078] The reverb suppression factor of each time-frequency unit of the initial speech signal in the time-frequency domain is determined according to the reverb dominant area and the speech dominant area;

[0079] The reverb suppression factors of all time-frequency units are calculated to obtain the reverb suppression gain of the initial speech signal in the time-frequency domain.

[0080] It should be noted that the frequency domain reverb energy spectrum in the present application represents the overall distribution characteristics of all reverb energies of the initial speech signal in the time-frequency domain.

[0081] In a specific implementation, the reverb energy distribution of each frequency band is merged to obtain a full frequency domain reverb energy spectrum, which can be implemented in the following way, that is, first, the reverb energy distribution of each frequency band is sorted according to its corresponding frequency range on the vertical axis (for example, arranged from low frequency to high frequency), second, the time axis of all frequency bands is calibrated, third, the linear interpolation method is used to fill the frequency gaps between adjacent frequency bands, then, the reverb energy proportion data of each frequency band at the same time point is integrated according to the frequency order to form a two-dimensional matrix with time as the horizontal axis, frequency as the vertical axis, and data point value representing the reverb energy proportion of the corresponding time-frequency unit, finally, the two-dimensional matrix is visualized (for example, the color depth represents the proportion), and the spectrum obtained by visualizing is used as the full frequency domain reverb energy spectrum.

[0082] It should be noted that the speech dominant area in the present application represents a set of time-frequency units in which the speech signal energy dominates the total energy (the sum of speech energy and reverb energy) in the time-frequency domain, the reverb dominant area represents a set of time-frequency units in which the reverb energy dominates the total energy in the time-frequency domain, and the time-frequency unit represents the basic unit for dividing the signal in time-frequency analysis, which is a rectangular area defined by the time interval and the frequency bandwidth in the time-frequency two-dimensional plane.

[0083] In a specific implementation, the reverberation energy map and the power spectrum of the initial speech signal are compared and analyzed to obtain the reverberation dominant region and the speech dominant region, which can be achieved in the following manner: first, the reverberation energy map and the power spectrum of the initial speech signal are mapped to the same time-frequency grid (the time and frequency resolutions remain consistent), so that each time-frequency unit can correspond to compare the reverberation energy value of the reverberation energy map and the total energy value of the power spectrum; second, the proportion of the reverberation energy in each time-frequency unit to the total energy is calculated, and a threshold judgment method commonly used in the prior art is adopted (set the proportion of the reverberation to be greater than 50% as the limit), when the proportion of the reverberation energy in a certain time-frequency unit exceeds the threshold, it is determined that the time-frequency unit is a reverberation dominant region, otherwise, when the proportion of the reverberation energy is lower than the threshold, it is determined that it is a speech dominant region.

[0084] It should be noted that the reverberation suppression factor in the present application represents a parameter for quantifying the strength of the reverberation attenuation in the time-frequency unit in the time-frequency domain.

[0085] In a specific implementation, the reverberation suppression factor of each time-frequency unit of the initial speech signal in the time-frequency domain can be determined according to the reverberation dominant region and the speech dominant region, which can be achieved in the following manner: first, for the time-frequency unit of the speech dominant region, a reverberation suppression factor close to 1 (for example, 0.8-1.0) is used, second, for the time-frequency unit of the reverberation dominant region, a smaller reverberation suppression factor (for example, 0.2-0.5) is used, and finally, for the transition zone (time-frequency unit with energy proportion close to the threshold) between the two types of regions, a Gaussian smoothing algorithm is used to generate a reverberation suppression factor between the above two.

[0086] It should be noted that the reverberation suppression gain in the present application represents a parameter for quantifying the degree of attenuation of the reverberation energy of the initial speech signal in each time-frequency unit in the time-frequency domain.

[0087] In a specific implementation, the gain of the reverberation suppression factor of all time-frequency units is calculated to obtain the reverberation suppression gain of the initial speech signal in the time-frequency domain, which can be achieved in the following manner: first, the reverberation suppression factor of all time-frequency units is directly used as linear gain, that is, the value of the reverberation suppression factor is the energy proportion that needs to be preserved in the time-frequency unit (for example, the reverberation suppression factor is 0.2, which means only 20% of the energy is preserved, and the reverberation suppression factor is 0.9, which means 90% of the energy is preserved), second, the reverberation suppression factor of each time-frequency unit is multiplied with the time-frequency amplitude of the initial speech signal in the corresponding time-frequency unit, and all the multiplied values are used as the suppression gain value of each time-frequency unit, thereby completing the gain application, and finally, all time-frequency units after the multiplication form an overall gain distribution, and the obtained overall gain distribution is used as the reverberation suppression gain of the initial speech signal in the time-frequency domain.

[0088] In step 103, the initial speech signal is reverb suppressed according to the reverb suppression gain to obtain a reverb-suppressed speech signal.

[0089] In some embodiments, the initial speech signal can be reverb suppressed according to the reverb suppression gain to obtain a reverb-suppressed speech signal by the following steps:

[0090] The reverb threshold value in each frequency band is determined according to the reverb energy distribution in the frequency band.

[0091] The time-frequency coefficient of each time-frequency unit of the initial speech signal in the time-frequency domain is suppressed by the reverb suppression gain and the reverb threshold value in each frequency band, and a reverb-suppressed speech signal is obtained.

[0092] It should be noted that the reverb threshold value in the present application represents the critical energy value that needs to be focused on for suppressing the reverb energy of the time-frequency unit in the frequency band.

[0093] In specific implementation, the reverb threshold value in each frequency band can be determined according to the reverb energy distribution in the frequency band by the following method, that is, for each frequency band, the reverb energy values of multiple typical time points (such as energy of different stages such as speech gap and speech tail) are uniformly sampled from the reverb energy distribution in the frequency band, then the median of these sampling values is calculated, and the median is multiplied by an empirical adjustment coefficient (such as 1.2 times, which can be flexibly set according to the degree of interference of reverb on speech in the actual scene), and the multiplied value is taken as the reverb threshold value in the frequency band, thereby obtaining the reverb threshold value in each frequency band.

[0094] In specific implementation, the time-frequency coefficient of each time-frequency unit of the initial speech signal in the time-frequency domain can be suppressed by the reverb suppression gain and the reverb threshold value in each frequency band, and a reverb-suppressed speech signal can be obtained by the following method, that is, first, the time-frequency coefficient of each time-frequency unit of the initial speech signal in the time-frequency domain is extracted by short-time Fourier transform, second, for each time-frequency unit, the suppression gain value of the time-frequency unit is obtained from the reverb suppression gain, and the suppression gain value of the time-frequency unit is multiplied by the time-frequency coefficient on the time-frequency unit, thereby completing the basic attenuation, and for the time-frequency coefficient after the basic attenuation, it is judged whether the reverb energy of the frequency band to which the time-frequency coefficient belongs still exceeds the reverb threshold value in the frequency band, if it exceeds, it means that the reverb suppression is insufficient, and the amplitude of the time-frequency coefficient is multiplied by a reinforcement attenuation coefficient (such as 0.3-0.5), if it does not exceed, the basic attenuation result is maintained, finally, the amplitudes of all time-frequency units after processing are recombined with the original phase, and the time-domain waveform is converted back by inverse short-time Fourier transform, thereby obtaining a reverb-suppressed speech signal.

[0095] In step 104, time-frequency mask estimation is performed on the initial speech signal to obtain a time-frequency mask of the initial speech signal, and residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain is determined according to the time-frequency mask and an interference covariance matrix of a target sound source.

[0096] In some embodiments, time-frequency mask estimation performed on the initial speech signal to obtain a time-frequency mask of the initial speech signal can be implemented by the following steps:

[0097] Detecting, by a pre-trained voice activity detection model, a speech presence probability of each time-frequency unit of the initial speech signal in the time-frequency domain;

[0098] Determining the time-frequency mask of the initial speech signal according to all speech presence probabilities.

[0099] In a specific implementation, detecting, by a pre-trained voice activity detection model, a speech presence probability of each time-frequency unit of the initial speech signal in the time-frequency domain can be implemented in the following manner, that is:

[0100] It should be noted that the pre-trained voice activity detection model in the present application is a machine learning model trained by large-scale labeled data (containing speech segments, silence segments, noise segments, and diversified scene samples), and its core function is to determine whether there is speech in the input signal and the start and end time of the speech. In the prior art, a deep learning architecture combining a convolutional neural network and a long short-term memory network is usually used, wherein the convolutional neural network layer can extract local time-frequency features of the signal (such as a speech energy pattern of a specific frequency combination), and the long short-term memory network layer can capture the context association in the time dimension (for example, the continuity before and after the speech). The voice activity detection model usually inputs features such as the Mel spectrum and the short-time energy obtained by preprocessing the speech signal, and outputs a speech presence probability (0-1 range) of each time unit (or time-frequency unit). After training, the voice activity detection model can be generalized to different noise and reverberation environments, so as to accurately distinguish speech and non-speech components.

[0101] It should also be noted that the speech presence probability in the present application represents a numerical index quantifying the possibility of containing valid speech components in each time-frequency unit.

[0102] In specific implementation, the probability of speech presence in each time-frequency unit of the initial speech signal in the time-frequency domain can be detected by a pre-trained speech activity detection model in the following way: First, the initial speech signal is divided into frames and windowed, and the time spectrum (containing amplitude information of each time frame and frequency point) is obtained by short-time Fourier transform. Then, it is converted into the input format of the speech activity detection model (for example, extracting Mel spectrum features, mapping the frequency axis to a Mel scale that conforms to human auditory perception, and taking 40-80 dimensional features). Second, the processed features are input into the pre-trained speech activity detection model. The speech activity detection model outputs a value in the range of 0-1 for each time-frequency unit, and the obtained value is used as the probability of speech presence in each time-frequency unit.

[0103] It should be noted that the time-frequency mask described in this application represents a weight matrix that marks the proportion of speech components in the time-frequency domain of the initial speech signal.

[0104] In specific implementation, the time-frequency mask of the initial speech signal can be determined based on the probability of all speech occurrences in the following way: First, divide the time-frequency units into intervals (e.g., 0-0.3, 0.3-0.7, 0.7-1.0) according to the distribution characteristics of the speech occurrence probabilities. Each interval corresponds to a different mask value level. Then, assign a lower mask value (e.g., 0.1-0.2) to the time-frequency units in the 0-0.3 interval, a medium mask value (e.g., 0.4-0.6) to the time-frequency units in the 0.3-0.7 interval, and a higher mask value (e.g., 0.8-1.0) to the time-frequency units in the 0.7-1.0 interval. The final generated time-frequency mask is completely aligned with the time-frequency structure of the signal. Combine the mask values ​​of all time-frequency units according to their corresponding time-frequency positions to form a matrix, and use the obtained matrix as the time-frequency mask of the initial speech signal.

[0105] In some embodiments, reference Figure 3 As shown in the figure, this is a flowchart illustrating the determination of residual interference energy in some embodiments of this application. In this embodiment, the determination of the residual interference energy of the initial speech signal in each time-frequency unit in the time-frequency domain based on the time-frequency mask and the interference covariance matrix of the target sound source can be achieved by the following steps:

[0106] First, in step 1041, the interference covariance matrix of the target sound source is determined;

[0107] Secondly, in step 1042, the effective region of the initial speech signal in the time-frequency domain is identified according to the time-frequency mask;

[0108] Then, in step 1043, residual interference analysis is performed on the effective region of the initial speech signal in the time-frequency domain using the interference covariance matrix of the target sound source, so as to obtain the residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain.

[0109] It should be noted that the interference covariance matrix in the present application represents the statistical correlation characteristics between the frequency components of the interference signal in the time-frequency domain in the environment where the target sound source is located.

[0110] In a specific implementation, the interference covariance matrix of the target sound source can be determined in the following manner: first, input the initial speech signal into a pre-trained speech activity detection model, analyze the time-frequency characteristics of the initial speech signal through the speech activity detection model, and output the speech activity state of each time frame, i.e., “voiced speech” (containing target speech energy) or “unvoiced speech” (without target speech energy). Based on the output result, mark all the time periods of “unvoiced speech”, and regard the signals corresponding to all the time periods of “unvoiced speech” as non-occurred speech signals (such signals are generated during the user's silent period, mainly containing reverberation residues, environmental noise, and other pure interference components without effective speech energy mixed in). Extract the complex time-frequency coefficients of each time frame and each frequency point from all the non-occurred speech signals, integrate all the complex time-frequency coefficients into a two-dimensional data matrix according to the “time-frequency” dimension, and regard the matrix as an interference sample. Then, input the interference sample, calculate the linear correlation degree between different time-frequency units by using the sample covariance formula, and regard the generated matrix as the interference covariance matrix of the target sound source.

[0111] In a pure interference scenario without target speech (for example, only environmental noise or other sound sources exist during a certain period of time), the collected interference signal is framed, windowed, and subjected to short-time Fourier transform to obtain complex time-frequency coefficients of the interference signal in the time-frequency domain (arranged as a matrix according to the frequency dimension, with each row corresponding to a frequency point and each column corresponding to a time frame). Then, based on the calculation method of the covariance matrix in the prior art, the time-frequency coefficient matrix is statistically averaged according to the time dimension. Specifically, the cross-correlation values of the time-frequency coefficients of the interference signals at different frequency points are calculated to form a square matrix with the same number of dimensions as the frequency points (the diagonal elements are the energies of the interference signals at the frequency points, and the non-diagonal elements are the correlations of the interference signals at different frequency points). The obtained square matrix is regarded as the interference covariance matrix of the target sound source.

[0112] It should be noted that the effective area in the present application represents a set of time-frequency units in the time-frequency domain that are determined to be dominated by target speech components.

[0113] In a specific implementation, the identification of the valid area of the initial speech signal in the time-frequency domain according to the time-frequency mask can be implemented in the following manner: first, based on the statistical distribution characteristics (for example, the probability density function of the mask value) of the time-frequency mask, a plurality of determination thresholds (for example, a low threshold of 0.3 and a high threshold of 0.7) are determined; then, all time-frequency units are traversed, and the time-frequency units with a mask value greater than or equal to 0.7 are marked as “strong valid units”, the time-frequency units with a mask value greater than 0.3 and less than 0.7 are marked as “weak valid units”, and the weak valid units with a mask value less than 0.3 are marked as “invalid units”; subsequently, the “strong valid unit” area is expanded by using the dilation operation in the morphological filtering, so as to include the adjacent “weak valid units”; finally, the continuous time-frequency area set obtained after the expansion is taken as the valid area of the initial speech signal in the time-frequency domain.

[0114] It should be noted that the residual interference energy in the present application represents the interference energy caused by the leakage of reverberation sound energy to the time-frequency units in the time-frequency domain of the initial speech signal to the frequency components of the initial speech signal.

[0115] In a specific implementation, the residual interference analysis of the valid area of the initial speech signal in the time-frequency domain by using the interference covariance matrix of the target sound source can be implemented in the following manner: first, the complex time-frequency coefficients of all time-frequency units in the valid area are extracted; then, the interference energy projection calculation is performed on the complex time-frequency coefficients of each time-frequency unit in the valid area by using the interference covariance matrix of the target sound source, specifically, the quadratic form operation (that is, the conjugate transpose of the complex time-frequency coefficient multiplied by the interference covariance matrix and then multiplied by the complex time-frequency coefficient) is performed on the complex time-frequency coefficient and the interference covariance matrix, and the component obtained after the quadratic form operation is taken as the energy component originating from the interference in the time-frequency unit, and the obtained energy component is taken as the residual interference energy of the time-frequency unit in the time-frequency domain of the initial speech signal, so as to obtain the residual interference energy of each time-frequency unit in the time-frequency domain of the initial speech signal.

[0116] In step 105, the speech signal after reverberation suppression is subjected to harmonic enhancement by all the residual interference energies, and a speech enhanced signal is obtained.

[0117] In some embodiments, the speech signal after reverberation suppression can be subjected to harmonic enhancement by all the residual interference energies in the following steps:

[0118] The interference area corresponding to the harmonic structure of the speech signal after reverberation suppression in the time-frequency domain is identified by all the residual interference energies;

[0119] The amplitude of the interference area corresponding to the harmonic structure in the time-frequency domain is enhanced, and a speech enhanced signal is obtained.

[0120] It should be noted that the interference region in the present application represents the region where the acoustic crosstalk in the reverberation environment overlaps with the harmonic structure of the speech signal in the time-frequency domain after the reverberation suppression.

[0121] In a specific implementation, the interference region corresponding to the harmonic structure of the speech signal in the time-frequency domain after the reverberation suppression can be identified by all the residual interference energy in the following manner: first, frame the speech signal after the reverberation suppression, and then convert each signal frame into a time-frequency matrix in the time-frequency domain by short-time Fourier transform (each element in the time-frequency matrix corresponds to the signal energy of a specific time frame and frequency point, and completely corresponds to the time-frequency unit of the initial speech signal in the time and frequency coordinates); second, match the residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain to the time-frequency matrix according to the same coordinates to form the interference energy data of each time-frequency unit; third, generate an interference energy distribution heat map by mapping the interference energy value to a color gradient from low to high; fourth, use the inherent characteristics of the speech signal after the reverberation suppression, i.e., the speech is generated by the periodic vibration of the vocal cords, and the frequency components contain integer multiples of the fundamental frequency (i.e., harmonics), which appear as continuous energy bands parallel to the time axis in the time-frequency domain (for example, when the fundamental frequency is 150 Hz, 300 Hz, 450 Hz, etc. will form corresponding harmonic bands); fifth, extract the fundamental frequency from the initial speech signal by the autocorrelation method, and then determine the theoretical frequency range of each order of harmonic bands; and finally, superimpose the interference energy heat map and the theoretical frequency range of each order of harmonic bands, filter out the time-frequency units that are both within the harmonic band range and satisfy the condition that the residual interference energy exceeds a preset threshold (the threshold can be set as the average of all residual interference energy), and form a set of time-frequency units, which are the interference region corresponding to the harmonic structure of the speech signal in the time-frequency domain after the reverberation suppression.

[0122] In a specific implementation, the amplitude of the interference region corresponding to the harmonic structure in the time-frequency domain is enhanced, and then the speech enhancement signal is obtained by using the following method, that is, first, the reference features of the harmonic structure are extracted from the initial speech signal, the time-frequency matrix of the non-interference region is obtained by short-time Fourier transform, the amplitude distribution law of each order harmonic (fundamental frequency and integer multiple frequency) in the normal state is counted, including the amplitude change gradient (for example, the amplitude difference of adjacent frames is not more than 15%) of the same harmonic in the continuous time frame and the energy ratio relationship between different harmonics; second, for the identified harmonic structure interference region, the corresponding reference features are matched according to the harmonic order to which the interference region belongs: for the interference region of the continuous time frame, the sliding window average method is used, the same order harmonic amplitude mean of the non-interference region in front and back 3-5 frames is taken as the reference, and the amplitude in the interference region is adjusted to 90%-95% of the reference by linear interpolation, for the isolated interference unit, the same order harmonic amplitude average of the adjacent 2-3 non-interference frequency points in the same time frame is referred to, and the amplitude in the interference unit is increased to more than 85% of the same order harmonic amplitude average; finally, the enhanced interference region time-frequency data and the time-frequency data of the non-interference region are fused, that is, the time-frequency units at the boundary of the interference region (that is, 1-2 columns of time frames or 1-2 rows of frequency points adjacent to the non-interference region) are processed by weighted smoothing, and the weight is distributed according to the distance (for example, the time-frequency unit closer to the interference region is given higher weight of the enhanced data, and the time-frequency unit closer to the non-interference region is given higher weight of the original non-interference data), for the non-boundary region, the enhanced interference region data and the original non-interference region data are directly retained, then the inverse short-time Fourier transform is performed on the fused complete time-frequency matrix, and the time domain signal obtained by the inverse short-time Fourier transform is taken as the speech enhancement signal.

[0123] In addition, another aspect of the present application provides a speech control system in some embodiments, which comprises a speech enhancement unit, wherein the speech enhancement unit is configured to Figure 4 The figure is a structural schematic diagram of a speech enhancement unit according to some embodiments of the present application, which comprises an acquisition module 401, a processing module 402 and an execution module 403, which are described as follows:

[0124] The acquisition module 401 is mainly used for acquiring the initial speech signal in the reverberation environment in the present application;

[0125] The processing module 402 is used for determining the reverberation energy distribution in each frequency band according to the time-frequency spectrum of the initial speech signal and the frequency response attenuation curve in each frequency band, and determining the reverberation suppression gain of the initial speech signal in the time-frequency domain through all the reverberation energy distributions in the present application;

[0126] It should be noted that the processing module 402 in this application is also used to perform reverberation suppression on the initial speech signal according to the reverberation suppression gain to obtain a reverberation-suppressed speech signal;

[0127] It should be noted that the processing module 402 described in this application is also used to perform time-frequency mask estimation on the initial speech signal to obtain the time-frequency mask of the initial speech signal, and to determine the residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain based on the time-frequency mask and the interference covariance matrix of the target sound source.

[0128] The execution module 403 in this application is mainly used to perform harmonic enhancement on the reverberation-suppressed speech signal through all residual interference energy to obtain a speech enhancement signal.

[0129] In addition, this application also provides a computer device, the computer device including a memory and a processor, the memory storing code, and the processor being configured to acquire the code and execute the above-described speech enhancement method in a reverberation environment.

[0130] In some embodiments, reference Figure 5 The figure is a schematic diagram of the structure of a computer device implementing a speech enhancement method in a reverberant environment according to some embodiments of this application. The speech enhancement method in a reverberant environment in the above embodiments can be implemented through... Figure 5 The computer device shown is used to implement this, and the computer device 500 includes at least one processor 501, a communication bus 502, a memory 503, and at least one communication interface 504.

[0131] The processor 501 may be a general-purpose central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more devices used to control the execution of the speech enhancement method in the reverberation environment of this application.

[0132] The communication bus 502 can be used to transmit information between the aforementioned components.

[0133] The memory 503 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM), or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magnetic disk storage or other magnetic storage devices, or any other medium capable of storing desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited to this. The memory 503 can exist independently, and is connected to the processor 501 through the communication bus 502. The memory 503 can also be integrated with the processor 501.

[0134] The memory 503 is configured to store program codes for implementing the solutions of the present application, and the processor 501 is configured to control the execution of the program codes. The processor 501 is configured to execute the program codes stored in the memory 503. The program codes can include one or more software modules. The methods described in the above method embodiments can be implemented by the processor 501 and one or more software modules in the program codes in the memory 503.

[0135] The communication interface 504 is configured to communicate with other devices or communication networks, such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc., using any transceiver-like device.

[0136] In specific implementations, as an embodiment, the computer device can include multiple processors, each of which can be a single-CPU processor or a multi-CPU processor. The processor herein can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0137] The computer device described above can be a general-purpose computer device or a special-purpose computer device. In a specific implementation, the computer device can be a desktop computer, a laptop computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. The embodiments of the present application do not limit the type of the computer device.

[0138] In addition, the present application also provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the voice enhancement method in a reverberation environment.

[0139] Although the preferred embodiments of the present application have been described, those skilled in the art who are informed of the basic inventive concept can make further changes and modifications to the embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0140] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.

Claims

1. A speech enhancement method in a reverberant environment, used for speech enhancement in a speech control system, characterized in that, The method includes the following steps: Acquire the initial speech signal under reverberation conditions; The reverberation energy distribution in each frequency band is determined based on the time-frequency spectrum of the initial speech signal and the frequency response attenuation curves in each frequency band. The reverberation suppression gain of the initial speech signal in the time-frequency domain is determined by all the reverberation energy distributions. The initial speech signal is subjected to reverberation suppression based on the reverberation suppression gain to obtain a reverberation-suppressed speech signal. The initial speech signal is subjected to time-frequency mask estimation to obtain the time-frequency mask of the initial speech signal. The residual interference energy of the initial speech signal in each time-frequency unit in the time-frequency domain is determined based on the time-frequency mask and the interference covariance matrix of the target sound source. The speech signal is enhanced by harmonic enhancement of the reverberation-suppressed speech signal using all residual interference energy.

2. The method of claim 1, wherein, Determining the reverberation energy distribution in each frequency band based on the time-spectrum diagram of the initial speech signal and the frequency response attenuation curves in each frequency band specifically includes: The initial speech signal is subjected to time-frequency transformation to obtain the time-frequency spectrum of the initial speech signal; The reverberation time in each frequency band is determined by the frequency response decay curves in each frequency band. The reverberation energy distribution in each frequency band is determined based on the time-frequency spectrum and the reverberation time in each frequency band.

3. The method of claim 2, wherein, Determining the reverberation time in each frequency band using the frequency response decay curves specifically includes: For each frequency band; Determine the frequency response attenuation curve within the frequency band; Trend extraction is performed on the frequency response attenuation curves within the frequency band, and the effective segments in each frequency response attenuation curve in which energy decreases monotonically with time are selected. Calculate the time interval during which the energy drops from the initial peak value to the energy attenuation reference value within the effective band, and use the obtained time interval as the reverberation time within the frequency band.

4. The method of claim 1, wherein, Determining the reverberation suppression gain of the initial speech signal in the time-frequency domain by analyzing all reverberation energy distributions specifically includes: The reverberation energy distributions in each frequency band are merged to obtain a full-frequency domain reverberation energy spectrum. By comparing and analyzing the reverberation energy spectrum with the power spectrum of the initial speech signal, the reverberation-dominant region and the speech-dominant region are obtained. The reverberation suppression factor of the initial speech signal in each time-frequency unit in the time-frequency domain is determined based on the reverberation dominance region and the speech dominance region. The reverberation suppression factor of all time-frequency units is calculated to obtain the reverberation suppression gain of the initial speech signal in the time-frequency domain.

5. The method of claim 1, wherein, The initial speech signal is subjected to reverberation suppression based on the reverberation suppression gain to obtain a reverberation-suppressed speech signal, specifically including: The reverberation threshold value in each frequency band is determined based on the reverberation energy distribution in each frequency band. The reverberation suppression gain and the reverberation threshold value in each frequency band are used to suppress the time-frequency coefficients of the initial speech signal in each time-frequency unit in the time-frequency domain, thereby obtaining the reverberation-suppressed speech signal.

6. The method of claim 1, wherein, Performing time-frequency mask estimation on the initial speech signal to obtain the time-frequency mask specifically includes: The probability of speech presence in each time-frequency unit of the initial speech signal in the time-frequency domain is detected by a pre-trained speech activity detection model. Determine a time-frequency mask of the initial speech signal according to all speech existence probabilities.

7. The method of claim 1, wherein, Determine residual interference energy of each time-frequency unit of the initial speech signal in a time-frequency domain according to the time-frequency mask and an interference covariance matrix of a target sound source specifically includes: Determine an interference covariance matrix of a target sound source; Identify a valid area of the initial speech signal in a time-frequency domain according to the time-frequency mask; Perform interference leakage estimation on the valid area of the initial speech signal in the time-frequency domain through the interference covariance matrix of the target sound source to obtain the residual interference energy of each time-frequency unit of the initial speech signal in the time-frequency domain.

8. The method of claim 1, wherein, Perform harmonic enhancement on the speech signal after reverberation suppression through all residual interference energies to obtain a speech enhanced signal specifically includes: Identify an interference area corresponding to a harmonic structure of the speech signal after reverberation suppression in a time-frequency domain through all residual interference energies; Perform amplitude enhancement on the interference area corresponding to the harmonic structure in the time-frequency domain to obtain the speech enhanced signal.

9. The method of claim 1, wherein, The reverberation environment represents an acoustic environment in which, in addition to direct sound, sound waves propagate in a closed or semi-closed space, multiple reflections and scattering occur on the surfaces of walls, floors, ceilings and objects in the space to form a superimposed sound field.

10. A voice control system comprising a voice enhancement unit, characterized in that The speech enhancement unit includes: An acquisition module configured to acquire an initial speech signal in a reverberation environment; A processing module configured to determine reverberation energy distribution in each frequency band according to a time-frequency spectrogram of the initial speech signal and a frequency response attenuation curve in each frequency band, and determine reverberation suppression gain of the initial speech signal in a time-frequency domain through all reverberation energy distributions; The processing module is further configured to perform reverberation suppression on the initial speech signal according to the reverberation suppression gain to obtain a speech signal after reverberation suppression; The processing module is further configured to perform time-frequency mask estimation on the initial speech signal to obtain a time-frequency mask of the initial speech signal, and determine residual interference energy of each time-frequency unit of the initial speech signal in a time-frequency domain according to the time-frequency mask and an interference covariance matrix of a target sound source; An execution module configured to perform harmonic enhancement on the speech signal after reverberation suppression through all residual interference energies to obtain a speech enhanced signal.

Citation Information

Patent Citations

  • Speech enhancement method applied to dual-microphone array

    CN105788607A

  • Speech enhancement method and system, computer equipment and storage medium

    CN110503972A