A target speech extraction method based on video information assistance
By combining video information-assisted methods, the spatial orientation and direct reverb ratio of the potential speaker are estimated, and a hybrid minimum mean square distortion-free response beamformer is constructed, which solves the problems of poor robustness and limited interference suppression capabilities in adaptive beamforming in traditional microphone array systems, and achieves efficient extraction and interference suppression of target speech.
Patent Information
- Application Number
- CN202211147202.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-20
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-09-20
AI Technical Summary
Traditional microphone array systems have problems such as poor robustness, large computing volume and limited interference suppression in adaptive beam formation, especially in indoor environments of multiple speakers, which are difficult to effectively extract target voice.
Combined with video information assistance, by estimating the spatial orientation and direct reverb ratio of the potential speaker, a hybrid minimum mean square distortion-free response beamformer is constructed, and the beamforming preprocessing and covariance matrix estimation is used to improve the accuracy of target speech extraction and interference suppression effect.
While extracting the target voice, it significantly reduces interference and spatial reverb in other directions, improves the robustness and computing efficiency of adaptive beam formation, and is suitable for communication systems with cameras such as smart large screens and smart glasses.
Smart Images

Figure CN115472151B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of speech data processing, and in particular to a target speech extraction method based on video information assistance. Background Art
[0002] Currently, commonly used microphone array systems use microphone signals to perform sound source localization and beamforming to extract target speech signals. Depending on whether the beamformer coefficients depend on the data itself, commonly used microphone array beamformers can be divided into fixed beamformers and adaptive beamformers. Fixed beamformers offer low computational complexity and high robustness, but their interference suppression capabilities are limited. Adaptive beamformers can create strong nulls at interfering sound sources, effectively suppressing interference, but they suffer from high computational complexity and poor robustness.
[0003] To improve the robustness of adaptive beamforming in practical applications, common improvements include diagonal loading and noise covariance matrix reconstruction. In practice, empirical values are often used to determine the diagonal loading coefficients. However, excessive diagonal loading can affect interference suppression, while insufficient loading can lead to self-cancellation. Reconstructing the noise covariance matrix can also be used within a microphone array to reduce the target signal component in the covariance matrix. However, this method requires scanning the entire space and calculating the power output in each direction, which is computationally intensive and places very high demands on hardware for practical applications.
[0004] Therefore, traditional adaptive beamforming methods have certain limitations in practical applications. Summary of the Invention
[0005] With the prevalence of camera-enabled communication systems such as smart screens and smart glasses, more and more audio and video fusion methods are being developed to improve beamforming performance. This application proposes a target speech extraction method based on video information, which can effectively extract the target speech while suppressing other interference and reverberation.
[0006] This application provides a method for extracting target speech based on video information assistance, including:
[0007] A microphone array system with a camera is used to acquire video information and voice signals; the voice signals are acquired by microphone elements in the microphone array system; the video information is acquired by a camera in the microphone array system; the video information includes spatial orientation information of a potential speaker;
[0008] Using two microphone array elements at different spacings to form multiple array element pairs, the speech signal is divided into multiple frequency bands with continuous frequencies in the frequency domain; combining the spatial orientation information of the potential speaker, the direct-to-reverberation ratio of the target speech in the received signal at each frequency point within the multiple frequency bands is estimated based on the microphone array element pairs at different spacings;
[0009] Performing beamforming preprocessing on the frequency domain signal of the speech signal in combination with the spatial orientation information of the potential speaker; determining the presence probability of the target speech in the received signal at each frequency point using the energy ratio of the frequency domain signal before and after the beamforming preprocessing and the direct-to-reverberation ratio of the target speech in the received signal at each frequency point;
[0010] In combination with the spatial orientation information of the potential speaker, a hybrid minimum mean square distortionless response beamformer is constructed based on at least the direct-to-reverberation ratio of the target speech in the received signals at each frequency point and the existence probability of the target speech in the received signals at each frequency point.
[0011] In a possible implementation, the target speech extraction method further includes filtering the frequency domain signal of the speech signal using the hybrid minimum mean square distortionless response beamformer to obtain the frequency domain signal of the target speech.
[0012] In a possible implementation, the target speech extraction method further includes performing an inverse short-time Fourier transform on the frequency domain signal of the target speech to obtain a time domain signal; and performing windowing and smoothing on the time domain signal to obtain the time domain signal of the target speech.
[0013] In one possible implementation, before determining the probability of the presence of the target speech in the received signals at each frequency point using the energy ratio of the frequency domain signals before and after the beamforming preprocessing, the energy ratio needs to be corrected to obtain a new energy ratio; the correction is used to offset the attenuation error caused by the increase in the frequency of the target speech energy.
[0014] In a possible implementation, the target speech extraction method further includes:
[0015] The video information also includes spatial orientation information of an interfering sound source; a data-independent covariance matrix is constructed based on a steering vector of the interfering sound source; the steering vector of the interfering sound source is determined based on the spatial orientation information of the interfering sound source and the spacing between microphone array elements in the microphone array system;
[0016] Performing diagonally loaded super-directional fixed beam preprocessing on the frequency domain signal of the speech signal according to the spatial orientation information of the interfering sound source to obtain a beam output signal in the direction of the interfering sound source; the direction of the interfering sound source is determined according to the spatial orientation information of the potential interfering sound source:
[0017] According to the beam output signal in the direction of the interference sound source and the data-independent covariance matrix, a new noise covariance matrix is constructed:
[0018] A data-dependent covariance matrix is constructed according to the frequency domain signal of the speech signal; the data-dependent covariance matrix does not consider the spatial orientation information of the potential speaker and the spatial orientation information of the interfering sound source.
[0019] In one possible implementation, a hybrid minimum mean square distortionless response beamformer is constructed based on at least the direct-to-reverberation ratio of the target speech in the received signals at each frequency point and the existence probability of the target speech in the received signals at each frequency point, including:
[0020] Calculating the mixing coefficients and diagonal loading coefficients of the mixing covariance matrix according to the direct-to-reverberation ratio of the target speech in the received signals at each frequency point and the existence probability of the target speech in the received signals at each frequency point;
[0021] Constructing a mixed noise covariance matrix according to the data-dependent covariance matrix, the new noise covariance matrix, and the mixing coefficient of the mixed covariance matrix according to a weighted algorithm;
[0022] constructing a diagonally loaded noise covariance matrix according to the diagonal loading coefficients and the hybrid noise covariance matrix;
[0023] The hybrid minimum mean square distortionless response beamformer is constructed according to the diagonally loaded noise covariance matrix in combination with the spatial orientation information of the potential speaker.
[0024] In one possible implementation, the mixing coefficients of the mixing covariance matrix are calculated, including:
[0025] Determining a calculation form for a mixing coefficient of a mixing covariance matrix based on a comparison relationship between a first parameter and a second parameter dividing the reverberation intensity and the direct-reverberation ratio of the target speech in the received signals at each frequency point; determining the mixing coefficient of the mixing covariance matrix according to the calculation form using the direct-reverberation ratio of the target speech in the received signals at each frequency point and the probability of the target speech in the received signals at each frequency point;
[0026] The first parameter and the second parameter for dividing the reverberation intensity are determined based on empirical values.
[0027] In a possible implementation, calculating the mixing coefficients of the mixing covariance matrix further includes performing inter-frame smoothing and inter-frequency smoothing on the mixing coefficients of the mixing covariance matrix to obtain smoothed mixing coefficients of the mixing covariance matrix.
[0028] In one possible implementation, calculating the diagonal loading coefficient includes,
[0029] Determining a calculation form for a diagonal loading coefficient based on a comparison relationship between a first parameter and a second parameter dividing the reverberation intensity and a direct-reverberation ratio of the target speech in the received signals at each frequency point; and determining the diagonal loading coefficient according to the calculation form using a probability of the target speech in the received signals at each frequency point.
[0030] The first parameter and the second parameter for dividing the reverberation intensity are determined based on empirical values.
[0031] In one possible implementation, a mixed noise covariance matrix is constructed, including:
[0032] Processing the data-dependent covariance matrix and the noise covariance matrix to obtain normalized data-dependent covariance matrix and noise covariance matrix;
[0033] A weighted algorithm is executed on the normalized data-related covariance matrix and the noise covariance matrix according to the mixing coefficient of the mixed covariance matrix to construct a mixed noise covariance matrix.
[0034] This application uses video information to provide the spatial orientation of the target sound source, adopts array signal processing technology to estimate the direct-to-reverberation ratio of the target speech, estimates the spatial acoustic characteristics such as the target speech and noise energy size, and constructs a hybrid minimum mean square distortionless response beamformer based on the above information, thereby achieving the effect of improving the suppression of interference while extracting the target speech. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is a schematic diagram of a sound pickup scenario provided by an embodiment of the present application;
[0036] Figure 2 This is a flow chart of a method for extracting target speech based on video information assistance provided by an embodiment of the present application;
[0037] Figure 3 This is a flowchart of an algorithm for a beamforming auxiliary parameter estimation module based on video information provided by an embodiment of the present application;
[0038] Figure 4 This is a spectrogram and a corresponding target speech existence probability comparison diagram provided in an embodiment of the present application;
[0039] Figure 5 This is a flowchart of an interference noise covariance matrix estimation module algorithm combined with DOA information provided in an embodiment of the present application;
[0040] Figure 6This is a flowchart of an algorithm for constructing a covariance matrix and a data-dependent covariance matrix module based on spatial information provided by an embodiment of the present application;
[0041] Figure 7 This is a flow chart of a method for constructing a hybrid minimum mean square distortionless response beamformer provided in an embodiment of the present application;
[0042] Figure 8 is a weighting parameter for estimating the received signal spectrogram and the hybrid noise covariance matrix provided in the embodiment of the present application;
[0043] Figure 9 This is a flowchart of an algorithm for calculating diagonal loading coefficients and constructing an MDVR beamformer module provided in an embodiment of the present application;
[0044] Figure 10 This is a flowchart of a target speech extraction module algorithm based on video information assistance provided by an embodiment of the present application;
[0045] Figure 11 This is the test environment and software selection interface provided by the embodiment of the present application;
[0046] Figure 12 This is a spectrum comparison chart of the recorded data and the output results provided in the embodiment of the present application. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.
[0048] In the description of the embodiments of the present application, words such as "exemplary," "for example," or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary," "for example," or "for example" in the embodiments of the present application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary," "for example," or "for example" is intended to present the relevant concepts in a concrete manner.
[0049] In the description of the embodiments of the present application, the term "and / or" is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, B exists alone, and A and B exist at the same time. In addition, unless otherwise specified, the term "plurality" means two or more. For example, "multiple systems" refers to two or more systems, and "multiple screen terminals" refers to two or more screen terminals.
[0050] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly identifying the technical features being referred to. Thus, features specified as "first" or "second" may explicitly or implicitly include one or more of such features. The terms "include," "comprising," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.
[0051] The method proposed in this application is applicable to microphone array systems that can provide video-assisted information, such as smart large screens in intelligent research and conference systems. Without loss of generality, this application uses a uniform array with camera functions as an example. The principles of other arrays are similar and will not be described separately.
[0052] Figure 1 This is a schematic diagram of a sound pickup scenario provided by this application. Figure 1 As shown, the sound pickup device is a uniform microphone array with camera functionality. The initial speech signal received by the microphone array includes the target signal, interference signals, and ambient noise signals. The video information collected by the camera includes the spatial orientation information of the potential speaker and the spatial orientation information of the interference signals.
[0053] Place the microphone array in three-dimensional space, with the array center coinciding with the origin O. Without loss of generality, number the M elements of the microphone array from left to right as m_1, m_2, …, m_M. Arrange the elements in a row with uniform spacing, and denote the element spacing as d0.
[0054] In three-dimensional space, there are G sound sources from the far field to The direction of the incident sound is toward the microphone array. Without loss of generality, Ω0 is set as the target sound source direction, Ω1~Ω g Set as the direction of the interference sound source. and θ0 represent the depression angle and horizontal angle of the target sound source from the far field, respectively. and θ g represent the depression angle and horizontal angle of the g-th interference sound source incident from the far field.
[0055] In one example, the microphone array extracts the target signal based on the initial speech signal received by the microphone array elements, and performs sound source localization and beamforming. Microphone array beamformers can be divided into fixed beamformers and adaptive beamformers. Commonly used fixed beamformers include Delay and Sum Beamformer (DSB) and Superdirective Beamformer (SDB). Adaptive beamformers include Minimum Variance Distortionless Response (MVDR) beamformer and Linearly Constrained Minimum Variance (LCMV) beamformer.
[0056] Beamforming algorithms often require the location information of the sound source. However, in practical applications, due to the influence of reverberation and the non-steady-state characteristics of speech, sound source localization methods that rely solely on microphone reception are difficult, especially for multi-speaker localization in indoor environments.
[0057] Based on this, the present invention provides a method for extracting target speech based on video-assisted information. This method combines video-assisted information with direct-to-reverberation ratio estimation and spatial energy preprocessing in a microphone array. Furthermore, a hybrid spatial filtering beamformer is constructed based on this auxiliary information, thereby extracting the target direction sound source and suppressing noise. The method includes the following steps:
[0058] (1) Using the video information from the camera to find the spatial orientation of the potential speaker, the orientation information is used to estimate the direct-to-reverberation ratio and perform beam preprocessing for the subsequent design of the hybrid beamformer;
[0059] (2) Designing a covariance matrix based on spatial information according to the potential speaker's position, designing a data-dependent covariance matrix according to the microphone's received signal, and mixing them according to the estimated direct-reverberation ratio and acoustic characteristics such as the energy of each sound source to obtain a new hybrid covariance matrix;
[0060] (3) In order to further reduce the "self-cancellation" phenomenon of the adaptive beamformer caused by the target signal mixed in the covariance matrix, this application designs a target speech presence probability function, which is used to automatically increase the diagonal loading coefficient when there is a target, and reduce the diagonal loading coefficient when there is no target or the target signal energy is weak, so that the noise can be effectively suppressed and the target speech can be retained according to the actual situation.
[0061] The actual experimental results show that the solution proposed in this application can significantly reduce interference from other directions and spatial reverberation while retaining the target speech, and has important application value.
[0062] Figure 2 This is a flowchart of a target speech extraction method based on video information assistance provided by this application.
[0063] like Figure 2 Said method includes the following steps S201 to S204.
[0064] In step S201, a microphone array system with a camera is used to obtain video information and voice signals; the voice signals are collected by microphone elements in the microphone array system; the video information is collected by the camera in the microphone array system; the video information includes spatial orientation information of a potential speaker.
[0065] Assume that the speech time domain signal received by the microphone array is x(t), and x(t) is subjected to short-time Fourier transform (STFT) to obtain the lth frame, N FFT The kth spectral component of the point FFT:
[0066]
[0067] Among them, X m (k, l) represents the received signal of the mth microphone, S d (k, l) and S g (k, l) represent the target signal value and the g-th interference signal value respectively, and Denote the steering vectors of the target signal and the g-th interference signal respectively, and W(k, l) denotes the noise signal. Assume that the target sound source is from the far field with The steering vector of the target signal is:
[0068]
[0069] Among them, p m =[p x,m , P y,m , p z,m ] T is the coordinate of the mth microphone in the three-dimensional Cartesian coordinate system, m = 1, 2, 3...M is the microphone number, c is the speed of sound, f k is the frequency corresponding to the kth spectral component. Similarly, the steering vector of the gth interference sound source is There are also similar expressions, which will not be described in detail.
[0070] In one example, a suitable hybrid spatial filter beamformer needs to be designed to filter the received signal X(k, l) to obtain a frequency domain signal of the target speech enhancement.
[0071] For example, the received signal X(k, l) is filtered using the M×1 dimensional complex weight vector w(k, l) to obtain the enhanced signal value Y(k, l):
[0072] Y(k, l) = w H (k, l)X(k, l) (3)
[0073] Finally, the output signals of all frequency points are subjected to inverse short-time Fourier transform (ISTFT) to obtain the time domain output signal of the target speech.
[0074] In practical applications, the accuracy of noise covariance matrix estimation using adaptive beamforming methods directly impacts adaptive beamforming performance. Traditional adaptive beamforming algorithms (such as MVDR) can experience "self-cancellation" in situations with mismatched steering vectors and low snapshots, resulting in target signal suppression and distortion. Furthermore, adaptive beamformers suffer from high computational complexity and poor robustness.
[0075] In order to improve the performance of the adaptive beamforming algorithm, the embodiment of the present application estimates the auxiliary parameters of the beamforming based on the video information. Figure 3 FIG. 1 is a flowchart of an algorithm for a beamforming auxiliary parameter estimation module based on video information provided in an embodiment of the present application. The estimation of auxiliary parameters is divided into the following two steps:
[0076] (1) Based on the azimuth information of the potential speech signal in the video, the direct-to-reverberation ratio is calculated using array combinations of different spacings to determine the severity of the reverberation;
[0077] (2) Beam preprocessing is performed in a specific direction based on the azimuth information of the potential interfering noise signal to determine whether there is an audio signal in that direction, thereby effectively improving the subsequent noise covariance estimation performance and providing auxiliary information for the diagonal loading design.
[0078] The above two-step auxiliary parameter estimation process is reflected in step S202 to step S203.
[0079] In step S202, multiple array element pairs are formed using two microphone array elements at different intervals to divide the speech signal into multiple continuous frequency bands in the frequency domain. In combination with the spatial orientation information of the potential speaker, the direct-to-reverberation ratio of the target speech in the received signal at each frequency point within the multiple frequency bands is estimated based on the microphone array pairs at different intervals.
[0080] In practical applications, the direction of a potential speaker can be determined using video information provided by devices such as cameras. This directional information can then be used to estimate the direct-to-reverberation ratio for subsequent processing. First, multiple microphone pairs consisting of two different microphones are selected. Based on the potential spatial target direction Ω0, the corresponding two-dimensional pointing direction θ0 for the two-microphone case is calculated. Based on this target direction information, a probability function for the direct-to-reverberation ratio of the target component in the received signal corresponding to different frequencies is designed:
[0081]
[0082] Among them, Re(·) is the real part operation, C nn (k) is the theoretical diffuse field noise covariance matrix corresponding to the kth frequency point, C ss (k) is the direct sound covariance matrix of the kth frequency point corresponding to the target θ0 direction, C xx (k, l) is the coherence function calculated from the array received signal. The expressions of the above three parameters are:
[0083] C ss (k) = exp(j2πf k / c·dsinθ0) (5)
[0084]
[0085]
[0086] in, is the cross-correlation covariance matrix of the received signals of the u-th microphone and the v-th microphone corresponding to the k-th frequency point of the l-th frame, is the autocorrelation covariance matrix of the signal received by the u-th microphone, is the autocorrelation covariance matrix of the vth microphone, α is the smoothing factor, and its typical value is 0.68. d is the spacing between microphones, f k is the frequency corresponding to the kth frequency point, and c is the speed of sound.
[0087] In one example, the element spacing d0 of a 12-element uniform microphone linear array is 2 cm. Taking into account the effects of frequency resolution and frequency aliasing, microphone pairs with different apertures are used in different frequency bands to calculate the corresponding direct-reverberation ratio probability function. Generally speaking, in order to improve resolution, the larger the aperture, the better. However, as the aperture increases, "spatial aliasing" is more likely to occur. Therefore, the frequency corresponding to half a wavelength is generally used as the upper frequency limit to avoid "aliasing". Taking two microphones with a spacing of d as an example, the corresponding frequency upper limit is f = c / (2d0), so the frequency division is performed as follows:
[0088] (1) 0Hz to c / (2(M-1)d0)Hz, the received signals of the 1st mic and the Mth mic are used for calculation, and the direct-to-reverberation ratio probability function P of the target signal component is calculated using the method of formulas (4)-(7) cdr (k, l), spacing d = (M-1)d0;
[0089] (2)c / (2(M-1)d0)Hz to c / (2(M-2)d0)Hz, the received signals of the 1st mic and the M-1th mic are used for calculation, and the direct-to-reverberation ratio probability function P of the target signal component is calculated using the method of formulas (4)-(7) cdr (k, l), spacing d = (M-2)d0;
[0090] (3) Similarly, the qth frequency band is from c / (2(M-q+1)d0)Hz to c / (2(Mq)d0)Hz, and the received signals of the 1st mic and the M-q+1th mic are used for calculation. The direct-to-reverberation ratio probability function P of the target signal component is calculated using the method of formulas (4)-(7) cdr (k, l), spacing d = (Mq)d0.
[0091] (4) When the frequency is higher than c / (2d0) Hz, aliasing cannot be avoided due to aperture limitation, and the reverberation ratio probability function P can only be directly reached by using the target signal component corresponding to the spacing d = d0. cdr (k, l) is calculated.
[0092] The above method can be used to estimate the direct-to-reverberation ratio in a specified direction as accurately as possible without aliasing.
[0093] The direct reverberation ratio estimation method based on video auxiliary information proposed in this application has the following advantages: First, for multiple potential sound source positions, the directional information is used to estimate the direct reverberation ratio corresponding to the direction, so that the reverberation degree corresponding to the sound sources in different directions can be determined. For example, in indoor conditions, the reverberation of sound sources at long distances is generally larger, and the reverberation of sound sources at close distances is generally smaller. By judging the reverberation degree of the directional target sound source, it can be used for the subsequent estimation of the noise covariance matrix; secondly, video information can only provide the direction of the potential speaker, but due to the discontinuity of speech, whether the potential target is speaking and the speaking volume cannot be obtained from the video information. Therefore, it is necessary to perform beam preprocessing on a specific direction based on the directional information of the potential speech signal to determine whether there is an audio signal in this direction in different time periods. In this way, the energy information of the potential speaker can be monitored in real time for the subsequent estimation of the noise covariance matrix.
[0094] In step S203, the frequency domain signal of the speech signal is preprocessed by beamforming in combination with the spatial orientation information of the potential speaker; the probability of the target speech in the received signal at each frequency point is determined by using the energy ratio of the frequency domain signal before and after beamforming preprocessing and the direct-to-reverberation ratio of the target speech in the received signal at each frequency point.
[0095] To further improve the robustness of the adaptive beamformer and reduce computational complexity, a directional preprocessing approach is used to calculate the signal energy at each frequency point in the direction of the potential target. This determines the presence of audio signals in that direction and monitors them in real time. This is then used to estimate the noise covariance matrix and design beamloading parameters. The preprocessing method is only used for subsequent parameter control. To reduce computational complexity, this method employs a super-directional fixed beamforming method based on diagonal loading.
[0096] Specifically, before using the energy ratio of the frequency domain signals before and after beamforming preprocessing to determine the probability of the target speech in the received signals at each frequency point, the energy ratio needs to be corrected to obtain a new energy ratio; the above correction is used to offset the attenuation error caused by the increase in the target speech energy with the frequency.
[0097] The specific calculation formula is:
[0098]
[0099] in, is the steering vector corresponding to the M microphones pointing to the target direction Ω0, is the diffuse field noise covariance matrix of the M microphones corresponding to the kth frequency point. In practical applications, in order to reduce the white noise amplification phenomenon, diagonal loading can be appropriately performed, I M is an M×M dimensional unit matrix, ε0(k) is the diagonal loading coefficient, which is usually an empirical value. The beam output signal is expressed as Y(k, l) = w sd (k)X(k, l).
[0100] Then, the energy ratio before and after the beam output is used, and combined with the direct-to-reverberation ratio function, to construct a new probability function for the existence of the target speech under reverberation conditions. When the reverberation is low, the direct sound has a clear advantage over the reverberation sound energy. If there is an audio signal in the target direction, the energy ratio before and after the beam output is larger. If there is no target signal or other interference sounds are large, the ratio is smaller. However, in the case of large reverberation, the reverberation can be regarded as a sound signal incident from all directions, and the direct sound in the target direction does not have an advantage over the reverberation sound energy. At this time, regardless of whether there is a target signal or not, the energy ratio before and after the beam output is relatively low. It is inaccurate to use only this ratio to judge whether there is a target signal. In order to effectively represent the probability of the existence of the target speech, the present application constructs a new probability function for the existence of the target speech under reverberation conditions. The specific steps are as follows:
[0101] (1) Take the average of the received signal amplitudes of the M microphones at the kth frequency point in the lth frame, denoted as mean(|X(k, l)|), and calculate its ratio to the beam output signal amplitude rt(k, l) = |w sd (k)X(k,l)| / mean(|X(k,l)|);
[0102] (2) The direct-to-reverberation ratio probability function P calculated using formulas (4)-(7) cdr (k, l), combined with the ratio rt(k, l), a new target speech existence probability function is designed. For the frequency point within the qth frequency band:
[0103] P tar (k, l) = P cdr (k, l) · rt(k, l) b(q) (9)
[0104] Among them, considering that the higher the frequency, the higher the signal energy attenuation, the probability of the target speech represented by rt(k, l) may be underestimated when the frequency is higher. Therefore, b(q) is used here for correction, specifically:
[0105]
[0106] In one example, if Figure 4 The spectrogram and the corresponding target speech presence probability comparison diagram show the target speech presence probability function obtained with a 12-element linear array with a 2cm element spacing, with the target at 90 degrees (i.e., directly in front of the array) and the interference at 135 degrees. Using the 0-4kHz range as an example, the results are compared for a single target speech, a single interference speech, and both simultaneously. The target speech is spoken from 0-33s, the interference speech from 33s to 57s, and both simultaneously from 57s to 95s. Figure 4 The upper figure is the spectrogram of the first channel received signal, and the lower figure is the target speech existence probability function calculated by formula (9). From the experimental results, it can be seen that the method of this application can obtain a higher P value at the time-frequency point corresponding to the target speech. tar (k, l), rather than the P corresponding to the time-frequency point of the target speech tar Therefore, relying on the potential directional information provided by the video system, the method proposed in this application can effectively determine the target speech signal component and the non-target signal component. This information will help estimate the noise covariance matrix in the subsequent beamforming.
[0107] In step S204, a hybrid minimum mean square distortionless response beamformer is constructed based on at least the direct-to-reverberation ratio of the target speech in the received signals at each frequency point and the probability of the target speech in the received signals at each frequency point, in combination with the spatial orientation information of the potential speaker.
[0108] To construct a hybrid minimum mean square distortionless response (MMSR) beamformer, a preliminary noise covariance matrix filter must be constructed. A detailed analysis is as follows: Currently, when commonly used adaptive beamformers (e.g., MVDR beamformers) enhance target speech, if the noise covariance matrix includes the target speech component, this can easily lead to "self-cancellation" of the target speech. In practical applications, the ideal noise covariance matrix is generally unknown. To accurately estimate the noise covariance matrix from the received signal, scholars at home and abroad have proposed numerous methods, such as voice activity detection based on direction of arrival (DOA) estimation. However, in practical applications, DOA determination is not very accurate, prone to false alarms and missed detections; and the Complex Gaussian Mixture Model (CGMM) method, which is computationally intensive and difficult to meet the real-time requirements of applications.
[0109] In practical applications, the spatial characteristics of noise are often complex, and may be dominated by direct sound or reverberant sound. When direct sound is dominant, an ideal noise covariance matrix based on azimuth information can often be designed, and a deep null can be designed in the interference direction using a fixed beamforming method. This method can avoid including target components in the covariance matrix, but it can only achieve a good nulling beam effect when the designed null direction matches the actual noise direction and the interference sound is dominated by direct sound. When diffuse field noise, such as reverberant noise, is dominant, the noise source energy is no longer primarily based on the physical direction of the noise source, but comes from various reflection directions such as walls. These reflection directions are unknown and time-varying. Therefore, an adaptive beamformer is often required to achieve a good noise suppression effect. However, in this case, the target signal components may be mixed into the noise covariance matrix estimate, causing distortion of the target signal.
[0110] In order to solve the above problems, Figure 5 As shown, this application proposes an interference noise covariance matrix estimation process combined with DOA information, which is divided into the following two steps:
[0111] (1) Designing a covariance matrix based on spatial information according to the video information, and using the received signal to design the covariance matrix corresponding to the data-dependent beamformer;
[0112] (2) According to the estimated spatial acoustic characteristics such as the direct-to-reverberation ratio of the target sound source and the energy of each potential noise source, a new mixed covariance matrix is obtained.
[0113] above Figure 5 In step (1) of the process shown in the figure, the covariance matrix constructed based on spatial information is designed according to the video information, and the covariance matrix corresponding to the data-dependent beamformer is designed using the received signal; the specific module algorithm is as follows Figure 6 As shown in .
[0114] Figure 6 This is a flowchart of the algorithm for constructing a covariance matrix and a data-dependent covariance matrix module based on spatial information provided by an embodiment of the present application. The technical solution for constructing a covariance matrix and a data-dependent covariance matrix based on spatial information is as follows:
[0115] The video information also includes spatial orientation information of the interference sound source; a data-independent covariance matrix is constructed based on the steering vector of the interference sound source; and the steering vector of the interference sound source is determined based on the spatial orientation information of the interference sound source and the spacing between microphone array elements in the microphone array system.
[0116] According to the spatial orientation information of the interfering sound source, the frequency domain signal of the speech signal is preprocessed with a diagonally loaded super-directional fixed beam to obtain a beam output signal in the direction of the interfering sound source; the direction of the interfering sound source is determined according to the spatial orientation information of the potential interfering sound source.
[0117] A new noise covariance matrix is constructed based on the beam output signal in the direction of the interference sound source and the data-independent covariance matrix.
[0118] A data-dependent covariance matrix is constructed based on the frequency domain signal of the speech signal; the data-dependent covariance matrix does not consider the spatial orientation information of the potential speaker and the spatial orientation information of the interfering sound source.
[0119] Specifically, based on the video information, it is assumed that there are G interference directions where potential interference sound sources may exist, and the corresponding spatial angle is Ω g , g = 1, ..., G, for each direction Ω g , the corresponding direction Ω g The direction of the steering vector is recorded as Then the noise covariance matrix based on spatial information in this direction can be constructed as:
[0120]
[0121] For the gth potential interference direction, its energy needs to be estimated to construct a new covariance matrix. A similar form of formula 9) is used, which corresponds to the super-directional fixed beamformer w based on diagonal loading at the kth frequency point. g (k), and get the beam output Yg (k, l):
[0122]
[0123]
[0124] in, Indicates pointing to Ω g Steering vector of direction, Y g (k, l) points to Ω g Using equations (11)-(13), a new noise estimation covariance matrix is constructed:
[0125]
[0126] Among them, λ g (k, l)=| Y g(k,l)| 2 Represents the energy of the g-th beam output signal.
[0127] If the directional information is not considered, the covariance matrix of the received signal of the data-dependent beamformer is estimated as follows:
[0128] R e (k, l) = (1-β)R e (k, l-1)+βX(k, l)X H (k, l) (15)
[0129] Among them, β is the smoothing factor.
[0130] In one example, the present application performs simulation using β=0.85 as an example.
[0131] above Figure 5 In step (2) of the process shown, the new mixed covariance matrix is obtained by mixing the estimated direct-to-reverberation ratio, the energy of each potential noise source and other spatial acoustic characteristics. The specific steps are as follows: Figure 7 As described in steps S701-S702.
[0132] Figure 7 This is a flowchart of a method for constructing a hybrid minimum mean square distortionless response beamformer provided in an embodiment of the present application, including steps S701 to S704, which is a further detailed description of the technical solution of step S204 above.
[0133] In step S701, the mixing coefficients and diagonal loading coefficients of the mixing covariance matrix are calculated according to the direct-to-reverberation ratio of the target speech in the received signals at each frequency point and the existence probability of the target speech in the received signals at each frequency point.
[0134] In step S702, a mixed noise covariance matrix is constructed according to the data-dependent covariance matrix, the new noise covariance matrix, and the mixing coefficient of the mixed covariance matrix using a weighted algorithm.
[0135] In one example, calculating the mixing coefficients of the mixing covariance matrix includes determining a calculation form of the mixing coefficients of the mixing covariance matrix based on a comparison relationship between a first parameter and a second parameter that demarcate reverberation intensity and a direct-to-reverberation ratio of a target speech in a received signal at each frequency point; and determining the mixing coefficients of the mixing covariance matrix according to the calculation form using the direct-to-reverberation ratio of the target speech in the received signal at each frequency point and the probability of existence of the target speech in the received signal at each frequency point.
[0136] The first parameter and the second parameter for dividing the reverberation intensity are determined based on empirical values.
[0137] In one example, calculating the mixing coefficients of the mixing covariance matrix further includes performing inter-frame smoothing and inter-frequency smoothing on the mixing coefficients of the mixing covariance matrix to obtain smoothed mixing coefficients of the mixing covariance matrix.
[0138] In one example, a mixed noise covariance matrix is constructed, including:
[0139] The data-related covariance matrix and the noise covariance matrix are processed to obtain the normalized data-related covariance matrix and the noise covariance matrix:
[0140] The normalized data-related covariance matrix and the noise covariance matrix are subjected to a weighting algorithm according to the mixing coefficient of the mixed covariance matrix to construct a mixed noise covariance matrix.
[0141] In practical applications, in order to better design the noise covariance matrix in the adaptive filter, this application adopts a hybrid noise covariance matrix estimation method based on auxiliary parameters. The principle and process are as follows:
[0142] (1) For the signal of the kth frequency point in the lth frame, use formula (4) to determine the indoor reverberation level P cdr (k, l), use formula (9) to determine the probability function P of the target speech tar (k, l);
[0143] (2) In order to reduce the impact of signal energy, prevent the two R w (k, l) and R e The large energy difference between (k, l) leads to the failure of the weighted algorithm. and Normalization is performed to accurately reflect the spatial characteristics of the noise covariance matrix. The specific solution is as follows:
[0144]
[0145] in, for The modulus of the i-th element on the diagonal. Similarly, The specific solution is as follows:
[0146]
[0147] in, for The modulus of the i-th element on the diagonal.
[0148] According to the judgment result in process (1) and Perform weighted mixing to reduce the noise level in high signal-to-noise ratio and low reverberation conditions. The proportion of components, thereby reducing the target speech components in the estimated covariance matrix to reduce speech distortion; in the case of low signal-to-noise ratio and high reverberation, The ratio of components, that is, retaining more actual noise signal components, thereby improving noise suppression performance. For the kth frequency point, it is judged to belong to the qth frequency band, and the weighting method is as follows:
[0149]
[0150]
[0151] Among them, η1 and η2 represent the boundaries of reverberation intensity, with typical values of η1 = 0.3 and η2 = 0.7. The design idea of the weighting factor κ is as follows: In the case of small and medium reverberation, the sound source is mostly direct sound, so P tar (k, l) is more accurate, at this time P tar (k, l) is used to determine whether there is a target signal, and P cdr The smaller (k, l) is, the greater the P tar The larger the reverberation, the more reliable the result of (k, l); when the reverberation is large, it means that the incoming wave direction is more complicated and chaotic. Even if there is no speech signal in the target direction, it is possible that the reflection from other directions will cause P tar (k, l) is large, so P tar (k, l) The judgment is not accurate enough, and its credibility should be appropriately reduced. tar (k, l) and P cdr (k, l) jointly determine; when the reverberation is particularly large, P tar The feasibility of (k, l) is poor, that is, it has basically no reference value, so it is completely based on P cdr (k, l) is used for weighted control.
[0152] In one example, calculating the mixing coefficients of the mixing covariance matrix further includes performing inter-frame smoothing and inter-frequency smoothing on the mixing coefficients of the mixing covariance matrix to obtain smoothed mixing coefficients of the mixing covariance matrix.
[0153] That is, in order to improve the robustness and continuity of the processed speech, κ(k, l) can be smoothed between frames and between frequencies, that is,
[0154]
[0155]
[0156] Where L0 is the number of smoothed frames, typically 3, and K0 is the number of smoothed frequencies, typically 2.
[0157] Figure 8 κ(k, l) is the weighting parameter for estimating the received signal spectrogram and hybrid noise covariance matrix provided in the embodiment of the present application. It shows the received signal spectrogram and the corresponding weighting factor κ(k, l) calculated by formula (19) under different noise types under room reverberation conditions using a 12-element linear array with a spacing of 2 cm. The target speech is located at 90 degrees, the noise is located at 135 degrees, and the noise types are white noise (0-7s), speech interference (7-14s), and engine noise (14-24s). Figure 8 The upper figure is the received signal spectrogram, and the lower figure is the weighting factor κ(k, l) calculated by formula (19). Figure 8 The results show that the weighted parameter design method proposed in this application can obtain a larger value under the target condition, so that the actual signal The components of the target component are relatively small, that is, the actual signal is included as little as possible in the estimated noise covariance matrix. When there is no target or the target component is small, a smaller weighting parameter is obtained, thereby increasing the actual signal component in the noise covariance matrix. The above results show that the weighting method proposed in this application can effectively reduce the target signal component that may be mixed in the noise covariance estimation matrix.
[0158] In one example, calculating the diagonal loading coefficient includes determining a calculation form of the diagonal loading coefficient based on a comparison relationship between a first parameter and a second parameter that demarcate reverberation intensity and a direct-to-reverberation ratio of a target speech in a received signal at each frequency point; and determining the diagonal loading coefficient according to the calculation form using a probability of the target speech in the received signal at each frequency point.
[0159] The first parameter and the second parameter for dividing the reverberation intensity are determined based on empirical values.
[0160] In order to further reduce the "self-cancellation" phenomenon of the adaptive beamformer caused by the target signal mixed in the covariance matrix, this application designs an automatic loading method based on the direct-reverberation ratio and the specific direction beam output, which can automatically increase the diagonal loading coefficient when there is a target, and reduce the diagonal loading coefficient when there is no target or the target signal energy is weak.
[0161] This application designs the following diagonal loading method:
[0162]
[0163] Among them, η1 and η2 represent the boundaries of reverberation intensity, with typical values of η1 = 0.3 and η2 = 0.7. μ(k, l) design idea: In the case of small reverberation, it completely depends on P tar (k, l) judgment results, but in the case of medium and high reverberation, P tar The accuracy of the (k, l) judgment result is not high, and a lower limit needs to be set to control it. In this application, the reverberation lower limit η1 and the lower limit η2 in the high reverberation case are set to 0.2 and 0.5 respectively.
[0164] In step S703, a diagonally loaded noise covariance matrix is constructed according to the diagonal loading coefficients and the hybrid noise covariance matrix.
[0165] The specific design method is as follows:
[0166]
[0167] Where μ(k, l) represents the diagonal loading coefficient. In practical applications, for the noise covariance matrix When the target signal components are large, the diagonal loading coefficient should be increased. When the target signal components are small or absent, the loading coefficient should be reduced or even not used.
[0168] In step S704, a hybrid minimum mean square distortionless response beamformer is constructed according to the diagonally loaded noise covariance matrix in combination with the spatial orientation information of the potential speaker.
[0169] Specifically, the classic MVDR beamformer is used for filtering to obtain the beam output of the kth frequency point in the lth frame:
[0170]
[0171]
[0172] Figure 9 This is a flowchart of the algorithm for calculating the diagonal loading coefficient and constructing the MDVR beamformer module provided by the embodiment of the present application. Figure 7 The technical solutions of steps S701 to S704 are modularly expressed.
[0173] In one example, after executing step S204, the method further includes filtering the frequency domain signal of the speech signal using a hybrid minimum mean square distortionless response beamformer to obtain a frequency domain signal of the target speech, performing an inverse short-time Fourier transform on the frequency domain signal of the target speech to obtain a time domain signal, and performing windowing and smoothing on the time domain signal to obtain a time domain signal of the target speech.
[0174] In one example, the output signal Y(k, l) is subjected to inverse Fourier transform (ISTFT), windowing, and smoothing to obtain a final time-domain output signal y(t).
[0175] Figure 10 This is a flowchart of a target speech extraction module algorithm based on video information assistance provided by an embodiment of the present application. The specific process is as follows:
[0176] (1) Process the received signal in different frequency bands, use different microphone pairs and combine the directional information of the target signal, and estimate the direct reverberation ratio P of the target component in the received signal at each frequency point according to equations (4)-(7) cdr (k, l);
[0177] (2) According to the potential target direction information, fixed beamforming preprocessing is performed and the target speech existence probability function P is calculated according to formula (9). tar (k, l);
[0178] (3) Based on the spatial information of the potential interfering sound source provided by the video and the received signal, two covariance matrices are constructed using equations (14) and (15) respectively, and normalized using equations (16)-(17);
[0179] (4) Using equations (18)-(19) to weight the two normalized covariance matrices in step (3), the mixed noise covariance matrix is obtained;
[0180] (5) Substitute the mixed noise covariance matrix into equations (24) and (25) to obtain the final filter coefficients and beam output signal.
[0181] Finally, the proposed method is compared with super-directional fixed beamforming and the traditional MVDR method based on diagonal loading. The MVDR method uses formula (15) to estimate the covariance matrix. The array is a 12-element linear array with a spacing of 2 cm. In the experiment, the interference source is located 2 meters in front of the array at an angle of 90 degrees. The target sound source is located 2 meters to the side of the array at an angle of approximately 20 degrees to the interference sound source. The signal-to-noise ratio is -5 dB.
[0182] Figure 11The test environment and software selection interface are given, where circular area 1 is the target direction and circular area 2 is the interference direction.
[0183] Figure 12 This is a spectrogram comparison diagram of the recorded data and the output results provided by the embodiment of the present application. The spectrogram comparison of the beam output results after processing using the method of the present application and the traditional MVDR method based on diagonal loading is based on the video information. From top to bottom, they are the original received signal, the super-directional fixed beam output, the beam output of the traditional MVDR method based on diagonal loading, and the output result of the present application. It can be seen from the results that the super-directional fixed beam forming algorithm has a low noise reduction, and the traditional MVDR method based on diagonal loading has a relatively obvious speech eating phenomenon. The method proposed in this application not only has a high noise reduction, but also has a small distortion of the voice signal. The sound pickup effect under low signal-to-noise ratio conditions has obvious advantages over the other two methods.
[0184] This application proposes a video-assisted beamforming method that can effectively extract the target speech and suppress other interference and reverberation. Based on the key issues commonly encountered in practical applications, innovative solutions are proposed, which are specifically reflected in the following aspects:
[0185] (1) Although video information can provide possible speech signal orientations, the discontinuous characteristics and reverberation characteristics of speech will affect the subsequent noise covariance matrix estimation. Therefore, a direct-to-reverberation ratio and spatial energy preprocessing method based on video auxiliary information is proposed to assist in the subsequent noise covariance matrix estimation. First, based on the orientation information of the potential speech signal in the video, the direct-to-reverberation ratio is calculated using array combinations of different apertures to determine the severity of the reverberation. Second, based on the orientation information of the potential speech signal, beam preprocessing is performed on a specific direction to determine whether there is an audio signal in that direction at each moment. The above method can effectively determine the degree of reverberation and whether there is an audio signal in the direction of interest, thereby effectively improving the subsequent noise covariance estimation performance and providing auxiliary information for the design of diagonal loading.
[0186] (2) A hybrid noise covariance matrix estimation method based on the direct-reverberation ratio and spatial energy preprocessing results is proposed. First, a covariance matrix based on spatial information is designed according to the orientation of the potential interfering sound source, and a data-dependent covariance matrix is designed according to the microphone received signal. Then, the covariance matrix is mixed according to the estimated direct-reverberation ratio and the energy of each sound source, thereby obtaining a new hybrid covariance matrix. Actual experimental results show that the hybrid noise covariance matrix estimation method proposed in this application can reduce the target speech component in the estimated covariance matrix under high signal-to-noise ratio conditions, thereby reducing speech distortion; and still retain more actual noise signal components under low signal-to-noise ratio conditions, thereby improving noise suppression performance.
[0187] (3) In order to further reduce the "self-cancellation" phenomenon of the adaptive beamformer caused by the target signal mixed in the covariance matrix, this application designs a target speech existence probability function, which is jointly determined by the direct reverberation ratio in item (1) and the specific directional beam output result. It can automatically increase the diagonal loading coefficient when there is a target, and reduce the diagonal loading coefficient when there is no target or the target signal energy is weak, so that the noise can be effectively suppressed and the target speech can be retained according to the actual situation.
[0188] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0189] It is understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not intended to limit the scope of the embodiments of the present application. It should be understood that in the embodiments of the present application, the order of the sequence numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0190] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of this application in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of this application should be included in the scope of protection of this application.
Claims
1. A target speech extraction method based on video information assistance, characterized in that: include: Use a microphone array system with a camera to obtain video information and voice signals; The speech signal is collected by the microphone elements in the microphone array system; The video information is collected by a camera in the microphone array system; the video information includes spatial orientation information of a potential speaker; Multiple array element pairs are formed by using two microphone array elements with different spacings to divide the speech signal into multiple frequency bands with continuous frequencies in the frequency domain. In combination with the spatial orientation information of the potential speaker, estimating the direct-to-reverberation ratio of the target speech in the received signal at each frequency point within the multiple frequency bands according to microphone array pairs with different spacings; Performing beamforming preprocessing on the frequency domain signal of the speech signal in combination with the spatial orientation information of the potential speaker; determining the presence probability of the target speech in the received signal at each frequency point using the energy ratio of the frequency domain signal before and after the beamforming preprocessing and the direct-to-reverberation ratio of the target speech in the received signal at each frequency point; In combination with the spatial orientation information of the potential speaker, a hybrid minimum mean square distortionless response beamformer is constructed based on at least the direct-to-reverberation ratio of the target speech in the received signals at each frequency point and the existence probability of the target speech in the received signals at each frequency point.
2. The method according to claim 1, characterized in that The method further includes filtering the frequency domain signal of the speech signal using the hybrid minimum mean square distortionless response beamformer to obtain the frequency domain signal of the target speech.
3. The method according to claim 2, characterized in that The method further includes performing an inverse short-time Fourier transform on the frequency domain signal of the target speech to obtain a time domain signal; and performing windowing and smoothing on the time domain signal to obtain the time domain signal of the target speech.
4. The method according to claim 1, wherein Before determining the existence probability of the target speech in the received signals at each frequency point using the energy ratio of the frequency domain signals before and after the beamforming preprocessing, it is also necessary to correct the energy ratio to obtain a new energy ratio; The correction is used to offset the attenuation error caused by the target speech energy increasing with the frequency.
5. The method according to claim 1, wherein Also includes, The video information also includes spatial orientation information of an interfering sound source; a data-independent covariance matrix is constructed based on a steering vector of the interfering sound source; the steering vector of the interfering sound source is determined based on the spatial orientation information of the interfering sound source and the spacing between microphone array elements in the microphone array system; Performing diagonally loaded super-directional fixed beam preprocessing on the frequency domain signal of the speech signal according to the spatial orientation information of the interfering sound source to obtain a beam output signal in the direction of the interfering sound source; the direction of the interfering sound source is determined according to the spatial orientation information of the potential interfering sound source; Constructing a new noise covariance matrix according to the beam output signal in the direction of the interference sound source and the data-independent covariance matrix; A data-dependent covariance matrix is constructed according to the frequency domain signal of the speech signal; the data-dependent covariance matrix does not consider the spatial orientation information of the potential speaker and the spatial orientation information of the interfering sound source.
6. The method according to claim 5, characterized in that The hybrid minimum mean square distortionless response beamformer is constructed based on at least the direct-to-reverberation ratio of the target speech in the received signals at each frequency point and the existence probability of the target speech in the received signals at each frequency point, including: Calculating the mixing coefficients and diagonal loading coefficients of the mixing covariance matrix according to the direct-to-reverberation ratio of the target speech in the received signals at each frequency point and the existence probability of the target speech in the received signals at each frequency point; Constructing a mixed noise covariance matrix according to the data-dependent covariance matrix, the new noise covariance matrix, and the mixing coefficient of the mixed covariance matrix according to a weighted algorithm; constructing a diagonally loaded noise covariance matrix according to the diagonal loading coefficients and the hybrid noise covariance matrix; The hybrid minimum mean square distortionless response beamformer is constructed according to the diagonally loaded noise covariance matrix in combination with the spatial orientation information of the potential speaker.
7. The method according to claim 6, characterized in that The calculation of the mixing coefficient of the mixing covariance matrix includes: Determining a calculation form for a mixing coefficient of a mixing covariance matrix based on a comparison relationship between a first parameter and a second parameter dividing the reverberation intensity and the direct-reverberation ratio of the target speech in the received signals at each frequency point; determining the mixing coefficient of the mixing covariance matrix according to the calculation form using the direct-reverberation ratio of the target speech in the received signals at each frequency point and the presence probability of the target speech in the received signals at each frequency point; The first parameter and the second parameter for dividing the reverberation intensity are determined based on empirical values.
8. The method according to claim 7, characterized in that The calculating of the mixing coefficients of the mixing covariance matrix further includes performing inter-frame smoothing and inter-frequency smoothing on the mixing coefficients of the mixing covariance matrix to obtain the smoothed mixing coefficients of the mixing covariance matrix.
9. The method according to claim 6, characterized in that The calculation of the diagonal loading coefficient includes: Determining a calculation form for a diagonal loading coefficient based on a comparison relationship between a first parameter and a second parameter dividing the reverberation intensity and a direct-reverberation ratio of the target speech in the received signals at each frequency point; and determining the diagonal loading coefficient according to the calculation form using a probability of the target speech in the received signals at each frequency point. The first parameter and the second parameter for dividing the reverberation intensity are determined based on empirical values.
10. The method according to claim 6, characterized in that The construction of the mixed noise covariance matrix includes: Processing the data-dependent covariance matrix and the noise covariance matrix to obtain normalized data-dependent covariance matrix and noise covariance matrix; A weighted algorithm is executed on the normalized data-related covariance matrix and the noise covariance matrix according to the mixing coefficient of the mixed covariance matrix to construct a mixed noise covariance matrix.
Citation Information
Patent Citations
Speech signal processing method, device and system
CN105679328A
Two-channel beam forming speech enhancement method based on noise mixed coherence
CN105869651A