A method and device for extracting target speech based on masked beamforming
By using masked beamforming, combined with speech pre-separation and beamforming, and optimizing the MVDR beamformer with amplitude scaling mask, the problem of suppressing co-directional interference sources is solved, thereby improving the auditory quality of target speech extraction and speech recognition performance.
Patent Information
- Application Number
- CN202411981356.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-12-30
AI Technical Summary
Existing beamforming-based target speech extraction methods cannot effectively suppress interference sources when the interference speaker's speech and the target speaker's speech come from the same direction, making the suppression of co-directional interference sources a critical problem that urgently needs to be solved.
A masking beamforming-based approach is adopted, which combines microphone array pickup, speech pre-separation and beamforming steps with speaker encoder and mask estimator. It utilizes minimum variance distortionless response beamformer and amplitude scaling mask optimization to achieve two-stage processing of speech separation and beamforming, thereby reducing amplitude and phase distortion in target speech extraction.
It effectively solves the problem of suppressing co-directional interference sources, improves the auditory quality and speech recognition performance of target speech, and reduces amplitude and phase distortion in target source extraction.
Smart Images

Figure CN119943087B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of acoustic signal processing, and in particular to a method and apparatus for extracting target speech based on masking beamforming. Background Technology
[0002] Target speech extraction (TSE) refers to extracting the speech of the target speaker in a multi-speaker scenario, thereby providing strong front-end technical support for subsequent tasks such as high-quality voice communication, speech recognition, and human-computer voice interaction. Current TSE methods can be mainly divided into two categories: masking-based methods and beamforming-based methods.
[0003] Masking-based target speaker speech extraction methods primarily learn mask information from registered speaker speech and observed signals through a neural network constructed using encoding and decoding modules, thereby acquiring the target speaker's speech. Examples include speaker extraction (SpEx) networks and targeted voice separation filters (VoiceFilters). While numerous studies have demonstrated the effectiveness of these methods in extracting target speaker speech, the nonlinear distortion of the network and the effective recovery of phase information remain significant challenges. Therefore, many studies combine speech separation with beamforming to mitigate the amplitude and phase distortion caused by the separation network, resulting in beamforming-based time-domain audio separation networks (Beam-TasNet) and beamforming-guided time-domain audio separation networks (Beam-guided TasNet).
[0004] Beamforming-based methods primarily utilize the spatial filtering characteristics of beamformers to suppress interfering speaker speech and reduce speech distortion. This is often achieved through adaptive beamformers, such as Linearly Constrained Minimum Variance (LCMV) beamformers, Minimum Variance Distortionless Response (MVDR) beamformers, and generalized sidelobe cancellers (GSC). Among these adaptive beamformers, the MVDR beamformer has gained widespread research and application due to its ability to minimize the filtering result of the non-target signal covariance matrix and constrain the target direction to be distortion-free, its simple principle, and its flexible optimization capabilities. However, when the interfering speaker's speech and the target speaker's speech originate from the same direction, the beamformer's main lobe will apply the same filtering process to both sources, resulting in the inability to effectively suppress the interfering source. This characteristic of beamformers makes suppressing co-directional interfering sources a crucial problem that urgently needs to be addressed in current beamforming design.
[0005] Minimum Variance Distortionless Response (MVDR) beamformers and generalized sidelobe cancellers (GSCs) are examples of adaptive beamformers. Among these, the MVDR beamformer has gained widespread research and application due to its ability to minimize the filtering result of the non-target signal covariance matrix and constrain the target direction to be distortion-free, its simple principle, and its flexible optimization capabilities. However, when the interfering speaker's speech and the target speaker's speech come from the same direction, the beamformer's main lobe will perform the same filtering process on both sources, resulting in the inability to effectively suppress the interfering source. This characteristic of beamformers makes suppressing co-directional interfering sources a crucial problem that urgently needs to be addressed in current beamforming design. Summary of the Invention
[0006] This application discloses a method and apparatus for target speech extraction based on masking beamforming.
[0007] In a first aspect, this application discloses a target speech extraction method based on masking beamforming, the method comprising:
[0008] The microphone array pickup step involves acquiring multi-channel observation signals through a microphone array.
[0009] The first stage of speech pre-separation involves extracting the feature vector of the target speech by the speaker encoder through the registered speaker speech information, extracting the coding features of each observed signal by the speech encoder, fusing the feature vector and the coding features and obtaining the mask information of all speech sources by the mask estimator, wherein the all speech sources include target speech sources and non-target speech sources, and obtaining the pre-separated temporal speech information through masking processing and decoding mapping.
[0010] The first-stage beamforming target speech separation step uses pre-separated temporal speech information to determine whether the speech sources are in the same direction, and calculates the minimum variance distortion-free weights of the corresponding beamforming based on the judgment result and mask information to achieve speech separation.
[0011] The second-stage speech pre-separation step uses the time-domain speech as an auxiliary input to the target speech separation network and extracts feature vectors. It then combines the feature vectors of the registered speaker's speech and the feature vectors of the separated target speech to construct new fusion features. The mask information is estimated more accurately through a mask estimator and a decoding mapping.
[0012] The second-stage beamforming speech separation step, based on the mask information of the second-stage pre-separation, repeats the first-stage beamforming speech separation process until the final target speech extraction is completed.
[0013] In one exemplary embodiment of this application, the first-stage processing step of the method further includes:
[0014] The multi-channel observation signals from the microphone are processed based on a preset speech encoder to obtain the coding features of the observed speech signals.
[0015] Based on a preset mask estimator, the mask information of all speech sources is estimated by taking the feature vector of the target speech and the coding features of the observed signal as input.
[0016] The encoded features of the observed signal are multiplied with the mask information of all speech sources to obtain the separation features of each speech source, and then mapped by a speech decoder to obtain the separated speech of each speech source.
[0017] Based on the speech separation of each speech source and the preset generalized feature value decomposition method, the steering vector of each speech source is calculated respectively;
[0018] Based on the set discrimination threshold, it is determined whether the target speech source and the non-target speech source are in the same direction;
[0019] Based on the discrimination results, an amplitude scaling mask is obtained to optimize the enhancement results of minimum variance distortionless response beamforming.
[0020] In one exemplary embodiment of this application, the method further includes:
[0021] Based on the registered speech, the feature vector of the registered speech is obtained by using a Bi-directional Long Short-Term Memory (BLSTM) network, activation function, linear layer, and mean pooling module based on the preset speaker encoder.
[0022] The preset microphone array is either a uniform linear array or a uniform circular array.
[0023] In one exemplary embodiment of this application, the method further includes:
[0024] The observed signals from each microphone channel are nonlinearly transformed using a 1×1 convolution, parameterized activation function, and normalization layer (Norm) based on a fully convolutional temporal audio separation network.
[0025] Use depthwise separable convolution instead of standard convolution to reduce the number of parameters;
[0026] The coded features of the observed speech signal are obtained by adding the signal to the input signal after one-dimensional convolution.
[0027] In one exemplary embodiment of this application, the method further includes:
[0028] The feature vector and the encoded features of the observed speech signal are concatenated along the channel dimension to obtain an intermediate representation between the target speaker's speech and the observed signal;
[0029] The spliced features are normalized and nonlinear transformation is performed using 1×1 convolution. Based on the preset convolutional coding module, the features of the speaker's speech are obtained.
[0030] After parameterized activation function, 1×1 convolution and sigmoid activation function, mask information of target speech coding features is obtained.
[0031] In one exemplary embodiment of this application, the method further includes:
[0032] The preset calculation methods include short-time Fourier transform, covariance matrix multiplication, and generalized eigenvalue decomposition.
[0033] In one exemplary embodiment of this application, the method further includes:
[0034] Based on the separated speech sources of each channel, the steering vector of each speech source is obtained through generalized eigenvalue decomposition. The angle between each steering vector is calculated by the error weighted minimization method to determine whether the non-target speech source is in the same direction as the target speech source.
[0035] Secondly, this application discloses a target speech extraction device based on masking beamforming, the device comprising:
[0036] The microphone array pickup module is used to acquire multi-channel observation signals through a microphone array;
[0037] The first-stage processing module is used to extract the feature vector of the target speech by the speaker encoder through the registered speaker speech information, extract the coding features of each observed signal by the speech encoder, obtain the mask information of all speech sources by fusing the feature vector and the coding features and the mask estimator, the total speech sources including target speech sources and non-target speech sources, obtain pre-separated temporal speech information through masking processing and decoding mapping; and determine whether the speech sources are in the same direction by using the pre-separated temporal speech information, and calculate the weight of the minimum variance distortionless corresponding beamforming according to the judgment result and the mask information to achieve speech separation.
[0038] The two-stage processing module is used to use the time-domain speech as an auxiliary input to the target speech separation network and extract feature vectors. It then combines the feature vectors of the registered speaker's speech and the feature vectors of the separated target speech to construct new fused features. The module estimates the mask information more accurately through a mask estimator and a decoding map. Based on the pre-separated mask information, the module repeats the beamforming speech separation process of the first stage until the final target speech extraction is completed.
[0039] Thirdly, this application discloses an electronic device comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to perform the method as described in any of the preceding aspects.
[0040] Fourthly, this application discloses a non-transitory computer-readable storage medium, wherein when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the methods described in any of the preceding aspects.
[0041] Fifthly, this application discloses a computer program product in which, when the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to perform the method described in any of the preceding aspects.
[0042] The target speech extraction method of this application includes: a microphone array pickup step, which acquires multi-channel observation signals; a first-stage processing step, which obtains the feature vector of the target speech and the encoded features of the observation signals based on the speaker encoder and the speech encoder, obtains the masks of the target speech source and non-target speech source through feature fusion and mask estimator, obtains time-domain separated speech through masking processing and decoding mapping, and achieves further separation of the sound source through mask-based minimum variance distortionless response beamforming to obtain time-domain speech; and a second-stage processing step, which uses the time-domain speech as an auxiliary input of the target speech separation network to construct new fusion features, and repeats the first-stage processing step until the target speech extraction is completed.
[0043] This application fuses the feature vector of the registered target speech with the coded features of the observed signal to achieve speech separation capable of identifying the target speech source. Combining the speech separation results, it obtains the steering vector through GEVD (Generalized Eigenvalue Decomposition) to determine whether the target source and competing sources are in the same direction. It then optimizes the MVDR main lobe without distortion using an amplitude scaling mask between in-direction sources, suppressing competing speech sources in the same direction while ensuring effective recovery of phase information. A two-stage reinforcement training method is employed, and a multi-objective training loss function is constructed using SNR, MSE, and SI-SDR, effectively improving the auditory quality and speech recognition performance of the extracted target speech. This application combines target source extraction with speech separation and beamforming, effectively reducing amplitude and phase distortion in target source extraction. The amplitude scaling mask of in-direction speech sources is used to optimize the MVDR beamformer, effectively solving the difficulty of effectively eliminating competing speech sources in scenarios with in-direction speech sources. Attached Figure Description
[0044] Figure 1 This is a flowchart of the steps of a target speech extraction method based on masking beamforming according to this application.
[0045] Figure 2 This is a two-stage target speaker speech extraction framework diagram of a target speech extraction method based on masking beamforming proposed in this application.
[0046] Figure 3 This is a structural diagram of the target speech separation network of a target speech extraction method based on masking beamforming according to this application.
[0047] Figure 4 This is a flowchart of MVDR beamforming extraction of target speech, a target speech extraction method based on masked beamforming, according to this application.
[0048] Figure 5This is a detailed structural diagram of each module of the target speech separation network of the target speech extraction method based on masking beamforming in this application.
[0049] Figure 6 This is a structural block diagram of a target speech extraction device based on masking beamforming according to this application.
[0050] Figure 7 This is a block diagram of an electronic device according to this application.
[0051] Figure 8 This is a block diagram of a computer-readable storage medium according to this application. Detailed Implementation
[0052] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0053] Reference Figure 1 The diagram illustrates a flowchart of a target speech extraction method based on masking beamforming, which can be applied to electronic devices. Specifically, the method may include the following steps:
[0054] Microphone array pickup step S110: Acquire multi-channel observation signals through microphone array;
[0055] In the first stage of speech pre-separation step S120, the feature vector of the target speech is extracted by the speaker encoder through the registered speaker speech information, the coding features of each observation signal are extracted by the speech encoder, the feature vector and the coding features are fused and the mask information of all speech sources is obtained by the mask estimator, the all speech sources include target speech sources and non-target speech sources, and the pre-separated temporal speech information is obtained through masking processing and decoding mapping.
[0056] In the first stage of beamforming target speech separation step S130, the speech sources are judged to be in the same direction by pre-separated temporal speech information, and the weight of the corresponding beamforming with minimum variance and no distortion is calculated according to the judgment result and mask information to achieve speech separation.
[0057] In the second stage of speech pre-separation step S140, the time-domain speech is used as the auxiliary input of the target speech separation network and feature vectors are extracted. The feature vectors of the registered speaker's speech and the feature vectors of the separated target speech are combined to construct new fusion features. The mask information is estimated more accurately through the mask estimator and decoding mapping.
[0058] In the second-stage beamforming speech separation step S150, based on the mask information of the second-stage pre-separation, the first-stage beamforming speech separation process is repeated until the final target speech extraction is completed.
[0059] The target speech extraction method based on masking beamforming in this application includes: a microphone array pickup step, in which multi-channel observation signals are acquired through a microphone array; a first-stage processing step, in which feature vectors of the target speech and encoded features of the observation signals are obtained based on a speaker encoder and a speech encoder, masks of the target speech source and non-target speech source are obtained through feature fusion and a mask estimator, temporally separated speech is obtained through masking processing and decoding mapping, and the sound source is further separated through masking-based minimum variance distortionless response beamforming to obtain temporal speech; and a second-stage processing step, in which the temporal speech is used as an auxiliary input to the target speech separation network to construct new fusion features, and the first-stage processing step is repeated until the target speech extraction is completed. This application fuses the feature vector of the registered target speech with the coded features of the observed signal to achieve speech separation capable of identifying the target speech source. Combining the speech separation results, it obtains the steering vector through GEVD to determine whether the target source and competing sources are in the same direction. It then optimizes the MVDR main lobe without distortion using an amplitude scaling mask between in-direction sources, suppressing competing speech sources in the same direction while ensuring effective recovery of phase information. A two-stage reinforcement training method is employed, and a multi-objective training loss function is constructed using SNR, MSE, and SI-SDR, effectively improving the auditory quality and speech recognition performance of the extracted target speech. This application combines target source extraction with speech separation and beamforming, effectively reducing amplitude and phase distortion in target source extraction. The amplitude scaling mask of in-direction speech sources is used to optimize the MVDR beamformer, effectively solving the difficulty of effectively eliminating competing speech sources in scenarios with in-direction speech sources.
[0060] Example 1:
[0061] In this example embodiment, speech separation refers to separating multiple speech source signals in a multi-speech source scenario using certain processing techniques, such as beamforming, independent vector analysis, and neural network mapping.
[0062] A mask generally refers to the ratio of the target signal to the noisy signal under certain conditions. It is usually estimated by a supervised training of a neural network model using an ideal mask of a certain form, such as an ideal binary mask, an ideal amplitude scaling mask, or an ideal complex scaling mask. It can be used to achieve speech enhancement, speech separation, and target speech extraction.
[0063] Minimum variance distortionless response beamforming is a spatial filtering technique that achieves adaptive interference source suppression and noise reduction by forming a distortionless filtering main lobe in the desired direction, attenuating filtering side lobes in the undesired direction, and adaptively forming null traps in the direction of interference sources.
[0064] Target speech extraction refers to extracting the speech of the speaker of interest from the observation signal collected by the microphone in noisy, reverberant, and multi-source scenarios.
[0065] In the microphone array pickup step S110, multi-channel observation signals can be acquired through the microphone array.
[0066] In this example embodiment, the first-stage processing step of the method further includes:
[0067] The array observation signal is processed based on a preset speech encoder to obtain the coding features of the observed speech signal;
[0068] Based on a preset mask estimator, the mask information of all speech sources is estimated by taking the feature vector of the target speech and the coding features of the observed signal as input.
[0069] The encoded features of the observed signal are multiplied with the mask information of all speech sources to obtain the separation features of each speech source, and then mapped by the speech decoder to obtain the separated speech of each speech source.
[0070] Based on the speech separation of each speech source and the preset generalized feature value decomposition method, the steering vector of each speech source is calculated respectively;
[0071] Based on the set discrimination threshold, it is determined whether the target speech source and the non-target speech source are in the same direction;
[0072] Based on the discrimination results, an amplitude scaling mask is obtained to optimize the enhancement results of minimum variance distortionless response beamforming.
[0073] In this example embodiment, the method further includes:
[0074] Based on the registered speech, the feature vector of the registered speech is obtained using a bidirectional long short-term memory network based on a preset speaker encoder, the ReLU activation function, a linear layer, and an average pooling module.
[0075] The microphone array of the preset number of microphones is a uniform linear array or a uniform circular array.
[0076] In this example embodiment, the method further includes:
[0077] The observation signal from the preset microphone is nonlinearly transformed using 1×1 convolution, parameterized activation function PReLU, and normalization layer based on the fully convolutional temporal audio separation network Conv-TasNet.
[0078] Use depthwise separable convolution instead of standard convolution to reduce the number of parameters;
[0079] The coded features of the observed speech signal are obtained by adding the signal to the input signal after one-dimensional convolution.
[0080] In this example embodiment, the method further includes:
[0081] The feature vector and the encoded features of the observed speech signal are concatenated along the channel dimension to obtain an intermediate representation between the target speaker's speech and the observed signal;
[0082] The spliced features are normalized and nonlinear transformation is performed using 1×1 convolution. Based on the preset convolutional coding module, the features of the speaker's speech are obtained.
[0083] After applying the parameterized activation function PReLU, 1×1 convolution, and the activation function sigmoid, the mask information of the target speech coding features is obtained.
[0084] In this example embodiment, the method further includes:
[0085] The preset calculation methods include short-time Fourier transform, covariance matrix multiplication, and generalized eigenvalue decomposition.
[0086] In this example embodiment, the method further includes:
[0087] Based on the generalized eigenvalue decomposition of each separated speech source, the steering vector of each speech source is obtained, and the angle between each steering vector is calculated by the error weighted minimization method to determine whether the non-target speech source is in the same direction as the target speech source.
[0088] In the first-stage processing step S120, the feature vector of the target speech and the coding features of the observed signal can be obtained based on the speaker encoder and the speech encoder. The masks of the target speech source and non-target speech source are obtained through feature fusion and mask estimator. The temporal-domain separated speech is obtained through masking processing and decoding mapping.
[0089] In the first-stage processing step S130, the sound source can be reseparated by masking-based minimum variance distortionless response beamforming to obtain time-domain speech.
[0090] In the second-stage processing step S150, the time-domain speech can be used as an auxiliary input to the target speech separation network to construct new fusion features, and the first-stage processing steps can be repeated until the target speech extraction is completed.
[0091] In the embodiments of this example, the target speech extraction method of this application with joint separation masking optimization minimum variance distortionless response beamforming mainly solves the following problems: First, based on the registered target speaker speech, a speech separation network oriented towards target speaker speech extraction is constructed to obtain the target speech source and its multi-channel separation results, and this is combined with MVDR beamforming to reduce amplitude and phase distortion in target speech extraction; Second, based on the separation results, it is determined whether the target speech source and non-target speech source are in the same direction, and the amplitude ratio mask of the main lobe response constraint is calculated to optimize the MVDR beamformer, thereby achieving effective extraction of the target speaker speech under the condition of competing speech sources in the same direction; Third, a two-stage learning architecture is adopted to jointly train the separation and MVDR beamforming processes, and multi-target loss is used for joint training to achieve high-quality auditory and speech recognition target speaker extraction.
[0092] Example 2:
[0093] In this example embodiment, the target speaker speech extraction method of this application with joint separation masking optimization minimum variance distortionless response beamforming has the following overall technical solution: Figure 2 As shown, the system mainly consists of a microphone array for sound pickup and a two-stage target speech separation and beamforming enhancement. In the first-stage target speech separation and beamforming enhancement, multiple speech sources in multiple channels are separated by combining the target speaker's registered speech and the observed signal, using M target speech separators (where the first separated speech source is the target speech source). Then, the MVDR beamformer is remodeled using a target speech source in-direction amplitude scaling mask, thereby reducing target speech distortion while effectively suppressing in-direction competing speech sources. Similar to Beam-guided TasNet, the separation results of the first-stage reference channel are used as auxiliary inputs to the second-stage separator to guide it to achieve better speech separation. Since a large number of iterations would increase the difficulty of network training and model parameters, this application only adopts a two-stage processing structure and does not perform multiple iterations of the second-stage processing.
[0094] In the one-stage and two-stage processing, the overall structure of the separation network is as follows: Figure 3As shown in the diagram, since the target speech information is incorporated and the separated target speech source is known, there is no speech source substitution problem; therefore, this model is named the Target Speech Separation Network. In the first-stage processing, a speaker encoder and a speech encoder are first used to obtain the feature vector of the target speech and the encoded features of the observed signal. Then, feature fusion and a mask estimator are used to obtain the masking of the target speech source and the non-target speech source. Finally, the temporal separated speech is obtained through masking processing and decoding mapping. In the second-stage processing, the separated speech from the first-stage reference channel is used as an auxiliary input to the target speech separation network to guide it to better achieve speech separation, while the rest of the structure remains unchanged.
[0095] In the first and second stage processing, based on the separated signals of the target speech source and non-target speech sources, the MVDR beamforming process for extracting the target speech is as follows: Figure 4 As shown, firstly, the non-target signal is solved based on the separated target speech. Then, the steering vector of each speech source is obtained through generalized eigenvalue decomposition, and the angle between each steering vector is calculated by the error weighted minimization method to determine whether the non-target speech source is in the same direction as the target speech source. Next, based on the determination of whether the speech sources are in the same direction, the amplitude scaling mask is calculated, and the MVDR beamformer is modeled and optimized by combining the covariance matrix of the non-target speech signal and the steering vector to estimate the filtering weights. Finally, the observed signal is filtered based on the filtering weight vector to extract the target speech signal.
[0096] The specific technical solution is achieved through the following steps:
[0097] S01. Obtain the feature vector of the registered speech. Based on the registered speech x(t), extract cues from the target speaker's speech using a speaker encoder. The speaker encoder structure adopts the SpEx speaker coding model, such as... Figure 5 As shown, the network consists of a bidirectional long short-term memory network, the ReLU activation function, linear layers, and average pooling, and is used to extract feature vectors from the registered speaker's speech. Specifically, the BLSTM has 256 neurons in both the forward and backward propagation, which can combine the contextual information of the speech to extract the target speaker's feature vector; the ReLU activation function is used to improve the network's non-linear mapping ability and accelerate network training; the linear layers are used to map the target speaker's feature vector; and average pooling is used to reduce the dimensionality of the feature vector.
[0098] S02. Acquire multi-channel observation signals. Using a microphone array (uniform linear array or uniform circular array) with M microphones, acquire the array observation signals, specifically the observation signal y from the m-th microphone. m (t) can be expressed by equation (1) as follows:
[0099]
[0100] Where m = 1, 2, ..., M are microphone indices, Q is the number of speakers, q = 1, 2, ..., Q are speaker indices; Γ is the duration of 60 dB reverberation decay (i.e., RT). 60 ), τ=1,2,...,Γ are the reverberation time indices, h q,m (τ) is the transfer function from the q-th speech source to the m-th microphone at time t–τ; s q,orig (t) represents the speech signal emitted by the q-th speech source at time t, with the subscript "q". orig " represents the original spoken speech signal; n m (t) represents the noise signal collected by the m-th microphone at time t. Based on equation (1), the multi-channel observation signal vector at time t can be expressed as: superscript " T " indicates the transpose of a matrix or vector. Represents the space of real numbers;
[0101] S03. Obtain the coding features of the observed speech signal. The speech encoder processes the signal y of the m-th channel. m (t) is processed to obtain the encoded features of the speech in that channel. This patent application does not make any innovative design to the speech encoder; it adopts the Conv-TasNet speech encoder structure, such as... Figure 5 As shown. First, 1×1 convolution, parameterized activation function (PReLU), and normalization layer are used to apply the algorithm to y. m (t) A nonlinear transformation is performed; then, depthwise separable convolution is used instead of standard convolution to reduce the number of parameters; finally, the encoded features are obtained by adding the one-dimensional convolution to the input signal. This encoding structure based on convolution operation is also called 1-D Conv, and is used for subsequent mask estimation.
[0102] S04. Estimate the masking features of the speech source. The feature vectors of the registered speech and the speech coding features are fused and fed into the mask estimator to estimate the speech mask features of each speaker. Given the excellent performance of Conv-TasNet in mask separation, its convolutional model is used to estimate the mask features, with the structure as follows: Figure 5 As shown, firstly, the feature vector and speech coding features are concatenated along the channel dimension to obtain an intermediate representation between the target speaker's speech and the observed signal; then, the concatenated features are normalized and subjected to a 1×1 convolution for nonlinear transformation, and then fed into a large number of... Figure 5The convolutional coding module (1-D Conv) shown implements a function similar to a deep temporal convolutional network (TCN) to acquire speaker speech features. Finally, after PReLU, 1×1 convolution, and the sigmoid activation function, the mask information of the speech coding features is obtained. Note that, unlike Conv-TasNet, the feature vector of the target speech is incorporated here, thereby prioritizing the target speech mask and achieving the purpose of distinguishing target speech from non-target speech.
[0103] S05. Obtain the separated signals. Multiply the speech coding features by the estimated target speech mask information to obtain the separation features of each speech source. After mapping by the speech decoder, the separated speech signals of each speech source can be obtained. j,q,m (t), where the subscripts j = 1, 2 indicate the number of processing stages. Note that due to the inclusion of the target speech feature vector in the mask estimator, the target speech separation result is preferentially ranked as the first separation result, i.e., s. j,1,m (t) represents the separated target speech. The speech decoder is similar to the speech encoder, such as... Figure 5 As shown.
[0104] S06. Estimate the speech source steering vector. Repeat the target speech extraction processing of S01 to S05 for each channel of the microphone array to obtain the multi-channel separated signals {s} of Q speech sources. j,1 (t), ..., s j,q (t), ..., s j,Q (t)}. Separation results of the q-th speech source according to equation (2) Perform STFT to obtain frequency domain results
[0105] S j,q (k,l)=STFT{s j,q (t)} (2)
[0106] Where STFT{} denotes the short-time Fourier transform, Represents the space of complex numbers.
[0107] Calculate S according to formula (3) j,q The covariance matrix of (k,l)
[0108]
[0109] Where B is the number of snapshots, used to make R j,q (k,l) is a positive definite matrix.
[0110] Applying GEVD to equation (3) using equation (4), we obtain the result from the largest eigenvalue λ.max Corresponding feature vector Obtain the guide vector
[0111]
[0112] Where GEVD{} represents the generalized eigenvalue decomposition operation. By repeating the steps of equations (2) to (4) for the separated signals of Q speech sources, the steering vector of each speech source can be obtained.
[0113] S07. Determine whether the target speech and non-target speech have the same direction. The pre-estimated target speech source steering vector is a1(k,l). The angle between the q-th speech source steering vector and the target speech source steering vector can be calculated using equation (5) as follows:
[0114]
[0115] Here, arccos{} represents the arccosine operation.
[0116] Due to the steering vector angle θ of each frequency band 1,q The inconsistency between (k,l) may lead to discriminative inconsistencies, meaning there is an error between the angle estimated by equation (5) and the actual angle. This can cause inconsistencies in the MVDR beamforming filtering models of different frequency bands in subsequent steps, resulting in distortion of the target speech. To address this issue, this patent application employs a modeling method that minimizes the weighted error of the multi-band steering vector angle. First, let θ be the actual angle between the target source in the current frame and the steering vector of the q-th sound source. 1,q,real (l), then the included angle θ of each frequency band 1,q The error of (k,l) can be modeled as follows:
[0117] θ 1,q,error (k,l)=θ 1,q (k,l)-θ 1,q,real (l) (6)
[0118] Assume θ 1,q,error If (k,l) follows a zero-mean Gaussian distribution, then we can apply θ 1,q (k,l) is used for weighted error minimization modeling. [6] Assume θ 1,q,error There are two reasons why (k,l) follows a zero-mean Gaussian distribution: First, the zero-mean Gaussian distribution assumption has general significance for cases with both positive and negative errors; second, the Gaussian distribution can assign smaller weights to larger errors, thereby reducing the impact of larger errors on the true angle θ. 1,q,real (l) The impact in the estimation. Therefore, θ can be... 1,q,error The probability distribution f of (k,l) 1,q,error (k,l) is modeled as follows:
[0119]
[0120] in, Let K be the variance of the current frame error, K be the number of frequency bands, and k be the index of the frequency band.
[0121] Based on equation (7), the error weighting can be defined as follows:
[0122]
[0123] Based on equations (6) and (9), the optimization function that minimizes the weighted error can be modeled as follows:
[0124]
[0125] Where, θ 1,q,est (l) is θ 1,q,real (l) estimate.
[0126] Based on the estimation results of equation (10), the co-directionality of the target speech source and the non-target speech source can be determined according to equation (11):
[0127]
[0128] Where, θ threshold To determine the threshold, a value typically ranges from 0.1° to 3°, depending on the severity of the requirement. q (k,l) represents the discrimination result. A value of 0 indicates that the target speech source and the non-target speech source are not in the same direction, while a value of 1 indicates that they are in the same direction.
[0129] S08. Calculate the MVDR beamformer weights. The MVDR beamformer is used after the separation network to reduce the nonlinear distortion of the target speech and better recover the phase information. Combining the separation results of Equation (2), the steering vector estimated by Equation (4), and the decision results of Equation (11), the MVDR beamformer is modeled as follows:
[0130]
[0131] in, and Let be the beamformer weight vector and the estimated non-target signal covariance matrix, respectively, in the j-th stage of processing. R is the amplitude scaling mask constraint factor. j,non (k,l) and It can be estimated using equations (13) and (14) respectively.
[0132]
[0133] In equation (13), The STFT result of the observed signal. In equation (14), p = {1, q|υ q (k,l)=1} is the set of indices for speech sources in the same direction, and p∈p is the index of the source in the same direction.
[0134] From equation (14), it can be seen that when p = {1}, i.e. υ q (k,l)=0 indicates that there are no speech sources in the same direction. Equation (12) then becomes the classic MVDR beamformer:
[0135]
[0136] When p = {1, q, ...}, i.e. υ q (k,l)=1 indicates the existence of a speech source in the same direction. Attenuation constraints are formed in the direction of the main lobe, that is, the distortion-free constraints of the main lobe are masked, so as to suppress competing speech sources in the same direction while ensuring the recovery of phase information.
[0137] Using the Lagrange multiplier method, the solution to equation (12) is:
[0138]
[0139] S09. Obtain the enhanced results of amplitude scaling mask-optimized MVDR beamforming. The weights W in equation (11) are... j (k,l) is used to observe the signal Y(k,l), and by performing the inverse STFT (ISTFT), the time-domain enhanced target speech s can be obtained. en (t), as shown in equations (17) and (18):
[0140]
[0141] s j,en (t)=ISTFT{S j,en (k, l)} (18)
[0142] Among them, S j,en (k,l) represents the frequency domain enhancement signal of the target speech, and ISTFT{} denotes the inverse STFT operation.
[0143] S01 to S09 complete the first stage of processing; the next step is the second stage of processing.
[0144] S10, Second-stage target speech separation. In the second-stage target speech separation process, the registered speaker's speech, the reference channel separation result, and the observation signal are sent together to... Figure 2 In the target speech separation network shown, where s1,1 (t)...s 1,Q (t) represent the first to Qth speech sources separated in the first-stage reference channel. Compared with the structure of the first-stage target speech separation network, only the input information has changed, and the rest of the structure remains the same. Therefore, the second-stage target speech separation steps will not be described again.
[0145] S11. Two-stage MVDR beamforming speech enhancement. The enhancement steps for the two-stage MVDR beamforming speech enhancement based on amplitude scaling mask optimization are the same as those in S06 to S09, and will not be described again here.
[0146] S12. Joint training of the first and second stage processing. In order to achieve good results in both auditory perception and speech recognition of the final extracted target speech, a multi-objective loss function is adopted in the training, as shown in equations (19) to (22):
[0147] L=η1L SNR +η2L MSE +(1-η1-η2)L SI-SDR (19)
[0148]
[0149] Where, the number of processing stages J = 2; SINR{} represents the calculation of signal-to-interference-plus-noise ratio (SINR), α SINR =0.65 and β SINR =1-α SINR α is the adjustment parameter for SINR loss during the separation and beamforming stages; MSE{} represents the calculation of the mean square error (MSE), and α MSE =0.5 and β MSE =1-α MSE α represents the adjustment parameter for MSE loss during the separation and beamforming stages; SISDR{} denotes the calculation of the scale-invariant signal-to-distortion ratio (SI-SDR), and α SI-SDR =0.7 and β SI-SDR =1-α SI-SDRη1 = 0.25 and η2 = 0.4 are the adjustment parameters for SI-SDR loss in the separation and beamforming stages; η1 = 0.25 and η2 = 0.4 are the smoothing results between different loss functions. Among them, SINR can be used to improve the performance of network separation in denoising and eliminating competing sources, which is beneficial to improving auditory quality; MSE can reduce the spectral distortion of separated and enhanced speech, which is beneficial to speech recognition tasks; SI-SDR can reduce the distortion of target speech, which is beneficial to improving auditory quality and speech recognition performance.
[0150] In this example embodiment, the target speech extraction and training method of the present application, which uses a joint separation mask to optimize minimum variance distortionless response beamforming, is as follows: the feature vector of the registered target speech is fused with the coded features of the observed signal to achieve speech separation that can identify the target speech source; the steering vector is obtained through GEVD based on the speech separation result to determine whether the target source and the competing source are in the same direction, and the amplitude ratio mask between the in-direction sources is used to optimize the distortionless constraint of the MVDR main lobe, so as to suppress the in-direction competing speech sources while ensuring the effective recovery of phase information; a two-stage reinforcement training is adopted, and a multi-objective training loss function is constructed using SNR, MSE, and SI-SDR, which can effectively improve the auditory quality and speech recognition performance of the extracted target speech.
[0151] In the embodiments of this example, compared with the prior art, the beneficial effects of the patented technology of this application are as follows: First, combining target source extraction with speech separation and beamforming can effectively reduce amplitude and phase distortion in target source extraction; Second, using the amplitude ratio mask of the same-direction speech source to optimize the MVDR beamformer effectively solves the difficulty that the MVDR beamformer cannot effectively eliminate competing speech sources in the same-direction speech source scenario.
[0152] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions involved are not necessarily required by this application.
[0153] Reference Figure 6 This diagram illustrates a structural block diagram of a target speech extraction device based on masking beamforming, comprising a microphone array pickup module 210, a first-stage processing module 220, and a second-stage processing module 230, wherein:
[0154] Microphone array pickup module 210 is used to acquire multi-channel observation signals through a microphone array;
[0155] The first-stage processing module 220 is used to extract the feature vector of the target speech by the speaker encoder through the registered speaker speech information, extract the coding features of each observed signal by the speech encoder, obtain the mask information of all speech sources by fusing the feature vector and the coding features and by the mask estimator, wherein the all speech sources include target speech sources and non-target speech sources, obtain pre-separated temporal speech information through masking processing and decoding mapping; and determine whether the speech sources are in the same direction by using the pre-separated temporal speech information, and calculate the weight of the minimum variance distortionless corresponding beamforming according to the determination result and the mask information to achieve speech separation.
[0156] The second-stage processing module 230 is used to use the time-domain speech as an auxiliary input to the target speech separation network and extract feature vectors, combine the feature vectors of the registered speaker's speech and the feature vectors of the separated target speech to construct new fusion features, estimate mask information more accurately through a mask estimator and decoding mapping, and repeat the first-stage beamforming speech separation process based on the pre-separated mask information until the final target speech extraction is completed.
[0157] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0158] Optionally, this application also provides an electronic device, including: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the various processes of the above method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0159] This application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0160] Figure 7 This is a block diagram illustrating an electronic device 800. For example, the electronic device 800 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0161] Reference Figure 7The electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0162] Processing component 802 typically controls the overall operation of electronic device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.
[0163] Memory 804 is configured to store various types of data to support the operation of device 800. Examples of such data include instructions for any application or method operating on electronic device 800, contact data, phonebook data, messages, images, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0164] Power supply component 806 provides power to various components of electronic device 800. Power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 800.
[0165] Multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0166] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when electronic device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.
[0167] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0168] Sensor assembly 814 includes one or more sensors for providing state assessments of various aspects of electronic device 800. For example, sensor assembly 814 may detect the on / off state of device 800, the relative positioning of components such as the display and keypad of electronic device 800, changes in position of electronic device 800 or a component of electronic device 800, the presence or absence of user contact with electronic device 800, orientation or acceleration / deceleration of electronic device 800, and temperature changes of electronic device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0169] Communication component 816 is configured to facilitate wired or wireless communication between electronic device 800 and other devices. Electronic device 800 can access wireless networks based on communication standards, such as WiFi, carrier networks (such as 2G, 3G, 4G, or 5G), or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast operation information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0170] Short-range communication. For example, NFC modules can be based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0171] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0172] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of an electronic device 800 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0173] Figure 8 This is a block diagram illustrating a computer-readable storage medium 1900. For example, the computer-readable storage medium 1900 can be provided as a server.
[0174] Reference Figure 8 The computer-readable storage medium 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by memory 1932 for storing instructions executable by the processing component 1922, such as an application program. The application program stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.
[0175] The computer-readable storage medium 1900 may also include a power supply component 1926 configured to perform power management of the computer-readable storage medium 1900, a wired or wireless network interface 1950 configured to connect the computer-readable storage medium 1900 to a network, and an input / output (I / O) interface 1958. The computer-readable storage medium 1900 can operate on an operating system stored in memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.
[0176] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0177] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0178] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
[0179] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0180] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0181] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0182] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0183] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0184] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0185] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A target speech extraction method based on masking beamforming, characterized in that, The method includes: The microphone array pickup step involves acquiring multi-channel observation signals through a microphone array. The first stage of speech pre-separation involves extracting the feature vector of the target speech by the speaker encoder through the registered speaker speech information, extracting the coding features of each observed signal by the speech encoder, fusing the feature vector and the coding features and obtaining the mask information of all speech sources by the mask estimator, wherein the all speech sources include target speech sources and non-target speech sources, and obtaining the pre-separated temporal speech information through masking processing and decoding mapping. The first-stage beamforming target speech separation step uses pre-separated temporal speech information to determine whether the speech sources are in the same direction, and calculates the minimum variance distortion-free corresponding beamforming weights based on the judgment result and mask information to achieve speech separation; The second-stage speech pre-separation step uses the time-domain speech as an auxiliary input to the target speech separation network and extracts feature vectors. It then combines the feature vectors of the registered speaker's speech and the feature vectors of the separated target speech to construct new fusion features. The mask information is estimated more accurately through a mask estimator and a decoding mapping. The second-stage beamforming speech separation step, based on the mask information of the second-stage pre-separation, repeats the first-stage beamforming speech separation process until the final target speech extraction is completed.
2. The method as described in claim 1, characterized in that, The first-stage processing step of the method further includes: The multi-channel observation signals from the microphone are processed based on a preset speech encoder to obtain the coding features of the observed speech signals. Based on a preset mask estimator, the mask information of all speech sources is estimated by taking the feature vector of the target speech and the coding features of the observed signal as input. The encoded features of the observed signal are multiplied with the mask information of all speech sources to obtain the separation features of each speech source, and then mapped by the speech decoder to obtain the separated speech of each speech source. Based on the speech separation of each speech source and the preset generalized feature value decomposition method, the steering vector of each speech source is calculated respectively; Based on the set discrimination threshold, it is determined whether the target speech source and the non-target speech source are in the same direction; Based on the discrimination results, an amplitude scaling mask is obtained to optimize the enhancement results of minimum variance distortionless response beamforming.
3. The method as described in claim 2, characterized in that, The method further includes: Based on the registered speech, the feature vector of the registered speech is obtained through a bidirectional long short-term memory network, activation function, linear layer, and average pooling module based on the preset speaker encoder. The preset microphone array is either a uniform linear array or a uniform circular array.
4. The method as described in claim 3, characterized in that, The method further includes: A 1×1 convolution, parameterized activation function, and normalization layer based on a fully convolutional temporal audio separation network are used to perform nonlinear transformations on the observed signals of each microphone channel. Depth-separable convolution is used to replace standard convolution to reduce the number of parameters; The coded features of the observed speech signal are obtained by adding the signal to the input signal after one-dimensional convolution.
5. The method as described in claim 4, characterized in that, The method further includes: The feature vector and the encoded features of the observed speech signal are concatenated along the channel dimension to obtain an intermediate representation between the target speaker's speech and the observed signal; The spliced features are normalized and nonlinear transformation is performed using 1×1 convolution. Based on the preset convolutional coding module, the features of the speaker's speech are obtained. After parameterized activation function, 1×1 convolution and sigmoid activation function, mask information of target speech coding features is obtained.
6. The method as described in claim 5, characterized in that, The method further includes: The preset calculation methods include short-time Fourier transform, covariance matrix multiplication, and generalized eigenvalue decomposition.
7. The method as described in claim 2, characterized in that, The method further includes: Based on the separated speech sources of each channel, the steering vector of each speech source is obtained through generalized eigenvalue decomposition. The angle between each steering vector is calculated by the error weighted minimization method to determine whether the non-target speech source is in the same direction as the target speech source.
8. A target speech extraction device based on masking beamforming, characterized in that, The device includes: The microphone array pickup module is used to acquire multi-channel observation signals through a microphone array; The first-stage processing module is used to extract the feature vector of the target speech by the speaker encoder through the registered speaker speech information, extract the coding features of each observed signal by the speech encoder, obtain the mask information of all speech sources by fusing the feature vector and the coding features and the mask estimator, the total speech sources including target speech sources and non-target speech sources, obtain pre-separated temporal speech information through masking processing and decoding mapping; and determine whether the speech sources are in the same direction by using the pre-separated temporal speech information, and calculate the weight of the minimum variance distortionless corresponding beamforming according to the judgment result and the mask information to achieve speech separation. The two-stage processing module is used to use the time-domain speech as an auxiliary input to the target speech separation network and extract feature vectors. It then combines the feature vectors of the registered speaker's speech and the feature vectors of the separated target speech to construct new fused features. The module estimates the mask information more accurately through a mask estimator and a decoding map. Based on the pre-separated mask information, the module repeats the beamforming speech separation process of the first stage until the final target speech extraction is completed.
9. An electronic device, characterized in that, include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Target voice separation method and system based on cross-modal loss
CN118016093A
Target speaker voice extraction method and device
CN119007728A