Target voice extraction method and device based on masking beam forming

By introducing masking technology into the beamformer, mask information is obtained by fusion of the registered speaker's voice and the observed signal, and beamforming weight is calculated, the problem of suppressing the same-directional interference sound source is solved, and the quality of target speech extraction and speech recognition performance are improved.

CN119943087AActive Publication Date: 2025-05-06CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411981356.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

At this stage, the beamformer cannot effectively suppress the interfering sound source when the same-direction interferes with the target sound source, resulting in poor voice separation effect.

Method used

The mask-based minimum variance without distortion response beamforming method is adopted, and the speech separation is achieved by fusing the characteristic vector of the speaker's voice with the coded features of the observed signal, mask information of the voice source is obtained, and the weight of beam formation is calculated based on the mask information to achieve speech separation.

Benefits of technology

It effectively reduces the amplitude and phase distortion in target speech extraction, and optimizes the MVDR beamformer in the same-way competitive voice source scenario, significantly improving the extraction quality and speech recognition performance of target speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943087A_ABST
    Figure CN119943087A_ABST
Patent Text Reader

Abstract

The invention provides a target voice extraction method and device based on masking beam forming, and the method comprises the steps: carrying out the pickup of a microphone array, and collecting a multi-channel observation signal through the microphone array; in the first stage, feature vectors of target voice and coding features of observation signals are obtained based on a speaker encoder and a voice encoder, masks of a target voice source and a non-target voice source are obtained through feature fusion and a mask estimator, and time domain separation voice is obtained through masking processing and decoding mapping; re-separation of a sound source is realized through masking-based minimum variance undistorted response beam forming, and time domain voice is obtained; and in the second stage, the time domain voice in the first stage is used as auxiliary input of the target voice separation network, a new fusion feature is constructed, and the processing step in the first stage is repeated until target voice extraction is completed. Amplitude and phase distortion in target source extraction can be effectively reduced, and the difficulty that an MVDR beam former cannot effectively eliminate a competitive voice source in a same-direction voice source scene is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of acoustic signal processing, and in particular to a method and device for extracting target speech based on masking beamforming. Background Art

[0002] Target speech extraction (TSE) refers to extracting the speech of the target speaker of interest in a multi-speaker scenario, thereby providing strong front-end technical support for subsequent high-quality voice communication, speech recognition, human-computer voice interaction and other tasks. At present, target speech extraction methods can be mainly divided into two categories: masking-based methods and beamforming-based methods.

[0003] The target speaker speech extraction method based on masking mainly learns the mask information based on the registered speaker speech and the observation signal through the neural network constructed by the encoding and decoding modules to obtain the target speaker speech, such as speaker extraction (SpEx) network, targeted voice separation filter (VoiceFilter), etc. Although a large number of studies have shown that such methods are very effective in the task of extracting target speaker speech information, the nonlinear distortion of the network and the effective recovery of phase information have always been a difficult problem to solve. Therefore, many studies combine speech separation with beamforming to improve the amplitude and phase distortion caused by the separation network, such as beamformingbased time-domain audio separation network (Beam-TasNet), beamforming guided time-domain audio separation network (Beam-guided TasNet), etc.

[0004] The beamforming-based method mainly uses the spatial filtering characteristics of the beamformer to suppress the interfering speaker's speech and reduce speech distortion. It is often implemented through an adaptive beamformer, such as the Linearly Constrained Minimum Variance (LCMV) beamformer, the Minimum Variance Distortionless Response (MVDR) beamformer, and the generalized sidelobe canceller (GSC). Among these adaptive beamformers, the MVDR beamformer has been widely studied and applied because it minimizes the filtering results of the covariance matrix of non-target signals and constrains the target direction to be distortion-free. Its principle is simple and can be flexibly optimized. However, when the interfering speaker's speech and the target speaker's speech come from the same direction, the main lobe of the beamformer will perform the same filtering processing on the interfering sound source and the target sound source, resulting in the inability to effectively suppress the interfering sound source. Due to this characteristic of the beamformer, the suppression of co-directional interfering sound sources has become an important issue that needs to be broken through in the current beamforming design.

[0005] Among these adaptive beamformers, the MVDR beamformer has been widely studied and applied because it minimizes the filtering results of the covariance matrix of non-target signals and constrains the target direction to be distortion-free. Its principle is simple and can be flexibly optimized. However, when the interfering speaker's voice and the target speaker's voice come from the same direction, the main lobe of the beamformer will perform the same filtering processing on the interfering sound source and the target sound source, resulting in the inability to effectively suppress the interfering sound source. Due to this characteristic of the beamformer, the suppression of co-directional interfering sound sources has become an important issue that needs to be broken through in the current beamforming design. Summary of the invention

[0006] The present application shows a method and device for extracting target speech based on masked beamforming.

[0007] In a first aspect, the present application shows a target speech extraction method based on masked beamforming, the method comprising:

[0008] A microphone array sound pickup step, collecting multi-channel observation signals through the microphone array;

[0009] In the first stage, the speech pre-separation step is to extract the feature vector of the target speech by the speaker encoder through the registered speaker speech information, extract the coding features of each observation signal by the speech encoder, and obtain the mask information of all speech sources by fusing the feature vector and the coding features and using the mask estimator, wherein the all speech sources include the target speech source and the non-target speech source, and obtain the pre-separated time domain speech information through masking processing and decoding mapping;

[0010] In the first stage, the target speech separation step of beamforming is to determine whether the speech source is in the same direction through the pre-separated time domain speech information, and calculate the weight of the minimum variance distortion-free corresponding beamforming based on the judgment result and mask information to achieve speech separation;

[0011] In the second stage of speech pre-separation, the time domain speech is used as an auxiliary input of the target speech separation network and a feature vector is extracted, a new fusion feature is constructed by combining the feature vector of the registered speaker's speech and the feature vector of the separated target speech, and mask information is estimated more accurately through a mask estimator and a decoding map;

[0012] The second-stage beamforming speech separation step repeats the first-stage beamforming speech separation process based on the mask information of the second-stage pre-separation until the final target speech extraction is completed.

[0013] In an exemplary embodiment of the present application, a stage of processing steps of the method further includes:

[0014] Processing the microphone multi-channel observation signal based on a preset speech encoder to obtain coding features of the observation speech signal;

[0015] Based on a preset mask estimator, taking the feature vector of the target speech and the coding feature of the observation signal as input, estimating the mask information of all speech sources;

[0016] Multiplying the coded features of the observation signal with the mask information of all speech sources to obtain the separation features of each speech source, and mapping through a speech decoder to obtain the separated speech of each speech source;

[0017] Based on the separated speech of each speech source and the preset generalized eigenvalue decomposition method, the steering vector of each speech source is calculated respectively;

[0018] Based on the set discrimination threshold, it is judged whether the target speech source and the non-target speech source are in the same direction;

[0019] According to the discrimination results, an amplitude ratio mask is obtained to optimize the enhancement result of minimum variance distortionless response beamforming.

[0020] In an exemplary embodiment of the present application, the method further includes:

[0021] According to the registered speech, the feature vector of the registered speech is obtained based on the bidirectional long short-term memory (BLSTM) network, activation function, linear layer (linear), and mean pooling module of the preset speaker encoder;

[0022] The preset microphone array is a uniform linear array or a uniform circular array.

[0023] In an exemplary embodiment of the present application, the method further includes:

[0024] Based on the 1×1 convolution, parameterized activation function, and normalization layer (Norm) of the fully convolutional time-domain audio separation network, the observation signals of each microphone channel are nonlinearly transformed;

[0025] Use depth-wise separable convolution instead of standard convolution to reduce the number of parameters;

[0026] The encoding features of the observed speech signal are obtained by adding it to the input signal after one-dimensional convolution.

[0027] In an exemplary embodiment of the present application, the method further includes:

[0028] Concatenating the feature vector with the encoded features of the observed speech signal according to the channel dimension to obtain an intermediate representation between the target speaker's speech and the observed signal;

[0029] The concatenated features are normalized and nonlinearly transformed using 1×1 convolution, and the features of the speaker's speech are obtained based on a preset convolutional coding module;

[0030] After parameterizing the activation function, 1×1 convolution and activation function sigmoid, the mask information of the target speech coding features is obtained.

[0031] In an exemplary embodiment of the present application, the method further includes:

[0032] The preset calculation method includes short-time Fourier transform, covariance matrix multiplication, and generalized eigenvalue decomposition operations.

[0033] In an exemplary embodiment of the present application, the method further includes:

[0034] According to the separated speech sources of each channel, the steering vector of each speech source is obtained through generalized eigenvalue decomposition, and the angle between the steering vectors is calculated by the error weighted minimization method to determine whether the non-target speech source is in the same direction as the target speech source.

[0035] In a second aspect, the present application shows a target speech extraction device based on masked beamforming, the device comprising:

[0036] A microphone array pickup module is used to collect multi-channel observation signals through a microphone array;

[0037] A first-stage processing module is used to extract a feature vector of a target speech through a speaker encoder through the registered speaker speech information, extract coding features of each observation signal through a speech encoder, obtain mask information of all speech sources through a mask estimator by fusing the feature vector and the coding features, the all speech sources include target speech sources and non-target speech sources, obtain pre-separated time-domain speech information through masking processing and decoding mapping; and, determine whether the speech sources are in the same direction through the pre-separated time-domain speech information, and calculate the weight of minimum variance distortion-free corresponding beamforming according to the determination result and the mask information to achieve speech separation;

[0038] The second-stage processing module is used to use the time-domain speech as an auxiliary input of the target speech separation network and extract feature vectors, construct new fusion features by combining the feature vectors of the registered speaker's speech and the feature vectors of the separated target speech, and more accurately estimate the mask information through the mask estimator and the decoding map; and, based on the pre-separated mask information, repeatedly perform a stage of beamforming speech separation process until the final target speech extraction is completed.

[0039] In a third aspect, the present application shows an electronic device, which includes: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the method described in any of the above aspects.

[0040] In a fourth aspect, the present application illustrates a non-temporary computer-readable storage medium, which, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute a method as described in any of the above aspects.

[0041] In a fifth aspect, the present application illustrates a computer program product. When instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to execute the method described in any of the above aspects.

[0042] The target speech extraction method of the present application comprises: a microphone array sound pickup step, which collects multi-channel observation signals; a first-stage processing step, which obtains the feature vector of the target speech and the coding features of the observation signal based on a speaker encoder and a speech encoder, obtains the masks of the target speech source and the non-target speech source through feature fusion and a mask estimator, obtains the time-domain separated speech through masking processing and decoding mapping, and realizes the re-separation of the sound source through masking-based minimum variance distortion-free response beamforming to obtain the time-domain speech; and a second-stage processing step, which uses the time-domain speech as the auxiliary input of the target speech separation network, constructs a new fusion feature, and repeatedly executes the first-stage processing step until the target speech extraction is completed.

[0043] This application integrates the feature vector of the registered target speech with the coding feature of the observed signal to achieve speech separation that can identify the target speech source; combined with the speech separation result, the steering vector is obtained through GEVD (Generalized Eigenvalue Decomposition) to determine whether the target source and the competing source are in the same direction, and the amplitude ratio mask between the same-direction sources is used to optimize the MVDR main lobe distortion-free constraint, while ensuring the effective recovery of phase information and suppressing the same-direction competing speech source; two-stage intensive training is adopted, and SNR, MSE, and SI-SDR are used to construct a multi-objective training loss function, which can effectively improve the auditory quality and speech recognition performance of the extracted target speech. This application can combine target source extraction with speech separation and beamforming, which can effectively reduce the amplitude and phase distortion in target source extraction; the amplitude ratio mask of the same-direction speech source is used to optimize the MVDR beamformer, which effectively solves the difficulty that the MVDR beamformer cannot effectively eliminate the competing speech source in the same-direction speech source scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 This is a flowchart of the steps of a target speech extraction method based on masked beamforming of the present application.

[0045] Figure 2 This is a two-stage target speaker speech extraction framework diagram of a target speech extraction method based on masked beamforming in the present application.

[0046] Figure 3 It is a structural diagram of a target speech separation network of a target speech extraction method based on masked beamforming in the present application.

[0047] Figure 4 It is a flowchart of extracting target speech by MVDR beamforming of a target speech extraction method based on masked beamforming of the present application.

[0048] Figure 5This is a schematic diagram of the detailed structure of each module of the target speech separation network of a target speech extraction method based on masked beamforming in the present application.

[0049] Figure 6 This is a structural block diagram of a target speech extraction device based on masked beamforming in the present application.

[0050] Figure 7 It is a block diagram of an electronic device of the present application.

[0051] Figure 8 It is a block diagram of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0052] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0053] Reference Figure 1 , shows a flowchart of the steps of a target speech extraction method based on masked beamforming of the present application, which can be applied to electronic devices, wherein the method specifically may include the following steps:

[0054] Microphone array sound pickup step S110, collecting multi-channel observation signals through the microphone array;

[0055] In the first stage, the speech pre-separation step S120 is to extract the feature vector of the target speech by the speaker encoder through the registered speaker speech information, extract the coding features of each observation signal by the speech encoder, obtain the mask information of all speech sources by fusing the feature vector and the coding features and by the mask estimator, wherein the all speech sources include the target speech source and the non-target speech source, and obtain the pre-separated time domain speech information through masking processing and decoding mapping;

[0056] The first stage of beamforming target speech separation step S130 determines whether the speech source is in the same direction through the pre-separated time domain speech information, and calculates the weight of the minimum variance distortion-free corresponding beamforming according to the determination result and the mask information to achieve speech separation;

[0057] The second stage speech pre-separation step S140 uses the time domain speech as an auxiliary input of the target speech separation network and extracts a feature vector, combines the feature vector of the registered speaker speech and the feature vector of the separated target speech to construct a new fusion feature, and estimates the mask information more accurately through a mask estimator and a decoding map;

[0058] The second stage beamforming speech separation step S150 repeats the first stage beamforming speech separation process based on the mask information of the second stage pre-separation until the final target speech is extracted.

[0059] The target speech extraction method based on masked beamforming of the present application comprises: a microphone array pickup step, in which a multi-channel observation signal is collected through a microphone array; a first-stage processing step, in which a feature vector of a target speech and a coding feature of an observation signal are obtained based on a speaker encoder and a speech encoder, masks of a target speech source and a non-target speech source are obtained through feature fusion and a mask estimator, time-domain separated speech is obtained through masking processing and decoding mapping, and sound source re-separation is achieved through masked minimum variance distortion-free response beamforming to obtain time-domain speech; and a second-stage processing step, in which the time-domain speech is used as an auxiliary input of a target speech separation network, a new fusion feature is constructed, and the first-stage processing step is repeatedly executed until the target speech extraction is completed. This application integrates the feature vector of the registered target speech with the coding feature of the observed signal to achieve speech separation that can identify the target speech source; combines the speech separation result with GEVD to obtain the steering vector, determines whether the target source and the competing source are in the same direction, and uses the amplitude ratio mask between the same-direction sources to optimize the MVDR main lobe distortion-free constraint, suppressing the same-direction competing speech source while ensuring the effective recovery of phase information; adopts two-stage intensive training, and uses SNR, MSE, and SI-SDR to construct a multi-objective training loss function, which can effectively improve the auditory quality and speech recognition performance of the extracted target speech. This application can combine target source extraction with speech separation and beamforming, which can effectively reduce the amplitude and phase distortion in target source extraction; the amplitude ratio mask of the same-direction speech source is used to optimize the MVDR beamformer, which effectively solves the difficulty that the MVDR beamformer cannot effectively eliminate the competing speech source in the same-direction speech source scenario.

[0060] Embodiment 1:

[0061] In the embodiment of this example, speech separation refers to separating multiple speech source signals in a multi-speech source scenario through certain processing techniques, such as beamforming, independent vector analysis, neural network mapping, etc.

[0062] Mask generally refers to the ratio of the target signal to the noisy signal in a certain form. It is usually estimated by training a neural network model in a supervised manner with the help of a certain form of ideal mask, such as an ideal binary mask, an ideal amplitude ratio mask, an ideal complex ratio mask, etc. It can be used to achieve speech enhancement, speech separation and target speech extraction.

[0063] Minimum variance distortionless response beamforming is a spatial filtering technology that achieves adaptive interference source suppression and noise reduction by forming a main lobe of distortionless filtering in the desired direction, forming side lobes of attenuation filtering in the undesired direction, and adaptively forming a zero point in the direction of the interference sound source.

[0064] Target speech extraction refers to extracting the speech of the speaker of interest from the observation signal collected by the microphone in scenes with noise, reverberation, and multiple sound sources.

[0065] In the microphone array sound pickup step S110, a multi-channel observation signal may be collected by the microphone array.

[0066] In the embodiment of this example, the first-stage processing steps of the method further include:

[0067] Processing the array observation signal based on a preset speech encoder to obtain coding features of the observation speech signal;

[0068] Based on a preset mask estimator, taking the feature vector of the target speech and the coding feature of the observation signal as input, estimating the mask information of all speech sources;

[0069] Multiplying the coded features of the observation signal with the mask information of all speech sources to obtain the separation features of each speech source, and mapping through a speech decoder to obtain the separated speech of each speech source;

[0070] Based on the separated speech of each speech source and the preset generalized eigenvalue decomposition method, the steering vector of each speech source is calculated respectively;

[0071] Based on the set discrimination threshold, it is judged whether the target speech source and the non-target speech source are in the same direction;

[0072] According to the discrimination results, an amplitude ratio mask is obtained to optimize the enhancement result of minimum variance distortionless response beamforming.

[0073] In the embodiment of this example, the method further includes:

[0074] According to the registered speech, the feature vector of the registered speech is obtained based on the bidirectional long short-term memory network, activation function ReLU, linear layer, and average pooling module of the preset speaker encoder;

[0075] The microphone array of the preset number of microphones is a uniform linear array or a uniform circular array.

[0076] In the embodiment of this example, the method further includes:

[0077] Based on the 1×1 convolution, parameterized activation function PReLU, and normalization layer of the fully convolutional time-domain audio separation network Conv-TasNet, the observation signal of the preset microphone is nonlinearly transformed;

[0078] Use depth-wise separable convolution instead of standard convolution to reduce the number of parameters;

[0079] The encoding features of the observed speech signal are obtained by adding it to the input signal after one-dimensional convolution.

[0080] In the embodiment of this example, the method further includes:

[0081] Concatenating the feature vector with the encoded features of the observed speech signal according to the channel dimension to obtain an intermediate representation between the target speaker's speech and the observed signal;

[0082] The concatenated features are normalized and nonlinearly transformed using 1×1 convolution, and the features of the speaker's speech are obtained based on a preset convolutional coding module;

[0083] After parameterizing the activation function PReLU, 1×1 convolution and activation function sigmoid, the mask information of the target speech coding features is obtained.

[0084] In the embodiment of this example, the method further includes:

[0085] The preset calculation method includes short-time Fourier transform, covariance matrix multiplication, and generalized eigenvalue decomposition operations.

[0086] In the embodiment of this example, the method further includes:

[0087] According to the generalized eigenvalue decomposition of each separated speech source, the steering vector of each speech source is obtained, and the angle between each steering vector is calculated by the error weighted minimization method to determine whether the non-target speech source is in the same direction as the target speech source.

[0088] In the first-stage processing step S120, the feature vector of the target speech and the coding features of the observation signal can be obtained based on the speaker encoder and the speech encoder, the masks of the target speech source and the non-target speech source can be obtained through feature fusion and mask estimator, and the time-domain separated speech can be obtained through masking processing and decoding mapping;

[0089] In the first-stage processing step S130, the sound source can be re-separated by masking-based minimum variance distortionless response beamforming to obtain time-domain speech.

[0090] In the second-stage processing step S150, the time-domain speech can be used as an auxiliary input of the target speech separation network to construct a new fusion feature, and the first-stage processing steps are repeatedly performed until the target speech extraction is completed.

[0091] In the embodiment of this example, the target speech extraction method of the joint separation masked optimized minimum variance distortionless response beamforming of the present application mainly solves the following problems: first, according to the registered target speaker speech, a speech separation network for target speaker speech extraction is constructed, and the multi-channel separation results of the target speech source and its speech source are obtained, and combined with the MVDR beamforming to reduce the amplitude and phase distortion in the target speech extraction; second, according to the separation result, it is determined whether the target speech source and the non-target speech source are in the same direction, and the amplitude ratio mask of the main lobe response constraint is calculated, and the MVDR beamformer is optimized to achieve effective extraction of the target speaker speech under the condition of the same-direction competing speech source; third, a two-stage learning architecture is adopted to jointly train the separation and MVDR beamforming processing processes, and multi-target loss joint training is adopted to achieve high-quality hearing and speech recognition target speaker extraction.

[0092] Embodiment 2:

[0093] In the embodiment of this example, the target speaker speech extraction method of the joint separation masking optimization minimum variance distortion-free response beamforming of the present application has an overall technical solution as follows: Figure 2 As shown, it is mainly composed of microphone array pickup and two-stage target speech separation and beamforming enhancement. In the first-stage target speech separation and beamforming enhancement, the target speaker's registered speech and observation signal are combined, and multi-channel multiple speech sources are separated by M target speech separators (the first separated speech source is the target speech source), and then the MVDR beamformer is remodeled by the target speech source same-direction amplitude ratio mask, thereby reducing the target speech distortion while achieving effective suppression of the same-direction competing speech source. Similar to Beam-guided TasNet, the separation result of the first-stage reference channel is used as an auxiliary input of the second-stage separator to guide the separator to better achieve speech separation. Since a large number of iterations will increase the difficulty of network training and model parameters, this application only adopts a two-stage processing structure and does not perform multiple iterations on the two-stage processing.

[0094] In the one-stage and two-stage processing, the overall structure of the separation network is as follows Figure 3As shown in the figure. Since the target speech information is integrated, the separated target speech source is known and there is no speech source replacement problem, so the model is named the target speech separation network. In the first stage processing, the speaker encoder and the speech encoder are first used to obtain the feature vector of the target speech and the encoding features of the observation signal, and then the masking of the target speech source and the non-target speech source is obtained through feature fusion and mask estimator. Finally, the time domain separated speech is obtained through masking processing and decoding mapping. In the second stage processing, the separated speech of the first stage reference channel is used as the auxiliary input of the target speech separation network to guide the target speech separation network to better achieve speech separation, while the rest of the structure remains unchanged.

[0095] In the one-stage and two-stage processing, based on the separated signals of the target speech source and the non-target speech source, the process of MVDR beamforming to extract the target speech is as follows: Figure 4 As shown. First, the non-target signal is solved according to the separated target speech; then, the steering vector of each speech source is obtained through generalized eigenvalue decomposition according to each separated speech source, and the angle between each steering vector is calculated by the error weighted minimization method to determine whether the non-target speech source is in the same direction as the target speech source; then, according to the determination result of whether the speech source is in the same direction, the amplitude ratio mask is calculated, and the MVDR beamformer is modeled and optimized by combining the covariance matrix of the non-target speech signal and the steering vector to estimate the filtering weight; finally, the observed signal is filtered according to the filtering weight vector to extract the target speech signal.

[0096] The specific technical solution is implemented through the following steps:

[0097] S01. Obtain the feature vector of the registered speech. According to the registered speech x(t), the clues of the target speaker's speech are extracted through the speaker encoder. The speaker encoder structure adopts the speaker encoding model of SpEx, such as Figure 5 As shown in the figure, it is composed of bidirectional long short-term memory network, activation function ReLU, linear layer, average pooling and other modules, which are used to extract the feature vector of the registered speaker's voice. Among them, the number of neurons in the forward and backward propagation of BLSTM is 256, which can extract the feature vector of the target speaker in combination with the context information of the voice; the activation function ReLU is used to improve the nonlinear mapping ability of the network and speed up the training of the network; the linear layer is used to map the feature vector of the target speaker; and the average pooling is used to reduce the dimension of the feature vector.

[0098] S02, obtain multi-channel observation signals. The array observation signal is collected through a microphone array (uniform linear array or uniform circular array) with M microphones. The observation signal y of the mth microphone is m (t) can be expressed by formula (1) as follows:

[0099]

[0100] Where m = 1, 2, ..., M is the microphone index, Q is the number of speakers, q = 1, 2, ..., Q is the speaker index; Γ is the time it takes for the reverberation to decay by 60 dB (i.e., RT 60 ), τ=1,2,...,Γ is the reverberation time index, h q,m (τ) is the transfer function from the qth speech source to the mth microphone at time t–τ; s q,orig (t) is the speech signal emitted by the qth speech source at time t, and the subscript “ orig " represents the original voiced speech signal; n m (t) is the noise signal collected by the mth microphone at time t. Based on formula (1), the multi-channel observation signal vector at time t can be expressed as Superscript " T " represents the transpose of a matrix or vector, represents the real number space;

[0101] S03, obtain the coding features of the observed speech signal. The speech encoder encodes the signal y of the mth channel m (t) is processed to obtain the coding features of the channel speech. This patent application does not make innovative designs for the speech encoder, but adopts the Conv-TasNet speech encoder structure, such as Figure 5 First, we use 1×1 convolution, parameterized activation function (PReLU), and normalization layer to transform y m (t) performs nonlinear transformation; then, depthwise separable convolution is used to replace standard convolution to reduce the number of parameters; finally, the encoding feature is obtained by adding it to the input signal after one-dimensional convolution. This encoding structure based on convolution operation is also called 1-D Conv and is used for subsequent mask estimation.

[0102] S04. Estimate the masking features of the speech source. The feature vector of the registered speech is fused with the speech coding features and sent to the mask estimator to estimate the mask features of each speaker's speech. In view of the excellent performance of Conv-TasNet in masking separation, its convolutional model is used to estimate the mask features. The structure is as follows: Figure 5 First, the feature vector and the speech coding feature are concatenated according to the channel dimension to obtain the intermediate representation between the target speaker's speech and the observed signal; then, the concatenated features are normalized and nonlinearly transformed using 1×1 convolution, and then sent to a large number of Figure 5The convolutional coding module (1-D Conv) shown in the figure implements functions similar to those of a deep temporal convolutional network (TCN) to obtain the characteristics of the speaker's speech; finally, after PReLU, 1×1 convolution and activation function sigmoid, the mask information of the speech coding characteristics is obtained. Note that unlike Conv-TasNet, the feature vector of the target speech is incorporated here to prioritize the target speech mask and achieve the purpose of distinguishing the target speech from the non-target speech.

[0103] S05. Obtain separation signal. Multiply the speech coding feature with the estimated target speech mask information to obtain the separation feature of each speech source, and then map it through the speech decoder to obtain the separation speech signal of each speech source. j,q,m (t), where j = 1, 2 in the subscript indicates the number of processing stages. Note that due to the addition of the feature vector of the target speech in the mask estimator, the separation result of the target speech is prioritized in the position of the first separation result, i.e., s j,1,m (t) is the separated target speech. The speech decoder is similar to the speech encoder, such as Figure 5 shown.

[0104] S06: Estimate the speech source steering vector. Repeat the target speech extraction process of S01 to S05 for each channel of the microphone array to obtain the multi-channel separation signal {s j,1 (t),…,s j,q (t),…,s j,Q (t)}. The separation result of the qth speech source according to formula (2) is Do STFT and get frequency domain results

[0105] S j,q (k,l)=STFT{s j,q (t)} (2)

[0106] Where STFT{} represents short-time Fourier transform, Represents complex space.

[0107] Calculate S according to formula (3): j,q The covariance matrix of (k,l)

[0108]

[0109] Where B is the number of snapshots, which is used to make R j,q (k,l) is a positive definite matrix.

[0110] According to formula (4), GEVD is performed on formula (3), and the maximum eigenvalue λmax The corresponding eigenvector Get Steering Vector

[0111]

[0112] Wherein, GEVD{} represents the generalized eigenvalue decomposition operation. By repeating the steps of equations (2) to (4) for the separated signals of Q speech sources, the steering vector of each speech source can be obtained.

[0113] S07, determine whether the target speech and the non-target speech have the same direction problem. The estimated target speech source steering vector is preset as a1(k, l), and the angle between the qth speech source steering vector and the target speech source steering vector can be calculated according to formula (5) as follows:

[0114]

[0115] Among them, arccos{} represents the arccosine operation.

[0116] Since the steering vector angle θ of each frequency band 1,q (k, l) are inconsistent, so there may be a problem of inconsistent judgment, that is, there is an error between the angle estimated by formula (5) and the actual angle, which will lead to inconsistency in the MVDR beamforming filter models of different frequency bands in the subsequent steps, causing distortion of the target speech. In order to solve this problem, the patent application adopts a modeling method of weighted minimization of the multi-band steering vector angle error to solve it. First, let the actual angle between the current frame target source and the qth sound source steering vector be θ 1,q,real (l), then the angle θ of each frequency band is 1,q The error of (k,l) can be modeled as follows:

[0117] θ 1,q,error (k,l)=θ 1,q (k,l)-θ 1,q,real (l) (6)

[0118] Assume θ 1,q,error (k,l) obeys zero-mean Gaussian distribution, then θ 1,q (k,l) weighted error minimization modeling [6] Assume θ 1,q,error There are two reasons why (k,l) obeys zero-mean Gaussian distribution: first, for the case of positive and negative errors, the zero-mean Gaussian distribution assumption has general significance; second, Gaussian distribution can give smaller weights to larger errors, thereby reducing the large error in the true angle θ 1,q,real (l) The impact of the estimation. Therefore, θ 1,q,error The distribution probability f of (k,l) 1,q,error (k,l) is modeled as follows:

[0119]

[0120] in, is the variance of the current frame error, K is the number of frequency bands, and k is the index of the frequency band number.

[0121] Based on formula (7), the error weight can be defined as follows:

[0122]

[0123] Based on equations (6) and (9), the optimization function for minimizing the weighted error can be modeled as follows:

[0124]

[0125] Among them, θ 1,q,est (l) is θ 1,q,real (l) Estimates.

[0126] Based on the estimation result of formula (10), whether the target speech source and the non-target speech source have the same direction can be judged according to formula (11):

[0127]

[0128] Among them, θ threshold It is the judgment threshold, which is generally between 0.1° and 3° according to the severity. q (k, l) is the discrimination result. When the value is 0, it means that the target speech source and the non-target speech source are not in the same direction, and when it is 1, it means they are in the same direction.

[0129] S08. Calculate the MVDR beamformer weights. The MVDR beamformer is used after the separation network to reduce the nonlinear distortion of the target speech and better restore the phase information. Combining the separation result of formula (2), the steering vector estimated by formula (4) and the decision result of formula (11), the MVDR beamformer model is as follows:

[0130]

[0131] in, and are the beamformer weight vector and the estimated non-target signal covariance matrix in the j-th stage processing, respectively, is the amplitude ratio mask constraint factor. j,non (k,l) and They can be estimated by equation (13) and equation (14) respectively.

[0132]

[0133] In formula (13), is the STFT result of the observed signal. In formula (14), p = {1, q | q (k,l)=1} is the index set of the same-direction speech sources, and p∈p is the same-direction source index.

[0134] It can be seen from formula (14) that when p = {1}, that is, υ q (k,l)=0, there is no speech source in the same direction. Then equation (12) becomes the classic MVDR beamformer:

[0135]

[0136] When p={1,q,...}, that is, υ q (k,l)=1 There is a speech source in the same direction. An attenuation constraint is formed in the main lobe direction, that is, the distortion-free constraint of the main lobe is masked, so that while ensuring the recovery of phase information, competing speech sources in the same direction are suppressed.

[0137] Using the Lagrange multiplier method, the solution of equation (12) can be obtained as:

[0138]

[0139] S09, obtain the enhancement result of the amplitude ratio mask optimized MVDR beamforming. j (k, l) is used to observe the signal Y(k, l), and perform the inverse STFT (ISTFT) to obtain the target speech s enhanced in the time domain en (t), as shown in equations (17) and (18):

[0140]

[0141] s j,en (t) = ISTFT{S j,en (k, l)} (18)

[0142] Among them, S j,en (k, l) is the frequency domain enhanced signal of the target speech, and ISTFT{} represents the inverse STFT operation.

[0143] S01 to S09 complete the first stage of processing, and the next step is the second stage of processing.

[0144] S10, two-stage target speech separation. In the two-stage target speech separation process, the registered speaker speech, the reference channel separation result, and the observation signal are sent to the Figure 2 In the target speech separation network shown, s1,1 (t)...s 1,Q (t) represent the 1st to Qth speech sources of the first-stage reference channel separation respectively. Compared with the structure of the first-stage target speech separation network, only the input information is changed, and the rest of the structure remains unchanged, so the second-stage target speech separation steps will not be described again.

[0145] S11, two-stage MVDR beamforming speech enhancement The enhancement steps of the two-stage speech enhancement based on amplitude ratio mask optimization MVDR beamforming are the same as S06 to S09, and will not be described again here.

[0146] S12, joint training of the first and second stage processing. In order to make the final extracted target speech have better effects in both hearing and speech recognition, a multi-objective loss function is used in the training, as shown in equations (19) to (22):

[0147] L=η1L SNR +η2L MSE +(1-η1-η2)L SI-SDR (19)

[0148]

[0149] Wherein, the number of processing stages J=2; SINR{} represents the calculation of signal to interference-plus-noise ratio (SINR), α SINR =0.65 and β SINR =1-α SINR is the adjustment parameter of SINR loss in the separation stage and beamforming stage; MSE{} represents the calculation of mean square error (MSE), α MSE = 0.5 and β MSE =1-α MSE is the adjustment parameter of the MSE loss in the separation stage and the beamforming stage; SISDR{} represents the calculation of the scale-invariant signal-to-distortion ratio (SI-SDR), α SI-SDR = 0.7 and β SI-SDR =1-α SI-SDRis the adjustment parameter of SI-SDR loss in the separation stage and beamforming stage; η1=0.25 and η2=0.4 are the smoothing results between different loss functions. Among them, SINR can be used to improve the performance of denoising and removing competing sources during network separation, which is beneficial to improving the hearing quality; MSE can reduce the spectral distortion of separated and enhanced speech, which is beneficial to speech recognition tasks; SI-SDR can reduce the distortion of the target speech, which is beneficial to improving the hearing quality and speech recognition performance.

[0150] In the embodiment of this example, the target speech extraction and training method of the joint separation mask optimized minimum variance distortionless response beamforming of the present application: the feature vector of the registered target speech is integrated with the coding feature of the observed signal to achieve speech separation that can identify the target speech source; the steering vector is obtained through GEVD in combination with the speech separation result to determine whether the target source and the competing source are in the same direction, and the amplitude ratio mask between the co-directional sources is used to optimize the MVDR main lobe distortion-free constraint, while ensuring the effective recovery of phase information, while suppressing the co-directional competing speech source; two-stage intensive training is adopted, and SNR, MSE, and SI-SDR are used to construct a multi-objective training loss function, which can effectively improve the auditory quality and speech recognition performance of the extracted target speech.

[0151] In the embodiment of this example, compared with the prior art, the beneficial effects of the patent technology of this application are as follows: First, the target source extraction is combined with speech separation and beamforming, which can effectively reduce the amplitude and phase distortion in the target source extraction; Second, the amplitude ratio mask of the co-directional speech source is used to optimize the MVDR beamformer, which effectively solves the difficulty that the MVDR beamformer cannot effectively eliminate the competing speech source in the co-directional speech source scenario.

[0152] It should be noted that, for the method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions involved are not necessarily required by the present application.

[0153] Reference Figure 6 , shows a structural block diagram of a target speech extraction device based on masked beamforming of the present application, the device includes a microphone array pickup module 210, a first-stage processing module 220 and a second-stage processing module 230, wherein:

[0154] A microphone array pickup module 210 is used to collect multi-channel observation signals through a microphone array;

[0155] The first-stage processing module 220 is used to extract the feature vector of the target speech by the speaker encoder through the registered speaker speech information, extract the coding features of each observation signal through the speech encoder, obtain the mask information of all speech sources by fusing the feature vector and the coding features and using the mask estimator, wherein the all speech sources include the target speech source and the non-target speech source, obtain the pre-separated time domain speech information through masking processing and decoding mapping; and determine whether the speech sources are in the same direction through the pre-separated time domain speech information, and calculate the weight of the minimum variance distortion-free corresponding beamforming according to the determination result and the mask information to achieve speech separation;

[0156] The second-stage processing module 230 is used to use the time-domain speech as an auxiliary input of the target speech separation network and extract feature vectors, construct new fusion features by combining the feature vectors of the registered speaker's speech and the feature vectors of the separated target speech, and more accurately estimate the mask information through the mask estimator and the decoding mapping; and, based on the pre-separated mask information, repeatedly perform the first-stage beamforming speech separation process until the final target speech extraction is completed.

[0157] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0158] Optionally, an embodiment of the present application further provides an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the various processes of the above-mentioned method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0159] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, each process of the above method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0160] Figure 7 800 is a block diagram of an electronic device 800 shown in the present application. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0161] Reference Figure 7, the electronic device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .

[0162] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0163] The memory 804 is configured to store various types of data to support operations on the device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, images, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0164] The power supply component 806 provides power to the various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 800.

[0165] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and the rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0166] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0167] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: home button, volume button, start button, and lock button.

[0168] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the device 800, the relative positioning of components, such as the display and keypad of the electronic device 800, and the sensor assembly 814 can also detect the position change of the electronic device 800 or a component of the electronic device 800, the presence or absence of contact between the user and the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0169] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast operation information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0170] Short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0171] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0172] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by a processor 820 of an electronic device 800 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0173] Figure 8 19 is a block diagram of a computer-readable storage medium 1900 shown in the present application. For example, the computer-readable storage medium 1900 may be provided as a server.

[0174] Reference Figure 8 , the computer-readable storage medium 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions, such as an application, that can be executed by the processing component 1922. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.

[0175] The computer readable storage medium 1900 may also include a power supply component 1926 configured to perform power management of the computer readable storage medium 1900, a wired or wireless network interface 1950 configured to connect the computer readable storage medium 1900 to a network, and an input / output (I / O) interface 1958. The computer readable storage medium 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™ or the like.

[0176] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0177] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0178] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

[0179] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the present application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0180] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0181] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0182] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0183] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0184] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard drives, ROM, RAM, magnetic disks, or optical disks.

[0185] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A target speech extraction method based on masked beamforming, characterized in that: The method comprises: A microphone array sound pickup step, collecting multi-channel observation signals through the microphone array; In the first stage, the speech pre-separation step is to extract the feature vector of the target speech by the speaker encoder through the registered speaker speech information, extract the coding features of each observation signal by the speech encoder, obtain the mask information of all speech sources by fusing the feature vector and the coding features and using the mask estimator, wherein the all speech sources include the target speech source and the non-target speech source, and obtain the pre-separated time domain speech information through masking processing and decoding mapping; In the first stage, the target speech separation step of beamforming is to determine whether the speech source is in the same direction through the pre-separated time domain speech information, and calculate the weight of the minimum variance distortion-free corresponding beamforming based on the judgment result and mask information to achieve speech separation; In the second stage of speech pre-separation, the time domain speech is used as an auxiliary input of the target speech separation network and a feature vector is extracted, a new fusion feature is constructed by combining the feature vector of the registered speaker's speech and the feature vector of the separated target speech, and mask information is estimated more accurately through a mask estimator and a decoding map; The second-stage beamforming speech separation step repeats the first-stage beamforming speech separation process based on the mask information of the second-stage pre-separation until the final target speech extraction is completed.

2. The method according to claim 1, characterized in that The first-stage processing steps of the method also include: Processing the microphone multi-channel observation signal based on a preset speech encoder to obtain coding features of the observation speech signal; Based on a preset mask estimator, taking the feature vector of the target speech and the coding feature of the observation signal as input, estimating the mask information of all speech sources; Multiplying the coded features of the observation signal with the mask information of all speech sources to obtain the separation features of each speech source, and mapping through a speech decoder to obtain the separated speech of each speech source; Based on the separated speech of each speech source and the preset generalized eigenvalue decomposition method, the steering vector of each speech source is calculated respectively; Based on the set discrimination threshold, it is judged whether the target speech source and the non-target speech source are in the same direction; According to the discrimination results, an amplitude ratio mask is obtained to optimize the enhancement result of minimum variance distortionless response beamforming.

3. The method according to claim 2, characterized in that The method further comprises: According to the registered speech, a feature vector of the registered speech is obtained based on a bidirectional long short-term memory network, an activation function, a linear layer, and an average pooling module of a preset speaker encoder; The preset microphone array is a uniform linear array or a uniform circular array.

4. The method according to claim 3, characterized in that The method further comprises: Based on the 1×1 convolution, parameterized activation function, and normalization layer of the fully convolutional time-domain audio separation network, the observation signals of each microphone channel are nonlinearly transformed; Use depth-wise separable convolution instead of standard convolution to reduce the number of parameters; The encoding features of the observed speech signal are obtained by adding it to the input signal after one-dimensional convolution.

5. The method according to claim 4, characterized in that The method further comprises: Concatenating the feature vector with the encoded features of the observed speech signal according to the channel dimension to obtain an intermediate representation between the target speaker's speech and the observed signal; The concatenated features are normalized and nonlinearly transformed using 1×1 convolution, and the features of the speaker's speech are obtained based on a preset convolutional coding module; After parameterizing the activation function, 1×1 convolution and activation function sigmoid, the mask information of the target speech coding features is obtained.

6. The method according to claim 5, characterized in that The method further comprises: The preset calculation method includes short-time Fourier transform, covariance matrix multiplication, and generalized eigenvalue decomposition operations.

7. The method according to claim 2, characterized in that The method further comprises: According to the separated speech sources of each channel, the steering vector of each speech source is obtained through generalized eigenvalue decomposition, and the angle between the steering vectors is calculated by the error weighted minimization method to determine whether the non-target speech source is in the same direction as the target speech source.

8. A target speech extraction device based on masked beamforming, characterized in that: The device comprises: A microphone array pickup module is used to collect multi-channel observation signals through a microphone array; A first-stage processing module is used to extract a feature vector of a target speech through a speaker encoder through the registered speaker speech information, extract coding features of each observation signal through a speech encoder, obtain mask information of all speech sources through a mask estimator by fusing the feature vector and the coding features, the all speech sources include target speech sources and non-target speech sources, obtain pre-separated time-domain speech information through masking processing and decoding mapping; and, determine whether the speech sources are in the same direction through the pre-separated time-domain speech information, and calculate the weight of minimum variance distortion-free corresponding beamforming according to the determination result and the mask information to achieve speech separation; The second-stage processing module is used to use the time-domain speech as an auxiliary input of the target speech separation network and extract feature vectors, construct new fusion features by combining the feature vectors of the registered speaker's speech and the feature vectors of the separated target speech, and more accurately estimate the mask information through the mask estimator and the decoding map; and, based on the pre-separated mask information, repeatedly perform a stage of beamforming speech separation process until the final target speech extraction is completed.

9. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method according to any one of claims 1 to 7 when executed by the processor.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Voice processing method and device and electronic equipment

    CN112466327A

  • Target speaker voice extraction method, system and device based on comparative learning and medium

    CN115910039A

  • Voice extraction method and system based on hourglass structure and self-attention mechanism

    CN116665655A

  • Target voice separation method and system based on cross-modal loss

    CN118016093A

  • Target speaker voice extraction method and device

    CN119007728A