Target speech extraction method based on block feature fusion and dual-path Transform

By introducing block feature fusion and dual-path Transformer into the target speech extraction technology, the problem of target speech extraction in multi-person speaking scenarios is solved, and a more efficient and purer speech extraction effect is achieved.

CN119993180AActive Publication Date: 2025-05-13ZHEJIANG UNIV OF TECH

Patent Information

Application Number
CN202510085354.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The prior art is difficult to effectively extract the voice of the target speaker in scenarios where multiple people speak at the same time, and the time domain method uses the reference voice characteristics for limited use, so the purity of the reconstruction results is not high.

Method used

A target speech extraction method based on block feature fusion and dual-path Transformer is adopted. By convolutionally encoding the mixed audio and reference speech, block feature fusion is performed to improve the utilization of reference features, and the fusion features are processed using a dual-path Transformer to generate an extraction mask to achieve accurate recognition and extraction of the target speaker's voice.

Benefits of technology

The performance of the target speaker's voice extraction is improved, the amount of network parameters is reduced, the extraction operation speed is improved, and the better extraction effect is achieved, reducing the presence of noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993180A_ABST
    Figure CN119993180A_ABST
Patent Text Reader

Abstract

A target speech extraction method based on block feature fusion and a dual-path Transform comprises the steps that firstly, a block feature fusion method is utilized to enable each segment of block to obtain complete reference speech features of a target speaker, and the utilization rate of the features is improved; performing global and local modeling on the signal by using a dual-path Transform, and efficiently extracting voice information of a target speaker; and finally, pre-training the target speaker reference voice feature extraction network to improve the training efficiency, and jointly optimizing each component in the whole network through multi-task learning to enable the extraction effect to be optimal. According to the method, voice extraction of the target speaker is achieved, compared with an existing extraction method, the model parameter quantity is lower, the extraction speed is higher, and higher extraction purity is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of audio processing, and in particular relates to a target speech extraction method based on block feature fusion and dual-path Transformer. Background Art

[0002] With the rapid development of artificial intelligence, voice interaction has become an important way of human-computer interaction, which can effectively improve the user experience. Most of the existing voice interaction technologies are based on the premise of a single speaker. In scenarios where multiple people speak at the same time, the performance will deteriorate sharply. The traditional method is to use speech separation technology to achieve the separation of multiple speakers, but the speech separation method requires the acquisition or estimation of the number of speakers in the mixed speech in advance, which is often unknown in the actual environment. When the number of speakers in the environment increases, the computing power requirements of speech separation technology will also increase. Therefore, a speech processing technology is needed to separate the voice of the target speaker from the complex background environment.

[0003] The human brain can focus auditory attention on specific sounds by blocking out background interference sounds. Target speech extraction is a technology that imitates human auditory attention. This technology is based on neural networks and provides a reference speech of a speaker that has not been seen during the training process. The system guides attention to the target speaker based on this reference speech, simulating human autonomous attention, thereby extracting the speaker's target speech from the mixed speech. In previous studies, a common method is to extract in the frequency domain, extract the frequency domain information of the reference audio and the mixed audio as the system input, generate a mask that only allows the target speaker's voice to pass, apply the mask to the mixed audio to complete the extraction, and finally convert it into a time domain signal output. Zmolikova K et al. performed short-time Fourier transform (STFT) processing on the input speech as the input of the proposed SpeakerBeam model, and used a bidirectional long short-term memory network to extract the target speech, which can be used to extract longer and more developed mixed speech. Wang Q et al. divided the system into two parts: reference speech feature extraction and frequency domain mask generation network, and used convolutional neural network (CNN) combined with long short-term memory network to improve the extraction effect of frequency domain method. However, the frequency domain method has window effects and phase estimation problems during signal reconstruction. For this reason, some scholars have begun to study the extraction of target speakers in the time domain. Chenglin Xu et al. proposed a multi-scale time domain speaker extraction network SpEX, which uses a time domain convolutional network (TCN) module to model and extract mixed audio, achieving better results than frequency domain extraction and verifying the feasibility of their method. They also further improved the reference audio information extraction and proposed SpEx+, which uses the reference audio time domain feature information instead of the spectrogram to achieve complete time domain speaker extraction.

[0004] In summary, directly extracting the target speaker's voice from the time domain signal has become the mainstream direction at present, but the existing methods have limited utilization of the characteristics of the reference speech, and the reconstructed extraction results are generally pure and still have certain noises. Solving the above problems and further improving the performance of target speaker voice extraction are of great significance to the practical application of artificial intelligence. Summary of the invention

[0005] In order to overcome the shortcomings of the existing technology, the present invention provides a target speech extraction method based on block feature fusion and dual-path Transformer. In view of the problem of target speaker speech extraction, the present invention proposes a block feature fusion method to improve the utilization rate of the target speaker reference features. At the same time, the present invention also proposes the introduction of a dual-path Transformer network for target speaker extraction to further improve the purity of target speech reconstruction. Finally, the present invention achieves better target speaker speech extraction by pre-training the target speaker reference speech feature extraction network and using multi-task learning to optimize the overall network.

[0006] The technical solution adopted by the present invention to solve its technical problem is:

[0007] A target speech extraction method based on block feature fusion and dual-path Transformer includes the following steps:

[0008] Step 1: Convolutionally encode the mixed audio, convert the speech information into feature representation, encode the speech using a trainable convolution kernel, optimize the feature representation through training to achieve the best effect, input the mixed speech signal y, and output the mixed speech feature matrix Y;

[0009] Step 2: Convolutionally encode the target speaker's reference speech and extract its feature sequence. Multi-scale convolution can obtain features of different resolutions and simplify feature expression by reducing feature dimensions. Input the target speaker's reference speech Output target speaker speech features

[0010] Step 3: Divide the mixed audio features and reference speech features into blocks and perform block-level feature fusion to improve the utilization of reference features and enable the extraction process to more accurately identify the speaker's voice; input the mixed speech feature matrix Y and the target speaker's speech features Output feature fusion matrix E;

[0011] Step 4: Use the dual-path Transformer to process the fused features, calculate multi-head attention on the fused features, perform local and global modeling, so as to achieve accurate recognition and extraction of the target speaker, and finally generate a mask matrix that only allows the target speaker's voice information to pass through, input the feature fusion matrix E, and output the target speaker mask Mask;

[0012] Step 5: Modulate the mixed audio features and the extracted mask, filter out the speech features that are not related to the speaker, decode the modulated response features, restore the time domain signal from the features, obtain the target speech, input the mixed speech coding feature matrix Y and the target speaker mask Mask, and output the extracted target speaker speech

[0013] Step 6: Pre-train the reference feature extraction network to speed up the training convergence speed, use multi-task learning to jointly optimize the entire network, adjust the network to the best state, input the target speaker's speech feature Y, the target speaker's speech extraction result Output pre-trained reference feature extraction network and multi-task learning loss.

[0014] Furthermore, in step 1, the mixed speech signal Input a one-dimensional convolution layer with N convolution kernels. The size of each convolution kernel is L, and its step length is L / 2. The encoding of the mixed speech sequence can be expressed as:

[0015]

[0016] Wherein, k represents the frame sequence number, k∈{0,1,2,...,K}, K=2(TL) / L+1 is the total number of frames under the step size L / 2, y represents the frame sequence, L is the convolution kernel size, U represents the convolution kernel matrix, N is the number of convolution kernels, i and j represent the matrix element numbers, i = 0, 1, ..., N-1, j = 0, 1, ..., K-1, ReLU represents the activation function f(x) = max(x, 0);

[0017] Perform encoding calculation to obtain the feature matrix containing mixed audio information

[0018] Furthermore, the process of step 2 is:

[0019] Step 2.1: Perform time-domain multi-scale convolution coding on the reference speech to convert the time-domain information into feature representation. Reference speech feature extraction is based on the speech coding module of the previous process. The reference speech s is passed through three parallel convolution encoders. The convolution kernel size of each convolution layer is L1, L2, and L3 respectively. Convolution kernels of different scales encode the reference speech with different resolutions. Multi-scale speech feature S' = [S1, S2, S3];

[0020] Step 2.2: Use the residual and pooling modules to process multi-scale speech features, reduce the feature dimension to extract the feature sequence, and the reference speech feature S' will pass through N R A stacked residual module to obtain appropriate speaker feature representation;

[0021] The residual module consists of two layers of CNN with a convolution kernel size of 1*1, a one-dimensional maximum pooling layer and a residual connection. The batch normalization BN layer and the parameter ReLU (PReLU) activation function normalize and nonlinearly transform the output of each layer of CNN; the kernel size of the 1-D maximum pooling layer is 1*3, which is used to remove silence and compress the time series to 1 / 3 of the original; finally, point convolution is used to adjust the number of channels and average pooling is performed to obtain the reference feature coefficients Where D is the number of channels, that is, the predetermined feature dimension.

[0022] The process of step three is:

[0023] Step 3.1: Perform 50% overlapping cuts on each channel of the mixed speech feature and reorganize the blocks into a new feature matrix. The method of segmentation in the feature sequence dimension is adopted to segment the long feature sequence of length K into several small segments of length F, where each small segment overlaps with the previous segment by 50%, and finally can be divided into M small segments. The chunking operation is defined as Chunk(-), and the chunking result can be expressed as

[0024] Step 3.2: Fuse the mixed speech features after segmentation with the reference features of the target speaker, and add complete reference feature information to each segment. To solve the problems of low reference feature utilization, high hardware resource occupancy and slow training speed, a segment feature fusion method is proposed. In the production In the process, the feature dimension is set to the same length as the block, that is, D = F. Will By exchanging the dimensions of The P(-) operation is a dimension swap, which copies the second and third dimensions. The size is the same as E S Consistent, that is The C(-) operation is a copy operation, and the broadcast mechanism of pytorch can implement this operation more efficiently; finally, the feature matrix Y of the same size is obtained c and Then use the addition method to achieve feature fusion, that is After copying and adding, the final feature matrix used for extraction is obtained. The dotted box represents a small block of a feature channel, which has obtained all the speech features of the target speaker and improved the utilization rate of the features.

[0025] The process of step 4 is as follows:

[0026] Step 4.1: Use the dual-path Transformer to process the fused features, identify and extract the target speaker's voice information. In order to effectively learn the order information of the sequence in the dual-path structure and promote the convergence of the model, the long short-term memory network is introduced. The calculation formula is as follows:

[0027]

[0028] Among them, Z X Represents the input data of the Transformer module, Z M represents the intermediate variable after multi-head attention calculation, and O represents the output after the Transformer module operation;

[0029] The dual-path Transformer structure is a local and global Transformer. The two modules have the same structure and are used for local and global information modeling respectively. The modeling of different parts is achieved through dimension exchange. Both the local and global modules use a residual structure, that is, the input E is divided into two paths, and then added again after processing. The local and global processing is repeated B times in total to obtain the speech extraction output E o ;

[0030] Step 4.2: Restore the feature blocks processed by the dual-path Transformer to generate the extraction mask Mask, the output of the dual-path Transformer Y after block division with mixed speech feature matrix Y c The dimensions are consistent. In order to perform mask modulation with Y, the inverse operation of the block is required, which is defined as M(-). Mask is a mask matrix that only allows the target speaker to pass, and its dimension is consistent with the mixed speech feature matrix Y.

[0031] The process of step five is:

[0032] Step 5.1: Mask modulation of the mixed speech feature matrix. For the mixed speech coding feature matrix Y, according to the reference features of the target speaker provided, a mask matrix Mask is trained to allow only the target speaker's voice to pass through. The voices of non-target speakers (interfering speakers, noise) will be filtered out as much as possible. This process is called mask modulation. The result S is the modulation response of the target speaker, and its calculation formula is as follows:

[0033]

[0034] Where Y represents the hybrid speech coding feature matrix, Mask represents the target speaker mask matrix, S represents the modulation response of the target speaker, represents element-wise multiplicative modulation;

[0035] Step 5.2: Feature decoding of the modulation response is performed to obtain the extracted speech. Target speech decoding is to convert the speech features from a two-dimensional feature matrix into one-dimensional speech information through an inverse convolution operation. This process is also called target speech decoding. The target speech decoder is composed of a transposed convolution. The transposed convolution does not have fixed parameters initially, but is composed of learnable parameters. After training, the optimal decoding parameters are obtained to optimize the extraction task. The calculation formula is as follows:

[0036]

[0037] Where S represents the modulation response of the target speaker, represents the extracted time-domain speech signal of the target speaker, and Tconv represents the transposed convolution operation.

[0038] The process of step six is:

[0039] Step 6.1: Pre-train the reference feature extraction network so that the network can distinguish the identity of the speaker based on the sound. The function of the reference feature extraction network is to extract the speech features of the target speaker in the reference audio. This speech feature can be used to identify different speakers to confirm the target speaker from multiple voices. The extracted speech features are passed through a linear layer and a normalized exponential layer to obtain the probability of speaker prediction. The cross-entropy loss is used as the loss function to train and optimize the network. The calculation formula is as follows:

[0040]

[0041] Where i represents the speaker number, i∈{1,2,…,N s},N s is the total number of speakers in the training dataset, p i represents the true category label of the i-th speaker, represents the predicted speaker probability;

[0042] Step 6.2: Use multi-task learning methods to train multiple modules of the network and optimize the extraction performance. The target speaker extraction network is a network composed of multiple modules. Each module has its own function. It is necessary to combine various modules to complete the speech extraction task. Using multi-task learning training methods and training various modules of the network at the same time can make the network reach the best state. The pre-trained reference speech feature extraction module will also be optimized at the same time to adapt to the speech extraction task; at the same time, due to the pre-training of this module, the multi-task learning training process of the entire network will be accelerated;

[0043] The evaluation of speech extraction effect is to compare the similarity between the extracted speech and the original speech before mixing. The closer they are, the better the extraction effect. SI-SDR is an indicator used to calculate the similarity between two signals. s is the original signal before mixing, that is, the correct answer. is the reconstructed signal after the network. There is a certain distance between these two vectors, so Decompose into perpendicular to s and parallel to s Then the calculation formula of SI-SDR is:

[0044]

[0045] Calculate the extracted time domain speech The SI-SDR of the original speech s is used as an indicator of the network extraction effect, with the goal of minimizing the signal reconstruction error, namely SI-SDR loss:

[0046] ML-loss=αSI-SDR-loss+(1-α)CE-loss (7)

[0047] Among them, α represents the weight of multi-task learning loss, SI-SDR-loss represents the SI-SDR loss value, and CE-loss represents the cross entropy loss value.

[0048] The technical concept of the present invention is: in the scene where multiple people are speaking, the human-computer voice interaction function is difficult to apply. The target speaker voice extraction can extract the target voice from the mixed voice of multiple people by obtaining the reference voice information of the target. Compared with traditional voice separation, it is more targeted to the multi-person voice scene. Existing extraction methods are mostly based on frequency domain information, and there are problems such as window effect and phase estimation during signal reconstruction. The method of the present invention first encodes the speech in the time domain, and adopts the block feature fusion method to improve the utilization rate of the target reference feature information. Then, a dual-path Transformer is used to process the fused information, calculate multi-head attention, model the voice information locally and globally, identify the target speaker from the mixed voice, generate an extraction mask, and reconstruct the target voice through mask modulation and decoding. Finally, pre-training and multi-task learning are used to train and optimize the entire network.

[0049] The beneficial effects of the present invention are mainly manifested in: reducing the amount of network parameters, improving the extraction operation speed, and achieving better extraction effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 This is a system structure diagram of the target speech extraction method based on block feature fusion and dual-path Transformer.

[0051] Figure 2 It is a flow chart of hybrid speech coding.

[0052] Figure 3 This is the convolutional coding calculation principle diagram.

[0053] Figure 4 It is a reference speech feature extraction flow chart.

[0054] Figure 5 It is the structure diagram of the residual and pooling module.

[0055] Figure 6 It is a feature matrix block flow chart.

[0056] Figure 7 It is a flowchart of block feature fusion.

[0057] Figure 8 It is the Transformer structure diagram.

[0058] Fig. 9 It is a dual-path Transformer structure diagram.

[0059] Fig.10 It is a reference feature extraction network pre-training flowchart.

[0060] Fig.11 This is the SI-SDR calculation principle diagram. Fig.12 These are some experimental effect diagrams of the target speaker’s speech extraction experiment. DETAILED DESCRIPTION

[0061] The present invention will be further described below in conjunction with the accompanying drawings.

[0062] Reference Figures 1 to 12 , a target speech extraction method based on block feature fusion and dual-path Transformer, comprising the following steps:

[0063] Step 1: Convolutionally encode the mixed audio and convert the speech information into feature representation. Use a trainable convolution kernel to encode the speech and optimize the feature representation through training to achieve the best effect. Input the mixed speech signal y and output the mixed speech feature matrix Y. The implementation method is as follows:

[0064] Mixed speech signal Enter a one-dimensional convolutional layer, such as Figure 2 As shown in the figure, the convolution layer has N convolution kernels, each of which has a size of L, and its step length is L / 2. The encoding of the mixed speech sequence can be expressed as:

[0065]

[0066] Wherein, k represents the frame sequence number, k∈{0,1,2,...,K}, K=2(TL) / L+1 is the total number of frames under the step size L / 2, y represents the frame sequence, L is the convolution kernel size, U represents the convolution kernel matrix, N is the number of convolution kernels, i and j represent the matrix element numbers, i = 0, 1, ..., N-1, j = 0, 1, ..., K-1, and ReLU represents the activation function f(x) = max(x, 0).

[0067] Encoding calculation as Figure 3 As shown, the final feature matrix containing mixed audio information is obtained

[0068]

[0069] Step 2: Convolutionally encode the target speaker's reference speech and extract its feature sequence. Multi-scale convolution can obtain features of different resolutions and simplify feature expression by reducing feature dimensions. Input the target speaker's reference speech Output target speaker speech features The process of this step is as follows Figure 4 , the process is:

[0070] Step 2.1: Perform time-domain multi-scale convolution coding on the reference speech to convert the time-domain information into feature representation. Reference speech feature extraction is based on the speech coding module of the previous process. The reference speech s is passed through three parallel convolution encoders. The convolution kernel size of each convolution layer is L1, L2, and L3 respectively. Convolution kernels of different scales encode the reference speech with different resolutions. Multi-scale speech features S' = [S1, S2, S3].

[0071] Step 2.2: Use the residual and pooling modules to process multi-scale speech features and reduce the feature dimension to extract feature sequences. The reference speech feature S' will be processed by N R stacked residual modules to obtain appropriate speaker feature representation. The residual module structure is as follows Figure 5 shown.

[0072] The residual module consists of two layers of CNN with a convolution kernel size of 1*1, a 1D max pooling layer and a residual connection. The batch normalization (BN) layer and the parametric ReLU (PReLU) activation function normalize and nonlinearly transform the output of each CNN layer. The kernel size of the 1-D max pooling layer is 1*3, which is used to remove silence and compress the time series to 1 / 3 of the original. Finally, point convolution is used to adjust the number of channels and average pooling is performed to obtain the reference feature coefficients. Where D is the number of channels, that is, the predetermined feature dimension.

[0073] Step 3: Divide the mixed audio features and reference speech features into blocks and perform block-level feature fusion to improve the utilization of reference features and enable the extraction process to more accurately identify the speaker's voice. Output feature fusion matrix E, the process is:

[0074] Step 3.1: Cut each channel of the mixed speech feature by 50% overlap and reorganize it into a new feature matrix. Figure 6 As shown, for the mixed speech feature matrix The method of segmentation in the feature sequence dimension is adopted to segment the long feature sequence of length K into several small segments of length F, where each small segment overlaps with the previous segment by 50%, and finally can be divided into M small segments. The chunking operation is defined as Chunk(-), and the chunking result can be expressed as

[0075] Step 3.2: Fuse the mixed speech features after segmentation with the reference features of the target speaker, and add complete reference feature information to each segment. In order to solve the problems of low reference feature utilization, high hardware resource occupancy and slow training speed, the present invention proposes a method for segment feature fusion, such as Figure 7 As shown. For the reference speech feature In the production In the process, the feature dimension is set to the same length as the block, that is, D = F. Will By exchanging the dimensions of The P(-) operation is a dimension swap, which copies the second and third dimensions. The size is the same as E S Consistent, that is The C(-) operation is a copy operation, and the broadcast mechanism of pytorch can implement this operation more efficiently. Finally, the feature matrix Y of the same size is obtained c and Then use the addition method to achieve feature fusion, that is The specific process is shown in the figure. After copying and adding, the feature matrix used for extraction is obtained. The dotted box represents a small block of a feature channel, which has obtained all the speech features of the target speaker and improved the utilization rate of the features.

[0076] Step 4: Use the dual-path Transformer to process the fused features, calculate multi-head attention on the fused features, and perform local and global modeling to achieve accurate recognition and extraction of the target speaker, and finally generate a mask matrix that only allows the target speaker's voice information to pass. Input the feature fusion matrix E and output the target speaker mask Mask. The process is:

[0077] Step 4.1: Use the dual-path Transformer to process the fused features, identify and extract the target speaker's speech information. In order to effectively learn the order information of the sequence in the dual-path structure and promote the convergence of the model, the method of the present invention introduces a long short-term memory network, the structure of which is as follows: Figure 8 As shown, the calculation formula is as follows:

[0078]

[0079] Among them, Z X Represents the input data of the Transformer module, Z M represents the intermediate variable after multi-head attention calculation, and O represents the output after the Transformer module operation.

[0080] The dual-path Transformer structure is as follows Fig. 9 As shown in the figure, the main structure is the local and global Transformer. The two modules have the same structure and are used for local and global information modeling respectively. The modeling of different parts is achieved through dimension exchange. Both the local and global modules use a residual structure, that is, the input E is divided into two paths, which are added again after processing. The local and global processing is repeated B times in total to obtain the speech extraction output E o .

[0081] Step 4.2: Restore the feature blocks processed by the dual-path Transformer to generate the extraction mask Mask. Output of the dual-path Transformer Y after block division with mixed speech feature matrix Y c In order to perform mask modulation with Y, the inverse operation of the block is required, which is defined as M(-), so Mask is a mask matrix that only allows the target speaker to pass, and its dimension is consistent with the mixed speech feature matrix Y.

[0082] Step 5: Modulate the mixed audio features and the extracted mask to filter out speech features that are not related to the speaker. Decode the modulated response features, restore the time domain signal from the features, and obtain the target speech. Input the mixed speech coding feature matrix Y and the target speaker mask Mask, and output the extracted target speaker speech. The process is:

[0083] Step 5.1: Mask modulation of the mixed speech feature matrix. For the mixed speech coding feature matrix Y, according to the reference features of the target speaker provided, a mask matrix Mask is trained to allow only the target speaker's voice to pass through. The voices of non-target speakers (interfering speakers, noise) will be filtered out as much as possible. This process is called mask modulation. The result S is the modulation response of the target speaker, and its calculation formula is as follows:

[0084]

[0085] Where Y represents the hybrid speech coding feature matrix, Mask represents the target speaker mask matrix, S represents the modulation response of the target speaker, Represents element-wise multiplicative modulation.

[0086] Step 5.2: Decode the modulation response to obtain the extracted speech. Target speech decoding is to convert the speech features from a two-dimensional feature matrix into one-dimensional speech information through an inverse convolution operation. This process is also called target speech decoding. The target speech decoder is composed of transposed convolution. The transposed convolution does not have fixed parameters initially, but is composed of learnable parameters. After training, the optimal decoding parameters are obtained to optimize the extraction task. The calculation formula is as follows:

[0087]

[0088] Where S represents the modulation response of the target speaker, represents the extracted time-domain speech signal of the target speaker, and Tconv represents the transposed convolution operation.

[0089] Step 6: Pre-train the reference feature extraction network to speed up the training convergence. Use multi-task learning to jointly optimize the entire network and adjust the network to the best state. Input the target speaker's speech feature Y and the target speaker's speech extraction result Output pre-trained reference feature extraction network and multi-task learning loss. The process is:

[0090] Step 6.1: Pre-train the reference feature extraction network so that the network can distinguish the identity of the speaker based on the sound. The function of the reference feature extraction network is to extract the speech features of the target speaker in the reference audio. This speech feature can be used to identify different speakers and confirm the target speaker from multiple voices. The pre-training process of the feature extraction network is as follows: Fig.10 As shown in the figure, the extracted speech features are passed through the linear layer and the normalized exponential layer to obtain the probability of speaker prediction. The cross-entropy loss is used as the loss function to train and optimize the network. The calculation formula is as follows:

[0091]

[0092] Where i represents the speaker number, i∈{1,2,…,N s},N s is the total number of speakers in the training dataset, p i represents the true category label of the i-th speaker, represents the predicted speaker probability.

[0093] Step 6.2: Use multi-task learning methods to train multiple modules of the network and optimize the extraction performance. The target speaker extraction network is a network composed of multiple modules, each of which has its own function, and it is necessary to combine the modules to complete the speech extraction task. Using the multi-task learning training method, training each module of the network at the same time can make the network reach the best state. The pre-trained reference speech feature extraction module will also be optimized at the same time to adapt to the speech extraction task. At the same time, due to the pre-training of this module, the multi-task learning training process of the entire network will be accelerated.

[0094] The evaluation of speech extraction effect is generally to compare the similarity between the extracted speech and the original speech before mixing. The closer the similarity, the better the extraction effect. SI-SDR is an indicator used to calculate the similarity between two signals. Fig.11 As shown, s is the original signal before mixing, that is, the correct answer, is the reconstructed signal after the network. There is a certain distance between these two vectors, so Decompose into perpendicular to s and parallel to s Then the calculation formula of SI-SDR is:

[0095]

[0096] Calculate the extracted time domain speech The SI-SDR of the original speech s is used as an indicator of the network extraction effect, with the goal of minimizing the signal reconstruction error, namely SI-SDR loss:

[0097] ML-loss=αSI-SDR-loss+(1-α)CE-loss (7)

[0098] Among them, α represents the weight of multi-task learning loss, SI-SDR-loss represents the SI-SDR loss value, and CE-loss represents the cross entropy loss value.

[0099] In order to verify the performance of the method of the present invention, a comparative experiment was conducted. The dataset used in the experiment is WSJ0-2mix-extr, which is a 2-person mixed speech dataset, in which the speech signal is resampled to 8kHz sampling rate based on the WSJ0 speech database. The dataset is divided into three subsets: training set (20,000 sentences), validation set (5,000 sentences), and test set (3,000 sentences). Each sentence data contains three sentences: mixed speech, original speech, and reference speech.

[0100] The loss weight α of multi-task learning is set to 0.5, the encoder convolution kernel length L = 4, the encoder step size stride = L / 2 = 2, the number of encoder convolution kernels N = 64, the number of multi-head attention H = 4, and the number of dual-path transformer cycles B = 6. The total number of training epochs = 100, batch size = 4, initial learning rate = 0.125, using the warm-up learning strategy, and using Adam as the optimizer.

[0101] Experiment 1: Target speaker voice extraction experiment. The test set has 3,000 samples. Some experimental results are as follows: Fig.12 As shown, on the left is the time domain diagram of the audio (horizontal axis time, vertical axis amplitude), and on the right is the spectrogram of the audio (horizontal axis time, vertical axis frequency). From top to bottom, they are the mixed speech y, the original audio of the target speaker s, and the extracted result and reference audio The time domain graph of the mixed speech is the superposition of the two voices, and the spectrogram is relatively messy. The reference speech and the original speech have no connection except for the speaker. Compared with the mixed audio, the extraction result Only the target speaker's voice is retained in the mixed speech. From the time domain graph and the spectrogram, it can be found that the interfering voice is reduced and the extraction purity is high. Compared with the original audio s, the extraction results are basically reconstructed in each pronunciation, without extra noise, and the spectrogram is basically completely consistent, achieving good extraction performance.

[0102] Experiment 2: Performance comparison experiment. The same test set of 3000 samples was used to test the existing methods, and the average SI-SDR was used as the performance indicator. The results are shown in Table 1. It can be seen that this method not only reduces the number of model parameters, but also improves the extraction effect by 1.43dB compared with the previous time domain method, and has the best performance.

[0103] Model Methods Time-frequency domain processing Parameter quantity (M) SI-SDR(dB) Unprocessed - - 2.5 SpeakerBeam Frequency Domain 19.30 9.22 VoiceFilter Frequency Domain 18.90 12.4 SpEx Time Domain 10.80 14.6 SpEx+ Time Domain 11.10 18.2 Method of the present invention Time Domain 4.39 19.63

[0104] Table 1

[0105] The contents described in the embodiments of this specification are merely enumerations of implementation forms of the inventive concept and are for illustrative purposes only. The protection scope of the present invention should not be considered to be limited to the specific forms described in this embodiment, and the protection scope of the present invention also extends to equivalent technical means that can be thought of by ordinary technicians in this field based on the inventive concept.

Claims

1. A target speech extraction method based on block feature fusion and dual-path Transformer, characterized in that: The method comprises the following steps: Step 1: Convolutionally encode the mixed audio, convert the speech information into feature representation, encode the speech using a trainable convolution kernel, optimize the feature representation through training to achieve the best effect, input the mixed speech signal y, and output the mixed speech feature matrix Y; Step 2: Convolutionally encode the target speaker's reference speech and extract its feature sequence. Multi-scale convolution can obtain features of different resolutions and simplify feature expression by reducing feature dimensions. Input the target speaker's reference speech s and output the target speaker's speech features. Step 3: Divide the mixed audio features and reference speech features into blocks and perform block-level feature fusion to improve the utilization of reference features and enable the extraction process to more accurately identify the speaker's voice; input the mixed speech feature matrix Y and the target speaker's speech features Output feature fusion matrix E; Step 4: Use the dual-path Transformer to process the fused features, calculate multi-head attention on the fused features, perform local and global modeling, so as to achieve accurate recognition and extraction of the target speaker, and finally generate a mask matrix that only allows the target speaker's voice information to pass through, input the feature fusion matrix E, and output the target speaker mask Mask; Step 5: Modulate the mixed audio features and the extracted mask, filter out the speech features that are not related to the speaker, decode the modulated response features, restore the time domain signal from the features, obtain the target speech, input the mixed speech coding feature matrix Y and the target speaker mask Mask, and output the extracted target speaker speech Step 6: Pre-train the reference feature extraction network to speed up the training convergence speed, use multi-task learning to jointly optimize the entire network, adjust the network to the best state, input the target speaker's speech feature Y, the target speaker's speech extraction result Output pre-trained reference feature extraction network and multi-task learning loss.

2. The target speech extraction method based on block feature fusion and dual-path Transformer as claimed in claim 1, characterized in that: In the step 1, the mixed speech signal Input a one-dimensional convolution layer with N convolution kernels. The size of each convolution kernel is L, and its step length is L / 2. The encoding of the mixed speech sequence can be expressed as: Wherein, k represents the frame sequence number, k∈{0,1,2,...,K}, K=2(TL) / L+1 is the total number of frames under the step size L / 2, y represents the frame sequence, L is the convolution kernel size, U represents the convolution kernel matrix, N is the number of convolution kernels, i and j represent the matrix element numbers, i = 0, 1, ..., N-1, j = 0, 1, ..., K-1, ReLU represents the activation function f(x) = max(x, 0); Perform encoding calculation to obtain the feature matrix containing mixed audio information 3. The target speech extraction method based on block feature fusion and dual-path Transformer as described in claim 1 or 2, characterized in that: The process of step 2 is: Step 2.1: Perform time-domain multi-scale convolution coding on the reference speech to convert the time-domain information into feature representation. The reference speech feature extraction is based on the speech coding module of the previous process. The reference speech s is passed through three parallel convolution encoders. The convolution kernel size of each convolution layer is L1, L2, and L3 respectively. Convolution kernels of different scales encode the reference speech with different resolutions. Multi-scale speech feature S' = [S1, S2, S3]; Step 2.2: Use the residual and pooling modules to process multi-scale speech features, reduce the feature dimension to extract the feature sequence, and the reference speech feature S' will pass through N R A stacked residual module to obtain appropriate speaker feature representation; The residual module consists of two layers of CNN with a convolution kernel size of 1*1, a one-dimensional maximum pooling layer and a residual connection. The batch normalization BN layer and the parameter ReLU activation function normalize and nonlinearly transform the output of each layer of CNN; the kernel size of the 1-D maximum pooling layer is 1*3, which is used to remove silence and compress the time series to 1 / 3 of the original; finally, point convolution is used to adjust the number of channels and average pooling is performed to obtain the reference feature coefficients Where D is the number of channels, that is, the predetermined feature dimension.

4. The target speech extraction method based on block feature fusion and dual-path Transformer as claimed in claim 1 or 2, characterized in that: The process of step three is: Step 3.1: Perform 50% overlapping cuts on each channel of the mixed speech feature and reorganize the blocks into a new feature matrix. The method of segmentation in the feature sequence dimension is adopted to segment the long feature sequence of length K into several small segments of length F, where each small segment overlaps with the previous segment by 50%, and finally can be divided into M small segments. The chunking operation is defined as Chunk(-), and the chunking result can be expressed as Step 3.2: Fuse the mixed speech features after segmentation with the reference features of the target speaker, and add complete reference feature information to each segment. To solve the problems of low reference feature utilization, high hardware resource occupancy and slow training speed, a segment feature fusion method is proposed. In the production In the process, the feature dimension is set to the same length as the block, that is, D = F. Will By exchanging the dimensions of The P(-) operation is a dimension swap, which copies the second and third dimensions. The size is the same as E S Consistent, that is The C(-) operation is a copy operation, and the broadcast mechanism of pytorch can implement this operation more efficiently; finally, the feature matrix Y of the same size is obtained c and Then use the addition method to achieve feature fusion, that is After copying and adding, the final feature matrix used for extraction is obtained. The dotted box represents a small block of a feature channel, which has obtained all the speech features of the target speaker and improved the utilization rate of the features.

5. The target speech extraction method based on block feature fusion and dual-path Transformer as claimed in claim 1 or 2, characterized in that: The process of step 4 is as follows: Step 4.1: Use the dual-path Transformer to process the fused features, identify and extract the target speaker's voice information. In order to effectively learn the order information of the sequence in the dual-path structure and promote the convergence of the model, the long short-term memory network is introduced. The calculation formula is as follows: Among them, Z X Represents the input data of the Transformer module, Z M represents the intermediate variable after multi-head attention calculation, and O represents the output after the Transformer module operation; The dual-path Transformer structure is a local and global Transformer. The two modules have the same structure and are used for local and global information modeling respectively. The modeling of different parts is achieved through dimension exchange. Both the local and global modules use a residual structure, that is, the input E is divided into two paths, which are added again after processing. The local and global processing is repeated B times in total to obtain the speech extraction output E o ; Step 4.2: Restore the feature blocks processed by the dual-path Transformer to generate the extraction mask Mask, the output of the dual-path Transformer Y after block division with mixed speech feature matrix Y c The dimensions are consistent. In order to perform mask modulation with Y, the inverse operation of the block is required, which is defined as M(-). Mask is a mask matrix that only allows the target speaker to pass, and its dimension is consistent with the mixed speech feature matrix Y.

6. The target speech extraction method based on block feature fusion and dual-path Transformer as claimed in claim 1 or 2, characterized in that: The process of step five is: Step 5.1: Perform mask modulation on the mixed speech feature matrix. For the mixed speech coding feature matrix Y, according to the provided target speaker reference features, a mask matrix Mask is trained to allow only the target speaker's voice to pass through. The voices of non-target speakers will be filtered out as much as possible. This process is called mask modulation. The result S is the modulation response of the target speaker, and its calculation formula is as follows: Where Y represents the hybrid speech coding feature matrix, Mask represents the target speaker mask matrix, S represents the modulation response of the target speaker, represents element-wise multiplicative modulation; Step 5.2: Feature decoding is performed on the modulation response to obtain the extracted speech. Target speech decoding is to convert the speech features from a two-dimensional feature matrix into one-dimensional speech information through an inverse convolution operation. This process is also called target speech decoding. The target speech decoder is composed of transposed convolution. The transposed convolution does not have fixed parameters initially, but is composed of learnable parameters. After training, the optimal decoding parameters are obtained to optimize the extraction task. The calculation formula is as follows: Where S represents the modulation response of the target speaker, represents the extracted time-domain speech signal of the target speaker, and Tconv represents the transposed convolution operation.

7. The target speech extraction method based on block feature fusion and dual-path Transformer as claimed in claim 1 or 2, characterized in that: The process of step six is: Step 6.1: Pre-train the reference feature extraction network so that the network can distinguish the identity of the speaker based on the sound. The extracted speech features are passed through the linear layer and the normalized exponential layer to obtain the probability of speaker prediction. The cross entropy loss is used as the loss function to train and optimize the network. The calculation formula is as follows: Where i represents the speaker number, i∈{1,2,…,N s },N s is the total number of speakers in the training dataset, p i represents the true category label of the i-th speaker, represents the predicted speaker probability; Step 6.2: Use multi-task learning methods to train multiple modules of the network and optimize the extraction performance. The target speaker extraction network is a network composed of multiple modules. Each module has its own function. It is necessary to combine various modules to complete the speech extraction task. Using the multi-task learning training method, training various modules of the network at the same time can make the network reach the best state; the pre-trained reference speech feature extraction module will also be optimized at the same time to adapt to the speech extraction task; at the same time, due to the pre-training of this module, the multi-task learning training process of the entire network will be accelerated; The evaluation of speech extraction effect is to compare the similarity between the extracted speech and the original speech before mixing. The closer they are, the better the extraction effect. SI-SDR is an indicator used to calculate the similarity between two signals. s is the original signal before mixing, that is, the correct answer. is the reconstructed signal after the network. There is a certain distance between these two vectors, so Decompose into perpendicular to s and parallel to s Then the calculation formula of SI-SDR is: Calculate the extracted time domain speech The SI-SDR of the original speech s is used as an indicator of the network extraction effect, with the goal of minimizing the signal reconstruction error, namely SI-SDRloss: ML-loss=αSI-SDR-loss+(1-α)CE-loss (7) Among them, α represents the weight of multi-task learning loss, SI-SDR-loss represents the SI-SDR loss value, and CE-loss represents the cross entropy loss value.

Citation Information

Patent Citations

  • Speech enhancement method fusing Transform and U-net network

    CN114141238A

  • Voice data processing method and device, equipment and storage medium

    CN114913870A

  • Single-channel voice separation method based on deep learning

    CN116612779A

  • Single-channel voice separation method based on feature fusion

    CN118782072A

  • Target voice extraction method and device based on multi-reference clue fusion

    CN119229875A

Cited By

  • Music source extraction method, device and product based on reference audio and MIDI guidance

    CN121148413A

  • Voice signal denoising method and device, electronic equipment and program product

    CN121306167A