Target speaker extraction method and system based on multi-scale multi-modal alignment network
By using a multi-scale, multi-modal alignment network, the problems of prior knowledge dependence and insufficient speech features in target speaker extraction methods are solved. Cross-modal data alignment and parameter optimization are achieved, thereby improving the performance of target speaker extraction.
Patent Information
- Application Number
- CN202510290875.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-03-12
AI Technical Summary
Existing target speaker extraction methods rely on prior knowledge of the target speaker, resulting in insufficient speech feature extraction and difficulty in multimodal fusion, leading to high computational complexity and poor performance.
A multi-scale, multi-modal alignment network is employed to obtain speech embeddings at different time scales through multi-scale encoding. This is combined with multi-directional deep encoding to extract rich speech embeddings. Furthermore, the modal alignment unit minimizes the distance between EEG features and speech embeddings during network training. A loss function is constructed using noise contrast estimation loss and scale-invariant signal distortion ratio loss to achieve cross-modal data alignment and parameter adjustment.
It reduces the difficulty of multimodal fusion, improves the overall performance of target speaker extraction, realizes cross-modal data alignment and network parameter optimization, and enhances the effect of speech feature extraction.
Smart Images

Figure CN120126454B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of target speaker extraction, and more particularly, to: 1. a target speaker extraction method based on a multi-scale multi-modal alignment network; and 2. a target speaker extraction system based on a multi-scale multi-modal alignment network. BACKGROUND
[0002] The cocktail party problem refers to the ability of humans to track a specific conversation in a multi-speaker social environment, which is known as selective auditory attention. Target speaker extraction (TSE) is a basic task in speech signal processing, aiming to extract target speech using auxiliary cues to effectively solve the cocktail party problem. Traditional TSE methods usually rely on pre-recorded speech from the target speaker as a cue, but this method requires prior knowledge of the target speaker.
[0003] Research on auditory attention decoding (AAD) has revealed the correlation between brain neural activity and speech stimuli, opening up new avenues for TSE technology. Researchers explore how to use speech-related information contained in electroencephalogram (EEG) to guide the extraction of target speech, thereby avoiding the need for prior speaker information. Early methods focused on reconstructing the speech envelope from EEG data and matching it with the sound source to identify the target speech, but this method is computationally complex due to the need to separate all sources. Recent research has used EEG data directly as a reference cue and used a common convolutional encoder to extract speech features from speech data, which simplifies the process but increases the difficulty of subsequent multi-modal fusion, and the extraction of speech features is not sufficient. SUMMARY
[0004] Therefore, it is necessary to provide a target speaker extraction method and system based on a multi-scale multi-modal alignment network to address the problems of insufficient speech feature extraction and difficulty in multi-modal fusion in existing TSE methods.
[0005] The present application adopts the following technical solutions:
[0006] In a first aspect, the present application discloses a target speaker extraction method based on a multi-scale multi-modal alignment network, comprising:
[0007] Step one, obtaining a mixed speech Mixture in a multi-speaker scene and electroencephalogram (EEG) data generated by Mixture stimulation;
[0008] Step two, down-sampling Mixture to reduce computational complexity to obtain mixed speech Mixture';
[0009] The EEG is pre-processed to remove noise and then up-sampled to match the Mixture' to obtain the electroencephalogram data EEG';
[0010] Step three, inputting the Mixture' and the EEG' into the trained multi-scale multi-modal alignment network to obtain the target speaker voice
[0011] The multi-scale multi-modal alignment network includes a voice encoding unit, an electroencephalogram encoding unit, a speaker extraction unit, a voice decoding unit, and a modal alignment unit.
[0012] The voice encoding unit is configured to obtain voice embeddings X1-X4 at four time scales from the Mixture through multi-scale encoding, and aggregate X1-X4 into a voice embedding X c , and extract more rich voice embeddings from X1-X4 through multi-directional deep encoding Then, the voice embedding X is dimensionally adjusted to obtain a voice embedding X The electroencephalogram encoding unit is configured to extract an electroencephalogram feature E from the EEG', and pad E to obtain an electroencephalogram feature E with the same time dimension as X The speaker extraction unit is configured to process X and E to obtain a target voice mask M, and multiply M with X c to obtain a target voice feature S. The voice decoding unit is configured to reconstruct X
[0013] The modal alignment unit is configured to align the outputs of the voice encoding unit and the electroencephalogram encoding unit based on contrast learning at the time steps during network training, and calculate a noise contrast estimation loss L InfoNCE .
[0014] The loss function L total used by the multi-scale multi-modal alignment network during network training is:
[0015] L total = L SI-SDR + α * L InfoNCE ;
[0016] In the formula, L SI-SDR represents a scale-invariant signal distortion ratio loss constructed based on the output of the voice decoding unit; L InfoNCE represents the noise contrast estimation loss; and α represents a loss weight coefficient.
[0017] The target speaker extraction method based on the multi-scale multi-modal alignment network implements the method or process according to the embodiments of the present disclosure.
[0018] In a second aspect, the present application discloses a target speaker extraction system based on a multi-scale multi-modal alignment network, which uses the target speaker extraction method based on the multi-scale multi-modal alignment network disclosed in the first aspect.
[0019] The target speaker extraction system based on the multi-scale multi-modal alignment network comprises a data acquisition module, a pre-processing module and a network processing module.
[0020] The data acquisition module is configured to acquire a mixture voice Mixture in a multi-speaker scene and electroencephalogram data EEG generated by the Mixture. The pre-processing module is configured to, on one hand, perform down-sampling processing on the Mixture to reduce the amount of calculation and obtain a mixture voice Mixture'; and on the other hand, perform pre-processing on the EEG to remove noise, and then perform corresponding up-sampling to match the Mixture', and obtain electroencephalogram data EEG'. The network processing module is configured to input the Mixture' and the EEG' into the trained multi-scale multi-modal alignment network for processing, and obtain target speaker voice
[0021] The target speaker extraction system based on the multi-scale multi-modal alignment network implements the method or process according to the embodiments of the present application.
[0022] In a third aspect, the present application discloses a computer program product comprising a computer program. The computer program, when executed by a processor, implements the steps of the target speaker extraction method based on the multi-scale multi-modal alignment network disclosed in the first aspect.
[0023] Compared with the prior art, the present application has the following beneficial effects:
[0024] The present application constructs a multi-scale multi-modal alignment network for target speaker extraction. On one hand, the multi-scale multi-modal alignment network acquires voice embeddings of different time scales through multi-scale encoding, and extracts more rich voice embeddings through multi-directional deep encoding. On the other hand, the multi-scale multi-modal alignment network introduces a modal alignment unit based on contrast learning, which minimizes the distance between the electroencephalogram features and the voice embeddings at the same time step during network training, and constructs a noise contrast estimation loss L InfoNCE The scale-invariant signal distortion ratio loss L SI-SDR The loss function L used in the entire network total is not only capable of realizing alignment of cross-modal data and reducing the difficulty of multi-modal fusion, but also capable of adjusting the overall parameters of the network, thereby ensuring and improving the overall performance of the network in target speaker extraction. BRIEF DESCRIPTION OF DRAWINGS
[0025] In order to describe the technical solutions of the embodiments of the present application or the prior art more clearly, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description only constitute some embodiments of the present application, and all other embodiments obtained by those of ordinary skill in the art without creative effort based on these accompanying drawings also belong to the protection scope of the present application.
[0026] Figure 1 A data flow diagram of the target speaker extraction method based on the multi-scale multi-modal alignment network provided in Embodiment 1 of the present application;
[0027] Figure 2 A structure diagram of the speech encoding unit in Embodiment 1 of the present application; Figure 1
[0028] Figure 3 A structure diagram of the group state space layer in Embodiment 1 of the present application; Figure 2
[0029] Figure 4 A structure diagram of the electroencephalogram decoding unit in Embodiment 1 of the present application; Figure 1
[0030] Figure 5 A structure diagram of the speaker extraction unit and the speech decoding unit in Embodiment 1 of the present application; Figure 1
[0031] Figure 6 A structure diagram of the modal fusion layer in Embodiment 1 of the present application; Figure 1
[0032] Figure 7 A structure diagram of the modal alignment layer in Embodiment 1 of the present application. Figure 1 DETAILED DESCRIPTION The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments only constitute some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort belong to the protection scope of the present application.
[0033] It should be noted that when a component is referred to as being “mounted on” another component, it can be directly on the other component or there can be a middle component. When a component is referred to as being “disposed on” another component, it can be directly disposed on the other component or there can be a middle component. When a component is referred to as being “fixed on” another component, it can be directly fixed on the other component or there can be a middle component.
[0034] It should be noted that when a component is referred to as being “mounted on” another component, it can be directly on the other component or there can be a middle component. When a component is referred to as being “disposed on” another component, it can be directly disposed on the other component or there can be a middle component. When a component is referred to as being “fixed on” another component, it can be directly fixed on the other component or there can be a middle component.
[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0036] Embodiment 1
[0037] Firstly, for the two problems mentioned in the background art: through analysis, it is found that:
[0038] 1. Although the electroencephalogram data and the speech data are input into the model of the TSE method at the same time, they are not completely synchronized in fact - the sound stimulus (i.e. the speech data) must first pass through the ear and the auditory center, and then reach the auditory cortex, thereby generating the electroencephalogram data. Previous studies only focused on the effectiveness of modal fusion, ignoring the time alignment between speech and electroencephalogram, thus increasing the difficulty of multi-modal fusion.
[0039] 2. The common convolutional encoder structure design has room for improvement and cannot fully extract speech features to adapt to electroencephalogram data.
[0040] Therefore, in view of the above two problems, the embodiment 1 provides a target speaker extraction method based on a multi-scale multi-modal alignment network. Referring to Figure 1 , which shows the data flow diagram of the method, which actually also shows the flowchart of the method.
[0041] As Figure 1 shown, a target speaker extraction method based on a multi-scale multi-modal alignment network comprises:
[0042] Step 1, obtaining a mixed speech Mixture in a multi-speaker scene, and electroencephalogram data EEG generated by the Mixture stimulation.
[0043] It should be noted that Mixture and EEG are data of different sampling rates, so subsequent step 2 related operations are required.
[0044] Step 2, down-sampling Mixture to reduce the amount of calculation, to obtain mixed speech Mixture';
[0045] The EEG is first pre-processed to remove noise, and then up-sampled to match Mixture', to obtain electroencephalogram data EEG'.
[0046] Among them, the method for pre-processing EEG includes band-pass filtering, independent component analysis, and re-reference, etc.
[0047] It should be noted that:
[0048] 1, Mixture'∈R B×1×T ; B represents batch size, 1 represents the number of channels of Mixture', and T represents time series length.
[0049] 2, EEG'∈R B×C×T ; C represents the number of channels of EEG'.
[0050] Step three, input Mixture' and EEG' into the trained multi-scale multi-modal alignment network to obtain target speaker voice
[0051] One of the cores of the application is the structure design of the multi-scale multi-modal alignment network, and the second core is the loss function L used in the training of the multi-scale multi-modal alignment network total ; Both of them can balance the processing effect and the consumption of computing resources.
[0052] Referring to Figure 1 , first look at the overall architecture of the multi-scale multi-modal alignment network, which includes: speech encoding part, electroencephalogram encoding part, speaker extraction part, speech decoding part and modal alignment part.
[0053] It should be noted that:
[0054] In the process of training the multi-scale multi-modal alignment network to obtain the trained multi-scale multi-modal alignment network (referred to as the training stage), the speech encoding part, the electroencephalogram encoding part, the speaker extraction part, the speech decoding part and the modal alignment part are all working.
[0055] And in step three (referred to as the prediction stage), for the trained multi-scale multi-modal alignment network, the modal alignment part is not working, and the rest of the network is working.
[0056] (I) First, the specific structure of each part of the network (except the modal alignment part) in the prediction stage is described:
[0057] Ⅰ, the speech encoding part is used to obtain the speech embedding X1~X4 under four time scales from Mixture through multi-scale encoding, and aggregate X1~X4 into speech embedding X c , and extract more rich speech embedding from X1~X4 through multi-directional deep encoding Then adjust the dimension of to obtain speech embedding
[0058] Referring to Figure 2The speech encoding part includes: 1 multi-scale encoding layer, 1 pre-aggregation layer, 1 deep encoding layer, and 1 dimension adjustment layer.
[0059] 101、The multi-scale encoding layer aims to obtain speech embeddings X1-X4 under 4 time scales from Mixture. As shown in Figure 2 The multi-scale encoding layer can be designed based on multi-time window, which can be designed to include: 4 one-dimensional convolution layers and 4 ReLU activation function layers.
[0060] In the multi-scale encoding layer:
[0061] The first one-dimensional convolution layer is used to encode Mixture' through one-dimensional convolution of a short time window;
[0062] The first ReLU activation function layer is used to process the output of the first one-dimensional convolution layer through ReLU activation function to obtain speech embedding X1 under a short time scale;
[0063] The second one-dimensional convolution layer is used to encode Mixture' through one-dimensional convolution of a sub-short time window;
[0064] The second ReLU activation function layer is used to process the output of the second one-dimensional convolution layer through ReLU activation function to obtain speech embedding X2 under a sub-short time scale;
[0065] The third one-dimensional convolution layer is used to encode Mixture' through one-dimensional convolution of a sub-long time window;
[0066] The third ReLU activation function layer is used to process the output of the third one-dimensional convolution layer through ReLU activation function to obtain speech embedding X3 under a sub-long time scale;
[0067] The fourth one-dimensional convolution layer is used to encode Mixture' through one-dimensional convolution of a long time window;
[0068] The fourth ReLU activation function layer is used to process the output of the fourth one-dimensional convolution layer through ReLU activation function to obtain speech embedding X4 under a long time scale.
[0069] That is, in order to obtain more comprehensive speech features, the multi-scale encoding layer first applies 4 parallel one-dimensional convolution layers with different convolution kernel sizes (the lengths of their convolution kernels are short, sub-short, sub-long, and long, corresponding to 4 time scales) to obtain global and local speech information under 4 time windows, and then introduces nonlinearity through ReLU activation function to obtain more rich speech features.
[0070] Of course, the above process of the multi-scale encoding layer can be described in formula as:
[0071]
[0072] In the formula, ReLU(.) represents the ReLU activation function processing; Conv1d i (.) represents the processing of the i-th one-dimensional convolutional layer; N represents the number of speech channels; T s X represents i The size of the time dimension.
[0073] 102. The front-end aggregation layer aims to aggregate X1 to X4 into speech embeddings X. c This is in preparation for the subsequent speaker extraction section.
[0074] like Figure 2 As shown, the pre-aggregation layer can be designed to include: one pre-splitting layer and one one-dimensional convolutional layer.
[0075] In the front aggregation layer:
[0076] The front-end splicing layer is used to splice X1 to X4;
[0077] One-dimensional convolutional layers are used to perform one-dimensional convolution processing on the output of the preceding stitching layer to obtain X. c .
[0078] Of course, the above process of the pre-stitching layer can be described by a formula:
[0079]
[0080] In the formula, Conv1d(.) represents the one-dimensional convolution process; concat(.) represents the concatenation operation.
[0081] 103. The deep coding layer aims to extract richer speech embeddings from X1 to X4.
[0082] like Figure 2 As shown, the deep coding layer can be designed to include: 4-dimensional extension layers, 1 post-concatenation layer, and 2 grouped state space layers.
[0083] In deep coding layers:
[0084] The i-th dimension expansion layer is used to expand X. i Perform dimensional expansion to obtain the expanded speech embedding Y i ;
[0085] The post-stitching layer is used to stitch Y1 to Y4 together to obtain Y. c ;
[0086] The first grouped state space layer is used for Y cstate space modeling is performed;
[0087] The second group state space layer is used to perform state space modeling on the output of the first group state space layer to obtain
[0088] That is, the deep encoding layer first dimensionally expands the four different time scale speech embeddings, then regards them as data after grouping of the channels of an image, and then performs two rounds of state space modeling.
[0089] Of course, the above process of the deep encoding layer can be described in a formula as follows:
[0090]
[0091] In the formula, expand(.) represents the dimension expansion process; concat(.) represents the concatenation process; and GroupMamba(.) represents the state space modeling process.
[0092] It should be noted that each round of state space modeling is performed by first using four parameter sharing layers (referred to as VSSSBlock, the core component of which is Mamba) to cut and expand modeling from four different directions - which is equivalent to modeling long-range dependencies in different time windows with linear complexity to extract deeper and more complex speech features, then using attention to adaptively determine different weights for feature fusion to perform weighted optimization, and finally performing a series of processes such as layer normalization, linear transformation, residual connection, etc. to reduce the loss of speech information.
[0093] Referring to Figure 3 , the group state space layer includes: one separation layer, four parameter sharing layers, one sub-concatenation layer, one layer normalization layer, one feedforward layer, one product layer, and two superposition layers.
[0094] For the group state space layer, it has one input IN and one output OUT.
[0095] As shown in Figure 3 , in any group state space layer:
[0096] The separation layer is used to re-separate the input IN of the group state space layer into four speech embeddings in1-in4;
[0097] The first parameter sharing layer is used to cut and expand model in1 from left to right;
[0098] The second parameter sharing layer is used to cut and expand model in2 from top to bottom;
[0099] The third parameter sharing layer is used to cut and expand model in3 from bottom to top;
[0100] The 4th parameter sharing layer is used to chunk and unfold modeling in4 from right to left;
[0101] The weight calculation layer is used to calculate the attention weight W of IN;
[0102] The sub-splicing layer is used to splice the outputs of the 1st parameter sharing layer, the 2nd parameter sharing layer, the 3rd parameter sharing layer and the 4th parameter sharing layer;
[0103] The product layer is used to multiply the output of the sub-splicing layer with W;
[0104] The 1st superposition layer is used to superimpose the output of the product layer with IN;
[0105] The layer normalization layer is used to perform layer normalization on the output of the 1st superposition layer;
[0106] The feedforward layer is used to perform linear transformation on the output of the layer normalization layer;
[0107] The 2nd superposition layer is used to superimpose the output of the feedforward layer with IN to obtain the output OUT of the grouping state space layer.
[0108] Of course, the above process of the grouping state space layer can be described by formula as follows:
[0109] chunk(IN)={in1,in2,in3,in4};
[0110]
[0111] In the formula, chunk(.) represents the separation process; VSSSBlock(.) represents the chunking and unfolding modeling process; CrossScan i represents the i-th processing direction (i=1-4 respectively corresponds to from left to right, from top to bottom, from bottom to top, from right to left); ⊙ represents multiplication; LN(.) represents layer normalization; FFN(.) represents linear transformation process.
[0112] It should be noted that the IN of the 1st grouping state space layer is Y c , and in1-in4 are Y1-Y4 respectively.
[0113] 104、The dimension adjustment layer is designed to adjust the dimension of .
[0114] As shown in Figure 2 , the dimension adjustment layer can be designed to include: 1 linear layer, 1 dimension compression layer, 1 one-dimensional convolution layer.
[0115] In the dimension adjustment layer:
[0116] The linear layer is configured to perform linear processing on the output of the first graph convolution layer; The dimension compression layer is configured to perform dimension compression on the output of the linear layer;
[0117] The one-dimensional convolution layer is configured to perform one-dimensional convolution on the output of the dimension compression layer to obtain the deep time feature
[0118]
[0119] Of course, the above process of the dimension adjustment layer can be described by formula as follows:
[0120]
[0121] In the formula, Conv1d(.) represents the one-dimensional convolution processing process; squeeze(.) represents the dimension compression process; linear(.) represents the linear processing process; N e represents the number of EEG channels.
[0122] II. The EEG encoding unit is configured to first extract the EEG feature E from the EEG', and then fill E to obtain the EEG feature with the same time dimension as the EEG'
[0123] Referring to Figure 4 , the EEG encoding unit includes: 3 graph convolution layers, 1 channel normalization layer, 3 residual blocks, 1 one-dimensional convolution layer, and 1 padding layer.
[0124] As shown in Figure 4 , in the EEG encoding unit:
[0125] The first graph convolution layer is configured to perform graph convolution processing on the EEG';
[0126] The second graph convolution layer is configured to perform graph convolution processing on the output of the first graph convolution layer;
[0127] The third graph convolution layer is configured to perform graph convolution processing on the output of the second graph convolution layer;
[0128] The channel normalization layer is configured to perform channel normalization processing on the output of the third graph convolution layer;
[0129] The first residual block is configured to perform deep time feature extraction on the output of the channel normalization layer;
[0130] The second residual block is configured to perform deep time feature extraction on the output of the first residual block;
[0131] The third residual block is configured to perform deep time feature extraction on the output of the second residual block;
[0132] The one-dimensional convolution layer is used to perform one-dimensional convolution processing on the output of the third residual block to obtain E;
[0133] The padding layer is used to pad E in the time dimension to obtain
[0134] Of course, the above process of the electroencephalogram encoding unit can be described by formula as follows:
[0135]
[0136] In the formula, T e <T s ; T e represents the time dimension size of E; GCN(.) represents the graph convolution processing process; CN(.) represents the channel normalization processing process; Res(.) represents the residual block processing process; Conv1d(.) represents the one-dimensional convolution processing process; and Padding(.) represents the padding operation.
[0137] It should be noted that the graph convolution layer is suggested to use Chebyshev polynomials as the convolution kernel to simplify the calculation.
[0138] The residual block can be designed as a ResNet network, or can be designed to include two one-dimensional convolution layers, each followed by a batch normalization layer, a PReLU activation function layer, and after the second batch normalization layer, a residual connection is introduced, and after the second PReLU activation function layer, a maximum pooling layer is followed.
[0139] Therefore, the processing process of the residual block can be described by formula as follows:
[0140] N out =Maxpool(PReLU(N in +BN(Conv1d(PReLU(BN(Conv1d(N in )))))));
[0141] In the formula, Maxpool(.) represents the maximum pooling processing process; BN(.) represents the batch normalization processing process; PReLU(.) represents the PReLU activation function; Conv1d(.) represents the one-dimensional convolution processing process; N in represents the input of the residual block; and N out represents the output of the residual block.
[0142] III, the speaker extraction unit is used to first process and to obtain the target speech mask M, and then multiply M and X c to obtain the target speech feature S.
[0143] From the above records, we can know that: That is, both are identical in three dimensions, so subsequent operations can be performed.
[0144] See Figure 5 The speaker extraction unit includes: one modality fusion layer, one separation network layer, and one product layer.
[0145] like Figure 5 As shown, in the speaker extraction section:
[0146] 301. The modality fusion layer is used to combine multiple layers of cross-attention mechanism. and Fuse into feature vector X fuse .
[0147] In general, and Through three layers of information fusion processing streams (including speech processing stream and EEG processing stream) consisting of cross-attention, residual connections, and group normalization, bidirectional coupling of features is achieved during the cross-attention process to obtain X. fuse Thus X fuse It also includes both speech and EEG information, as well as their complementary features.
[0148] It should be noted that in the speech processing stream, the query Q for cross-attention is speech data, while the key K and value V are both EEG data; the opposite is true in the EEG processing stream.
[0149] See Figure 6 The modality fusion layer includes: 3 attention processing layers, 2 external stacking layers, 1 splicing layer, and 1 one-dimensional convolutional layer.
[0150] In the modality fusion layer:
[0151] The first attention processing layer is used to process data through a cross-attention mechanism. and Processing to obtain speech embedding and EEG characteristics
[0152] The second attention processing layer is used to process attention through a cross-attention mechanism. and Processing to obtain speech embedding and EEG characteristics
[0153] The third attention processing layer is used to process data through a cross-attention mechanism. and Processing to obtain speech embedding and EEG characteristics
[0154] The first external stacking layer is configured to stack ;
[0155] The second external stacking layer is configured to stack ;
[0156] The splicing layer is configured to splice the output of the first external stacking layer and the output of the second external stacking layer;
[0157] The one-dimensional convolution layer is configured to perform one-dimensional convolution processing on the output of the splicing layer to obtain X fuse .
[0158] Of course, the above process of the modal fusion layer can be described in a formula as follows:
[0159]
[0160] In the formula, concat(.) represents the splicing processing process; and Conv1d(.) represents the one-dimensional convolution processing process.
[0161] It should be noted that the attention processing layer can adopt a cross-attention network design, and can refer to Figure 6 and be designed to include: 2 cross-attention layers, 2 internal stacking layers, and 2 group normalization layers.
[0162] The attention processing layer has 2 inputs Input1-Input2 and 2 outputs Output1-Output2.
[0163] The first cross-attention layer is configured to perform cross-attention calculation by taking Input1 as a query Q and Input2 as a key K and a value V.
[0164] The first internal stacking layer is configured to stack Input1 and the output of the first cross-attention layer.
[0165] The first group normalization layer is configured to perform group normalization processing on the output of the first internal stacking layer to obtain Output1.
[0166] The second cross-attention layer is configured to perform cross-attention calculation by taking Input2 as a query Q and Input1 as a key K and a value V.
[0167] The second internal stacking layer is configured to stack Input1 and the output of the second cross-attention layer.
[0168] The second group normalization layer is configured to perform group normalization processing on the output of the second internal stacking layer to obtain Output2.
[0169] So we have:
[0170] Input1 of the 1st attention processing layer is Input2 is Output1 is Output2 is
[0171] Input1 of the 2nd attention processing layer is Input2 is Output1 is Output2 is
[0172] Input1 of the 3rd attention processing layer is Input2 is Output1 is Output2 is
[0173] Of course, the above process of the attention processing layer can be described in formula as:
[0174] Output1 = GN (Input1 + CrossAttention (Input1, Input2, Input2)) ;
[0175] Output2 = GN (Input2 + CrossAttention (Input2, Input1, Input1)) ;
[0176] In the formula, CrossAttention(.) represents the cross-attention calculation process; GN(.) represents the group normalization processing process.
[0177] 302, the separation network layer is used to block X fuse and respectively process the fusion features in and between the blocks to obtain M.
[0178] The separation network layer can adopt DPRNN (dual path recurrent) network, conv-tasnet network, sepformer network, mossformer network, etc. In this embodiment 1, it is recommended to adopt DPRNN network, which can effectively capture the time dependence compared with the convolutional neural network, and is more suitable for processing X fuse This kind of sequential data.
[0179] It should be noted that,
[0180] 303, the product layer is used to multiply M with X cS = S * S
[0181] That is,
[0182] IV, the speech decoding part is used to reconstruct from S
[0183] Referring to Figure 5 , the speech decoding part includes: 1 one-dimensional transpose convolution layer.
[0184] As Figure 5 shown in the speech decoding part:
[0185] The one-dimensional transpose convolution layer is used to one-dimensionally transpose convolution S to obtain
[0186] Of course, the above process of the speech decoding part can be described by formula:
[0187]
[0188] In the formula, deconv1D(.) represents the one-dimensional transpose convolution layer processing process.
[0189] (II) Then, the specific structure of each part of the network in the training stage is described:
[0190] In the training stage, a sample data set with real labels (which includes mixed speech samples and electroencephalogram data samples) is used; After the sample data set is preprocessed according to step two, it is divided into a training set and a validation set; Based on the training set, the multi-scale multi-modal alignment network is trained for multiple rounds, and based on the validation set, the multi-scale multi-modal alignment network after each round of training is verified until the best performance of the multi-scale multi-modal alignment network is selected and used as the trained multi-scale multi-modal alignment network.
[0191] Wherein, the loss function L total used by the multi-scale multi-modal alignment network in network training is:
[0192] L total = L SI-SDR + α * L InfoNCE ;
[0193] In the formula, L SI-SDR represents the scale-invariant signal distortion ratio loss constructed based on the output of the speech decoding part; L InfoNCE represents the noise contrast estimation loss constructed based on the output of the speech encoding part and the electroencephalogram encoding part; and α represents the loss weight coefficient.
[0194] 1. For the speech encoding part, its working mode is similar to that in the prediction stage, only the input is changed to sample data.
[0195] The same applies to the brain electrical coding unit, the speaker extraction unit, and the speech decoding unit, and will not be described again.
[0196] L SI-SDR That is, the speech decoding unit output is compared with the real label target speech, and the formula is:
[0197]
[0198] In the formula, X target represents the scale-invariant target speaker speech, X res represents the residual speech; represents the output of the speech decoding unit; s0 represents the real label target speech.
[0199] 2. The modal alignment unit works in the training stage, which is used to align the outputs of the speech coding unit and the brain electrical coding unit in time steps based on contrastive learning when the network is trained, and calculate the noise contrast estimation loss L InfoNCE .
[0200] Referring to Figure 7 , the modal alignment unit includes a dimension transformation layer and a contrastive learning layer.
[0201] In the modal alignment unit:
[0202] The dimension transformation layer is used to compress the dimensions of the outputs of the speech coding unit and the brain electrical coding unit, respectively;
[0203] The contrastive learning layer is used to calculate L InfoNCE based on the outputs of the dimension transformation layer.
[0204] Specifically, referring to the above description, the specifications of the outputs of the speech coding unit and the brain electrical coding unit are three-dimensional— Then, the dimension transformation layer compresses the outputs of the speech coding unit and the brain electrical coding unit into two dimensions— That is, the dimension transformation layer has two outputs: 1. Speech sequence (i.e., the output of the compressed speech coding unit); 2. Electroencephalogram sequence (i.e., the output of the compressed brain electrical coding unit).
[0205] For the dimension transformation layer, the speech sequence output can be represented as {x1, x2, …, x n}, and the electroencephalogram sequence can be represented as {e1, e2, …, e n}; n represents the sequence length. The electroencephalogram segment e j is regarded as a query vector q j , and the speech segment x j corresponding to it in the time sequence is considered as a positive sample, and the two form a positive sample pair (qj ,x j + Simultaneously, other speech segments from the same batch (x) g (g≠j) is considered as q j From the negative samples, we obtain n-1 negative sample pairs (q j ,x g - ).
[0206] To reduce modal misalignment at the temporal level and achieve contrastive learning, the contrastive learning layer uses a noisy contrastive estimation loss L. InfoNCE The goal of optimization is to maximize the similarity between positive samples and minimize the similarity between negative samples.
[0207] Among them, L InfoNCE The expression is:
[0208]
[0209] In the formula, x k represents any sample in the current batch; sim(.) represents the cosine similarity calculation between two samples; τ represents the temperature coefficient.
[0210] In summary, L InfoNCE L SI-SDR Construct L total The parameters of the entire network are jointly optimized to ensure time synchronization, thereby guaranteeing and improving the overall performance of the network in target speaker extraction.
[0211] Finally, it should be noted that:
[0212] The background section mentions two problems with existing models: the first is that they only address the surface result, and discovering the underlying cause—the lack of temporal alignment between speech and EEG—is difficult; the second problem, also a surface result, is extracting more comprehensive and richer speech features. Therefore, designing a method that addresses both problems is even more challenging.
[0213] Simulation verification
[0214] This embodiment 1 demonstrates the simulation verification of the method proposed above:
[0215] The existing 6 networks (Mixture, BESD, UBESD, BASEN, NeuroHeed, MSFNet) are introduced, together with the multi-scale multi-modal alignment network used in the method (referred to as ours), for performance comparison on three brain-controlled target speaker extraction datasets (Cocktail Party, AVED, MM-AAD) - the selection of indicators is SDR (the larger the better), SI-SDR (the larger the better), STOI (the larger the better), ESTOI (the larger the better), PESQ (the larger the better), and the results are shown in Table 1.
[0216] Table 1 Network performance comparison
[0217]
[0218] As can be seen from Table 1, the experiments on the above two known datasets show that the multi-scale multi-modal alignment network used in the method is optimal in the five indicators, proving the effectiveness and superiority of the method.
[0219] Embodiment 2
[0220] This embodiment 2 provides a target speaker extraction system based on a multi-scale multi-modal alignment network, which uses the target speaker extraction method based on a multi-scale multi-modal alignment network disclosed in embodiment 1.
[0221] The target speaker extraction system based on a multi-scale multi-modal alignment network comprises a data acquisition module, a pre-processing module and a network processing module.
[0222] The data acquisition module is used to acquire mixed speech Mixture in a multi-speaker scene and electroencephalogram data EEG generated by Mixture stimulation. The pre-processing module is used to downsample Mixture to reduce computational complexity and obtain mixed speech Mixture'; on the other hand, EEG is pre-processed to remove noise and then upsampled to match Mixture', obtaining EEG'. The network processing module is used to input Mixture', EEG' into the trained multi-scale multi-modal alignment network for processing to obtain target speaker speech
[0223] Since the system uses the target speaker extraction method based on a multi-scale multi-modal alignment network in embodiment 1, it also has the same effect, which is not repeated here.
[0224] Embodiment 3
[0225] The embodiment 3 discloses a computer device, comprising a memory and a processor, the memory stores a computer program, and the processor executes the computer program to realize the steps of the target speaker extraction method based on the multi-scale multi-modal alignment network disclosed in the embodiment 1.
[0226] The embodiment 3 further discloses a readable storage medium, which stores computer program instructions, and the computer program instructions are read and executed by a processor to perform the steps of the target speaker extraction method based on the multi-scale multi-modal alignment network disclosed in the embodiment 1.
[0227] The embodiment 3 further discloses a computer program product, comprising a computer program. The computer program is executed by a processor to realize the steps of the target speaker extraction method based on the multi-scale multi-modal alignment network disclosed in the embodiment 1.
[0228] The above-mentioned embodiments only express several embodiments of the present application, which are described in detail and specifically, but cannot be understood as the limitation of the patent scope of the present application. It should be pointed out that, for ordinary skilled in the art, several modifications and improvements can be made without departing from the concept of the present application, which all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A method for target speaker extraction based on a multi-scale, multi-modal alignment network, characterized in that, include: Step 1: Acquire mixed speech data (Mixture) and EEG data generated by Mixture stimulation in multi-speaker scenarios. Step 2: Downsample the Mixture to reduce computation and obtain the mixed speech Mixture'. The EEG data is first preprocessed to remove noise, and then upsampled accordingly to match the Mixture, thus obtaining the EEG data. Step 3: Input Mixture and EEG into the trained multi-scale multimodal alignment network for processing to obtain the target speaker's speech. The multi-scale, multi-modal alignment network includes: The speech coding unit first obtains speech embeddings X1 to X4 at four time scales from the Mixture through multi-scale coding, and then aggregates X1 to X4 into a speech embedding X. c Then, through multi-directional deep coding, richer speech embeddings are extracted from X1 to X4. Next to Perform dimensional adjustments to obtain speech embedding The EEG coding unit is used to first extract the EEG feature E from EEG', and then fill in E to obtain the corresponding EEG feature E. EEG features with the same time dimension Speaker extraction section, which is used to first extract the speaker's information. and Process to obtain the target speech mask M, then combine M with X c Multiply the features to obtain the target speech feature S; The speech decoding unit is used to decode and reconstruct the speech from S. as well as The modality alignment unit is used to align the outputs of the speech encoder and EEG encoder at time steps based on contrastive learning during network training, and to calculate the noise contrastive estimation loss L. InfoNCE ; Among them, the loss function L used in the multi-scale multimodal alignment network during network training is total for: L total =L SI-SDR +a*L InfoNCE ; In the formula, L SI-SDR L represents the scale-invariant signal distortion ratio loss constructed based on the output of the speech decoding unit; InfoNCE α represents the loss in noise comparison estimation; α represents the loss weighting coefficient.
2. The target speaker extraction method based on a multi-scale multimodal alignment network according to claim 1, characterized in that, The speech coding unit includes: one multi-scale coding layer, one pre-aggregation layer, one deep coding layer, and one dimension adjustment layer; The multi-scale coding layer comprises four one-dimensional convolutional layers and four ReLU activation function layers. Within the multi-scale coding layer: the first one-dimensional convolutional layer encodes the Mixture' using a one-dimensional convolution with a short time window; the first ReLU activation function layer processes the output of the first one-dimensional convolutional layer using the ReLU activation function to obtain the short-time-scale speech embedding X1; the second one-dimensional convolutional layer encodes the Mixture' using a one-dimensional convolution with a second-shortest time window; and the second ReLU activation function layer processes the output of the second one-dimensional convolutional layer using the ReLU activation function. The process is performed to obtain the speech embedding X2 at the next shortest time scale; the third one-dimensional convolutional layer is used to encode Mixture' through one-dimensional convolution with a next long time window; the third ReLU activation function layer is used to process the output of the third one-dimensional convolutional layer through the ReLU activation function to obtain the speech embedding X3 at the next long time scale; the fourth one-dimensional convolutional layer is used to encode Mixture' through one-dimensional convolution with a long time window; the fourth ReLU activation function layer is used to process the output of the fourth one-dimensional convolutional layer through the ReLU activation function to obtain the speech embedding X4 at the long time scale; The pre-aggregation layer includes: one pre-stitching layer and one one-dimensional convolutional layer; in the pre-aggregation layer: the pre-stitching layer is used to stitch X1 to X4 together; the one-dimensional convolutional layer is used to perform one-dimensional convolution on the output of the pre-stitching layer to obtain X. c ; The deep coding layer consists of: 4 dimension expansion layers, 1 post-concatenation layer, and 2 grouped state space layers; in the deep coding layer: the i-th dimension expansion layer is used to process X i Perform dimensional expansion to obtain the expanded speech embedding Y i i∈{1,2,3,4}; the post-concatenation layer is used to concatenate Y1~Y4 to obtain Y c The first grouped state space layer is used for Y c State-space modeling is performed; the second grouped state-space layer is used to perform state-space modeling on the output of the first grouped state-space layer to obtain... The dimension adjustment layer consists of: one linear layer, one dimension compression layer, and one one-dimensional convolutional layer; within the dimension adjustment layer: the linear layer is used to... Linear processing is performed; the dimensionality compression layer is used to compress the dimensionality of the output of the linear layer; the one-dimensional convolutional layer is used to perform one-dimensional convolution on the output of the dimensionality compression layer to obtain...
3. The target speaker extraction method based on a multi-scale multimodal alignment network according to claim 2, characterized in that, The grouped state space layer includes: 1 separation layer, 4 parameter sharing layers, 1 sub-stitching layer, 1 layer normalization layer, 1 feedforward layer, 1 product layer, and 2 stacking layers; In any grouped state space layer: The separation layer is used to re-separate the input IN of the grouped state space layer into four speech embeddings in1 to in4; The first parameter-sharing layer is used to slice in1 from left to right and expand it for modeling. The second parameter-sharing layer is used to slice in2 from top to bottom and expand it for modeling; The third parameter-sharing layer is used to slice in3 from bottom to top and expand it for modeling. The fourth parameter-sharing layer is used to slice in4 from right to left and expand it for modeling. The weight calculation layer is used to calculate the attention weight W of IN; The sub-stitching layer is used to stitch together the outputs of the first parameter sharing layer, the second parameter sharing layer, the third parameter sharing layer, and the fourth parameter sharing layer; The product layer is used to multiply the output of the sub-stitching layer with W; The first stacking layer is used to stack the output of the product layer with IN; The layer normalization layer is used to perform layer normalization processing on the output of the first stacked layer; The feedforward layer is used to perform a linear transformation on the output of the layer normalization layer; The second overlay layer is used to overlay the output of the feedforward layer with IN to obtain the output OUT of the grouped state space layer; Among them, IN of the first grouped state space layer is Y. c in1 to in4 are Y1 to Y4 respectively.
4. The target speaker extraction method based on a multi-scale multimodal alignment network according to claim 1, characterized in that, The EEG coding unit includes: 3 graph convolutional layers, 1 channel normalization layer, 3 residual blocks, 1 one-dimensional convolutional layer, and 1 filler layer; In the EEG coding section: The first graph convolutional layer is used to perform graph convolution processing on EEG'; The second graph convolutional layer is used to perform graph convolution processing on the output of the first graph convolutional layer; The third graph convolutional layer is used to perform graph convolution processing on the output of the second graph convolutional layer; The channel normalization layer is used to normalize the channel output of the third graph convolutional layer; The first residual block is used to extract deep temporal features from the output of the channel normalization layer; The second residual block is used to extract deep temporal features from the output of the first residual block; The third residual block is used to extract deep temporal features from the output of the second residual block; A one-dimensional convolutional layer is used to perform one-dimensional convolution on the output of the third residual block to obtain E; The fill layer is used to fill E in the time dimension to obtain 5. The target speaker extraction method based on a multi-scale multimodal alignment network according to claim 1, characterized in that, The speaker extraction unit consists of: one modality fusion layer, one separation network layer, and one product layer; In the speaker extraction section: Modality fusion layer is used to combine multiple layers of cross-attention mechanism. and Fuse into feature vector X fuse ; Separate network layers are used for X fuse The blocks are divided, and the fusion features are processed separately from within and between blocks to obtain M; Product layers are used to connect M and X c Multiply them to get S.
6. The target speaker extraction method based on a multi-scale multimodal alignment network according to claim 5, characterized in that, The modality fusion layer includes: 3 attention processing layers, 2 external stacking layers, 1 splicing layer, and 1 one-dimensional convolutional layer; In the modality fusion layer: The first attention processing layer is used to process data through a cross-attention mechanism. and Processing to obtain speech embedding and EEG characteristics The second attention processing layer is used to process attention through a cross-attention mechanism. and Processing to obtain speech embedding and EEG characteristics The third attention processing layer is used to process data through a cross-attention mechanism. and Processing to obtain speech embedding and EEG characteristics The first external overlay layer is used to... Superimpose; The second external overlay layer is used to... Superimpose; The splicing layer is used to... The outputs of the first external overlay layer and the outputs of the second external overlay layer are spliced together; One-dimensional convolutional layers are used to perform one-dimensional convolution processing on the output of the splicing layer to obtain X. fuse .
7. The target speaker extraction method based on a multi-scale multimodal alignment network according to claim 1, characterized in that, The speech decoding unit includes: one one-dimensional transposed convolutional layer; In the voice decoding department: A one-dimensional transposed convolutional layer is used to perform a one-dimensional transposed convolution on S to obtain 8. The target speaker extraction method based on a multi-scale multimodal alignment network according to claim 1, characterized in that, The modal alignment layer includes: one dimensional transformation layer and one contrast learning layer; In the modal alignment part: The dimension transformation layer is used to perform dimension compression on the output of the speech coding unit and the output of the EEG coding unit, respectively. The contrastive learning layer is used to calculate L based on the output of the dimension transformation layer. InfoNCE .
9. A target speaker extraction system based on a multi-scale, multi-modal alignment network, characterized in that, It uses the target speaker extraction method based on a multi-scale multimodal alignment network as described in any one of claims 1-8; The target speaker extraction system based on a multi-scale, multi-modal alignment network includes: The data acquisition module is used to acquire mixed speech data (Mixture) in multi-speaker scenarios, as well as EEG data generated by Mixture stimulation. The pre-processing module is used to downsample the Mixture to reduce the amount of computation and obtain the mixed speech Mixture'; on the other hand, it is used to preprocess the EEG to remove noise and then upsample it to match the Mixture' to obtain the EEG data'. as well as The network processing module is used to process the Mixture and EEG inputs into the trained multi-scale multimodal alignment network to obtain the target speaker's speech.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the target speaker extraction method based on a multi-scale multimodal alignment network as described in any one of claims 1-8.