Voice separation method with noise perception based on comparative learning and causal attention mechanism

By introducing a speech separation method of the causal attention and noise perception comparison learning module, the model parameter quantity and calculation efficiency are optimized, real-time speech separation on terminal devices is realized, and the separation effect is improved in complex noise environments.

CN120299472APending Publication Date: 2025-07-11SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510355845.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing deep learning speech separation model has high dependence on computing resources, which is difficult to apply in real time on terminal devices, and insufficient processing of complex background noise, limiting the online real-time application and separation effect.

Method used

A speech separation method based on contrast learning and causal attention mechanism is adopted. Through the encoder, mask network and decoder framework, a causal attention module and a noise-perceived contrast learning module are introduced to optimize the model parameter quantity and calculation efficiency, and the speech signal is separated in real time.

Benefits of technology

In the noise-containing mixed speech separation task, the separation effect is better than that of mainstream algorithms, and it has real-time application, which reduces the computing resource requirements and improves the processing ability of noise characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299472A_ABST
    Figure CN120299472A_ABST
Patent Text Reader

Abstract

The invention discloses a voice separation method with noise perception based on comparative learning and a causal attention mechanism, and the method comprises the following steps: carrying out the modeling of a pure source voice signal and a noise-containing mixed voice based on an encoder, and generating the learning domain feature representation of the source voice signal and the noise-containing mixed voice; mask estimation is carried out on the source voice signals and the noise signals based on a mask network, and estimation feature representations of different sound sources are generated; using learning domain feature representation of the source voice, the estimated voice and the estimated noise to optimize overall network parameters, and constructing a contrast learning module for realizing noise perception; and constructing a decoding network, and restoring a corresponding time domain estimation signal based on estimation feature representation to realize voice separation. According to the method, the real-time applicability is achieved while the separation effect exceeding the mainstream algorithm is obtained in the noise-containing mixed voice separation task with low parameter quantity and higher calculation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech separation, and in particular to a speech separation method with noise perception based on contrast learning and causal attention mechanism. Background Art

[0002] The speech separation task can be roughly divided into directions such as multi-speaker separation, noise reduction, and reverberation removal. Its goal is to separate each speech from the mixed speech where multiple sound sources overlap and may be mixed with noise, so as to help extract the target speech. As the front-end part of the speech signal processing system, the result of speech separation directly affects the performance of subsequent downstream tasks such as speech recognition. Moreover, with the continuous upgrade of hardware devices, the speech interaction function is increasingly integrated into daily devices. The effective utilization of speech when sound sources are overlapped and affected by noise can greatly improve the user experience, and speech separation is the key technology to solve this problem.

[0003] With the development of deep learning technology, the deep learning-based speech separation method has exceeded the traditional method in terms of performance and other aspects and has become the mainstream in the field of speech separation. Among the deep learning-based methods, different from the past separation methods based on fixed time-frequency analysis means, Luo Yi et al. first put the processing domain of speech separation onto the pure time domain, greatly exceeding the performance of time-frequency methods. Since the separation is achieved in the time domain, this network is called TasNet. The following year, Luo Yi et al. proposed Cov-TasNet, which directly models the speech signal using a convolutional network and achieved better performance. The proposed encoder-mask network-decoder framework had a profound impact on subsequent research. Facing long sequences, Luo Yi et al. proposed a dual-path recurrent neural network DPRNN in the mask network based on the encoder-mask network-decoder, emphasizing the importance of long-term dependence for long sequence modeling. The dual-path processing framework of DPRNN also serves as the basic framework for subsequent speech separation research. The trend of a series of subsequent research is to continuously improve the separation performance of deep neural networks by gradually increasing the model and more and more data volume. This trend has also led to the popularity of large-scale neural networks such as HuBERT and wav2vec2.0 in the field of speech processing.

[0004] However, the expansion of the scale of neural network models has led to an increased dependence on computing resources, which has increased the inference cost. Moreover, it is difficult to deploy ultra-large models on terminal devices, and relying on cloud processing of data may cause privacy and security issues. In addition, most existing separation models do not introduce a causal mechanism and are only applicable to offline scenarios. They cannot separate the speech data of each frame in real time while the sound signal is incoming, which limits the online real-time application on the device side. Finally, the existing speech separation models still lack in dealing with noise features. Especially in complex background noise, the separation effect of the model often fails to meet the requirements of practical applications. Making better use of noise features has become the key bottleneck for performance improvement. Summary of the Invention

[0005] In order to overcome the defects and deficiencies of the existing technology, the present invention provides a speech separation method based on contrast learning and causal attention mechanism with noise perception. While regarding noise as a sound source, the present invention introduces a causal attention module and a noise perception contrast learning module. With a relatively small number of parameters and high computational efficiency, this network achieves a separation effect superior to mainstream algorithms in the task of separating noisy mixed speech, and at the same time has real-time applicability.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] The present invention provides a speech separation method based on contrast learning and causal attention mechanism with noise perception, including the following steps:

[0008] Construct an encoder to model the pure source speech signal and the noisy mixed speech signal based on the encoder, and generate the learning domain feature representation of the source speech signal and the learning domain feature representation of the noisy mixed speech.

[0009] Construct a mask network to estimate the masks for the source speech signal and the noise signal, and multiply the output of the mask network element-wise with the learning domain feature representation of the noisy mixed speech to obtain the estimated feature representations of different sound sources.

[0010] Construct a contrast learning module. For the source speech feature representations of each sound source, the corresponding estimated speech feature representations of each sound source, and the estimated noise feature representations, respectively select positive samples, query samples, and negative samples, calculate the contrast loss, and after weighting this loss, superimpose the scale-invariant signal-to-noise ratio loss to obtain the overall loss function of the network, which is used as the joint optimization objective for model training. Calculate the gradients of each loss function during the backpropagation process and jointly optimize all parameters.

[0011] Construct a decoding network to restore the corresponding time-domain estimated signal based on the estimated feature representation to achieve speech separation.

[0012] As a preferred technical solution, the encoder uses a one-dimensional convolutional network for each source speech signal s kModel the time-domain mixed speech signal composed of the clean speech signal and the noise signal to obtain the feature representation of the noisy mixed speech in the learning domain and use it as the input of the mask network;

[0013] Meanwhile, the encoder also models each source speech signal separately to obtain the corresponding feature representation in the learning domain where F represents the dimension of the feature vector of h, and T′ is the time length after convolution.

[0014] As a preferred technical solution, the mask network includes a feature block processing module and a causal attention module;

[0015] The feature block processing module divides the feature representation of the noisy mixed speech along the time dimension into blocks, and then concatenates the divided results along the block dimension to obtain a feature vector h′, and inputs it into the causal attention module;

[0016] The causal attention module obtains the latent feature representation of the speech based on the segmented Transformer network and the memory Transformer network;

[0017] The feature block processing module performs block reshaping and mask generation on the latent feature representation of the speech.

[0018] As a preferred technical solution, the causal attention module has a total of L layers, where the structures of the first L - 1 layers are: the segmented Transformer network is combined with the mean calculation in the time dimension and the memory Transformer network through skip connections, and the structure of the Lth layer only contains the segmented Transformer network;

[0019] Set the input of the lth layer to the feature vector h l ′, then the processing of this layer is expressed as:

[0020] I l1 = segTransformer(h l ′)

[0021]

[0022] I l3 = memTransformer(I l2 )

[0023] I l4 = I l1 + I l3

[0024] where segTransformer(·), The segmented Transformer network, taking the mean in the time dimension, and the memory Transformer network are denoted as memTransformer(·) respectively;

[0025] The feature vector h l ′ is obtained as a feature vector through the segmented Transformer For the feature vector I l1 The average is taken in the time dimension to obtain the feature vector I l2 and the feature vector I l1 is added to the memory cell I l3 to obtain the feature vector I l4 The feature vector I l4 is used as the input of the (l + 1)-th layer, and after iterative processing, the latent feature representation h″ of the speech is obtained.

[0026] As a preferred technical solution, the feature block processing module performs block reshaping and mask generation on the latent feature representation of the speech. The latent feature representation of the speech passes through PReLU and a one-dimensional convolutional layer, and through a reshaping operation of re-splicing on the time axis, the time-step features h″′ corresponding to K source sound sources and noise are obtained, and then through the ReLU non-linear function for estimation, the masks source corresponding to each of the K sound sources and the mask m n corresponding to the noise are obtained.

[0027] As a preferred technical solution, in the Transformer network of the causal attention module, the input of the Transformer network based on causal attention is set as the feature vector z, and a causal mask is generated for the feature vector z;

[0028] The processing process entering the Transformer network is expressed as:

[0029] z′ = z + e pos

[0030] z″ = Multiheadattention(norm(z′))

[0031] z″′ = z′ + z″ + FFN(norm(z′ + z″))

[0032] z″″ = permute(BatchNorm(permute(z″′)))

[0033] z″″′ = z + z″″

[0034] where e pos, norm(·), Multiheadattention(·), FFN(·), permute(·), and BatchNorm(·) represent relative position encoding, layer normalization, multi-head attention mechanism, feed-forward neural network, transpose operation for swapping dimensions, and batch normalization, respectively;

[0035] Add the relative position encoding e to the feature vector z pos Obtain the feature vector z′. Applying layer normalization and the multi-head attention mechanism to the feature vector z′ can obtain the feature vector z″;

[0036] Apply layer normalization, feed-forward neural network, and two residual connections respectively connected to the feature vector z′ and the feature vector z″ to the feature vector z′ and the feature vector z″ in sequence to obtain the feature vector z″′;

[0037] Transpose and swap the dimensions of the feature vector z″′ and then normalize it, and then transform it back to the original dimension to obtain the feature vector z″″. Through the skip connection, the final output feature vector z″″′ of the network is obtained.

[0038] As a preferred technical solution, the contrastive learning module includes a sampler and a resampler;

[0039] The sampler is a sampling network obtained by cascading two-dimensional convolution, ReLU, and two-dimensional convolution. The resampler is composed of a fully connected layer, ReLU, and a fully connected layer in cascade.

[0040] As a preferred technical solution, the sampler performs multiple random samplings on the source speech feature representation, the corresponding estimated speech feature representation of the source speech, and the estimated noise feature representation respectively. The same rule followed each time is: randomly take 1 local block at the same position of the feature maps of the source speech feature representation, the corresponding estimated speech feature representation of the source speech, and the estimated noise feature representation, and at the same time take M - 1 other local blocks at other positions of the feature map of the estimated noise feature representation;

[0041] The sampling process of the sampler and the resampling process of the resampler are expressed as:

[0042]

[0043] Among them, respectively represent the positive sample, query sample, and negative sample obtained by the sampler from and The i-th sampling. H is the size of the convolutional kernel in the sampler;

[0044] K s All obtained by Projected into a three-dimensional embedding space through a reshaper, while using L2 normalization, and finally obtaining a set of positive samples, query samples, and negative sample feature vectors for contrastive learning

[0045] Positive samples, query samples, and negative sample feature vectors obtained by the reshaper Calculate the contrastive loss

[0046] As a preferred technical solution, calculate the contrastive loss, weight this loss and then superimpose the scale-invariant signal-to-noise ratio loss to obtain the overall loss function of the network, which is specifically expressed as:

[0047] Loss total =Loss SISNR +βLoss cll

[0048]

[0049] Among them, Loss total is the overall loss function of the network, Loss SI-SNR represents the scale-invariant signal-to-noise ratio loss, Loss cll represents the contrastive loss for K source sound sources, β represents the loss weight parameter, K source represents the number of sound sources, s k represents the source speech signal, is its corresponding time-domain estimated signal, <·,·> represents the inner product of two vectors, p positive represents the probability of the positive sample, Loss ce (s k ) represents the cross-entropy loss, represents the positive sample, represents the query sample, represents the negative sample, j∈[1,M], τ represents the temperature coefficient, which is used to adjust the smoothness of the similarity distribution

[0050] As a preferred technical solution, construct a decoding network, and restore the corresponding time-domain estimated signal based on the estimated feature representation to achieve speech separation, specifically including:

[0051] The decoder decodes and restores the estimated feature representations of the source speech signal and the noise respectively through one-dimensional transposed convolution to obtain the time-domain estimated source speech signal and the noise estimated signal. Among them, the estimated feature representations of the source speech signal and the noise are obtained by performing element-wise multiplication on the noisy mixed speech feature representation and the masks corresponding to the sound sources respectively The noise corresponding mask m n is obtained through the above operation

[0052] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0053] Different from the feature block processing method of the traditional dual-path processing flow, the present invention does not perform inter-block overlap on the blocks after block division, and independently processes each small block, reducing the computational amount;

[0054] In the causal attention module of the mask network, a new separation network constructed based on the causal attention mechanism adopts means such as causal masks, normalization, and skip connections to optimize the traditional Transformer-based separation network structure, reduce the model parameter quantity, and optimize the model efficiency;

[0055] Regarding noise as an equivalent sound source, and after the mask network and before the separator in the separation framework of "encoder - mask network - decoder", sampling the feature representations of the clean source speech signal, the separated estimated source speech signal, and the estimated noise signal, and applying the contrast learning strategy to add the contrast loss to the overall network loss to participate in the parameter update after the backpropagation gradient calculation, and using the noise features to achieve noise suppression and optimize the separation performance. While achieving separation results superior to mainstream algorithms in the noisy mixed speech separation task with a relatively low parameter quantity, the present invention also has real-time applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 It is a schematic diagram of the implementation framework of the speech separation method with noise perception based on contrast learning and causal attention mechanism of the present invention;

[0057] Figure 2 It is a schematic diagram of the mask network framework of the present invention;

[0058] Figure 3 It is a schematic diagram of the implementation process of the causal attention module of the present invention;

[0059] Figure 4 It is a schematic diagram of the Transformer network framework based on the causal attention mechanism of the present invention;

[0060] Figure 5 It is a schematic diagram of the framework of the noise perception contrast learning module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0062] Embodiment

[0063] As Figure 1As shown, this embodiment provides a speech separation method with noise perception based on contrastive learning and causal attention mechanism, including the following steps:

[0064] S1: construct an encoder, model the pure source speech signal and the noisy mixed speech based on the encoder, and generate the learning domain feature representation of the source speech signal and the learning domain feature representation of the noisy mixed speech;

[0065] In this embodiment, the encoder uses a one-dimensional convolutional network to encode each pure source speech signal s k The time domain mixed speech signal consisting of a noise signal n Modeling, k∈{1,2,…,K source}, and obtain the learning domain feature representation of noisy mixed speech And used as the input of the mask network. Among them, K source is the number of sound sources in the pure source speech signal. F represents the feature vector dimension of h, which is the same as the number of convolution kernels in the one-dimensional convolution layer. T′ is the time length after convolution, which is related to the convolution kernel size and step size. At the same time, the encoder also performs a separate k Modeling, get s k Learning domain feature representation k∈{1,2,…,K source}, as the input of the subsequent noise-aware contrastive learning module.

[0066] S2: Construct a mask network to perform mask estimation on the pure source speech signal and the noise, and multiply the output of the mask network by the learning domain feature representation of the noisy mixed speech element by element to obtain the estimated feature representation of different sound sources;

[0067] In this embodiment, the mask network takes the encoder output h as input, extracts the acoustic features of the speech, and finally generates K source The masks corresponding to each sound source and the noise corresponding mask m n .

[0068] The mask network includes a feature block processing module and a causal attention module. The feature block processing module is responsible for dividing the speech signal feature representation into blocks as the input of the causal attention module. At the same time, the output of the causal attention module is reshaped and masked. The causal attention module extracts the acoustic features of speech and generates corresponding potential feature representations for different sound sources in the mixed speech.

[0069] like Figure 2 As shown in the figure, the specific implementation process of the combination of the feature block processing module and the causal attention module is as follows:

[0070] The feature block processing module first represents the noisy mixed speech feature Partition along the time dimension into chunks of length C, with adjacent chunks having no overlap, and then concatenate the partition results along the chunk dimension to obtain a feature vector

[0071]

[0072] where N C is the number of chunks after partitioning, which depends on the input speech length and the chunk length. h′ is used as the input to the causal attention module for subsequent processing.

[0073] As Figure 3 shown, the causal attention module has a total of L layers (L ≥ 2), where the structure of the first L - 1 layers is: the segmented Transformer, through skip connections, combines the mean calculation in the time dimension and the memory Transformer. And the structure of the L-th layer only contains the segmented Transformer. Assume the input of the l-th layer is the feature vector h l ′. Then the processing of this layer can be described as the following process:

[0074] I l1 = segTransformer(h l ′)

[0075]

[0076] I l3 = memTransformer(I l2 )

[0077] I l4 = I l1 + I l3

[0078] In the above formulas, segTransformer(·), and memTransformer(·) represent the segmented Transformer, taking the mean in the time dimension, and the memory Transformer respectively. h l ′ can obtain a feature vector through the segmented Transformer Then, take the average of I l1 in the time dimension to obtain the feature vector I l2 , which is used as a cross-segment memory unit for storing and transmitting hidden states. I l2 is then processed by the memory Transformer to obtain the feature vector I l3 . When adding I l1 to the memory unit I l3 , I l3It will broadcast the features along the time axis to I which should have been the output without skip connections l1 and finally obtain I l4 。I l4 will be used as the input of the (l + 1)-th layer. In the causal attention module, the above process is repeated to finally process the input h′ into the potential feature representation of the speech and redeliver it to the feature block processing module. K in the h″ dimension source corresponds to the number of sound sources in the clean source speech signal s k , and “+1” means regarding the noise as a sound source as well.

[0079] As Figure 2 shown, after the causal attention module outputs h″, the feature block processing continues. h″ passes through a PReLU and a one-dimensional convolutional layer, and then through a reshaping operation of splicing again on the time axis to obtain h″′. That is, the time step features corresponding to K source sound sources and the noise. The PReLU expression is as follows:

[0080]

[0081] where α is a learnable parameter. After that, the time step features h″′ corresponding to K source sound sources and the noise are estimated through the ReLU non-linear function, and finally the masks source corresponding to each of the K sound sources and the mask m corresponding to the noise n are obtained.

[0082] As Figure 4 shown, in the causal attention module, the segmented Transformer and the memory Transformer have the same model structure, both of which are Transformer networks based on the causal attention mechanism:

[0083] Assume that the input of the Transformer network based on causal attention is the feature vector z. Before officially entering the network, the network will generate a causal mask (Casual Mask, CM) for z to prepare for the subsequent attention mechanism calculation. The processing process after entering the network is as follows:

[0084] z′ = z + e pos

[0085] z″ = Multiheadattention(norm(z′))

[0086] z″′ = z′ + z″ + FFN(norm(z′ + z″))

[0087] z″″ = permute(BatchNorm(permute(z″′)))

[0088] z″″′ = z + z″″

[0089] In the above formulas, e pos , norm(·), Multiheadattention(·), FFN(·), permute(·), and BatchNorm(·) represent relative position encoding, layer normalization, multi-head attention mechanism, feed-forward neural network, transpose operation for swapping dimensions, and batch normalization respectively. First, relative position encoding e pos is added to z to obtain the feature vector z′; layer normalization and the multi-head attention mechanism are used on z′ to obtain the feature vector z″; layer normalization, feed-forward neural network, and two residual connections respectively connected to z′ and z″ are sequentially used on z′ and z″ to obtain the feature vector z″′; then, to improve the performance and training stability of the model, z″′ is transposed to swap dimensions, normalized, and then transformed back to the original dimension to obtain the feature vector z″″; finally, the network's final output feature vector z″″′ is obtained through skip connection.

[0090] In the Transformer network based on the causal attention mechanism, the multi-head attention mechanism linearly transforms z′ into Q (query), K (key), and V (value) using three matrices with different parameters, and then calculates the non-linear relationship between Q, K, and V through scaled dot-product attention:

[0091]

[0092] where n head is the number of attention heads; d k is the dimension of the feature vector; CM is the causal mask generated before z officially enters the network, which is an upper triangular matrix: in the matrix, the part from the upper left to the lower right, except for the main diagonal, is negative infinity, and the diagonal and below are 0. By performing element-wise addition with this upper triangular matrix and then processing through the Softmax activation function, the positions of negative infinity will approach 0, removing the information of the future part. This method of calculating attention scores can help mask future information, limit the range of attention, and ensure the causality of attention.

[0093] S3: Utilize the learning domain feature representations of the clean source speech, estimated speech, and estimated noise to optimize the overall network parameters and implement the noise-aware contrast learning module;

[0094] As Figure 5 shown, the contrast learning module for K sourceEach sound source among the sound sources performs the same process, and for each sound source, the pure source speech feature representation is obtained through a sampler and a resampler. The corresponding estimated speech feature representation of each sound source And the estimated noise feature representation Positive samples, query samples, and negative samples are taken respectively, and the contrast loss is calculated using the three of them. After weighting this loss, the scale-invariant signal-to-noise ratio loss is superimposed to obtain the overall loss function of the network, and this is used as the joint optimization objective for model training. During the backpropagation process, the gradients of each loss function are calculated, and the gradient descent method is used to finally jointly optimize all the parameters of the network.

[0095] As Figure 5 shown, the sampler of the contrast learning module is a sampling network obtained by cascading a two-dimensional convolution, ReLU, and a two-dimensional convolution. The resampler is composed of a fully connected layer, ReLU, and a fully connected layer cascaded together.

[0096] For each of the K source sound sources s k , the sampler respectively performs operations on its pure source speech feature representation The corresponding estimated speech feature representation of this source speech And the estimated noise feature representation A total of K s random samplings are performed, and the same rule followed for each sampling is: randomly take 1 local block at the same position in the feature maps of and respectively, and at other positions in the feature map of the estimated noise feature representation , take another M - 1 local blocks. The sampling process of the sampler and the resampling process can be described as follows:

[0097]

[0098] Among them, respectively represent the positive sample, query sample, and negative sample obtained by the sampler from the i-th sampling from and , and H is the size of the convolution kernel in the sampler. All the s samples obtained from the K times of sampling are projected into a three-dimensional embedding space through the resampler, and at the same time, L2 normalization is used, and finally a set of positive sample, query sample, and negative sample feature vectors for contrast learning are obtained It should be clear that for each of the K source sound sources s k the above process is executed, that is, finally there will be K sourceGroup of samples. In each group of samples, when regarding the positive sample pairs as the "correct category" and all negative sample pairs as the "wrong category", a classification task of M + 1 is constituted.

[0099] Positive sample, query sample, and negative sample feature vectors obtained by the reshaper The process of calculating the contrast loss and obtaining the overall loss function of the network based on this loss is as follows:

[0100] The general form of the traditional cross-entropy loss when applied to a classification task is:

[0101]

[0102] where y t is the one-hot encoding of the true label, and p t is the predicted probability. In contrastive learning, for the sound source s k , only the positive samples are the true labels, that is, only the one-hot encoding of the positive samples is 1, and the one-hot encodings of the negative samples are all 0. Therefore, when calculating the cross-entropy loss at this time, only the probabilities of the positive samples need to be calculated.

[0103] In , since normalization has been completed by the reshaper, the cosine similarity between the positive sample and the query sample can be expressed as The cosine similarity between each negative sample and the query sample can be expressed as Inputting the similarities of each sample into the Softmax function, the probability p positive of the positive sample can be obtained:

[0104]

[0105] where τ is an additional introduced temperature coefficient used to adjust the smoothness of the similarity distribution. A smaller τ can enhance the contrast intensity. For a single sound source s k , finally, the positive sample probability is maximized through the following cross-entropy loss:

[0106]

[0107] Since the mixed speech contains K source sound sources, and the action mechanism of the contrastive learning module for each sound source is the same as that of the above-mentioned sound source s k , so further, the contrast loss Loss source for the overall K cll sound sources can be obtained:

[0108]

[0109] This loss enables the model to distinguish the local features of speech and noise when using the gradient descent algorithm for parameter update during backpropagation, so as to suppress noise in the separated signal and achieve noise perception.

[0110] In addition, adding Loss cll plus the scale-invariant signal-to-noise ratio loss Loss source calculated for K +1 sound sources (regarding noise as one of the sound sources, using to represent the noise signal n, ) to obtain the overall loss function of the network: SI-SNR

[0111]

[0112] Loss total SISNR = Loss SISNR + βLoss cll

[0113] where <·,·> represents the inner product of two vectors, s k represents the source speech signal, is its corresponding time-domain estimated signal; Loss total is the overall loss function of the network; β is a parameter manually adjusted to balance Loss SI-SNR and Loss total in the overall loss.

[0114] S4: Construct a decoding network, use the estimated feature representation to recover the corresponding time-domain estimated signal, and achieve speech separation;

[0115] In this embodiment, the decoder decodes and recovers the estimated feature representations of the clean source speech signal and noise respectively through one-dimensional transposed convolution to obtain the time-domain estimated source speech signal and the noise estimated signal to achieve speech separation. Among them, is obtained by performing element-wise multiplication of the noisy mixed speech feature representation with the mask m n n respectively.

[0116] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A voice separation method with noise perception based on contrastive learning and causal attention mechanism, characterized in that It includes the following steps: Construct an encoder to model the clean source speech signal and the noisy mixed speech signal based on the encoder, and generate the learned domain feature representation of the source speech signal and the learned domain feature representation of the noisy mixed speech; Construct a mask network to estimate the masks for the source speech signal and the noise signal, multiply the output of the mask network element-wise with the learned domain feature representation of the noisy mixed speech to obtain the estimated feature representations of different sound sources; Construct a contrastive learning module. For the source speech feature representations of each sound source, the corresponding estimated speech feature representations of each sound source, and the estimated noise feature representations, select positive samples, query samples, and negative samples respectively, calculate the contrastive loss, and after weighting this loss, superimpose the scale-invariant signal-to-noise ratio loss to obtain the overall loss function of the network, which is used as the joint optimization objective for model training. Calculate the gradients of each loss function during the backpropagation process to jointly optimize all parameters; Construct a decoding network to restore the corresponding time-domain estimated signal based on the estimated feature representation to achieve speech separation.

2. The method for speech separation with noise perception based on contrastive learning and causal attention mechanism according to claim 1, characterized in that The encoder models the time-domain mixed speech signal composed of each source speech signal s k and the noise signal through a one-dimensional convolutional network to obtain the feature representation of the noisy mixed speech in the learning domain and uses it as the input to the mask network; Meanwhile, the encoder also models each source speech signal separately to obtain the corresponding feature representations in the learning domain. Where F represents the dimensionality of the feature vector of h, and T' is the time length after convolution.

3. The speech separation method based on contrastive learning and causal attention mechanism with noise perception according to claim 1, wherein The mask network includes a feature block processing module and a causal attention module; The feature block processing module divides the learned domain feature representation of the noisy mixed speech along the time dimension into blocks, and then concatenates the divided results along the block dimension to obtain a feature vector h′, and inputs it into the causal attention module; The causal attention module obtains the latent feature representation of the speech based on the segmental Transformer network and the memory Transformer network; The feature block processing module performs block reshaping and mask generation on the latent feature representation of the speech.

4. The speech separation method based on contrastive learning and causal attention mechanism with noise perception according to claim 3, wherein The causal attention module has a total of L layers. Among them, the structures of the first L - 1 layers are as follows: the segmental Transformer network is combined with the mean calculation in the time dimension and the memory Transformer network through skip connections. The structure of the L-th layer only contains the segmental Transformer network; Set the input of the $l$-th layer as the feature vector $\mathbf{h}^{(\ell)}$, then the processing of this layer is represented as: l ' I l1 = segTransformer(h l ′) I l3 = memTransformer(I l2 ) I l4 = I l1 + I l3 Among them, segTransformer(·), and memTransformer(·) represent a segmented Transformer network, taking the mean in the time dimension, and a memory Transformer network, respectively; Feature vector h l 'The feature vector is obtained by the segmented Transformer For the feature vector I l1 Take the average in the time dimension to obtain the feature vector I l2 , and the feature vector I l1 Add it to the memory cell I l3 to obtain the feature vector I l4 , and the feature vector I l4 is used as the input of the (l + 1)-th layer, and the latent feature representation h″ of the speech is obtained after iterative processing.

5. The method for speech separation with noise perception based on contrastive learning and causal attention mechanism according to claim 3, characterized in that The feature block processing module performs block reshaping and mask generation on the latent feature representation of speech. The latent feature representation of speech passes through PReLU and a one-dimensional convolutional layer, and through a reshaping operation of re-splicing on the time axis, the time-step features h″′ corresponding to K source sound sources and noise are obtained, and then estimated through the ReLU non-linear function to obtain K source masks corresponding to each of the sound sources and the mask m corresponding to the noise n .

6. The method for speech separation with noise perception based on contrastive learning and causal attention mechanism according to claim 3, wherein In the Transformer network of the causal attention module, it is set that the input of the Transformer network based on causal attention is the feature vector z, and a causal mask is generated for the feature vector z; The processing process entering the Transformer network is expressed as: z′ = z + e pos z″ = Multiheadattention(norm(z′)) z″′ = z′ + z″ + FFN(norm(z′ + z″)) z″″ = permute(BatchNorm(permute(z″′))) z″″′ = z + z″″ Among them, e pos , norm(·), Multiheadattention(·), FFN(·), permute(·), and BatchNorm(·) represent relative position encoding, layer normalization, multi-head attention mechanism, feed-forward neural network, transpose operation for swapping dimensions, and batch normalization, respectively; Add relative position encoding e to the feature vector z pos Obtain the feature vector z'. Applying layer normalization and the multi-head attention mechanism to the feature vector z' can obtain the feature vector z''; The layer normalization, feed-forward neural network, and two residual connections connected to the feature vector z′ and the feature vector z″ are successively used for the feature vector z′ and the feature vector z″ to obtain the feature vector z″′; The dimension of the feature vector z″′ is transposed and exchanged for normalization, and then transformed back to the original dimension to obtain the feature vector z″″, and the final output feature vector z″″′ of the network is obtained through a skip connection.

7. The voice separation method based on contrastive learning and causal attention mechanism with noise perception according to claim 1, characterized in that The contrastive learning module includes a sampler and a reshaper; The sampler is a sampling network cascaded by a two-dimensional convolution, ReLU, and a two-dimensional convolution. The reshaper is composed of a fully connected layer, ReLU, and a fully connected layer cascaded in sequence.

8. The method for speech separation with noise perception based on contrastive learning and causal attention mechanism according to claim 7, characterized in that The sampler performs multiple random samplings on the source speech feature representation, the corresponding estimated speech feature representation of the source speech, and the estimated noise feature representation. The same rule followed each time is: randomly take 1 local block at the same position of the feature maps of the source speech feature representation, the corresponding estimated speech feature representation of the source speech, and the estimated noise feature representation, and at the same time, take M - 1 other local blocks at other positions of the feature map of the estimated noise feature representation; The sampling process of the sampler and the reshaping process of the resampler are expressed as: Among them, respectively represent the positive sample, query sample, and negative sample obtained by the sampler from h sk , and in the i-th sampling. H is the size of the convolution kernel in the sampler; K s All obtained by subsampling For i ∈ [1, K s , j ∈ [1, M], project to a three-dimensional embedding space through a reshaprer, while using L2 normalization, and finally obtain a set of positive samples, query samples, and negative sample feature vectors for contrastive learning Positive sample, query sample, and negative sample feature vectors obtained using a reshaper Calculate the contrastive loss.

9. The speech separation method based on contrastive learning and causal attention mechanism with noise perception according to claim 1, characterized in that Calculate the contrastive loss, and after weighting this loss, superimpose the scale-invariant signal-to-noise ratio loss to obtain the overall loss function of the network, which is specifically expressed as: Loss total = Loss SISNR + βLoss cll Among them, Loss total is the overall loss function of the network, Loss SI-SNR represents the scale-invariant signal-to-noise ratio loss, Loss cll represents the contrast loss for K source sound sources, β represents the loss weight parameter, K source represents the number of sound sources, s k represents the source speech signal, is its corresponding time-domain estimated signal, <·,·> represents the inner product of two vectors, p positive represents the probability of the positive sample, Loss ce (s k ) represents the cross-entropy loss, represents the positive sample, represents the query sample, represents the negative sample, j ∈ [1, M], τ represents the temperature coefficient, which is used to adjust the smoothness of the similarity distribution.

10. The method for speech separation with noise perception based on contrastive learning and causal attention mechanism according to claim 1, wherein Construct a decoding network to recover the corresponding time-domain estimated signal based on the estimated feature representation and achieve speech separation, specifically including: The decoder decodes and recovers the estimated feature representations of the source speech signal and the noise respectively through one-dimensional transposed convolution to obtain the time-domain estimated source speech signal and the noise estimation signal. Among them, the estimated feature representations of the source speech signal and the noise are obtained by performing element-wise multiplication on the noisy mixed speech feature representation and the masks corresponding to the sound sources respectively. The noise corresponding mask m n is obtained by the above operation.

Citation Information

Cited By

  • Blind noise reduction method for MEMS multi-sensor self-contrast learning

    CN121614752A

  • A blind denoising method of MEMS multi-sensor self-contrast learning

    CN121614752B

  • Remote underwater acoustic system-oriented end-to-end lightweight neural network time domain noise and interference suppression method

    CN122293216A