Sound separation and target sound extraction method based on unified architecture

Through the unified architecture of separation backbone network, attractor network and multimodal cue processing network, the problem of difficult flexible switching of sound separation and target sound extraction methods in existing technologies is solved, and efficient adaptability and robustness are improved in complex sound scenes.

CN120748429AActive Publication Date: 2025-10-03SHANGHAI JIAOTONG UNIV

Patent Information

Application Number
CN202510940403.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-10-03
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

Existing sound separation and target sound extraction methods are usually designed independently for a single task, making it difficult to switch flexibly within the same model. They also perform poorly in complex sound scenes, especially when the number of sound sources is uncertain.

Method used

It adopts a separation backbone network, attractor network and multimodal cue processing network based on a unified architecture, and through a two-stage training strategy and multimodal cue fusion, it achieves flexible switching between sound separation and target sound extraction, and can automatically estimate the number of sound sources and process multiple modal cues.

Benefits of technology

It achieves flexible execution of sound separation and target sound extraction tasks in the same model, improves adaptability and robustness in complex sound scenes, improves sound separation performance by 1.3dB and target sound extraction performance by 1.6dB, and is suitable for a variety of complex sound scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748429A_ABST
    Figure CN120748429A_ABST
Patent Text Reader

Abstract

The invention discloses a sound separation and target sound extraction method based on a unified architecture, and relates to the technical field of audio signal processing, and the method comprises the steps: in a first training stage, employing an attractor network to estimate the number of sound sources in a mixed audio signal, and generating attractor embedding; the attractors are embedded and input into the separation backbone network, and separated sound is generated; in the second training stage, the multi-mode clue processing network is adopted to process clue information of multiple modes, and clue embedding is generated; randomly selecting attractor embedding or clue embedding as input for separating the backbone network; in the training stage, an alignment loss function between attractor embedding and clue embedding is calculated; in the reasoning stage, if there is no clue information, a sound separation task is executed; and if one to three pieces of clue information exist, executing a target sound extraction task. The task is flexibly selected and executed according to the input clue condition, better adaptability and flexibility are achieved, and the method is suitable for complex sound scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio signal processing, and in particular to a method for sound separation and target sound extraction based on a unified architecture. Background Art

[0002] In the field of audio signal processing, sound separation (SS) and target sound extraction (TSE) are two key technologies. Sound separation aims to isolate different sound sources from a mixed audio signal, while target sound extraction focuses on extracting specific sounds of interest to the user from the mixed audio.

[0003] In complex acoustic environments, universal sound separation can provide high-quality input for subsequent audio analysis tasks. Universal Sound Separation (USS) has become a core task, aiming to separate any type of sound source (including speech, music, environmental sounds, and instrumental sounds) from a complex audio mixture without relying on a predefined number or type of sound sources. This flexibility greatly expands the application range of sound processing, making it indispensable in fields such as environmental monitoring and multimedia content analysis.

[0004] Despite recent progress in sound separation models, existing methods still face several limitations. For example, many models require the number of sound sources in a mixed audio to be predefined before inference, which limits their applicability in real-world scenarios where the number of sound sources is often uncertain.

[0005] To address these issues, target sound extraction leverages prior knowledge (i.e., cues) of the target sound from an unknown number of sources in a mixed audio. For example, the AudioSep model demonstrates significant performance by using natural language descriptions as extraction cues; APT uses audio samples as cues for TSE tasks; and PixelPlayer leverages visual information to extract the target sound.

[0006] There are still two shortcomings in practical applications: first, pre-existing clues may be of low quality or cannot be found at all in practical applications, which will reduce the performance of target sound extraction and even extract the wrong target; second, when the number of sound sources is fixed, methods based on target sound extraction usually perform worse than sound separation methods.

[0007] In addition, existing sound separation and target sound extraction methods are usually designed independently for a single task and are difficult to switch flexibly in the same model, which limits their application in complex sound scenes.

[0008] Therefore, technicians in this field are committed to developing a sound separation and target sound extraction method based on a unified architecture, which can flexibly choose to perform SS or TSE tasks according to the input clues, has better adaptability and flexibility, and is suitable for a variety of complex sound scenes. Summary of the Invention

[0009] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is how to flexibly switch between sound separation and specific sound extraction in the same model.

[0010] To achieve the above objectives, the present invention provides a method for sound separation and target sound extraction based on a unified architecture, wherein the method is implemented based on a separation backbone network, an attractor network, and a multimodal cue processing network;

[0011] It includes a training phase and an inference phase; the training phase includes a first training phase and a second training phase;

[0012] In the first training phase, the attractor network is used to estimate the number of sound sources in the mixed audio signal, generate an attractor embedding, and calculate the sound source counting loss; the attractor embedding is input into the separation backbone network to generate separated sounds, and the sound separation loss is calculated;

[0013] In the second training phase, the multimodal cue processing network is used to process cue information of multiple modalities to generate cue embeddings; the attractor embedding or the cue embedding is randomly selected as the input of the separation backbone network; if the cue embedding is selected as the input of the separation backbone network, all seven modal combinations are used for equal probability training;

[0014] During the training phase, if the input of the separation backbone network is the attractor embedding, a matching order between the separated audio and the real audio is calculated when the sound separation loss function is minimized, and an alignment loss function is calculated between the attractor embedding and the clue embedding based on the matching order to narrow the semantic space distance between the attractor and clue representations;

[0015] In the inference stage, if there is no clue information, the sound separation task is performed; if there are 1 to 3 clue information, the target sound extraction task is performed.

[0016] Furthermore, the separation backbone network includes:

[0017] An encoder for generating hidden features of the mixed audio signal;

[0018] A separator for estimating a mask for each sound source, and multiplying the latent feature by the mask corresponding to each sound source element-wise to obtain a latent representation;

[0019] The decoder, consisting of transposed convolutional layers, is symmetrical to the encoder and is used to estimate the separated sound sources.

[0020] Furthermore, the separator is used to estimate a mask for each sound source, including:

[0021] After passing the hidden features through a LayerNorm layer and a linear layer, they are split into overlapping blocks of size K = 250 and a 50% overlap rate;

[0022] A dual-path module is used to integrate Transformer blocks within and between blocks;

[0023] Perform sequence aggregation, use learnable weights to represent time regions with different dominant sounds in different subspaces, generate frame-level embedding W, and input W into the attractor network or the multimodal cue processing network;

[0024] The output V of the dual-path module is element-wise multiplied by the source sound representation of the attractor network or the multimodal cue processing network to form the input of the three-path module; the three-path module extends the dual-path design by adding a cross-channel Transformer block to capture the relationship between channels;

[0025] The output of the three-path module is sequentially passed through a parameterized ReLU layer, overlap-add, a gated output layer with two linear layers, and a linear layer with a ReLU activation function to generate a mask.

[0026] Furthermore, the attractor network includes an LSTM encoder and an LSTM decoder;

[0027] The LSTM encoder updates its hidden state by the following formula and cell status

[0028]

[0029] Among them, the hidden state and cell state of the LSTM encoder are initialized to zero vectors: and

[0030] The LSTM decoder estimates the attractor embedding by the following formula:

[0031]

[0032] At each step, the hidden state of the LSTM decoder is As an attractor of sound category s, its dimension D is consistent with the frame-level embedding W tThe hidden state and cell state of the LSTM decoder are initialized by the final state of the LSTM encoder: and

[0033] Furthermore, the attractor network further includes:

[0034] The existence probability of the attractor embedding is calculated through a fully connected layer with a sigmoid activation function, as follows:

[0035]

[0036] in, and are the trainable weights and bias parameters of the fully connected layer; a s Represents the estimated representation attractor of the sth sound in the mixed audio signal;

[0037] Each p exi Compare with the predefined threshold θ; if p exi >θ, then the attractor embedding is considered to exist.

[0038] Furthermore, the multimodal clue processing network encodes input clues of three modalities, namely text, video, and sound, into a unified D-dimensional space through a dedicated encoder; adopts a multimodal splicing strategy to form a unified multimodal clue; and uses a clue fusion module based on the attention mechanism to align and fuse the multimodal clues in time, thereby extracting global clue information from them and obtaining the clue embedding.

[0039] Furthermore, the dedicated encoder converts different modalities into D-dimensional embeddings, including:

[0040] A text encoder that converts natural language descriptions into text embeddings using the pre-trained DistilBERT model Among them, T t is the number of word tokens;

[0041] Video encoder, which uses the pre-trained Swin Transformer to process video frames and generate video embedding vectors Among them, T v is the number of frames;

[0042] Sound encoder, which maps the one-hot encoded sound event labels to sound embedding vectors through a linear layer

[0043] Furthermore, through a multimodal splicing strategy, the text embedding vector, video embedding vector, and sound embedding vector are spliced ​​to form a unified multimodal clue:

[0044]

[0045] Furthermore, the attention mechanism-based clue fusion module includes:

[0046] Using the output W of sequence aggregation as the query and U as the key and value, the clues are fused through the multi-head attention mechanism to obtain the fused clue embedding:

[0047]

[0048] Among them, MultiHeadAttention(Q,K,V) represents the multi-head attention mechanism, where Q, K, and V represent the query key and value respectively.

[0049] Furthermore, the alignment loss function is defined as follows:

[0050]

[0051] in, It is obtained by finding the optimal permutation between the attractor embedding and the cue embedding and calculating the mean squared error loss; Given N attractors and N cues, and the optimal permutation π obtained from the final separation PIT loss, it is obtained by averaging the InfoNCE loss for each pair of corresponding elements.

[0052] Compared with the prior art, the present invention has at least the following beneficial technical effects:

[0053] 1. The present invention integrates the sound separation (SS) and target sound extraction (TSE) tasks into a unified architecture, which can flexibly select to perform SS or TSE tasks according to the input clues. It has better adaptability and flexibility and is applicable to a variety of complex sound scenes.

[0054] 2. The present invention can automatically estimate the number of sound sources in mixed audio through an attractor network without manual intervention or pre-defined number. This makes it more practical in complex acoustic environments and can handle any number of sound source mixtures. In the sound separation task, the present invention improves the performance by 1.3dB compared to the existing single-task SS baseline model, significantly improving the separation effect in complex sound scenes. Moreover, compared to the baseline model Sepformer, it does not require the pre-defined number of sound sources and can separate mixed audio signals with an uncertain number of sound sources.

[0055] 3. The present invention supports multimodal cue input and can fuse these cues into a unified feature space through an attention mechanism. This multimodal processing capability significantly enhances robustness and adaptability in practical applications, especially when cue quality is low or some modalities are missing. In the target sound extraction task, the present invention improves the performance of the existing single-task TSE baseline model by 1.6dB and demonstrates stronger robustness and adaptability under multimodal cue fusion. Even when only single-modal cues are provided, it can achieve excellent target sound extraction performance.

[0056] 4. The present invention adopts a two-stage training strategy. The first stage focuses on the sound separation task, and the second stage performs joint training with the target sound extraction task. This strategy not only improves adaptability, but also further optimizes the performance of target sound extraction without compromising separation performance by calculating the alignment loss function of attractor embedding and cue embedding.

[0057] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 It is a schematic diagram of the overall structure of a preferred embodiment of the present invention;

[0059] Figure 2 1 is a schematic structural diagram of a separator according to a preferred embodiment of the present invention;

[0060] Figure 3 1 is a schematic diagram of a multimodal clue processing network structure according to a preferred embodiment of the present invention;

[0061] Figure 4 It is a visual proof diagram of the attractor and clue of a preferred embodiment of the present invention;

[0062] Figure 5 This is a graph evaluating the inference speed under different numbers of sound sources according to a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0063] The following describes several preferred embodiments of the present invention with reference to the accompanying drawings to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.

[0064] This embodiment provides a method for sound separation and target sound extraction based on a unified architecture, such as Figure 1 As shown, its architecture includes a separation backbone network, an attractor network, and a multimodal clue processing network.

[0065] The separation backbone network, employing an encoder-separator-decoder architecture, is responsible for feature extraction and sound separation in mixed audio signals. The attractor network dynamically estimates the number of sound sources in the mixed audio signal and provides attractor embeddings. The multimodal cue processing network processes multimodal cues such as text, video, and audio tags to generate unified cue embeddings.

[0066] 1. Separate the backbone network

[0067] Given a single-channel mixed audio signal from J sound sources (where L is an arbitrary length), the goal of the separation backbone network is to estimate the number of sound sources based on the mixed audio signal x and the attractor network Estimate each sound source signal

[0068] The encoder uses a 1D convolutional layer and ReLU activation function to generate hidden features of the mixed audio signal. The separator estimates a mask for each sound source By multiplying X element-wise by m j , get the estimated hidden representation The decoder consists of transposed convolutional layers, symmetrical to the encoder, to estimate the separated sound sources.

[0069] The separators use SepFormer based on time domain and BSRNN model based on frequency domain respectively. Here we mainly introduce the method based on SepFormer. Figure 2 As shown in the figure, a structure similar to SepFormer is used as the separator, where U, V, W, Y, and Z are intermediate features; N is the time dimension of the mixed audio feature; K is a constant, which is the size of each chunk during the chunking process; J is the number of sound sources in the mixed audio; and C is the time dimension of the feature after chunking.

[0070] First, the hidden feature X passes through a LayerNorm layer and a linear layer, and is split into overlapping blocks of size K = 250 with an overlap rate of 50%.

[0071] These overlapping blocks are then fed into a dual-path module that integrates intra-block and inter-block Transformer blocks.

[0072] Subsequently, sequence aggregation is performed, using learnable weights to represent temporal regions with different dominant sounds in different subspaces, and the output W is fed into an attractor network or a multimodal cue processing network.

[0073] Furthermore, the output V of the dual-path module is combined with the source sound representation from the attractor network or the multimodal cue processing network Multiply element-wise to form the input of the three-path module The three-path module extends the two-path design by adding a cross-channel Transformer block to capture the relationship between channels.

[0074] Finally, the output of the three-path module is The mask is generated by passing through a parameterized ReLU layer, overlap addition (OVA), a gated output layer with two linear layers, and a linear layer with a ReLU activation function.

[0075] 2. Attractor Network

[0076] Attractor networks are used to estimate the number of different sound categories and attractor embeddings from mixed audio signals, thereby achieving sound separation without knowing the number of sound sources in advance.

[0077] The attractor network adopts LSTM encoder-decoder framework to embed the frame level into W t Convert to a global attractor.

[0078] The LSTM encoder updates its hidden state by the following formula and cell status

[0079]

[0080] Among them, the hidden state and cell state of the LSTM encoder are initialized to zero vectors: and

[0081] The LSTM decoder estimates the attractor embedding via the following formula:

[0082]

[0083] At each step, the hidden state of the LSTM decoder is As the attractor embedding of the sound category s, its dimension D is the same as the frame-level embedding W t The dimensions are consistent.

[0084] The hidden state and cell state of the LSTM decoder are initialized by the final state of the LSTM encoder: and

[0085] The probability of the attractor embedding is calculated through a fully connected layer with a sigmoid activation function, as follows:

[0086]

[0087] in, and are the trainable weights and bias parameters of the fully connected layer, a s Represents the estimated representation attractor of the sth sound in the mixed audio signal Mixture. exi Compare with the predefined threshold θ. If p exi >θ, the attractor is considered to exist; otherwise, it is considered that there are no more sound sources in the mixed audio signal.

[0088] 3. Multimodal Cue Processing Network

[0089] The multimodal cue processing network takes cues from multiple modalities (e.g. text, video, sound) as input and generates a comprehensive cue embedding. Figure 3 As shown in the figure, each modality is first encoded into a unified D-dimensional space using a dedicated encoder. Then, a multimodal concatenation strategy is adopted to form a unified multimodal cue. Finally, a cue fusion module based on the attention mechanism is used to temporally align and fuse these cues, thereby extracting global cue information and obtaining a cue embedding.

[0090] A dedicated encoder is used to convert different modalities into D-dimensional embeddings for target sound extraction, including:

[0091] Text encoder, using the pre-trained DistilBERT model, converts natural language descriptions into text embedding vectors Among them, T t is the number of word tokens.

[0092] Video encoder, which uses the pre-trained Swin Transformer to process video frames and generate video embedding vectors Where T v is the number of frames.

[0093] Sound encoder, which maps the one-hot encoded sound event labels to sound embedding vectors through a linear layer

[0094] The multimodal splicing strategy forms a unified multimodal clue by splicing the text embedding vector, video embedding vector and sound embedding vector embedding:

[0095]

[0096] The clue fusion module based on the attention mechanism uses the output W of sequence aggregation as the query and U as the key and value. It fuses the clues through the multi-head attention mechanism to obtain the fused clue embedding:

[0097]

[0098] Among them, MultiHeadAttention(Q,K,V) represents the multi-head attention mechanism, where Q, K, and V represent the query key and value respectively. Obviously, the fused clue embedding C u Same length T as the sound embedding vector Q a .

[0099] 4. Training and Inference Strategies

[0100] This embodiment adopts a unique training and reasoning strategy, including a training phase and a reasoning phase; the training phase includes a first training phase and a second training phase.

[0101] In the first training stage, an attractor network is used to estimate the number of sound sources in the mixed audio signal, generate attractor embeddings, and calculate the sound source counting loss; the attractor embeddings are input into the separation backbone network to generate separated sounds, and the sound separation loss is calculated.

[0102] In the second training phase, a multimodal cue processing network is used to process cue information from multiple modalities (text, video, and sound) to generate cue embeddings; attractor embeddings (30% probability) or cue embeddings (70% probability) are randomly selected as input to the separation backbone network.

[0103] If cue embeddings are selected as the input to the separation backbone network, all seven modality combinations are trained with equal probability. The seven modality combinations include three unimodal combinations (text, video, and sound), three bimodal combinations (text+video, text+sound, and video+sound), and one trimodal combination (text+video+sound).

[0104] During training, when the attractor embedding is input to the separation network, the matching order between the separated audio and the real audio is calculated to minimize the sound separation loss function. Based on the matching order, the alignment loss function is calculated between the attractor embedding and the clue embedding to narrow the semantic space distance between the attractor and clue representations. This strategy enhances the adaptability and robustness of this embodiment and significantly reduces the training time.

[0105] In the inference stage, if there is no clue information, the sound separation task is performed; if there are 1 to 3 clue information, the target sound extraction task is performed.

[0106] 5. Objective Function

[0107] The objective function of this embodiment is a combination of three different loss functions: sound separation loss function Sound source counting loss function and the alignment loss function between attractor embedding and cue embedding

[0108] In the multi-source sound separation task, the signal-to-noise ratio (SNR) combined with permutation invariant training (PIT) is used as the objective function, and the sound separation loss function can be expressed as:

[0109]

[0110] Among them, Π K represents all possible permutations, π is a permutation map, K is the number of target sound sources, T is the length of the signal, is the value of the kth estimated sound source at time step t under permutation π, and s k (t) is the value of the target sound source at time step t.

[0111] The sound source counting loss function uses cross entropy to measure the accuracy of the model's estimation of the number of sound sources in the mixed audio signal. The formula is as follows:

[0112]

[0113] Among them, y i is the true label, and p exi is the predicted probability.

[0114] Alignment loss function It is defined as follows:

[0115]

[0116] By finding the optimal permutation between the attractor embedding and the cue embedding and calculating the mean squared error (MSE) loss, it can be expressed as:

[0117]

[0118] Among them, π represents a permutation, Π K is the set of all possible permutations, and D is the dimension of the embedding. π(m),i represents the i-th value of the m-th attractor under the permutation π, and c m,i is the i-th value of the m-th cue. M is the total number of cue embeddings or attractor embeddings.

[0119] Given N attractors and N clues and the optimal permutation π obtained from the final separation PIT loss, InfoNCE loss It is obtained by calculating the average of the InfoNCE loss for each pair of corresponding elements:

[0120]

[0121] in, is the embedding vector of the i-th attractor, is the embedding vector of the jth cue, π(i) is the index of the cue that best matches the ith attractor according to the separation PIT loss, and τ is the temperature adjustment coefficient.

[0122] In actual application, this embodiment has good performance indicators, including:

[0123] 1. Sound separation performance

[0124] The framework used in this embodiment is the USE framework. Sepformer and BSRNN are benchmark models for sound separation, and USE-S and USE-B are models implemented under the USE framework using Sepformer and BSRNN as separators, respectively. Stage 1 only trains separation tasks, and stage 2 jointly trains SS and TSE tasks based on stage 1, and has the ability to separate sounds and extract target sounds. nMix means that each Mixture in the data set is a mixture of n sounds. Traditional Sepformer and BSRNN models can only train data with a fixed number of Mixes separately, while this embodiment can train data with an unfixed number of Mixes (here, 2Mix and 3Mix data can be trained at the same time), thanks to the attractor network's ability to estimate the number of sound sources.

[0125] Table 1 shows that the USE framework significantly outperforms existing Sepformer and BSRNN models in sound separation performance. USE-S stage 1 achieves SNR improvements of 8.7 and 7.9 dB on the 2Mix Seen (in-domain) and Unseen (out-of-domain) datasets, respectively, and 6.4 and 5.2 dB on the 3Mix Seen dataset, respectively, surpassing Sepformer by 1.3 dB. USE-S stage 2 further trains SS and TSE simultaneously, achieving not only a slight improvement in separation metrics but also a slight improvement while achieving TSE. USE-B stage 1 achieves SNR improvements of 8.7 and 7.9 dB on the 2Mix Seen and Unseen datasets, respectively, and 6.4 and 5.2 dB on the 3Mix Seen dataset. USE-B stage 2 further trains SS and TSE simultaneously, achieving not a slight improvement in separation metrics while achieving TSE.

[0126] Table 1 Comparison of signal-to-noise ratio improvement in sound separation tasks (unit: dB)

[0127]

[0128] 2. Sound source counting accuracy

[0129] As shown in Table 2, in the 3Mix scenario, USE's sound source counting accuracy is approximately 85.24%, and slightly insufficient in the 2Mix scenario. This can be attributed to the fact that each audio track in the AudioSet dataset may still contain multiple sound events even if sound event detection (SED) technology is used to isolate fragments. This may lead to a decrease in the accuracy of sound source number estimation under the 2Mix condition. However, this level of accuracy shows that the model can still reliably estimate the number of sound sources in mixed audio, providing strong support for the sound separation task.

[0130] Table 2 Comparison of sound source counting accuracy

[0131] Counting accuracy 2Mix 3Mix Within the domain 79.74% 84.42% Overseas 74.68% 85.24%

[0132] 3. Target sound extraction performance

[0133] As shown in Table 3, USE performs exceptionally well in the target sound extraction task under multimodal cues. For example, on the 2Mix Seen dataset (Seen) and Unseen (Unseen), USE-B achieves 8.9dB and 8.8dB SNR improvements, respectively, outperforming the existing DCCRN model by 29.0% and 35.4%. Even with only a single modality, USE maintains high performance, demonstrating its robustness and adaptability.

[0134] Table 3 Comparison of target sound extraction signal-to-noise ratio improvement under multimodal clue conditions (unit: dB)

[0135]

[0136] like Figure 4 As shown in Figure 1, during inference, four 3Mix mixed samples were randomly selected, two of which were from the Seen dataset and the other two from the Unseen dataset. t-SNE visualization of the attractor and clue in each mixed sample was performed.

[0137] like Figure 4 As shown, different types of sounds have a certain spatial distance in the t-SNE diagram, indicating that they are separable to a certain extent in the feature space.

[0138] Furthermore, attractors and cues of the same sound type exhibit similarity in feature space, meaning that if one is missing, the other can serve as an effective substitute.

[0139] As shown in Table 4, the extraction performance of USE-S (based on Sepformer) and USE-B (based on BSRNN) is compared with the models trained only for the target sound extraction task.

[0140] Table 4 Comparison of target sound extraction signal-to-noise ratio improvement (unit: dB)

[0141]

[0142] The results show that the extraction performance of USE-S and USE-B does not degrade after joint training with sound separation and target sound extraction.

[0143] In fact, they even achieved slight improvements in performance on the target sound extraction task. USE-S achieved SNR improvements of 8.5 and 8.1 dB on the 2Mix Seen and Unseen datasets, respectively, and 6.5 and 5.9 dB on the 3Mix Seen and Unseen datasets, respectively, a slight improvement over Sepformer. USE-B achieved SNR improvements of 8.9 and 8.8 dB on the 2Mix Seen and Unseen datasets, respectively, and 6.3 and 5.0 dB on the 3Mix Seen and Unseen datasets, respectively, a significant improvement over BSRNN.

[0144] Furthermore, these models are still able to perform the sound separation task without clues using the EDA (Attractor Network) module.

[0145] As shown in Table 5, the performance of different general sound extraction models on the target sound extraction task is also compared. Most of these models use the same SED pruning strategy as the processing method of this embodiment to process the AudioSet dataset.

[0146] However, these models may differ in details, so these results can only be used as a reference. Nevertheless, it is clear from the table that the USE-B model proposed in this example performs best on the TSE task, and its performance is comparable to the separately trained BSRNN+ClueNet model.

[0147] Table 5 Comparison of signal-to-noise ratio improvements of general sound extraction models (unit: dB)

[0148] Model 2Mix 3Mix MAE(audio) 5.6 / USS(audio) 5.6 / LASS(text) 6.8 / AudioSep(text) 7.7 / DCCRN(text+tag+video) 6.9 / Sepformer(text+tag+video) 8.5 6.5 USE-S(text+tag+video) 8.5 6.5 BSRNN (text+tag+video) 8.4 5.9 USE-B (text+tag+video) 8.9 6.3

[0149] 4. General sound separation performance

[0150] As shown in Table 6, a BSRNN-based USE-B model is also trained on a large-scale general audio dataset, which includes a general dataset mainly remixed from VGGSound (2 to 6 mixed sources), AudioSet (2 to 3 mixed sources), FUSS (2 to 4 mixed sources), and Musan (music) + Librispeech (speech) (2 to 5 mixed sources).

[0151] Table 6 Comparison of the SNR improvement of USE-B in sound separation on the composite dataset (unit: dB)

[0152]

[0153] During the evaluation, two cases are considered: one where the number of sound sources to be separated is known, and the other where the number of sound sources is unknown and needs to be estimated using the EDA module.

[0154] First, we can see that USE-B is able to adapt to mixtures of varying numbers of sound sources and performs well on a variety of datasets. Second, we found that methods using the EDA module to predict the number of sound sources are generally comparable to methods based on a fixed number of preset sound sources, and in some cases even outperform them. Finally, USE-B not only successfully completes the separation task, but also performs well in the target sound extraction task.

[0155] As shown in Table 7, the USE-B model is also compared with other competitive models in terms of scale-invariant signal-to-noise ratio improvement (SI-SNRi) on the reverberation-free FUSS test set. It can be seen that the USE-B model achieves excellent separation performance in scenarios with two mixed sources, three mixed sources, and four mixed sources.

[0156] Table 7 Comparison of scale-invariant signal-to-noise ratio improvement (unit: dB)

[0157] Model 2Mix 3Mix 4Mix TDCN++ 11.2 11.6 7.4 USE-B 12.8 13.1 11.9

[0158] like Figure 5 As shown, the inference speed for extracting audio signals from one to six sound sources from mixed audio was also evaluated. It was observed that during inference, the computational complexity increases linearly with the number of sound sources. Notably, even when inferring up to six sound sources, the computational requirement in GFLOPS remains below 30. This ensures the real-time performance and high efficiency of the USE-B model during inference.

[0159] The present invention is applicable to a variety of practical application scenarios, including but not limited to:

[0160] 1. Environmental monitoring: separating specific sound sources (such as animal calls, traffic noise, etc.) from complex environmental sounds.

[0161] 2. Multimedia content analysis: extracting target sounds from video or audio content, even in real-world scenarios such as conference rooms, to improve user experience.

[0162] 3. Intelligent voice assistant, which extracts user voice commands in noisy environments and improves the accuracy of voice recognition.

[0163] 4. Audio editing and creation, helping audio engineers quickly separate and edit different sound components in audio.

[0164] The present invention is based on existing deep learning frameworks (such as ESPnet-SE) and is easily integrated with other audio processing systems. In addition, its multimodal clue processing capabilities allow for flexible expansion of input modalities according to specific needs, further enhancing the applicability of the present invention.

[0165] The present invention focuses on computational efficiency in its design, employing a highly efficient encoder-separator-decoder architecture and a lightweight multimodal processing module. It can run on resource-constrained devices, such as embedded systems or mobile devices, and has high practical value.

[0166] The above describes in detail the preferred embodiments of the present invention. It should be understood that those skilled in the art can make numerous modifications and variations based on the concepts of the present invention without inventive effort. Therefore, any technical solutions that can be derived by those skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.

Claims

1. A method for sound separation and target sound extraction based on a unified architecture, characterized in that: The method is implemented based on a separation backbone network, an attractor network, and a multimodal clue processing network; It includes a training phase and an inference phase; the training phase includes a first training phase and a second training phase; In the first training phase, the attractor network is used to estimate the number of sound sources in the mixed audio signal, generate an attractor embedding, and calculate the sound source count loss; Embedding the attractor into the separation backbone network to generate separated sounds and calculating the sound separation loss; In the second training phase, the multimodal cue processing network is used to process cue information of multiple modalities to generate cue embeddings; the attractor embedding or the cue embedding is randomly selected as the input of the separation backbone network; if the cue embedding is selected as the input of the separation backbone network, all seven modal combinations are used for equal probability training; During the training phase, if the input of the separation backbone network is the attractor embedding, a matching order between the separated audio and the real audio is calculated when the sound separation loss function is minimized, and an alignment loss function is calculated between the attractor embedding and the clue embedding based on the matching order to narrow the semantic space distance between the attractor and clue representations; In the inference stage, if there is no clue information, the sound separation task is performed; if there are 1 to 3 clue information, the target sound extraction task is performed.

2. The method for sound separation and target sound extraction based on a unified architecture as claimed in claim 1, wherein: The separation backbone network includes: An encoder for generating hidden features of the mixed audio signal; A separator for estimating a mask for each sound source, and multiplying the latent feature by the mask corresponding to each sound source element-wise to obtain a latent representation; The decoder, consisting of transposed convolutional layers, is symmetrical to the encoder and is used to estimate the separated sound sources.

3. The method for sound separation and target sound extraction based on a unified architecture as claimed in claim 2, wherein: The separator is used to estimate a mask for each sound source and includes: After passing the hidden features through a LayerNorm layer and a linear layer, they are split into overlapping blocks of size K = 250 and a 50% overlap rate; A dual-path module is used to integrate Transformer blocks within and between blocks; Perform sequence aggregation, use learnable weights to represent time regions with different dominant sounds in different subspaces, generate frame-level embedding W, and input W into the attractor network or the multimodal cue processing network; The output V of the dual-path module is element-wise multiplied by the source sound representation of the attractor network or the multimodal cue processing network to form the input of the three-path module; the three-path module extends the dual-path design by adding a cross-channel Transformer block to capture the relationship between channels; The output of the three-path module is sequentially passed through a parameterized ReLU layer, overlap-add, a gated output layer with two linear layers, and a linear layer with a ReLU activation function to generate a mask.

4. The method for sound separation and target sound extraction based on a unified architecture as claimed in claim 1, wherein: The attractor network includes an LSTM encoder and an LSTM decoder; The LSTM encoder updates its hidden state by the following formula and cell status Among them, the hidden state and cell state of the LSTM encoder are initialized to zero vectors: and The LSTM decoder estimates the attractor embedding by the following formula: At each step, the hidden state of the LSTM decoder is As an attractor of sound category s, its dimension D is consistent with the frame-level embedding W t The hidden state and cell state of the LSTM decoder are initialized by the final state of the LSTM encoder: and 5. The method for sound separation and target sound extraction based on a unified architecture as claimed in claim 4, wherein: The attractor network further includes: The existence probability of the attractor embedding is calculated through a fully connected layer with a sigmoid activation function, as follows: in, and are the trainable weights and bias parameters of the fully connected layer; a s Represents the estimated representation attractor of the sth sound in the mixed audio signal; Each p exi Compare with the predefined threshold θ; if p exi >θ, then the attractor embedding is considered to exist.

6. The method for sound separation and target sound extraction based on a unified architecture as claimed in claim 1, wherein: The multimodal clue processing network encodes input clues in the three modalities of text, video, and sound into a unified D-dimensional space through a dedicated encoder; adopts a multimodal splicing strategy to form a unified multimodal clue; and uses a clue fusion module based on the attention mechanism to align and fuse the multimodal clues in time, thereby extracting global clue information from them and obtaining the clue embedding.

7. The method for sound separation and target sound extraction based on a unified architecture as claimed in claim 6, wherein: The specialized encoder converts different modalities into D-dimensional embeddings, including: A text encoder that converts natural language descriptions into text embeddings using the pre-trained DistilBERT model Among them, T t is the number of word tokens; Video encoder, which uses the pre-trained Swin Transformer to process video frames and generate video embedding vectors Among them, T v is the number of frames; Sound encoder, which maps the one-hot encoded sound event labels to sound embedding vectors through a linear layer 8. The method for sound separation and target sound extraction based on a unified architecture as claimed in claim 7, wherein: Through the multimodal splicing strategy, the text embedding vector, video embedding vector, and sound embedding vector are spliced ​​to form a unified multimodal clue:

9. The method for sound separation and target sound extraction based on a unified architecture as claimed in claim 8, wherein: The attention mechanism-based clue fusion module includes: Using the output W of sequence aggregation as the query and U as the key and value, the clues are fused through the multi-head attention mechanism to obtain the fused clue embedding: Among them, MultiHeadAttention(Q,K,V) represents the multi-head attention mechanism, where Q, K, and V represent the query key and value respectively.

10. The method for sound separation and target sound extraction based on a unified architecture as claimed in claim 1, wherein: The alignment loss function is defined as follows: in, It is obtained by finding the optimal arrangement between the attractor embedding and the cue embedding and calculating the mean square error loss; Given N attractors and N cues, and the optimal permutation π obtained from the final separation PIT loss, it is obtained by averaging the InfoNCE loss for each pair of corresponding elements.

Citation Information

Patent Citations

  • Sound source localization and sound source separation method and system based on dual consistent network

    CN113850246A

  • Target voice separation method and system based on cross-modal loss

    CN118016093A

  • Audio separation model training method and device, equipment and storage medium

    CN120071953A

  • Target speaker extraction method and system based on multi-scale multi-modal alignment network

    CN120126454A

  • Sound source separation and localization device, method and program

    JP2014021315A

Cited By

  • Universal audio separation method and system based on dual-path heterogeneous collaboration

    CN122337234A