Audio data separation method, device, electronic device and storage medium

By adopting a convolutional self-attention mechanism model based on a codec structure in the audio separation network, the mixed audio is separated and the problem of inefficient separation in the prior art is solved, and higher separation accuracy and efficiency are achieved.

CN114446318BActive Publication Date: 2025-06-10BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210120055.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-07
Publication Date
2025-06-10
Estimated Expiration
2042-02-07

AI Technical Summary

Technical Problem

The prior art is inefficient in separating target audio information from mixed audio, especially when the target audio information has contextual relevance, the limitations of the convolutional network and the high complexity of the recurrent neural network lead to separation accuracy and inefficiency.

Method used

An audio separation network based on a convolutional self-attention mechanism model based on a codec structure is adopted to transform the audio data to be processed, spectrum features are extracted, and the predicted spectrum features are obtained through the encoder, attention module and decoder to obtain the predicted spectrum features of each target audio information.

Benefits of technology

It improves the accuracy and effect of audio separation, reduces network complexity, improves separation efficiency, and can simultaneously capture local characteristics and global dependency information of audio data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114446318B_ABST
    Figure CN114446318B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, apparatus, electronic device, and storage medium for audio data separation. The method includes: performing transformation processing on the audio data to be processed to obtain spectral features corresponding to the audio data to be processed, where the audio data to be processed includes target audio information of multiple audio types; performing separation processing on the spectral features through an audio separation network to obtain predicted spectral features corresponding to each of the target audio information in the audio data to be processed, where the audio separation network is a convolutional self-attention mechanism model based on an encoding-decoding structure; performing inverse transformation processing on the predicted spectral features corresponding to each of the target audio information to obtain each of the target audio information in the audio data to be processed. Using the present disclosure can improve the efficiency and accuracy of audio separation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of multimedia technologies, and in particular, to an audio data separation method, apparatus, electronic device, and storage medium. Background Art

[0002] In daily life, there is a need to separate a fixed type of sound from mixed audio. For example, the audio data collected by a microphone usually consists of multiple sound sources, such as speech, music, and background noise. The mixing of these sound signals will reduce the audio quality. For example, if there is singing in the audio, it will interfere with the speech recognition of the audio, resulting in speech recognition failure or reduced speech recognition accuracy.

[0003] In related technologies, a convolutional network or a recurrent neural network can be used for audio data separation, that is, to separate various target audio information from mixed audio. For example, to separately separate target audio information such as speech, music, and background noise from mixed audio.

[0004] When the target audio information is audio information with context relevance, due to the fixed receptive field of the convolutional network, the convolutional network lacks the ability to capture global dependencies, resulting in low separation accuracy and poor separation effect. Although the recurrent neural network can capture the long-term dependence relationship of audio signals, the complexity of such a network is very high, and the separation efficiency of audio data is low. Summary of the Invention

[0005] The present disclosure provides an audio data separation method, apparatus, electronic device, and storage medium to at least solve the problem of low separation efficiency in related technologies. The technical solution of the present disclosure is as follows:

[0006] According to a first aspect of an embodiment of the present disclosure, an audio data separation method is provided, including:

[0007] Performing transformation processing on the audio data to be processed to obtain the spectral features corresponding to the audio data to be processed, where the audio data to be processed includes target audio information of multiple audio types;

[0008] Performing separation processing on the spectral features through an audio separation network to obtain the predicted spectral features corresponding to each target audio information in the audio data to be processed, where the audio separation network is a convolutional self-attention mechanism model based on an encoder-decoder structure;

[0009] Performing inverse transformation processing on the predicted spectral features corresponding to each target audio information to obtain each target audio information in the audio data to be processed.

[0010] In one embodiment, the audio separation network includes an encoder, an attention module, and a plurality of decoders corresponding to the multiple audio types one by one. Separating the spectral features through the audio separation network to obtain the predicted spectral features corresponding to each target audio information in the audio data to be processed includes:

[0011] Encoding the spectral features through the encoder to obtain the audio features of the audio data to be processed;

[0012] Extracting features from the audio features through the attention module to obtain the attention features corresponding to each target audio information;

[0013] Respectively decoding the attention features corresponding to each target audio information through the decoders corresponding to each target audio information to obtain the predicted spectral features corresponding to each target audio information.

[0014] In one embodiment, both the encoder and the decoder are constructed by a convolutional self-attention mechanism model. The convolutional self-attention mechanism model includes a feature excitation layer, and the feature excitation layer is used to perform feature learning from the spatial dimension and the convolutional channel dimension.

[0015] In one embodiment, the attention module includes a plurality of attention mechanisms corresponding to the multiple audio types one by one. The attention mechanism includes a convolutional module and a first feature normalization module. Extracting features from the audio features through the attention module to obtain the attention features corresponding to each target audio information includes:

[0016] For any one of the attention mechanisms, extracting features from the audio features through the convolutional module in the attention mechanism to obtain the initial attention features of the target audio information corresponding to the attention mechanism;

[0017] Normalizing the initial attention features through the first feature normalization module in the attention mechanism to obtain the normalized initial attention features, and fusing the normalized initial attention features with the audio features to obtain the attention features of the target audio information corresponding to the attention mechanism.

[0018] In one embodiment, before separating the spectral features through the audio separation network to obtain the predicted spectral features corresponding to each target audio information in the audio data to be processed, the method further includes:

[0019] Obtain a sample group, where the sample group includes the sample audio data and the annotation information of the sample audio data, and the annotation information of the sample audio data includes the sample target audio information of N audio types used to constitute the sample audio data;

[0020] Train an initial audio separation network using the sample group in the training set to obtain the audio separation network.

[0021] In one embodiment, the training of the initial audio separation network using the sample group in the training set to obtain the audio separation network includes:

[0022] Perform transformation processing on the sample audio data in the sample group to obtain the sample spectral features of the sample audio data;

[0023] Perform separation processing on the sample spectral features through the initial audio separation network to obtain the sample spectral features corresponding to each sample target audio information in the sample audio data;

[0024] Perform inverse transformation processing on the sample spectral features corresponding to each sample target audio information to obtain multiple predicted target audio information;

[0025] Determine the separation loss of the initial audio separation network according to the multiple predicted target audio information and the multiple sample target audio information corresponding to the sample audio data;

[0026] Train the initial audio separation network according to the separation loss to obtain the audio separation network.

[0027] In one embodiment, the determining the separation loss of the initial audio separation network according to the multiple predicted target audio information and the multiple sample target audio information corresponding to the sample audio data includes:

[0028] Determine the first loss and the second loss of the initial audio separation network according to each predicted target audio information and each sample target audio information corresponding to the sample audio data, where the first loss is used to characterize the difference between the predicted target audio information and the sample target audio information, and the second loss is used to characterize the difference between the predicted target audio information;

[0029] Perform fusion processing on the first loss and the second loss to obtain the separation loss of the initial audio separation network.

[0030] According to the second aspect of the embodiments of the present disclosure, there is provided an audio data separation device, including:

[0031] A first transformation unit, configured to perform a transformation process on the audio data to be processed to obtain spectral features corresponding to the audio data to be processed, where the audio data to be processed includes target audio information of multiple audio types;

[0032] A first separation unit, configured to perform a separation process on the spectral features through an audio separation network to obtain predicted spectral features corresponding to each of the target audio information in the audio data to be processed, where the audio separation network is a convolutional self-attention mechanism model based on an encoding-decoding structure;

[0033] A first inverse transformation unit, configured to perform an inverse transformation process on the predicted spectral features corresponding to each of the target audio information to obtain each of the target audio information in the audio data to be processed.

[0034] In one embodiment, the audio separation network includes an encoder, an attention module, and a plurality of decoders corresponding to the multiple audio types one by one, and the separation unit is further configured to perform:

[0035] Encoding the spectral features through the encoder to obtain audio features of the audio data to be processed;

[0036] Extracting features from the audio features through the attention module to obtain attention features corresponding to each of the target audio information;

[0037] Respectively decoding the attention features corresponding to each of the target audio information through the decoders corresponding to each of the target audio information to obtain predicted spectral features corresponding to each of the target audio information.

[0038] In one embodiment, both the encoder and the decoder are constructed through a convolutional self-attention mechanism model, and the convolutional self-attention mechanism model includes a feature excitation layer, and the feature excitation layer is used for feature learning from the spatial dimension and the convolutional channel dimension.

[0039] In one embodiment, the attention module includes a plurality of attention mechanisms corresponding to the multiple audio types one by one, the attention mechanism includes a convolutional module and a first feature normalization module, and the separation unit is further configured to perform:

[0040] For any one of the attention mechanisms, extracting features from the audio features through the convolutional module in the attention mechanism to obtain initial attention features of the target audio information corresponding to the attention mechanism;

[0041] The initial attention features are normalized by the first feature normalization module in the attention mechanism to obtain normalized initial attention features, and the normalized initial attention features are fused with the audio features to obtain the attention features of the target audio information corresponding to the attention mechanism.

[0042] In one embodiment, the apparatus further includes:

[0043] An acquisition unit, configured to execute acquiring a sample group, the sample group including the sample audio data and the annotation information of the sample audio data, and the annotation information of the sample audio data including sample target audio information of multiple audio types for constituting the sample audio data;

[0044] A second transformation unit, configured to execute transforming the sample audio data in the sample group to obtain the sample spectrum features of the sample audio data;

[0045] A second separation unit, configured to execute separating the sample spectrum features through an initial audio separation network to obtain the sample spectrum features corresponding to each sample target audio information in the sample audio data;

[0046] A second inverse transformation unit, configured to execute inverse-transforming the sample spectrum features corresponding to each sample target audio information to obtain multiple predicted target audio information;

[0047] A determination unit, configured to execute determining the separation loss of the initial audio separation network according to the multiple predicted target audio information and the N sample target audio information corresponding to the sample audio data;

[0048] A training unit, configured to execute training the initial audio separation network according to the separation loss to obtain the audio separation network.

[0049] In one embodiment, the determination unit is further configured to execute:

[0050] Determining a first loss and a second loss of the initial audio separation network according to each predicted target audio information and each sample target audio information corresponding to the sample audio data, where the first loss is used to characterize the difference between the predicted target audio information and the sample target audio information, and the second loss is used to characterize the difference between the predicted target audio information;

[0051] Performing a fusion process on the first loss and the second loss to obtain the separation loss of the initial audio separation network.

[0052] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement any one of the audio data separation methods provided in the first aspect.

[0053] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute any one of the audio data separation methods provided in the first aspect.

[0054] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, which includes instructions, when the instructions are executed by a processor of an electronic device, enabling the electronic device to execute any one of the audio data separation methods provided in the first aspect.

[0055] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0056] The embodiments of the present disclosure provide an audio data separation method, apparatus, electronic device and storage medium, which can perform transformation processing on the audio data to be processed to obtain the spectral features corresponding to the audio data to be processed, wherein the audio data to be processed includes target audio information of multiple audio types. Further, the spectral features can be separated through an audio separation network to obtain the predicted spectral features corresponding to each target audio information in the audio data to be processed, wherein the audio separation network is a convolutional self-attention mechanism model based on an encoder-decoder structure. Still further, inverse transformation processing can be performed on the predicted spectral features corresponding to each target audio information to obtain each target audio information in the audio data to be processed. Based on the audio data separation method, apparatus, electronic device and storage medium provided by the embodiments of the present disclosure, audio separation of the audio data to be processed can be realized through an audio separation network constructed by a convolutional self-attention mechanism model based on an encoder-decoder structure, and each target audio information in the audio data to be processed can be obtained. Since the audio separation network constructed by the convolutional self-attention mechanism model based on the encoder-decoder structure can simultaneously obtain the local features and global dependence information of the audio data to be processed, the embodiments of the present disclosure improve the accuracy of audio separation and the audio separation effect. Moreover, since the audio separation does not use a recurrent neural network, the network complexity can be reduced and the audio separation efficiency can be improved.

[0057] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings

[0058] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an undue limitation of the present disclosure.

[0059] Figure 1 It is a flowchart of a method for separating audio data shown according to an exemplary embodiment.

[0060] Figure 2 It is a structural block diagram of an audio separation network shown according to an exemplary embodiment.

[0061] Figure 3 It is a flowchart of step 104 shown according to an exemplary embodiment.

[0062] Figure 4 It is a structural block diagram of a convolutional self-attention mechanism model shown according to an exemplary embodiment.

[0063] Figure 5 It is a structural block diagram of an attention module shown according to an exemplary embodiment.

[0064] Figure 6 It is a flowchart of step 304 shown according to an exemplary embodiment.

[0065] Figure 7 It is a flowchart of a method for separating audio data shown according to an exemplary embodiment.

[0066] Figure 8 It is a flowchart of step 710 shown according to an exemplary embodiment.

[0067] Figure 9 It is a block diagram of an audio data separation device shown according to an exemplary embodiment.

[0068] Figure 10 It is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners

[0069] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0070] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present disclosure are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0071] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties.

[0072] In one embodiment, as Figure 1 shown, an audio data separation method is provided. In this embodiment, an example is given where this method is applied to a terminal. It can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0073] In step 102, the audio data to be processed is subjected to a transformation process to obtain the spectral characteristics corresponding to the audio data to be processed, where the audio data to be processed includes target audio information of multiple audio types.

[0074] In the embodiments of the present disclosure, the audio data to be processed can be mixed audio data including multiple sound signals, which at least includes target audio information of multiple audio types. In the embodiments of the present disclosure, no specific limitation is made on the number of audio types included in the audio data to be processed. Exemplarily, the audio data to be processed can include 3 audio types: speech, music, and background noise. Then, the target audio information in the audio data to be processed includes: speech audio information of the speech type, music audio information of the music type, and audio information of the background noise of the background noise type.

[0075] For example, the spectral characteristics corresponding to the audio data to be processed can be obtained by subjecting the audio data to be processed to a transformation process. Exemplarily, the short-time Fourier transform can be used to transform the audio data to be processed to obtain the spectral characteristics corresponding to the audio data to be processed. In the embodiments of the present disclosure, no specific limitation is made on the method of the transformation process for obtaining the spectral characteristics. Any method that can transform the audio data to be processed into the corresponding spectral characteristics is applicable to the embodiments of the present disclosure. For example: wavelet transform, non-linear transform based on neural network, Mel-spectrum feature transform and other transformation methods.

[0076] In step 104, the spectral features are separated through an audio separation network to obtain the predicted spectral features corresponding to each target audio information in the audio data to be processed, where the audio separation network is a convolutional self-attention mechanism model based on an encoder-decoder structure.

[0077] In the embodiments of the present disclosure, the audio separation network can be pre-trained. The audio separation network can be a network constructed based on a conformer (Convolution-augmented Transformer for Speech Recognition) model with an encoder-decoder structure. This audio separation network can be used to separate the audio data to be processed. Its input information is the spectral features of the audio data to be processed, and the output information is the predicted spectral features corresponding to multiple target audio information in the audio data to be processed. Among them, the training process of the audio separation network is not elaborated in the embodiments of the present disclosure, and any neural network training method is applicable to the training of the audio separation network.

[0078] In one example, the output of each decoder can also be fused with the audio features (for example: matrix multiplication, splicing, non-linear transformation, etc. The specific fusion method is not specifically limited in the embodiments of the present disclosure and will not be specially described in the following embodiments) to obtain the predicted spectral features corresponding to multiple target audio information in the audio data to be processed.

[0079] After obtaining the spectral features of the audio data to be processed, the spectral features of the audio data to be processed can be input into the audio separation network for separation processing to obtain the predicted spectral features corresponding to each target audio information in the audio data to be processed. For example: when the audio data to be processed includes three audio types of speech, music, and background noise, after separating the spectral features of the audio data to be processed through the audio separation network, the predicted spectral feature 1 corresponding to the speech audio information, the predicted spectral feature 2 corresponding to the music audio information, and the predicted spectral feature 3 corresponding to the background noise audio information in the audio data to be processed can be obtained.

[0080] In step 106, inverse transformation processing is performed on the predicted spectral features corresponding to each target audio information to obtain each target audio information in the audio data to be processed.

[0081] In the embodiments of the present disclosure, after obtaining the predicted spectral features corresponding to each target audio information in the audio data to be processed, the inverse transformation processing can be respectively performed on the predicted spectral features corresponding to each target audio information to obtain each target audio information in the audio data to be processed. Among them, the inverse transformation processing method is the inverse transformation of the transformation processing performed on the audio data to be processed as described above. Exemplarily, in the case where the transformation processing is the short-time Fourier transform, the inverse transformation processing can be the inverse Fourier transform, that is, the inverse Fourier transform can be used to perform the inverse transformation processing on each predicted spectral feature, and each target audio information in the audio data to be processed can be obtained. At this time, the separation of multiple target audio information from the audio data to be processed is completed.

[0082] Exemplarily, still taking the foregoing example as an example, the inverse transformation processing can be performed on the predicted spectral feature 1 to obtain the speech audio information in the audio data to be processed, the inverse transformation processing can be performed on the predicted spectral feature 2 to obtain the music audio information in the audio data to be processed, and the inverse transformation processing can be performed on the predicted spectral feature 3 to obtain the background noise audio information in the audio data to be processed.

[0083] The embodiments of the present disclosure provide an audio data separation method, which can perform transformation processing on the audio data to be processed to obtain the spectral features corresponding to the audio data to be processed. Among them, the audio data to be processed at least includes target audio information of multiple audio types. Further, the spectral features can be separated through an audio separation network to obtain the predicted spectral features corresponding to each target audio information in the audio data to be processed. Among them, the audio separation network is a network constructed based on a convolutional self-attention mechanism model with a decoding structure. Still further, the inverse transformation processing can be performed on the predicted spectral features corresponding to each target audio information to obtain each target audio information in the audio data to be processed. Based on the audio data separation method provided by the embodiments of the present disclosure, the audio separation of the audio data to be processed can be realized through an audio separation network constructed based on a convolutional self-attention mechanism model with a decoding structure, and each target audio information in the audio data to be processed can be obtained. Since the audio separation network constructed based on a convolutional self-attention mechanism model with a decoding structure can simultaneously obtain the local features and global dependency information of the audio data to be processed, the embodiments of the present disclosure improve the accuracy of audio separation, improve the audio separation effect, and since the audio separation does not use a recurrent neural network, the network complexity can be reduced and the audio separation efficiency can be improved.

[0084] In an exemplary embodiment, as Figure 2 shown, the audio separation network includes an encoder, an attention module, and a plurality of decoders corresponding one-to-one to multiple audio types (exemplarily, Figure 2 the case of 3 audio types is shown), as Figure 3As shown, in step 104, the spectral features are separated by an audio separation network to obtain the predicted spectral features corresponding to each target audio information in the audio data to be processed, which may include:

[0085] In step 302, the spectral features are encoded by an encoder to obtain the audio features of the audio data to be processed;

[0086] In step 304, the audio features are feature-extracted by an attention module to obtain the attention features corresponding to each target audio information;

[0087] In step 306, the attention features corresponding to each target audio information are decoded by the decoders corresponding to each target audio information to obtain the predicted spectral features corresponding to each target audio information.

[0088] In the embodiments of the present disclosure, the audio separation network may include an encoder, an attention module, and multiple decoders, and the multiple decoders correspond to multiple audio types one by one. Among them, the encoder is used to extract the audio features of the audio data to be processed from the spectral features of the audio data to be processed. The attention module is used to respectively extract the features useful for the target audio information of each audio type from the audio features of the audio data to be processed, that is, to extract the attention features corresponding to each audio type in the audio features, and input them into the decoders corresponding to each target audio information. The decoder can be used to perform feature extraction on the attention features of the corresponding audio type, so as to obtain the predicted spectral features of the target audio information of the audio type, and then through the inverse transformation process of the predicted spectral features of each target audio information, each target audio information can be obtained.

[0089] Exemplarily, still taking the above example as an example, the audio separation network includes an encoder, an attention module, and decoder 1, decoder 2, and decoder 3. After transforming the audio data to be processed into spectral features, the spectral features can be input into the encoder for encoding processing to output the audio features of the audio data to be processed. The audio features are used as input information to be input into the attention module for feature extraction, and the attention feature 1 corresponding to the speech audio information, the attention feature 2 corresponding to the music audio information, and the attention feature 3 corresponding to the background noise audio information can be obtained respectively. Further, decoder 1 is used to decode the attention feature 1 corresponding to the speech audio information to obtain the predicted spectral feature 1 corresponding to the speech audio information, decoder 2 is used to decode the attention feature 2 corresponding to the music audio information to obtain the predicted spectral feature 2 corresponding to the music audio information, and decoder 3 is used to decode the attention feature 3 corresponding to the background noise audio information to obtain the predicted spectral feature 3 corresponding to the background noise audio information.

[0090] Furthermore, inverse transformation processing can be respectively performed on the predicted spectral feature 1, the predicted spectral feature 2, and the predicted spectral feature 3 to obtain speech audio information, music audio information, and background noise audio information.

[0091] In the audio data separation method provided by the embodiments of the present disclosure, the audio separation network includes a plurality of decoders corresponding one by one to multiple audio types. That is, for the target audio information of any audio type, the embodiments of the present disclosure can use a separate decoder to decode the attention features corresponding to the target audio information of different audio types, which can further improve the accuracy of audio separation and the effect of audio separation.

[0092] In an exemplary embodiment, both the encoder and the decoder are constructed by a convolutional self-attention mechanism model. The convolutional self-attention mechanism model includes a feature excitation layer, and the feature excitation layer is used to perform feature learning from the spatial dimension and the convolutional channel dimension.

[0093] In the embodiments of the present disclosure, the convolutional self-attention mechanism model (conformer model) may include a feature excitation layer, and the feature excitation layer can perform feature learning from the spatial dimension and the convolutional channel dimension through the feature excitation layer to amplify useful features and suppress useless features, thereby improving the feature extraction accuracy of the convolutional self-attention mechanism model.

[0094] Exemplarily, referring to Figure 4 As shown, the convolutional self-attention mechanism model may include a first fully connected module, a multi-head attention module, a multi-convolution module, and a second fully connected module. Among them, the multi-convolution module is sequentially composed of a first feature normalization layer, a first convolutional layer, a gated linear activation layer, a second convolutional layer, a feature excitation layer, a second feature normalization layer, a Swish activation layer, and a third convolutional layer.

[0095] Referring to Figure 2 As shown, in the embodiments of the present disclosure, both the encoder and the decoder can be obtained by stacking a plurality of convolutional self-attention mechanism models. As Figure 4 As shown, the first fully connected module can be connected to the second feature normalization module. After the feature (in the encoder, this feature can be a spectral feature, and in the decoder, this feature can be an attention feature) is input into the first fully connected module for full connection processing, it is output to the second feature normalization module for normalization processing, and the input feature can be further fused to obtain a first feature. The processing process of this first feature can be referred to as shown in the following formula (1).

[0096]

[0097] Among them, Let \(x\) represent the first feature, \(z\) represent the input feature of the fully-connected module, \(FFN()\) represent the fully-connected module, and \(layernorm()\) represent the feature normalization module.

[0098] Referring to Figure 4 As shown, the multi-head attention module can be connected to the third feature normalization module. As the input information of the multi-head attention module, after being input into the multi-head attention module for feature extraction, it can be output to the third feature normalization module for normalization processing, and further fuse the first feature to obtain the second feature. The processing process of the second feature can be referred to as shown in the following formula (2).

[0099]

[0100] Among them, \(z'\) is used to represent the second feature, and \(selfattention()\) is used to represent the multi-head attention module.

[0101] Referring to Figure 4 As shown, the multi-convolution module is connected to the fourth feature normalization module. \(z'\) is used as the input information of the multi-convolution module. After being input into the multi-convolution module for feature processing, it can be output to the fourth feature normalization module for normalization processing, and further fuse the first feature to obtain the third feature. The processing process of the third feature can be referred to as shown in the following formula (3).

[0102] \(z'' = layernorm(conv(z') + z')\) Formula (3)

[0103] Among them, \(z''\) is used to represent the third feature, and \(conv()\) is used to represent the multi-convolution module.

[0104] Within the multi-convolution module, the second feature can be normalized by the first feature normalization layer to obtain the normalized second feature (1). The normalized second feature (1) is input into the first convolutional layer for convolutional processing to obtain the convolved second feature (1). The convolved second feature (1) is input into the gated linear activation layer to obtain the activated second feature (1), and the activated second feature (1) is input into the second convolutional layer to obtain the convolved second feature (2). The convolved second feature (2) is input into the feature excitation layer for feature learning to obtain the learned second feature. The feature excitation layer can be a convolutional network for feature learning from the spatial dimension and the convolutional channel dimension, which can amplify useful features and suppress useless features. The learned second feature is input into the second feature normalization layer for normalization to obtain the normalized second feature (2). The normalized second feature (2) is input into the Swish activation layer for activation processing to obtain the activated second feature (2), and the activated second feature (2) is input into the third convolutional layer for sequential convolutional processing to obtain the output information of this multi-convolution module.

[0105] Referring to Figure 4 As shown, the second fully-connected module is connected to the fifth feature normalization module. "z" is used as the input information of the second fully-connected module. After being input into this second fully-connected module for fully-connected processing, it can be output to the fifth feature normalization module for normalization processing, and further fuse the third feature to obtain the output information of the conformer module. The processing process of this output information can be referred to as shown in the following formula (IV).

[0106]

[0107] Among them, output is used to represent the output information of the conformer module.

[0108] It should be noted that the above first fully-connected module and second fully-connected module can both be implemented by fully-connected layers. The multi-head attention module can be implemented by the multi-head attention mechanism. The first feature normalization layer and the second feature normalization layer in the multi-convolution module can both be implemented by normalization functions. The gated linear activation layer and the Swish activation layer can both be implemented by activation functions. The first convolutional layer, the second convolutional layer, and the third convolutional layer can all be implemented by convolutional neural networks. The feature excitation layer can be implemented by a convolutional network.

[0109] The audio data separation method provided by the embodiments of the present disclosure, in which both the encoder and the decoder in the audio separation network are composed of stacked convolutional self-attention mechanism models. Since the convolutional self-attention mechanism model can perform feature learning from the spatial dimension and the convolutional channel dimension through the feature excitation layer to amplify useful features and suppress useless features, the separation accuracy of the audio separation network can be improved, and the separation effect can be enhanced.

[0110] In an exemplary embodiment, referring to Figure 5 as shown, the attention module may include a plurality of attention mechanisms corresponding one-to-one to various audio types ( Figure 5 the case of 3 audio types is shown), the attention mechanism includes a convolutional module and a first feature normalization module. Referring to Figure 6 as shown, in step 304, the audio features are subjected to feature extraction through the attention module to obtain the attention features corresponding to each target audio information, which can be specifically implemented through the following steps:

[0111] In step 602, for any attention mechanism, the audio features are subjected to feature extraction through the convolutional module in the attention mechanism to obtain the initial attention features of the target audio information corresponding to the attention mechanism;

[0112] In step 604, the initial attention features are subjected to normalization processing through the first feature normalization module in the attention mechanism to obtain the normalized initial attention features, and the normalized initial attention features are fused with the audio features to obtain the attention features of the target audio information corresponding to the attention mechanism.

[0113] In the embodiments of the present disclosure, the attention module may include a plurality of attention mechanisms, and thus the feature information of each audio type in the audio features can be extracted respectively through the plurality of attention mechanisms to obtain the attention features corresponding to each target audio information.

[0114] Among them, the attention module may include a convolutional module and a first feature normalization module. After the encoder encodes the spectral features corresponding to the audio data to be processed to obtain audio features, the audio features can be respectively input into the convolutional modules of each attention mechanism for non-linear processing to extract the initial attention features of the audio type corresponding to the attention mechanism from the audio features (the initial attention features may be feature maps, and the present disclosure does not specifically limit the manifestation form of the initial attention features).

[0115] Further, the initial attention features obtained by the convolutional module can be input into the first feature normalization module for normalization processing to obtain the normalized initial attention features, and the normalized initial attention features can be used as the attention features of the target audio information of the audio type corresponding to the attention mechanism.

[0116] Exemplarily, still taking the foregoing example as an example, the target audio information includes speech audio information, music information, and background noise audio information. Correspondingly, the attention module includes attention mechanism 1, attention mechanism 2, and attention mechanism 3. After obtaining the audio features corresponding to the audio data to be processed, the audio features can be input into attention mechanism 1, attention mechanism 2, and attention mechanism 3 respectively for feature extraction. Attention mechanism 1 extracts features related to the speech audio information from the audio features. After fusing with the audio features, the attention feature 1 of the speech audio information can be obtained. Attention mechanism 2 extracts features related to the music audio information from the audio features. After fusing with the audio features, the attention feature 2 of the music audio information can be obtained. Attention mechanism 3 extracts features related to the background noise audio information from the audio features. After fusing with the audio features, the attention feature 3 of the background noise audio information can be obtained.

[0117] Further, the attention feature 1 can be input into the corresponding decoder 1 for decoding, and the predicted spectral feature 1 corresponding to the speech audio information in the audio data to be processed can be obtained. The attention feature 2 can be input into the corresponding decoder 2 for decoding, and the predicted spectral feature 2 corresponding to the music audio information in the audio data to be processed can be obtained. The attention feature 3 can be input into the corresponding decoder 3 for decoding, and the predicted spectral feature 3 corresponding to the background noise audio information in the audio data to be processed can be obtained. And inverse transformation processing is respectively performed on the predicted spectral feature 1, the predicted spectral feature 2, and the predicted spectral feature 3, and the speech audio information, the music audio information, and the background noise audio information in the audio data to be processed can be obtained.

[0118] In the audio data separation method provided by the embodiments of the present disclosure, the attention module in the audio separation network can respectively extract the features of each target audio information from the audio features corresponding to the audio data to be processed through multiple attention mechanisms, so as to further decode the attention features of each target audio information through multiple decoders, which can further improve the separation accuracy of the audio separation network and improve the separation effect.

[0119] In an exemplary embodiment, referring to Figure 5 as shown, the attention module may further include a dimension increasing module, and the attention mechanism may further include a dimension decreasing module. In step 304, when the attention module extracts features from the audio features to obtain the attention features corresponding to each target audio information, it may further include:

[0120] The dimension of the audio features is expanded through the dimension increasing module to obtain the expanded audio features;

[0121] In step 602, the convolutional module in the attention mechanism is used to extract features from the audio features, and the initial attention features of the target audio information corresponding to the attention mechanism are obtained, which may include:

[0122] The convolutional module in the attention mechanism is used to extract features from the upsampled audio features, and the initial attention features of the target audio information corresponding to the attention mechanism are obtained;

[0123] In step 604, the second feature normalization module in the attention mechanism is used to normalize the initial attention features and fuse them with the audio features to obtain the attention features of the target audio information, which may include:

[0124] The normalization module in the attention mechanism is used to normalize the initial attention features to obtain the normalized attention features;

[0125] The dimensionality reduction module in the attention mechanism is used to perform dimensionality reduction on the normalized attention features to obtain the attention features of the target audio information corresponding to the attention mechanism.

[0126] In the embodiments of the present disclosure, before the attention module extracts features from the audio features, the upsampling module may be used to expand the dimensions of the audio features to obtain more refined audio features. The upsampling module may be implemented by a convolutional neural network, and any convolutional neural network that can perform dimensionality expansion is applicable to the embodiments of the present disclosure. The embodiments of the present disclosure do not make specific limitations on the upsampling module.

[0127] After the upsampling module expands the dimensions of the audio features, the audio data processed by each attention mechanism is the upsampled audio features. After each attention mechanism extracts features from the upsampled audio features, the dimensionality reduction module needs to perform dimensionality reduction on the obtained normalized attention features (the normalized initial attention features) to obtain the attention features of each target audio information.

[0128] In the audio data separation method provided by the embodiments of the present disclosure, the attention module in the audio separation network may expand the dimensions of the audio features corresponding to the audio data to be processed through the upsampling module, so that the information of the audio features is more refined. Then, the attention mechanism can more accurately extract the attention features corresponding to each target audio information, further improving the separation accuracy of the audio separation network and the separation effect.

[0129] In an exemplary embodiment, referring to Figure 7 As shown, before step 104, where the audio separation network separates the spectral features to obtain the predicted spectral features corresponding to each target audio information in the audio data to be processed, the method may further include:

[0130] In step 702, a sample group is obtained. The sample group includes sample audio data and annotation information of the sample audio data. The annotation information of the sample audio data includes sample target audio information of multiple audio types used to constitute the sample audio data;

[0131] In step 704, the sample audio data in the sample group is subjected to transformation processing to obtain the sample spectral features of the sample audio data;

[0132] In step 706, the sample spectral features are separated through an initial audio separation network to obtain the sample spectral features corresponding to each sample target audio information in the sample audio data;

[0133] In step 708, inverse transformation processing is performed on the sample spectral features corresponding to each sample target audio information to obtain multiple predicted target audio information;

[0134] In step 710, according to the multiple predicted target audio information and the multiple sample target audio information corresponding to the sample audio data, the separation loss of the initial audio separation network is determined;

[0135] In step 712, the initial audio separation network is trained according to the separation loss to obtain an audio separation network.

[0136] In the embodiments of the present disclosure, a training set for training the audio separation network can be pre-constructed. Exemplarily, sample target audio information of multiple audio types can be collected, and multiple sample target audio information of different audio types can be mixed to obtain sample audio data. Exemplarily, taking the audio types including speech, music, and background noise as an example, multiple speech audio information, multiple music audio information, and multiple background noise audio information can be collected. By fusing 1 speech audio information, 1 music audio information, and 1 background noise audio information, sample audio data can be obtained. By analogy, multiple sample audio data can be obtained. Each sample audio data and the multiple sample target audio information constituting the sample audio data form a sample group, and multiple sample groups can be used to construct a training set.

[0137] After obtaining the training set, the initial audio separation network can be trained using the sample groups in the training set to obtain the audio separation network. In the embodiments of the present disclosure, the sample audio data in the diverse group can be transformed to obtain the sample spectral features of the sample audio data. Then, the sample spectral features of the sample audio data are used as the input information of the initial audio separation network. The network structure of the initial audio separation network can refer to the relevant descriptions in the foregoing embodiments, and will not be elaborated herein. After the initial audio separation network separates the sample spectral features, the sample spectral features corresponding to each sample target audio information in the sample audio data can be obtained. Then, inverse transformation processing can be performed on each sample spectral feature to obtain each predicted target audio information in the sample audio data.

[0138] After obtaining multiple predicted target audio information, the separation loss of the initial audio separation network can be obtained through the difference between the multiple predicted target audio information and the multiple sample target audio information. Then, after adjusting the network parameters of the initial audio separation network through the separation loss, the adjusted initial audio separation network can be continuously iteratively trained until the separation loss of the initial audio separation network meets the training requirements (for example: the separation loss is less than a preset loss threshold), and the training is stopped to obtain the audio separation network.

[0139] The audio data separation method provided by the embodiments of the present disclosure can pre-train an audio separation network constructed based on a convolutional self-attention mechanism model with an encoder-decoder structure, and then perform audio separation on the audio data to be processed through the audio separation network, which can improve the audio separation efficiency and the audio separation result.

[0140] In an exemplary embodiment, referring to Figure 8 as shown, in the above step 710, according to the multiple predicted target audio information and the multiple sample target audio information corresponding to the sample audio data, determining the separation loss of the initial audio separation network can be implemented through the following steps:

[0141] In step 802, according to each predicted target audio information and each sample target audio information corresponding to the sample audio data, the first loss and the second loss of the initial audio separation network are determined. The first loss is used to characterize the difference between the predicted target audio information and the sample target audio information, and the second loss is used to characterize the difference between the predicted target audio information.

[0142] In step 804, the first loss and the second loss are fused to obtain the separation loss of the initial audio separation network.

[0143] In the embodiments of the present disclosure, the first loss of the initial audio separation network can be obtained through the difference between each predicted target audio information and the corresponding sample target audio information of each predicted target audio information. And the second loss of the initial audio separation network can be obtained through the difference between each predicted target audio information and the sample target audio information corresponding to other predicted target audio information. After fusing the first loss and the second loss, the total separation loss of the initial audio separation network is obtained.

[0144] Among them, the first loss is a loss based on amplitude, and the second loss is used to measure the difference between the predicted target audio information and other target audio information except the sample target audio information corresponding to the predicted target audio information. The greater the difference, the better the separation effect. In the embodiments of the present disclosure, the determination methods of the first loss and the second loss are not specifically limited. For example, loss functions such as the Loss loss function, L1 regularization, and L2 regularization are all applicable to the embodiments of the present disclosure.

[0145] Exemplarily, the determination processes of the first loss, the second loss, and the separation loss can refer to the following formulas (five), (six), and (seven).

[0146]

[0147]

[0148]

[0149] Among them, represents the first loss, represents the second loss, represents the separation loss, i and j respectively represent the labels of the predicted target audio information, X i represents the i-th predicted target audio information, S i represents the sample target audio information corresponding to the i-th predicted target audio information, S j represents the j-th sample target audio information, and λ is a hyperparameter used to balance each loss term in the loss function so that these two losses are on the same scale numerically.

[0150] The first loss based on amplitude realizes intra-class compactness, and the second loss based on inter-class improves inter-class separability, which can enhance the discrimination ability of the audio separation network for the separation output categories (speech, music, and noise). Therefore, based on the audio separation method provided by the embodiments of the present disclosure, training the audio separation network through the first loss and the second loss can improve the accuracy of the audio separation network.

[0151] To enable those skilled in the art to better understand the embodiments of the present disclosure, the following will illustrate the embodiments of the present disclosure through specific examples.

[0152] In the embodiments of the present disclosure, an audio separation network is pre-trained. The audio separation network may include: (1) an encoder: processing the spectral features of the audio data to be processed to generate an audio feature representation of the audio data to be processed. (2) an attention module: based on the audio features, selecting and extracting beneficial features for each audio track of each audio type, and connecting the encoder and the decoder. (3) a plurality of decoders: respectively processing the audio features selected by the attention module for each audio type and generating predicted spectral features of the separated target audio information. That is, after the audio data to be processed is subjected to a transformation process, the spectral features enter the audio separation network and then enter the encoder - attention module - decoder in sequence, and finally generate the predicted spectral features of the separated target audio information for each audio type. After an inverse transformation process, the separated target audio information for each audio type can be obtained.

[0153] Exemplarily, first, the audio data to be processed is converted into spectral features by using the short-time Fourier transform. The encoder is used to extract the basic audio features. The input of the encoder is the spectral features of the audio data to be processed, and the output is the audio features of the audio data to be processed, as shown in the following formula (VIII).

[0154] E O = Encoder(|Y(t,f)|) Formula (VIII)

[0155] where |Y(t,f)| is the spectral features of the audio data to be processed, and E o represents the audio features output by the encoder. The encoder is composed of a plurality of Conformer modules.

[0156] The attention module is used to dynamically focus on the features that are only useful for different audio types. The processing process of the attention module is shown in the following formula (IX).

[0157] A 1 ,A 2 ,A 3 = Attention(E O ) Formula (IX)

[0158] where A 1 ,A 2 ,A 3 are respectively used to represent the attention features of the target audio information of three different audio types (in this example, they can represent the attention features corresponding to speech, music, and background noise), and Attention(E o ) represents the attention module.

[0159] The attention module can use a convolutional network to expand the dimensions of the audio features of the audio data to be processed, so as to obtain a fine feature representation. Then, three convolutional layers are further used to perform a non-linear transformation on the feature representation, and three audio types, namely speech, music, and background noise, are respectively modeled to learn the importance of each audio type in the feature representation, obtain the initial attention features of different audio types, and further apply a sigmoid network for normalization processing. Then, matrix multiplication is performed between the audio features and the initial attention features to obtain the normalized attention features of each audio type. The normalized attention features contain rich information selected from the encoder. Finally, a convolutional network is used for feature dimensionality reduction to obtain the attention features of each audio type.

[0160] Since different audio types, namely speech, music, and background noise, have different acoustic features, three decoders are used to model each audio type respectively, as shown in the following formula (ten).

[0161] |X i (t,f)| = Decoder i ×|Y(t,f)|, i = 1, 2, 3 Formula (ten)

[0162] where, |X i (t,f)| is the predicted spectral feature of each audio type. The decoder for each predicted audio feature consists of multiple Conformer modules, and Decoder i represents the i-th decoder. Then, the inverse Fourier transform is used to transform the predicted spectral feature to obtain the target audio signal.

[0163] In the embodiments of the present disclosure, the Conformer structure is extended to the audio separation task. The Conformer structure can simultaneously process the local features of audio data and capture the global dependencies of the audio data signal. Therefore, in the embodiments of the present disclosure, the local and global information of the audio data to be processed is obtained through the Conformer structure, which can improve the performance of audio data separation.

[0164] Considering the different characteristics of audio information of different audio types, embodiments of the present disclosure propose an audio separation network composed of a Conformer structure and based on an encoding and decoding framework. Specifically, the same encoder is shared in the audio separation network, the attention module selects useful features for each audio type, and separate decoders generate separated spectral features for each audio type. During the process of training the audio separation network, the audio separation network can be trained by a first loss and a second loss (for specific descriptions, refer to the relevant descriptions of the foregoing embodiments). The first loss can enhance intra-class compactness, and the second loss can increase the distance between the target audio information and the audio information of other audio types, increasing the distance between the separated outputs, enhancing the discrimination ability of the network, and alleviating the problem of misclassification.

[0165] It should be understood that although Figures 1 - 8 the steps in the flowchart of Figures 1 - 8 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,

[0166] It can be understood that the same / similar parts among the various embodiments of the methods in this specification can be referred to each other. Each embodiment focuses on the differences from other embodiments. For the relevant parts, refer to the descriptions of other method embodiments.

[0167] Figure 9 is a block diagram of an audio data separation device shown according to an exemplary embodiment. Referring to Figure 9 the figure, the device includes a first transformation unit 902, a first separation unit 904, and a first inverse transformation unit 906.

[0168] The first transformation unit 902 is configured to perform a transformation process on the audio data to be processed to obtain spectral features corresponding to the audio data to be processed, where the audio data to be processed includes target audio information of multiple audio types;

[0169] The first separation unit 904 is configured to perform a separation process on the spectral features through an audio separation network to obtain predicted spectral features corresponding to each of the target audio information in the audio data to be processed, where the audio separation network is a convolutional self-attention mechanism model based on an encoding and decoding structure;

[0170] The first inverse transformation unit 906 is configured to perform an inverse transformation process on the predicted spectral features corresponding to each of the target audio information to obtain each of the target audio information in the audio data to be processed.

[0171] An embodiment of the present disclosure provides an audio data separation device, which can perform a transformation process on audio data to be processed to obtain spectral features corresponding to the audio data to be processed, where the audio data to be processed includes target audio information of multiple audio types. Further, the spectral features can be separated through an audio separation network to obtain predicted spectral features corresponding to each target audio information in the audio data to be processed, where the audio separation network is a convolutional self-attention mechanism model based on an encoder-decoder structure. Still further, an inverse transformation process can be performed on the predicted spectral features corresponding to each target audio information to obtain each target audio information in the audio data to be processed. Based on the audio data separation device provided by the embodiment of the present disclosure, audio separation of the audio data to be processed can be realized through an audio separation network constructed by a convolutional self-attention mechanism model based on an encoder-decoder structure, and each target audio information in the audio data to be processed can be obtained. Since the audio separation network constructed by the convolutional self-attention mechanism model based on the encoder-decoder structure can simultaneously obtain local features and global dependence information of the audio data to be processed, the embodiment of the present disclosure improves the accuracy of audio separation, improves the audio separation effect, and since the audio separation does not use a recurrent neural network, the network complexity can be reduced and the audio separation efficiency can be improved.

[0172] In an exemplary embodiment, the audio separation network includes an encoder, an attention module, and a plurality of decoders corresponding to the multiple audio types one by one. The first separation unit 904 is further configured to perform:

[0173] Encoding the spectral features through the encoder to obtain audio features of the audio data to be processed;

[0174] Extracting features from the audio features through the attention module to obtain attention features corresponding to each of the target audio information;

[0175] Decoding the attention features corresponding to each of the target audio information through the decoders corresponding to each of the target audio information to obtain predicted spectral features corresponding to each of the target audio information.

[0176] In an exemplary embodiment, both the encoder and the decoder are constructed by a convolutional self-attention mechanism model, and the convolutional self-attention mechanism model includes a feature excitation layer, and the feature excitation layer is used for feature learning from the spatial dimension and the convolutional channel dimension.

[0177] In an exemplary embodiment, the attention module includes a plurality of attention mechanisms corresponding one-to-one to the multiple audio types. The attention mechanism includes a convolution module and a first feature normalization module. The first separation unit 604 is further configured to perform:

[0178] For any one of the attention mechanisms, perform feature extraction on the audio features through the convolution module in the attention mechanism to obtain the initial attention features of the target audio information corresponding to the attention mechanism;

[0179] Perform normalization processing on the initial attention features through the first feature normalization module in the attention mechanism to obtain the normalized initial attention features, and fuse the normalized initial attention features with the audio features to obtain the attention features of the target audio information corresponding to the attention mechanism.

[0180] In an exemplary embodiment, the apparatus further includes:

[0181] Obtain a training set, where the training set includes a plurality of sample groups. The sample group includes the sample audio data and the annotation information of the sample audio data. The annotation information of the sample audio data includes the sample target audio information of the multiple audio types used to constitute the sample audio data;

[0182] A second transformation unit, configured to perform transformation processing on the sample audio data in the sample group to obtain the sample spectral features of the sample audio data;

[0183] A second separation unit, configured to perform separation processing on the sample spectral features through the initial audio separation network to obtain the sample spectral features corresponding to each sample target audio information in the sample audio data;

[0184] A second inverse transformation unit, configured to perform inverse transformation processing on the sample spectral features corresponding to each sample target audio information to obtain a plurality of predicted target audio information;

[0185] A determination unit, configured to determine the separation loss of the initial audio separation network according to the plurality of predicted target audio information and the plurality of sample target audio information corresponding to the sample audio data;

[0186] A training unit, configured to train the initial audio separation network according to the separation loss to obtain the audio separation network.

[0187] In an exemplary embodiment, the determination unit is further configured to perform:

[0188] Determine a first loss and a second loss of the initial audio separation network according to each of the predicted target audio information and each of the sample target audio information corresponding to the sample audio data, where the first loss is used to characterize the difference between the predicted target audio information and the sample target audio information, and the second loss is used to characterize the difference between the predicted target audio information;

[0189] Perform a fusion process on the first loss and the second loss to obtain a separation loss of the initial audio separation network.

[0190] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0191] Figure 10 It is a block diagram of an electronic device 1000 for an audio data separation method shown according to an exemplary embodiment. For example, the electronic device 1000 can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0192] Refer to Figure 10 , the electronic device 1000 may include one or more of the following components: a processing component 1002, a memory 1004, a power supply component 1006, a multimedia component 1008, an audio component 1010, an input / output (I / O) interface 1012, a sensor component 1014, and a communication component 1016.

[0193] The processing component 1002 generally controls the overall operation of the electronic device 1000, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 1002 may include one or more processors 1020 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 1002 may include one or more modules to facilitate the interaction between the processing component 1002 and other components. For example, the processing component 1002 may include a multimedia module to facilitate the interaction between the multimedia component 1008 and the processing component 1002.

[0194] The memory 1004 is configured to store various types of data to support the operation of the electronic device 1000. Examples of such data include instructions for any application or method operating on the electronic device 1000, contact data, phone book data, messages, pictures, videos, etc. The memory 1004 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, optical disk, or graphene memory.

[0195] The power supply component 1006 provides power to various components of the electronic device 1000. The power supply component 1006 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 1000.

[0196] The multimedia component 1008 includes a screen that provides an output interface between the electronic device 1000 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 1008 includes a front camera and / or a rear camera. When the electronic device 1000 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0197] The audio component 1010 is configured to output and / or input audio signals. For example, the audio component 1010 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 1000 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 1004 or transmitted via the communication component 1016. In some embodiments, the audio component 1010 further includes a speaker for outputting audio signals.

[0198] The I / O interface 1012 provides an interface between the processing component 1002 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons may include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.

[0199] The sensor assembly 1014 includes one or more sensors for providing status assessment of various aspects for the electronic device 1000. For example, the sensor assembly 1014 can detect the on / off state of the electronic device 1000, the relative positioning of components, such as the display and keypad of the electronic device 1000. The sensor assembly 1014 can also detect changes in the position of the electronic device 1000 or components of the electronic device 1000, the presence or absence of user contact with the electronic device 1000, the orientation or acceleration / deceleration of the device 1000, and temperature changes of the electronic device 1000. The sensor assembly 1014 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 1014 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 1014 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0200] The communication component 1016 is configured to facilitate communication between the electronic device 1000 and other devices in a wired or wireless manner. The electronic device 1000 can access a wireless network based on communication standards, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 1016 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1016 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0201] In an exemplary embodiment, the electronic device 1000 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above methods.

[0202] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as the memory 1004 including instructions, and the above instructions can be executed by the processor 1020 of the electronic device 1000 to complete the above methods. For example, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, magnetic tape, a floppy disk, and an optical data storage device, etc.

[0203] In an exemplary embodiment, a computer program product is further provided. The computer program product includes instructions that can be executed by the processor 1020 of the electronic device 1000 to implement the above method.

[0204] It should be noted that the above-mentioned apparatus, electronic device, computer-readable storage medium, computer program product, etc. may also include other implementation manners according to the description of the method embodiments. The specific implementation manners can refer to the description of the relevant method embodiments and will not be elaborated here one by one.

[0205] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed herein. The specification and embodiments are to be considered as exemplary only, and the true scope and spirit of the present disclosure are pointed out by the claims.

[0206] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. An audio data separation method, characterized in that, comprising: performing transformation processing on the audio data to be processed to obtain the spectral features corresponding to the audio data to be processed, wherein the audio data to be processed includes target audio information of multiple audio types; performing separation processing on the spectral features through an audio separation network to obtain the predicted spectral features corresponding to each target audio information in the audio data to be processed; performing inverse transformation processing on the predicted spectral features corresponding to each target audio information to obtain each target audio information in the audio data to be processed; wherein, the audio separation network includes an encoder, an attention module, and multiple decoders corresponding to the multiple audio types one by one, and the performing separation processing on the spectral features through the audio separation network to obtain the predicted spectral features corresponding to each target audio information in the audio data to be processed includes: performing encoding processing on the spectral features through the encoder to obtain the audio features of the audio data to be processed; performing feature extraction on the audio features through the attention module to obtain the attention features corresponding to each target audio information; respectively performing decoding processing on the attention features corresponding to the target audio information of each audio type through the decoders corresponding to each audio type to obtain the predicted spectral features corresponding to each target audio information; the encoder and the decoder are both constructed by a convolutional self-attention mechanism model, and the convolutional self-attention mechanism model includes a feature excitation layer, and the feature excitation layer is used for feature learning from the spatial dimension and the convolutional channel dimension.

2. The method according to claim 1, characterized in that, the attention module includes multiple attention mechanisms corresponding to the multiple audio types one by one, the attention mechanism includes a convolutional module and a first feature normalization module, and the performing feature extraction on the audio features through the attention module to obtain the attention features corresponding to each target audio information includes: for any one of the attention mechanisms, performing feature extraction on the audio features through the convolutional module in the attention mechanism to obtain the initial attention features of the target audio information corresponding to the attention mechanism; performing normalization processing on the initial attention features through the first feature normalization module in the attention mechanism to obtain the normalized initial attention features, and fusing the normalized initial attention features with the audio features to obtain the attention features of the target audio information corresponding to the attention mechanism.

3. The method according to claim 1 or 2, characterized in that, before the performing separation processing on the spectral features through the audio separation network to obtain the predicted spectral features corresponding to each target audio information in the audio data to be processed, the method further includes: obtaining a sample group, the sample group includes sample audio data and the annotation information of the sample audio data, and the annotation information of the sample audio data includes sample target audio information of the multiple audio types used to constitute the sample audio data; Perform transformation processing on the sample audio data in the sample group to obtain the sample spectral features of the sample audio data; Perform separation processing on the sample spectral features through an initial audio separation network to obtain the sample spectral features corresponding to each sample target audio information in the sample audio data; Perform inverse transformation processing on the sample spectral features corresponding to each sample target audio information to obtain multiple predicted target audio information; Determine the separation loss of the initial audio separation network according to the multiple predicted target audio information and the multiple sample target audio information corresponding to the sample audio data; Train the initial audio separation network according to the separation loss to obtain the audio separation network.

4. The method according to claim 3, wherein, the determining the separation loss of the initial audio separation network according to the multiple predicted target audio information and the multiple sample target audio information corresponding to the sample audio data includes: Determine a first loss and a second loss of the initial audio separation network according to each predicted target audio information and each sample target audio information corresponding to the sample audio data, where the first loss is used to characterize the difference between the predicted target audio information and the sample target audio information, and the second loss is used to characterize the difference between the predicted target audio information; Perform fusion processing on the first loss and the second loss to obtain the separation loss of the initial audio separation network.

5. An audio data separation device, wherein, comprising: A first transformation unit configured to perform transformation processing on the audio data to be processed to obtain the spectral features corresponding to the audio data to be processed, wherein the audio data to be processed includes target audio information of multiple audio types; A first separation unit configured to perform separation processing on the spectral features through an audio separation network to obtain the predicted spectral features corresponding to each target audio information in the audio data to be processed, wherein the audio separation network is a convolutional self-attention mechanism model based on an encoder-decoder structure; A first inverse transformation unit configured to perform inverse transformation processing on the predicted spectral features corresponding to each target audio information to obtain each target audio information in the audio data to be processed; wherein the audio separation network includes an encoder, an attention module, and multiple decoders corresponding to the multiple audio types one by one, and the first separation unit is further configured to perform: Encode the spectral features through the encoder to obtain the audio features of the audio data to be processed; Extract features of the audio features through the attention module to obtain the attention features corresponding to each target audio information; Decode the attention features corresponding to the target audio information of each audio type through the decoders corresponding to each audio type respectively to obtain the predicted spectral features corresponding to each target audio information; The encoder and the decoder are both constructed by the convolutional self-attention mechanism model, and the convolutional self-attention mechanism model includes a feature excitation layer, which is used to perform feature learning from the spatial dimension and the convolutional channel dimension.

6. The apparatus according to claim 5, wherein, the attention module includes a plurality of attention mechanisms corresponding to the multiple audio types one by one, the attention mechanism includes a convolutional module and a first feature normalization module, and the first separation unit is further configured to perform: For any one of the attention mechanisms, perform feature extraction on the audio features through the convolutional module in the attention mechanism to obtain an initial attention feature of the target audio information corresponding to the attention mechanism; Perform normalization processing on the initial attention feature through the first feature normalization module in the attention mechanism to obtain a normalized initial attention feature, and fuse the normalized initial attention feature with the audio features to obtain an attention feature of the target audio information corresponding to the attention mechanism.

7. The apparatus according to claim 5 or 6, wherein, the apparatus further includes: an acquisition unit configured to acquire a sample group, the sample group including sample audio data and annotation information of the sample audio data, and the annotation information of the sample audio data includes sample target audio information of the multiple audio types used to form the sample audio data; a second transformation unit configured to perform transformation processing on the sample audio data in the sample group to obtain sample spectral features of the sample audio data; a second separation unit configured to perform separation processing on the sample spectral features through an initial audio separation network to obtain sample spectral features corresponding to each sample target audio information in the sample audio data; a second inverse transformation unit configured to perform inverse transformation processing on the sample spectral features corresponding to each sample target audio information to obtain a plurality of predicted target audio information; a determination unit configured to determine a separation loss of the initial audio separation network according to the plurality of predicted target audio information and the plurality of sample target audio information corresponding to the sample audio data; a training unit configured to train the initial audio separation network according to the separation loss to obtain the audio separation network.

8. The apparatus according to claim 7, wherein, the determination unit is further configured to perform: Determine a first loss and a second loss of the initial audio separation network according to each predicted target audio information and each sample target audio information corresponding to the sample audio data, the first loss is used to characterize the difference between the predicted target audio information and the sample target audio information, and the second loss is used to characterize the difference between the predicted target audio information; Perform fusion processing on the first loss and the second loss to obtain a separation loss of the initial audio separation network.

9. An electronic device, wherein, it includes: a processor; a memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the audio data separation method according to any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the audio data separation method according to any one of claims 1 to 4.

11. A computer program product, the computer program product includes instructions, characterized in that when the instructions are executed by a processor of an electronic device, the electronic device is enabled to execute the audio data separation method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Audio processing method and device, terminal and storage medium

    CN113113040A

  • Multi-person voice separation method and training method of voice separation model

    CN113744753A