Audio separation method and device, equipment, storage medium and program product
By introducing vocalprint perception and feature fusion methods into audio separation technology, the problem of difficulty in separating vocal and instrumental components in complex music environments is solved, and a more efficient audio separation effect is achieved.
Patent Information
- Application Number
- CN202510036609.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-09
AI Technical Summary
In complex music environments, traditional sound source separation technology has limited effects. Neural network-based audio separation technology also has challenges when separating instrument sound components from human voices, especially when the instrument sound components are similar in spectral structure to human voices, the model is prone to be separated by mistake, resulting in unsatisfactory separation of separation effects.
A method of audio separation based on voiceprint perception is proposed. By extracting the mixed audio, the voiceprint features of human voices are extracted, and fusions them with the original features to generate a vocal mask and a background mask, thereby achieving accurate separation of human voice and background sound.
By extracting and fusion of voiceprint features, the vocal components can be more accurately identified, the separation effect of vocals and background sounds can be improved, and the audio separation quality can be improved.
Smart Images

Figure CN119993190A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing technology, and in particular to an audio separation method, device, equipment, storage medium and program product. Background Art
[0002] In the field of music production and audio processing, the separation of human voice and background sound has always been an important research topic. Traditional sound source separation technology often relies on spectral analysis and statistical models, but the effect is limited in complex music environments. With the rapid development of deep learning technology, audio separation technology based on neural networks has gradually been applied, but it still faces certain challenges in separating human voice and background sound in music data, especially when the spectral structure of instrument sound components and human voice is very similar, which will cause the model to mistakenly separate the instrument sound components into the human voice, and the separation effect is not ideal. Summary of the invention
[0003] Based on the above technical status, the present application proposes an audio separation method, device, equipment, storage medium and program product, which can improve the separation effect of human voice and background sound.
[0004] The first aspect of the present application provides an audio separation method, comprising:
[0005] Performing feature extraction processing on the mixed audio to obtain a first audio feature; the mixed audio includes human voice audio and background audio;
[0006] Extracting a voiceprint feature of a human voice from the first audio feature to obtain a second audio feature including the voiceprint feature;
[0007] fusing the first audio feature and the second audio feature to obtain a third audio feature;
[0008] Based on the third audio feature, human voice audio and background audio are extracted from the mixed audio.
[0009] In some implementations, extracting human voice audio and background audio from the mixed audio based on the third audio feature includes:
[0010] generating a vocal mask and a background mask based on the third audio feature;
[0011] Based on the vocal mask and the background mask, vocal audio and background audio are extracted from the mixed audio.
[0012] In some implementations, performing feature extraction processing on the mixed audio to obtain a first audio feature, extracting a voiceprint feature of a human voice from the first audio feature to obtain a second audio feature including the voiceprint feature, fusing the first audio feature and the second audio feature to obtain a third audio feature, and generating a human voice mask and a background mask based on the third audio feature, including:
[0013] Inputting the audio features of the mixed audio into the audio separation model to obtain a voice mask and a background mask output by the audio separation model;
[0014] Wherein, the audio separation model includes:
[0015] A first feature modeling module is used to perform feature extraction processing on the mixed audio to obtain a first audio feature;
[0016] A voiceprint perception module, configured to extract a voiceprint feature of a human voice from the first audio feature to obtain a second audio feature including the voiceprint feature;
[0017] A second feature modeling module, configured to fuse the first audio feature and the second audio feature to obtain a third audio feature;
[0018] A mask estimation module is used to generate a vocal mask and a background mask based on the third audio feature.
[0019] In some implementations, performing feature extraction processing on the mixed audio to obtain the first audio feature includes:
[0020] Dividing the audio features of the mixed audio into sub-bands to obtain multiple sub-band audio features;
[0021] The multiple sub-band audio features are jointly modeled in the time dimension and the frequency dimension to obtain a first audio feature.
[0022] In some implementations, the performing joint modeling of the multiple sub-band audio features in terms of time dimension and frequency dimension to obtain the first audio feature includes:
[0023] Performing context joint modeling processing in a time dimension on the multiple sub-band audio features to obtain a first global feature;
[0024] Performing dimensionality reduction processing on the first global feature, and superimposing the first global feature after the dimensionality reduction processing with the multiple sub-band audio features to obtain a second global feature;
[0025] Performing context joint modeling processing in a frequency dimension on the second global feature to obtain a third global feature;
[0026] A dimension reduction process is performed on the third global feature, and the third global feature after the dimension reduction process is superimposed on the second global feature to obtain a first audio feature.
[0027] In some implementations, the sub-band division of the audio features of the mixed audio to obtain a plurality of sub-band audio features includes:
[0028] According to the relationship that the bandwidth of the divided sub-band audio feature is proportional to the frequency of the sub-band audio feature, the audio feature of the mixed audio is divided into sub-bands to obtain multiple sub-band audio features.
[0029] In some implementations, extracting a voiceprint feature of a human voice from the first audio feature to obtain a second audio feature including the voiceprint feature includes:
[0030] Inputting the first audio feature into a voiceprint perception model, and the voiceprint perception model extracting a voiceprint feature of a human voice from the first audio feature to obtain a second audio feature including the voiceprint feature;
[0031] The voiceprint perception model is obtained by extracting voiceprint features from input audio feature samples, and performing speaker recognition training based on the extracted audio features containing the voiceprint features.
[0032] In some implementations, the fusing the first audio feature and the second audio feature to obtain a third audio feature includes:
[0033] concatenating the first audio feature and the second audio feature to obtain a concatenated feature;
[0034] The concatenated features are jointly modeled in terms of time dimension and frequency dimension to obtain a third audio feature.
[0035] In some implementations, the performing joint modeling of the splicing feature in the time dimension and the frequency dimension to obtain the third audio feature includes:
[0036] Performing context joint modeling processing in a time dimension on the spliced features to obtain a first fusion feature;
[0037] Performing dimensionality reduction processing on the first fused feature, and superimposing the first fused feature after the dimensionality reduction processing with the splicing feature to obtain a second fused feature;
[0038] Performing context joint modeling processing in the frequency dimension on the second fused feature to obtain a third fused feature;
[0039] Perform dimensionality reduction processing on the third fusion feature, and superimpose the third fusion feature after the dimensionality reduction processing with the second fusion feature to obtain a third audio feature.
[0040] In some implementations, generating a vocal mask and a background mask based on the third audio feature includes:
[0041] Extracting features corresponding to respective sub-bands from the third audio features; the respective sub-bands are determined by dividing the audio features of the mixed audio into sub-bands;
[0042] Based on the features corresponding to each sub-band, determining the estimation results of the vocal mask and the background mask corresponding to each sub-band;
[0043] The estimation results of the vocal mask and the background mask corresponding to each sub-band are combined to determine the vocal mask and the background mask.
[0044] A second aspect of the present application provides an audio separation device, comprising:
[0045] A first feature processing unit is used to perform feature extraction processing on the mixed audio to obtain a first audio feature; the mixed audio includes human voice audio and background audio;
[0046] a voiceprint feature extraction unit, configured to extract a voiceprint feature of a human voice from the first audio feature to obtain a second audio feature including the voiceprint feature;
[0047] A second feature processing unit, configured to fuse the first audio feature and the second audio feature to obtain a third audio feature;
[0048] The audio separation unit is used to extract human voice audio and background audio from the mixed audio based on the third audio feature.
[0049] A third aspect of the present application provides an electronic device, including a memory and a processor;
[0050] The memory is connected to the processor and is used to store programs;
[0051] The processor is used to implement the above-mentioned audio separation method by running the program in the memory.
[0052] In a fourth aspect, the present application proposes a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned audio separation method is implemented.
[0053] A fifth aspect of the present application provides a computer program product, comprising computer program instructions, which, when executed by a processor, enable the processor to perform the above-mentioned audio separation method.
[0054] The audio separation method proposed in the present application extracts voiceprint features from the first audio feature of the mixed audio, and after fusing the second audio feature containing the voiceprint features with the first audio feature, separates the user audio from the fused third audio feature. The third audio feature used for audio separation contains the voiceprint features of the human voice in the mixed audio, which is conducive to accurately identifying the human voice audio component from the mixed audio based on the third audio feature, thereby achieving more accurate separation of the human voice and background sound, and improving the quality of audio separation. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0056] Figure 1 A flowchart of an audio separation method provided in an embodiment of the present application.
[0057] Figure 2 A structural diagram of an audio separation model provided in an embodiment of the present application.
[0058] Figure 3 A schematic diagram of the structure of another audio separation model provided in an embodiment of the present application.
[0059] Figure 4 A schematic diagram of the structure of a sub-band division module provided in an embodiment of the present application.
[0060] Figure 5 A schematic diagram of the structure of the intra-subband and inter-subband modeling modules provided in an embodiment of the present application.
[0061] Figure 6 A schematic diagram of the structure of two intra-sub-band and inter-sub-band modeling modules connected in series provided in an embodiment of the present application.
[0062] Figure 7 A schematic diagram of the structure of the voiceprint perception model provided in an embodiment of the present application.
[0063] Figure 8 A schematic diagram of the structure of a mask estimation module provided in an embodiment of the present application.
[0064] Fig. 9A schematic diagram of the structure of an audio separation device provided in an embodiment of the present application.
[0065] Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0066] The technical solution of the embodiment of the present application is applicable to audio separation application scenarios, especially to application scenarios of separating human voice audio and background audio from mixed audio. The technical solution of the embodiment of the present application can more accurately separate human voice audio and background audio from mixed audio.
[0067] In the field of music production and audio processing, the separation of human voice and background sound has always been an important research topic. Traditional sound source separation technology often relies on spectrum analysis and statistical models, but its effect is limited in complex music environments.
[0068] With the rapid development of deep learning technology, audio separation technology based on neural networks has gradually been applied, but it still faces certain challenges in separating vocals and background sounds in music data. The existing neural network-based music vocal and background sound separation method only relies on data-driven methods to implicitly learn the differences between vocals and background sounds to complete the separation task. However, for unseen data, due to generalization problems, the separation effect is poor, especially when the instrument sound components and vocals are very similar in spectral structure, the model will mistakenly separate the instrument sound components into vocals, resulting in unsatisfactory separation effect.
[0069] In response to the above technical problems, this application proposes an audio separation method based on voiceprint perception. This method takes into account the uniqueness and stability of voiceprint features, which is conducive to helping the model identify the human voice component in music and effectively separate it, thereby achieving more accurate separation of human voice and background sound.
[0070] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0071] The audio separation method proposed in the embodiments of the present application can be exemplarily applied to audio processing devices or systems, such as computers, servers, workstations, smart terminals, handheld terminals, wearable devices and other devices with audio processing functions, or to systems composed of these devices and the audio processing systems within these devices, and can be specifically executed by processors in the above-mentioned devices or systems.
[0072] The present application embodiment first proposes an audio separation method, see Figure 1 As shown, the method includes:
[0073] S101. Perform feature extraction processing on the mixed audio to obtain a first audio feature.
[0074] The mixed audio includes vocal audio and background audio, that is, the mixed audio can be obtained by mixing the vocal audio and the background audio.
[0075] When the mixed audio is obtained, the time domain signal of the mixed audio is firstly extracted. For example, the signal of the mixed audio is normalized to the maximum amplitude, and the signal amplitude of the mixed audio is normalized to the range of [-1, 1]. Then, the normalized mixed audio is converted into a spectrogram through short-time Fourier transform (STFT), and then the logarithmic power spectrum (lps) feature is obtained by the following formula: lps = ln (Y2), where Y represents the spectrogram of the mixed audio.
[0076] The above-mentioned logarithmic power spectrum feature can be used as the first audio feature of the mixed audio. In other embodiments, the above-mentioned logarithmic power spectrum feature can be further processed, such as performing convolution processing based on context interaction, further extracting deep features from the above-mentioned logarithmic power spectrum feature, and using the processed features as the first audio feature. In addition, audio feature extraction can also be performed on the above-mentioned mixed audio through other audio feature extraction methods, and the extracted audio features can be used as the first audio features.
[0077] S102: extracting a voiceprint feature of a human voice from the first audio feature to obtain a second audio feature including the voiceprint feature.
[0078] Specifically, the voiceprint feature of human voice refers to the voiceprint feature of human voice reflected by the human voice audio in the mixed audio. The voiceprint feature can be used to reflect the voiceprint characteristics of the speaker who generates the human voice audio, and can be used to identify the human voice audio.
[0079] Commonly used voiceprint features include short-time energy, zero-crossing rate, Mel-frequency cepstral coefficients, etc.
[0080] Exemplarily, in an embodiment of the present application, a voiceprint feature of a human voice is extracted from a first audio feature through a voiceprint extraction model to obtain a second audio feature including the voiceprint feature.
[0081] That is, the above-mentioned first audio feature is input into a pre-trained voiceprint extraction model, and the voiceprint extraction model obtains the second audio feature including the voiceprint feature by extracting the voiceprint feature of the human voice from the first audio feature.
[0082] The above-mentioned voiceprint extraction model may adopt a Gaussian mixture model GMM, a support vector machine SVM, a deep neural network DNN or other model architectures. When the above-mentioned voiceprint extraction model is trained, the audio features of the sample audio from which the voiceprint features are to be extracted are input into the voiceprint extraction model, the voiceprint extraction model extracts the voiceprint features from the audio features, and the speaker is identified based on the voiceprint features output by the voiceprint extraction model. When the speaker can be accurately identified based on the voiceprint features output by the voiceprint extraction model, it can be considered that the voiceprint extraction model can accurately extract the speaker's voiceprint features.
[0083] It should be noted that the embodiment of the present application extracts the voiceprint features of human voices from the first audio features, which is to extract the voiceprint features of the audios of each speaker contained in the mixed audio. That is, if the mixed audio contains the human voices of several speakers, the voiceprint features of the human voices corresponding to the several speakers are extracted, and the second audio feature containing the voiceprint features finally obtained is the second audio feature containing the voiceprint features of each speaker in the mixed audio.
[0084] S103: Fuse the first audio feature and the second audio feature to obtain a third audio feature.
[0085] Specifically, after extracting the voiceprint feature from the first audio feature, the second audio feature including the voiceprint feature is fused with the original first audio feature to obtain the third audio feature.
[0086] The above fusion may be performed by concatenating and then interacting with the context features to obtain the third audio feature. Alternatively, the above fusion may be performed by superimposing the corresponding position features of the second audio feature and the first audio feature.
[0087] S104: Extract human voice audio and background audio from the mixed audio based on the third audio feature.
[0088] Specifically, the third audio feature obtained through the above processing includes both the complete mixed audio feature (first audio feature) and the voiceprint feature (second audio feature), so it can be used to separate the human voice audio and the background audio of the mixed audio.
[0089] In some embodiments, a vocal mask and a background mask are first generated based on the third audio feature.
[0090] Then, based on the above-mentioned vocal mask and background mask, the vocal audio and background audio are extracted from the mixed audio.
[0091] Exemplarily, the third audio feature can be input into a pre-trained audio separation model, such as BSRNN, etc. The model processes the third audio feature to separate the vocal mask and the background mask therefrom, and then multiplies the vocal mask and the background mask with the spectrogram of the mixed audio respectively to obtain the spectrum of the vocal audio and the spectrum of the background audio, and then performs inverse Fourier transform on the spectrum of the vocal audio and the spectrum of the background audio to obtain the vocal audio and the background audio in the time domain.
[0092] From the above introduction, it can be seen that the audio separation method proposed in the embodiment of the present application extracts the voiceprint feature from the first audio feature of the mixed audio, and after fusing the second audio feature containing the voiceprint feature with the first audio feature, separates the fused third audio feature from the user audio. The third audio feature used for audio separation contains the voiceprint feature of the human voice in the mixed audio, which is conducive to accurately identifying the human voice audio component from the mixed audio based on the third audio feature, thereby achieving more accurate separation of the human voice and background sound, and improving the quality of audio separation.
[0093] In another embodiment of the present application, an audio separation model is also proposed, through which the audio separation method proposed in the above embodiment of the present application can be executed.
[0094] See also Figure 2 As shown, the audio separation model includes a first feature modeling module, a voiceprint perception module, a second feature modeling module and a mask estimation module.
[0095] The first feature modeling module is used to perform feature extraction processing on the mixed audio to obtain a first audio feature;
[0096] A voiceprint perception module, configured to extract a voiceprint feature of a human voice from the first audio feature to obtain a second audio feature including the voiceprint feature;
[0097] A second feature modeling module, configured to fuse the first audio feature and the second audio feature to obtain a third audio feature;
[0098] A mask estimation module is used to generate a vocal mask and a background mask based on the third audio feature.
[0099] Next, in combination with the specific structure of each module of the above-mentioned audio separation model, the specific implementation process of the audio separation method proposed in the embodiment of the present application is introduced.
[0100] Before the mixed audio is input into the audio separation model for processing, the mixed audio signal is first normalized to the maximum amplitude, and the signal amplitude of the mixed audio is normalized to the range of [-1, 1]. Then, the normalized mixed audio is converted into a spectrogram through short-time Fourier transform (STFT), and then the logarithmic power spectrum (lps) feature is obtained based on the spectrogram. Figure 2 As shown, the logarithmic power spectrum (lps) feature is used as the audio feature of the mixed audio and input into the audio separation model for processing.
[0101] In some embodiments disclosed, see Figure 3 As shown, the above-mentioned first feature modeling module includes a sub-band division module and an intra-sub-band and inter-sub-band modeling module.
[0102] The sub-band division module is used to divide the audio features of the mixed audio into sub-bands to obtain multiple sub-band audio features.
[0103] The intra-sub-band and inter-sub-band modeling modules are used to perform joint modeling of the multiple sub-band audio features in the time dimension and the frequency dimension to obtain the first audio feature.
[0104] In some embodiments, see Figure 4 As shown, the above subband division module is a linear transformation unit, which consists of K normalization layers and K fully connected layers with an output dimension of N.
[0105] The sub-band division module divides the audio features of the mixed audio into sub-bands to obtain multiple sub-band audio features, specifically:
[0106] First, the sub-band division module divides the LPS features of the mixed audio into sub-bands according to a predefined bandwidth to obtain K sub-band LPS features.
[0107] The above predefined bandwidth is And satisfy F is the frequency range of the mixed audio.
[0108] Then, the obtained K sub-band lps features are passed through K normalization layers and fully connected layers respectively to obtain K sub-band audio features.
[0109] In other embodiments, when the audio features of the mixed audio are divided into sub-bands, the audio features of the mixed audio are divided into sub-bands according to the relationship that the bandwidth of the divided sub-band audio features is proportional to the frequency of the sub-band audio features, so as to obtain a plurality of sub-band audio features. That is, the audio features of the mixed audio are divided into sub-bands in a non-uniform manner in which the bandwidth of the sub-band divided in the low frequency band is narrower and the bandwidth of the sub-band divided in the high frequency band is wider, and the non-uniform sub-bands are unified into the same dimension N for merging to obtain the respective sub-band audio features.
[0110] In another embodiment, see Figure 5 As shown, the above-mentioned intra-subband and inter-subband modeling module is a modeling unit composed of a GRU, which consists of a layer normalization layer, a bidirectional GRU layer in the time dimension, a fully connected layer, a layer normalization layer, a bidirectional GRU layer in the frequency dimension, and a fully connected layer.
[0111] Based on the structure of the above-mentioned intra-subband and inter-subband modeling modules, the intra-subband and inter-subband modeling modules jointly model the multiple sub-band audio features in the time dimension and the frequency dimension to obtain the first audio feature, which specifically includes the following processing A1-A4:
[0112] A1. Use the layer normalization layer and the bidirectional GRU layer in the time dimension to perform context-joint modeling processing on the multiple sub-band audio features in the time dimension, and the K sub-bands share the same bidirectional GRU layer to obtain a first global feature.
[0113] A2. Use the fully connected layer to perform a dimensionality reduction operation on the first global feature obtained in step A1, keep the input channel and the channel of the sub-band audio feature unchanged, and then superimpose the first global feature after the dimensionality reduction processing with the multiple sub-band audio features to obtain a second global feature.
[0114] A3. Use the layer normalization layer and the bidirectional GRU layer in the frequency dimension to perform context joint modeling processing on the second global feature obtained in step A2 in the sub-band frequency dimension, and share the same bidirectional GRU layer in the time dimension to obtain a third global feature.
[0115] A4. Perform a dimensionality reduction operation on the third global feature obtained in step A3 using the fully connected layer, keep the input channel the same as the channel of the second global feature described in A2, and then superimpose the third global feature after the dimensionality reduction processing with the second global feature to obtain a first audio feature.
[0116] In another embodiment, the above-mentioned intra-subband and inter-subband modeling modules may be a series combination of two or more intra-subband and inter-subband modeling modules, such as Figure 6 As shown, the two intra-subband and inter-subband modeling modules can be connected in series to jointly model the audio features. In this way, through multiple intra-subband and inter-subband modeling processes, the first audio feature can contain more feature information, that is, the first audio feature can be more accurate.
[0117] In another embodiment, the voiceprint perception module extracts the voiceprint feature of the human voice from the first audio feature to obtain the second audio feature including the voiceprint feature, which is specifically implemented by a voiceprint perception model.
[0118] That is, the first audio feature is input into the voiceprint perception model, and the voiceprint perception model extracts the voiceprint feature of the human voice from the first audio feature to obtain the second audio feature including the voiceprint feature;
[0119] The voiceprint perception model is obtained by extracting voiceprint features from input audio feature samples, and performing speaker recognition training based on the extracted audio features containing the voiceprint features.
[0120] Specifically, when training the above-mentioned voiceprint perception model, the audio features of the sample audio from which the voiceprint features are to be extracted are input into the voiceprint perception model, the voiceprint perception model extracts the voiceprint features from the audio features, and performs speaker recognition based on the voiceprint features output by the voiceprint perception model. When the voiceprint features output by the voiceprint perception model can accurately identify the speaker, it can be considered that the voiceprint perception model can accurately extract the speaker's voiceprint features.
[0121] In another embodiment, the voiceprint perception model is formed by stacking three identical residual convolutional coding units. Figure 7 As shown, the residual convolutional coding unit consists of a two-dimensional convolutional layer, a layer normalization layer, a ReLU activation function, a FSMN (Feedforward Sequential Memory Networks) unit, a layer normalization layer and a residual module. Based on the structure of the above-mentioned voiceprint perception model, the voiceprint perception model extracts the voiceprint feature of the human voice from the first audio feature to obtain the second audio feature containing the voiceprint feature, which specifically includes the following steps B1-B4:
[0122] B1. The two-dimensional convolution layer performs a convolution operation on the first audio feature through a set of two-dimensional convolution filters of size k×k to obtain a local feature vector.
[0123] B2, the layer normalization layer and the ReLU activation function process the local feature vector obtained in step B1 to obtain a second feature vector.
[0124] B3. Use the FSMN unit of size h to process the second feature vector obtained in step B2 to obtain a third feature vector.
[0125] B4. Use the layer normalization layer and the residual module to process the third feature vector and superimpose the processing result with the first audio feature to obtain a second audio feature containing a voiceprint feature.
[0126] In another embodiment, when the second feature modeling module fuses the first audio feature and the second audio feature, the first audio feature and the second audio feature are first spliced to obtain a spliced feature, and then the spliced feature is jointly modeled in the time dimension and the frequency dimension to obtain a third audio feature.
[0127] That is, the first audio feature and the second audio feature are first concatenated to merge the two. Then, the concatenated feature is contextually jointly modeled in the time dimension and the frequency dimension, so as to fully integrate the context feature information in the concatenated feature from the time dimension and the frequency dimension to obtain the third audio feature.
[0128] In another embodiment, the second feature modeling module can be used Figure 5 The structural form of the intra-subband and inter-subband modeling modules shown. Based on the above structural form, when the second feature modeling module performs joint modeling of the splicing feature in the time dimension and the frequency dimension, it specifically includes the processing steps C1-C4:
[0129] C1. Use a layer normalization layer and a bidirectional GRU layer in the time dimension to perform contextual joint modeling on the concatenated features of the first audio feature and the second audio feature in the time dimension. The K sub-bands share the same bidirectional GRU layer to obtain the first fusion feature.
[0130] C2. Use a fully connected layer to perform dimensionality reduction processing on the first fusion feature obtained in step C1, keep the input channel and the spliced feature channel after the first audio feature and the second audio feature are spliced unchanged, and then superimpose the first fusion feature after the dimensionality reduction processing with the spliced feature to obtain the second fusion feature.
[0131] C3. Use the layer normalization layer and the bidirectional GRU layer in the frequency dimension to perform context joint modeling on the second fused feature obtained in step C2 in the sub-band frequency dimension. The time dimension shares the same bidirectional GRU layer to obtain the third fused feature.
[0132] C4. Use a fully connected layer to perform dimensionality reduction processing on the third fusion feature obtained in step C3, keep the input channel the same as the channel of the second fusion feature obtained in step C2, and then superimpose the third fusion feature after dimensionality reduction processing with the second fusion feature to obtain a third audio feature.
[0133] In another embodiment, the second feature modeling module adopts Figure 6 The two intra-subband and inter-subband modeling modules are shown in series, that is, the second feature modeling module is composed of two layers of intra-subband and inter-subband modeling modules.
[0134] Based on this structural form, the first-layer intra-sub-band and inter-sub-band modeling module performs the above-mentioned C1-C4 processing on the spliced features after the first audio feature and the second audio feature are spliced. Then, the first-layer intra-sub-band and inter-sub-band modeling module superimposes the third fusion feature after the dimensionality reduction processing with the second fusion feature by executing the above-mentioned step C4, and uses the superimposed fusion feature as the input of the second-layer intra-sub-band and inter-sub-band modeling module. The second-layer intra-sub-band and inter-sub-band modeling module processes the superimposed fusion feature according to the above-mentioned steps C1-C4 to finally obtain the third audio feature.
[0135] In this way, multiple intra-sub-band and inter-sub-band modeling processes can be performed so that the third audio feature finally obtained contains more feature information, that is, the third audio feature is more accurate.
[0136] In other embodiments, the second feature modeling module may also be composed of more layers of intra-sub-band and inter-sub-band modeling modules, so as to achieve more in-depth feature extraction and obtain more accurate third audio features.
[0137] In another embodiment, see Figure 8 As shown in Figure 1, the mask estimation module consists of K multi-layer perception units, each of which consists of a normalization layer, two fully connected layers, a ReLU activation function, and a sigmoid activation function.
[0138] Based on the above structure, the mask estimation module generates a human voice mask and a background mask based on the third audio feature, specifically including the following processing steps D1-D4:
[0139] D1. Split the third audio feature into K sub-bands to obtain features corresponding to each sub-band. The sub-band division method of the third audio feature is the same as the sub-band division method of the audio feature of the mixed audio described in the above embodiment.
[0140] D2, layer normalization layer normalizes the features corresponding to the K sub-bands obtained in step D1 to obtain K normalized features.
[0141] D3. Use K fully connected layers and ReLU activation function to process the K normalized features obtained in step D2 to obtain K hidden layer feature vectors.
[0142] D4, using K fully connected layers and sigmoid activation function to process the K hidden layer feature vectors obtained in step D3, and obtain K mask values, each of which includes a vocal mask and a background mask. Finally, the estimated results of the vocal mask and background mask corresponding to each of the K sub-bands are merged to obtain the final vocal mask and background mask.
[0143] Corresponding to the above-mentioned audio separation method, the present application embodiment also provides an audio separation device, see Fig. 9 As shown, the device comprises:
[0144] The first feature processing unit 100 is used to perform feature extraction processing on the mixed audio to obtain a first audio feature; the mixed audio includes human voice audio and background audio;
[0145] A voiceprint feature extraction unit 110, configured to extract a voiceprint feature of a human voice from the first audio feature to obtain a second audio feature including the voiceprint feature;
[0146] A second feature processing unit 120, configured to fuse the first audio feature and the second audio feature to obtain a third audio feature;
[0147] The audio separation unit 130 is used to extract human voice audio and background audio from the mixed audio based on the third audio feature.
[0148] In some implementations, the audio separation unit 130 extracts the human voice audio and the background audio from the mixed audio based on the third audio feature, including:
[0149] generating a vocal mask and a background mask based on the third audio feature;
[0150] Based on the vocal mask and the background mask, vocal audio and background audio are extracted from the mixed audio.
[0151] In some implementations, the first feature processing unit 100 performs feature extraction processing on the mixed audio to obtain a first audio feature, the voiceprint feature extraction unit 110 extracts the voiceprint feature of the human voice from the first audio feature to obtain a second audio feature including the voiceprint feature, the second feature processing unit 120 fuses the first audio feature and the second audio feature to obtain a third audio feature, and the audio separation unit 130 generates a human voice mask and a background mask based on the third audio feature, including:
[0152] Inputting the audio features of the mixed audio into the audio separation model to obtain a voice mask and a background mask output by the audio separation model;
[0153] Wherein, the audio separation model includes:
[0154] A first feature modeling module is used to perform feature extraction processing on the mixed audio to obtain a first audio feature;
[0155] A voiceprint perception module, configured to extract a voiceprint feature of a human voice from the first audio feature to obtain a second audio feature including the voiceprint feature;
[0156] A second feature modeling module, configured to fuse the first audio feature and the second audio feature to obtain a third audio feature;
[0157] A mask estimation module is used to generate a vocal mask and a background mask based on the third audio feature.
[0158] In some implementations, the first feature processing unit 100 performs feature extraction processing on the mixed audio to obtain the first audio feature, including:
[0159] Dividing the audio features of the mixed audio into sub-bands to obtain multiple sub-band audio features;
[0160] The multiple sub-band audio features are jointly modeled in the time dimension and the frequency dimension to obtain a first audio feature.
[0161] In some implementations, the first feature processing unit 100 performs joint modeling of the multiple sub-band audio features in the time dimension and the frequency dimension to obtain the first audio feature, including:
[0162] Performing context joint modeling processing in a time dimension on the multiple sub-band audio features to obtain a first global feature;
[0163] Performing dimensionality reduction processing on the first global feature, and superimposing the first global feature after the dimensionality reduction processing with the multiple sub-band audio features to obtain a second global feature;
[0164] Performing context joint modeling processing in a frequency dimension on the second global feature to obtain a third global feature;
[0165] A dimension reduction process is performed on the third global feature, and the third global feature after the dimension reduction process is superimposed on the second global feature to obtain a first audio feature.
[0166] In some implementations, the first feature processing unit 100 performs sub-band division on the audio features of the mixed audio to obtain a plurality of sub-band audio features, including:
[0167] According to the relationship that the bandwidth of the divided sub-band audio feature is proportional to the frequency of the sub-band audio feature, the audio feature of the mixed audio is divided into sub-bands to obtain multiple sub-band audio features.
[0168] In some implementations, the voiceprint feature extraction unit 110 extracts the voiceprint feature of the human voice from the first audio feature to obtain the second audio feature including the voiceprint feature, including:
[0169] Inputting the first audio feature into a voiceprint perception model, and the voiceprint perception model extracting a voiceprint feature of a human voice from the first audio feature to obtain a second audio feature including the voiceprint feature;
[0170] The voiceprint perception model is obtained by extracting voiceprint features from input audio feature samples, and performing speaker recognition training based on the extracted audio features containing the voiceprint features.
[0171] In some implementations, the second feature processing unit 120 fuses the first audio feature and the second audio feature to obtain a third audio feature, including:
[0172] concatenating the first audio feature and the second audio feature to obtain a concatenated feature;
[0173] The concatenated features are jointly modeled in terms of time dimension and frequency dimension to obtain a third audio feature.
[0174] In some implementations, the second feature processing unit 120 performs joint modeling of the splicing feature in the time dimension and the frequency dimension to obtain a third audio feature, including:
[0175] Performing context joint modeling processing in a time dimension on the spliced features to obtain a first fusion feature;
[0176] Performing dimensionality reduction processing on the first fused feature, and superimposing the first fused feature after the dimensionality reduction processing with the splicing feature to obtain a second fused feature;
[0177] Performing context joint modeling processing in the frequency dimension on the second fused feature to obtain a third fused feature;
[0178] Perform dimensionality reduction processing on the third fusion feature, and superimpose the third fusion feature after the dimensionality reduction processing with the second fusion feature to obtain a third audio feature.
[0179] In some implementations, the audio separation unit 130 generates a vocal mask and a background mask based on the third audio feature, including:
[0180] Extracting features corresponding to respective sub-bands from the third audio features; the respective sub-bands are determined by dividing the audio features of the mixed audio into sub-bands;
[0181] Based on the features corresponding to each sub-band, determining the estimation results of the vocal mask and the background mask corresponding to each sub-band;
[0182] The estimation results of the vocal mask and the background mask corresponding to each sub-band are combined to determine the vocal mask and the background mask.
[0183] The audio separation device provided in this embodiment belongs to the same application concept as the audio separation method provided in the above embodiments of this application, and can execute the audio separation method provided in any of the above embodiments of this application, and has the corresponding functional modules and beneficial effects of the execution method. For technical details not fully described in this embodiment, please refer to the specific processing content of the audio separation method provided in the above embodiments of this application, and will not be repeated here.
[0184] The functions implemented by the above-mentioned units may be implemented by the same or different processors respectively, which is not limited in the embodiments of the present application.
[0185] It should be understood that the units in the above devices can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor calls the instructions stored in the memory to implement any of the above methods or realize the functions of each unit of the device, wherein the processor can be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory can be a memory in the device or a memory outside the device. Alternatively, the units in the device can be implemented in the form of hardware circuits, and the functions of some or all units can be realized by designing the hardware circuits. The hardware circuit can be understood as one or more processors; for example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are realized by designing the logical relationship of the components in the circuit; for another example, in another implementation, the hardware circuit can be implemented by PLD, taking FPGA as an example, which can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by the configuration file, so as to realize the functions of some or all of the above units. All units of the above devices can be implemented in the form of a processor calling software, or in the form of hardware circuits, or in part by a processor calling software, and the remaining part is implemented in the form of hardware circuits.
[0186] In an embodiment of the present application, a processor is a circuit with the ability to process signals. In one implementation, the processor may be a circuit with the ability to read and run instructions, such as a CPU, a microprocessor, a GPU, or a DSP; in another implementation, the processor may implement certain functions through the logical relationship of a hardware circuit, and the logical relationship of the hardware circuit is fixed or reconfigurable, such as a hardware circuit implemented by an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, DPU, etc.
[0187] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.
[0188] In addition, all or part of the units in the above device can be integrated together, or can be implemented independently. In one implementation, these units are integrated together and implemented in the form of a SOC. The SOC may include at least one processor for implementing any of the above methods or implementing the functions of each unit of the device. The type of the at least one processor may be different, for example, including a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.
[0189] Another embodiment of the present application also provides an electronic device, see Fig.10 As shown, the device includes:
[0190] Memory 200 and processor 210;
[0191] The memory 200 is connected to the processor 210 and is used to store programs;
[0192] The processor 210 is used to implement the audio separation method disclosed in any of the above embodiments by running the program stored in the memory 200.
[0193] Specifically, the electronic device may further include: a bus, a communication interface 220 , an input device 230 and an output device 240 .
[0194] The processor 210, the memory 200, the communication interface 220, the input device 230 and the output device 240 are connected to each other via a bus.
[0195] A bus may include a pathway that transfers information between components of a computer system.
[0196] The processor 210 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the scheme of the present invention. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0197] The processor 210 may include a main processor, and may also include a baseband chip, a modem, and the like.
[0198] The memory 200 stores a program for executing the technical solution of the present invention, and may also store an operating system and other key services. Specifically, the program may include a program code, and the program code includes a computer operation instruction. More specifically, the memory 200 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk storage, a flash, and the like.
[0199] The input device 230 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor.
[0200] Output device 240 may include a device that allows information to be output to a user, such as a display screen, a printer, a speaker, etc.
[0201] The communication interface 220 may include any device such as a transceiver to communicate with other devices or communication networks, such as Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.
[0202] The processor 210 executes the program stored in the memory 200 and calls other devices, which can be used to implement each step of any audio separation method provided in the above embodiments of the present application.
[0203] An embodiment of the present application also proposes a chip, which includes a processor and a data interface. The processor reads and runs a program stored in a memory through the data interface to execute the audio separation method introduced in any of the above embodiments. The specific processing process and its beneficial effects can be found in the introduction to the embodiment of the above audio separation method.
[0204] In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the audio separation method described in any of the above embodiments of this specification.
[0205] The computer program product may be written in any combination of one or more programming languages to write program codes for performing the operations of the embodiments of the present application, including object-oriented programming languages, such as Java, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0206] In addition, an embodiment of the present application may also be a storage medium on which a computer program is stored, and the computer program is executed by a processor to execute the steps of the audio separation method described in any of the above embodiments of this specification.
[0207] For the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0208] It should be noted that each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0209] The steps in the methods of each embodiment of the present application can be adjusted in order, combined and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.
[0210] The modules and sub-modules in the devices and terminals of the various embodiments of the present application can be combined, divided and deleted according to actual needs.
[0211] In the several embodiments provided in the present application, it should be understood that the disclosed terminals, devices and methods can be implemented in other ways. For example, the terminal embodiments described above are only schematic, for example, the division of modules or submodules is only a logical function division, and there may be other division methods in actual implementation, for example, multiple submodules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.
[0212] The modules or submodules described as separate components may or may not be physically separated, and the components of the modules or submodules may or may not be physical modules or submodules, that is, they may be located in one place, or they may be distributed on multiple network modules or submodules. Some or all of the modules or submodules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0213] In addition, each functional module or submodule in each embodiment of the present application may be integrated into one processing module, or each module or submodule may exist physically separately, or two or more modules or submodules may be integrated into one module. The above-mentioned integrated modules or submodules may be implemented in the form of hardware or in the form of software functional modules or submodules.
[0214] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0215] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly by hardware, software units executed by a processor, or a combination of the two. The software units may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0216] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0217] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An audio separation method, characterized in that: include: Performing feature extraction processing on the mixed audio to obtain a first audio feature; the mixed audio includes human voice audio and background audio; Extracting a voiceprint feature of a human voice from the first audio feature to obtain a second audio feature including the voiceprint feature; fusing the first audio feature and the second audio feature to obtain a third audio feature; Based on the third audio feature, human voice audio and background audio are extracted from the mixed audio.
2. The method according to claim 1, characterized in that The extracting human voice audio and background audio from the mixed audio based on the third audio feature includes: generating a vocal mask and a background mask based on the third audio feature; Based on the vocal mask and the background mask, vocal audio and background audio are extracted from the mixed audio.
3. The method according to claim 2, characterized in that The method includes performing feature extraction processing on the mixed audio to obtain a first audio feature, extracting a voiceprint feature of a human voice from the first audio feature to obtain a second audio feature including the voiceprint feature, fusing the first audio feature and the second audio feature to obtain a third audio feature, and generating a human voice mask and a background mask based on the third audio feature, including: Inputting the audio features of the mixed audio into the audio separation model to obtain a voice mask and a background mask output by the audio separation model; Wherein, the audio separation model includes: A first feature modeling module is used to perform feature extraction processing on the mixed audio to obtain a first audio feature; A voiceprint perception module, configured to extract a voiceprint feature of a human voice from the first audio feature to obtain a second audio feature including the voiceprint feature; A second feature modeling module, configured to fuse the first audio feature and the second audio feature to obtain a third audio feature; A mask estimation module is used to generate a vocal mask and a background mask based on the third audio feature.
4. The method according to any one of claims 1 to 3, characterized in that The step of performing feature extraction processing on the mixed audio to obtain a first audio feature includes: Dividing the audio features of the mixed audio into sub-bands to obtain multiple sub-band audio features; The multiple sub-band audio features are jointly modeled in the time dimension and the frequency dimension to obtain a first audio feature.
5. The method according to claim 4, characterized in that The step of performing joint modeling of the time dimension and the frequency dimension on the multiple sub-band audio features to obtain the first audio feature includes: Performing context joint modeling processing in a time dimension on the multiple sub-band audio features to obtain a first global feature; Performing dimensionality reduction processing on the first global feature, and superimposing the first global feature after the dimensionality reduction processing with the multiple sub-band audio features to obtain a second global feature; Performing context joint modeling processing in a frequency dimension on the second global feature to obtain a third global feature; A dimension reduction process is performed on the third global feature, and the third global feature after the dimension reduction process is superimposed on the second global feature to obtain a first audio feature.
6. The method according to claim 4, characterized in that The sub-band division of the audio features of the mixed audio to obtain a plurality of sub-band audio features includes: According to the relationship that the bandwidth of the divided sub-band audio feature is proportional to the frequency of the sub-band audio feature, the audio feature of the mixed audio is divided into sub-bands to obtain multiple sub-band audio features.
7. The method according to any one of claims 1 to 3, characterized in that The step of extracting the voiceprint feature of a human voice from the first audio feature to obtain the second audio feature including the voiceprint feature includes: Inputting the first audio feature into a voiceprint perception model, and the voiceprint perception model extracting a voiceprint feature of a human voice from the first audio feature to obtain a second audio feature including the voiceprint feature; The voiceprint perception model is obtained by extracting voiceprint features from input audio feature samples, and performing speaker recognition training based on the extracted audio features containing the voiceprint features.
8. The method according to any one of claims 1 to 3, characterized in that: The fusing the first audio feature and the second audio feature to obtain a third audio feature includes: concatenating the first audio feature and the second audio feature to obtain a concatenated feature; The concatenated features are jointly modeled in terms of time dimension and frequency dimension to obtain a third audio feature.
9. The method according to claim 8, characterized in that The joint modeling of the splicing feature in the time dimension and the frequency dimension to obtain the third audio feature includes: Performing context joint modeling processing in a time dimension on the spliced features to obtain a first fusion feature; Performing dimensionality reduction processing on the first fused feature, and superimposing the first fused feature after the dimensionality reduction processing with the splicing feature to obtain a second fused feature; Performing context joint modeling processing in the frequency dimension on the second fused feature to obtain a third fused feature; Perform dimensionality reduction processing on the third fusion feature, and superimpose the third fusion feature after the dimensionality reduction processing with the second fusion feature to obtain a third audio feature.
10. The method according to claim 2 or 3, characterized in that: The generating a human voice mask and a background mask based on the third audio feature comprises: Extracting features corresponding to respective sub-bands from the third audio features; the respective sub-bands are determined by dividing the audio features of the mixed audio into sub-bands; Based on the features corresponding to each sub-band, determining the estimation results of the vocal mask and the background mask corresponding to each sub-band; The estimation results of the vocal mask and the background mask corresponding to each sub-band are combined to determine the vocal mask and the background mask.
11. An audio separation device, characterized in that: include: A first feature processing unit is used to perform feature extraction processing on the mixed audio to obtain a first audio feature; the mixed audio includes human voice audio and background audio; a voiceprint feature extraction unit, configured to extract a voiceprint feature of a human voice from the first audio feature to obtain a second audio feature including the voiceprint feature; A second feature processing unit, configured to fuse the first audio feature and the second audio feature to obtain a third audio feature; The audio separation unit is used to extract human voice audio and background audio from the mixed audio based on the third audio feature.
12. An electronic device, characterized in that: including memory and processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the audio separation method according to any one of claims 1 to 10 by running the program in the memory.
13. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the audio separation method according to any one of claims 1 to 10 is implemented.
14. A computer program product, characterized in that The method comprises computer program instructions, which, when executed by a processor, enable the processor to perform the audio separation method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Voice data separation method and device, equipment and storage medium
CN113470688A
Multi-person voice separation method and device based on voiceprint features and medium
CN113990344A
Voice separation method based on voiceprint features
CN115240702A
Speaker recognition method and device, electronic equipment and computer readable storage medium
CN115482824A
Directional voice separation method based on deep neural network
CN116030824A