Audio processing method, audio processing device and storage medium
By applying diffusion model and noise compensation technology in audio processing, combining large language model and semantic embedding layer, the problem of unsatisfactory audio separation in the existing technology is solved, and efficient sound source separation in multi-sound source scenarios is achieved.
Patent Information
- Application Number
- CN202311540425.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-20
AI Technical Summary
In the prior art, the audio separation effect is not ideal, and it is difficult to process overlapping sound signals, especially in multi-sound source scenarios, it is difficult to separate audio of unknown sound source types.
The audio processing method based on diffusion model and noise compensation is adopted, and the text description information is processed through a large language model to obtain the semantic embedding layer, combined with the audio feature information to perform noise addition processing, the sound source separation is used for diffusion model, and the audio recovery process is performed through phase information.
The sound source separation effect is achieved in complex scenarios and multi-sound source scenarios, which can effectively separate the target sound source and reduce interference from other sound sources, especially when dealing with unknown sound source types, which shows high separation accuracy.
Smart Images

Figure CN120020947A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of audio processing, and particularly to an audio processing method, an audio processing device, and a storage medium. Background Art
[0002] The multi-source sound separation technology refers to separating the signals of different sound sources from a mixed signal, which is an important research direction in the field of signal processing. In the field of audio signal processing, the multi-source separation technology has become an important foundation for applications such as speech recognition, speech enhancement, and speech synthesis.
[0003] In the related art, the audio separation effect is not ideal, the separated audio is not clear, and it is difficult to process overlapping sound signals. Summary of the Invention
[0004] To overcome the problems existing in the related art, the present disclosure provides an audio processing method, an audio processing device, and a storage medium.
[0005] According to the first aspect of the embodiments of the present disclosure, an audio processing method is provided, including: obtaining an audio to be processed, and obtaining text description information of the audio to be extracted; obtaining first feature information, and obtaining the phase information of the audio to be processed, where the first feature information is the audio feature information of the audio to be processed; according to the text description information and the first feature information, obtaining second feature information through a diffusion model, where the second feature information is the audio feature information of the audio to be extracted; performing audio restoration processing on the second feature information according to the phase information, and determining the audio obtained by the audio restoration processing as the target audio.
[0006] In an implementation manner, the obtaining of the first feature information includes: converting the time-domain signal of the audio to be processed into a frequency-domain signal, and obtaining a spectrogram corresponding to the frequency-domain signal; performing encoding processing on the spectrogram to obtain the first feature information.
[0007] In an implementation manner, the obtaining of the second feature information through the diffusion model according to the text description information and the first feature information includes: processing the text description information through a large language model to obtain a semantic embedding layer; determining the number and recognition type of the audio to be extracted according to the text description information; performing noise addition processing on the first feature information according to the number and recognition type of the audio to be extracted and the semantic embedding layer to obtain noise-added feature information; performing diffusion processing on the first feature information through the diffusion model according to the noise-added feature information and the semantic embedding layer to obtain the second feature information.
[0008] In one implementation, the recognition types of the audio to be extracted include known types and unknown types. The audio of the known type is the same as the audio type of the audio in the diffusion model training data, and the audio of the unknown type is different from the audio type of the audio in the diffusion model training data. The number of the audio to be extracted includes one or more. The adding noise processing of the first feature information according to the number and recognition type of the audio to be extracted and the semantic embedding layer to obtain the noise-added feature information includes: in response to the number being one and the recognition type of the single audio to be extracted being a known type, performing adding noise processing on the first feature information according to the semantic embedding layer to obtain the noise-added feature information; in response to the number being one and the recognition type of the single audio to be extracted being an unknown type, obtaining zero-shot feature information, and performing adding noise processing on the first feature information according to the semantic embedding layer and the zero-shot feature information to obtain the noise-added feature information; in response to the number being multiple and the recognition types of all the multiple audio to be extracted being known types, obtaining noise compensation information, and successively performing adding noise processing on the first feature information according to the semantic embedding layer and the noise compensation information to obtain multiple noise-added feature information; in response to the number being multiple and there being audio of an unknown type among the multiple audio to be extracted, obtaining zero-shot feature information and noise compensation information, and successively performing adding noise processing on the first feature information according to the semantic embedding layer, the zero-shot feature information, and the noise compensation information to obtain multiple noise-added feature information.
[0009] In one implementation, the obtaining of the zero-shot feature information includes: obtaining zero-shot audio of the same audio type according to the audio type of the audio to be extracted of the unknown type; converting the time-domain signal of the zero-shot audio into a frequency-domain signal, and obtaining the spectrogram corresponding to the frequency-domain signal; performing encoding processing on the spectrogram to obtain the zero-shot feature information.
[0010] In one implementation, the noise compensation information is information that is updated successively. The successively performing adding noise processing on the first feature information according to the semantic embedding layer and the noise compensation information to obtain multiple noise-added feature information includes: when performing the first adding noise processing, performing adding noise processing on the first feature information according to the semantic embedding layer to obtain the noise-added feature information; after completing the first adding noise processing, successively obtaining updated noise compensation information, and successively performing adding noise processing on the first feature information according to the semantic embedding layer and the successively updated noise compensation information to obtain multiple noise-added feature information; determining the noise-added feature information obtained when performing the first adding noise processing and the multiple noise-added feature information obtained by the successive adding noise processing as the noise-added feature information of the audio to be obtained.
[0011] In one implementation, the noise compensation information is information updated sequentially. The first feature information is sequentially denoised according to the semantic embedding layer, the zero-shot feature information, and the noise compensation information to obtain multiple pieces of denoised feature information, including: when performing the first denoising process, the first feature information is denoised according to the semantic embedding layer and the zero-shot feature information to obtain denoised feature information, or the first feature information is denoised according to the semantic embedding layer to obtain denoised feature information; after the first denoising process is completed, updated noise compensation information is sequentially obtained, and the first feature information is sequentially denoised according to the semantic embedding layer, the zero-shot feature information, and the sequentially updated noise compensation information to obtain multiple pieces of denoised feature information, or after the first denoising process is completed, updated noise compensation information is sequentially obtained, and the first feature information is sequentially denoised according to the semantic embedding layer and the sequentially updated noise compensation information to obtain multiple pieces of denoised feature information; the denoised feature information obtained when performing the first denoising process and the multiple pieces of denoised feature information obtained after completing the sequential denoising process are determined as the denoised feature information of the audio to be obtained.
[0012] In one implementation, the first feature information is diffused through a diffusion model according to the denoised feature information and the semantic embedding layer to obtain second feature information, including: in response to the number of audio to be extracted being single, the first feature information is diffused through a diffusion model according to the denoised feature information and the semantic embedding layer to obtain a single second feature information; in response to the number of audio to be extracted being multiple, the first feature information is sequentially diffused through a diffusion model according to the semantic embedding layer and the sequentially obtained denoised feature information to sequentially obtain multiple second feature information.
[0013] In one implementation, the method further includes: in response to the denoised feature information not being information obtained based on zero-shot feature information, the first feature information is denoised and diffused through a diffusion model according to the denoised feature information and the semantic embedding layer to eliminate the feature information corresponding to the denoised feature information in the first feature information, obtaining second feature information; in response to the denoised feature information being information obtained based on zero-shot feature information, the first feature information is mapped and diffused through a diffusion model according to the denoised feature information and the semantic embedding layer to retain the feature information corresponding to the denoised feature information in the first feature information, obtaining second feature information.
[0014] In one implementation, the audio restoration process of the second feature information according to the phase information, and determining the audio obtained from the audio restoration process as the target audio includes: in response to the number of audio to be extracted being multiple, performing audio restoration processing on the second feature information obtained successively according to the phase information, and determining the multiple audios obtained from the successive audio restoration processing as the target audio; in response to the number of audio to be extracted being single, performing audio restoration processing on the obtained second feature information according to the phase information, and determining the single audio obtained from the audio restoration processing as the target audio.
[0015] In one implementation, the noise compensation information is obtained in the following manner: according to the audio obtained from the successive audio restoration processing, successively extracting the feature information of the audio, and successively updating the noise compensation information according to the successively extracted feature information.
[0016] According to the second aspect of the embodiments of the present disclosure, there is provided an audio processing device, including: an acquisition unit, configured to acquire the audio to be processed, acquire the text description information of the audio to be extracted, acquire the first feature information, and acquire the phase information of the audio to be processed, where the first feature information is the audio feature information of the audio to be processed; a processing unit, configured to obtain the second feature information through a diffusion model according to the text description information and the first feature information, where the second feature information is the audio feature information of the audio to be extracted; a determination unit, configured to perform audio restoration processing on the second feature information according to the phase information, and determine the audio obtained from the audio restoration processing as the target audio.
[0017] In one implementation, the acquisition unit obtains the first feature information in the following manner: converting the time-domain signal of the audio to be processed into a frequency-domain signal, and obtaining the spectrogram corresponding to the frequency-domain signal; performing encoding processing on the spectrogram to obtain the first feature information.
[0018] In one implementation, the processing unit obtains the second feature information through a diffusion model according to the text description information and the first feature information in the following manner: processing the text description information through a large language model to obtain a semantic embedding layer; determining the number and recognition type of the audio to be extracted according to the text description information; performing noise addition processing on the first feature information according to the number and recognition type of the audio to be extracted and the semantic embedding layer to obtain noise-added feature information; performing diffusion processing on the first feature information through a diffusion model according to the noise-added feature information and the semantic embedding layer to obtain the second feature information.
[0019] In one implementation, the recognition types of the audio to be extracted include known types and unknown types. The audio of the known type is the same as the audio type of the audio in the diffusion model training data, and the audio of the unknown type is different from the audio type of the audio in the diffusion model training data. The number of the audio to be extracted includes single or multiple. The processing unit performs noise addition processing on the first feature information according to the number and recognition type of the audio to be extracted and the semantic embedding layer to obtain noise-added feature information in the following manner: In response to the number being single and the recognition type of the single audio to be extracted being a known type, perform noise addition processing on the first feature information according to the semantic embedding layer to obtain the noise-added feature information; In response to the number being single and the recognition type of the single audio to be extracted being an unknown type, obtain zero-shot feature information, and perform noise addition processing on the first feature information according to the semantic embedding layer and the zero-shot feature information to obtain the noise-added feature information; In response to the number being multiple and the recognition types of all the multiple audio to be extracted being known types, obtain noise compensation information, and perform noise addition processing on the first feature information successively according to the semantic embedding layer and the noise compensation information to obtain multiple noise-added feature information; In response to the number being multiple and there being audio of an unknown type among the multiple audio to be extracted, obtain zero-shot feature information and noise compensation information, and perform noise addition processing on the first feature information successively according to the semantic embedding layer, the zero-shot feature information, and the noise compensation information to obtain multiple noise-added feature information.
[0020] In one implementation, the processing unit obtains zero-shot feature information in the following manner: Obtain zero-shot audio of the same audio type according to the audio type of the audio to be extracted of the unknown type; Convert the time-domain signal of the zero-shot audio into a frequency-domain signal, and obtain the spectrogram corresponding to the frequency-domain signal; Perform encoding processing on the spectrogram to obtain the zero-shot feature information.
[0021] In one implementation, the noise compensation information is information that is updated successively. The processing unit performs noise addition processing on the first feature information successively according to the semantic embedding layer and the noise compensation information to obtain multiple noise-added feature information in the following manner: When performing the first noise addition processing, perform noise addition processing on the first feature information according to the semantic embedding layer to obtain noise-added feature information; After completing the first noise addition processing, successively obtain updated noise compensation information, and perform successive noise addition processing on the first feature information according to the semantic embedding layer and the successively updated noise compensation information to obtain multiple noise-added feature information; Determine the noise-added feature information obtained when performing the first noise addition processing and the multiple noise-added feature information obtained by successive noise addition processing as the noise-added feature information of the audio to be obtained.
[0022] In one implementation, the noise compensation information is sequentially updated information. The processing unit uses the following method to sequentially perform noise addition processing on the first feature information according to the semantic embedding layer, the zero-shot feature information, and the noise compensation information to obtain multiple pieces of noise-added feature information: When performing the first noise addition processing, perform noise addition processing on the first feature information according to the semantic embedding layer and the zero-shot feature information to obtain noise-added feature information; after completing the first noise addition processing, sequentially obtain updated noise compensation information, and perform sequential noise addition processing on the first feature information according to the semantic embedding layer, the zero-shot feature information, and the sequentially updated noise compensation information to obtain multiple pieces of noise-added feature information; determine the noise-added feature information obtained when performing the first noise addition processing and the multiple pieces of noise-added feature information obtained after completing the sequential noise addition processing as the noise-added feature information of the audio to be obtained.
[0023] In one implementation, the processing unit uses the following method to perform diffusion processing on the first feature information through a diffusion model according to the noise-added feature information and the semantic embedding layer to obtain second feature information: In response to the number of audio to be extracted being single, perform diffusion processing on the first feature information through a diffusion model according to the noise-added feature information and the semantic embedding layer to obtain a single second feature information; in response to the number of audio to be extracted being multiple, perform sequential diffusion processing on the first feature information through a diffusion model according to the semantic embedding layer and the sequentially obtained noise-added feature information to sequentially obtain multiple second feature information.
[0024] In one implementation, the processing unit is further configured to: in response to the noise-added feature information not being information obtained based on zero-shot feature information, perform denoising diffusion processing on the first feature information through a diffusion model according to the noise-added feature information and the semantic embedding layer to eliminate the feature information corresponding to the noise-added feature information in the first feature information and obtain second feature information; in response to the noise-added feature information being information obtained based on zero-shot feature information, perform mapping diffusion processing on the first feature information through a diffusion model according to the noise-added feature information and the semantic embedding layer to retain the feature information corresponding to the noise-added feature information in the first feature information and obtain second feature information.
[0025] In one implementation, the determining unit performs audio restoration processing on the second feature information according to the phase information in the following manner, and determines the audio obtained by the audio restoration processing as the target audio: in response to the number of audio to be extracted being multiple, performing audio restoration processing on the second feature information obtained successively according to the phase information, and determining the multiple audios obtained by the successive audio restoration processing as the target audio; in response to the number of audio to be extracted being single, performing audio restoration processing on the obtained second feature information according to the phase information, and determining the single audio obtained by the audio restoration processing as the target audio.
[0026] In one implementation, the noise compensation information is obtained by the processing unit in the following manner: according to the audio obtained by successive audio restoration processing, successively extracting the feature information of the audio, and successively updating the noise compensation information according to the successively extracted feature information.
[0027] According to the third aspect of the embodiments of the present disclosure, there is provided an audio processing device, including: a processor; and a memory for storing processor-executable instructions; wherein, the processor is configured to: execute the audio processing method described in the first aspect or any one of the implementations of the first aspect.
[0028] According to the fourth aspect of the embodiments of the present disclosure, there is provided a storage medium storing instructions, which when executed by a processor, enable the processor to execute the audio processing method described in the first aspect or any one of the implementations of the first aspect.
[0029] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: obtaining the audio to be processed, and obtaining the phase information and audio feature information of the audio to be processed. According to the text description information of the audio to be extracted and the audio feature information of the audio to be processed, the audio feature information corresponding to the audio to be extracted is obtained through a diffusion model. Performing audio restoration processing on the audio feature information corresponding to the audio to be extracted according to the phase information of the audio to be processed to obtain the target audio. Through the present disclosure, guided by the text information, source separation is performed through a diffusion model, ensuring the stability of the separation effect, so that the obtained audio comes from the target sound source without interference from other sound sources.
[0030] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.
[0032] Figure 1 It is a flowchart of an audio processing method shown according to an exemplary embodiment.
[0033] Figure 2 It is a flowchart of a method for obtaining first feature information shown according to an exemplary embodiment.
[0034] Figure 3 It is a flowchart of a method for obtaining second feature information shown according to an exemplary embodiment.
[0035] Figure 4 It is a flowchart of a method for obtaining noise-added feature information shown according to an exemplary embodiment.
[0036] Figure 5 It is a flowchart of a method for obtaining zero-shot feature information shown according to an exemplary embodiment.
[0037] Figure 6 It is a flowchart of a method for obtaining multiple noise-added feature information shown according to an exemplary embodiment.
[0038] Figure 7 It is a flowchart of a method for obtaining multiple noise-added feature information shown according to another exemplary embodiment.
[0039] Figure 8 It is a flowchart of a method for obtaining second feature information shown according to an exemplary embodiment.
[0040] Figure 9 It is a flowchart of a method for obtaining second feature information shown according to an exemplary embodiment.
[0041] Figure 10 It is a flowchart of a method for obtaining a target audio shown according to an exemplary embodiment.
[0042] Figure 11 It is a flowchart of a method for obtaining noise compensation information shown according to an exemplary embodiment.
[0043] Figure 12 It is a flowchart of an audio processing method shown according to an exemplary embodiment of the present disclosure.
[0044] Figure 13 It is a flowchart of an audio processing method shown according to another exemplary embodiment of the present disclosure.
[0045] Figure 14 It is a flowchart of an audio processing method shown according to another exemplary embodiment of the present disclosure.
[0046] Figure 15It is a flowchart of an audio processing method shown according to another exemplary embodiment of the present disclosure.
[0047] Figure 16 It is a block diagram of an audio processing apparatus shown according to an exemplary embodiment.
[0048] Figure 17 It is a block diagram of an apparatus for an audio processing apparatus shown according to an exemplary embodiment. Detailed implementation manners
[0049] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure.
[0050] The audio processing method provided by the embodiments of the present disclosure is applied to a scenario of obtaining an audio corresponding to a target sound source from multi-source audio.
[0051] The multi-source separation technology refers to separating the signals of different sound sources from a mixed signal, which is an important research direction in the field of signal processing. In the field of speech signal processing, the multi-source separation technology has become an important basis for applications such as speech recognition, speech enhancement, and speech synthesis.
[0052] In the related art, the multi-source separation technology mainly includes the following: The multi-source separation technology based on Independent Component Analysis (ICA), a statistical method, which decomposes the mixed audio signal into multiple independent components through independent component analysis of the mixed audio signal, thereby realizing multi-source separation. The multi-source separation technology based on Blind Source Separation (BSS), a signal processing method, which separates the mixed signal into multiple independent source signals through blind source separation of the mixed signal, thereby realizing multi-source separation. The multi-source separation technology based on deep learning: a neural network-based method, such as multi-channel speech separation based on Convolutional Neural Networks (CNN), speech separation based on Recurrent Neural Network (RNN), etc., which separates the mixed signal into multiple independent source signals through training a deep neural network, thereby realizing multi-source separation. The multi-source separation technology based on array signal processing: an array signal processing-based method, which collects multiple microphone signals in an array and uses array signal processing technology to separate the mixed signal into multiple independent source signals, thereby realizing multi-source separation.
[0053] The multi-source separation technology in the above related technologies has the following deficiencies: the separation effect is not ideal. In a complex acoustic environment, the clarity of the separated speech obtained from the multi-source audio is insufficient, and the intelligibility of the separated audio is insufficient. It is difficult to separate the sound sources of overlapping audio signals in a multi-source scenario. In the case of multiple sound sources, when multiple sound sources emit audio simultaneously, the audio corresponding to the multiple sound sources will overlap, making it difficult to separate and obtain the audio corresponding to the selected sound source. It is difficult to separate the audio of unknown sound source types: when there is audio corresponding to an unknown sound source in the multi-source audio, the multi-source separation model will make mistakes and cannot separate the audio corresponding to the unknown sound source. Moreover, due to the interference of the unknown sound source, the sound source separation results of other known sound sources are inaccurate.
[0054] In view of this, an audio processing method based on a diffusion model and noise reduction compensation proposed in the present disclosure can achieve the functions of multi-source separation or single-source separation. Guided by a large language model through text prompts for sound source separation, it can selectively control the content corresponding to the selected sound source in the multi-source audio separation. Through the combination of zero-shot audio and the guidance of corresponding text prompts, and the diffusion separation mechanism of the diffusion model, the function of separating unknown sound sources can be achieved, ensuring the sound source separation effect in complex scenarios and multi-source scenarios.
[0055] The audio processing method in the present disclosure mainly relates to the fields of sound source separation technology, multi-source separation and blind source separation, and related artificial intelligence technologies. It is a controllable multi-source separation method based on a diffusion model and noise compensation, and can be extended in technical fields such as speech separation and target speaker separation.
[0056] Figure 1 It is a flowchart of an audio processing method shown according to an exemplary embodiment. As Figure 1 shown, the method includes steps S101 to S104.
[0057] In step S101, the audio to be processed is obtained, and the text description information of the audio to be extracted is obtained.
[0058] In step S102, the first feature information is obtained, and the phase information of the audio to be processed is obtained.
[0059] Among them, the first feature information is the audio feature information of the audio to be processed.
[0060] In step S103, according to the text description information and the first feature information, the second feature information is obtained through the diffusion model.
[0061] Among them, the second feature information is the audio feature information of the audio to be extracted.
[0062] In step S104, audio restoration processing is performed on the second feature information according to the phase information, and the audio obtained through the audio restoration processing is determined as the target audio.
[0063] In the embodiments of the present disclosure, the audio to be processed is an audio containing multiple sound sources, and the audio corresponding to multiple sound sources will overlap. In the present disclosure, the text description information of the audio to be extracted is taken as the description information of the sound source type, which is used to guide the extraction of the target audio (the audio to be extracted) in the subsequent audio processing process. The text description information can describe a single sound source or multiple sound sources, and is respectively used for single sound source extraction and multiple sound sources extraction. The text description information can be regarded as an audio extraction instruction in the form of an interactive text. The present disclosure extracts the audio corresponding to the specified sound source based on the text description information, reducing the operation threshold and facilitating user operation.
[0064] In the embodiments of the present disclosure, the diffusion model cannot directly process audio information, so it is necessary to perform feature extraction processing on the audio to be processed, and process the audio features of the audio to be processed through the diffusion model. It can be understood that after the diffusion model processes the audio features of the audio, the output information is also a type of audio feature, rather than the audio to be extracted. Therefore, when the present disclosure extracts the audio features of the audio to be processed, the phase information of the audio to be processed is also obtained. When obtaining the audio features output by the diffusion model, the audio features output by the diffusion model are restored in combination with the phase information, so as to obtain the target audio.
[0065] In the embodiments of the present disclosure, the guidance of sound source separation is performed through text prompts, and the audio content corresponding to the selected sound source in the multi-source audio can be selectively controlled. The function of sound source separation is performed through the diffusion separation mechanism of the diffusion model to ensure the sound source separation effect in complex scenarios and multi-source scenarios.
[0066] It can be understood that the diffusion model is mainly used to process image information. Therefore, when extracting the feature information of the audio to be processed, the feature information of the audio to be processed is represented in the form of an image, which is convenient for the diffusion model to perform diffusion separation processing. The following embodiments of the present disclosure will illustrate the method for obtaining the first feature information.
[0067] Figure 2 is a flowchart of a method for obtaining first feature information shown according to an exemplary embodiment. As Figure 2 shown, the method includes steps S201 to S202.
[0068] In step S201, the time-domain signal of the audio to be processed is converted into a frequency-domain signal, and the spectrogram corresponding to the frequency-domain signal is obtained.
[0069] In step S202, the spectrogram is encoded to obtain the first feature information.
[0070] In the embodiments of the present disclosure, the directly obtained audio to be processed is a time-domain signal. When obtaining the first feature information corresponding to the audio to be processed, the time-domain signal needs to be converted. Through Fourier transform, the time-domain signal is converted into a frequency-domain signal. Then, the frequency-domain signal is represented in the form of a spectrogram, realizing the representation of the time-domain audio signal in the form of an image, which is convenient for denoising processing through a diffusion model (such as Stable Diffusion, SD). Further, to ensure local modeling of the spectrogram, the spectrogram needs to be convolved. A two-layer convolutional layer (Conv2D) with all parameters of the spectrogram initialized to 1 is used as the audio feature corresponding to the audio to be processed and for transition. It can be understood that in the present disclosure, the feature dimension of the spectrogram output after convolution is the same as the dimension of the original spectrogram, that is, the number of feature vectors of the original spectrogram is the same as that of the spectrogram after convolution.
[0071] In the embodiments of the present disclosure, through the conversion of the original audio to be processed from a time-domain signal to a frequency-domain signal, the conversion of the frequency-domain signal to a spectrogram, and the convolution processing of the spectrogram, the feature information of the audio to be processed is represented in the form of an image, which is convenient for the diffusion model to perform diffusion processing on it, and finally the target audio is obtained.
[0072] In the embodiments of the present disclosure, the text description information guides the diffusion model to extract the feature information of the target audio during the process of the diffusion model processing the first feature information. Among them, the text description information cannot be directly input into the diffusion model and needs to be further transformed. And the processing flow of the diffusion model for the first feature information is a "denoising" process, so the initialization noise corresponding to the first feature information needs to be obtained to facilitate the diffusion model to execute the denoising process. The following embodiments of the present disclosure illustrate the method for obtaining the second feature information.
[0073] Figure 3 is a flowchart of a method for obtaining second feature information shown according to an exemplary embodiment. As Figure 3 shown, the method includes steps S301 to S304.
[0074] In step S301, the text description information is processed by a large language model to obtain a semantic embedding layer.
[0075] In step S302, according to the text description information, the number and recognition type of the audio to be extracted are determined.
[0076] In step S303, according to the number and recognition type of the audio to be extracted and the semantic embedding layer, the first feature information is subjected to noise addition processing to obtain noise-added feature information.
[0077] In step S304, according to the noise-added feature information and the semantic embedding layer, the diffusion model is used to perform diffusion processing on the first feature information to obtain second feature information.
[0078] In the embodiments of the present disclosure, the large language model (LLM) is used to encode the text description information to obtain the semantic embedding layer corresponding to the text description information, that is, the large language model LLM is used to perform convolutional processing on the text description information, and the output of the last layer of the model architecture (Transformer) for sequence conversion is used as the semantic embedding layer, guiding the diffusion model to extract the feature information (second audio information) of the target audio during the process of the diffusion model processing the first feature information. It can be understood that in the case of not inputting the semantic embedding layer from the outside, the diffusion model will use the text-encoder module to encode the input text (that is, the text description information of the audio to be processed), and autonomously generate the semantic embedding layer to guide the denoising process of the diffusion model for the first feature information. And using the text encoding module built into the diffusion model to obtain the semantic embedding layer and sending the semantic embedding layer to the CrossAttention part of the Unet in the diffusion model as a condition to guide the denoising operation of the Unet will increase the computational cost compared to using the large language model. Therefore, the present disclosure uses the large language model to encode the text description information of the audio to be processed to obtain the semantic embedding layer, no longer uses the text encoder, and replaces the role of the text encoder. In summary, the present disclosure uses the large language model to obtain the semantic embedding layer, guiding the diffusion model to extract the feature information (second audio information) of the target audio during the process of the diffusion model processing the first feature information, achieving the effect of reducing computational expenses.
[0079] In the embodiments of the present disclosure, according to the sound source information of the audio to be extracted included in the text description information, the number of the audio to be extracted and the type of each audio to be extracted can be determined, and then different noise-adding strategies can be adopted according to the number of the audio to be extracted and the type of each audio to be extracted, so as to flexibly respond to various audio extraction instructions issued by the user to meet the user's needs.
[0080] In the embodiments of the present disclosure, the semantic embedding layer obtained by the large language model and the first feature information (the feature information corresponding to the audio to be processed) are sent to the "noise generation" module together to generate initial noise, which is used as the feature of the target sound source to be separated for initialization and plays an auxiliary role in the denoising process of the diffusion model. Guide the diffusion model to obtain the feature information corresponding to the target audio.
[0081] In the embodiments of the present disclosure, for text data, a large language model is used as a feature mapping generation method to perform feature mapping on text description information (such as "human voice, dog barking, movie playing sound, singing voice", etc.) to obtain a semantic embedding layer. The large language model in the present disclosure is optional, and open source models such as ChatGPT, ChatGLM, and Moss can be used as the large language model for extracting the semantic embedding layer.
[0082] In the embodiments of the present disclosure, both the number of audio to be extracted and the type of audio to be extracted included in the text description information are influencing conditions for obtaining the noise-added feature information. The recognized types of audio to be extracted in the present disclosure include known types and unknown types. The audio of the known type is the same as the audio type in the diffusion model training data, and the audio of the unknown type is different from the audio type in the diffusion model training data. The number of audio to be extracted includes single or multiple. According to whether the number of audio to be extracted is one or more, and whether the type of audio to be extracted is a known type or an unknown type, different noise-added feature information extraction strategies are selected respectively. The following embodiments of the present disclosure illustrate the method for obtaining the noise-added feature information.
[0083] Figure 4 It is a flowchart of a method for obtaining noise-added feature information shown according to an exemplary embodiment. As Figure 4 shown, the method includes step S401, step S402A, step S402B, step S402C, and step S402D.
[0084] In step S401, according to the text description information, determine the number and recognized type of the audio to be extracted.
[0085] In step S402A, in response to the number being single and the recognized type of the single audio to be extracted being a known type, perform noise addition processing on the first feature information according to the semantic embedding layer to obtain the noise-added feature information.
[0086] In the embodiments of the present disclosure, when the number of audio to be extracted is single and of a known type, only a single diffusion process can be performed based on the guidance of the noise-added feature information through the diffusion model to obtain the second feature information (the feature information corresponding to the target audio). Since there is no secondary diffusion process, there is no need to update the noise-added feature information, and the problem of repeatedly obtaining the same second feature through multiple diffusion processes will not occur. Therefore, the noise-added feature information can be obtained by performing single noise addition processing on the first feature information in combination with the semantic embedding layer.
[0087] In step S402B, in response to the number being single and the recognized type of the single audio to be extracted being an unknown type, obtain zero-shot feature information, and perform noise addition processing on the first feature information according to the semantic embedding layer and the zero-shot feature information to obtain the noise-added feature information.
[0088] In the embodiments of the present disclosure, when the number of audio to be extracted is single and the type is unknown, only a single noise addition process is required to obtain noise-added feature information. However, since there is no training audio of the same audio type as the unknown audio in the training data of the diffusion model and the unknown type of audio cannot be recognized, it is necessary to introduce zero-shot audio of an audio type different from the training data of the diffusion model to obtain noise-added feature information, so as to guide the diffusion model through the diffusion process to obtain the feature information of the unknown type of audio to be extracted (target audio) based on the noise-added feature information.
[0089] In step S402C, in response to the number being multiple and the recognized types of the multiple audio to be extracted being known types, noise compensation information is obtained, and the first feature information is successively subjected to noise addition processing according to the semantic embedding layer and the noise compensation information to obtain multiple pieces of noise-added feature information.
[0090] In the embodiments of the present disclosure, when the audio types of the audio to be extracted are all known types, there is no need to introduce zero-shot feature information. When the number of audio to be extracted is multiple, the first feature information will be processed through the diffusion model in a loop to obtain the feature information corresponding to each of the multiple audio to be extracted one by one.
[0091] In step S402D, in response to the number being multiple and there being audio with an unknown recognized type among the multiple audio to be extracted, zero-shot feature information and noise compensation information are obtained, and the first feature information is successively subjected to noise addition processing according to the semantic embedding layer, the zero-shot feature information, and the noise compensation information to obtain multiple pieces of noise-added feature information.
[0092] In the embodiments of the present disclosure, when the number of audio to be extracted is multiple and there is unknown type audio, the first feature information needs to be processed through the diffusion model in a loop, and the feature information corresponding to each of the multiple audio to be extracted is input one by one, and zero-shot feature information is introduced for the unknown type of audio.
[0093] In the embodiments of the present disclosure, different noise-added feature information extraction strategies are respectively selected according to different combinations of the number of audio to be extracted and the type of audio to be extracted. It is possible to perform sound source extraction on diverse original audio (audio to be processed) and meet the corresponding sound source extraction requirements of the diverse needs of users. The application range and flexibility of the audio processing method of the present disclosure are guaranteed.
[0094] It can be understood that the noise-added feature information obtained based on the zero-shot feature information is used to guide the diffusion model to obtain the feature information of the unknown type of audio. There is a corresponding relationship between the zero-shot feature information and the feature information of the unknown type of audio, and correspondingly, there is also a corresponding relationship between the zero-shot audio used to obtain the zero-shot feature information and the unknown type of audio. In the embodiments of the present disclosure, the method for obtaining the zero-shot feature information is described.
[0095] Figure 5 is a flowchart of a method for obtaining zero-shot feature information shown according to an exemplary embodiment. As Figure 5 shown, the method includes steps S501 to S503.
[0096] In step S501, according to the audio type of the audio to be extracted with an unknown type, a zero-shot audio of the same audio type is obtained.
[0097] In step S502, the time-domain signal of the zero-shot audio is converted into a frequency-domain signal, and a spectrogram corresponding to the frequency-domain signal is obtained.
[0098] In step S503, the spectrogram is encoded to obtain zero-shot feature information.
[0099] In the embodiments of the present disclosure, the zero-shot feature information is obtained in the same way as the first feature information (i.e., the feature information of the audio to be processed), that is, the zero-shot audio in the form of a time-domain signal is subjected to a conversion process, and through Fourier transform, the time-domain signal is converted into a frequency-domain signal, and then the frequency-domain signal is represented in the form of a spectrogram, and the spectrogram is subjected to a convolution process to obtain zero-shot feature information.
[0100] In the embodiments of the present disclosure, the zero-shot audio for obtaining zero-shot feature information corresponds to the audio to be extracted with an unknown type, and has the same recognition type and audio type as the audio to be extracted with an unknown type. The sound source signal that has never appeared in the diffusion model training is the zero-shot audio, and has an audio type different from the diffusion model training data. It can be understood that various types of sound sources are used in the training process of the diffusion model, and the sound source data of the same type also meets a certain order of magnitude. The diffusion model is trained with the above various types of sound source data, so that the diffusion model has the ability to separate the sound sources of the types. That is, in the inference process of the diffusion model, the diffusion model has the ability to separate the sound sources of the types that have been seen (i.e., the multiple types corresponding to the training data). When it comes to the sound sources that have not been seen (i.e., other types of sound sources outside the multiple types corresponding to the training data), it cannot be separated. In the present disclosure, any audio emitted by an unknown type of sound source outside the diffusion model training set is called zero-shot audio. By obtaining the zero-shot audio corresponding to the unknown audio to be extracted, zero-shot feature information is obtained, and the noise-added feature information is obtained according to the zero-shot feature information, so as to guide the diffusion model to obtain the feature information of the unknown audio in the diffusion process of the diffusion model, and realize the sound source separation for the unknown audio.
[0101] In an exemplary embodiment of the present disclosure, when there is unknown audio in the audio to be extracted, a 5-second zero-shot audio of the same audio type as the unknown audio can be obtained, the corresponding zero-shot feature information can be extracted, and it can participate in the subsequent processing flow to extract the unknown audio.
[0102] The zero-shot audio in the embodiments of the present disclosure is optional and can be the actually collected audio or the "virtual noise" generated by the large language model, that is, the audio generated according to the sound source characteristics (the sound source characteristics of the unknown audio in the audio to be processed) simulated by the text information. Its purpose is to generate the characteristic signal of the sound source to be separated (the sound source emitting the unknown audio) according to the reference text (text description information), and to realize the denoising operation according to the reference text to obtain the unknown audio specified by the reference text.
[0103] In the embodiments of the present disclosure, when there are multiple audios to be extracted, noise compensation needs to be performed each time a single audio is extracted, the noise-added feature information is updated, and the next audio is obtained according to the updated noise-added feature information. The following embodiments of the present disclosure further illustrate the method for obtaining noise information.
[0104] Figure 6 It is a flowchart of a method for obtaining multiple pieces of noise-added feature information shown according to an exemplary embodiment. As Figure 6 shown, the method includes steps S601 to S603.
[0105] In step S601, when performing the first noise addition process, the first feature information is subjected to noise addition processing according to the semantic embedding layer to obtain the noise-added feature information.
[0106] In step S602, after the first noise addition process is completed, the updated noise compensation information is obtained successively, and the first feature information is subjected to successive noise addition processing according to the semantic embedding layer and the successively updated noise compensation information to obtain multiple pieces of noise-added feature information.
[0107] In step S603, the noise-added feature information obtained when performing the first noise addition process and the multiple pieces of noise-added feature information obtained by successive noise addition processing are determined as the noise-added feature information of the audio to be obtained.
[0108] In the embodiments of the present disclosure, when the number of audio to be extracted is multiple, the diffusion model is used to repeatedly process the first feature information, and the feature information corresponding to each of the multiple audio to be extracted is input one by one. When extracting one of the multiple audio to be extracted, there is no need to obtain the feature information corresponding to the already extracted audio again. In this case, noise compensation is performed based on the already extracted audio, and the noise-added feature information is updated. Then, based on the updated noise-added feature information, the feature information of the next audio to be extracted is extracted. After the next audio to be extracted is successfully extracted, new noise compensation information is obtained again, and then the process of updating the noise-added feature information to extracting a single audio to be extracted is repeated until all the audio to be extracted are obtained.
[0109] In the embodiments of the present disclosure, by continuously obtaining feature compensation information and updating the noise-added feature information, the extraction of multiple audio to be extracted one by one is realized, and the repeated acquisition of the feature information of a single audio is avoided.
[0110] It can be understood that when the number of audio to be extracted is multiple and there are unknown type audio among the multiple audio to be extracted, the above processes of introducing zero-shot feature information and updating the noise-added feature information need to be executed simultaneously. The following embodiments of the present disclosure further illustrate the method for obtaining the noise-added feature information.
[0111] Figure 7 is a flowchart of a method for obtaining multiple pieces of noise-added feature information shown according to another exemplary embodiment. As Figure 7 shown, the method includes steps S701 to S703.
[0112] In step S701, when performing the first noise addition process, the first feature information is subjected to noise addition processing according to the semantic embedding layer and the zero-shot feature information to obtain the noise-added feature information, or the first feature information is subjected to noise addition processing according to the semantic embedding layer to obtain the noise-added feature information.
[0113] In step S702, after the first noise addition process is completed, updated noise compensation information is obtained successively. According to the semantic embedding layer, the zero-shot feature information, and the successively updated noise compensation information, the first feature information is subjected to successive noise addition processing to obtain multiple pieces of noise-added feature information, or after the first noise addition process is completed, updated noise compensation information is obtained successively. According to the semantic embedding layer and the successively updated noise compensation information, the first feature information is subjected to successive noise addition processing to obtain multiple pieces of noise-added feature information.
[0114] In step S703, the noise-added feature information obtained when performing the first noise addition process and the multiple pieces of noise-added feature information obtained after the successive noise addition process are determined as the noise-added feature information of the audio to be obtained.
[0115] In the embodiments of the present disclosure, when the number of audio to be extracted is multiple and there are unknown type audio among the multiple audio to be extracted. For the multiple unknown type audio existing in the multiple audio to be extracted, zero-shot audio of the same audio type is respectively obtained, and the corresponding multiple zero-shot feature information is obtained. The noise-added feature information is obtained according to the obtained multiple zero-shot feature information to guide the subsequent feature information extraction. And when each single audio to be extracted is extracted, the noise compensation information is obtained, the noise-added feature information is updated, and the next audio to be extracted is obtained until the extraction of the audio to be extracted is completed.
[0116] In the embodiments of the present disclosure, by obtaining zero-shot audio corresponding to the unknown audio to be extracted, obtaining zero-shot feature information, and obtaining noise-added feature information according to the zero-shot feature information, the diffusion model is guided to obtain the feature information of the unknown audio during the diffusion process of the diffusion model, so as to realize the sound source separation of the unknown audio. And by continuously obtaining feature compensation information and updating the noise-added feature information, the extraction of multiple audio to be extracted one by one is realized, avoiding repeated acquisition of the feature information of a single audio.
[0117] In the embodiments of the present disclosure, the number of audio to be extracted not only affects the generation process of the noise-added feature information, but also further affects the process of obtaining the second feature information (the audio feature information of the audio to be extracted) according to the noise-added feature information. The following embodiments of the present disclosure illustrate the method for obtaining the second feature information.
[0118] Figure 8 is a flowchart of a method for obtaining second feature information shown according to an exemplary embodiment. As Figure 8 shown, the method includes step S801, step S802A and step S802B.
[0119] In step S801, a semantic embedding layer is obtained, and the number of audio to be extracted is determined.
[0120] In step S802A, in response to the number of audio to be extracted being single, according to the noise-added feature information and the semantic embedding layer, the first feature information is diffusely processed by the diffusion model to obtain a single second feature information.
[0121] In step S802B, in response to the number of audio to be extracted being multiple, according to the semantic embedding layer and the sequentially obtained noise-added feature information, the first feature information is sequentially diffusely processed by the diffusion model to sequentially obtain multiple second feature information.
[0122] In the embodiments of the present disclosure, when the number of audio to be extracted is one, the audio feature information (second feature information) of the audio to be extracted can be obtained through a single diffusion process by a diffusion model. That is, the Unet in the diffusion model is used to perform step-by-step noise reduction processing on the feature information (first feature information) of the audio to be extracted. Combining the noise-added feature information and the guidance of the semantic embedding layer, the feature information of the audio to be extracted in the audio feature information to be extracted is retained, and the feature information of other audio is eliminated. Finally, the diffusion model outputs the processing result to complete the extraction of the second feature information.
[0123] In the embodiments of the present disclosure, when the number of audio to be extracted is multiple, the diffusion model is used to perform successive diffusion processing to obtain multiple audio to be extracted. That is, according to the different noise-added feature information obtained each time, combined with the semantic embedding layer and the first feature information, the audio feature information of each audio to be extracted is extracted one by one to complete the acquisition of the second feature information.
[0124] In the embodiments of the present disclosure, different second feature information extraction strategies are adopted for single or multiple audio to be extracted, meeting the diverse needs of users.
[0125] Figure 9 It is a flowchart of a method for obtaining second feature information shown according to an exemplary embodiment. As Figure 9 shown, the method includes step S901, step S902A, and step S902B.
[0126] In step S901, noise-added feature information is obtained, and a semantic embedding layer is obtained.
[0127] In step S902A, in response to the noise-added feature information being information obtained based on zero-shot feature information, according to the noise-added feature information and the semantic embedding layer, the diffusion model performs diffusion processing on the first feature information by mapping, retaining the feature information corresponding to the noise-added feature information in the first feature information to obtain the second feature information.
[0128] In step S902B, in response to the noise-added feature information not being information obtained based on zero-shot feature information, according to the noise-added feature information and the semantic embedding layer, the diffusion model performs denoising diffusion processing on the first feature information to eliminate the feature information corresponding to the noise-added feature information in the first feature information to obtain the second feature information.
[0129] In the embodiments of the present disclosure, when the noise-added feature information is not information obtained based on zero-shot feature information, the diffusion model combines the noise-added feature information input into the diffusion model and the semantic embedding layer to perform noise reduction processing on the first feature information, and the second feature information is obtained by predicting the noise and subtracting the noise from the first feature information.
[0130] In an embodiment of the present disclosure, when the noise-added feature information is information obtained based on zero-shot feature information, a method opposite to the above "denoising" method is used to obtain the second feature information. That is, after obtaining the noise-added feature information based on the zero-shot feature information and sending the noise-added feature information into the diffusion model, the zero-shot feature information is used as a noise reference, and the second feature information to be separated is generated and output from the first feature information in a mapped manner through mapping.
[0131] In an embodiment of the present disclosure, for the audio to be extracted of known types and the audio to be extracted of unknown types, different feature extraction methods are respectively used through the diffusion model for extraction, so as to realize the extraction of the audio corresponding to the unknown sound source in the multi-source audio.
[0132] In an embodiment of the present disclosure, different target audio generation strategies are respectively adopted for the number of audio to be extracted (single or multiple) to meet the diverse needs of users.
[0133] Figure 10 It is a flowchart of a method for obtaining a target audio shown according to an exemplary embodiment. As Figure 10 shown, the method includes step S1001, step S1002A, and step S1002B.
[0134] In step S1001, the number of audio to be extracted is determined, and the second feature information is obtained.
[0135] In step S1002A, in response to the number of audio to be extracted being multiple, audio restoration processing is successively performed on the successively obtained second feature information according to the phase information, and the multiple audios obtained by the successive audio restoration processing are determined as the target audio.
[0136] In step S1002B, in response to the number of audio to be extracted being single, audio restoration processing is performed on the obtained second feature information according to the phase information, and the single audio obtained by the audio restoration processing is determined as the target audio.
[0137] In the embodiments of the present disclosure, when the number of audio to be extracted is multiple, multiple audio restoration processes are performed according to the obtained multiple second feature information to obtain multiple audio to be extracted (target audio). When the number of audio to be extracted is single, a single audio restoration process is performed according to the obtained single second feature information to obtain a single audio to be extracted (target audio). Among them, the feature restoration process of the second feature information is opposite to the acquisition process of the first feature information, that is, the second feature information is deconvolved (decoded) according to the phase information of the original audio (audio to be processed) to obtain the spectrogram of the audio to be extracted, and then the audio to be extracted in the form of a frequency-domain signal is obtained. Furthermore, through the conversion process of the frequency-domain signal - time-domain signal, the audio to be extracted in the form of a time-domain signal is obtained. After all the audio to be extracted are obtained, the extraction of the target audio is completed.
[0138] In an exemplary embodiment of the present disclosure, the first feature information, the semantic embedding layer, and the noise-added feature information are used as input features and fed into a diffusion model for denoising operations. By controlling the noise generation parameters, the reference audio, and the inherent parameters of the diffusion model, source separation is performed from the mixed signal (first feature information) in the input signal. For the image generated by the diffusion model, it is restored through a "spectrum restoration" module to obtain a "separated source spectrum" (audio to be extracted) with the same dimension as the input, and an inverse Fourier transform is performed in combination with the phase characteristics of the original signal (audio to be processed) to restore and output the time-domain signal, obtaining the separation result. Among them, the "spectrum restoration" module is composed of an inverse convolutional kernel (DeConv2D) initialized with 2 fully parameterized layers symmetric to the "feature extraction" module.
[0139] In the embodiments of the present disclosure, after obtaining a single audio to be extracted, noise compensation needs to be performed based on the extracted audio to obtain multiple other low-quality extracted audio. The following embodiments of the present disclosure illustrate the method for obtaining noise compensation information.
[0140] Figure 11 It is a flowchart of a method for obtaining noise compensation information shown in an exemplary embodiment. As Figure 11 shown, the method includes steps S1101 to step S1102.
[0141] In step S1101, according to the audio obtained by successive audio restoration processing, the feature information of the audio is successively extracted.
[0142] In step S1102, according to the successively extracted feature information, the noise compensation information is successively updated.
[0143] In an embodiment of the present disclosure, when the number of audio to be extracted is multiple, when obtaining one of the multiple audio to be extracted, the audio feature information of the obtained audio is extracted, and the corresponding noise compensation information is obtained according to the audio feature information. When obtaining a new audio to be extracted next time, the corresponding compensation information is re-obtained to update the noise compensation information.
[0144] In an embodiment of the present disclosure, in the scenario of multi-source audio (including full-source separation), that is, it is necessary to separate multiple single-source information contained in the current mixed audio. When the present technical solution proposes an audio processing method for source separation, for the separated audio (one of the multiple audio to be extracted) first generated by the diffusion model, the acoustic feature information of the separated audio is obtained by the same "feature extraction" method as the feature information (first feature information) of the current mixed audio (audio to be processed), and is passed to the "noise compensation" module for processing to obtain the noise compensation information. "Noise compensation" uses the principle of spectral subtraction. By performing spectral subtraction based on the original data features (the previously obtained noise-added feature information), the features after spectral subtraction are compressed by the encoder part of the variational auto-encoder (VAE) of the diffusion model to complete the update of the noise-added feature information. The compressed features (the updated noise-added feature information) are concatenated in the channel dimension (Concat) into new features and sent into the diffusion model. Through the above method, the second separated data (another one of the multiple audio to be extracted) can be generated. In the subsequent successive separation process, the above process is carried out in an accumulative manner until it does not match the data to be separated (audio to be extracted) described by the semantic embedding layer, that is, the process of generating the separated data (multiple audio to be extracted) ends.
[0145] In an embodiment of the present disclosure, the corresponding feature compensation information is continuously obtained through the obtained audio to be extracted, and the noise-added feature information is updated to realize the extraction of multiple audio to be extracted one by one, avoiding repeated acquisition of the audio feature information for a single audio.
[0146] In an exemplary embodiment of the present disclosure, extract a single known type of audio from the multi-source audio, such as Figure 12As shown in the flowchart of the audio processing method, the method specifically includes the following steps: perform feature extraction and phase information acquisition on the input audio to obtain the audio feature information and phase information of the input audio; while obtaining the audio feature information, process the input text through a large language model to obtain a semantic embedding layer; combine the semantic embedding layer to perform noise addition processing (i.e., noise generation) on the audio feature information to obtain noise-added feature information; after obtaining the noise-added feature information and the semantic embedding layer, input the noise-added feature information, the semantic embedding layer, and the audio feature information into a diffusion model to obtain the target feature information corresponding to the separated audio; perform spectrum restoration processing on the target feature information based on the phase information of the input audio to obtain the separated audio, and complete the extraction of a single known audio.
[0147] In an exemplary embodiment of the present disclosure, extract a single unknown type of audio from multi-source audio, such as Figure 13 As shown in the flowchart of the audio processing method, the method specifically includes the following steps: perform feature extraction and phase information acquisition on the input audio to obtain the audio feature information and phase information of the input audio; while obtaining the audio feature information, perform feature extraction on zero-shot audio to obtain zero-shot feature information, and process the input text through a large language model to obtain a semantic embedding layer; combine the semantic embedding layer and the zero-shot feature information to perform noise addition processing (i.e., noise generation) on the audio feature information to obtain noise-added feature information; after obtaining the noise-added feature information and the semantic embedding layer, input the noise-added feature information, the semantic embedding layer, and the audio feature information into a diffusion model to obtain the target feature information corresponding to the separated audio; perform spectrum restoration processing on the target feature information based on the phase information of the input audio to obtain the separated audio, and complete the extraction of a single known audio.
[0148] In an exemplary embodiment of the present disclosure, extract multiple known types of audio from multi-source audio, such as Figure 14As shown in the flowchart of the audio processing method, the method specifically includes the following steps: When extracting the first separated audio, perform feature extraction and phase information acquisition on the input audio to obtain the audio feature information and phase information of the input audio; while obtaining the audio feature information, process the input text through a large language model to obtain a semantic embedding layer; combine the semantic embedding layer to perform noise addition processing (i.e., noise generation) on the audio feature information to obtain noise-added feature information; after obtaining the noise-added feature information and the semantic embedding layer, input the noise-added feature information, the semantic embedding layer, and the audio feature information into a diffusion model to obtain the target feature information corresponding to the separated audio; perform spectral restoration processing on the target feature information based on the phase information of the input audio to obtain the separated audio, perform feature extraction on the obtained separated audio to extract the audio feature information of the separated audio, and perform noise compensation according to the audio feature information of the separated audio, and update the noise-added feature information by combining the audio feature information of the separated audio, the audio feature information of the input audio, and the semantic embedding layer; obtain the next separated audio based on the updated noise-added feature information, and repeat the process from noise compensation to spectral restoration until all the separated audio indicated by the text is obtained, and complete the extraction of multiple known audio.
[0149] In an exemplary embodiment of the present disclosure, multiple audio are extracted from a multi-source audio, and there is an audio of an unknown type among the multiple audio to be extracted. As Figure 15 As shown in the flowchart of the audio processing method, the method specifically includes the following steps: When extracting the first separated audio, perform feature extraction and phase information acquisition on the input audio to obtain the audio feature information and phase information of the input audio; while obtaining the audio feature information, perform feature extraction on the zero-shot audio to obtain zero-shot feature information (in response to the audio to be separated being an unknown audio), and process the input text through a large language model to obtain a semantic embedding layer; combine the semantic embedding layer and the zero-shot feature information (in response to the audio to be separated being an unknown audio) to perform noise addition processing (i.e., noise generation) on the audio feature information to obtain noise-added feature information; after obtaining the noise-added feature information and the semantic embedding layer, input the noise-added feature information, the semantic embedding layer, and the audio feature information into a diffusion model to obtain the target feature information corresponding to the separated audio; perform spectral restoration processing on the target feature information based on the phase information of the input audio to obtain the separated audio, perform feature extraction on the obtained separated audio to extract the audio feature information of the separated audio, and perform noise compensation according to the audio feature information of the separated audio, and update the noise-added feature information by combining the audio feature information of the separated audio, the zero-shot feature information (in response to the audio to be separated being an unknown audio), the audio feature information of the input audio, and the semantic embedding layer; obtain the next separated audio based on the updated noise-added feature information, and repeat the process from noise compensation to spectral restoration until all the separated audio indicated by the text is obtained, and complete the extraction of multiple known audio.
[0150] In the embodiments of the present disclosure, obtaining a semantic embedding layer based on interactive text and a large language model and guiding the audio separation process can achieve the separation of a predetermined target sound source according to different requirements for separation effects (separating audio corresponding to a single sound source or separating audio corresponding to multiple sound sources) and different separation types (separating audio of a known type or unknown type). In the process of gradually separating multiple sound sources in the present disclosure, the denoising effect of the diffusion model is improved by using the separated audio data as noise compensation, and the quality of the obtained separated audio is improved. In the case of multiple speakers or overlapping multiple sound sources, according to the sound source selection mechanism of the interactive text, the diffusion model is combined to perform single sound source separation processing or multiple sound source separation, ensuring the audio separation effect and stability and ensuring the quality of the generated audio. The present disclosure obtains a semantic embedding layer based on a large language model without using the text encoder in the diffusion model, which improves the sound source separation effect, reduces the algorithm complexity, and improves the scalability and customizability of the audio processing method in the present disclosure, that is, the algorithm can be shared with the large language models built in various terminal devices (such as mobile phones, tablets, etc.), improving the adaptation efficiency of software and hardware and the technical utilization rate.
[0151] Based on the same concept, an audio processing apparatus 100 is further provided in the embodiments of the present disclosure.
[0152] It can be understood that, in order to implement the above functions, the audio processing apparatus 100 provided in the embodiments of the present disclosure includes the corresponding hardware structures and / or software modules for executing each function. Combining the units and algorithm steps of the examples disclosed in the embodiments of the present disclosure, the embodiments of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiments of the present disclosure.
[0153] Figure 16 is a block diagram of an audio processing apparatus 100 shown according to an exemplary embodiment. Referring to Figure 16 , the apparatus includes an acquisition unit 101, a processing unit 102, and a determination unit 103.
[0154] The acquisition unit 101 is configured to acquire the audio to be processed, acquire the text description information of the audio to be extracted, acquire the first feature information, and acquire the phase information of the audio to be processed.
[0155] Wherein, the first feature information is the audio feature information of the audio to be processed
[0156] The processing unit 102 is configured to obtain the second feature information through a diffusion model according to the text description information and the first feature information.
[0157] Among them, the second feature information is the audio feature information of the audio to be extracted.
[0158] The determination unit 103 is configured to perform audio restoration processing on the second feature information according to the phase information, and determine the audio obtained by the audio restoration processing as the target audio.
[0159] In one implementation, the acquisition unit 101 acquires the first feature information in the following manner: converts the time-domain signal of the audio to be processed into a frequency-domain signal, and acquires the spectrogram corresponding to the frequency-domain signal. Performs encoding processing on the spectrogram to obtain the first feature information.
[0160] In one implementation, the processing unit 102 obtains the second feature information from the text description information and the first feature information through a diffusion model in the following manner: processes the text description information through a large language model to obtain a semantic embedding layer. Determines the number and recognition type of the audio to be extracted according to the text description information. Performs noise addition processing on the first feature information according to the number and recognition type of the audio to be extracted and the semantic embedding layer to obtain noise-added feature information. Performs diffusion processing on the first feature information through a diffusion model according to the noise-added feature information and the semantic embedding layer to obtain the second feature information.
[0161] In one implementation, the recognition types of the audio to be extracted include known types and unknown types. The audio of the known type has the same audio type as the audio in the diffusion model training data, and the audio of the unknown type has a different audio type from the audio in the diffusion model training data. The number of the audio to be extracted includes single or multiple. The processing unit 102 performs noise addition processing on the first feature information according to the number and recognition type of the audio to be extracted and the semantic embedding layer to obtain noise-added feature information in the following manner: in response to the number being single and the recognition type of the single audio to be extracted being a known type, performs noise addition processing on the first feature information according to the semantic embedding layer to obtain noise-added feature information. In response to the number being single and the recognition type of the single audio to be extracted being an unknown type, obtains zero-shot feature information, and performs noise addition processing on the first feature information according to the semantic embedding layer and the zero-shot feature information to obtain noise-added feature information. In response to the number being multiple and the recognition types of the multiple audio to be extracted being all known types, obtains noise compensation information, and performs noise addition processing on the first feature information successively according to the semantic embedding layer and the noise compensation information to obtain multiple noise-added feature information. In response to the number being multiple and there being an audio with an unknown type among the multiple audio to be extracted, obtains zero-shot feature information and noise compensation information, and performs noise addition processing on the first feature information successively according to the semantic embedding layer, the zero-shot feature information and the noise compensation information to obtain multiple noise-added feature information.
[0162] In one implementation, the processing unit 102 obtains zero-shot feature information in the following manner: According to the audio type of the audio to be extracted with an unknown type, zero-shot audio of the same audio type is obtained. The time-domain signal of the zero-shot audio is converted into a frequency-domain signal, and a spectrogram corresponding to the frequency-domain signal is obtained. The spectrogram is encoded to obtain zero-shot feature information.
[0163] In one implementation, the noise compensation information is information that is updated successively. The processing unit 102 successively adds noise to the first feature information according to the semantic embedding layer and the noise compensation information to obtain multiple pieces of noise-added feature information in the following manner: When performing the first noise addition, the first feature information is added with noise according to the semantic embedding layer to obtain noise-added feature information. After the first noise addition is completed, updated noise compensation information is successively obtained, and the first feature information is successively added with noise according to the semantic embedding layer and the successively updated noise compensation information to obtain multiple pieces of noise-added feature information. The noise-added feature information obtained when performing the first noise addition and the multiple pieces of noise-added feature information obtained by successive noise addition are determined as the noise-added feature information of the audio to be obtained.
[0164] In one implementation, the noise compensation information is information that is updated successively. The processing unit 102 successively adds noise to the first feature information according to the semantic embedding layer, the zero-shot feature information, and the noise compensation information to obtain multiple pieces of noise-added feature information in the following manner: When performing the first noise addition, the first feature information is added with noise according to the semantic embedding layer and the zero-shot feature information to obtain noise-added feature information, or the first feature information is added with noise according to the semantic embedding layer to obtain noise-added feature information; after the first noise addition is completed, updated noise compensation information is successively obtained, and the first feature information is successively added with noise according to the semantic embedding layer, the zero-shot feature information, and the successively updated noise compensation information to obtain multiple pieces of noise-added feature information, or after the first noise addition is completed, updated noise compensation information is successively obtained, and the first feature information is successively added with noise according to the semantic embedding layer and the successively updated noise compensation information to obtain multiple pieces of noise-added feature information; the noise-added feature information obtained when performing the first noise addition and the multiple pieces of noise-added feature information obtained after successive noise addition are determined as the noise-added feature information of the audio to be obtained.
[0165] In one implementation, the processing unit 102 performs diffusion processing on the first feature information through a diffusion model according to the noise-added feature information and the semantic embedding layer in the following manner to obtain the second feature information: In response to the number of audio to be extracted being single, diffusion processing is performed on the first feature information through the diffusion model according to the noise-added feature information and the semantic embedding layer to obtain a single second feature information. In response to the number of audio to be extracted being multiple, successive diffusion processing is performed on the first feature information through the diffusion model according to the semantic embedding layer and the successively obtained noise-added feature information to successively obtain multiple second feature information.
[0166] In one implementation, the processing unit 102 is further configured to: In response to the noise-added feature information not being information obtained based on zero-shot feature information, perform denoising diffusion processing on the first feature information through the diffusion model according to the noise-added feature information and the semantic embedding layer to eliminate the feature information corresponding to the noise-added feature information in the first feature information and obtain the second feature information. In response to the noise-added feature information being information obtained based on zero-shot feature information, perform mapping diffusion processing on the first feature information through the diffusion model according to the noise-added feature information and the semantic embedding layer to retain the feature information corresponding to the noise-added feature information in the first feature information and obtain the second feature information.
[0167] In one implementation, the determination unit 103 performs audio restoration processing on the second feature information according to the phase information in the following manner and determines the audio obtained by the audio restoration processing as the target audio: In response to the number of audio to be extracted being multiple, successive audio restoration processing is performed on the successively obtained second feature information according to the phase information, and the multiple audio obtained by the successive audio restoration processing are determined as the target audio. In response to the number of audio to be extracted being single, audio restoration processing is performed on the obtained second feature information according to the phase information, and the single audio obtained by the audio restoration processing is determined as the target audio.
[0168] In one implementation, the noise compensation information is obtained by the processing unit 102 in the following manner: The feature information of the audio is successively extracted according to the audio obtained by the successive audio restoration processing, and the noise compensation information is successively updated according to the successively extracted feature information.
[0169] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0170] Figure 17 It is a block diagram of a device 200 for audio processing shown according to an exemplary embodiment. The device 200 may be provided as a terminal. For example, the device 200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0171] Referring to FIG.. The apparatus 200 may include one or more of the following components: a processing component 202, a memory 204, a power component 206, a multimedia component 208, an audio component 210, an input / output (I / O) interface 212, a sensor component 214, and a communication component 216.
[0172] The processing component 202 generally controls the overall operation of the apparatus 200, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 202 may include one or more processors 220 to execute instructions to complete all or part of the steps of the above-described methods. In addition, the processing component 202 may include one or more modules to facilitate the interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate the interaction between the multimedia component 208 and the processing component 202.
[0173] The memory 204 is configured to store various types of data to support the operation of the apparatus 200. Examples of such data include instructions for any application or method operating on the apparatus 200, contact data, phone book data, messages, pictures, videos, etc. The memory 204 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0174] The power component 206 provides power to the various components of the apparatus 200. The power component 206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the apparatus 200.
[0175] The multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0176] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC) that is configured to receive external audio signals when the device 200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 204 or transmitted via the communication component 216. In some embodiments, the audio component 210 further includes a speaker for outputting audio signals.
[0177] The I / O interface 212 provides an interface between the processing component 202 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0178] The sensor component 214 includes one or more sensors for providing a status assessment of various aspects of the device 200. For example, the sensor component 214 can detect the on / off state of the device 200, the relative positioning of components, such as the display and the keypad of the device 200. The sensor component 214 can also detect a change in the position of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200, and the temperature change of the device 200. The sensor component 214 can include a proximity sensor that is configured to detect the presence of nearby objects without any physical contact. The sensor component 214 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 214 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0179] The communication component 216 is configured to facilitate communication between the device 200 and other devices in a wired or wireless manner. The device 200 can access a communication standard-based wireless network, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0180] In an exemplary embodiment, the device 200 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0181] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 204 including instructions, and the above instructions can be executed by the processor 220 of the device 200 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0182] It can be understood that "a plurality of" in the present disclosure means two or more, and other quantifiers are similar thereto. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The singular forms of "a", "the", and "said" are also intended to include the plural forms unless the context clearly indicates otherwise.
[0183] It can be further understood that the terms "first", "second", etc. are used to describe various information, but this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other and do not represent a specific order or degree of importance. In fact, the expressions such as "first" and "second" can be used interchangeably. For example, without departing from the scope of the present disclosure, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information.
[0184] It can be further understood that the orientation or positional relationship indicated by terms such as "center", "longitudinal", "lateral", "front", "rear", "upper", "lower", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing this embodiment and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation.
[0185] It can be further understood that unless otherwise specified, "connection" includes direct connection without other components between the two, and also includes indirect connection with other elements between the two.
[0186] It can be further understood that although the operations are described in a specific order in the drawings in the embodiments of the present disclosure, it should not be understood as requiring these operations to be performed in the specific order or serial order shown, or requiring all the operations shown to obtain the desired result. In a specific environment, multitasking and parallel processing may be advantageous.
[0187] Those skilled in the art will readily think of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of this solution, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure.
[0188] It should be understood that the present disclosure is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. An audio processing method, characterized in that: include: Obtain the audio to be processed and obtain the text description information of the audio to be extracted; Acquire first characteristic information, and acquire phase information of the audio to be processed, wherein the first characteristic information is audio characteristic information of the audio to be processed; According to the text description information and the first feature information, second feature information is obtained through a diffusion model, where the second feature information is audio feature information of the audio to be extracted; The second characteristic information is subjected to audio restoration processing according to the phase information, and the audio obtained by the audio restoration processing is determined as the target audio.
2. The method according to claim 1, characterized in that The obtaining of the first characteristic information includes: Convert the time domain signal of the audio to be processed into a frequency domain signal, and obtain a spectrum diagram corresponding to the frequency domain signal; The spectrum diagram is encoded to obtain the first feature information.
3. The method according to claim 1, characterized in that The obtaining second feature information by using a diffusion model according to the text description information and the first feature information includes: Processing the text description information through a large language model to obtain a semantic embedding layer; Determine the quantity and identification type of the audio to be extracted according to the text description information; According to the quantity and recognition type of the audio to be extracted and the semantic embedding layer, performing noise processing on the first feature information to obtain noise-added feature information; According to the noise-added feature information and the semantic embedding layer, the first feature information is diffused by a diffusion model to obtain second feature information.
4. The method according to claim 3, characterized in that The identification type of the audio to be extracted includes a known type and an unknown type, the audio of the known type is the same as the audio type of the audio in the diffusion model training data, the audio of the unknown type is different from the audio type of the audio in the diffusion model training data, and the number of the audio to be extracted includes a single or a plurality of audios; The step of performing noise processing on the first feature information according to the quantity and recognition type of the audio to be extracted and the semantic embedding layer to obtain the noise-added feature information includes: In response to the number being single and the identification type of the single audio to be extracted being a known type, performing noise processing on the first feature information according to the semantic embedding layer to obtain the noise-added feature information; In response to the number being single and the identification type of the single audio to be extracted being an unknown type, obtaining zero-sample feature information, and performing noise processing on the first feature information according to the semantic embedding layer and the zero-sample feature information to obtain the noise-added feature information; In response to the number being multiple and the identification types of the multiple audios to be extracted being all known types, obtaining noise compensation information, and performing noise addition processing on the first feature information one by one according to the semantic embedding layer and the noise compensation information to obtain multiple pieces of noise addition feature information; In response to the fact that the number is multiple and there is audio of an unknown type among the multiple audios to be extracted, zero-sample feature information and noise compensation information are obtained, and the first feature information is noised one by one according to the semantic embedding layer, the zero-sample feature information and the noise compensation information to obtain multiple pieces of noisy feature information.
5. The method according to claim 4, characterized in that The obtaining of zero-sample feature information includes: According to the audio type of the unknown type of audio to be extracted, obtaining zero-sample audio of the same audio type; Convert the time domain signal of the zero-sample audio into a frequency domain signal, and obtain a spectrogram corresponding to the frequency domain signal; The spectrum graph is encoded to obtain the zero-sample feature information.
6. The method according to claim 4, characterized in that The noise compensation information is information that is updated successively. The step of performing noise processing on the first feature information one by one according to the semantic embedding layer and the noise compensation information to obtain multiple pieces of noise-added feature information includes: When performing the first noise addition process, performing noise addition process on the first feature information according to the semantic embedding layer to obtain the noise addition feature information; After completing the first noise addition process, successively obtaining updated noise compensation information, and performing noise addition processes on the first feature information one by one according to the semantic embedding layer and the successively updated noise compensation information to obtain multiple pieces of noise addition feature information; The noise feature information obtained when performing the first noise addition process and multiple pieces of noise feature information obtained by successive noise addition processes are determined as the noise feature information of the audio to be obtained.
7. The method according to claim 4, characterized in that The noise compensation information is information that is updated successively. The step of performing noise processing on the first feature information one by one according to the semantic embedding layer, the zero-sample feature information and the noise compensation information to obtain multiple pieces of noise-added feature information includes: When performing the first noise addition process, performing noise addition process on the first feature information according to the semantic embedding layer and the zero-sample feature information to obtain the noise addition feature information, or Performing noise processing on the first feature information according to the semantic embedding layer to obtain noisy feature information; After completing the first noise addition process, successively obtaining updated noise compensation information, and performing noise addition processes on the first feature information one by one according to the semantic embedding layer, the zero-sample feature information and the successively updated noise compensation information to obtain multiple pieces of noise addition feature information, or After the first noise addition process is completed, updated noise compensation information is successively obtained, and according to the semantic embedding layer and the successively updated noise compensation information, the first feature information is successively subjected to noise addition process to obtain multiple pieces of noise addition feature information; The noise feature information obtained when performing the first noise addition process and the multiple noise feature information obtained after completing the successive noise addition processes are determined as the noise feature information of the audio to be obtained.
8. The method according to claim 3, characterized in that The step of performing diffusion processing on the first feature information through a diffusion model according to the noise-added feature information and the semantic embedding layer to obtain second feature information includes: In response to the number of the audio to be extracted being a single one, performing diffusion processing on the first feature information through a diffusion model according to the noise-added feature information and the semantic embedding layer to obtain a single second feature information; In response to the number of the audios to be extracted being multiple, the first feature information is successively diffused through a diffusion model according to the semantic embedding layer and the noise-added feature information obtained one by one, so as to obtain multiple second feature information one by one.
9. The method according to claim 8, characterized in that The method further comprises: In response to the fact that the noisy feature information is not information obtained based on zero-sample feature information, performing denoising diffusion processing on the first feature information through a diffusion model according to the noisy feature information and the semantic embedding layer, eliminating feature information corresponding to the noisy feature information in the first feature information, and obtaining second feature information; In response to the fact that the noisy feature information is information obtained based on zero-sample feature information, a diffusion process is performed on the first feature information by mapping it through a diffusion model according to the noisy feature information and the semantic embedding layer, and feature information in the first feature information corresponding to the noisy feature information is retained to obtain second feature information.
10. The method according to claim 1, characterized in that The performing audio restoration processing on the second feature information according to the phase information, and determining the audio obtained by the audio restoration processing as the target audio, includes: In response to the number of audios to be extracted being multiple, performing audio restoration processing on the second characteristic information obtained successively according to the phase information, and determining the multiple audios obtained by the audio restoration processing as the target audio; In response to the number of audios to be extracted being single, audio restoration processing is performed on the second characteristic information obtained according to the phase information, and the single audio obtained by the audio restoration processing is determined as the target audio.
11. The method according to claim 4, characterized in that The noise compensation information is obtained in the following manner: According to the audio obtained by the successive audio restoration processing, the characteristic information of the audio is extracted one by one, and according to the successively extracted characteristic information, the noise compensation information is updated one by one.
12. An audio processing device, characterized in that: include: An acquisition unit, used to acquire the audio to be processed, acquire text description information of the audio to be extracted, acquire first feature information, and acquire phase information of the audio to be processed, wherein the first feature information is audio feature information of the audio to be processed; A processing unit, configured to obtain second feature information through a diffusion model according to the text description information and the first feature information, wherein the second feature information is audio feature information of the audio to be extracted; A determination unit is used to perform audio restoration processing on the second characteristic information according to the phase information, and determine the audio obtained by the audio restoration processing as the target audio.
13. The device according to claim 12, characterized in that The acquisition unit acquires the first characteristic information in the following manner: Convert the time domain signal of the audio to be processed into a frequency domain signal, and obtain a spectrum diagram corresponding to the frequency domain signal; The spectrum diagram is encoded to obtain the first feature information.
14. The device according to claim 12, characterized in that The processing unit obtains the second feature information by using a diffusion model according to the text description information and the first feature information in the following manner: Processing the text description information through a large language model to obtain a semantic embedding layer; Determine the quantity and identification type of the audio to be extracted according to the text description information; According to the quantity and recognition type of the audio to be extracted and the semantic embedding layer, performing noise processing on the first feature information to obtain noise-added feature information; According to the noise-added feature information and the semantic embedding layer, the first feature information is diffused by a diffusion model to obtain second feature information.
15. The device according to claim 14, characterized in that The identification type of the audio to be extracted includes a known type and an unknown type, the audio of the known type is the same as the audio type of the audio in the diffusion model training data, the audio of the unknown type is different from the audio type of the audio in the diffusion model training data, and the number of the audio to be extracted includes a single or a plurality of audios; The processing unit performs noise addition processing on the first feature information according to the quantity and recognition type of the audio to be extracted and the semantic embedding layer to obtain the noise addition feature information in the following manner: In response to the number being single and the identification type of the single audio to be extracted being a known type, performing noise processing on the first feature information according to the semantic embedding layer to obtain the noise-added feature information; In response to the number being single and the identification type of the single audio to be extracted being an unknown type, obtaining zero-sample feature information, and performing noise processing on the first feature information according to the semantic embedding layer and the zero-sample feature information to obtain the noise-added feature information; In response to the number being multiple and the identification types of the multiple audios to be extracted being all known types, obtaining noise compensation information, and performing noise addition processing on the first feature information one by one according to the semantic embedding layer and the noise compensation information to obtain multiple pieces of noise addition feature information; In response to the fact that the number is multiple and there is audio of an unknown type among the multiple audios to be extracted, zero-sample feature information and noise compensation information are obtained, and the first feature information is noised one by one according to the semantic embedding layer, the zero-sample feature information and the noise compensation information to obtain multiple pieces of noisy feature information.
16. The device according to claim 15, characterized in that The processing unit obtains zero-sample feature information in the following manner: According to the audio type of the unknown type of audio to be extracted, obtaining zero-sample audio of the same audio type; Convert the time domain signal of the zero-sample audio into a frequency domain signal, and obtain a spectrogram corresponding to the frequency domain signal; The spectrum graph is encoded to obtain the zero-sample feature information.
17. The device according to claim 15, characterized in that The noise compensation information is information that is updated successively. The processing unit performs noise processing on the first feature information one by one according to the semantic embedding layer and the noise compensation information to obtain multiple pieces of noise-added feature information in the following manner: When performing the first noise addition process, performing noise addition process on the first feature information according to the semantic embedding layer to obtain the noise addition feature information; After completing the first noise addition process, successively obtaining updated noise compensation information, and performing noise addition processes on the first feature information one by one according to the semantic embedding layer and the successively updated noise compensation information to obtain multiple pieces of noise addition feature information; The noise feature information obtained when performing the first noise addition process and multiple pieces of noise feature information obtained by successive noise addition processes are determined as the noise feature information of the audio to be obtained.
18. The device according to claim 15, characterized in that The noise compensation information is information that is updated successively. The processing unit performs noise addition processing on the first feature information one by one according to the semantic embedding layer, the zero-sample feature information and the noise compensation information to obtain multiple pieces of noise addition feature information in the following manner: When performing the first noise addition process, performing noise addition process on the first feature information according to the semantic embedding layer and the zero-sample feature information to obtain the noise addition feature information, or Performing noise processing on the first feature information according to the semantic embedding layer to obtain noisy feature information; After completing the first noise addition process, successively obtaining updated noise compensation information, and performing noise addition processes on the first feature information one by one according to the semantic embedding layer, the zero-sample feature information and the successively updated noise compensation information to obtain multiple pieces of noise addition feature information, or After the first noise addition process is completed, updated noise compensation information is successively obtained, and according to the semantic embedding layer and the successively updated noise compensation information, the first feature information is successively subjected to noise addition process to obtain multiple pieces of noise addition feature information; The noise feature information obtained when performing the first noise addition process and the multiple noise feature information obtained after completing the successive noise addition processes are determined as the noise feature information of the audio to be obtained.
19. The device according to claim 14, characterized in that The processing unit performs diffusion processing on the first feature information through a diffusion model according to the noise-added feature information and the semantic embedding layer to obtain second feature information in the following manner: In response to the number of the audio to be extracted being a single one, performing diffusion processing on the first feature information through a diffusion model according to the noise-added feature information and the semantic embedding layer to obtain a single second feature information; In response to the number of the audios to be extracted being multiple, the first feature information is successively diffused through a diffusion model according to the semantic embedding layer and the noise-added feature information obtained one by one, so as to obtain multiple second feature information one by one.
20. The device according to claim 19, characterized in that The processing unit is also used for: In response to the fact that the noisy feature information is not information obtained based on zero-sample feature information, performing denoising diffusion processing on the first feature information through a diffusion model according to the noisy feature information and the semantic embedding layer, eliminating feature information corresponding to the noisy feature information in the first feature information, and obtaining second feature information; In response to the fact that the noisy feature information is information obtained based on zero-sample feature information, a diffusion process is performed on the first feature information by mapping it through a diffusion model according to the noisy feature information and the semantic embedding layer, and feature information in the first feature information corresponding to the noisy feature information is retained to obtain second feature information.
21. The device according to claim 12, characterized in that The determining unit performs audio restoration processing on the second feature information according to the phase information in the following manner, and determines the audio obtained by the audio restoration processing as the target audio: In response to the number of audios to be extracted being multiple, performing audio restoration processing on the second characteristic information obtained successively according to the phase information, and determining the multiple audios obtained by the audio restoration processing as the target audio; In response to the number of audios to be extracted being single, audio restoration processing is performed on the second characteristic information obtained according to the phase information, and the single audio obtained by the audio restoration processing is determined as the target audio.
22. The device according to claim 15, characterized in that The noise compensation information is obtained by the processing unit in the following manner: According to the audio obtained by the successive audio restoration processing, the characteristic information of the audio is extracted one by one, and according to the successively extracted characteristic information, the noise compensation information is updated one by one.
23. An audio processing device, characterized in that: include: processor: a memory for storing processor-executable instructions; Wherein, the processor is configured to: execute the audio processing method according to any one of claims 1 to 11.
24. A storage medium, characterized in that The storage medium stores instructions, and when the instructions in the storage medium are executed by a processor, the processor is enabled to execute the audio processing method according to any one of claims 1 to 11.