Audio generation model training method, audio generation method and device

The audio generation model is trained through the latent spatial stream matching and self-cross attention mechanism, and the accuracy and efficiency problems when generating two-channel stereo are solved, achieving efficient generation of high-quality stereo audio.

CN120356454AActive Publication Date: 2025-07-22BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510856929.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-22
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

When the existing audio generation model generates two-channel stereo, the lack of sufficient training data leads to poor accuracy of the generated audio content and low training efficiency.

Method used

The audio generation model is trained using the method of potential spatial stream matching. By obtaining the scale spectrum of the two-channel stereo data and adding noise, combining self-attention and cross-attention mechanisms, a sample sequence is generated using descriptive text or video features, and multiple iterative processing is performed to generate two-channel stereo data.

Benefits of technology

It improves the accuracy and training efficiency of generated content, and can enhance the sense of space and reality while ensuring audio quality, and enhance user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356454A_ABST
    Figure CN120356454A_ABST
Patent Text Reader

Abstract

The invention provides a training method of an audio generation model, and an audio generation method and device, and belongs to the technical field of audio. Through the training process, training of the generative model can be carried out by adopting a potential spatial stream matching mode, in the model training process, rapid convergence of prediction is guided under the condition of describing a text or a video, and on the basis of a self-attention mechanism of the generative model, a cross attention mechanism is combined, so that the prediction efficiency is improved. According to the method, simpler expression can be achieved in the form of predicting the proportion spectrum, so that the training efficiency can be improved while the accuracy of the generated content is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of audio technology, and particularly to a training method for an audio generation model, an audio generation method, and an apparatus. Background Art

[0002] With the rapid development of multimedia technology, audio processing technology has also been continuously advancing, especially in the fields of text-to-audio (TTA) and video-to-audio (VTA). The applications of these technologies are extensive and are reflected in many aspects such as online education, entertainment, and virtual reality. However, in the pursuit of a higher immersion experience, the limitations of traditional mono audio output have gradually emerged, especially in expressing the sense of space and direction. Therefore, how to generate stereo audio has become a current research hotspot.

[0003] Current solutions usually directly train a sound generation model through video conditions or text conditions during the TTA or VTA process, attempting Figure 1 to directly generate stereo audio with azimuth information end-to-end. Although this method can meet basic needs to a certain extent, due to the lack of sufficient stereo audio for training, the generated audio often has poor content accuracy. Therefore, there is an urgent need for a training method for an audio generation model that can improve accuracy. Summary of the Invention

[0004] The present disclosure provides a training method for an audio generation model, an audio generation method, and an apparatus, which can improve the training efficiency while ensuring the accuracy of the generated content. The technical solutions of the present disclosure are as follows: According to one aspect of the embodiments of the present disclosure, a training method for an audio generation model is provided, including: Obtaining training data, where the training data includes a plurality of stereo audio data, and the training data further includes at least one of a sample video from which the plurality of stereo audio data is derived and a description text of the sample video; Obtaining a ratio spectrum of each of the stereo audio data, where the ratio spectrum represents the ratio of the difference to the sum of the spectra of the left and right channel data in the stereo audio data; Adding noise to a plurality of the ratio spectra based on a Gaussian distribution; Generating a plurality of first sample sequences based on each of the sample videos and the stereo audio data, where the first sample sequences represent the sample videos, the average spectrum of the stereo audio data, and the ratio spectra after adding noise; Generating a second sample sequence based on at least one of the sample video and the description text of the sample video; Using the first sample sequence as the main input of the model and the second sample sequence as the auxiliary input of the model, the audio generation model processes the first sample sequence and the second sample sequence to obtain a predicted vector field corresponding to each of the first sample sequences; Based on the noise-added ratio spectrum, the predicted vector field corresponding to each of the first sample sequences, and the Gaussian distribution, the audio generation model is trained.

[0005] In some embodiments, the audio generation model includes a self-attention layer and a cross-attention layer. The query vector, key vector, and value vector of the self-attention layer are all obtained based on the first sample sequence. The query vector of the cross-attention layer is obtained based on the first sample sequence, and the key vector and value vector of the cross-attention layer are both obtained based on the second sample sequence.

[0006] In some embodiments, the adding noise to multiple ratio spectra includes: Based on the Gaussian distribution and a scaling coefficient selected from 0 to 1, multiple ratio spectra are added with noise.

[0007] In some embodiments, the generating multiple first sample sequences based on the respective sample videos and the stereo data includes: Based on the sum and difference between the spectra of the left and right channel data of the stereo data, the ratio spectrum is generated; The mean of the spectra of the left and right channel data of the stereo data is output as the average spectrum of the stereo data; The video features of the sample video are extracted, and the video features are aligned with the spectra of the left and right channel data to obtain the aligned video features; The ratio spectrum, the average spectrum, and the aligned video features are concatenated in dimension to obtain the first sample sequence.

[0008] In some embodiments, the generating the second sample sequence based on at least one of the video features of the sample video and the description text features of the video includes: If there is a description text for the sample video, the description text features of the sample video are extracted through a text encoder; The video features of the sample video and the description text features of the sample video are concatenated in length to obtain the second sample sequence.

[0009] According to another aspect of the embodiments of the present disclosure, there is provided a training device for an audio generation model, including: A first data acquisition unit, configured to acquire training data, where the training data includes a plurality of stereo data, and the training data further includes at least one of a sample video from which the plurality of stereo data is derived and a description text of the sample video; A ratio spectrum acquisition unit, configured to acquire a ratio spectrum of each of the stereo data, where the ratio spectrum represents a ratio of a difference to a sum of spectra of left and right channel data in the stereo data; A noise addition unit, configured to add noise to the plurality of ratio spectra based on a Gaussian distribution; A first sequence generation unit, configured to generate a plurality of first sample sequences based on each of the sample videos and the stereo data, where the first sample sequences represent the sample videos, an average spectrum of the stereo data, and the ratio spectra after adding noise; A second sequence generation unit, configured to generate a second sample sequence based on at least one of the sample video and the description text of the sample video; A training unit, configured to use the first sample sequences as the main input of the model and the second sample sequences as the auxiliary input of the model, and process the first sample sequences and the second sample sequences through an audio generation model to obtain a predicted vector field corresponding to each of the first sample sequences; and train the audio generation model based on the ratio spectra after adding noise, the predicted vector fields corresponding to the first sample sequences, and the Gaussian distribution.

[0010] In some embodiments, the audio generation model includes a self-attention layer and a cross-attention layer, The query vector, key vector, and value vector of the self-attention layer are all obtained based on the first sample sequences, the query vector of the cross-attention layer is obtained based on the first sample sequences, and the key vector and value vector of the cross-attention layer are both obtained based on the second sample sequences.

[0011] In some embodiments, the noise addition unit is configured to add noise to the plurality of ratio spectra based on the Gaussian distribution and a scaling coefficient selected from 0 to 1.

[0012] In some embodiments, the first sequence generation unit is configured to generate the ratio spectrum based on a sum and a difference between spectra of left and right channel data of the stereo data; Output an average spectrum of the stereo data as an average of spectra of left and right channel data of the stereo data; Extract video features of the sample video, align the video features with spectra of the left and right channel data to obtain aligned video features; Concatenate the ratio spectrum, the average spectrum, and the aligned video features in dimension to obtain the first sample sequence.

[0013] In some embodiments, the second sequence generation unit is configured to, if there is a description text for the sample video, extract the description text features of the sample video through a text encoder; Concatenate the video features of the sample video and the description text features of the sample video in length to obtain the second sample sequence.

[0014] According to another aspect of the embodiments of the present disclosure, there is provided an audio generation method, including: Obtain monophonic audio data; Perform multiple rounds of iterative processing on the monophonic audio data based on an audio generation model to obtain a target ratio spectrum, where the audio generation model is trained based on a latent space flow matching algorithm, and the target ratio spectrum represents the ratio of the difference to the sum of the spectra of the left and right channel data in the stereo data to be generated; Generate left and right channel data based on the target ratio spectrum and the monophonic audio data; Obtain the stereo data corresponding to the monophonic audio data based on the left and right channel data.

[0015] In some embodiments, the performing multiple rounds of iterative processing on the monophonic audio data based on the audio generation model to obtain a target ratio spectrum includes: In each round of iteration, input the ratio spectrum obtained in the previous round of iteration into the audio generation model for processing to obtain an intermediate ratio spectrum and a vector field, sample the vector field based on an ODE (Ordinary Differential Equations) solver, update the intermediate ratio spectrum based on the sampling points, and output the ratio spectrum and the vector field obtained in this round of iteration.

[0016] According to another aspect of the embodiments of the present disclosure, there is provided an audio generation device, including: A second data acquisition unit configured to obtain monophonic audio data; A ratio spectrum generation unit configured to perform multiple rounds of iterative processing on the monophonic audio data based on an audio generation model to obtain a target ratio spectrum, where the audio generation model is trained based on a latent space flow matching algorithm, and the target ratio spectrum represents the ratio of the difference to the sum of the spectra of the left and right channel data in the stereo data to be generated; A left and right channel generation unit configured to generate left and right channel data based on the target ratio spectrum and the monophonic audio data; A stereo generation unit configured to obtain two-channel stereo data corresponding to the mono audio data based on the left and right channel data.

[0017] In some embodiments, the ratio spectrum generation unit is configured to, in each iteration process, input the ratio spectrum obtained in the previous iteration into an audio generation model for processing to obtain an intermediate ratio spectrum and a vector field, sample the vector field based on an ODE solver, update the intermediate ratio spectrum based on the sampling points, and output the ratio spectrum and the vector field obtained in this iteration.

[0018] According to another aspect of the embodiments of the present disclosure, there is provided an electronic device, which includes: One or more processors; A memory for storing executable program code of the processor; Wherein, the processor is configured to execute the program code to implement the above-mentioned training method or audio generation method of the audio generation model.

[0019] According to another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the program code in the computer-readable storage medium is executed by a processor of an electronic device, enabling the electronic device to execute the above-mentioned training method or audio generation method of the audio generation model.

[0020] According to another aspect of the embodiments of the present disclosure, there is provided a computer program product, including computer programs / instructions, which implement the above-mentioned training method or audio generation method of the audio generation model when executed by a processor.

[0021] Through the above training process, the generative model can be trained in a way of latent space flow matching. During the model training process, not only is the prediction guided to converge quickly based on the descriptive text or video, but also the cross-attention mechanism is combined on the basis of the self-attention mechanism of the generative model, so as to achieve a simpler expression in the form of predicting the ratio spectrum, thereby improving the training efficiency while ensuring the accuracy of the generated content.

[0022] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.

[0024] Figure 1 It is a schematic diagram of an implementation environment shown according to an exemplary embodiment.

[0025] Figure 2 It is a flowchart of a method for training an audio generation model shown according to an exemplary embodiment.

[0026] Figure 3 It is a schematic structural diagram of model training shown according to an exemplary embodiment.

[0027] Figure 4 It is a flowchart of a method for generating audio shown according to an exemplary embodiment.

[0028] Figure 5 It is a schematic diagram of stereo generation shown according to an exemplary embodiment.

[0029] Figure 6 It is a block diagram of an apparatus for training an audio generation model shown according to an exemplary embodiment.

[0030] Figure 7 It is a block diagram of an apparatus for generating audio shown according to an exemplary embodiment.

[0031] Figure 8 It is a block diagram of a terminal shown according to an exemplary embodiment.

[0032] Figure 9 It is a block diagram of a server shown according to an exemplary embodiment. Detailed implementation manners

[0033] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0035] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the task execution progress, objects, etc. involved in this disclosure are all obtained under full authorization.

[0036] LFM (latent flow matching) is a method for training generative models. It efficiently generates new samples by learning the optimal transport path (flow) between data distributions in the latent space.

[0037] VTA (Video-to-Audio) is a model that generates corresponding audio (such as ambient sound or dialogue) from video. By learning the association between vision and sound, it realizes cross-modal generation from vision to audition.

[0038] Text-to-Audio (TTA) is a technology that generates corresponding audio (such as ambient sound effects, music, or speech) from text. By modeling the relationship between language and sound, it realizes cross-modal generation from text to auditory content.

[0039] Wav-VAE is a model based on the variational autoencoder (VAE) for audio generation and reconstruction at the raw waveform level. It realizes high-quality sound modeling by learning the latent representation of audio data.

[0040] Figure 1 It is a schematic diagram of an implementation environment shown according to an exemplary embodiment. Taking the electronic device being provided as a terminal as an example, refer to Figure 1 This implementation environment specifically includes: a terminal 101 and a server 102. The terminal 101 and the server 102 can be directly or indirectly connected through wired or wireless communication methods, which are not limited in this disclosure.

[0041] The terminal 101 is at least one of devices such as a smart phone, a smart watch, a desktop computer, a laptop computer, an MP3 player, an MP4 player, and a laptop portable computer. An application program supporting model training runs on the terminal 101. The user can log in to the application program through the terminal 101 to obtain the services provided by the application program. For example, the user can perform model training through the application program on the terminal 101. Of course, the terminal 101 can also be an application program running with support for calling an audio generation model, and can generate stereophonic audio data by calling the audio generation model.

[0042] The terminal 101 generally refers to one of multiple terminals. In this embodiment, the terminal 101 is used as an example for illustration. Those skilled in the art can understand that the number of the above terminals can be more or less. For example, the above terminals can be several, or dozens or hundreds of the above terminals, or a larger number. The embodiments of the present disclosure do not limit the number and device type of the terminals.

[0043] The server 102 is at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. The server 102 is used to provide background services for model training. The server 102 can be connected to the terminal 101 through a wireless network or a wired network. The server 102 can provide database services and the like for the terminal. Of course, other types of services can also be provided, such as data preprocessing and the like. In some embodiments, the number of the above servers can be more or less, and the embodiments of the present disclosure do not limit this. Of course, the server 102 also includes other functional servers to provide more comprehensive and diverse services. Of course, in other embodiments, the server 102 can also be used as a carrier for model training to implement the model training process of the terminal 101 as described above, so as to provide the trained audio generation model for the terminal to use.

[0044] Figure 2 is a flowchart of a method for training an audio generation model provided by an embodiment of the present application. Figure 3 is a schematic structural diagram of model training shown according to an exemplary embodiment. Combining Figure 2 and Figure 3 , the method includes the following steps.

[0045] In step 201, training data is obtained. The training data includes multiple dual-channel stereo data, and the training data also includes at least one of the sample video from which the multiple dual-channel stereo data is sourced and the description text of the sample video.

[0046] For the training data, it can include multiple batches of training data. Each batch of training data includes multiple dual-channel stereo data, and moreover, the training data can also include at least one of the sample video from which the dual-channel stereo data is sourced and the description text. For the model training process, each time a batch of training data can be input to perform an iterative process, and thus the parameters of the audio generation model are adjusted once.

[0047] In step 202, the ratio spectrum of the dual-channel stereo data of each sample video is obtained. The ratio spectrum represents the ratio of the difference to the sum of the spectra of the left and right channel data in the dual-channel stereo data.

[0048] In the embodiments of the present application, the spectra are respectively extracted for the left and right channel data in the stereo data. The spectra can be in the form of Mel spectra or linear spectra, etc. After the spectra are extracted, the ratio spectra can be obtained based on the spectra of the left and right channel data respectively. For example, taking the Mel spectrum as an example, the formula for the ratio spectrum can be expressed as Formula 1 below.

[0049]

[0050] Subratio is used to represent the ratio spectrum, is used to represent the left channel data in the stereo data, is used to represent the right channel data in the stereo data.

[0051] The ratio spectrum calculated by the method of Formula 1 has a value between -1 and 1, which is convenient for model prediction. Moreover, the size of the ratio spectrum is the same as that of the single-channel spectrum, and the prediction difficulty is lower than predicting the two-channel spectrum simultaneously, greatly reducing the calculation difficulty.

[0052] In step 203, multiple ratio spectra are noise-added based on the Gaussian distribution.

[0053] In some embodiments, noise-adding multiple ratio spectra includes: noise-adding multiple ratio spectra based on the Gaussian distribution and a scaling coefficient selected from 0 to 1.

[0054] It should be noted that in the embodiments of the present application, LFM (Latent Flow Matching) is used to train the audio generation model, which is equivalent to learning a continuous probability flow so as to be able to convert a simple prior distribution (such as the standard Gaussian distribution) into a complex data distribution. During training, a "noise-adding - denoising" process is adopted, gradually adding noise to the input until it becomes pure noise; while the inference process of the trained audio generation model can be regarded as a process of denoising the input. Therefore, it is necessary to add noise to the samples during the training process to obtain the model input.

[0055] For the data sampled from the latent space distribution the flow matching model defines a probability path from the Gaussian distribution to in the form of an ordinary differential equation (ODE): where is the time-dependent vector field, and The probability path describes the probability distribution of the latent variable at time such that ​ In the embodiments of the present application, an optimal transmission path is adopted, and this path is defined by a linear transformation between the data distribution and the Gaussian distribution as formula two below.

[0056] Formula two:

[0057] Wherein, is a small constant, is the above-mentioned ratio spectrum, t is a time randomly sampled according to the cosine distribution in the interval i.e., the scaling coefficient, is the Gaussian distribution, is the ratio spectrum after adding noise.

[0058] In addition, using the Optimal Transport technology to construct the probability path can further accelerate the training speed and improve the generalization ability of the model.

[0059] In step 204, based on each sample video and the two-channel stereo data, a plurality of first sample sequences are generated, and the first sample sequences represent the average spectrum of the sample video, the two-channel stereo data, and the ratio spectrum after adding noise.

[0060] Among them, when extracting the video features of the sample video, it can be performed by a video encoder. The most basic clip model can be used as the video encoder to reduce the overall training cost. It can be understood that if the training data does not include the sample video, then the video features of the sample video do not need to be spliced, but directly filled with preset features.

[0061] In some embodiments, when generating the first sample sequence, the following process can be adopted: output the mean of the spectra of the left and right channel data of the two-channel stereo data as the average spectrum of the two-channel stereo data; extract the video features of the sample video, align the video features with the spectra of the left and right channel data to obtain the aligned video features; splice the ratio spectrum, the average spectrum, and the aligned video features in dimensions to obtain the first sample sequence.

[0062] Among them, the video features may originally have a small dimension, which may lead to inability to splice. In order to achieve the purpose of splicing, the video features and the spectra can be aligned, so as to achieve the purpose of splicing. This alignment method can be obtained by interpolating the video features, and the embodiments of the present application do not limit this.

[0063] In step 205, a second sample sequence is generated based on at least one of the video features of the sample video and the description text features of the video.

[0064] In step 205, if there is descriptive text in the sample video, the descriptive text features of the sample video are extracted through a text encoder; the video features of the sample video and the descriptive text features of the sample video are concatenated in length to obtain a second sample sequence. Among them, the text encoder can adopt the most basic T5 model to reduce the complexity during training. For video features, dimensionality reduction can be performed on them to achieve concatenation with the descriptive text features and avoid phenomena such as excessive length. It can be understood that if the training data does not include the sample video or descriptive text, the second sample sequence can be obtained based on the included features. In some embodiments, if necessary, the features with insufficient length can be padded to obtain the second sample sequence with the required length.

[0065] In step 206, using the first sample sequence as the main input of the model and the second sample sequence as the auxiliary input of the model, the audio generation model processes the first sample sequence and the second sample sequence to obtain the predicted vector field corresponding to each first sample sequence.

[0066] In some embodiments, the audio generation model includes a self-attention layer and a cross-attention layer. The query vector, key vector, and value vector of the self-attention layer are all obtained based on the first sample sequence. The query vector of the cross-attention layer is obtained based on the first sample sequence, and the key vector and value vector of the cross-attention layer are both obtained based on the second sample sequence. By adopting the method of combining the self-attention mechanism and the cross-attention mechanism for processing, it is possible to take into account the long-range dependencies within the sequence, establish associations at different time steps and feature dimensions, capture the long-term dependencies in the data, thereby better estimating the vector field, and it is also possible to consider the interaction between the different input sequences to enhance the guiding effect through the processing of multimodal information, so as to generate a more accurate vector field.

[0067] It should be noted that the above first sample sequence, as the main input of the audio generation model, participates in the calculations of each layer of the audio generation model, while the second sample sequence, as the auxiliary input, only participates in the calculations in the cross-attention layer, mainly used to generate the key vector and value vector in the cross-attention layer. Since the second sample sequence includes at least one of the descriptive text features and video features of the sample video, it is possible to consider the multimodal situation and guide the diffusion process to converge, thereby enhancing the controllability of the generation.

[0068] Among them, the audio generation model is a generative model. It should be noted that this audio generation model can be understood as a velocity field, which can obtain the vector field predicted by the model based on the noisy sample after random noise addition. Its input is the first sample sequence as the noisy sample and the second sample sequence as the constraint condition, and its output is the predicted vector field. Additionally, the t, which will also be used as the input of the transformer. For example, the audio generation model can be DiT (Diffusion Transformer), whose basic architecture is a transformer. The transformer includes an input layer, a self-attention layer, a cross-attention layer, and an output layer. Among them, the input layer can perform processing such as encoding and position encoding on the input. The self-attention layer can execute the self-attention mechanism based on the first sample sequence processed by the input layer. Correspondingly, it can also include a feed-forward network, etc., for further processing of the output. After the processing of the self-attention layer, the cross-attention layer executes the cross-attention mechanism based on the output of the self-attention layer, with the second sample sequence as the driving condition, integrates the conditional information into the processing process of the model to obtain cross-attention features, and through the processing of the output layer, outputs a predicted vector field. It should be noted that technicians can control the dimensions of the model, the depth of the model, and other parameters according to actual needs, so as to control the size of the model. This application embodiment does not make any limitations in this regard.

[0069] In step 207, the audio generation model is trained based on the noise-added ratio spectrum, the predicted vector fields corresponding to each first sample sequence, and the Gaussian distribution.

[0070] It should be noted that during the training process of the audio generation model, the following objective function is minimized to learn the vector field at time : Objective function:

[0071] Among them, is the parametric representation of the audio generation model, refers to finding the parameters that minimize the expression , is the expectation with respect to time and , represents the predicted vector field, and is the target vector field at time .

[0072] The goal of flow matching is to learn a vector field function such that starting from any initial point and flowing along this vector field, it can finally reach the target distribution. In the embodiments of this application, that is, the Gaussian distribution. Therefore, the purpose of the above objective function is to minimize the mean square error between the predicted vector field and the target vector field, and it can refer to the predicted vector field and time The difference between the target vector fields at [location] and the optimization objective are used to update the parameters of the model through the backpropagation algorithm, enabling the model to gradually learn the correct vector field, thereby achieving an accurate mapping from the noise distribution to the data distribution, and making the vector field predicted by the model as accurately describe the true flow of the data as possible.

[0073] Through the above training process, the generative model can be trained using the latent space flow matching method. During the model training process, not only is the prediction guided to converge quickly by conditioning on the description text or video, but also the cross-attention mechanism is combined based on the self-attention mechanism of the generative model, so as to achieve a simpler expression in the form of a predicted ratio spectrum, thereby being able to improve the training efficiency while ensuring the accuracy of the generated content.

[0074] In addition, the improvement of the training efficiency of the present application not only solves the problem of insufficient training data currently, but also provides the ability to flexibly adjust the sound characteristics in different scenarios.

[0075] Figure 4 is a flowchart of an audio generation method provided by an embodiment of the present application. Refer to Figure 4 and the method includes the following steps.

[0076] In step 401, monophonic audio data is obtained.

[0077] In the embodiment of the present application, the monophonic audio data can refer to any kind of audio data, and the embodiment of the present application does not make any limitation thereto. In some embodiments, if there is a source video for the monophonic audio data, the source video can also be obtained as part of the input.

[0078] That is, the input sequence is obtained by splicing the spectrum of the monophonic audio data and the video features of the source video. Of course, in order to obtain an input sequence of the same length, a sequence with a preset length can also be spliced. In addition, in the embodiment of the present application, the audio generation model also has the input of the first sample sequence and the second sample sequence as shown in Figure 2 where the above input sequence is equivalent to the first sample sequence, and the second sample sequence as the condition can also include at least one of the video features of the source video of the monophonic audio data and the description text features. Its processing process in the model is the same as the above steps and will not be elaborated here.

[0079] In step 402, the monophonic audio data is iteratively processed multiple times based on the audio generation model to obtain a target ratio spectrum. The audio generation model is trained based on the latent space flow matching algorithm, and the target ratio spectrum represents the ratio of the difference to the sum of the spectra of the left and right channel data in the stereo data to be generated.

[0080] In an embodiment of the present application, the audio generation model can obtain the above-mentioned target ratio spectrum through multiple rounds of iterative processing. In each round of iteration, the ratio spectrum obtained in the previous round of iteration is input into the audio generation model for processing to obtain an intermediate ratio spectrum and a vector field. Based on the ODE solver, sampling is performed on the vector field, and the intermediate ratio spectrum is updated based on the sampling points, and the ratio spectrum and the vector field obtained in this round of iteration are output.

[0081] In the above process, an update process of the output is involved to generate a new input for the model. For a certain iteration process, the ratio spectrum obtained in the previous round of iteration is input into the audio generation model for processing. The audio generation model is also a velocity field model, which can output a corresponding vector field while outputting the ratio spectrum. The ODE solver calculates the sample position or sample difference at the next time t2 using the vector field based on the current sample position and the current time t1, and updates the intermediate ratio spectrum based on the sample position or sample difference, so as to be used as the input for the next round of iteration. Through continuous iteration, the ODE solver gradually performs iterative processing on the input monophonic audio data along the path defined by the vector field, thereby obtaining the target ratio spectrum. This process can be regarded as a process in which the vector field model and the ODE solver gradually denoise the input ratio spectrum in reverse chronological order of time, and finally obtain the target ratio spectrum at time 0, realizing the generation of data.

[0082] In step 403, left and right channel data are generated based on the target ratio spectrum and the monophonic audio data.

[0083] In an embodiment of the present application, when the target ratio spectrum is obtained, taking the mel spectrum as an example of the spectrum, the mel spectra of the left and right channels can be calculated using the following formulas three and four.

[0084] Formula three:

[0085] Formula four :

[0086] Wherein, mel is the mel spectrum of the monophonic audio data, is the generated left channel data, is the generated right channel data.

[0087] Among them, based on the target ratio spectrum and the monophonic audio data, the spectra of the left and right channels can be generated, and then the generated spectra can be converted into waveform audio, thereby obtaining the left and right channel data. This waveform has stronger anti-noise ability than wav-vae. When converting the waveform, lightweight vocoders such as BigVGAN, HiFi-GAN, Parallel WaveGAN, nsf-HifiGAN, or Vocos can be used to ensure real-time performance and sound quality.

[0088] In step 404, based on the left and right channel data, the stereo data corresponding to the monophonic audio data is obtained.

[0089] For the left and right channel data, by concatenating their waveforms in the dimension, the stereo data of the two channels can be obtained. This process can be reflected by Figure 5 the process shown.

[0090] In the technical solution provided by the embodiment of the present application, an audio generation model trained by using the latent space flow matching algorithm is utilized. Taking the description of text or video as a condition, the given monophonic audio data is processed to obtain a ratio spectrum that can represent the left and right channel data, so as to ensure the accuracy of the audio content while retaining the quality of the original audio, and enhancing the sense of space and realism of the user experience.

[0091] During the inference process, through experiments, it can be proved that the audio generation method adopted by the embodiment of the present application uses the Euler ODE solver to generate the required samples from the Gaussian distribution in only 10 steps, greatly improving the inference efficiency.

[0092] Figure 6 is a block diagram of a training device for an audio generation model shown according to an exemplary embodiment. Refer to Figure 6 , the training device for the audio generation model includes: A first data acquisition unit 601, configured to acquire training data, where the training data includes a plurality of stereo data of two channels, and the training data further includes at least one of a sample video from which the plurality of stereo data of two channels are derived and a description text of the sample video; A ratio spectrum acquisition unit 602, configured to acquire the ratio spectrum of each of the stereo data of two channels, where the ratio spectrum represents the ratio of the difference to the sum of the spectra of the left and right channel data in the stereo data of two channels; A noise addition unit 603, configured to add noise to the plurality of ratio spectra based on the Gaussian distribution; The first sequence generation unit 604 is configured to generate a plurality of first sample sequences based on the respective sample videos and the two-channel stereo data, where the first sample sequences represent the sample videos, the average spectrum of the two-channel stereo data, and the proportion spectrum after adding noise; The second sequence generation unit 605 is configured to generate a second sample sequence based on at least one of the sample video and the description text of the sample video; The training unit 606 is configured to use the first sample sequence as the main input of the model and the second sample sequence as the auxiliary input of the model, and process the first sample sequence and the second sample sequence through an audio generation model to obtain a predicted vector field corresponding to each of the first sample sequences; based on the proportion spectrum after adding noise, the predicted vector field corresponding to each of the first sample sequences, and the Gaussian distribution, train the audio generation model.

[0093] In some embodiments, the audio generation model includes a self-attention layer and a cross-attention layer. The query vector, key vector, and value vector of the self-attention layer are all obtained based on the first sample sequence. The query vector of the cross-attention layer is obtained based on the first sample sequence, and the key vector and value vector of the cross-attention layer are both obtained based on the second sample sequence.

[0094] In some embodiments, the noise addition unit is configured to add noise to a plurality of the proportion spectra based on a Gaussian distribution and a scaling coefficient selected from 0 to 1.

[0095] In some embodiments, the first sequence generation unit is configured to generate the proportion spectrum based on the sum and difference between the spectra of the left and right channel data of the two-channel stereo data; Output the mean value of the spectra of the left and right channel data of the two-channel stereo data as the average spectrum of the two-channel stereo data; Extract the video features of the sample video, align the video features with the spectra of the left and right channel data to obtain the aligned video features; Concatenate the proportion spectrum, the average spectrum, and the aligned video features in dimension to obtain the first sample sequence.

[0096] In some embodiments, the second sequence generation unit is configured to, if there is description text for the sample video, extract the description text features of the sample video through a text encoder; Concatenate the video features of the sample video and the description text features of the sample video in length to obtain the second sample sequence.

[0097] It should be noted that when training the audio generation model provided in the above embodiments, only the division of the above functional units is used for illustration. In practical applications, the above functions can be allocated to different functional units according to needs, that is, the internal structure of the electronic device is divided into different functional units to complete all or part of the functions described above. In addition, the audio generation model training device provided in the above embodiments and the audio generation method embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments and will not be elaborated here.

[0098] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0099] According to another aspect of the embodiments of the present disclosure, an audio generation device is provided. Refer to Figure 7 , the device includes: A second data acquisition unit 701, configured to acquire monophonic audio data; A proportional spectrum generation unit 702, configured to perform multiple rounds of iterative processing on the monophonic audio data based on an audio generation model to obtain a target proportional spectrum, where the audio generation model is trained based on a latent space flow matching algorithm, and the target proportional spectrum represents the ratio of the difference to the sum of the spectra of the left and right channel data in the to-be-generated stereophonic data; A left and right channel generation unit 703, configured to generate left and right channel data based on the target proportional spectrum and the monophonic audio data; A stereophonic generation unit 704, configured to obtain the stereophonic data corresponding to the monophonic audio data based on the left and right channel data.

[0100] In some embodiments, the proportional spectrum generation unit is configured to, in each round of iteration, input the proportional spectrum obtained in the previous round of iteration into the audio generation model for processing to obtain an intermediate proportional spectrum and a vector field, sample the vector field based on an ODE solver, update the intermediate proportional spectrum based on the sampling points, and output the proportional spectrum and the vector field obtained in this round of iteration.

[0101] It should be noted that when generating audio by the audio generation device provided in the above embodiments, only the division of the above functional units is used for illustration. In practical applications, the above functions can be allocated to different functional units according to needs, that is, the internal structure of the electronic device is divided into different functional units to complete all or part of the functions described above. In addition, the audio generation device provided in the above embodiments and the audio generation method embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments and will not be elaborated here.

[0102] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0103] When the electronic device is provided as a terminal, Figure 8 is a block diagram of a terminal 800 shown according to an exemplary embodiment. The Figure 8 shows a structural block diagram of a terminal 800 provided by an exemplary embodiment of the present disclosure. The terminal 800 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer, or a desktop computer. The terminal 800 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.

[0104] Generally, the terminal 800 includes: a processor 801 and a memory 802.

[0105] The processor 801 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 801 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 801 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 801 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 801 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0106] The memory 802 may include one or more computer-readable storage media, which may be non-transitory. The memory 802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 802 is used to store at least one computer program, and the at least one computer program is used to be executed by the processor 801 to implement the method provided in the method embodiments of the present application.

[0107] In some embodiments, the terminal 800 may further optionally include: a peripheral device interface 803 and at least one peripheral device. The processor 801, the memory 802, and the peripheral device interface 803 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 803 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 807, and a power supply 808.

[0108] The peripheral device interface 803 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 801 and the memory 802. In some embodiments, the processor 801, the memory 802, and the peripheral device interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 801, the memory 802, and the peripheral device interface 803 may be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0109] The radio frequency circuit 804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 804 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 804 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. In some embodiments, the radio frequency circuit 804 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and the like. The radio frequency circuit 804 may communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 804 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.

[0110] The display screen 805 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 805 is a touch display screen, the display screen 805 also has the ability to collect touch signals on or above the surface of the display screen 805. The touch signals can be input to the processor 801 as control signals for processing. At this time, the display screen 805 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there may be one display screen 805, which is disposed on the front panel of the terminal 800; in other embodiments, there may be at least two display screens 805, which are respectively disposed on different surfaces of the terminal 800 or are in a foldable design; in other embodiments, the display screen 805 may be a flexible display screen, which is disposed on the curved surface or the folding surface of the terminal 800. Even further, the display screen 805 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 805 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0111] The camera module 806 is used to capture images or videos. In some embodiments, the camera module 806 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, so as to implement functions such as background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera module 806 may further include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. A two-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0112] The audio circuit 807 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 801 for processing, or input to the radio frequency circuit 804 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 800. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 807 may further include a headphone jack.

[0113] The power supply 808 is used to supply power to each component in the terminal 800. The power supply 808 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 808 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery may also be used to support fast charging technology.

[0114] Those skilled in the art can understand that Figure 8 the structure shown in

[0115] does not limit the terminal 800, and may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements.

[0116] The above computer device may also be implemented as a server. The structure of the server will be introduced below: Figure 9It is a schematic structural diagram of a server provided by an embodiment of the present application. The server 900 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 901 and one or more memories 902. Among them, at least one computer program is stored in the one or more memories 902, and the at least one computer program is loaded and executed by the one or more processors 901 to implement the methods provided in the above various method embodiments. Of course, the server 900 may also have components such as wired or wireless network interfaces, keyboards, and input / output interfaces for input / output. The server 900 may also include other components for implementing device functions, which will not be elaborated here.

[0117] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including a computer program. The above computer program can be executed by a processor to complete the methods in the above embodiments. For example, the computer-readable storage medium can be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), magnetic tapes, floppy disks, and optical data storage devices, etc.

[0118] In an exemplary embodiment, a computer program product or a computer program is also provided. The computer program product or the computer program includes program code, and the program code is stored in a computer-readable storage medium. The processor of the computer device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computer device executes the above method.

[0119] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0120] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A training method for an audio generation model, characterized in that, The method includes: Obtaining training data, where the training data includes a plurality of stereo audio data, and the training data further includes at least one of a sample video from which the plurality of stereo audio data is sourced and a description text of the sample video; Obtaining a ratio spectrum for each of the stereo audio data, where the ratio spectrum represents the ratio of the difference to the sum of the spectra of the left and right channel data in the stereo audio data; Adding noise to the plurality of ratio spectra based on a Gaussian distribution; Generating a plurality of first sample sequences based on each of the sample videos and the stereo audio data, where the first sample sequence represents the sample video, the average spectrum of the stereo audio data, and the ratio spectrum after adding noise; Generating a second sample sequence based on at least one of the sample video and the description text of the sample video; Using the first sample sequence as the main input of the model and the second sample sequence as the auxiliary input of the model, and processing the first sample sequence and the second sample sequence through an audio generation model to obtain a predicted vector field corresponding to each of the first sample sequences; Training the audio generation model based on the ratio spectrum after adding noise, the predicted vector field corresponding to each of the first sample sequences, and the Gaussian distribution.

2. The training method of the audio generation model according to claim 1, wherein The audio generation model includes a self-attention layer and a cross-attention layer. The query vector, key vector, and value vector of the self-attention layer are all obtained based on the first sample sequence, the query vector of the cross-attention layer is obtained based on the first sample sequence, and the key vector and value vector of the cross-attention layer are both obtained based on the second sample sequence.

3. The training method of the audio generation model according to claim 1, wherein The adding noise to the plurality of ratio spectra includes: Adding noise to the plurality of ratio spectra based on the Gaussian distribution and a scaling factor selected from between 0 and 1.

4. The training method of the audio generation model according to claim 1, characterized in that The generating a plurality of first sample sequences based on each of the sample videos and the stereo audio data includes: Generating the ratio spectrum based on the sum and difference between the spectra of the left and right channel data of the stereo audio data; Outputting the mean of the spectra of the left and right channel data of the stereo audio data as the average spectrum of the stereo audio data; Extracting video features of the sample video, aligning the video features with the spectra of the left and right channel data to obtain aligned video features; Concatenating the ratio spectrum, the average spectrum, and the aligned video features in dimension to obtain the first sample sequence.

5. The training method of the audio generation model according to claim 1, wherein The generating a second sample sequence based on at least one of the video features of the sample video and the description text features of the video includes: If there is a description text for the sample video, extracting the description text features of the sample video through a text encoder; Concatenating the video features of the sample video and the description text features of the sample video in length to obtain the second sample sequence.

6. An audio generation method, characterized in that, Includes: Obtaining monophonic audio data; Performing multiple rounds of iterative processing on the mono audio data based on an audio generation model to obtain a target ratio spectrum, where the audio generation model is trained based on a latent space flow matching algorithm, and the target ratio spectrum represents the ratio of the difference to the sum of the spectra of the left and right channel data in the stereo data to be generated; Generating left and right channel data based on the target ratio spectrum and the mono audio data; Obtaining the stereo data corresponding to the mono audio data based on the left and right channel data.

7. The audio generation method according to claim 6, wherein The performing multiple rounds of iterative processing on the mono audio data based on the audio generation model to obtain a target ratio spectrum includes: In each round of iteration, inputting the ratio spectrum obtained in the previous round of iteration into the audio generation model for processing to obtain an intermediate ratio spectrum and a vector field, sampling the vector field based on an ordinary differential equation (ODE) solver, updating the intermediate ratio spectrum based on the sampling points, and outputting the ratio spectrum and the vector field obtained in this round of iteration.

8. A training device for an audio generation model, characterized in that, Including: A first data acquisition unit configured to acquire training data, where the training data includes multiple stereo data, and the training data further includes at least one of the sample videos from which the multiple stereo data are sourced and the description text of the sample videos; A ratio spectrum acquisition unit configured to acquire the ratio spectrum of each of the stereo data, where the ratio spectrum represents the ratio of the difference to the sum of the spectra of the left and right channel data in the stereo data; A noise addition unit configured to add noise to the multiple ratio spectra based on a Gaussian distribution; A first sequence generation unit configured to generate multiple first sample sequences based on the respective sample videos and the stereo data, where the first sample sequences represent the sample videos, the average spectrum of the stereo data, and the ratio spectra after noise addition; A second sequence generation unit configured to generate a second sample sequence based on at least one of the sample video and the description text of the sample video; A training unit configured to use the first sample sequences as the main input of the model and the second sample sequences as the auxiliary input of the model, and process the first sample sequences and the second sample sequences through the audio generation model to obtain the predicted vector fields corresponding to the respective first sample sequences; Training the audio generation model based on the ratio spectra after noise addition, the predicted vector fields corresponding to the respective first sample sequences, and the Gaussian distribution.

9. An audio generation device, characterized in that, Including: A second data acquisition unit configured to acquire mono audio data; A ratio spectrum generation unit configured to perform multiple rounds of iterative processing on the mono audio data based on an audio generation model to obtain a target ratio spectrum, where the audio generation model is trained based on a latent space flow matching algorithm, and the target ratio spectrum represents the ratio of the difference to the sum of the spectra of the left and right channel data in the stereo data to be generated; A left and right channel generation unit configured to generate left and right channel data based on the target ratio spectrum and the mono audio data; A stereo generation unit configured to obtain the stereo data corresponding to the mono audio data based on the left and right channel data.

10. An electronic device, characterized in that, The electronic device includes: One or more processors; A memory for storing program code executable by the processor; Wherein, the processor is configured to execute the program code to implement the training method of the audio generation model according to any one of claims 1 to 5; or, the audio generation method according to claim 6 or 7.

11. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the training method of the audio generation model according to any one of claims 1 to 5; Or, the audio generation method according to claim 6 or 7.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the training method of the audio generation model according to any one of claims 1 to 5; or, the audio generation method according to claim 6 or 7.

Citation Information

Patent Citations

  • Audio three-dimensional method based on multi-attention audio-visual fusion

    CN113099374A

  • Method and apparatus for speech / music classification and core encoder selection in sound codec

    CN115428068A

  • Audio generation method and device

    CN117676449A

  • Method for generating surround channel audio

    US20170171683A1

  • Speech synthesis method and apparatus, and readable storage medium

    US20230075891A1