Audio generation method and device, computer program product and electronic equipment

By inputting text into multimodal large model and performing spatial transformation and variational encoder processing, the Mel frequency cepspectral features are generated and input into the vocoder, the problem of scarcity of audio-text pairing annotation data is solved, and the efficiency of text generation of audio and model application scenarios are improved.

CN120148466APending Publication Date: 2025-06-13CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411280662.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, the scarcity and difficulty of obtaining audio-text pairing labeling data lead to low training efficiency of text generation audio models.

Method used

By inputting text into the multimodal large model, obtaining the output vector, performing spatial transformation and inputting the variational encoder, generating the Mel frequency cepspectral feature, and inputting it into the vocoder to generate audio.

Benefits of technology

This method simplifies the process of converting text into audio, improves the efficiency of text generation audio, and solves the training data acquisition problem through implicit spatial transformation model, broadening the application scenarios and practicality of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148466A_ABST
    Figure CN120148466A_ABST
Patent Text Reader

Abstract

The invention relates to an audio generation method and device, a computer program product and electronic equipment, and relates to the technical field of computers.The method comprises the steps that a target text is obtained, the target text is input into a target multi-mode large model, and a first representation vector is obtained; obtaining a target spatial transformation model, and inputting the first expression vector into the target spatial transformation model to obtain a second expression vector; inputting the second representation vector into a target variational encoder to obtain a Mel-frequency cepstrum feature corresponding to the second representation vector; and inputting the Mel-frequency cepstrum feature into a target vocoder to obtain a target audio corresponding to the target text. The efficiency of generating the audio based on the text is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] The development of multi-modal large models has further enriched the application scenarios of text-to-audio generation. Although several open-source text-audio multi-modal large models have been released in the industry, the further development in this field faces a key bottleneck: the scarcity and difficulty of obtaining audio-text paired annotation data.

[0003] The annotation work of such data is not only costly but also often strictly restricted by copyright issues, resulting in the inability to widely share and effectively utilize even a large amount of training data, which poses a severe challenge to the training of text-to-audio models. At the same time, it also leads to low efficiency in generating high-quality and high-fluency videos through text.

[0004] Therefore, a new audio generation method is needed.

[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0006] The purpose of the present disclosure is to provide an audio generation method, an audio generation device, a computer program product, and an electronic device, so as to at least to some extent overcome the problem of low efficiency in generating audio through text caused by the limitations and defects of related technologies.

[0007] According to one aspect of the present disclosure, an audio generation method is provided, including:

[0008] Inputting text into a multi-modal large model to obtain a first output vector;

[0009] Inputting the first output vector into a spatial transformation model to obtain a second output vector;

[0010] Inputting the second output vector into a variational encoder to obtain Mel-frequency cepstral coefficients corresponding to the second output vector;

[0011] Inputting the Mel-frequency cepstral coefficients into a vocoder to obtain the audio corresponding to the text.

[0012] In an exemplary embodiment of the present disclosure, before inputting the first output vector into the spatial transformation model, the method further includes:

[0013] Obtaining audio data, preprocessing the audio data to obtain audio sample data;

[0014] Input the first Mel-frequency cepstral coefficients (MFCCs) of the audio sample data into the encoder of a preset variational autoencoder (VAE) to obtain a third output vector;

[0015] Input the third output vector into the decoder of the preset VAE to obtain second MFCCs;

[0016] Construct a loss function based on the first MFCCs and the second MFCCs, and train the preset VAE to obtain the VAE.

[0017] In an exemplary embodiment of the present disclosure, the obtaining of the audio data and the preprocessing of the audio data to obtain audio sample data include:

[0018] Convert the audio data to the same sampling rate or the same number of data bits to obtain first audio data;

[0019] Augment the first audio data to obtain the audio sample data.

[0020] In an exemplary embodiment of the present disclosure, the constructing of the loss function based on the first MFCCs and the second MFCCs and the training of the preset VAE to obtain the VAE include:

[0021] Determine the multi-scale short-time Fourier transform loss of the first MFCCs and the second MFCCs as the loss function;

[0022] Train the preset VAE based on the loss function to obtain the VAE.

[0023] In an exemplary embodiment of the present disclosure, after obtaining the VAE, the method further includes:

[0024] Input the first MFCCs of the audio sample data into the encoder of the VAE to obtain a third output vector;

[0025] Input the first MFCCs into the multi-modal large model to obtain a fourth output vector;

[0026] Train a preset implicit space transformation model through the third output vector and the fourth output vector to obtain the space transformation model.

[0027] In an exemplary embodiment of the present disclosure, the training of the preset implicit space transformation model through the third output vector and the fourth output vector to obtain the space transformation model includes:

[0028] Input the fourth output vector into the preset implicit space transformation model to obtain a fifth output vector;

[0029] Construct a loss function based on the fifth output vector and the third output vector, and train the preset implicit space transformation model based on the loss function to obtain the space transformation model.

[0030] In an exemplary embodiment of the present disclosure, after obtaining the variational encoder, the method further includes:

[0031] Input the third Mel frequency cepstral features of the text corresponding to the audio sample data and the first Mel frequency cepstral features of the audio sample data into the multimodal large model to obtain a sixth output vector;

[0032] Input the first Mel frequency cepstral features into the encoder of the variational encoder to obtain a third output vector;

[0033] Train the preset implicit space transformation model through the third output vector and the sixth output vector to obtain the space transformation model.

[0034] In an exemplary embodiment of the present disclosure, the training the preset implicit space transformation model through the third output vector and the sixth output vector to obtain the space transformation model includes:

[0035] Input the sixth output vector into the preset implicit space transformation model to obtain a seventh output vector;

[0036] Construct a loss function based on the seventh output vector and the third output vector, and train the preset implicit space transformation model based on the loss function to obtain the space transformation model.

[0037] In an exemplary embodiment of the present disclosure, the method further includes:

[0038] Input the first Mel frequency cepstral features of the audio sample data into a preset vocoder to obtain a predicted audio signal;

[0039] Construct a loss function based on the predicted audio signal and the actual audio signal of the audio sample data, and train the preset vocoder based on the loss function to obtain the vocoder.

[0040] According to one aspect of the present disclosure, there is provided an audio generation device, including:

[0041] A first vector acquisition module, configured to input text into a multimodal large model to obtain a first output vector;

[0042] A second vector acquisition module, configured to input the first output vector into a spatial transformation model to obtain a second output vector;

[0043] A feature acquisition module, configured to input the second output vector into a variational encoder to obtain Mel-frequency cepstral coefficients features corresponding to the second output vector;

[0044] An audio generation module, configured to input the Mel-frequency cepstral coefficients features into a vocoder to obtain audio corresponding to the text.

[0045] According to one aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the audio generation method described in any one of the above exemplary embodiments.

[0046] According to one aspect of the present disclosure, there is provided an electronic device, including:

[0047] A processor; and

[0048] A memory for storing executable instructions of the processor;

[0049] wherein the processor is configured to execute the audio generation method described in any one of the above exemplary embodiments by executing the executable instructions.

[0050] An audio generation method provided by an embodiment of the present disclosure includes: inputting text into a multi-modal large model to obtain a first output vector; inputting the first output vector into a spatial transformation model to obtain a second output vector; inputting the second output vector into a variational encoder to obtain Mel-frequency cepstral coefficients features corresponding to the second output vector; and inputting the Mel-frequency cepstral coefficients features into a vocoder to obtain audio corresponding to the text. On the one hand, by inputting text into a multi-modal large model to obtain a first output vector, inputting the first output vector into a spatial transformation model to obtain a second output vector, and inputting the second output vector into a variational encoder to obtain Mel-frequency cepstral coefficients features, since the implicit space difference between the variational encoder and the multi-modal large model is huge, it is easy to have a situation where it is difficult to converge during training. Through the spatial transformation model, the long-standing problem of obtaining training data in the field of text-to-audio generation is solved, and at the same time, the application scenario and practicability of the model are broadened. On the other hand, after obtaining Mel-frequency cepstral coefficients features corresponding to the text through the multi-modal large model and the spatial transformation model, inputting the Mel-frequency cepstral coefficients features into a vocoder to obtain audio corresponding to the text simplifies the process of converting text to audio and improves the efficiency of text-to-audio generation.

[0051] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings

[0052] The drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0053] Figure 1 A flowchart schematically showing a method for generating audio according to an exemplary embodiment of the present disclosure.

[0054] Figure 2 A flowchart schematically showing a method for generating audio before inputting a first output vector into a spatial transformation model according to an exemplary embodiment of the present disclosure.

[0055] Figure 3 A flowchart schematically showing a method for obtaining audio data, preprocessing the audio data, and obtaining audio sample data according to an exemplary embodiment of the present disclosure.

[0056] Figure 4 A flowchart schematically showing a method for training a preset variational encoder to obtain a variational encoder according to an exemplary embodiment of the present disclosure.

[0057] Figure 5 A flowchart schematically showing a method for generating audio after obtaining a variational encoder according to an exemplary embodiment of the present disclosure.

[0058] Figure 6 A flowchart schematically showing a method for training a preset implicit spatial transformation model with a third output vector and a fourth output vector to obtain a spatial transformation model according to an exemplary embodiment of the present disclosure.

[0059] Figure 7 A flowchart schematically showing a method for generating audio after obtaining a variational encoder according to an exemplary embodiment of the present disclosure.

[0060] Figure 8 A schematic diagram schematically showing the structure of a spatial transformation model according to an exemplary embodiment of the present disclosure.

[0061] Figure 9 A flowchart schematically showing a method for generating audio according to an exemplary embodiment of the present disclosure.

[0062] Figure 10 A block diagram schematically showing an audio generation device according to an exemplary embodiment of the present disclosure.

[0063] Figure 11Schematically illustrate an electronic device for implementing the above audio generation method according to an exemplary embodiment of the present disclosure. Detailed implementation manners

[0064] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will recognize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring the various aspects of the present disclosure.

[0065] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0066] As a cutting-edge topic in the field of artificial intelligence, the text-to-audio generation model aims to convert text information into vivid and realistic speech or music through computer technology. In recent years, with the rapid development of deep learning technology, the text-to-audio generation model has made remarkable progress and can already generate natural and fluent audio with high sound quality, greatly improving the experience and efficiency of human-computer interaction.

[0067] At the same time, the development of multimodal large models has further enriched the application scenarios of text-to-audio generation. Although several open-source text-audio multimodal large models have been released in the industry, the further development of this field faces a key bottleneck: the scarcity and difficulty of obtaining audio-text paired annotation data. The annotation work for this type of data is not only costly but also often strictly restricted by copyright issues, resulting in the inability to widely share and effectively utilize even a large amount of training data, which poses a severe challenge to the training of text-to-audio generation models. Therefore, how to effectively solve the data problem has become one of the key factors driving the development of this field.

[0068] Based on the above one or more problems, in this exemplary embodiment, an audio generation method is first provided, and this method can run on the server side; of course, those skilled in the art can also run the method of the present invention on other platforms according to requirements, and no special limitation is made in this exemplary embodiment. Refer to Figure 1 As shown, this audio generation method may include steps S110 - S140:

[0069] Step S110. Input the text into a multi-modal large model to obtain a first output vector;

[0070] Step S120. Input the first output vector into a spatial transformation model to obtain a second output vector;

[0071] Step S130. Input the second output vector into a variational encoder to obtain the Mel-frequency cepstral coefficients corresponding to the second output vector;

[0072] Step S140. Input the Mel-frequency cepstral coefficients into a vocoder to obtain the audio corresponding to the text.

[0073] For the above audio generation method, the text is input into a multi-modal large model to obtain a first output vector; the first output vector is input into a spatial transformation model to obtain a second output vector; the second output vector is input into a variational encoder to obtain the Mel-frequency cepstral coefficients corresponding to the second output vector; the Mel-frequency cepstral coefficients are input into a vocoder to obtain the audio corresponding to the text. On the one hand, when the text is input into a multi-modal large model to obtain a first output vector, the first output vector is input into a spatial transformation model to obtain a second output vector, and the second output vector is input into a variational encoder to obtain the Mel-frequency cepstral coefficients. Since the implicit space between the variational encoder and the multi-modal large model is extremely different, it is easy to have a situation where it is difficult to converge during training. Through the spatial transformation model, the long-existing problem of obtaining training data in the field of text-to-audio generation is solved, and at the same time, the application scenario and practicability of the model are broadened; on the other hand, after obtaining the Mel-frequency cepstral coefficients corresponding to the text through the multi-modal large model and the spatial transformation model, the Mel-frequency cepstral coefficients are input into a vocoder to obtain the audio corresponding to the text, which simplifies the process of converting text to audio and improves the efficiency of text-to-audio generation.

[0074] Hereinafter, each step involved in the audio generation method of the exemplary embodiment of the present disclosure will be explained and described in detail.

[0075] First, the application scenarios and purposes of the exemplary embodiments of the present disclosure are explained and described. Specifically, the exemplary embodiments of the present disclosure can be used to generate audio corresponding to text. Based on model training, audio data is preprocessed to generate audio sample data. The preset variational encoder is trained using the audio sample data to obtain a variational encoder. After completing the training of the preset variational encoder, the preset spatial transformation model is trained to obtain a spatial transformation model; finally, the preset vocoder is trained to obtain a vocoder. After the model training is completed, the text is input into the multi-modal large model to obtain a first output vector. The first output vector is input into the spatial transformation model to obtain a second output vector. The second output vector is input into the variational encoder to obtain the Mel-frequency cepstral coefficients corresponding to the second output vector, and the Mel-frequency cepstral coefficients are input into the vocoder to obtain the audio corresponding to the text.

[0076] Secondly, the audio generation algorithm framework includes: a variational encoder, a multi-modal large model, a spatial transformation model, and a vocoder. Among them, the variational encoder includes an encoder and a decoder. The encoder takes the Mel-frequency cepstral coefficients of the audio as input, and the output is the implicit embedding representation vector encoded from the input. The decoder takes the implicit embedding representation vector as input, and the output is the Mel-frequency cepstral coefficients predicted according to the input implicit embedding representation vector. The encoder and the decoder are jointly trained. The target multi-modal large model can be a text-audio multi-modal large model, which is a model that has been trained with a large amount of text-audio data. For example, the open-source large model CLAP takes the Mel-frequency cepstral coefficients of text or audio as input, and the output is the implicit embedding representation vector corresponding to the text or audio. The target multi-modal large model can map text and audio to the same-dimensional implicit embedding representation space, and the implicit embedding representation vectors generated by the relevant audio and text are also highly correlated. The spatial transformation model is used to convert the implicit embedding vector of the multi-modal large model into the implicit embedding representation vector generated by the variational encoder. The vocoder takes the Mel-frequency cepstral coefficients as input, and the output is the audio generated by the Mel-frequency cepstral coefficients.

[0077] Hereinafter, steps S110 - S140 are explained and described in detail.

[0078] In step S110, the text is input into the multi-modal large model to obtain a first output vector.

[0079] Among them, the text can be lyrics text or game narration text, and no specific limitation is imposed on the text in this exemplary embodiment. The multimodal large model can be a trained text-audio multimodal large model. This multimodal large model takes the Mel-frequency cepstral features of the text or the Mel-frequency cepstral features of the audio as input, and the output is the implicit embedding representation vector corresponding to the text or the audio. This multimodal large model can map the text and the audio to the implicit embedding representation space of the same dimension, and the implicit embedding representation vectors generated by the relevant audio and text are also highly correlated. For example, this multimodal large model can be the open-source large model CLAP, or it can be other open-source large models, and no specific limitation is imposed on this in this exemplary embodiment.

[0080] After obtaining the text, calculate the Mel-frequency cepstral features of the text, and input the Mel-frequency cepstral features into the multimodal large model to obtain a first output vector, which is an implicit embedding representation vector. After obtaining the implicit embedding representation vector, obtain a space transformation model, and input the first output vector into the space transformation model.

[0081] In step S120, input the first output vector into the space transformation model to obtain a second output vector.

[0082] The space transformation model is obtained by training a preset implicit space transformation model. Since the implicit spaces of the variational encoder and the multimodal large model are very different, it is easy to encounter difficult convergence situations during training. Therefore, this preset implicit space transformation model is used to convert the output vector of the multimodal large model into the representation vector generated by the variational encoder. The variational encoder is divided into two parts: an encoder and a decoder. Among them, the encoder takes the Mel-frequency cepstral features of the audio as input, and the output is the implicit embedding representation vector obtained by encoding the input. The decoder takes the implicit embedding representation vector as input, and the output is the Mel-frequency cepstral features predicted according to the input implicit embedding representation vector. The encoder and the decoder are jointly trained.

[0083] Reference Figure 2 As shown, before inputting the first output vector into the space transformation model, the method may further include:

[0084] Step S210. Obtain audio data, and preprocess the audio data to obtain audio sample data;

[0085] Step S220. Input the first Mel-frequency cepstral features of the audio sample data into the encoder of the preset variational encoder to obtain a third output vector;

[0086] Step S230. Input the third output vector into the decoder of the preset variational encoder to obtain the second Mel-frequency cepstral features;

[0087] Step S240. Construct a loss function using the first Mel-frequency cepstral coefficient (MFCC) feature and the second MFCC feature, and train the preset variational autoencoder to obtain the variational autoencoder.

[0088] Hereinafter, steps S210 - S240 will be further explained and described. Specifically, audio data is obtained. This audio data can be a voice recording or a music recording. In this exemplary embodiment, the audio data is not specifically limited. After obtaining the audio data, it is preprocessed to obtain audio sample data. The first MFCC feature of the audio sample data is obtained, and the first MFCC feature of the audio sample data is input into the encoder of the preset variational autoencoder to obtain a third output vector corresponding to the first MFCC feature of the audio sample data. This third output vector is an implicit embedding representation vector. After obtaining the third output vector, it is input into the decoder of the preset variational autoencoder, and the second MFCC feature is obtained through this decoder. A loss function is constructed using the first MFCC feature and the second MFCC feature, and the preset variational autoencoder is trained to obtain the variational autoencoder.

[0089] Further, referring to Figure 3 as shown, obtaining audio data and preprocessing the audio data to obtain audio sample data may include:

[0090] Step S310. Convert the audio data to the same sampling rate or the same number of data bits to obtain first audio data;

[0091] Step S320. Augment the first audio data to obtain the audio sample data.

[0092] Hereinafter, steps S310 and S320 will be further explained and described. Specifically, when the audio data is obtained, the preprocessing of the audio data includes converting the audio data to the same sampling rate or the same number of data bits to obtain first audio data. The first audio data can be augmented. This augmentation process includes accelerating or decelerating the first audio data, adding random noise, or introducing distortion, etc. In this exemplary embodiment, the augmentation process is not specifically limited.

[0093] Further, referring to Figure 4 as shown, constructing a loss function using the first MFCC feature and the second MFCC feature, and training the preset variational autoencoder to obtain the variational autoencoder may include:

[0094] Step S410. Determine the multi-scale short-time Fourier transform loss of the first Mel-frequency cepstral coefficient (MFCC) feature and the second MFCC feature as the loss function;

[0095] Step S420. Train the preset variational autoencoder based on the loss function to obtain the variational autoencoder.

[0096] Hereinafter, steps S410 and S420 will be further explained and described. Specifically, determine the multi-scale short-time Fourier transform loss of the first MFCC feature and the second MFCC feature as the loss function, and jointly train the encoder and decoder in the preset variational autoencoder based on this loss function to obtain the variational autoencoder.

[0097] After obtaining the variational autoencoder, the implicit space transformation model can be trained according to the multi-modal large model and this variational autoencoder to obtain the space transformation model. Refer to Figure 5 As shown, after obtaining the variational autoencoder, the method further includes:

[0098] Step S510. Input the first MFCC feature of the audio sample data into the encoder of the variational autoencoder to obtain a third output vector;

[0099] Step S520. Input the first MFCC feature into the multi-modal large model to obtain a fourth output vector;

[0100] Step S530. Train the preset implicit space transformation model through the third output vector and the fourth output vector to obtain the space transformation model.

[0101] Hereinafter, steps S510 - S530 will be further explained and described. Specifically, obtain the first MFCC feature of the audio sample data, input this first MFCC feature into the variational autoencoder to obtain a third output vector, and this third output vector is an implicit embedding representation vector. At the same time, input the first MFCC feature of the audio sample data into the multi-modal large model to obtain a fourth output vector, and this fourth output vector is an implicit embedding representation vector. Train the preset implicit space transformation model through the third output vector and the fourth output vector to obtain the space transformation model. When training the preset implicit space transformation model, the fourth output vector is used as the input, and the third output vector is used as the training target.

[0102] Refer to Figure 6 As shown, training the preset implicit space transformation model through the third output vector and the fourth output vector to obtain the space transformation model may include:

[0103] Step S610. Input the fourth output vector into the preset implicit space transformation model to obtain a fifth output vector;

[0104] Step S620. Construct a loss function according to the fifth output vector and the third output vector, and train the preset implicit space transformation model based on the loss function to obtain the space transformation model.

[0105] Hereinafter, steps S610 and S620 will be further explained and described. Specifically, input the fourth output vector into the preset implicit space transformation model to obtain a fifth output vector, and this fifth output vector is an implicit embedding representation vector; construct a loss function according to the fifth output vector and the third output vector, and the weight of the loss function can be α, and train the preset implicit space transformation model based on the constructed loss function to obtain the space transformation model.

[0106] In the present exemplary embodiment, when there is a small amount of audio-text paired data, the preset implicit space transformation model can also be trained according to the audio-text paired data. Refer to Figure 7 As shown, after obtaining the variational encoder, the method further includes:

[0107] Step S710. Input the third Mel-frequency cepstral coefficient (MFCC) features of the text corresponding to the audio sample data and the first MFCC features of the audio sample data into the multi-modal large model to obtain a sixth output vector;

[0108] Step S720. Input the first MFCC features into the encoder of the variational encoder to obtain a third output vector;

[0109] Step S730. Train the preset implicit space transformation model through the third output vector and the sixth output vector to obtain the space transformation model.

[0110] Hereinafter, steps S710 - S730 will be further explained and described. Specifically, obtain the third MFCC features of the text corresponding to the audio sample data and the first MFCC features of the audio sample data, input the third MFCC features into the multi-modal large model to obtain a sixth output vector, and this sixth output vector is an implicit embedding representation vector. Input the first MFCC features into the decoder of the variational encoder to obtain a third output vector, and train the preset implicit space transformation model according to the sixth output vector and the third output vector to obtain the space transformation model.

[0111] Further, training the preset implicit space transformation model through the third output vector and the sixth output vector to obtain the space transformation model includes:

[0112] Input the sixth output vector into the preset implicit space transformation model to obtain a seventh output vector;

[0113] Construct a loss function based on the seventh output vector and the third output vector, and train the preset implicit space transformation model based on the loss function to obtain the space transformation model.

[0114] Specifically, use the third output vector as the training target of the preset implicit space transformation model, use the sixth output vector as the input of the preset implicit space transformation model, input it into the preset implicit space transformation model to obtain a seventh output vector. This seventh output vector is an implicit embedding representation vector. Construct a loss function based on the seventh output vector and the third output vector. Among them, the weight of the loss function can be β, and β is at least twice that of α or more. Train the preset implicit space transformation model based on the loss function to obtain the space transformation model.

[0115] Furthermore, since the implicit spaces of the variational autoencoder and the multimodal large model are very different, it is easy to have a situation where it is difficult to converge during training. Therefore, the structure and training method of this space transformation model need to meet the following conditions:

[0116] 1) The space transformation model needs at least 2 convolutional modules Di as Figure 8 shown. After experimental verification, the convolutional module of this structure can effectively extract the feature information of the input target space transformation model. Compared with models of other structures, the value of the loss function can be smaller; each convolutional module includes C 1 layer, C 2 layer, C 3 layer, an activation function layer, and a strided convolutional layer; The C i layer includes an activation function layer F, a dilated convolutional layer C E , an activation function F, a 1×1 convolutional layer, and an exclusive OR operation layer;

[0117] 2) The activation function used between convolutional layers is f(x) = max(x, 0) - 0.02 * min(x, 0). This activation function can avoid the disappearance of gradients in the negative region and accelerate the convergence rate;

[0118] 3) When training the preset implicit space transformation model, it is also necessary to add the pre-trained multi-modal large model to jointly train. However, only the parameters of the last 2-3 layers of the multi-modal large model are open for training, and the parameters of the remaining layers are fixed and do not participate in training. This method can significantly accelerate the convergence speed of the preset implicit space transformation model. When updating the parameters through backpropagation, the learning rate of the multi-modal large model should be less than 1 / 10 of the learning rate of the implicit space transformation model. Through experiments, it is known that if the learning rate of the multi-modal large model is too high, it will reduce the generalization performance of the multi-modal large model, thus significantly reducing the diversity of the finally generated audio.

[0119] 4) When training with a small number of text-audio pairs at the same time, when calculating the loss, the weight β of the loss generated by the text should be significantly higher than the loss α generated by the audio, at least more than 2 times that of α. Using this weight setting can significantly improve the matching between the text and the audio. If the β weight is low, it cannot play a role in improving the matching between the text and the audio.

[0120] After training the variational encoder and the space transformation model, it is also necessary to train the vocoder. When training the vocoder, refer to Figure 9 , the method further includes:

[0121] Step S910. Input the first Mel-frequency cepstral feature of the audio sample data into the preset vocoder to obtain a predicted audio signal;

[0122] Step S920. Construct a loss function based on the predicted audio signal and the actual audio signal of the audio sample data, and train the preset vocoder based on the loss function to obtain the vocoder.

[0123] Hereinafter, steps S910 and S920 will be further explained and described. Specifically, the first Mel-frequency cepstral feature of the audio sample data is obtained, input into the preset vocoder to obtain a predicted audio signal, a loss function is constructed based on the output predicted audio signal and the actual audio signal of the input audio sample data, and the preset vocoder is trained based on this loss function to obtain the target vocoder.

[0124] In this exemplary embodiment, after training the variational encoder, the space transformation model, and the vocoder, the first output vector can be input into the space transformation model, and through this space transformation model, a second output vector is obtained.

[0125] In step S130, the second output vector is input into the variational encoder to obtain the Mel-frequency cepstral feature corresponding to the second output vector.

[0126] After obtaining the second output vector, input the second output vector into the decoder of the variational autoencoder to obtain the Mel-frequency cepstral coefficients (MFCCs) corresponding to the second output vector.

[0127] In step S140, input the Mel-frequency cepstral coefficients into the vocoder to obtain the audio corresponding to the text.

[0128] After obtaining the Mel-frequency cepstral coefficients, input the Mel-frequency cepstral coefficients into the vocoder to obtain the audio corresponding to the text.

[0129] The audio generation method provided by the embodiments of the present disclosure has at least the following advantages: on the one hand, by introducing an implicit space transformation model and utilizing the powerful generalization ability of the pre-trained multi-modal large model, it is directly applied to the text-to-audio synthesis process, significantly reducing the dependence on a large amount of text-audio paired data. Even when using only a small amount or no such paired data, efficient model training can be achieved, effectively solving the long-standing problem of obtaining training data in the field of text-to-audio generation and broadening the application scenarios and practicality of the model. On the other hand, after obtaining the Mel-frequency cepstral coefficients corresponding to the text through the multi-modal large model and the space transformation model, input the Mel-frequency cepstral coefficients into the vocoder to obtain the audio corresponding to the text, simplifying the process of converting text to audio and improving the efficiency of text-to-audio generation.

[0130] The embodiments of the present disclosure also provide an audio generation device. Referring to Figure 10 as shown, it may include: a first vector acquisition module 1010, a second vector acquisition module 1020, a feature acquisition module 1030, and an audio generation module 1040. Among them:

[0131] The first vector acquisition module 1010 is configured to input the text into the multi-modal large model to obtain a first output vector.

[0132] The second vector acquisition module 1020 is configured to input the first output vector into the space transformation model to obtain a second output vector.

[0133] The feature acquisition module 1030 is configured to input the second output vector into the variational autoencoder to obtain the Mel-frequency cepstral coefficients corresponding to the second output vector.

[0134] The audio generation module 1040 is configured to input the Mel-frequency cepstral coefficients into the vocoder to obtain the audio corresponding to the text.

[0135] The specific details of each module in the above audio generation device have been described in detail in the corresponding audio generation method, so they will not be elaborated here.

[0136] In an exemplary embodiment of the present disclosure, before inputting the first output vector into the spatial transformation model, the method further includes:

[0137] Obtain audio data, preprocess the audio data to obtain audio sample data;

[0138] Input the first Mel-frequency cepstral coefficients (MFCC) features of the audio sample data into the encoder of a preset variational autoencoder to obtain a third output vector;

[0139] Input the third output vector into the decoder of the preset variational autoencoder to obtain second Mel-frequency cepstral coefficients;

[0140] Construct a loss function based on the first Mel-frequency cepstral coefficients and the second Mel-frequency cepstral coefficients, and train the preset variational autoencoder to obtain the variational autoencoder.

[0141] In an exemplary embodiment of the present disclosure, the obtaining audio data, preprocessing the audio data to obtain audio sample data includes:

[0142] Convert the audio data to the same sampling rate or the same number of data bits to obtain first audio data;

[0143] Augment the first audio data to obtain the audio sample data.

[0144] In an exemplary embodiment of the present disclosure, the constructing a loss function based on the first Mel-frequency cepstral coefficients and the second Mel-frequency cepstral coefficients, and training the preset variational autoencoder to obtain the variational autoencoder includes:

[0145] Determine the multi-scale short-time Fourier transform loss of the first Mel-frequency cepstral coefficients and the second Mel-frequency cepstral coefficients as the loss function;

[0146] Train the preset variational autoencoder based on the loss function to obtain the variational autoencoder.

[0147] In an exemplary embodiment of the present disclosure, after obtaining the variational autoencoder, the method further includes:

[0148] Input the first Mel-frequency cepstral coefficients of the audio sample data into the encoder of the variational autoencoder to obtain a third output vector;

[0149] Input the first Mel-frequency cepstral coefficients into the multi-modal large model to obtain a fourth output vector;

[0150] Train the preset implicit space transformation model with the third output vector and the fourth output vector to obtain the space transformation model.

[0151] In an exemplary embodiment of the present disclosure, the training of the preset implicit space transformation model with the third output vector and the fourth output vector to obtain the space transformation model includes:

[0152] Input the fourth output vector into the preset implicit space transformation model to obtain a fifth output vector;

[0153] Construct a loss function based on the fifth output vector and the third output vector, and train the preset implicit space transformation model based on the loss function to obtain the space transformation model.

[0154] In an exemplary embodiment of the present disclosure, after obtaining the variational encoder, the method further includes:

[0155] Input the third Mel-frequency cepstral features of the text corresponding to the audio sample data and the first Mel-frequency cepstral features of the audio sample data into the multimodal large model to obtain a sixth output vector;

[0156] Input the first Mel-frequency cepstral features of the audio sample data into the encoder of the variational encoder to obtain a third output vector;

[0157] Train the preset implicit space transformation model with the third output vector and the sixth output vector to obtain the space transformation model.

[0158] In an exemplary embodiment of the present disclosure, the training of the preset implicit space transformation model with the third output vector and the sixth output vector to obtain the space transformation model includes:

[0159] Input the sixth output vector into the preset implicit space transformation model to obtain a seventh output vector;

[0160] Construct a loss function based on the seventh output vector and the third output vector, and train the preset implicit space transformation model based on the loss function to obtain the space transformation model.

[0161] In an exemplary embodiment of the present disclosure, the method further includes:

[0162] Input the first Mel-frequency cepstral features of the audio sample data into a preset vocoder to obtain a predicted audio signal;

[0163] Construct a loss function based on the predicted audio signal and the actual audio signal of the audio sample data, and train the preset vocoder based on the loss function to obtain the vocoder.

[0164] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0165] In addition, although the steps of the methods in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.

[0166] In an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided.

[0167] Those skilled in the art to which the present disclosure pertains can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to herein as "circuitry", "module", or "system".

[0168] The following refers to Figure 11 to describe the electronic device 1100 according to this embodiment of the present disclosure. Figure 11 The shown electronic device 1100 is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0169] As Figure 11 shown, the electronic device 1100 is presented in the form of a general-purpose computing device. The components of the electronic device 1100 may include, but are not limited to: the at least one processing unit 1110 described above, the at least one storage unit 1120 described above, a bus 1130 connecting different system components (including the storage unit 1120 and the processing unit 1110), and a display unit 1140.

[0170] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 1110, so that the processing unit 1110 executes the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section above of this specification. For example, the processing unit 1110 can execute steps such as Figure 1 shown in S110: Input the text into the multi-modal large model to obtain a first output vector; S120: Input the first output vector into the spatial transformation model to obtain a second output vector; S130: Input the second output vector into the variational encoder to obtain the Mel-frequency cepstral coefficients corresponding to the second output vector; S140: Input the Mel-frequency cepstral coefficients into the vocoder to obtain the audio corresponding to the text.

[0171] The storage unit 1120 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 11201 and / or a cache storage unit 11202, and may further include a read-only storage unit (ROM) 11203.

[0172] The storage unit 1120 may further include a program / utility 11204 having a set (at least one) of program modules 11205. Such program modules 11205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0173] The bus 1130 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0174] The electronic device 1100 can also communicate with one or more external devices 1200 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 1100, and / or communicate with any device that enables the electronic device 1100 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 1150. Moreover, the electronic device 1100 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 1160. As shown in the figure, the network adapter 1160 communicates with other modules of the electronic device 1100 through the bus 1130. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 1100, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0175] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by a manner of software combined with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0176] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above method of this specification is stored. In some possible implementation manners, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.

[0177] The program product for implementing the above method according to the embodiments of the present disclosure can adopt a portable compact disc read-only memory (CD-ROM) and include program code, and can run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.

[0178] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0179] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0180] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0181] The program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0182] In addition, the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, and are not for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes may be executed, for example, synchronously or asynchronously in multiple modules.

[0183] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

Claims

1. An audio generation method, characterized in that: include: Input the text into the multimodal large model to obtain the first output vector; Inputting the first output vector into a spatial transformation model to obtain a second output vector; Inputting the second output vector into a variational encoder to obtain a Mel-frequency cepstrum feature corresponding to the second output vector; The Mel-frequency cepstrum feature is input into a vocoder to obtain audio corresponding to the text.

2. The audio generation method according to claim 1, characterized in that: Before inputting the first output vector into the spatial transformation model, the method further comprises: Acquire audio data, and preprocess the audio data to obtain audio sample data; Inputting the first Mel-frequency cepstrum feature of the audio sample data into an encoder of a preset variational encoder to obtain a third output vector; Inputting the third output vector into a decoder of the preset variational encoder to obtain a second Mel-frequency cepstrum feature; A loss function is constructed using the first Mel-frequency cepstrum feature and the second Mel-frequency cepstrum feature, and the preset variational encoder is trained to obtain the variational encoder.

3. The audio generation method according to claim 2, characterized in that: The step of acquiring audio data and preprocessing the audio data to obtain audio sample data includes: Convert the audio data into the same sampling rate or the same data bit number to obtain first audio data; The first audio data is augmented to obtain the audio sample data.

4. The audio generation method according to claim 2, characterized in that: The step of constructing a loss function by using the first Mel-frequency cepstrum feature and the second Mel-frequency cepstrum feature, and training the preset variational encoder to obtain the variational encoder includes: Determine the multi-scale short-time Fourier transform loss of the first Mel-frequency cepstrum feature and the second Mel-frequency cepstrum feature as a loss function; The preset variational encoder is trained based on the loss function to obtain the variational encoder.

5. The audio generation method according to claim 2, characterized in that: After obtaining the variational encoder, the method further includes: Inputting the first Mel-frequency cepstrum feature of the audio sample data into an encoder of the variational encoder to obtain a third output vector; Inputting the first Mel-frequency cepstrum feature into the multimodal large model to obtain a fourth output vector; The preset implicit spatial transformation model is trained by using the third output vector and the fourth output vector to obtain the spatial transformation model.

6. The audio generation method according to claim 5, characterized in that: The step of training a preset implicit spatial transformation model by using the third output vector and the fourth output vector to obtain the spatial transformation model includes: Inputting the fourth output vector into the preset implicit spatial transformation model to obtain a fifth output vector; A loss function is constructed according to the fifth output vector and the third output vector, and the preset implicit spatial transformation model is trained based on the loss function to obtain the spatial transformation model.

7. The audio generation method according to claim 2, characterized in that: After obtaining the variational encoder, the method further includes: Inputting a third Mel-frequency cepstrum feature of the text corresponding to the audio sample data and a first Mel-frequency cepstrum feature of the audio sample data into the multimodal large model to obtain a sixth output vector; Inputting the first Mel-frequency cepstrum feature into an encoder of the variational encoder to obtain a third output vector; The preset implicit spatial transformation model is trained by using the third output vector and the sixth output vector to obtain the spatial transformation model.

8. The audio generation method according to claim 7, characterized in that: The step of training the preset implicit spatial transformation model by using the third output vector and the sixth output vector to obtain the spatial transformation model includes: Inputting the sixth output vector into the preset implicit spatial transformation model to obtain a seventh output vector; A loss function is constructed according to the seventh output vector and the third output vector, and the preset implicit spatial transformation model is trained based on the loss function to obtain the spatial transformation model.

9. The audio generation method according to claim 2, characterized in that: The method further comprises: Inputting the first Mel-frequency cepstrum feature of the audio sample data into a preset vocoder to obtain a predicted audio signal; A loss function is constructed according to the predicted audio signal and the actual audio signal of the audio sample data, and the preset vocoder is trained based on the loss function to obtain the vocoder.

10. An audio generating device, characterized in that: include: A first vector acquisition module, used for inputting text into the multimodal large model to obtain a first output vector; A second vector acquisition module, used for inputting the first output vector into a space transformation model to obtain a second output vector; A feature acquisition module, used for inputting the second output vector into a variational encoder to obtain a Mel-frequency cepstrum feature corresponding to the second output vector; The audio generation module is used to input the Mel-frequency cepstrum feature into a vocoder to obtain audio corresponding to the text.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

12. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 9 by executing the executable instructions.