A multi-modal music short film generation method based on audio, image and text

By building the TS model and MV diffusion model, the problems of data consistency and synchronization of three modes and above in multimodal generation are solved, and more comprehensive and comprehensive music short film generation capabilities are achieved, and creativity and understanding capabilities are improved.

CN117412144BActive Publication Date: 2025-05-27JINAN FALAI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311456758.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2025-05-27
Estimated Expiration
2043-11-03

AI Technical Summary

Technical Problem

The prior art mainly focuses on the dual-modal generation task in the field of multimodal generation, and it is difficult to ensure consistency and synchronization between data in three modes and above, and the creativity and comprehensive understanding capabilities need to be improved.

Method used

A multimodal music short film generation method based on audio, image and text is adopted. By building a TS model and MV diffusion model, the alignment and synchronization between audio, image and text is achieved to ensure the consistency and synchronization of multimodal data.

Benefits of technology

A more comprehensive and comprehensive multimodal music short film generation capability is achieved, ensuring consistency and synchronization between various modes, and improving the comprehensive understanding and creativity of music short films.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117412144B_ABST
    Figure CN117412144B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of music short film generation applications, and particularly relates to a multi-modal music short film generation method based on audio, images, and text. The present invention provides a multi-modal music short film generation method based on audio, images, and text, and constructs a text encoding structure, an audio encoding structure, a TS model, an MV diffusion model training architecture, and an MV diffusion model testing architecture. Through the established multi-modal deep learning model, a more comprehensive and integrated multi-modal music short film generation ability is achieved, ensuring the consistency and synchronization between various modalities, and being able to simultaneously process multiple related tasks such as audio, images, and text, thereby improving the comprehensive understanding and creation ability of music short films.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of music short video generation applications, and particularly relates to a multi-modal music short video generation method based on audio, images, and text. Background Art

[0002] In today's society, significant progress has been made in various fields such as audio processing, image processing, and natural language processing, mainly due to the improvement of computing power and the widespread application of deep learning technology. Currently, deep learning technology, especially the Transformer network module, has become the mainstream method for multi-modal generation, bringing breakthrough performance improvements to various audio, image, and text understanding tasks. These tasks include audio feature extraction, image generation, lyric analysis, etc., which integrate the processing and analysis of data of types such as text, images, and audio. In the multi-modal background art, data of different modalities can influence each other, providing more comprehensive and accurate information, enabling the system to better simulate the diversity of human perception and cognition. Although deep learning technology has achieved great success in the field of multi-modal generation, current research mainly focuses on solving dual-modal generation tasks, and there are still limitations in the research on three-modal and above, unable to ensure the consistency and synchronization between modalities, and both the creative ability and comprehensive understanding ability need to be improved. Summary of the Invention

[0003] In view of the technical problems existing in the generation of multi-modal music short videos, the present invention proposes a multi-modal music short video generation method based on audio, images, and text, which is reasonable in design, simple in method, strong in theory, and can achieve ensuring the consistency and synchronization between modalities and being more comprehensive and integrated.

[0004] To achieve the above object, the technical solution adopted by the present invention is: a multi-modal music short video generation method based on audio, images, and text, comprising the following steps:

[0005] S1. Data acquisition, S2. Build a TS model, S3. Train and test the TS model, S4. Build an MV diffusion model, S5. Train and test the MV diffusion model, and verify the alignment accuracy of the three modalities of image audio, text, and image, as well as the richness and diversity of music short videos through simulation experiments.

[0006] Preferably, the specific operation method of step S1 is:

[0007] S1-1. Obtain open-source music short video data, segment each sentence of music text and its accompanying music and pictures according to time to obtain text, audio, and image data, and at the same time ensure the alignment relationship among text, audio, and image;

[0008] S1-2. Select the text and audio data in S1-1 to construct a dataset , and divide the dataset in a ratio of 8:2 for training and testing the TS model. Select the audio, text, and image data in S1-1 to construct a dataset , and divide the dataset D in a ratio of 8:2 for training and testing the MVdiffusion model.

[0009] Preferably, the specific operation method of step S2 is as follows:

[0010] S2-1. Build the text encoding part, perform Token tokenization on the text description, and then perform Token encoding and position encoding. Use the CLIP text encoder as the main body, and freeze the pre-trained weights in the CLIP text encoder. Finally, obtain the hidden space feature vector V of the text t , where t represents the vector dimension;

[0011] S2-2. Build the audio encoding part, perform Patch audio slicing on the audio, and then perform Patch encoding and position encoding. Use 4 Transformer Encodes to construct the core of the audio encoder. Finally, obtain the hidden space feature vector V of the audio s , where s represents the vector dimension;

[0012] S2-3. At the same time, ensure that the hidden space feature vectors V s和 V t output by the audio and text encoders have the same dimension, that is ;

[0013] S2-4. Normalize the hidden space feature vectors output by the audio and text encoders respectively, and define them as: where is the normalized feature vector of the audio, is the normalized feature vector of the text, is the transpose of V s , is the transpose of V t . Take the similarity after logarithmic transformation of the two as the loss function of the TS model, and the formula is: where is the normalized feature vector of the audio, is the normalized feature vector of the text, and obtain the text and audio hidden space mapping model.

[0014] Preferably, the specific operation method of step S3 is as follows:

[0015] S3-1. Feed the training set in the data into the model. The learning rate change strategy for the training method adopts warmup, and train for 100 epochs. Finally, obtain the mapping relationship between audio and text, and get the trained audio-text latent space mapping model;

[0016] S3-2. Feed the test set in the data into the model, verify the accuracy of the mapping relationship between audio and text, save the audio encoding weights that meet the requirements, and obtain a general audio encoder.

[0017] Preferably, the specific operation method of step S4 is:

[0018] S4-1. Construct an MV diffusion training stage model, perform vectorization processing on audio, image, and text data. The audio passes through the audio encoder to obtain an audio vector, the image passes through the VAE encoder to obtain an image latent space vector, and the text passes through the CLIP text encoder to obtain a text vector;

[0019] S4-2. Add noise to the image latent space vector described in S4-1. The noise needs to conform to a Gaussian distribution and is used to construct the data required for the diffusion model. Then, send the three-modal vectors into Unet for training to predict the noise distribution;

[0020] S4-3. Compare the predicted noise with the real noise, and use weighted KL divergence and MSE as loss functions to measure the difference between the generated sample distribution and the real data distribution, and measure the difference between the generated sample and the real data. The loss function formula is: where, is the real sample data, is the generated sample data, N is the total amount of sample data, i is the sample label, , is the dynamic weight, ;

[0021] S4-4. Construct an MV diffusion test stage model. The difference from the training stage model is that the image latent space vector of Unet is replaced by noise. Unet can be iterated multiple times, and the output latent space vector needs to be restored to an image through the VAE decoder. Finally, obtain the MV diffusion test stage model.

[0022] Preferably, the specific operation method of step 5 is:

[0023] S5-1. During the training phase, use the MV diffusion training phase model. Freeze the VAE encoder, image encoder, and audio encoder, and only train the Unet. The loss function is the one obtained in S4-3. Adopt the warmup strategy for the learning rate and train for 100 rounds.

[0024] S5-2. During the testing phase, use the MV diffusion testing phase model. Freeze all weights for inference, and verify the alignment accuracy of the image-audio, text-image trinity and the richness and diversity of music shorts through simulation experiments.

[0025] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0026] The present invention provides a method for generating music shorts based on multi-modalities of audio, image, and text, and constructs a text encoding structure, an audio encoding structure, a TS model, an MV diffusion model training architecture, and an MV diffusion model testing architecture. Through the constructed multi-modal deep learning model, it realizes a more comprehensive and integrated multi-modal music short generation ability, ensures the consistency and synchronization between various modalities, and can simultaneously process multiple related tasks such as audio, image, and text, thereby improving the comprehensive understanding and creation ability of music shorts. Description of the Drawings

[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0028] Figure 1 Schematic diagram of the text encoding structure provided by the embodiment of the present invention;

[0029] Figure 2 Schematic diagram of the audio encoding structure provided by the embodiment of the present invention;

[0030] Figure 3 Schematic diagram of the TS model provided by the embodiment of the present invention;

[0031] Figure 4 Schematic diagram of the MV diffusion model training architecture provided by the embodiment of the present invention;

[0032] Figure 5 Schematic diagram of the MV diffusion model testing architecture provided by the embodiment of the present invention;

[0033] Figure 6Schematic diagram of the MV diffusion model effect provided by the embodiments of the present invention. Detailed implementation manners

[0034] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described below with reference to the accompanying drawings and embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.

[0035] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the present invention is not limited by the specific embodiments disclosed in the following specification.

[0036] Embodiment. Although deep learning technology has achieved great success in the field of multimodal generation, current research mainly focuses on solving bimodal generation tasks. There are still limitations in the research on trimodal and above, and it is impossible to ensure the consistency and synchronization between each modality, and both the creative ability and the comprehensive understanding ability need to be improved. For this reason, the present invention proposes a multimodal music short film generation method based on audio, image and text. First, data acquisition. Obtain a music short film data set, and according to the time relationship, divide the data set into an image data set, a text data set, and an audio data set in the form of three modality data pairs. Specifically, obtain open-source music short film data, divide each sentence of music text and its accompanying music and pictures according to time to obtain text, audio and image data, and at the same time ensure the alignment relationship among text, audio and image; select the text and audio data in S1-1 to construct a data set , and divide the data set in a ratio of 8:2 for training and testing the TS model. Select the audio, text, and image data in S1-1 to construct a data set , and divide the data set D in a ratio of 8:2 for training and testing the MV diffusion model.

[0037] Next, as Figure 3 shown, build a TS model. The text encoding part uses a CLIP text encoder to obtain a text and audio latent space mapping model. Specifically, as Figure 1 , Figure 2 shown, build the text encoding part, perform Token word segmentation on the text description, then perform Token encoding and position encoding on it, use the CLIP text encoder as the main body, and at the same time freeze the pre-trained weights in the CLIP text encoder, and finally obtain the latent space feature vector V of the text t, where \(t\) represents the vector dimension; build the audio encoding part, perform Patch slicing on the audio, then perform Patch encoding and positional encoding on it, and use 4 Transformer Encodes to construct the core of the audio encoder, finally obtaining the latent space feature vector \(V\) of the audio s , where \(s\) represents the vector dimension; at the same time, ensure that the latent space feature vectors \(V\) output by the audio and text encoders s和 V t have the same dimension, that is ; S2-4. Unify the latent space feature vectors output by the audio and text encoders respectively, and define them as: where is the unified feature vector of the audio, is the unified feature vector of the text, is the transpose of \(V\) s , is the transpose of \(V\) t , and use the similarity after logarithmic transformation of the two as the loss function of the TS model. The formula is: where is the unified feature vector of the audio, is the unified feature vector of the text.

[0038] Then, train and test the TS model, verify the accuracy of the text and audio latent space mapping, and obtain the audio encoder. Specifically, send the training set in the data into the model. The learning rate change strategy of the training method adopts warmup, train for 100 epochs, and finally obtain the mapping relationship between the audio and the text, obtaining the trained audio-text latent space mapping model; send the test set in the data into the model to verify the accuracy of the mapping relationship between the audio and the text, save the audio encoding weights that meet the requirements, and obtain a general audio encoder.

[0039] Subsequently, as Figure 4 , Figure 5Build the MV diffusion model as shown. For the audio encoding part, use the audio encoder obtained by S3. For the image encoding part, use the VAE encoder. For the text encoding part, use the CLIP text encoder. For the decoder, use the diffusion model. Finally, obtain the audio, text, and image multi-modal diffusion model. Specifically, build the MV diffusion training stage model, and vectorize the audio, image, and text data. The audio passes through the audio encoder to obtain the audio vector, the image passes through the VAE encoder to obtain the image latent space vector, and the text passes through the CLIP text encoder to obtain the text vector. Add noise to the image latent space vector described in S4-1. The noise needs to conform to the Gaussian distribution and is used to construct the data required for the diffusion model. Then, send the three-modal vectors into Unet for training to predict the noise distribution. Compare the predicted noise with the real noise, and use the weighted KL divergence and MSE as the loss functions to measure the difference between the generated sample distribution and the real data distribution, and to measure the difference between the generated samples and the real data. The loss function formula is: where is the real sample data, is the generated sample data, N is the total amount of sample data, i is the sample label, , is the dynamic weight, ; Build the MV diffusion test stage model. The difference from the training stage model is that the image latent space vector of Unet is replaced by noise. Unet can be iterated multiple times, and the output latent space vector needs to be restored to an image through the VAE decoder. Finally, obtain the MV diffusion test stage model.

[0040] Finally, as Figure 6 shown, train and test the MV diffusion model, and use simulation experiments to verify the alignment accuracy of the three modalities of image-audio, text, and image, as well as the richness and diversity of music shorts. Specifically, in the training stage, use the MV diffusion training stage model, freeze the pre-trained weights in the VAE encoder, image encoder, and audio encoder in the model, and only train the Unet model part. Use the weighted KL divergence and MSE in S4-3 as the loss functions, send the training set in data D into the model, and adopt the warmup strategy for the learning rate change strategy during training, and train for 100 epochs. In the test stage, use the MV diffusion test stage model, freeze all the weights in the model for inference, and use simulation experiments to verify the alignment accuracy of the three modalities of image-audio, text, and image, as well as the richness and diversity of music shorts.

[0041] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention in any other form. Any person skilled in the relevant art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A multi-modal music short film generation method based on audio, image and text, characterized in that, it includes the following steps: S1. Data acquisition, S2. Build a TS model, S3. Train and test the TS model, S4. Build an MV diffusion model, S5. Train and test the MV diffusion model, The specific operation method of step S1 is: S1-1. Obtain open-source music short film data, and segment each sentence of music text and its accompanying music and pictures according to time to obtain text, audio and image data, while ensuring the alignment relationship among text, audio and image; S1-2. Select the text and audio data in S1-1 to construct a dataset , and divide the dataset in an 8:2 ratio for training and testing the TS model. Select the audio, text, and image data in S1-1 to construct a dataset , and divide the dataset D in an 8:2 ratio for training and testing the MVdiffusion model; The specific operation method of step S2 is: S2-1. Build the text encoding part, perform Token tokenization on the text description, then perform Token encoding and position encoding on it. Use the CLIP text encoder as the main body, and at the same time freeze the pre-trained weights in the CLIP text encoder. Finally, obtain the hidden space feature vector V of the text t, t represents the vector dimension; S2-2. Build the audio encoding part, perform Patch audio slicing on the audio, and then perform Patch encoding and positional encoding on it. Use 4 Transformer Encodes to construct the core of the audio encoder, and finally obtain the latent space feature vector V of the audio s, s represents the vector dimension; S2-3. Meanwhile, ensure that the latent space feature vectors V output by the audio and text encoders s和 V t have the same dimension, that is ; S2-4. Uniformize the latent space feature vectors output by the audio and text encoders respectively, and define them as Among them, is the unified feature vector of the audio, is the unified feature vector of the text, is the transpose of V s , is the transpose of V t . The similarity after logarithmic transformation of the two is used as the loss function of the TS model, and the formula is: Among them, is the unified feature vector of the audio, is the unified feature vector of the text, and a text and audio latent space mapping model is obtained; The specific operation method of step S3 is: S3-1. Send the training set in the data into the model. The learning rate change strategy of the training method adopts warmup. Train for 100 epochs to finally obtain the mapping relationship between audio and text and get the trained audio-text latent space mapping model. S3-2. Send the test set in the data into the model to verify the accuracy of the mapping relationship between audio and text, save the audio encoding weights that meet the requirements, and obtain a general audio encoder; The specific operation method of step S4 is: S4-1. Build an MV diffusion training stage model, vectorize the audio, image and text data. The audio is obtained as an audio vector through an audio encoder, the image is obtained as an image latent space vector through a VAE encoder, and the text is obtained as a text vector through a CLIP text encoder; S4-2. Add noise to the image latent space vector described in S4-1. The noise needs to conform to a Gaussian distribution and is used to construct the data required for the diffusion model. Then, the three-modal vectors are sent to Unet for training to predict the noise distribution; S4-3. Compare the predicted noise with the real noise, and use weighted KL divergence and MSE as loss functions to measure the difference between the generated sample distribution and the real data distribution, and measure the difference between the generated sample and the real data. The loss function formula is: Among them, is the real sample data, is the generated sample data, N is the total amount of sample data, i is the sample label, , is the dynamic weight, ; S4-4. Build an MV diffusion test stage model. The difference from the training stage model is that the image latent space vector of Unet is replaced by noise. Unet can be iterated multiple times, and the output latent space vector needs to be restored to an image through a VAE decoder, and finally an MV diffusion test stage model is obtained.

2. The multi-modal music short film generation method based on audio, image and text according to claim 1, characterized in that, The specific operation method of step S5 is: S5-1. In the training stage, use the MV diffusion training stage model, freeze the VAE encoder, image encoder, and audio encoder, and only train Unet. The loss function is the loss function obtained in S4-3, and the learning rate adopts a warmup strategy, and train for 100 rounds; S5-2. In the test stage, use the MV diffusion test stage model.

Citation Information

Patent Citations

  • Cross-modal understanding and generating method and device based on multi-modal pre-training model

    CN113591902A

  • Multimodal fusion audio generation method and device based on diffusion model

    CN116884391A