Text generation music algorithm based on melody guidance
By encoding and aligning the data of music waveforms, melody and text descriptions, combining the melody vector database and diffusion process, high-quality music that conforms to text description and melody guidance is solved, and the problem that music generation in the prior art is difficult to preserve semantic integrity and inner harmony.
Patent Information
- Application Number
- CN202510117783.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing music generation technology is difficult to preserve the semantic integrity and inner harmony of the generated music at the same time, and cannot meet the high-quality needs of users.
By obtaining the data of music waveforms, melody and text descriptions, encode them separately, and aligning the encoded audio, melody and text representations in a unified vector space. A melody vector database is constructed based on melody representation, and the target melody vector is queryed using text representation, and a potential music representation that conforms to text description and melody guidance is generated in combination with the diffusion process, which is finally converted into playable music through a variational autoencoder and vocoder.
The semantic integrity and inner harmony of generated music are realized, and high-quality music that conforms to text description and melody guidance can be generated to meet the high-quality needs of users.
Smart Images

Figure CN119943011A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of music generation, and in particular to a text-generated music algorithm based on melody guidance. Background Art
[0002] Music generation has attracted much attention in the field of artificial intelligence and has been widely used in scenarios such as short video background music production. Currently, music generation mainly generates music by generating musical scores or generating audio. The first method regards music as a structured language and generates it in a sequence-to-sequence framework; the second method generates playable music by directly generating audio files that capture the original sound information. Since these methods do not align the melody with the text and waveform at the same time, they often fail to preserve the semantic integrity and intrinsic harmony of the generated music.
[0003] In the prior art, in order to preserve the semantic integrity and inherent harmony of generated music, some studies have proposed several music generation methods based on deep neural networks, including models based on recurrent neural networks (RNNs), models based on variational autoencoders (VAEs), models based on generative adversarial networks (GANs), and models based on Transformers, which are more suitable for generating musical scores rather than playable music. Therefore, the problems existing in the above models cannot meet the high-quality needs of users.
[0004] In view of this, a melody-guided text-generated music algorithm is needed. Summary of the invention
[0005] The embodiment of the present application provides a melody-guided text-based music generation algorithm, which is used to address the problem that the music generation model cannot meet the high-quality needs of users.
[0006] The first aspect of the embodiment of the present application provides a melody-guided text-based music generation algorithm, including:
[0007] Obtain data in three modalities: music waveform, melody, and text description through public datasets, and encode these three modal data respectively;
[0008] Align the encoded audio representation, melody representation, and text representation in a unified vector space;
[0009] Building a melody vector database based on the melody representation, and retrieving a target melody vector representation in the melody vector database using the text representation as a query condition;
[0010] The target melody vector representation and the text representation are used as fusion conditions to guide the diffusion process, generating a latent music representation that conforms to the text description and melody guidance;
[0011] Using a decoder in a variational autoencoder to preliminarily decode the latent music representation into a target mel-spectrogram;
[0012] The target mel-spectrogram is converted into playable music through a vocoder.
[0013] Furthermore, the three modal data of music waveform, melody and text description are obtained through the public data set, and the three modal data are encoded respectively, including:
[0014] After converting each music waveform into a Mel-spectrogram, HTS-AT is used to encode the Mel-spectrogram into an audio representation, the parameters of HTS-AT are fixed, and a multi-layer perceptron is trained after the HTS-AT module;
[0015] Use RoBERTa to encode the text description to get the text representation, fix the parameters of RoBERTa, and train a multi-layer perceptron after the RoBERTa module;
[0016] A randomly initialized small multilayer perceptron is used as a melody encoder to process the pitch and duration information of the melody, and a pooling strategy is used to convert melody marker sequences of different lengths into a uniform shape, which is then input into another multilayer perceptron to generate an updated melody representation.
[0017] Furthermore, the step of aligning the encoded audio representation, melody representation, and text representation in a unified vector space includes:
[0018] The contrast loss between the audio representation and the melody representation, the contrast loss between the audio representation and the text representation, and the contrast loss between the melody representation and the text representation are calculated respectively to obtain a contrast loss function;
[0019] Contrastive learning is performed within each batch, and the audio representation, melody representation, and text representation are aligned into a unified vector space through a batch gradient descent algorithm with the goal of minimizing the contrastive loss function.
[0020] Furthermore, the contrast loss between the audio representation and the melody representation, the contrast loss between the audio representation and the text representation, and the contrast loss between the melody representation and the text representation are calculated respectively to obtain a contrast loss function, including:
[0021]
[0022] Where: L is the total contrast loss, L am is the contrast loss between audio representation and melody representation, L ma is the contrast loss between melody representation and audio representation, L at is the contrast loss between audio representation and text representation, Lta is the contrast loss between text representation and audio representation, L mt is the contrast loss between melody representation and text representation, L tm is the contrast loss between text representation and melody representation.
[0023] Furthermore, the contrast loss L between the audio representation and the melody representation is am The expressions include:
[0024]
[0025] Where: N is the batch size, is the audio representation vector obtained after processing the i-th audio in the batch data, is the melody representation vector obtained after processing the i-th melody in the batch data, is the melody representation vector obtained after processing the j-th melody in the batch data, and τ is a learnable temperature parameter.
[0026] Furthermore, the step of constructing a melody vector database based on the audio representation and retrieving a target melody vector representation in the melody vector database using the text representation as a query condition includes:
[0027] The target melody vector representation is determined by using a hierarchical navigable small-world graph method. The expression of the target melody vector representation is:
[0028]
[0029] in: is the target melody vector representation, is the melody vector database, r is the element in the melody vector database, x t Represents the text.
[0030] Furthermore, the target melody vector representation and the text representation are used as fusion conditions to guide the diffusion process to generate a potential music representation that conforms to the text description and melody guidance, including:
[0031] The expression for fusing the target melody vector representation with the text representation is as follows:
[0032]
[0033] Among them: c is the fusion condition, and b∈R d′ is a learnable parameter, Represents the text x t and the target melody vector representation Concatenate them to get a vector of dimension 2d′, where d′ is the dimension.
[0034] Furthermore, the target melody vector representation and the text representation are used as fusion conditions to guide the diffusion process to generate a potential music representation that conforms to the text description and melody guidance, including:
[0035] The diffusion process includes a forward process and a reverse process. The forward process is a process of gradually adding noise to the original sample; the reverse process starts from the sample obtained from the forward process that obeys the normal distribution of the representation, and by gradually predicting and removing noise, the sample that conforms to the original data distribution is restored to generate a potential music representation that meets the requirements.
[0036] Furthermore, the expression of the forward process is as follows:
[0037]
[0038] Where: x n and x n―1 They represent the samples in the nth and n-1th steps of the diffusion process, respectively, and x 0 is the original sample, and Determines the distribution generated during the denoising process, For the known sample x at step n n , the original sample x 0 And the mean of the n-1th step samples under the fusion condition c, is the variance related parameter;
[0039] The calculation formula for noise prediction is as follows:
[0040]
[0041] Where: ∈ θ (x n ,n,c) is the predicted value of the noise at step n under condition c, ∈ θ (x n ,n) is the predicted value of the noise at the nth step without considering condition c, and w is the bootstrap parameter without classifier.
[0042] Furthermore, the encoder in the variational autodecoder is used to initially decode the potential music representation into a target mel-spectrogram, including:
[0043] The neural network structure of the encoder in the variational autodecoder is used to transform the vector information in the low-dimensional latent space into the target Mel-spectrogram.
[0044] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0045] The present invention obtains three modal data of music waveform, melody and text description from a public data set and encodes them separately, so that different types of data can be presented in a digital form that can be processed by the model; the encoded audio, melody and text representations are aligned in a unified vector space, so that information of different modalities can be compared, associated and fused in the same space, which greatly enhances the comprehensive processing ability of the model for multimodal information; a melody vector database is constructed based on melody representation, and retrieval is performed using text representation as a query condition, so that target melody vectors related to the semantics of text description can be found from massive melodies, and accurate screening of melody by text is achieved, so that the generated music can fit the intention conveyed by the text in terms of melody; the target melody is Vector representation and text representation are used as fusion conditions to perform a diffusion process to generate a potential music representation that conforms to the text description and melody guidance. The semantic information of the text and the characteristic information of the melody are combined here, so that the generated music representation not only conforms to the theme, emotion and other requirements of the text description at the macro level, but is also guided by the retrieved melody in the melody structure, laying the foundation for generating music with specific style and semantics; the variational autoencoder is used to optimize the music representation, which not only retains the basic acoustic characteristics of the music, but also conforms to the conditions set by the text and melody guidance. The vocoder can restore the frequency and amplitude information in the Mel spectrum graph to the time domain signal, generating high-quality playable music that conforms to the text description and has a beautiful melody. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 A schematic diagram of an embodiment of a melody-guided text-to-music algorithm in the present invention;
[0047] Figure 2 A schematic diagram of the structure of a melody-guided text-to-music algorithm in the present invention;
[0048] Figure 3 A schematic diagram of the precise alignment of the experiment in the present invention;
[0049] Figure 4 A schematic diagram of the experimental ablation study in the present invention;
[0050] Figure 5 A schematic diagram of the experimental melody guide in the present invention;
[0051] Figure 6 A schematic diagram of the experiment in the present invention for accurately understanding semantic information;
[0052] Figure 7 It is a schematic diagram of the experimental parameter analysis in the present invention;
[0053] Figure 8 A schematic diagram of the attitude towards artificial intelligence music in the experiment of the present invention;
[0054] Fig. 9 A schematic diagram of the recognition results of the experimental artificial intelligence music in the present invention. DETAILED DESCRIPTION
[0055] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "corresponding to" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0056] Embodiment 1
[0057] The implementation method in this embodiment can be implemented in the system, can be implemented in the server, and can also be implemented in the terminal, and the specific details are not clearly limited. The following will introduce the text-based music generation algorithm based on melody guidance in this application from the perspective of system implementation. Figure 1-Figure 2 , the method provided in the embodiment of the present application comprises the following steps:
[0058] S11. Obtain data of three modes, namely, music waveform, melody and text description, through a public data set, and encode the three modal data respectively;
[0059] In this embodiment, step S11 further includes the following:
[0060] 1. After converting each music waveform into a Mel-spectrogram, use HTS-AT to encode the Mel-spectrogram into an audio representation, fix the parameters of HTS-AT, and train a multi-layer perceptron after the HTS-AT module;
[0061] The original continuous music waveform signal is divided into multiple short time periods according to the human ear's perception of sound frequency. For each short time period, the short-time Fourier transform is used to convert it from the time domain to the frequency domain to obtain a spectrogram. After that, the Mel filter bank is applied to filter the spectrogram, converting the spectrum from the linear frequency scale to the Mel frequency scale, highlighting the frequency area that the human ear is sensitive to and compressing the insensitive area to obtain a Mel spectrum.
[0062] HTS-AT (Hierarchical Token-Semantic Audio Transformer) is used to extract key features from the mel-spectrogram and convert it into an audio representation. In this step, the parameters of HTS-AT are fixed to avoid changing the learned feature extraction method in subsequent training. Then, a multi-layer perceptron (MLP) is connected after the HTS-AT module. MLP consists of multiple neuron layers, and the layers are fully connected. By training this MLP, the audio representation output by HTS-AT can be further processed and feature transformed, such as weighted combination of features, dimensionality adjustment, etc., to better adapt to the subsequent model task requirements.
[0063] 2. Use RoBERTa to encode the text description to obtain the text representation, fix the parameters of RoBERTa, and train a multi-layer perceptron after the RoBERTa module;
[0064] RoBERTa (A Robustly Optimized BERT Pretraining Approach) is used to pre-train on large-scale text corpora based on the Transformer architecture, which can understand the semantic information in the text. The text description is input into the RoBERTa model, and the model will output a text representation that captures the semantic features of the text. Similar to processing audio representations, the parameters of RoBERTa are fixed to keep its learned semantic understanding ability unchanged. A multi-layer perceptron is trained after the RoBERTa module. This MLP processes the text representation output by RoBERTa, and performs feature transformation and information integration on the text representation by adjusting the weights and biases of neurons, such as enhancing the expression of certain semantic features, reducing dimensions, etc., in preparation for subsequent fusion with other modal information.
[0065] 3. A randomly initialized small multilayer perceptron is used as a melody encoder to process the pitch and duration information of the melody, and a pooling strategy is used to convert melody marker sequences of different lengths into a uniform shape, which is then input into another multilayer perceptron to generate an updated melody representation.
[0066] A small multilayer perceptron with random initialization is used as a melody encoder. The small MLP has a relatively simple structure and low computational complexity, which is suitable for processing a specific type of data such as melody. The MLP receives the pitch and duration information of the melody as input. The pitch information reflects the height of the note, and the duration information reflects the duration of the note. The MLP extracts features and performs preliminary transformations on this information, such as identifying and encoding the combination pattern of pitch and duration.
[0067] The pooling operation downsamples the melody token sequences of different lengths and converts them into a uniform shape. After that, the melody features of the uniform shape are input into another multi-layer perceptron. This MLP further processes and transforms the pooled melody features to generate an updated melody representation. This representation integrates the key features of the melody and can better integrate with other modal information, providing effective melody information for the music generation model.
[0068] S12. Aligning the encoded audio representation, melody representation, and text representation in a unified vector space;
[0069] In this embodiment, the comparative language-music pre-training CLMP is used to achieve the alignment of the audio representation, the melody representation and the text representation. Step S12 includes the following:
[0070] 1. Calculate the contrast loss between music representation and melody representation, the contrast loss between audio representation and text representation, and the contrast loss between melody representation and text representation respectively to obtain the contrast loss function;
[0071] The total contrast loss function is obtained by calculating the contrast loss separately, with the contrast loss L between the audio representation and the melody representation am For example, the calculation formula is as follows:
[0072]
[0073] Where: N is the batch size, is the audio representation vector obtained after processing the i-th audio in the batch data, is the melody representation vector obtained after processing the i-th melody in the batch data, is the melody representation vector obtained after processing the j-th melody in the batch data, and τ is a learnable temperature parameter.
[0074] Specifically, the calculation formulas for the contrast loss between the melody representation and the audio representation, the contrast loss between the audio representation and the text representation, the contrast loss between the text representation and the audio representation, the contrast loss between the melody representation and the text representation, and the contrast loss between the text representation and the melody representation are similar to the above L am .
[0075] Combining the above contrast losses, the total contrast loss function expression is as follows:
[0076]
[0077] Where: L is the total contrast loss, L am is the contrast loss between audio representation and melody representation, L ma is the contrast loss between melody representation and audio representation, Lat is the contrast loss between audio representation and text representation, L ta is the contrast loss between text representation and audio representation, L mt is the contrast loss between melody representation and text representation, L tm is the contrast loss between text representation and melody representation.
[0078] 2. Perform contrastive learning within each batch, and align the audio representation, melody representation, and text representation into a unified vector space through the batch gradient descent algorithm with the goal of minimizing the contrastive loss function.
[0079] In each batch, the contrastive loss function calculated above is used for contrastive learning. The core idea of contrastive learning is to make positive sample pairs closer in the unified vector space, while making the distance between negative sample pairs larger. During the training process, for each batch of data, by calculating the contrastive loss between different modal representations, the model learns how to adjust the vector representation of the representation so that similar representations are closer in the vector space, while dissimilar representations are further away.
[0080] In order to minimize the contrastive loss function, the batch gradient descent algorithm is used. The gradient of the contrastive loss function with respect to each representation is calculated in each iteration. The representation vector is continuously updated through multiple iterations of the batch gradient descent algorithm, so that the contrastive loss function is gradually reduced, and finally the audio representation, melody representation and text representation are aligned into a unified vector space. In this unified space, the representations of different modalities will be reasonably distributed according to their semantic similarity and internal connection, providing more consistent and fusible input for subsequent tasks such as music generation.
[0081] S13. Construct a melody vector database based on the melody representation, and retrieve the target melody vector representation in the melody vector database using the text representation as a query condition;
[0082] The target melody vector representation is determined by using a hierarchical navigable small-world graph method. The expression of the target melody vector representation is:
[0083]
[0084] in: is the target melody vector representation, is the melody vector database, r is the element in the melody vector database, x t Represents the text.
[0085] The purpose of the retrieval operation is to find the melody vector that is most semantically relevant to the input text description, and to provide melody guidance for subsequent music generation, so that the generated music not only conforms to the text description but also has a certain melodic basis, avoiding the generation of monotonous or disharmonious music clips.
[0086] S14. Use the target melody vector representation and the text representation as fusion conditions to guide the diffusion process, and generate a latent music representation that conforms to the text description and melody guidance;
[0087] The above steps have retrieved the target melody vector representation from the melody vector database, and there is a text representation as the query. In order to use these two pieces of information as conditions for the diffusion process, they need to be fused. The fusion expression is as follows:
[0088]
[0089] Among them: c is the fusion condition, and b∈R d′ is a learnable parameter, Represents the text x t and the target melody vector representation Concatenate them to get a vector of dimension 2d′, where d′ is the dimension.
[0090] Through this linear combination, the text representation and the target melody vector representation are fused to form a comprehensive conditional vector. This fusion takes into account the information of both text and melody, providing a unified condition for the diffusion process to ensure that the generated music is influenced by both text description and melody guidance.
[0091] Specifically, the diffusion process includes a forward process and a reverse process. The forward process is a process of gradually adding noise to the original sample; the reverse process starts from the sample obtained from the forward process that obeys the normal distribution of the representation, and by gradually predicting and removing noise, the sample that conforms to the original data distribution is restored to generate a potential music representation that meets the requirements.
[0092] The expression of the forward process is as follows:
[0093]
[0094] Where: x n and x n―1 They represent the samples in the nth and n-1th steps of the diffusion process, respectively, and x 0 is the original sample, and Determines the distribution generated during the denoising process, For the known sample x at step n n , the original sample x 0 And the mean of the n-1th step samples under the fusion condition c, is the variance related parameter;
[0095] The calculation formula for noise prediction is as follows:
[0096]
[0097] Where: ∈ θ (x n ,n,c) is the predicted value of the noise at step n under condition c, ∈ θ (x n ,n) is the predicted value of the noise at the nth step without considering condition c, and w is the bootstrap parameter without classifier.
[0098] In the inference stage, guided by the fusion condition, the noise samples are continuously denoised to finally generate the first music representation that conforms to the text description and melody guidance. In this process, the diffusion model gradually adjusts the generated music representation according to the text description information and melody guidance information in the fusion condition, so that it not only conforms to the theme, emotion, etc. described in the text, but also is guided by the melody, ensuring the coherence and harmony of the generated music in melody.
[0099] The above steps use the melody vector database to perform targeted melody retrieval, and fuse the retrieved target melody vector representation with the text representation as the condition of the diffusion process to generate a latent music representation. This process combines the information of text description and melody guidance, laying the foundation for further processing and generating playable music using modules such as variational autoencoders, and fully considers the characteristics of melody and text in the music generation process, which helps to generate better quality music works that better meet user expectations.
[0100] S15. Preliminarily decode the latent music representation into a target Mel-spectrogram using an encoder in a variational autodecoder;
[0101] In this embodiment, step S15 includes the following:
[0102] 1. Use the neural network structure of the encoder in the variational autodecoder to convert the vector information in the low-dimensional latent space into the target Mel-spectrogram;
[0103] The output of the latent diffusion module is a latent music representation, which is generated in a low-dimensional latent space through a step-by-step denoising process. It contains key feature information for generating music, including information about musical elements such as melody, rhythm, and timbre. However, this representation is an abstract vector form that cannot be directly perceived by people, and requires subsequent decoding operations to convert it into a playable music form.
[0104] like Figure 2 As shown, the variational autoencoder (VAE) here consists of an encoder and a decoder. In the model, the decoder part of the VAE receives the potential music representation x output by the potential diffusion module 0Perform preliminary decoding. The preliminary decoding process is to convert the vector information in the low-dimensional latent space into a Mel spectrum map x through the neural network structure in the decoder. mel Mel-spectrogram is a time-frequency representation commonly used in audio processing. It decomposes the audio signal in the time and frequency dimensions, and can more clearly show how the frequency characteristics of the audio change over time. This initial decoding process maps the music feature information in the potential representation to a time-frequency representation space that is closer to perceptible audio, preparing for further conversion into playable music.
[0105] S16. Convert the target mel-spectrogram into playable music through a vocoder.
[0106] Mel spectrograms contain the frequency and amplitude information of audio, but they cannot be played directly. Therefore, HiFi-GAN vocoder is used to convert the Mel spectrogram output by VAE decoder into playable music. The vocoder here is based on the principle of generative adversarial network. Through adversarial training of generator and discriminator, the spectrum information in Mel spectrogram is restored to time domain signal, and finally playable audio is obtained.
[0107] The present invention is verified and described in detail in combination with experiments below:
[0108] 1. Dataset: Two well-known music datasets, MusicCaps and MusicBench, are used. The two datasets contain approximately 5,000 samples and 42,000 samples respectively, both of which contain audio and corresponding description information, and are divided into training set, validation set and test set according to 80%:10%:10% and 90%:5%:5% respectively. Basic Pitch is used to extract the melody from the music waveform file and save it as a MIDI file, retaining only the pitch and duration information. Each note in the MIDI file is converted into a triplet to determine the duration of the pitch and the time before the next pitch starts. There are 128 pitches in total, and the continuous duration range within 6.3 seconds is divided into 512 partitions.
[0109] 2. Baseline model: To evaluate MG 2 Model performance, comparing it with several current representative models to measure MG 2 The quality of the model performance. These models can be divided into two categories: text-generated audio models and text-generated music models, including: MusicLM, AudioLDM, TANGO, MusicGen, AudioLDM 2, Mustango and FluxMusic.
[0110] 3. Experimental environment: For MG 2In the CLMP part, Adam is selected as the optimizer and the learning rate is set to 1×10 -5 , the batch size is 48, and the number of epochs is 90. For MG 2 The retrieval enhancement diffusion module and decoding module are constructed with AdamW as the optimizer and the learning rate is 1×10 -4 , batch size is 24, and compression level is 4. 100 and 500 denoising steps are used on MusicCaps and MusicBench datasets, respectively, and the weights of unconditional guidance are set to 3 and 3.5, respectively. CLMP is trained on an NVIDIA RTX 4090 24GB GPU, while the retrieval-enhanced diffusion module is trained on an NVIDIA A800 80GB GPU. During training, audio is first used as a query condition to improve the model's ability to generate music, and then text is used as a query condition to improve the model's ability to understand the semantic information of text input.
[0111] 4. Evaluation method: Frechette Audio Distance (FAD), Inception Score (IS) and Kulbeck-Leibler (KL) divergence are used as evaluation indicators. FAD uses VGGish as a music classifier to evaluate the similarity between the original music and the generated music. IS measures the quality and diversity of the generated samples at the same time, and KL divergence measures the distribution distance between the generated samples and the original samples.
[0112] 5. Music Generation Results: As shown in Table 1, the proposed MG 2 Specifically, based on the FAD metric, MG 2 The performance on MusicCaps is 1.30 higher than the strongest baseline AudioLDMFull and 0.75 higher on MusicBench. 2 It is trained on MusicCaps, which contains 4,000 tracks, and MusicBench, which contains 37,000 tracks. In contrast, AudioLDM-Full is trained on 29,510 hours of samples, which is about MG 2 In addition, due to the overlap of fine-tuning datasets, models such as TANGO and AudioLDM may have been exposed to part of the MusicCaps test set in our experimental setting. 2 Still, it achieved comparable performance on the IS index, further verifying its effectiveness. 2 The overall performance of MG is better than all these baseline models and uses less training data. The melody vector database as an external knowledge base is combined with the diffusion module to enable MG 2Becomes efficient with less training data. Implicit and explicit use of melodic guidance ensures that MG 2 The generated music is harmonious both between segments and within individual segments, resulting in less noise and more harmonious rhythms.
[0113] Table 1
[0114]
[0115] 6. Multimodal Alignment: To demonstrate the effectiveness of implicitly using melody information, the performance of alignment modules with and without melody information is compared here. As shown in Table 2, CLMP is fine-tuned on MusicCaps and MusicBench using the same checkpoints provided by CLAP; Recall and Mean Average Precision are used to evaluate the alignment performance. The results clearly show that the proposed CLMP significantly outperforms CLAP on both datasets and all metrics. For example, in terms of waveform and text alignment, CLMP achieves an R@1 score of 0.7414, which is 21.82% higher than CLAP. In addition, CLMP also shows excellent performance in aligning other modalities. This improvement may be due to the intrinsic connection between melody and waveform. As Figure 3 As shown in , the music spectrograms generated under text guidance and audio guidance are almost identical in each subgraph. This consistency comes from the semantic alignment between text and audio waveforms. In summary, the proposed trimodal alignment achieves excellent performance not only in experimental metrics but also in the actual music generation process. To demonstrate the effectiveness of explicitly using melody, a reduction study is conducted here on the MusicCaps and MusicBench datasets. Figure 4 As shown, MG 2 Performance is similar to that of MG without melody 2 Specifically, the melody part in Equation 10 is replaced by zero padding. The results clearly show that on both datasets, MG 2 The FAD and KL of MG have been improved, which proves that 2 The performance of the proposed method is better than that of the non-melody version. In terms of IS indicators, the performance on the MusicCaps dataset is improved, but it is reduced on the MusicBench dataset. Figure 5 The four cases in Figure 1 illustrate the impact of deleting the melody. It can be seen from the figure that the addition of melody information makes MG 2 It can improve the completeness, clarity, semantic accuracy and rhythmic coherence of generated music.
[0116] Table 2
[0117]
[0118] 7. Case study: To prove MG 2 The advantages in understanding semantic information were compared from different perspectives. Three or four almost identical prompts were used, but with slight differences in wording, resulting in subtle changes in semantic content. Figure 6 Part of MG 2 The spectrograms show similar waveforms, indicating that the same instrument—the violin—is used, but their clarity is different. The violin music played in the cozy little room is clearer due to the lack of echo, while the violin music played in the outdoor park and train station sounds more distorted and noisy, reflecting environmental factors such as echo and background noise, which can be clearly seen from the darker parts of the spectrogram. In part (b), it can be seen that MG 2 The semantic information related to the frequency requirement is effectively captured. For example, the cue of "high pitch" produces more high-frequency content compared to the cue of "low pitch", which shows that the model is able to adapt to different frequency characteristics. In part (c), the MG 2 Generate music consistent with the specified instrumental material. The different waveforms in the spectrogram correspond to the different instruments described in the prompt. Similarly, in part (d) MG 2 The generated music matches the desired instrument, as verified by the different waveforms in the spectrogram. Figure 7 As the number of sampling steps increases from 10 to 100, the FAD and KL indicators decrease, while the IS indicator improves, which indicates that MG 2 However, when the number of steps exceeds 100, the performance improvement brought by further increasing the number of steps becomes negligible, indicating that although increasing the number of steps can improve sampling quality, the benefits diminish when the number of steps becomes large. Figure 7 (b) shows the effect of different CFG values. As CFG increases from 1.5 to 3.5, the FAD and KL metrics decrease, while the IS metric increases, showing better overall performance. However, when CFG = 4.0, all metrics deteriorate, which may be due to the CFG mechanism balancing the prediction errors of conditional and unconditional information in the reverse process of diffusion. Although conditional information improves performance, too high a CFG value will reduce performance.
[0119] The following is combined with investigation and evaluation to verify the effect of the present invention:
[0120] 1. A questionnaire containing demographic questions and specific evaluation criteria was designed for three user groups to comprehensively evaluate the proposed MG 2The criteria cover four key aspects: recognizability, text relevance, satisfaction, quality, and market potential. Each user group provides unique insights based on their expertise: ordinary users evaluate music based on subjective satisfaction, reflecting personal preferences but lacking the technical perspective of professional quality assessment. Musicians provide professional insights into music quality but may not be able to assess commercial potential. Short video creators use their experience in content creation to assess market potential. In the demographic section, 162 subjects’ attitudes towards AI-generated music were collected, such as Figure 8 The results show that 88.4% of the subjects have a supportive or neutral attitude towards AI music, which indicates that AI has broad prospects in music generation.
[0121] 2. All 162 subjects were then asked to identify AI-generated music from a set of five music clips, including two MG 2 , and three human-generated works. Subjects could select "none", "one", or "more than one". Fig. 9 As shown, only 6.7% of the subjects successfully identified two AI-generated pieces from the five music clips. More than half of the participants incorrectly labeled all human-generated music as AI-generated music and failed to identify any AI-generated music. This result is in stark contrast to AI works in the image and text domains, where people often recognized AI content despite the high quality of the models and few obvious flaws. Further analysis of the selected works showed that AI-generated works received the fewest votes for "AI music", while human-generated tracks, such as Music 5 and Music 1, received the fewest votes. This suggests that MG 2 can effectively produce music that many listeners would consider artificial. This suggests that MG 2 The result is music that many listeners believe was created by humans, highlighting its high-quality and realistic output.
[0122] 2. The second part of the questionnaire shows the 2 Seven pieces of music were generated, along with corresponding prompts and a set of evaluation questions. Each participant heard the prompt and the generated music, and then answered questions tailored to their expertise, so that the generated music could be evaluated from different perspectives. However, all participants were asked to evaluate the consistency between the generated music and its prompt. Participants chose from five options: very consistent (VC), consistent (C), uncertain (NS), inconsistent (I), and inconsistent (VI), which are recorded as 5, 4, 3, 2, and 1 in Table 3, respectively. We calculated the numerical results by averaging the responses for each piece of music, and then the average score of all seven pieces represented the MG. 2As shown in Table 3, about three-quarters of the participants rated the music as "very consistent" or "consistent with the cue", indicating that the music was highly consistent with the cue. The overall mean score was 3.88, close to the "consistent" rating of 4, further demonstrating the high cue relevance recognized by most participants.
[0123] 3. 124 participants without professional music knowledge were randomly selected for the survey. This questionnaire, tailored for general users, included questions about the clarity, appeal, and freshness of the generated music. The results are shown in Table 4. On average, 17.71% and 42.94% of the respondents rated the generated music as "very consistent" and "consistent", respectively, which meets the required clarity, appeal, and freshness attributes.
[0124] Table 3
[0125]
[0126] Table 4
[0127]
[0128] 4. To evaluate the quality of generated music, we collected responses from 18 musicians with formal music education. These musicians answered professional questions tailored to evaluate the technical and artistic quality of the music. As shown in Table 5, more than half of the musicians, including the average scores of 12.80 and 42.76 in the last row, believed that the quality of the generated music was high. These feedbacks emphasized that MG 2 Professional feasibility of the results.
[0129] Table 5
[0130]
[0131] 4. The MG 2Market potential in real creative scenarios. The number of fans of these bloggers ranged from 1,000 to 20,000. The questionnaire was designed to assess their willingness to use generated music. The questionnaire used in the previous sections to evaluate the feasibility of music in short videos was designed to assess their willingness to use generated music editing and the bloggers' willingness to pay for it. Tables 6 and 7 show that 75.72% of bloggers believe that music is feasible in short video editing. In addition, 46.43% of bloggers expressed their willingness to pay for generated music. By asking bloggers whether they would use generated music in content creation, the practicality of music was evaluated from another perspective. Three types of music were shown - exciting, funny, and suspenseful music - each with three clips. Bloggers were asked to select zero to three clips for use in their creation. As shown in Table 8, more than 90% of bloggers said they would use at least one music clip in their content. The results reflected in these three tables demonstrate the multi-perspective MG 2 The practical value of this technology highlights its potential in practical applications such as video creation.
[0132] Table 6
[0133]
[0134] Table 7
[0135]
[0136] Table 8
[0137]
[0138] In summary, the melody-guided text-generated music algorithm proposed in the present invention, namely the MG 2 The model uses melody to guide the music generation process in both implicit and explicit ways. A large number of experiments were conducted to verify the proposed MG 2 In addition, a comprehensive human evaluation was performed to demonstrate the effectiveness of MG 2 Feasibility in real-world scenarios.
[0139] It is understandable that those skilled in the art can, under the guidance of the above embodiments, combine various implementation methods in the above embodiments to obtain technical solutions of multiple implementation methods.
[0140] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A melody-guided text-generated music algorithm, characterized in that: include: Obtain data in three modalities: music waveform, melody, and text description through public datasets, and encode these three modal data respectively; Align the encoded audio representation, melody representation, and text representation in a unified vector space; Building a melody vector database based on the melody representation, and retrieving a target melody vector representation in the melody vector database using the text representation as a query condition; The target melody vector representation and the text representation are used as fusion conditions to guide the diffusion process, generating a latent music representation that conforms to the text description and melody guidance; Using a decoder in a variational autoencoder to preliminarily decode the latent music representation into a target mel-spectrogram; The target mel-spectrogram is converted into playable music through a vocoder.
2. The melody-guided text-generated music algorithm according to claim 1 is characterized in that: The method of obtaining data of three modes, namely, music waveform, melody and text description, through a public data set and encoding the three modes of data respectively includes: After converting each music waveform into a Mel-spectrogram, HTS-AT is used to encode the Mel-spectrogram into an audio representation, the parameters of HTS-AT are fixed, and a multi-layer perceptron is trained after the HTS-AT module; Use RoBERTa to encode the text description to get the text representation, fix the parameters of RoBERTa, and train a multi-layer perceptron after the RoBERTa module; A randomly initialized small multilayer perceptron is used as a melody encoder to process the pitch and duration information of the melody, and a pooling strategy is used to convert melody marker sequences of different lengths into a uniform shape, which is then input into another multilayer perceptron to generate an updated melody representation.
3. The melody-guided text-generated music algorithm according to claim 1 is characterized in that: The step of aligning the encoded audio representation, melody representation, and text representation in a unified vector space includes: The contrast loss between the audio representation and the melody representation, the contrast loss between the audio representation and the text representation, and the contrast loss between the melody representation and the text representation are calculated respectively to obtain a contrast loss function; Contrastive learning is performed within each batch, and the audio representation, melody representation, and text representation are aligned into a unified vector space through a batch gradient descent algorithm with the goal of minimizing the contrastive loss function.
4. The melody-guided text-generated music algorithm according to claim 3 is characterized in that: The contrast loss between the audio representation and the melody representation, the contrast loss between the audio representation and the text representation, and the contrast loss between the melody representation and the text representation are calculated respectively to obtain a contrast loss function. include: Where: L is the total contrast loss, L am is the contrast loss between audio representation and melody representation, L ma is the contrast loss between melody representation and audio representation, L at is the contrast loss between audio representation and text representation, L ta is the contrast loss between text representation and audio representation, L mt is the contrast loss between melody representation and text representation, L tm is the contrast loss between text representation and melody representation.
5. The melody-guided text-generated music algorithm according to claim 4 is characterized in that: The contrast loss L between the audio representation and the melody representation am The expressions include: Where: N is the batch size, is the audio representation vector obtained after processing the i-th audio in the batch data, is the melody representation vector obtained after processing the i-th melody in the batch data, is the melody representation vector obtained after processing the j-th melody in the batch data, and τ is a learnable temperature parameter.
6. The melody-guided text-generated music algorithm according to claim 1 is characterized in that: The step of constructing a melody vector database based on the audio representation and retrieving a target melody vector representation in the melody vector database using the text representation as a query condition comprises: The target melody vector representation is determined by using a hierarchical navigable small-world graph method. The expression of the target melody vector representation is: in: is the target melody vector representation, is the melody vector database, r is the element in the melody vector database, x t Represents the text.
7. The melody-guided text-generated music algorithm according to claim 6 is characterized in that: The step of using the target melody vector representation and the text representation as fusion conditions to guide the diffusion process and generate a potential music representation that conforms to the text description and melody guidance includes: The expression for fusing the target melody vector representation with the text representation is as follows: Among them: c is the fusion condition, and b∈R d′ is a learnable parameter, Represents the text x t and the target melody vector representation Concatenate them to get a vector of dimension 2d′, where d′ is the dimension.
8. The melody-guided text-generated music algorithm according to claim 7 is characterized in that: The step of using the target melody vector representation and the text representation as fusion conditions to guide the diffusion process and generate a potential music representation that conforms to the text description and melody guidance includes: The diffusion process includes a forward process and a reverse process. The forward process is a process of gradually adding noise to the original sample; the reverse process starts from the sample obtained from the forward process that obeys the normal distribution of the representation, and by gradually predicting and removing noise, the sample that conforms to the original data distribution is restored to generate a potential music representation that meets the requirements.
9. The melody-guided text-generated music algorithm according to claim 8 is characterized in that: The expression of the forward process is as follows: Where: x n and x n―1 They represent the samples in the nth and n-1th steps of the diffusion process respectively, x0 is the original sample, and Determines the distribution generated during the denoising process, For the known sample x at step n n , the mean of the n-1th step sample under the original sample x0 and fusion condition c, is the variance related parameter; The calculation formula for noise prediction is as follows: Where: ∈ θ (x n ,n,c) is the predicted value of the noise at step n under condition c, ∈ θ (x n ,n) is the predicted value of the noise at the nth step without considering condition c, and w is the bootstrap parameter without classifier.
10. The melody-guided text-generated music algorithm according to claim 1 is characterized in that: The method of using an encoder in a variational autodecoder to preliminarily decode the potential music representation into a target mel-spectrogram comprises: The neural network structure of the encoder in the variational autodecoder is used to transform the vector information in the low-dimensional latent space into the target Mel-spectrogram.
Citation Information
Patent Citations
Music generation method based on variational auto-encoder and spectrogram transformation
CN118298782A