Long-time pure music generation method based on text driving

By designing a reference network and a denoising network, combining speech activity detection and the CLAP model, screening and training data sets, and generating long pure music, the problems of rigid melody and limited length in existing technologies are solved, and music generation with melody changes and controllable listening experience is achieved, which is suitable for fields such as augmented reality, virtual reality, and game development.

CN120600000APending Publication Date: 2025-09-05NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510639240.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing technologies have difficulty generating pure music with high precision, diversity, and duration, especially in the task of converting natural language descriptions to music generation. Moreover, the music melodies generated by existing models are rigid, making it difficult to meet the needs of fields such as augmented reality, virtual reality, and game development.

Method used

By designing the reference network Reference Net and the denoising network Denoise Net, combining voice activity detection, text vectorization, the CLAP model and the mutual attention mechanism, screening and training data sets, using Mel-spectrograms to generate long pure music, and using the neural vocoder HiFiGAN for audio reconstruction, we can achieve melody changes and controllable listening experience.

Benefits of technology

It realizes the generation of long-duration pure music with rich changes and controllable listening experience, and supports the generation of music of theoretically infinite duration. It is suitable for fields such as augmented reality, virtual reality, game development and video editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600000A_ABST
    Figure CN120600000A_ABST
Patent Text Reader

Abstract

The invention provides a text-driven long-time pure music generation method, which comprises the following steps of: screening and manufacturing a pure music data set with a text label by using an LP-MusicCaps model; designing a reference network (Reference Net), and generating a Mel spectrogram coherent with a given preorder music condition; designing a de-noising network Denoise Net, and generating a Mel spectrogram conforming to a given text condition; designing a mutual attention mechanism, and fusing attention features respectively captured by the reference network Reference Net and the denoising network Denoise Net; and converting the Mel spectrogram into a monaural music audio by using a HiFiGAN model. According to the method, potential relations between music and preorder music and between music and text description are learned at the same time, and therefore pure music generation which is consistent with user input conditions and can be infinitely long in theory is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a method for generating long-duration pure music based on text-driven operation. Background Art

[0002] Generating music audio based on personalized requirements is an important task in application areas such as augmented reality, virtual reality, game development, and video editing. Traditionally, audio generation is achieved through signal processing techniques. In recent years, generative models, either unconditional or conditioned on data from other modalities, have greatly changed the approach to this task. Liu et al. (Liu, X., Iqbal, T., Zhao, J., Huang, Q., Plumbley, MD, and Wang, W. Conditional sound generation using neural discrete time-frequency representation learning. In IEEE International Workshop on Machine Learning for Signal Processing, pp. 1–6. IEEE, 2021b) use a set of audio labels to achieve the generation from labels to audio. The number of such labels is usually small, such as the ten sound categories in the UrbanSound8K dataset (Salamon, J., Jacoby, C., and Bello, JPA dataset and taxonomy for urban sound research. In Proceedings of the 22nd ACM international conference on Multimedia, pp. 1041–1044, 2014), or several music genres; in contrast, natural language descriptions are more flexible and practical than audio labels because they can contain fine-grained descriptions of audio signals, such as pitch, instruments used, and time order. Compared to traditional methods, achieving high accuracy and diversity in generating music from natural language descriptions requires a large amount of high-quality paired data to train generative models. This data, including high-fidelity audio waveforms and detailed text descriptions, often constitutes only a small fraction of available datasets. Low-quality music waveforms and weakly labeled or unlabeled data are common problems in open-source datasets. Furthermore, existing methods struggle to generate long, varied music audio.

[0003] In recent years, music generation models based on text descriptions have attracted widespread attention. The purpose of this type of generation task is to create music clips corresponding to the input descriptive text or summary text. According to the different network structures adopted, the existing work paradigms can be divided into two categories: using language models to model quantized waveform representations, and using diffusion models to model the spectral characteristics of music. Copet et al. (Copet, J.; Kreuk, F.; Gat, I.;

[0004] Remez, T.; Kant, D.; Synnaeve, G.; Adi, Y.; and De′fossez, A. 2024. Simple and controllable music generation. Advances in Neural Information Processing Systems, 36) proposed the MUSICGEN framework, which contains a Transformer architecture (Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukaszKaiser,and Illia Polosukhin.Attention is all you need.In Advances in Neural Information Processing Systems,volume 30.Curran Associates,Inc.,2017) proposed an autoregressive decoder. This model is built on the quantized units output by the audio word segmenter. It compresses the audio into a low-frame-rate discrete representation and then achieves high-fidelity reconstruction based on the discrete representation. This work also explored multiple codebook interleaving strategies and selected the most effective one to exploit the inherent structure of the quantized unit representation output by the audio word segmenter. However, the design concept of the autoregressive model determines that the generated data sequence has a strong correlation with the historical value, and the model may generate music with relatively dull melodies.

[0005] On the other hand, with the goal of generating audio from natural language descriptions, Huang et al. (Huang, P.-Y.; Xu, H.; Li, J.; Baevski, A.; Auli, M.; Galuba, W.; Metze, F.; and Feichtenhofer, C. 2022. Masked autoencoders that listen. Advances in Neural Information Processing Systems, 35: 28708–28720) studied a simple extension of the image-based masked autoencoder MAE (K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022), first encoding the music spectrogram segments at a high mask rate and feeding only the unmasked tokens into the encoder layer. The decoder then rearranges and decodes the encoded context padded with masked tokens to reconstruct the input spectrogram, from which it learns a self-supervised representation. Liu et al. (Liu, Haohe and Chen, Zehua and Yuan, Yi and Mei, Xinhao and Liu, Xubo and Mandic, Danilo and Wang, Wenwu and Plumbley, Mark D. In Proceedings of the International Conference on Machine Learning (ICML), 2023) proposed AudioLDM, a TTA (Text-to-Audio) system built on a latent space that uses audio embeddings to train a latent diffusion model while providing text embeddings as conditions during the sampling process. By learning a latent representation of the audio signal without modeling cross-modal relationships, AudioLDM improves generation quality and computational efficiency. However, the length of the music or sound effect audio generated by these models is limited by the size of the spectrogram. Summary of the Invention

[0006] Purpose of the invention: The technical problem to be solved by the present invention is to address the deficiencies of the existing technology and provide a method for generating long-duration pure music based on text-driven methods, comprising the following steps:

[0007] Step 1: Using voice activity detection (VAD) and text vectorization techniques, we filter out pure music data without human voice from the existing dataset. We then construct a short-term dataset with a single sample length of 10 seconds and text annotations, and a long-term dataset with a single sample length of 20 seconds. The minimum generation unit is a 10-second music sample.

[0008] Step 2: Design and train a reference network, Reference Net, and randomly select text generated by the previous music in the long-term dataset, or select text generated by the automatic annotation model for the current music in the long-term dataset. Use the CLAP (Contrastive Language-Audio Pretraining) model to map the text generated by the previous music in the long-term dataset into an embedding vector, or map the text generated by the automatic annotation model for the current music in the long-term dataset into an embedding vector. Use the embedding vector as a guide condition to train the Reference Net so that the Reference Net can generate a mel-spectrogram. The Reference Net includes an encoder, a decoder, and an attention block.

[0009] Step 3: Design a denoising network, Denoise Net. Use the CLAP model to map the text annotations in the short-term dataset into embedding vectors, or map the text generated by the automatic annotation model for the current music in the long-term dataset into embedding vectors. Use the embedding vectors as guidance to train Denoise Net, and make Denoise Net and Reference Net have the same structure. Initialize the denoising network Denoise Net with the weights of the trained Reference Net.

[0010] Step 4: Design a mutual attention mechanism. Use the preceding music or silent audio in the long-term dataset as the condition for the Reference Net. During the training of the Denoise Net, inject the output of the attention block in the Reference Net into the attention block of the Denoise Net.

[0011] In step 5, the Mel-spectrograms generated more than twice are combined in the time dimension according to the required duration, and the neural vocoder HiFiGAN is used to convert the Mel-spectrograms into mono audio to achieve the generation of long-term pure music.

[0012] In step 1, for the long-term data set without text annotations, the root mean square (RMS) value of the audio is calculated as the energy index; the beats per minute (BPM) value of the audio is calculated as the rhythm index, and the energy threshold Th is taken with reference to signal processing theory and music theory. energy Set the rhythm threshold Th to 0.1 tempo The value is 120, and only samples with energy index less than the energy threshold and rhythm index less than the rhythm threshold are retained, so that the filtered music has a relatively slow rhythm and is not intense or harsh.

[0013] In step 1, for the short-term dataset, the cosine similarity index between the text annotation and the template description is calculated:

[0014] S cos =ω1S1+ω2S2+ω3S3,

[0015] Among them S cos is the final similarity score after weighted summation, S1, S2 and S3 represent the cosine similarity scores between the text annotation and the first template description, the second template description and the third template description respectively, ω1, ω2 and ω3 represent the weight of the first template description, the weight of the second template description and the weight of the third template description respectively, satisfying the conditions:

[0016] ω1+ω2+ω3=1,

[0017] In the present invention, ω1 is set to 0.5, ω2 is set to 0.25, and ω3 is set to 0.25, and the similarity threshold Th is set to similarity is 0.73, only the cosine similarity index greater than the similarity threshold Th is retained similarity We screened out the text and music pairing data that met the requirements, and in the subsequent training process, unified the sampling rate of all music data to 16kHz. We expanded the duration of the music data from 10 seconds to 10.24 seconds by filling the silent part, ensuring that the size of the spectrum graph obtained after the short-time Fourier transform was the same and reasonable.

[0018] In step 2, the encoder encodes the input data into latent variables, the decoder restores the generated latent variables to the same form as the input data, and the attention block captures the relationship between the guidance condition and the latent variables to guide the generation of the latent variables;

[0019] Only long-term datasets are used to train the Reference Net. First, the long-term pure music data with a single sample length of 20s is divided into pre-order music and current music. The first 10s of data are recorded as pre-order music and the second 10s of data are recorded as current music. The CLAP model is used to map the pre-order music into an embedding vector, or the automatic annotation model LP-MusicCaps is used to generate a corresponding text annotation for the current music, recorded as the current annotation, and then the current annotation is mapped to an embedding vector, and the text selection probability γ is set. text A value of 0.2 indicates that during the training of the Reference Net, 20% of the guidance conditions were text and 80% were music.

[0020] In the process of training the reference network Reference Net, signal processing technology is used to perform a short-time Fourier transform on the current music to obtain a mel-spectrogram, and a variational autoencoder (VAE) is used to map the mel-spectrogram into latent variables in a low-dimensional latent space. The latent variables are then used as the generation targets of the reference network Reference Net under the guidance conditions obtained from the previous music or the current annotation, and the previous music is used as the conditions guiding the generation. In particular, for the latent variables, before inputting the latent variables into the reference network, the numerical range of the latent variables needs to be linearly transformed to the interval [-1,1]. Then, based on the original range before the transformation, the output of the reference network is inversely transformed to the original range to ensure that the denoising process operates in the interval [-1,1]. The input of the reference network Reference Net is the guidance conditions and the original noise-free mel-spectrogram. At a given time t, the following mathematical relationship exists:

[0021]

[0022] where x t represents the image at time t, x0 represents the original Mel spectrum without noise; ∈ t is noise, is a pre-set weight parameter that satisfies the mathematical relationship:

[0023]

[0024] α t =1-β t ,

[0025]

[0026] where β t represents the noise coefficient at time t, α t It represents the intermediate result when the weight parameter at time t is derived from the noise coefficient at time t, α krepresents the intermediate result at time k, where the value range of k is [1, t]; β start is the initial noise coefficient, generally taken as 0.00085; β end is the termination noise coefficient, which is generally taken as 0.012; T is the total number of time steps, which is generally taken as 1000;

[0027] Under the given guidance conditions, the prediction target of the reference network Reference Net, that is, the output is:

[0028]

[0029] where v t is the combination of the noise at time t and the original mel-spectrogram latent variable.

[0030] In step 3, the structure of the denoising network Denoise Net is the same as that of the reference network Reference Net;

[0031] Both the reference network Reference Net and the denoising network Denoise Net include an encoder, a decoder and an attention block, wherein the encoder consists of three downsampling blocks containing an attention mechanism and one standard downsampling block, with a total of four layers; the decoder consists of three upsampling blocks containing an attention mechanism and one standard upsampling block, with a total of four layers; the attention blocks are located in the downsampling blocks (or upsampling blocks) containing the attention mechanism, with a total of six; the denoising network Denoise Net is trained using both short-term and long-term datasets, and the weights of the trained reference network Reference Net are used to initialize the denoising network Denoise Net, and only text annotations are used as guiding conditions for the denoising network Denoise Net. In particular, for short-term datasets, the current annotations are provided by the dataset, rather than by the automatic annotation model LP-MusicCaps.

[0032] In step 4, when training the Denoise Net, a mutual attention mechanism is introduced. For long-term datasets, the reference network Reference Net is guided by the embedding vector obtained based on the previous music; for short-term datasets, the reference network Reference Net is guided by the zero vector. The final output of the reference network Reference Net is not used. The output of each attention block in the reference network Reference Net is injected into the corresponding attention block of the denoising network Denoise Net, fusing the features captured by the reference network and the denoising network Denoise Net.

[0033] During training, the encoder, decoder, and attention blocks in the reference network Reference Net are frozen, and the first three layers of the encoder in the denoising network Denoise Net are frozen. Only the parameters in the last layer of the encoder, decoder, and attention blocks in the reference network Denoise Net are updated.

[0034] The output H of each attention block in the denoising network Denoise Net i , the feature fusion expression is as follows:

[0035]

[0036] Where i is a natural number with a maximum value of 6, H i represents the output of the i-th attention block, LN represents layer normalization, is the hidden state output by the i-th attention block in the reference network Reference Net, is the hidden state output by the i-th attention block in the denoising network Denoise Net, is the hidden state after fusion, ω is the fusion rate; the guiding condition of the denoising network Denoise Net is the embedding vector obtained based on the current annotation.

[0037] In step 5, for the trained denoising network Denoise Net and reference network Reference Net, during inference, multiple mel-spectrograms are iteratively generated based on the user-entered text description T and duration requirement N. All generated results are then spliced ​​together in the time dimension, and the complete mel-spectrogram is input into the neural vocoder to obtain monophonic audio, thus achieving the generation of long-term pure music.

[0038] In step 5, first, based on the input text description T, without enabling the mutual attention mechanism, the first audio segment is generated only through the denoising network Denoise Net. Then, the first audio segment is used as the guidance condition of the reference network Reference Net, and the mutual attention mechanism is enabled to generate the second audio segment through the denoising network Denoise Net until the Nth audio segment is generated, where N is a natural number.

[0039] After meeting the duration requirement (the duration requirement is input by the user), the two or more mel-spectrograms generated are spliced ​​in the time dimension to obtain the final complete mel-spectrogram. The complete mel-spectrogram is input into the neural vocoder HiFiGAN using the neural vocoder HiFiGAN, and mono audio is output to achieve the generation of long-term pure music that meets the input text description.

[0040] The present invention also provides an electronic device, comprising a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the method.

[0041] The present invention also provides a storage medium storing a computer program or instruction, which executes the steps of the method when the computer program or instruction is run on a computer.

[0042] Beneficial Effects: This invention proposes a feasible text-driven method for generating music of theoretically unlimited duration. This method enables the generation of long-duration pure music with rich melodic variations and controllable auditory quality. The proposed method has broad practical application in fields such as augmented reality, virtual reality, game development, and video editing, demonstrating its high practical value and promising future. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more apparent.

[0044] Figure 1 Flowchart of the method of the present invention.

[0045] Figure 2 Schematic diagram of the reference network Reference Net in an embodiment of the present invention.

[0046] Figure 3 Flowchart of the operation phase in an embodiment of the present invention.

[0047] Figure 4 This is a diagram showing the results in an embodiment of the present invention. DETAILED DESCRIPTION

[0048] like Figure 1 As shown, an embodiment of the present invention provides a method for generating long-duration pure music based on text-driven operation, comprising the following steps:

[0049] Step 1: Different data screening methods are used for existing unlabeled long-term music data and short-term music data with text annotations. For unlabeled long-term music data, the numpy library function and librosa library function are called to read the audio and extract the energy and rhythm features of the audio, and the energy threshold Th is taken. energy Set the rhythm threshold Th to 0.1 trmpo The value is 120, which filters out music with a slower rhythm and not too intense or harsh.

[0050] For short-term music data with text annotations, we perform filtering based on text annotations and audio features. Taking piano music as an example, we first perform character matching and retain only data containing the word "piano" in the text annotations. Then, we use VAD (Voice Activity Detection) technology to detect whether there are segments containing human voices in the audio, and only retain data without human voices. Then, we manually provide a template description and use text vectorization technology to calculate the cosine similarity between the text annotation and the template description. The calculation formula is:

[0051] S cos =ω0S0+ω1S1+ω2S2,

[0052] Among them S i (i=0,1,2) is the cosine similarity score between the text annotation and different template descriptions, S cos is the final similarity score after weighted summation, and the similarity threshold Th is taken similarity The sampling rate of all audio data is unified to 16kHz by using signal processing technology.

[0053] Step 2: Based on the U-Net framework and referring to the idea of ​​the potential diffusion model, design and train the reference network Reference Net (reference network), such as Figure 2 As shown. During training, short-term datasets are not used. Only long-term pure music data with a single sample length of 20s is randomly intercepted from the long-term dataset and divided into pre-order music and current music. The first 10s of data are recorded as pre-order music (pre-wav, indicating the waveform of the front part), and the last 10s of data are recorded as current music (cur-wav, current waveform, indicating the waveform of the current part). Then, the CLAP model is used to map the pre-order music to an embedding vector, or the automatic annotation model LP-MusicCaps is first used to generate a corresponding text annotation for the current music, recorded as the current annotation (cur-cap, current-caption, indicating the text description condition corresponding to the current music), and then the current annotation is mapped to an embedding vector, and the text selection probability γ is taken. textIt is 0.2, which means that during the training process, 80% of the guidance conditions are obtained by mapping the previous music, and 20% of the guidance conditions are obtained by mapping the current annotation, and are added to the U-Net framework through the cross-attention mechanism. During training, signal processing technology is used to perform a short-time Fourier transform on the current music to obtain a mel-spectrogram, and VAE (Variational Auto-Encoder) is used to map the mel-spectrogram into latent variables in a low-dimensional latent space. The latent variables obtained from the current music are then used as the generation target of the reference network Reference Net under the guidance conditions obtained from the previous music (or current annotation). This makes the reference network Reference Net focus on the connection between the previous music and the current music, while also having a certain ability to process text conditions. The input of the reference network Reference Net is:

[0054]

[0055] where x t represents the image latent variable at time t, x0 represents the noise-free original Mel spectrum map latent variable, ∈ t is the noise, which obeys Gaussian distribution, is a pre-set weight parameter. Under given conditions, the prediction target of the reference network Reference Net, that is, the output is:

[0056]

[0057] where v t It is a combination of the noise at time t and the latent variables of the original image.

[0058] Step 3: Design a denoising network, DenoiseNet, with the exact same structure as the Reference Net. Initialize the DenoiseNet with the weights of the trained Reference Net. During DenoiseNet training, both unlabeled long-duration music data and short-duration music data with text annotations are used. For the short-duration music data, the preceding music is silence; for the long-duration music data, the current annotation is generated by LP-MusicCaps based on the current music. Training follows a similar strategy to that used for training the Reference Net, using the latent variables derived from the current music as the generation targets and the embedding vectors obtained by mapping the text using the CLAP model as the guiding condition.

[0059] Step 4. In the process of training the denoising network Denoise Net, introduce the mutual attention mechanism. Freeze all parameters in the reference network Reference Net, use the previous music or silent audio in the long-term dataset as the condition of the reference network Reference Net, use the latent variables obtained from the current music as the input of the reference network Reference Net, do not use the final output of the reference network Reference Net, but inject the output of each attention block in the reference network Reference Net into the corresponding layer of the denoising network Denoise Net, and fuse the features captured by the two major networks. The main structure of the denoising network Denoise Net is the same as that of the reference network Reference Net. It consists of a downsampling module, an intermediate module, and an upsampling module. During training, freeze the downsampling module in the denoising network Denoise Net, that is, the first three layers of the encoder, and only update the parameters of other parts. For the output H of each attention block in the denoising network Denoise Net i (i=1,2,…), the feature fusion expression is as follows:

[0060]

[0061] Where LN represents layer normalization, is the hidden state output by each attention block in the reference network Reference Net, is the hidden state output by each attention block in the denoising network Denoise Net, is the hidden state after fusion, ω is the fusion rate, which is set to 0.6; since the reference network Reference Net focuses on the connection between the previous music and the current music, and the denoising network Denoise Net is guided by the text annotations of the dataset or the generated text annotations during training, and is added to the U-Net framework through the cross-attention mechanism, the features captured by the two networks are fused, which is conducive to enabling the complete model to generate pure music that is both coherent in the temporal dimension and consistent with the text description; larger data volumes and richer types of text descriptions are also conducive to improving the diversity of model generation results.

[0062] Step 5: After the complete model training is completed, input the text description and duration requirements to generate long pure music that meets the text description, such as Figure 3As shown in the figure, the model first generates the first audio segment based on the input text description, using only the denoising network (Denoise Net) without mutual attention. It then uses the first audio segment as a guide for the reference network (Reference Net), enabling mutual attention and generating the second audio segment through the denoising network (Denoise Net), and so on. Once the duration requirement is met, the multiple generated mel-spectrograms are spliced ​​together in the time dimension to obtain the final complete mel-spectrogram. This complete mel-spectrogram is then input into the neural vocoder (HiFiGAN) and output as monophonic audio, achieving the generation of long-duration pure music that matches the input text description.

[0063] Through the above steps, the present invention can generate long pure music that meets the text description according to the text description and the required duration, and the generated result is as follows: Figure 4 As shown, it has good consistency with the text description and coherence in the time dimension.

[0064] The present invention provides a method for generating long-duration pure music based on text-driven processing. There are many methods and approaches for implementing this technical solution. The above is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art may make various improvements and modifications without departing from the principles of the present invention, and such improvements and modifications should also be considered within the scope of protection of the present invention. Any components not specified in this embodiment may be implemented using existing technologies.

Claims

1. A method for generating long-duration pure music based on text-driven, characterized in that: The following steps are involved: Step 1: Use voice activity detection (VAD) and text vectorization technology to filter out pure music data without human voice from the dataset. Construct a short-term dataset with a single sample length of 10 seconds and text annotations, and a long-term dataset with a single sample length of 20 seconds. Use a music sample of 10 seconds as the minimum generation unit. Step 2: Design and train a reference network, Reference Net, and randomly select text generated by the previous music in the long-term dataset, or select text generated by the automatic annotation model for the current music in the long-term dataset. Use the CLAP model to map the text generated by the previous music in the long-term dataset into an embedding vector, or map the text generated by the automatic annotation model for the current music in the long-term dataset into an embedding vector. Use the embedding vector as a guide condition to train the Reference Net so that the Reference Net can generate a mel-spectrogram. The Reference Net includes an encoder, a decoder, and an attention block. Step 3: Design a denoising network, Denoise Net. Use the CLAP model to map the text annotations in the short-term dataset into embedding vectors, or map the text generated by the automatic annotation model for the current music in the long-term dataset into embedding vectors. Use the embedding vectors as guidance to train Denoise Net, and make Denoise Net and Reference Net have the same structure. Use the weights of the trained Reference Net to initialize the denoising network, Denoise Net. Step 4: Design a mutual attention mechanism. Use the preceding music or silent audio in the long-term dataset as the condition for the Reference Net. During the training of the Denoise Net, inject the output of the attention block in the Reference Net into the attention block of the Denoise Net. In step 5, the Mel-spectrograms generated more than twice are combined in the time dimension according to the required duration, and the neural vocoder HiFiGAN is used to convert the Mel-spectrograms into mono audio to achieve the generation of long-term pure music.

2. The method according to claim 1, characterized in that In step 1, for the long-term data set without text annotations, the root mean square (RMS) value of the audio is calculated as the energy index; the beats per minute (BPM) value of the audio is calculated as the rhythm index, and the energy threshold Th is set based on the signal processing theory and music theory. energy , set the rhythm threshold Th tempo , only retain the samples whose energy index is less than the energy threshold and whose rhythm index is less than the rhythm threshold, and obtain the filtered music.

3. The method according to claim 2, characterized in that In step 1, for the short-term dataset, the cosine similarity index between the text annotation and the template description is calculated: S cos =ω1S1+ω2S2+ω3S3, Among them S cos is the final similarity score after weighted summation, S1, S2 and S3 represent the cosine similarity scores between the text annotation and the first template description, the second template description and the third template description respectively, ω1, ω2 and ω3 represent the weight of the first template description, the weight of the second template description and the weight of the third template description respectively, satisfying the conditions: ω1+ω2+ω3=1, Set the similarity threshold Th similarity , only retain the cosine similarity index greater than the similarity threshold Th similarity The samples were screened to obtain text and music pairing data that met the requirements. In the subsequent training process, the sampling rate of all music data was unified to 16kHz, and the duration of the music data was extended from 10 seconds to 10.24 seconds by filling the silent part.

4. The method according to claim 3, characterized in that In step 2, the encoder encodes the input data into latent variables, the decoder restores the generated latent variables to the same form as the input data, and the attention block captures the relationship between the guidance condition and the latent variables to guide the generation of the latent variables; Only long-term datasets are used to train the Reference Net. First, the long-term pure music data with a single sample length of 20s is divided into pre-order music and current music. The first 10s of data are recorded as pre-order music and the second 10s of data are recorded as current music. The CLAP model is used to map the pre-order music into an embedding vector, or the automatic annotation model LP-MusicCaps is used to generate a corresponding text annotation for the current music, recorded as the current annotation, and then the current annotation is mapped to an embedding vector, and the text selection probability γ is set. text ; During the training of the reference network, signal processing techniques are used to perform a short-time Fourier transform on the current music to obtain a mel-spectrogram. The mel-spectrogram is then mapped into latent variables in a low-dimensional latent space using a variational autoencoder (VAE). The latent variables are then used as generation targets for the reference network under guidance from the previous music or the current annotation, with the previous music serving as the guiding generation condition. Before inputting the latent variables into the reference network, the latent variable numerical range needs to be linearly transformed to the interval [-1, 1]. Then, based on the original range before the transformation, the output of the reference network is back-transformed back to the original range. The input of the reference network Reference Net is the guidance condition and the noise-free original Mel spectrum map. At a given time t, there is the following mathematical relationship: where x t represents the image at time t, x0 represents the original Mel spectrum without noise; ∈ t is noise, is a pre-set weight parameter that satisfies the mathematical relationship: α t =1-β t , where β t represents the noise coefficient at time t, α t It represents the intermediate result when the weight parameter at time t is derived from the noise coefficient at time t, α k represents the intermediate result at time k, where the value range of k is [1, t]; β start is the initial noise coefficient; β end is the termination noise coefficient; T is the total number of time steps; Under the given guidance conditions, the prediction target of the reference network Reference Net, that is, the output is: where v t is the combination of the noise at time t and the original mel-spectrogram latent variable.

5. The method according to claim 4, characterized in that In step 3, the structure of the denoising network Denoise Net is the same as that of the reference network Reference Net; Both the reference network Reference Net and the denoising network Denoise Net include an encoder, a decoder and an attention block, wherein the encoder consists of three downsampling blocks containing an attention mechanism and one standard downsampling block, with a total of four layers; the decoder consists of three upsampling blocks containing an attention mechanism and one standard upsampling block, with a total of four layers; the attention blocks are located in the downsampling blocks or upsampling blocks containing the attention mechanism, with a total of six; the denoising network Denoise Net is trained using short-term and long-term datasets at the same time, and the weights of the trained reference network Reference Net are used to initialize the denoising network Denoise Net, and only text annotations are used as guiding conditions for the denoising network Denoise Net. For the short-term dataset, the current annotation is provided by the dataset.

6. The method according to claim 5, characterized in that In step 4, when training the denoising network DenoiseNet, a mutual attention mechanism is introduced. For long-term datasets, the reference network Reference Net is guided by the embedding vector obtained based on the previous music; for short-term datasets, the reference network Reference Net is guided by the zero vector. The final output of the reference network Reference Net is not used. The output of each attention block in the reference network Reference Net is injected into the corresponding attention block of the denoising network DenoiseNet, fusing the features captured by the reference network Reference Net and the denoising network DenoiseNet. During training, the encoder, decoder, and attention blocks in the reference network Reference Net are frozen, and the first three layers of the encoder in the denoising network Denoise Net are frozen. Only the parameters in the last layer of the encoder, decoder, and attention blocks in the reference network Denoise Net are updated. The output H of each attention block in the denoising network Denoise Net i , the feature fusion expression is as follows: Where i is a natural number with a maximum value of 6, H i represents the output of the i-th attention block, LN represents layer normalization, is the hidden state output by the i-th attention block in the reference network Reference Net, is the hidden state output by the i-th attention block in the denoising network Denoise Net, is the hidden state after fusion, ω is the fusion rate; the guiding condition of the denoising network Denoise Net is the embedding vector obtained based on the current annotation.

7. The method according to claim 6, characterized in that In step 5, for the trained denoising network Denoise Net and reference network Reference Net, during inference, the mel-spectrogram is first iteratively generated based on the user-entered text description T and duration requirement N. All generated results are then spliced ​​together in the time dimension, and the complete mel-spectrogram is input into the neural vocoder to obtain monophonic audio, thus achieving the generation of long-term pure music.

8. The method according to claim 7, characterized in that In step 5, first, based on the input text description T, without enabling the mutual attention mechanism, only the denoising network Denoise Net is used to generate the first audio segment. Then, the first audio segment is used as the guidance condition of the reference network Reference Net, with the mutual attention mechanism enabled, and the second audio segment is generated through the denoising network Denoise Net until the Nth audio segment is generated, where N is a natural number. After meeting the duration requirement, the two or more generated mel-spectrograms are spliced ​​together in the time dimension to obtain the final complete mel-spectrogram. The complete mel-spectrogram is then input into the neural vocoder HiFiGAN using the neural vocoder HiFiGAN, which outputs mono audio to achieve the generation of long-duration pure music that meets the input text description.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 8.

10. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 8 are executed.