Song generation method and device and electronic equipment

The lyrics and time position characteristics of songs are processed through the Diffusion Transformer model, and the song audio is generated and restored, which solves the problems of complexity and slow convergence speed of existing AI-generated music models, and realizes the effect of simplifying input parameters and accelerating the convergence of the generation model.

CN119964528APending Publication Date: 2025-05-09WONDERSHARE TECH (HUNAN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411958063.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing AI-generated music model architecture is complex and has slow convergence speed, requiring GPU cluster training, making it difficult to simplify the input parameters and accelerate the convergence speed of the generation model.

Method used

By obtaining the lyrics and time position characteristics of the song and inputting them into the Diffusion Transformer model, an intermediate expression of the song audio is generated, and the intermediate expression is finally restored to the target song audio.

Benefits of technology

It is implemented to generate songs required by users by simply entering the lyrics and time position characteristics of the song, simplifying the input parameters and speeding up the convergence speed of the generation model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964528A_ABST
    Figure CN119964528A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a song generation method. The method comprises the following steps: obtaining lyric features and time position features of a song; generating an intermediate expression of the song audio according to the lyric features and the time position features; and restoring the intermediate expression of the song audio into a target song audio. Embodiments of the invention provide a song generation method and apparatus, and an electronic device, which can generate a song required by a user only by inputting lyric features and time position features of the song, thereby simplifying input parameters and accelerating the convergence speed of a generation model. The embodiment of the invention further provides a song generation device and electronic equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of music editing, and more specifically, to a song generation method, device, and electronic device. Background Art

[0002] AI (Artificial Intelligence) music generation is an emerging field that has emerged in recent years with the rapid development of artificial intelligence technology. It uses advanced technologies such as deep learning and neural networks to learn and analyze a large amount of music data, master the basic laws and style characteristics of music, and thus be able to create music fragments or complete music works.

[0003] However, AI-generated music still faces some challenges and problems. For example, the existing AI-generated music model architecture is complex, converges slowly, and requires GPU cluster training. Summary of the invention

[0004] In response to the problems existing in the above-mentioned prior art, the embodiments of the present application provide a song generation method, device, and electronic device, which only require the input of the lyrics features and time position features of the song to generate the song required by the user, thereby simplifying the input parameters and accelerating the convergence speed of the generation model.

[0005] In a first aspect, an embodiment of the present application provides a song generation method, comprising the following steps:

[0006] Obtain lyrics features and time position features of the song;

[0007] Generate an intermediate expression of the song audio according to the lyrics feature and the time position feature; and

[0008] The intermediate expression of the song audio is restored to the target song audio.

[0009] Furthermore, the step of obtaining lyrics features and time position features of a song includes:

[0010] Encoding the lyrics of the song into an embedding vector to obtain lyrics features of the song; and

[0011] The time position of the song is encoded to obtain the time position feature of the song.

[0012] Furthermore, generating an intermediate expression of the song audio according to the lyrics feature and the time position feature includes:

[0013] The lyrics features and the time position features are input into a diffusion model to generate an intermediate representation of the song audio.

[0014] Furthermore, the step of inputting the lyrics feature and the time position feature into a diffusion model to generate an intermediate expression of the song audio includes:

[0015] The lyrics features and the time position features are added in the second dimension and input into the diffusion model to generate an intermediate representation of the song audio.

[0016] Furthermore, before inputting the lyrics feature and the time position feature into the diffusion model to generate an intermediate expression of the song audio, the method further includes:

[0017] The diffusion model is trained, and the loss function of the diffusion model training is to perform mean square error loss calculation on the output audio waveform.

[0018] Furthermore, the step of restoring the intermediate expression of the song audio to the target song audio includes:

[0019] The intermediate expression of the song audio is restored to an audio waveform to generate the target song audio.

[0020] Furthermore, the step of restoring the intermediate expression of the song audio to an audio waveform to generate the target song audio includes:

[0021] The intermediate expression of the song audio is restored to an audio waveform through the VAE model to generate the target song audio.

[0022] In a second aspect, the embodiment of the present application further provides a song generation device, including:

[0023] A feature acquisition module is used to acquire lyrics features and time position features of songs;

[0024] An intermediate expression generation module, used to generate an intermediate expression of the song audio according to the lyrics feature and the time position feature; and

[0025] The target audio generation module is used to restore the intermediate expression of the song audio into the target song audio recommendation.

[0026] In a third aspect, an embodiment of the present application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to implement the song generation method according to the first aspect described above when executing the program.

[0027] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, wherein the computer program is used to implement the song generation method according to the first aspect above.

[0028] The embodiments of the present application bring the following beneficial effects:

[0029] In the song generation method provided in the embodiment of the present application, the lyrics features and time position features of the song are first obtained, and based on the lyrics features and the time position features, an intermediate expression of the song audio is generated, and finally the intermediate expression of the song audio is restored to the target song audio. The song generation method provided in the embodiment of the present application only needs to input the lyrics features and time position features of the song to generate the song required by the user, thereby simplifying the input parameters and accelerating the convergence speed of the generation model. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.

[0031] Figure 1 A flowchart of a song generation method provided in an embodiment of the present application;

[0032] Figure 2 A structural block diagram of a song generation device provided in an embodiment of the present application;

[0033] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0034] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0035] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments described in the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this application.

[0036] In the specification and claims of this application and the above-mentioned drawings, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "multiple" means two or more. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to the specific circumstances.

[0037] Figure 1 is a flow chart of a song generation method according to an embodiment of the present application. Figure 1 As shown, the song generation method of the embodiment of the present application includes the following steps:

[0038] S101: Acquire lyrics features and time position features of a song;

[0039] The embodiment of the present application is based on Diffusion Transformer as the main model architecture. Diffusion Transformer is a generation framework that combines Transformer and diffusion model. It gradually converts data from the original state to the target state through a series of diffusion steps. The model will learn the distribution law of data in different states and generate new data samples accordingly. The uniqueness of Diffusion Transformer is that it can capture subtle differences between data, thereby generating more realistic and diverse data samples. In the song generation method provided in the embodiment of the present application, the lyrics features and time position features need to be input into the model through cross attention.

[0040] S102: Generate an intermediate expression of the song audio according to the lyrics feature and the time position feature; and

[0041] That is, the lyrics features and time position features are directly input into the cross attention, and high-quality audio intermediate representation is directly generated through DiffusionTransformer.

[0042] S103: Restore the intermediate expression of the song audio to the target song audio.

[0043] As described above, when the lyrics features and time position features of the song are obtained, an intermediate expression of the song audio is generated according to the lyrics features and the time position features, and finally the intermediate expression of the song audio is restored to the target song audio. The target song here includes vocal singing and accompaniment.

[0044] Therefore, in the song generation method provided in the embodiment of the present application, the lyrics features and time position features of the song are first obtained, and based on the lyrics features and the time position features, an intermediate expression of the song audio is generated, and finally the intermediate expression of the song audio is restored to the target song audio. The song generation method provided in the embodiment of the present application only needs to input the lyrics features and time position features of the song to generate the song required by the user, thereby simplifying the input parameters and accelerating the convergence speed of the generation model.

[0045] It should be noted that the solution adopted in the embodiment of the present application is based on an improved network architecture based on Stable Audio. Stable Audio is composed of four modules: text encoder, time position encoder, variational autoencoder (VAE) and diffusion model. The text encoder mainly encodes the input lyrics to obtain the feature embedding vector of the lyrics text. The time position encoder converts the given duration information into a continuous Fourier embedding vector. The differential autoencoder mainly restores the audio intermediate expression generated by the diffusion model into an audio waveform. The input of the diffusion model has two parts, one is random Gaussian noise, and the other is the cross-attention of the text or lyrics features and the time position features, and then generates the intermediate expression of the audio.

[0046] Of course, the Diffusion Transformer model provided in the embodiment of the present application can replace the Transformer with a U-Net network structure to achieve the same effect.

[0047] The embodiment of the present application is based on the song generation algorithm of Diffusion Transformer. The specific implementation process of the four modules is described as follows by way of example:

[0048] Suppose the input lyrics are "When will the bright moon appear? I ask the sky with wine in hand". The lyrics are encoded using the T5 (a large language model) model. They are first tokenized through T5 to obtain the unique label corresponding to the discrete Chinese characters. Here, it is assumed that "When will the bright moon appear? I ask the sky with wine in hand" is tokenized to obtain "76, 108, 512, 11, 2154, 776, 1982, 2345, 99, 654, 98". Then, T5 is used to encode these token sequences into embedding vectors. Suppose the feature dimension of the embedding vector is 512, that is, a 11x512 feature is obtained. Therefore, the features of the lyrics are obtained through the text encoder.

[0049] Assume that the input time position is 13.2, that is, the starting time of the lyrics "When will the bright moon appear? I raise my wine cup to ask the blue sky" in the whole song is 13.2 seconds. A simple time position encoder is designed here, which accepts floating-point numbers in a given range (minimum value is 0, maximum value is 512) and returns the continuous Fourier embedding of the provided floating-point numbers. The feature dimension obtained here is 256x512.

[0050] The lyrics features and time position features obtained above are added in the second dimension to obtain a new feature dimension of 267x512, which is input into the diffusion model. The diffusion model uses the Diffusion Transformer structure, which uses an open source model without modification. In the inference phase of the Diffusion Transformer model, random Gaussian mixed noise is input, and the cross attention input is the new features (267x512) above. Then the Diffusion Transformer model uses these inputs to obtain the intermediate expression of the audio.

[0051] The audio intermediate expression output by the Diffusion Transformer model is consistent with the encoder output of the VAE, so the audio intermediate expression can be restored to a waveform through the VAE decoder. The VAE module uses DAC (Descript Audio Codec), which can compress the audio waveform and restore the compressed features to a waveform through the decoder. The DAC uses an open source model and has not been modified.

[0052] Furthermore, the step of obtaining lyrics features and time position features of a song includes:

[0053] Encoding the lyrics of the song into an embedding vector to obtain lyrics features of the song; and

[0054] The time position of the song is encoded to obtain the time position feature of the song.

[0055] As mentioned above, the text encoder mainly encodes the input lyrics to obtain the feature embedding vector of the lyrics text to obtain the lyrics features of the song. The time position encoder converts the given duration information into a continuous Fourier embedding vector for input into the diffusion model for processing, thereby generating the target song audio required by the user.

[0056] Furthermore, generating an intermediate expression of the song audio according to the lyrics feature and the time position feature includes:

[0057] The lyrics features and the time position features are input into a diffusion model to generate an intermediate representation of the song audio.

[0058] Specifically, the input of the diffusion model has two parts, one is random Gaussian noise, and the other is the cross-attention of text or lyrics features and time position features, and then an intermediate representation of the audio is generated in order to generate the target song audio required by the user.

[0059] Furthermore, the step of inputting the lyrics feature and the time position feature into a diffusion model to generate an intermediate expression of the song audio includes:

[0060] The lyrics features and the time position features are added in the second dimension and input into the diffusion model to generate an intermediate representation of the song audio.

[0061] As described above, the lyrics features and the time position features are added in the second dimension to obtain a new feature dimension of 267x512, which is input into the diffusion model to generate an intermediate representation of the song audio.

[0062] Furthermore, before inputting the lyrics feature and the time position feature into the diffusion model to generate an intermediate expression of the song audio, the method further includes:

[0063] The diffusion model is trained, and the loss function of the diffusion model training is to perform mean square error loss calculation on the output audio waveform.

[0064] Specifically, before the lyrics feature and the time position feature are input into the diffusion model to generate the intermediate expression of the song audio, the expansion model needs to be trained. The specific training process is as follows:

[0065] The text encoder preferably uses an open source pre-trained model. Since the model is trained directly using lyrics, the language information covered by the text is relatively weak, so directly using the pre-trained model can better obtain the semantic information of the text.

[0066] The temporal position encoding module is trained together with the Diffusion Transformer model. This module has only two layers of networks. The first layer is nn.Embedding and the second layer is nn.Linear. The training process inputs the temporal position data (floating point numbers) and outputs the continuous Fourier embedding of floating point numbers with a feature dimension of 256x512.

[0067] The training process of Diffusion Transformer is a standard Diffusion process. The input of the Diffusion forward process is the audio features, and then Gaussian mixed noise is added to the features. The audio features with added noise are then input into the Transformer model. At the same time, the above lyrics text features and time position features are input into the cross attention. Through the powerful multimodal learning ability of Transformer, the denoised audio expression is output.

[0068] The training process of VAE is based on the open source model of DAC. The model is not modified here. The training data is the same as the data for training Diffusion Transformer. Finally, an audio compression reversible codec is obtained. When training this module, it does not need to rely on other modules and can be trained independently.

[0069] All the above modules are required when training Diffusion Transformer. Text Ecoder and VAE do not need to update parameters. The training loss function) performs MSE loss (mean square error loss) on the output audio waveform.

[0070] The song generation method provided in the embodiment of the present application uses the interval distance (FD, Fréchet Distance) to evaluate the similarity between the statistics of the audio set generated in the feature space and the reference audio set under the same training set and parameter settings. The FD value of AudioLDM2 is 170.31, while the FD value of the solution architecture provided in the embodiment of the present application is 103.66.

[0071] Furthermore, the step of restoring the intermediate expression of the song audio to the target song audio includes:

[0072] The intermediate expression of the song audio is restored to an audio waveform to generate the target song audio.

[0073] Specifically, unlike the AudioLDM2 architecture, which needs to learn a large audio-based language model, then use the large language model to generate an intermediate audio expression from the lyrics, and finally use a decoder to convert the intermediate expression into an audio waveform, the embodiment of the present application directly inputs the lyrics and position information into the cross-attention, directly generates high-quality audio intermediate expressions through the Diffusion Transformer, and then restores it to the audio waveform through the VAE, making the model training more stable.

[0074] Furthermore, the step of restoring the intermediate expression of the song audio to an audio waveform to generate the target song audio includes:

[0075] The intermediate expression of the song audio is restored to an audio waveform through the VAE model to generate the target song audio.

[0076] Specifically, the VAE model mainly restores the audio intermediate expression generated by the diffusion model into an audio waveform to generate the target song audio required by the user.

[0077] Figure 2 2 is a structural block diagram of the song generation device 200 provided in the embodiment of the present application. Figure 2 As shown, the song generation device 200 of the embodiment of the present application includes: a feature acquisition module 210, an intermediate expression generation module 220 and a target audio generation module 230, wherein:

[0078] The feature acquisition module 210 is used to acquire the lyrics features and time position features of the song;

[0079] An intermediate expression generation module 220, configured to generate an intermediate expression of the song audio according to the lyrics feature and the time position feature; and

[0080] The target audio generation module 230 is used to restore the intermediate expression of the song audio into the target song audio recommendation.

[0081] In the song generation device provided in the embodiment of the present application, the lyrics features and time position features of the song are first obtained, and based on the lyrics features and the time position features, an intermediate expression of the song audio is generated, and finally the intermediate expression of the song audio is restored to the target song audio. The song generation method provided in the embodiment of the present application only needs to input the lyrics features and time position features of the song to generate the song required by the user, thereby simplifying the input parameters and accelerating the convergence speed of the generation model.

[0082] It should be noted that the specific implementation method of the song generation device of the embodiment of the present application is similar to the specific implementation method of the song generation method of the embodiment of the present application. Please refer to the description of the method part for details and will not be repeated here.

[0083] Figure 3 Schematic diagram of the structure of an electronic device 300 according to an embodiment of the present application.

[0084] like Figure 3As shown, the electronic device 300 includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from the storage part 302 to a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The CPU 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0085] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, etc.; an output section 307 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed, so that a computer program read therefrom is installed into the storage section 308 as needed.

[0086] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a machine-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 309, and / or installed from a removable medium 311. When the computer program is executed by a central processing unit (CPU) 301, the above-mentioned functions defined in the electronic device of the present application are executed.

[0087] It should be noted that the computer-readable medium shown in the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electronic device, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0088] In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction-executing electronic device, apparatus, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in combination with an instruction-executing electronic device, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0089] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the processing receiving device, method and computer program product according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the aforementioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based electronic device that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0090] The units or modules involved in the embodiments described in this application may be implemented by software or hardware. The units or modules described may also be set in a processor, and the processor is used to implement the song generation method when executing the program:

[0091] Obtain lyrics features and time position features of the song;

[0092] Generate an intermediate expression of the song audio according to the lyrics feature and the time position feature; and

[0093] The intermediate expression of the song audio is restored to the target song audio.

[0094] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or may exist independently and not be installed in the electronic device. The above computer-readable storage medium stores one or more programs, and when the above programs are used by one or more processors to execute the song generation method described in the present application:

[0095] Obtain lyrics features and time position features of the song;

[0096] Generate an intermediate expression of the song audio according to the lyrics feature and the time position feature; and

[0097] The intermediate expression of the song audio is restored to the target song audio.

[0098] As another aspect, the present application further provides a computer program product, which may be included in the electronic device described in the above embodiment; or may exist independently without being installed in the electronic device. The above computer program product stores one or more programs, and when the above programs are used by one or more processors to execute the song generation method described in the present application:

[0099] Obtain lyrics features and time position features of the song;

[0100] Generate an intermediate expression of the song audio according to the lyrics feature and the time position feature; and

[0101] The intermediate expression of the song audio is restored to the target song audio.

[0102] The above description is only a preferred embodiment of the present application, and does not limit the patent scope of the present application. All equivalent structural changes made by using the contents of the present application specification and drawings under the application concept of the present application, or directly / indirectly used in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A song generation method, characterized in that: The following steps are involved: Obtain lyrics features and time position features of the song; Generate an intermediate expression of the song audio according to the lyrics feature and the time position feature; and The intermediate expression of the song audio is restored to the target song audio.

2. The song generation method according to claim 1, characterized in that: The step of obtaining lyrics features and time position features of a song includes: Encoding the lyrics of the song into an embedding vector to obtain lyrics features of the song; and The time position of the song is encoded to obtain the time position feature of the song.

3. The song generation method according to claim 1, characterized in that: The step of generating an intermediate expression of the song audio according to the lyrics feature and the time position feature comprises: The lyrics features and the time position features are input into a diffusion model to generate an intermediate representation of the song audio.

4. The song generation method according to claim 3, characterized in that: The step of inputting the lyrics feature and the time position feature into a diffusion model to generate an intermediate expression of the song audio includes: The lyrics features and the time position features are added in the second dimension and input into the diffusion model to generate an intermediate representation of the song audio.

5. The song generation method according to claim 4, characterized in that: Before inputting the lyrics feature and the time position feature into the diffusion model to generate an intermediate expression of the song audio, the method further includes: The diffusion model is trained, and the loss function of the diffusion model training is to perform mean square error loss calculation on the output audio waveform.

6. The song generation method according to claim 1, characterized in that: The step of restoring the intermediate expression of the song audio to the target song audio comprises: The intermediate expression of the song audio is restored to an audio waveform to generate the target song audio.

7. The song generation method according to claim 1, characterized in that: The step of restoring the intermediate expression of the song audio to an audio waveform to generate the target song audio comprises: The intermediate expression of the song audio is restored to an audio waveform through the VAE model to generate the target song audio.

8. A song generating device, characterized in that: include: A feature acquisition module is used to acquire lyrics features and time position features of songs; An intermediate expression generation module, used to generate an intermediate expression of the song audio according to the lyrics feature and the time position feature; and The target audio generation module is used to restore the intermediate expression of the song audio into the target song audio recommendation.

9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor is used to implement the song generation method according to any one of claims 1 to 7 when executing the program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is used to implement the song generation method according to any one of claims 1-7.

Citation Information

Cited By

  • Song generation method, song generation model training method, device and equipment

    CN119889256A

  • Song generation method, song generation model training method, device and equipment

    CN119889256B