Music generation method and device, electronic equipment and readable storage medium
By looping into text description and audio data to the music generation model, and retaining the audio data for the initial time period in the audio data, the problem of music quality decline and generation length limitation in the existing music generation methods is solved, and streaming generation of high-quality music is achieved.
Patent Information
- Application Number
- CN202510066231.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
The problems of music quality decline and generation length limitations in existing music generation methods.
By inputting text description and first audio data into the trained music generation model, second audio data is generated, and the audio data of the preset time period before the second audio data is set to the audio data of the initial time period, the step is performed cyclically until the generation time period reaches the preset time period.
It realizes the generation of high-quality music of any length without reducing the music quality, solving the problem of limited music generation length.
Smart Images

Figure CN119993099A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of computer technology, and in particular, relates to a music generation method, device, electronic device and readable storage medium. Background Art
[0002] The existing music generation method is to use the generative model to predict the next music clip based on the last predicted music clip. Specifically, the generative model uses the self-attention mechanism to infer the last predicted music clip and predict the next music clip.
[0003] However, this method of music generation is limited by the model architecture, resulting in a decline in music quality and a limited generation length. Summary of the invention
[0004] The embodiments of the present application provide a music generation method, device, electronic device, readable storage medium and computer program product, which can solve the problems of reduced quality and limited generation length of music generated by a generation model.
[0005] In a first aspect, an embodiment of the present application provides a music generation method, comprising:
[0006] Inputting the text description and the first audio data into a trained music generation model to obtain second audio data output by the trained music generation model, wherein the text description is a music style description;
[0007] The audio data of the first preset time period in the second audio data is set as the audio data of the initial time period. After obtaining the first audio data, return to the execution step: input the text description and the first audio data into the trained music generation model, obtain the second audio data output by the trained music generation model, until the generation time reaches the preset time;
[0008] The trained music generation model is used to generate and output the second audio data that conforms to the text description based on the first audio data.
[0009] In a second aspect, an embodiment of the present application provides a music generating device, including:
[0010] A first generation module is used to input the text description and the first audio data into a trained music generation model to obtain second audio data output by the trained music generation model, wherein the text description is a music style description;
[0011] A second generating module is used to set the audio data of the first preset time period in the second audio data as the audio data of the initial time period, and after obtaining the first audio data, return to the execution step of: inputting the text description and the first audio data into the trained music generation model, obtaining the second audio data output by the trained music generation model, until the generation time reaches the preset time;
[0012] The trained music generation model is used to generate and output the second audio data that conforms to the text description based on the first audio data.
[0013] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements a method as described in any one of the above-mentioned first aspects when executing the computer program.
[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method as described in any one of the above-mentioned first aspects is implemented.
[0015] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device executes any one of the methods described in the first aspect.
[0016] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0017] In an embodiment of the present application, a text description and first audio data are input into a trained music generation model to obtain second audio data output by the trained music generation model, wherein the text description is a music style description; the audio data of a preset time period before the second audio data is set as the audio data of an initial time period, and after obtaining the first audio data, the execution step is returned to: the text description and the first audio data are input into the trained music generation model to obtain the second audio data output by the trained music generation model, until the generation duration reaches a preset duration; the trained music generation model is used to generate and output second audio data that conforms to the text description based on the first audio data, so that the trained music generation model combines the audio data of the initial time period and the audio data output by the model last time to predict the second audio data, thereby ensuring the quality of the generated audio data and obtaining high-quality music of the required duration.
[0018] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 This is a first flow chart of a music generation method provided by an embodiment of the present application;
[0021] Figure 2 is a schematic diagram of displaying audio provided by an embodiment of the present application;
[0022] Figure 3 This is a second flow chart of the music generation method provided by an embodiment of the present application;
[0023] Figure 4 is a structural schematic diagram of a music generating device provided in one embodiment of the present application;
[0024] Figure 5 It is a schematic diagram of the structure of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0025] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0026] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0027] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0028] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0029] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0030] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0031] In one embodiment, Figure 1 This is a first flow chart of a music generation method provided by an embodiment of the present application. Figure 1 As shown, the method comprises:
[0032] S11: Input the text description and the first audio data into the trained music generation model to obtain the second audio data output by the trained music generation model, where the text description is a music style description.
[0033] The trained music generation model is used to generate and output second audio data that conforms to the text description based on the first audio data.
[0034] S12: The audio data of the previous preset time period in the second audio data is set as the audio data of the initial time period. After obtaining the first audio data, return to the execution step: input the text description and the first audio data into the trained music generation model, obtain the second audio data output by the trained music generation model, until the generation time reaches the preset time.
[0035] In the application, the first audio data includes the audio data of the initial time period and part of the audio of the last prediction result of the trained music generation model, which is the context information of the second audio data and serves as the pre-music for subsequent music reasoning. The trained music generation model predicts the following audio that is consistent with the structure of the first audio data based on the first audio data, and adjusts the content of the following audio based on the text description to obtain the second audio data.
[0036] The second audio data becomes the last prediction result of the trained music generation model. The audio data of the preset time period before the second audio data is replaced with the audio data of the initial time period to obtain the input data of the trained music generation model, that is, to obtain the first audio data. Then, step S11 is continued to realize continuous music inference until the generation time reaches the preset time.
[0037] The audio data of the initial time period is the audio data of the first preset time period in the initial audio data that is first input into the trained music generation model. The generation duration can be the total duration of the music or the duration of music generation.
[0038] For example, the initial audio data includes 1-6s audio, and the audio data of the initial time period includes 1-2s audio data. The first audio data includes 1s, 2s, 5s, 6s, 7s, 8s audio, and the second audio data includes 4-9s. The second audio data is processed to obtain the first audio data 1s, 2s, 6s, 7s, 8s, 9s required for the next prediction.
[0039] It is understandable that the existing music generation model makes inferences based on the audio data of the self-attention window length for each prediction. If the self-attention window length is directly moved to infer audio data that exceeds the window length, the music quality will decline rapidly. By analyzing the heat maps output by each layer of the attention layer, it is found that the attention layers of different layers maintain a large proportion of the weight of the initial input vector, and the SoftMax function causes the model to have attention convergence. Based on this, the results of the model output are processed, that is, the audio data of the initial time period is retained and the attention window is moved, so that the data input to the model each time for prediction includes the audio data of the initial time period and the audio data output by the model last time, and high-quality music can be output each time, and the problem of attention convergence can be solved without adding additional operations. In addition, the quality of music can be guaranteed without limiting the growth time, and low-quality music will not appear, so that users can obtain high-quality music of any required length, and realize streaming music generation.
[0040] By retaining the audio data of the initial time period and then moving the attention window, the music reasoning meets the consistency of the context structure and the process of repeatedly operating repeated fragments can be reduced, thereby reducing the delay in reasoning.
[0041] In one possible implementation, the text description received by the trained music generation model each time is the same or different.
[0042] In the application, when the user needs to change the music style, the text description of the input model is changed so that the trained music generation model can generate the music required by the user. At this time, the text description received by the trained music generation model is different from the text description received last time.
[0043] When the user does not need to change the music style, the previous text description continues to be used. At this time, the text description received by the trained music generation model is the same as the text description received last time.
[0044] Among them, music style includes BPM, instrument, mood, etc.
[0045] By changing the text description to change the music style, the user's demand for different styles of music can be met, and the music style can be switched at will.
[0046] In a possible implementation, after step S11, the method further includes:
[0047] The second audio data is played.
[0048] Figure 2 is a schematic diagram of displaying audio provided by an embodiment of the present application. Figure 2 As shown, each time the second audio data is obtained, it is pushed to the front-end page, supporting streaming generation of music, so that the user can be aware of and interact with the second audio data in a timely manner.
[0049] In this embodiment, a text description and first audio data are input into a trained music generation model to obtain second audio data output by the trained music generation model, wherein the text description is a music style description; audio data of a preset time period before the second audio data is set as audio data of an initial time period, and after obtaining the first audio data, the execution step is returned to: inputting the text description and the first audio data into the trained music generation model to obtain second audio data output by the trained music generation model, until the generation duration reaches a preset duration; the trained music generation model is used to generate and output second audio data that conforms to the text description based on the first audio data, so that the trained music generation model combines the audio data of the initial time period and the audio data output by the model last time to predict the second audio data, thereby ensuring the quality of the generated audio data and obtaining high-quality music of the required duration.
[0050] In one embodiment, the trained music generation model includes a trained text encoder, a trained Transformer decoder, a trained audio encoder, and a trained audio decoder.
[0051] Figure 3 1 is a second flow chart of the music generation method provided in one embodiment of the present application. Figure 3 As shown, step S11 includes:
[0052] S111: Input the text description into the trained text encoder, and encode the text description through the trained text encoder to obtain a text vector.
[0053] S112: Input the first audio data into the trained audio encoder, and encode the first audio data by the trained audio encoder to obtain a first audio vector.
[0054] In a possible implementation, after step S112, the method further includes:
[0055] S115: Using a compression algorithm, compress the first audio vector to obtain an audio discrete vector.
[0056] In the application, a music compression representation algorithm is used to discretize continuous audio using residual vector quantization to allocate it to the codeword closest to the Euclidean distance metric, and each codebook includes multiple codewords.
[0057] A hierarchical codebook structure is used through residual vector quantization to save codebook size and obtain discrete music representation with lower bit rate and high reconstruction quality.
[0058] S113: Input the text vector and the first audio vector into the trained Transformer decoder, and use the trained Transformer decoder to predict the second audio vector based on the text vector and the first audio vector using a multi-head attention mechanism.
[0059] In a possible implementation, correspondingly, step S113 includes:
[0060] S116: Input the text vector and the audio discrete vector into the trained Transformer decoder, and use the trained Transformer decoder to predict the second audio vector based on the text vector and the audio discrete vector using a multi-head attention mechanism.
[0061] In the application, the trained Transformer decoder includes multiple Transformer layers. In the Transformer layer, the audio discrete vector is processed by a masked multi-head attention layer and a linear layer. Then, the audio discrete vector and the text vector are cross-attended by a multi-head cross-attention layer, and then a second audio vector is output through a linear layer.
[0062] The text vector is introduced into the model as a condition so that the second audio vector predicted by the model matches the text description.
[0063] S114: Input the second audio vector to the trained audio decoder, decode the second audio vector through the trained audio decoder, and obtain and output second audio data.
[0064] In an application, the trained audio decoder asynchronously decodes the second audio vector to obtain second audio data.
[0065] This embodiment can generate second audio data that is more consistent with the structure of the first audio data and more matched with the text description based on the text vector and the first audio vector through the trained Transformer decoder.
[0066] In one embodiment, the music generation model is fine-tuned using LoRA, and a low-rank matrix is added to the music generation model parameters as an additional training parameter. Correspondingly, before step S11, the method further includes:
[0067] S21: Obtain training samples of various music categories, where the training samples include various audio samples and corresponding text annotations.
[0068] In a possible implementation, step S21 includes:
[0069] For various music categories, each audio sample of the music category is input into a trained audio understanding model, and the corresponding text annotation output by the trained audio understanding model is obtained to obtain a training sample.
[0070] In the application, the audio sample is input into the trained audio understanding model, and the trained audio understanding model provides a description of the music style for the audio sample to obtain text annotations.
[0071] Automatically annotate music through trained audio understanding models, reducing the cost of manual annotation.
[0072] In a possible implementation, the audio understanding model may be Qwen2Audio (AI speech model). The trained audio understanding model generates text annotations of audio samples according to the instructions.
[0073] For example, the instruction may include "Please describe the instrument, emotion, category, and pitch of this piece of music" and an answer example. The text is annotated as "This piece of music is a classical folk song played on guitar, with a tempo of 150 BPM..."
[0074] S22: Input the text annotation into the text encoder in the music generation model, encode the text annotation through the text encoder, and obtain a text annotation vector.
[0075] In the application, the text encoder is a pre-trained model. The text encoder encodes the text annotation to obtain the text annotation vector.
[0076] S23: Input the audio sample to the audio encoder in the music generation model, encode the audio sample through the audio encoder, and obtain an audio sample vector.
[0077] In the application, the audio encoder is a pre-trained model. The audio encoder encodes the audio sample through the audio encoder to obtain an audio sample vector.
[0078] S24: Input the text annotation vector and the audio sample vector into the Transformer decoder in the music generation model, and generate an audio prediction vector according to the text annotation vector and the audio sample vector through the Transformer decoder.
[0079] S25: Using the loss function of the music generation model, determine the value of the loss function according to the audio prediction vector and the actual audio vector.
[0080] S26: If the value of the loss function is greater than or equal to the preset loss value, update the low-rank matrix parameters of the Transformer decoder and return to the execution step: input the text annotation into the text encoder in the music generation model, encode the text annotation through the text encoder, and obtain the text annotation vector until the value of the loss function is less than the preset loss value.
[0081] S27: When the value of the loss function is less than the preset loss value, the trained music generation model is obtained.
[0082] In the application, LoRA is injected into the attention QKV matrix of the Transformer decoder. Specifically, the language model head matrix and the encoding and decoding mapping matrix injected into the attention QKV matrix are expressed as h=W0x+BAx, where W0 is the attention QKV matrix, that is, the original parameter matrix of the Transformer decoder, x is the input data, and B and A are the low-rank matrices of LoRA. Then initialize the matrix of LoRA, the A matrix is initialized using Kaiming (weight initialization), and the B matrix is initialized to zero.
[0083] The language model head matrix is the matrix of the linear layer that maps the Transformer decoder to the audio encoder input. The encoder-decoder mapping matrix is the matrix of the linear layer that maps the text encoder and audio encoder to the Transformer decoder input.
[0084] Then, the loss function is used to calculate the difference between the audio prediction vector and the actual audio vector to obtain the value of the loss function. The value of the loss function is back-propagated through the gradient to update the B and A low-rank matrices, and other parameters in the model are not updated. When the value of the loss function is less than the preset loss value, the model is fine-tuned and a trained music generation model is obtained.
[0085] The loss function may be a cross loss function. The actual audio vector is obtained by encoding real music.
[0086] This embodiment introduces low-rank matrices and updates low-rank matrices through LoRA to achieve model fine-tuning, which can quickly train the model without losing the generalization of the original model, so that the model supports music style generation of low-resource data, can generate music of various musical styles, and reduce the requirement for the length of audio samples through LoRA fine-tuning the model.
[0087] It should be understood that the order of execution of the steps in the above embodiments does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. In addition, the data collection in the above embodiments is compliant, and its use or implementation does not involve any infringement on the public interest.
[0088] Corresponding to the method described in the above embodiment, for the convenience of explanation, only the part related to the embodiment of the present application is shown.
[0089] In one embodiment, Figure 4 Schematic diagram of the structure of a music generating device provided by an embodiment of the present application. Figure 4 As shown, the device comprises:
[0090] The first generation module 10 is used to input the text description and the first audio data into the trained music generation model to obtain the second audio data output by the trained music generation model, and the text description is the music style description.
[0091] The second generation module 11 is used to set the audio data of the first preset time period in the second audio data as the audio data of the initial time period, and after obtaining the first audio data, return to the execution step: input the text description and the first audio data into the trained music generation model, and obtain the second audio data output by the trained music generation model, until the generation time reaches the preset time.
[0092] The trained music generation model is used to generate and output second audio data that conforms to the text description based on the first audio data.
[0093] In one embodiment, the first generation module is specifically used to input the text description into a trained text encoder, and encode the text description through the trained text encoder to obtain a text vector; input the first audio data into the trained audio encoder, and encode the first audio data through the trained audio encoder to obtain a first audio vector; input the text vector and the first audio vector into a trained Transformer decoder, and predict the second audio vector based on the text vector and the first audio vector through the trained Transformer decoder using a multi-head attention mechanism; input the second audio vector into the trained audio decoder, and decode the second audio vector through the trained audio decoder to obtain and output the second audio data; wherein the trained music generation model includes a trained text encoder, a trained Transformer decoder, a trained audio encoder, and a trained audio decoder.
[0094] In one embodiment, the first generation module is also used to compress the first audio vector using a compression algorithm to obtain an audio discrete vector; input the text vector and the audio discrete vector into a trained Transformer decoder, and use the trained Transformer decoder to predict the second audio vector based on the text vector and the audio discrete vector using a multi-head attention mechanism.
[0095] In one embodiment, the device further comprises:
[0096] A training module is used to obtain training samples of various music categories, where the training samples include audio samples and corresponding text annotations;
[0097] It is also used to input the text annotation into a text encoder in the music generation model, encode the text annotation through the text encoder, and obtain a text annotation vector;
[0098] It is also used to input the audio sample into the audio encoder in the music generation model, encode the audio sample through the audio encoder, and obtain the audio sample vector;
[0099] It is also used to input the text annotation vector and the audio sample vector into the Transformer decoder in the music generation model, and generate an audio prediction vector according to the text annotation vector and the audio sample vector through the Transformer decoder;
[0100] It is also used to use the loss function of the music generation model to determine the value of the loss function according to the audio prediction vector and the actual audio vector;
[0101] It is also used to update the low-rank matrix parameters of the Transformer decoder if the value of the loss function is greater than or equal to the preset loss value, and return to the execution step: input the text annotation to the text encoder in the music generation model, encode the text annotation through the text encoder, and obtain the text annotation vector until the value of the loss function is less than the preset loss value;
[0102] It is also used to obtain a trained music generation model when the value of the loss function is less than a preset loss value.
[0103] In one embodiment, the training module is further used to input each audio sample of each music category into a trained audio understanding model for each music category, obtain corresponding text annotations output by the trained audio understanding model, and obtain training samples.
[0104] Figure 5 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 5 As shown, the electronic device 2 of this embodiment includes: at least one processor 20 ( Figure 5 Only one is shown in the figure), a memory 21, and a computer program 22 stored in the memory 21 and executable on the at least one processor 20, wherein the processor 20 implements the steps of any of the above-mentioned method embodiments when executing the computer program 22.
[0105] The electronic device 2 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The electronic device 2 may include, but is not limited to, a processor 20 and a memory 21. Those skilled in the art will appreciate that Figure 5 It is only an example of the electronic device 2 and does not constitute a limitation on the electronic device 2. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.
[0106] The processor 20 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0107] In some embodiments, the memory 21 may be an internal storage unit of the electronic device 2, such as a hard disk or memory of the electronic device 2. In other embodiments, the memory 21 may also be an external storage device of the electronic device 2, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device 2. Further, the memory 21 may also include both an internal storage unit of the electronic device 2 and an external storage device. The memory 21 is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program. The memory 21 may also be used to temporarily store data that has been output or is to be output.
[0108] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0109] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0110] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0111] An embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0112] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the camera / terminal device, a recording medium, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electric carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some cases, the computer-readable medium cannot be an electric carrier signal and a telecommunication signal.
[0113] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0114] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0115] In the embodiments provided in the present application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0116] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0117] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A music generation method, characterized in that: include: Inputting the text description and the first audio data into a trained music generation model to obtain second audio data output by the trained music generation model, wherein the text description is a music style description; The audio data of the first preset time period in the second audio data is set as the audio data of the initial time period. After obtaining the first audio data, return to the execution step: input the text description and the first audio data into the trained music generation model, obtain the second audio data output by the trained music generation model, until the generation time reaches the preset time; The trained music generation model is used to generate and output the second audio data that conforms to the text description based on the first audio data.
2. The method according to claim 1, characterized in that The text description received by the trained music generation model each time may be the same or different.
3. The method according to claim 1 or 2, characterized in that: The trained music generation model includes a trained text encoder, a trained Transformer decoder, a trained audio encoder and a trained audio decoder; Inputting the text description and the first audio data into a trained music generation model to obtain second audio data output by the trained music generation model includes: Inputting the text description into the trained text encoder, and encoding the text description by the trained text encoder to obtain a text vector; Inputting the first audio data into the trained audio encoder, and encoding the first audio data by the trained audio encoder to obtain a first audio vector; Inputting the text vector and the first audio vector into the trained Transformer decoder, and predicting the second audio vector according to the text vector and the first audio vector by the trained Transformer decoder using a multi-head attention mechanism; The second audio vector is input into the trained audio decoder, and the second audio vector is decoded by the trained audio decoder to obtain and output the second audio data.
4. The method according to claim 3, characterized in that After encoding the first audio data by the trained audio encoder to obtain the first audio vector, the method further includes: Using a compression algorithm, compressing the first audio vector to obtain an audio discrete vector; The text vector and the audio discrete vector are input into the trained Transformer decoder, and the trained Transformer decoder uses a multi-head attention mechanism to predict a second audio vector based on the text vector and the audio discrete vector.
5. The method according to claim 4, characterized in that Before inputting the text description and the first audio data into the trained music generation model, the method further includes: Obtain training samples of various music categories, wherein the training samples include each audio sample and a corresponding text annotation; Inputting the text annotation into a text encoder in a music generation model, encoding the text annotation by the text encoder to obtain a text annotation vector; Input the audio sample into an audio encoder in the music generation model, and encode the audio sample by the audio encoder to obtain an audio sample vector; Inputting the text annotation vector and the audio sample vector into a Transformer decoder in the music generation model, and generating an audio prediction vector according to the text annotation vector and the audio sample vector through the Transformer decoder; Using the loss function of the music generation model, determining a value of the loss function according to the audio prediction vector and the actual audio vector; If the value of the loss function is greater than or equal to the preset loss value, the low-rank matrix parameters of the Transformer decoder are updated, and the execution step is returned to: the text annotation is input into a text encoder in the music generation model, and the text annotation is encoded by the text encoder to obtain a text annotation vector, until the value of the loss function is less than the preset loss value; When the value of the loss function is less than the preset loss value, the trained music generation model is obtained.
6. The method according to claim 5, characterized in that The method of obtaining training samples of multiple music categories includes: For various music categories, each of the audio samples of the music category is input into a trained audio understanding model, and the corresponding text annotation output by the trained audio understanding model is obtained to obtain the training sample.
7. A music generating device, characterized in that: include: A first generation module is used to input the text description and the first audio data into a trained music generation model to obtain second audio data output by the trained music generation model, wherein the text description is a music style description; A second generating module is used to set the audio data of the first preset time period in the second audio data as the audio data of the initial time period, and after obtaining the first audio data, return to the execution step of: inputting the text description and the first audio data into the trained music generation model, obtaining the second audio data output by the trained music generation model, until the generation time reaches the preset time; The trained music generation model is used to generate and output the second audio data that conforms to the text description based on the first audio data.
8. The device according to claim 7, characterized in that Also includes: A training module, used to obtain training samples of various music categories, wherein the training samples include audio samples and corresponding text annotations; Also used for inputting the text annotation into a text encoder in a music generation model, encoding the text annotation by the text encoder to obtain a text annotation vector; Also used for inputting the audio sample into an audio encoder in the music generation model, encoding the audio sample by the audio encoder to obtain an audio sample vector; It is also used to input the text annotation vector and the audio sample vector into a Transformer decoder in the music generation model, and generate an audio prediction vector according to the text annotation vector and the audio sample vector through the Transformer decoder; Also used for using the loss function of the music generation model to determine the value of the loss function according to the audio prediction vector and the actual audio vector; It is also used to update the low-rank matrix parameters of the Transformer decoder if the value of the loss function is greater than or equal to the preset loss value, and return to the execution step: input the text annotation to the text encoder in the music generation model, encode the text annotation by the text encoder, and obtain a text annotation vector until the value of the loss function is less than the preset loss value; It is also used to obtain the trained music generation model when the value of the loss function is less than the preset loss value.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.