Music generation method and device, electronic equipment and storage medium

By pre-training the initial music generation model and optimizing the GRPO algorithm, combined with a unified audio aesthetic evaluation tool, the problems of insufficient independent innovation and poor sound quality of existing music generation models are solved, and high-quality and stable music generation is achieved.

CN120673731APending Publication Date: 2025-09-19BEIJING HONGMIAN XIAOBING TECH CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202511143233.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing music generation models lack independent innovation capabilities, the sound quality of the generated music is poor, and the data requirements and annotation costs are high, making it difficult to meet the needs of professional music production and high-quality music appreciation.

Method used

By pre-training the initial music generation model, combining the GRPO algorithm and the unified audio aesthetics evaluation tool, optimizing the generation strategy, using the MusicFlow model to learn basic patterns and structures, introducing the Audiobox-Aesthetics tool for aesthetic evaluation, and optimizing the generation results.

Benefits of technology

It improves the artistry and auditory effects of music generation, reduces dependence on large amounts of training data, reduces annotation costs, ensures the stability and reliability of generated music, and avoids problems of model collapse and style deviation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673731A_ABST
    Figure CN120673731A_ABST
Patent Text Reader

Abstract

The invention provides a music generation method and device, electronic equipment and a storage medium, and relates to the technical field of music generation. In the training process of a target music generation model adopted by the method, an initial music generation model is pre-trained; the initial music generation model can learn the basic mode and the structure in the complete music fragment sample, and a foundation is laid for subsequent reinforcement learning optimization. By introducing a GRPO algorithm and a unified audio aesthetic evaluation tool, the pre-training model can be accurately optimized to obtain a target music generation model. The GRPO algorithm takes an aesthetic evaluation index of a unified audio aesthetic evaluation tool as a reward index, guides the pre-training model to adjust the generation strategy, can accurately measure the artistic value of the target music fragment generated by the target music generation model, enables the target music fragment to be more coordinated and graceful in melody, harmony, rhythm and other aspects, remarkably improves the quality, and improves the artistic value of the target music fragment. And the aesthetic standard is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of music generation, and in particular to a music generation method, device, electronic device and storage medium. Background Art

[0002] With the continuous development of artificial intelligence technology in the field of music creation, music generation has gradually become a research hotspot. With the rise of large-scale model technology, models such as MusicLM and MusicGen have emerged in the field of music generation. These models connect natural language and audio, generating music based on text prompts and specifying genres, instruments, and emotions.

[0003] However, while these models can generate music with a certain melody, harmony, and rhythm, the music they create often imitates and repeats the training data, lacking independent innovation capabilities and struggling to break through the traditional music framework to produce innovative works that are completely independent of existing data. Furthermore, due to the high complexity of audio, improving sound quality requires extremely high computing power, resulting in poor sound quality for the generated music, making it difficult to meet the demands of professional music production and high-quality music appreciation. Furthermore, the model training process requires a large amount of training data, and the annotation of music data requires certain professional skills. Therefore, the annotation cost of music data is relatively high, and the overall data volume is relatively small, resulting in poor model performance. Summary of the Invention

[0004] The present invention provides a music generation method, device, electronic device and storage medium to solve the defects existing in the related art.

[0005] The present invention provides a music generation method, comprising: Acquiring user input information, wherein the input information includes input text, or includes input text and input music clip; Inputting the input information into a target music generation model to obtain a target music segment output by the target music generation model that matches the input information; The target music generation model is trained based on the following steps: Taking the masked first music clip sample and the first text description sample corresponding to the first music clip sample as input, and the complete music clip sample corresponding to the first music clip sample as a label, pre-training the initial music generation model to obtain a pre-trained model; Using the pre-trained model as a strategy model, and applying the strategy model to the second text description sample corresponding to each second music clip sample in the training sample set, and obtaining multiple groups of generation results based on a group relative strategy optimization algorithm; Based on a unified audio aesthetic evaluation tool, the values ​​of each reward indicator corresponding to each group of generation results are determined, and based on the values ​​of each reward indicator corresponding to each group of generation results, the advantage value of each group of generation results is calculated. Based on the advantage value of each group of generation results, the strategy model is trained to obtain the target music generation model.

[0006] According to a music generation method provided by the present invention, the input music segment is masked, and the target music segment is a music segment filled with mask prediction results; Alternatively, the input music segment and the target music segment are continuous in time.

[0007] According to a music generation method provided by the present invention, determining the values ​​of each reward indicator corresponding to each group of generation results based on a unified audio aesthetics evaluation tool includes: Based on the values ​​of each reward indicator corresponding to each set of generated results, calculate the expected value of each reward indicator; Perform weighted summation on the expected values ​​of each reward indicator to obtain the total reward value of each group's generated results; Based on the total reward value, the average reward value and the reward value standard deviation of each group of generated results are calculated respectively, and based on the total reward value, the average reward value and the reward value standard deviation, the advantage value of each group of generated results is calculated.

[0008] According to a music generation method provided by the present invention, the target music generation model includes a semantic feature extraction module, a feature conversion module and a vocoder; Inputting the input information into a target music generation model to obtain a target music segment output by the target music generation model that matches the input information includes: Inputting the input information into the semantic feature extraction module, and having the semantic feature extraction module output semantic features; Inputting the semantic feature into the feature conversion module, and having the feature conversion module output the acoustic feature corresponding to the semantic feature; The acoustic features are input into the vocoder to obtain the target music segment output by the vocoder.

[0009] According to a music generation method provided by the present invention, the input information is input into a target music generation model, and a target music segment output by the target music generation model that matches the input information is obtained, which then includes: Obtaining a manual evaluation result of the target music clip; Based on the manual evaluation results, the generation strategy and structural parameters of the target music generation model are optimized.

[0010] According to a music generation method provided by the present invention, determining the values ​​of each reward indicator corresponding to each group of generation results based on a unified audio aesthetics evaluation tool includes: Determine the values ​​of the evaluation indicators for each group of generated results based on the unified audio aesthetics evaluation tool, and use the values ​​of the evaluation indicators for each group of generated results as the values ​​of the reward indicators corresponding to each group of generated results; The evaluation indicators include content enjoyment, content practicality, production complexity and production quality.

[0011] According to a music generation method provided by the present invention, the structure and parameters of the initial music generation model are determined based on music generation requirements.

[0012] The present invention also provides a music generating device, comprising: An input acquisition module, configured to acquire user input information, wherein the input information includes input text, or includes input text and input music clip; a music generation module, configured to input the input information into a target music generation model, and obtain a target music segment output by the target music generation model that matches the input information; The training module is used to train the target music generation model based on the following steps: Taking the masked first music clip sample and the first text description sample corresponding to the first music clip sample as input, and the complete music clip sample corresponding to the first music clip sample as a label, pre-training the initial music generation model to obtain a pre-trained model; Using the pre-trained model as a strategy model, and applying the strategy model to the second text description sample corresponding to each second music clip sample in the training sample set, and obtaining multiple groups of generation results based on a group relative strategy optimization algorithm; Based on a unified audio aesthetic evaluation tool, the values ​​of each reward indicator corresponding to each group of generation results are determined, and based on the values ​​of each reward indicator corresponding to each group of generation results, the advantage value of each group of generation results is calculated. Based on the advantage value of each group of generation results, the strategy model is trained to obtain the target music generation model.

[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the music generation method as described above is implemented.

[0014] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which implements any of the above-described music generation methods when executed by a processor.

[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-described music generation methods.

[0016] The present invention provides a music generation method, device, electronic device, and storage medium. During the training process, the target music generation model employed in this method pre-trains an initial music generation model, enabling the initial music generation model to learn the basic patterns and structures found in complete music clip samples, laying the foundation for subsequent reinforcement learning optimization. By introducing the GRPO algorithm and a unified audio aesthetics evaluation tool, the pre-trained model can be precisely optimized to obtain the target music generation model. The GRPO algorithm uses the aesthetic evaluation metrics of the unified audio aesthetics evaluation tool as reward indicators, guiding the pre-trained model to adjust its generation strategy. This accurately measures the artistic value of the target music clip generated by the target music generation model, resulting in a more harmonious and beautiful target music clip in terms of melody, harmony, and rhythm, significantly improving its quality and meeting aesthetic standards. Compared to traditional music generation methods, the target music clip generated using the target music generation model in the embodiments of the present invention is more appealing and appealing in terms of both artistry and auditory quality. Furthermore, the introduction of the GRPO algorithm not only improves the stability of the target music generation model's music generation process and enhances the quality and complexity of the target music clip, but also reduces reliance on large amounts of training data and reduces annotation costs. During the training process, through continuous learning and adjustment, the pre-trained model gradually converges to the optimal generation strategy, avoiding issues such as pattern collapse and stylistic deviations that can occur in traditional generation methods. Furthermore, the real-time scoring feedback from the unified audio aesthetics evaluation tool promptly identifies issues with the generated music, enabling timely adjustments and corrections to the pre-trained model, ensuring the stability and reliability of the resulting target music generation model. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 It is a flowchart of the music generation method provided by the present invention.

[0019] Figure 2 It is a structural schematic diagram of the music generating device provided by the present invention.

[0020] Figure 3 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0021] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0022] Since the music generation methods in the prior art have problems such as insufficient artistry and creativity, poor sound quality, and high data requirements and annotation costs, a music generation method is provided in an embodiment of the present invention to solve the above problems.

[0023] Figure 1 A flowchart of a music generation method provided in an embodiment of the present invention is shown in FIG. Figure 1 As shown, the method includes: S1, obtaining user input information, wherein the input information includes input text, or includes input text and input music clip; S2, inputting the input information into a target music generation model, and obtaining a target music segment output by the target music generation model that matches the input information; The target music generation model is trained based on the following steps: Taking the masked first music clip sample and the first text description sample corresponding to the first music clip sample as input, and the complete music clip sample corresponding to the first music clip sample as a label, pre-training the initial music generation model to obtain a pre-trained model; Using the pre-trained model as a strategy model, and applying the strategy model to the second text description sample corresponding to each second music clip sample in the training sample set, and obtaining multiple groups of generation results based on a group relative strategy optimization algorithm; Based on a unified audio aesthetic evaluation tool, the values ​​of each reward indicator corresponding to each group of generation results are determined, and based on the values ​​of each reward indicator corresponding to each group of generation results, the advantage value of each group of generation results is calculated. Based on the advantage value of each group of generation results, the strategy model is trained to obtain the target music generation model.

[0024] Specifically, the music generation method provided in the embodiment of the present invention is executed by a music generation device, which can be configured in a computer. The computer can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., which is not specifically limited here.

[0025] First, step S1 is executed to obtain user input information. The user input information can be input text or input text and input music clip.

[0026] The input text is a text description of the target music clip to be generated. It describes the target music clip and can include at least one of its length, style, rhythm, melody, harmony, emotion, and musical structure. Musical structure refers to the organization and logical relationships between the various parts of the target music clip, including elements such as paragraph division, thematic development, and repetition and contrast, which form the overall framework of the work (e.g., sonata form, three-part form, etc.). The input music clip is a reference music clip used to generate the target music clip. It can be a masked or unmasked music clip.

[0027] When the user's input information is input text, the purpose of the music generation method is to generate a target music clip from the input text; when the user's input information is input text and input music clip, if the input music clip carries a mask, the purpose of the music generation method is to use the input text to predict and fill the mask of the input music clip, generate a complete target music clip, and realize music filling. At this time, the target music clip is a music clip filled with the mask prediction result; if the input music clip does not carry a mask, the purpose of the music generation method is to use the input text to generate a target music clip after the input music clip, and realize the continuation of music. At this time, the input music clip and the target music clip are continuous in time.

[0028] Then, step S2 is executed to introduce a target music generation model, input the input information into the target music generation model, and the target music generation model outputs a target music segment that matches the input information based on the specific content of the input information.

[0029] The training steps of the target music generation model may include: The first music clip sample with a mask and the first text description sample corresponding to the first music clip sample are taken as input, and the complete music clip sample corresponding to the first music clip sample is taken as a label, and the initial music generation model is pre-trained to obtain a pre-trained model.

[0030] The initial music generation model can be a MusicFlow model, which provides a solid foundation for music creation. The MusicFlow model is an advanced flow-matching-based text-guided music generation model designed to efficiently generate high-quality music clips that match text descriptions.

[0031] A major advantage of the MusicFlow model lies in its efficient training and inference process, which is due to the high efficiency of the stream matching method. Compared with traditional generative models, the MusicFlow model is significantly smaller in size and faster inference speed, making music generation more efficient and practical. Furthermore, the MusicFlow model not only generates music from text but also performs tasks such as musical completion and continuation through contextual learning.

[0032] The purpose of pre-training is to enable the initial music generation model to learn the basic patterns and structures in the complete music clip samples, laying the foundation for subsequent reinforcement learning optimization.

[0033] The complete music clip samples can come from public music datasets, such as MusicNet, MAESTRO, etc., or can be music clips of a specific style or type collected and organized by users, which is not specifically limited here.

[0034] The pre-training process uses supervised learning, using complete music clip samples as labels. A masked first music clip sample and a first text description sample corresponding to the first music clip sample are input into the initial music generation model. The output of the initial music generation model is obtained. The loss is calculated using the output and the complete music clip sample. The structural parameters of the initial music generation model are iteratively optimized to minimize the loss until a preset number of iterations is reached or the loss converges, resulting in a pre-trained model. The preset number of iterations can be set as needed and is not specifically limited here.

[0035] Subsequently, the Group Relative Policy Optimization (GRPO) algorithm was introduced. GRPO is a reinforcement learning algorithm that uses learning to predict reward functions to guide the policy model to continuously optimize the generated policy. By optimizing the generated policy, the agent learns the optimal behavior strategy based on the reward signal fed back by the environment. The core idea of ​​the GRPO algorithm is to evaluate the advantage of the policy in a group of samples through a relative reward mechanism and update the policy model accordingly. Compared with the traditional Proximal Policy Optimization (PPO) algorithm, the GRPO algorithm simplifies its structure and directly obtains the reward signal through the use of rules or models, without the need for a separate value model, thereby improving training efficiency.

[0036] Furthermore, the GRPO algorithm gradually adjusts its generation strategy based on feedback from the current generation results. This eliminates the need for a large amount of perfect example data upfront and allows for continuous optimization of generated music sequences during training, making the learning process more efficient and reducing the initial need for large amounts of high-quality data.

[0037] The pretrained model is used as the policy model and applied to the second text description sample corresponding to each second music clip sample in the training sample set. The GRPO algorithm is used to generate multiple sets of results. The initial reference model has the same parameters as the pretrained model and remains frozen during the GRPO algorithm. Each flow step of the policy model outputs a probability distribution, and each set of generation results is obtained through sampling.

[0038] Afterwards, a unified audio aesthetics evaluation tool is introduced. The unified audio aesthetics evaluation tool can be Audiobox-Aesthetics. Audiobox-Aesthetics is a professional music aesthetics scoring tool that can quantitatively evaluate music works from multiple aesthetic perspectives. It can output evaluation indicators of four aesthetic dimensions: content enjoyment (CE), content usefulness (CU), production complexity (PC) and production quality (PQ).

[0039] The unified audio aesthetics evaluation tool determines the value of each reward metric for each set of generated results. Each evaluation metric can be used as a reward metric, and the value of each evaluation metric is also the value of a reward metric. Using the values ​​of each reward metric for each set of generated results, the advantage value of each set of generated results can be calculated. This advantage value is then used to construct an objective function, which is then used to optimize the structural parameters of the policy model. The KL loss can also be introduced when constructing the objective function.

[0040] By continuously iterating the above sampling, updating the values ​​of each reward indicator, the advantage value calculation process and the objective function construction process, the structural parameters of the strategy model are continuously optimized until the model can output the optimal generation strategy and obtain the target music generation model.

[0041] During the training process, the structural parameters of the policy model can also be saved periodically and the generated music clips can be evaluated to monitor the performance changes of the policy model.

[0042] The music generation method provided in the embodiments of the present invention uses a target music generation model. During the training process, the initial music generation model is pre-trained, enabling the initial music generation model to learn the basic patterns and structures in complete music clip samples, laying the foundation for subsequent reinforcement learning optimization. By introducing the GRPO algorithm and a unified audio aesthetics evaluation tool, the pre-trained model can be precisely optimized to obtain the target music generation model. The GRPO algorithm uses the aesthetic evaluation indicators of the unified audio aesthetics evaluation tool as reward indicators to guide the pre-trained model to adjust its generation strategy. This accurately measures the artistic value of the target music clip generated by the target music generation model, making the target music clip more harmonious and beautiful in terms of melody, harmony, and rhythm, significantly improving its quality and meeting aesthetic standards. Compared with traditional music generation methods, the target music clip generated by the target music generation model in the embodiments of the present invention is more attractive and appealing in terms of artistic and auditory effects. Moreover, the introduction of the GRPO algorithm not only improves the stability of the music generation process of the target music generation model and enhances the quality and complexity of the target music clip, but also reduces the reliance on large amounts of training data and reduces annotation costs. During the training process, through continuous learning and adjustment, the pre-trained model gradually converges to the optimal generation strategy, avoiding issues such as pattern collapse and stylistic deviations that can occur in traditional generation methods. Furthermore, the real-time scoring feedback from the unified audio aesthetics evaluation tool promptly identifies issues with the generated music, enabling timely adjustments and corrections to the pre-trained model, ensuring the stability and reliability of the resulting target music generation model.

[0043] On the basis of the above embodiment, the advantage value of each group of generation results is calculated based on the values ​​of each reward indicator corresponding to each group of generation results, including: Based on the values ​​of each reward indicator corresponding to each set of generated results, calculate the expected value of each reward indicator; Perform weighted summation on the expected values ​​of each reward indicator to obtain the total reward value of each group's generated results; Based on the total reward value, the average reward value and the reward value standard deviation of each group of generated results are calculated respectively, and based on the total reward value, the average reward value and the reward value standard deviation, the advantage value of each group of generated results is calculated.

[0044] Specifically, when calculating the advantage value of each group of generated results, we can first use the values ​​of each reward indicator corresponding to each group of generated results to calculate the expected value of each reward indicator, that is, Reward CE, Reward CU, Reward PC, Reward PQ=E[Audiobox-Aesthetics(Gen_Audio)]; Among them, Reward CE 、Reward CU 、Reward PC 、Reward PQ The expected values ​​for the content enjoyment index, content practicality index, production complexity index, and production quality index are shown in Figure 2. Gen_Audio represents each generated result, Audiobox-Aesthetics (Gen_Audio) represents the reward indicators corresponding to each generated result, and E represents the expectation.

[0045] Then, the expected value of each reward indicator is weighted and summed to obtain the total reward value of each group's generated results, that is: Reward=λ CE Reward CE +λ CU Reward CU +λ PC Reward PC +λ PQ Reward PQ ; Among them, Reward is the total reward value of each group’s generated results, λ CE ,λ CU ,λ PC ,λ PQ Reward CE 、Reward CU 、Reward PC 、Reward PQ The weight of .

[0046] After that, the total reward value is used to calculate the average reward value and reward value standard deviation of each group's generated results, and the total reward value, average reward value and reward value standard deviation are used to calculate the advantage value of each group's generated results. That is: A i =(Reward i -mean(Reward)) / std(Reward); Among them, A i The advantage value of generating the result for group i, Reward i The reward value for the generated result of group i can be obtained by generating the result Gen_Audio for group i. i The value of each reward indicator of Audiobox-Aesthetics (Gen_Audio i) is weighted summed up, mean(Reward) is the average reward value of the generated results of each group, and std(Reward) is the standard deviation of the reward value of the generated results of each group.

[0047] In an embodiment of the present invention, when calculating the advantage value of each group of generated results, the total reward value, average reward value and reward value standard deviation of each group of generated results are combined to calculate the advantage value of each group of generated results, which can ensure the rationality and accuracy of the advantage value and facilitate obtaining an accurate objective function.

[0048] Based on the above embodiment, the target music generation model includes a semantic feature extraction module, a feature conversion module and a vocoder; Inputting the input information into a target music generation model to obtain a target music segment output by the target music generation model that matches the input information includes: Inputting the input information into the semantic feature extraction module, and having the semantic feature extraction module output semantic features; Inputting the semantic feature into the feature conversion module, and having the feature conversion module output the acoustic feature corresponding to the semantic feature; The acoustic features are input into the vocoder to obtain the target music segment output by the vocoder.

[0049] Specifically, the target music generation model may include a semantic feature extraction module, a feature conversion module and a vocoder, the feature extraction module corresponds to the semantic modeling stage of the target music generation model, and the feature conversion module and the vocoder correspond to the acoustic modeling stage of the target music generation model.

[0050] When input information is input into the target music generation model, the input information can be input into a semantic feature extraction module, and the semantic feature extraction module outputs semantic features.

[0051] The semantic feature extraction module may include a self-supervised learning framework and a stream matching network. The self-supervised learning framework may be HuBERT, which is used to extract frame-level semantic features from the input music clip. The stream matching network is used to convert the input text into text semantic features.

[0052] When the user's input information is input text, the semantic features output by the semantic feature extraction module only include text semantic features. When the user's input information is input text and input music clips, the semantic features output by the semantic feature extraction module include frame-level semantic features and text semantic features.

[0053] The semantic features output by the semantic feature extraction module are input into the feature conversion module, which converts the input semantic features into low-level acoustic features. Here, the acoustic features may include detailed information such as volume and recording quality.

[0054] Finally, the acoustic features are input into the vocoder, which outputs the target music clip.

[0055] In the embodiment of the present invention, a specific processing process of the target music generation model is provided to achieve transparency of the input information processing process.

[0056] Based on the above embodiment, the step of inputting the input information into the target music generation model and obtaining the target music segment output by the target music generation model that matches the input information may then include: Obtaining a manual evaluation result of the target music clip; Based on the manual evaluation results, the generation strategy and structural parameters of the target music generation model are optimized.

[0057] Specifically, in embodiments of the present invention, after generating a target music clip, it can be periodically subjected to manual evaluation and professional testing. Music professionals and music enthusiasts are invited to perform auditory evaluations of the generated target music clip, providing scores and feedback based on aspects such as artistry, emotional expression, and innovation. Based on the manual evaluation results, the generation strategy and structural parameters of the target music generation model are further optimized, ensuring that the target music clips subsequently generated by the target music generation model better meet human aesthetic standards and actual listening needs.

[0058] Based on the above embodiment, the structure and parameters of the initial music generation model are determined based on music generation requirements.

[0059] Specifically, music generation requirements can include musical features such as 75 frames per second and 8 sound tokens per frame. Using these requirements, the structure and parameters of the initial music generation model, such as the number of neurons in the hidden layer, can be determined. This simplifies model initialization.

[0060] In summary, an embodiment of the present invention proposes a music generation method based on the GRPO algorithm. By combining the GRPO algorithm with the MusicFlow model and introducing Audiobox-Aesthetics as a tool for determining reward indicators, significant improvements in the quality and stability of music generation are achieved, providing an innovative, efficient and reliable solution for the field of music creation, which has broad application prospects and market value.

[0061] like Figure 2As shown, based on the above embodiment, an embodiment of the present invention provides a music generation device, including: An input acquisition module 21 is configured to acquire user input information, wherein the input information includes input text, or includes input text and input music clip; a music generation module 22 for inputting the input information into a target music generation model, and obtaining a target music segment output by the target music generation model that matches the input information; The training module 23 is used to train the target music generation model based on the following steps: Taking the masked first music clip sample and the first text description sample corresponding to the first music clip sample as input, and the complete music clip sample corresponding to the first music clip sample as a label, pre-training the initial music generation model to obtain a pre-trained model; Using the pre-trained model as a strategy model, and applying the strategy model to the second text description sample corresponding to each second music clip sample in the training sample set, and obtaining multiple groups of generation results based on a group relative strategy optimization algorithm; Based on a unified audio aesthetic evaluation tool, the values ​​of each reward indicator corresponding to each group of generation results are determined, and based on the values ​​of each reward indicator corresponding to each group of generation results, the advantage value of each group of generation results is calculated. Based on the advantage value of each group of generation results, the strategy model is trained to obtain the target music generation model.

[0062] Based on the above embodiment, in the music generation device provided in the embodiment of the present invention, the input music segment is masked, and the target music segment is a music segment filled with the mask prediction result; Alternatively, the input music segment and the target music segment are continuous in time.

[0063] On the basis of the above embodiment, in the music generation device provided in the embodiment of the present invention, the training module is specifically used for: Based on the values ​​of each reward indicator corresponding to each set of generated results, calculate the expected value of each reward indicator; Perform weighted summation on the expected values ​​of each reward indicator to obtain the total reward value of each group's generated results; Based on the total reward value, the average reward value and the reward value standard deviation of each group of generated results are calculated respectively, and based on the total reward value, the average reward value and the reward value standard deviation, the advantage value of each group of generated results is calculated.

[0064] On the basis of the above embodiment, in the music generation device provided in the embodiment of the present invention, the target music generation model includes a semantic feature extraction module, a feature conversion module and a vocoder; The music generation module is specifically used for: Inputting the input information into the semantic feature extraction module, and having the semantic feature extraction module output semantic features; Inputting the semantic feature into the feature conversion module, and having the feature conversion module output the acoustic feature corresponding to the semantic feature; The acoustic features are input into the vocoder to obtain the target music segment output by the vocoder.

[0065] Based on the above embodiment, the music generation device provided in the embodiment of the present invention further includes an optimization module for: Obtaining a manual evaluation result of the target music clip; Based on the manual evaluation results, the generation strategy and structural parameters of the target music generation model are optimized.

[0066] On the basis of the above embodiment, in the music generation device provided in the embodiment of the present invention, the training module is specifically used for: Determine the values ​​of the evaluation indicators for each group of generated results based on the unified audio aesthetics evaluation tool, and use the values ​​of the evaluation indicators for each group of generated results as the values ​​of the reward indicators corresponding to each group of generated results; The evaluation indicators include content enjoyment, content practicality, production complexity and production quality.

[0067] On the basis of the above-mentioned embodiment, in the music generation device provided in the embodiment of the present invention, the structure of the initial music generation model and its parameters are determined based on the music generation requirements.

[0068] Specifically, the functions of each module in the music generation device provided in the embodiment of the present invention correspond one-to-one to the operating procedures of each step in the above-mentioned method embodiment, and the effects achieved are also consistent. Please refer to the above-mentioned embodiment for details, and no further details will be given in the embodiment of the present invention.

[0069] Figure 3 An example of a physical structure diagram of an electronic device is shown below. Figure 3 As shown, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. The processor 810, the communications interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call logic instructions in the memory 830 to execute the music generation method provided in the above embodiments.

[0070] Furthermore, the logic instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the relevant art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0071] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the music generation method provided in the above embodiments.

[0072] In yet another aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program is used to implement the music generation methods provided in the aforementioned embodiments. The computer-readable storage medium may be either a non-transitory or a transient computer-readable storage medium, and is not specifically limited herein.

[0073] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0074] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A music generation method, characterized in that: include: Acquiring user input information, wherein the input information includes input text, or includes input text and input music clip; Inputting the input information into a target music generation model to obtain a target music segment output by the target music generation model that matches the input information; The target music generation model is trained based on the following steps: Taking the masked first music clip sample and the first text description sample corresponding to the first music clip sample as input, and the complete music clip sample corresponding to the first music clip sample as a label, pre-training the initial music generation model to obtain a pre-trained model; Using the pre-trained model as a strategy model, and applying the strategy model to the second text description sample corresponding to each second music clip sample in the training sample set, and obtaining multiple groups of generation results based on a group relative strategy optimization algorithm; Based on a unified audio aesthetic evaluation tool, the values ​​of each reward indicator corresponding to each group of generation results are determined, and based on the values ​​of each reward indicator corresponding to each group of generation results, the advantage value of each group of generation results is calculated. Based on the advantage value of each group of generation results, the strategy model is trained to obtain the target music generation model.

2. The music generation method according to claim 1, wherein: The input music clip is masked, and the target music clip is a music clip filled with the mask prediction result; Alternatively, the input music segment and the target music segment are continuous in time.

3. The music generation method according to claim 1, wherein: Determining the values ​​of each reward indicator corresponding to each set of generated results based on the unified audio aesthetics evaluation tool includes: Based on the values ​​of each reward indicator corresponding to each set of generated results, calculate the expected value of each reward indicator; Perform weighted summation on the expected values ​​of each reward indicator to obtain the total reward value of each group's generated results; Based on the total reward value, the average reward value and the reward value standard deviation of each group of generated results are calculated respectively, and based on the total reward value, the average reward value and the reward value standard deviation, the advantage value of each group of generated results is calculated.

4. The music generation method according to claim 1, wherein: The target music generation model includes a semantic feature extraction module, a feature conversion module and a vocoder; Inputting the input information into a target music generation model to obtain a target music segment output by the target music generation model that matches the input information includes: Inputting the input information into the semantic feature extraction module, and having the semantic feature extraction module output semantic features; Inputting the semantic feature into the feature conversion module, and having the feature conversion module output the acoustic feature corresponding to the semantic feature; The acoustic features are input into the vocoder to obtain the target music segment output by the vocoder.

5. The music generation method according to any one of claims 1 to 4, characterized in that: The step of inputting the input information into a target music generation model to obtain a target music segment output by the target music generation model that matches the input information comprises: Obtaining a manual evaluation result of the target music clip; Based on the manual evaluation results, the generation strategy and structural parameters of the target music generation model are optimized.

6. The music generation method according to any one of claims 1 to 4, characterized in that: Determining the values ​​of each reward indicator corresponding to each set of generated results based on the unified audio aesthetics evaluation tool includes: Determine the values ​​of the evaluation indicators for each group of generated results based on the unified audio aesthetics evaluation tool, and use the values ​​of the evaluation indicators for each group of generated results as the values ​​of the reward indicators corresponding to each group of generated results; The evaluation indicators include content enjoyment, content practicality, production complexity and production quality.

7. The music generation method according to any one of claims 1 to 4, characterized in that: The structure and parameters of the initial music generation model are determined based on music generation requirements.

8. A music generating device, characterized in that: include: An input acquisition module, configured to acquire user input information, wherein the input information includes input text, or includes input text and input music clip; a music generation module, configured to input the input information into a target music generation model, and obtain a target music segment output by the target music generation model that matches the input information; The training module is used to train the target music generation model based on the following steps: Taking the masked first music clip sample and the first text description sample corresponding to the first music clip sample as input, and the complete music clip sample corresponding to the first music clip sample as a label, pre-training the initial music generation model to obtain a pre-trained model; Using the pre-trained model as a strategy model, and applying the strategy model to the second text description sample corresponding to each second music clip sample in the training sample set, and obtaining multiple groups of generation results based on a group relative strategy optimization algorithm; Based on a unified audio aesthetic evaluation tool, the values ​​of each reward indicator corresponding to each group of generation results are determined, and based on the values ​​of each reward indicator corresponding to each group of generation results, the advantage value of each group of generation results is calculated. Based on the advantage value of each group of generation results, the strategy model is trained to obtain the target music generation model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the music generation method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the music generation method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Training method and training device of music generation model, storage medium and equipment

    CN115206269A

  • Music generation method and system, electronic equipment and medium

    CN116682399A

  • Song generation method and device, electronic equipment and storage medium

    CN116895266A

  • Text and audio alignment model construction method, method and device for generating music from text, equipment, medium and program product

    CN118800205A

  • Music generation method and device, electronic equipment and storage medium

    CN119400134A