Music generation method, music generation apparatus, and computer-readable storage medium

By providing guidance and adjustment functions on the user-interaction interface with the intelligent agent, the problem of music generated by machine learning models failing to meet user needs has been solved, enabling high-quality automatic music generation and adjustment, and improving the user experience.

WO2025222413A1PCT designated stage Publication Date: 2025-10-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/089617
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-24
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

In existing technologies, music generated by machine learning models cannot meet users' needs due to limitations in model performance and the accuracy of users' textual descriptions of the music.

Method used

A music generation method is provided, in which a user interacts with an intelligent agent and displays guiding information. The user generates a first musical work based on the prompts and can adjust it to generate a second musical work. This includes adjusting the basic attributes, style, type, lyrics, and vocals of the music.

Benefits of technology

It improves the quality of music generation and user experience, meets users' personalized needs, and achieves high-quality automatic music generation and adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024089617_30102025_PF_FP_ABST
    Figure CN2024089617_30102025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of computers, and relates to a music generation method, a music generation apparatus, and a computer-readable medium. The music generation method comprises: receiving first prompt information inputted by a user, wherein the first prompt information comprises description information of music; on the basis of the first prompt information, generating a first music work, wherein the first music work comprises first music; and on the basis of an adjustment operation of the user for the first music work, generating a second music work, wherein the adjustment operation comprises at least one of an adjustment operation for the first prompt information and an adjustment operation for the first music.
Need to check novelty before this filing date? Find Prior Art

Description

Music generation method, music generation device and computer-readable storage medium Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a music generation method, a music generation apparatus, a computer-readable storage medium, and a computer program product. Background Technology

[0002] Music generation technology is an important application in the field of computer technology. With the rapid development of artificial intelligence technology, natural language processing technology is widely used for music generation, utilizing machine learning models to automatically generate music data from text data.

[0003] However, due to limitations in the performance of machine learning models and the accuracy of users' textual descriptions of music, automatically generated music often fails to meet users' needs.

[0004] Summary of the Invention

[0005] In view of this, embodiments of this disclosure propose a music generation method, a music generation apparatus, a computer-readable storage medium, and a computer program product. After automatically generating music based on a text description, the user can make targeted adjustments to the generated music under the guidance of an intelligent agent, and regenerate the music based on the user's adjustments, thereby improving the quality of music generation and enhancing the user experience.

[0006] According to a first aspect of some embodiments of this disclosure, a music generation method is provided, comprising: displaying music generation guidance information on an interactive interface between a user and an intelligent agent in response to a user's music generation request; outputting a first musical work generated according to the first prompt information in response to the user's input of the guidance information, wherein the first prompt information includes music description information and the first musical work includes first music; and outputting a second musical work generated according to the user's adjustment of the first musical work in response to the user's adjustment, wherein the adjustment includes at least one of adjustment to the first prompt information and adjustment to the first music.

[0007] According to a second aspect of some embodiments of this disclosure, a music generation apparatus is provided, comprising: a display module configured to display guidance information for music generation on an interactive interface between a user and an intelligent agent in response to a user's music generation request; a first output module configured to output a first musical work generated based on the first prompt information input by the user based on the guidance information, wherein the first prompt information includes descriptive information of the music and the first musical work includes first music; and a second output module configured to output a second musical work generated based on an adjustment made by the user to the first musical work, wherein the adjustment includes at least one of adjustments to the first prompt information and adjustments to the first music.

[0008] According to a third aspect of some embodiments of the present disclosure, a music generation apparatus is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute any of the foregoing music generation methods based on instructions stored in the memory.

[0009] According to a fourth aspect of some embodiments of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements any of the aforementioned music generation methods.

[0010] According to a fifth aspect of some embodiments of the present disclosure, a computer program product is provided that, when the computer program product is run on a computer, causes the computer to implement any of the aforementioned music generation methods.

[0011] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0012] Other features and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0013] Preferred embodiments of the present disclosure are described below with reference to the accompanying drawings. The accompanying drawings, which are included to provide a further understanding of the present disclosure, and which, together with the following detailed description, are incorporated in and form a part of this specification and are used to explain the present disclosure. It should be understood that the drawings described below only relate to some embodiments of the present disclosure and are not intended to limit the present disclosure. In the drawings:

[0014] Figure 1 shows a flowchart of a music generation method according to some embodiments of the present disclosure;

[0015] Figure 2A shows an interactive interface for music generation according to some embodiments of the present disclosure;

[0016] Figure 2B shows an interactive interface for music generation according to other embodiments of the present disclosure;

[0017] Figure 2C shows an interactive interface for music generation according to some embodiments of the present disclosure;

[0018] Figure 3A shows the interactive interface after generating the first music according to some embodiments of the present disclosure;

[0019] Figure 3B shows the interactive interface after generating the first music according to some other embodiments of the present disclosure;

[0020] Figure 3C shows the interactive interface after generating the first music according to some embodiments of the present disclosure;

[0021] Figure 4A illustrates an interactive interface adapted for a musical work according to some embodiments of the present disclosure;

[0022] Figure 4B illustrates an interactive interface adapted for a musical work according to other embodiments of this disclosure;

[0023] Figure 4C shows an interactive interface adjusted for a musical work according to some embodiments of the present disclosure;

[0024] Figure 4D illustrates an interactive interface adjusted for a musical work according to some embodiments of the present disclosure;

[0025] Figure 5A illustrates a flowchart of generating a second musical work according to some embodiments of the present disclosure;

[0026] Figure 5B illustrates a flowchart of generating a second musical work according to some other embodiments of the present disclosure;

[0027] Figure 6A illustrates a flowchart of generating a second musical work based on editing controls according to some embodiments of the present disclosure;

[0028] Figure 6B illustrates a flowchart of generating a second musical work based on an editing control according to some other embodiments of the present disclosure;

[0029] Figure 7A shows an editing interface for a musical work according to some embodiments of the present disclosure;

[0030] Figure 7B shows an editing interface for a musical work according to other embodiments of the present disclosure;

[0031] Figure 8 shows a flowchart of a music generation method according to some other embodiments of the present disclosure;

[0032] Figure 9 shows a block diagram of a music generation apparatus according to some embodiments of the present disclosure;

[0033] Figure 10 shows a block diagram of a music generation apparatus according to other embodiments of the present disclosure;

[0034] Figure 11 shows a block diagram of an electronic device according to some embodiments of the present disclosure.

[0035] It should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not necessarily drawn to actual scale. The same or similar reference numerals are used in the various drawings to denote the same or similar parts. Therefore, once an item is defined in one drawing, it may not be discussed further in subsequent drawings. Detailed Implementation

[0036] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. However, it is obvious that the described embodiments are only some embodiments of this disclosure, and not all embodiments. The following description of the embodiments is merely illustrative and is in no way intended to limit this disclosure or its application or use. It should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.

[0037] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of components and steps set forth in these embodiments should be interpreted as merely exemplary and do not limit the scope of this disclosure.

[0038] As used in this disclosure, the term "comprising" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". Furthermore, as used in this disclosure, the term "including" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". Therefore, "comprising" and "including" are synonymous. The term "based on" means "at least partially based on".

[0039] Throughout this specification, the terms "one embodiment," "some embodiments," or "embodiment" mean that a specific feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of the invention. For example, the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; and the term "some embodiments" means "at least some embodiments." Furthermore, the appearance of the phrases "in one embodiment," "in some embodiments," or "in an embodiment" in various places throughout the specification does not necessarily refer to the same embodiment, but may refer to the same embodiment.

[0040] It should be noted that the concepts of "first," "second," etc., used in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to define the order of functions performed by these devices, modules, or units or their interdependencies. Unless otherwise specified, the concepts of "first," "second," etc., are not intended to imply that the objects described herein must be in a given temporal, spatial, rank, or any other given order.

[0041] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0042] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0043] The embodiments of this disclosure are described in detail below with reference to the accompanying drawings; however, this disclosure is not limited to these specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. Furthermore, in one or more embodiments, specific features, structures, or characteristics can be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.

[0044] In existing music generation technologies, the automatically generated music often fails to meet user needs due to limitations in the performance of machine learning models and the accuracy of users' textual descriptions of the music. To improve the quality of music generation and enhance user experience, this disclosure provides a novel music generation method. After automatically generating music based on textual descriptions, users can make targeted adjustments to the generated music under the guidance of an intelligent agent, and the music can be regenerated based on the user's adjustments.

[0045] Figure 1 shows a schematic flowchart of a music generation method according to some embodiments of the present disclosure.

[0046] As shown in Figure 1, the music generation method includes: Step S1, on the user-smart agent interaction interface, in response to the user's music generation request, displaying music generation guidance information; Step S2, in response to the user's input of first prompt information based on the guidance information, outputting a first musical work generated according to the first prompt information, wherein the first prompt information includes music description information and the first musical work includes first music; Step S3, in response to the user's adjustment to the first musical work, outputting a second musical work generated according to the adjustment, wherein the adjustment includes at least one of adjustment to the first prompt information and adjustment to the first music.

[0047] The music generation method in this embodiment can be executed on the client side or partially on the server side.

[0048] The first musical work can be a purely instrumental piece, without lyrics, consisting only of the first note, such as a piano piece or guitar piece. It can also be a vocal piece, including not only the first note but also the accompanying vocals. Vocal pieces usually also include lyrics, sung by the vocals.

[0049] The music generation method according to embodiments of the present disclosure will be further described below with reference to Figures 2A-2C. Figures 2A-2C respectively show the interactive interface for music generation according to different embodiments of the present disclosure.

[0050] As shown in Figure 2A, an input box is provided in the music generation interactive interface. In response to the user's music generation command entered in the input box via text or voice, guidance information for music generation is displayed on the user-agent interaction interface. Besides invoking the music generation model by inputting a music generation command, the user can also send a music generation request by triggering the music generation control shown at the bottom of the interactive interface in Figure 2A.

[0051] The music generation prompts, such as those shown in Figure 2A, ask users to describe the music they want to create. They can try describing the scene before them and their feelings, such as a California sunset, coconut trees, a sense of movement, a humid rainy season, or love and loss. Users can then input their own descriptions. These descriptions can include at least one of the following: the theme, the imagery, the emotional tone, or the genre. Examples of themes include homesickness, remembrance of friends, and celebration of success. Examples of imagery include sunsets, waves, rocks, streets, and buildings. Examples of emotional tones include sadness, joy, helplessness, emotion, and tranquility. Examples of genres include rock, jazz, pop, folk, and rap.

[0052] As shown in Figure 2A, users can input descriptions of music via text or voice. For example, "In the afterglow of the setting sun, the waves crash against the rocks, and the gentle singing of fishing boats can be heard in the distance. It is peaceful," which describes the imagery and emotional type of the music.

[0053] Of course, to help users input more efficient and accurate prompts, some suggestions can be provided on the interactive interface to facilitate user selection.

[0054] Figure 2B shows another interactive interface for music generation. Figure 2B also shows guidance information for music generation to guide the user to input descriptive information about the music.

[0055] As shown in Figure 2B, the guidance information includes at least one candidate prompt template. Each candidate prompt template corresponds to a complete music scene, and the information of each candidate prompt template includes at least one of the following: music theme, music visuals, music emotional type, and music style. The prompt information in a candidate prompt template selected by the user serves as part or all of the first prompt information.

[0056] As shown in Figure 2B, the interactive interface displays multiple candidate prompt templates for the user to choose from. For example, one candidate prompt template might include information such as "In the afterglow of the setting sun, waves crash against the rocks, and the soft singing of fishing boats drifts from afar; tranquility," reflecting the imagery and emotional tone of the music. Another candidate prompt template might include information such as Li Bai's "Drinking Alone Under the Moon," a rap piece, reflecting the theme and style of the music.

[0057] Since each candidate prompt template corresponds to a complete music scene, users can only select one. The prompt information from the selected candidate prompt template can be directly applied to the input box as the first prompt information. As shown in Figure 2A, if the user selects the candidate prompt template "In the afterglow of the setting sun, the waves crash against the rocks, and the soft singing of fishing boats can be heard in the distance, peaceful," then the corresponding prompt information will be directly applied to the input box.

[0058] Of course, users can further adjust the prompts based on the prompts in the candidate prompt templates, such as deleting or adding prompts in the input box. In this case, the prompts in a candidate prompt template selected by the user are only a part of the first prompt.

[0059] In other embodiments, as shown in FIG2C, the guidance information includes at least one candidate prompt, each candidate prompt including at least one of the following: music theme, music visuals, music emotional type, and music style; the one or more candidate prompts selected by the user serve as part or all of the first prompt.

[0060] As shown in Figure 2C, multiple candidate prompts are displayed on the interactive interface for the user to choose from. For example, the candidate prompts may include prompts reflecting the theme, imagery, emotional type, and style of the music, such as longing, sunset glow, waves, reefs, fishing boats, cheerful, tranquil, rock, and folk.

[0061] Users can select multiple candidate prompts and combine them. The selected prompts can be directly applied to the input box as the first prompt. For example, as shown in Figure 2C, users can select multiple candidate prompts such as "Sunset Glow," "Waves," "Reefs," "Fishing Boats," and "Tranquility." Users can also further adjust the prompts in the input box. For example, using [presumably a tool or method], the selected prompt combination "Sunset Glow, Waves, Reefs, Fishing Boats, Tranquility" can be modified to "In the sunset glow, waves crash against the reefs, and the gentle singing of fishing boats can be heard in the distance; tranquility." In this case, the multiple candidate prompts selected by the user are only a part of the first prompt.

[0062] The above, with reference to Figures 2A-2C, describes how, in step S1, guidance information for music generation is displayed in response to the user's music generation request, and how the first prompt information input by the user is received. The following, with reference to Figures 3A-4D, describes how, in step S2, the first musical piece is output based on the first prompt information, and how, in step S3, the second musical piece is output based on the user's adjustments to the first musical piece.

[0063] In step S2, a machine learning model can be used to generate the first music based on the first prompt information. The first music piece also includes a title and cover art. Similarly, the title and cover art can be generated using a machine learning model based on the first prompt information. As shown in Figures 2B and 2C, the S-style painting control at the bottom of the interactive interface can also be triggered to generate a cover art in the corresponding style. The S-style painting control can be configured based on user behavior or hotspot images.

[0064] Machine learning models include, for example, Large Language Models (LLMs) or other natural language processing (NLP) models. For instance, machine learning models such as Variational Autoencoders (VAEs), Generative Adversarial Networks (GANs), and Transformers can be used to achieve diverse and high-quality music generation. Figure 3A illustrates an interactive interface after generating the first piece of music according to some embodiments of this disclosure.

[0065] The interactive interface shown in Figure 3A outputs the generated first musical piece, including playing the first piece of music and displaying its title and cover art. For example, the title of the music is displayed as "Sunset Glow and Waves," and the cover art can be a picture of a sunset over the sea.

[0066] Figure 3A also shows guidance messages that help users adjust the generated music, such as "Here is the music generated for you. You can try providing a richer description to get new music." Users can regenerate music by changing the prompt message.

[0067] For example, users can select from the suggestions provided on the interactive interface shown in Figure 3A to trigger a regenerated control, which will then regenerate the music based on the new suggestion information. Alternatively, users can directly enter new suggestion information or modify the initial suggestion information in the input box.

[0068] Figure 3B shows the interactive interface after generating the first piece of music according to some other embodiments of the present disclosure. As shown in Figure 3B, based on the first prompt message "In the afterglow of the setting sun, the waves crash against the rocks, and the soft singing of fishing boats can be heard in the distance, peaceful," the user adds prompt message describing the style of the music, such as "pop music." That is, the user regenerates the music by entering the command "Music generation: In the afterglow of the setting sun, the waves crash against the rocks, and the soft singing of fishing boats can be heard in the distance, peaceful, pop music" on the interactive interface shown in Figure 3B. Figure 3B also shows some recommended prompt messages for the user to choose from.

[0069] Figure 3C shows an interactive interface after generating the first music according to some embodiments of the present disclosure. Similar to Figures 3A and 3B, the interactive interface shown in Figure 3C also outputs the generated first music work, including playing the first music, displaying the title and cover of the first music. Unlike Figures 3A and 3B, Figure 3C also shows some recommended adjustment information, including adjusting the basic attributes and / or style of the music, or adding lyrics.

[0070] As shown in Figure 3C, the interface for interacting with the intelligent agent displays recommended adjustment information such as "Faster tempo," "Change song style," and "Add lyrics." This recommended adjustment information can be configured as label-style controls. The user's selection of these controls is equivalent to inputting the corresponding recommended adjustment information.

[0071] In some embodiments, responding to user input regarding adjustments to the first musical work, outputting a second musical work generated based on the adjustments includes: receiving adjustment information from the user regarding basic attributes and / or style type of the first music, the basic attributes including rhythm, melody, timbre, and harmony; and outputting the second musical work generated based on the adjustment information, the second musical work comprising a second piece of music.

[0072] Figure 4A illustrates an interactive interface for adjusting a musical work according to some embodiments of the present disclosure. For example, a user selects the recommended adjustment information "faster tempo" on the interactive interface shown in Figure 3C to adjust the tempo of the first piece of music to obtain the second piece. Next, the user returns to the second musical work with the adjusted tempo on the user-agent interactive interface, as shown in Figure 4A. The second musical work also includes a title and cover art, which may be the same as or different from the title and cover art of the first piece of music.

[0073] Figure 4A also shows further guidance information, such as "The following is the music with the tempo adjusted for you. If you feel the tempo is not fast enough, you can continue to modify it using the following labels." For example, users can continue to select the recommended adjustment information "Faster tempo" to further adjust the tempo of the music.

[0074] Figure 4B illustrates an interactive interface for adjusting a musical work according to other embodiments of this disclosure. For example, a user selects the recommended adjustment information "Change Song Style" on the interactive interface shown in Figure 3C to adjust the song style of the first piece of music. Next, on the user's interactive interface with the intelligent agent, further guidance information can be returned, such as "Okay! Please describe the song style you want to me, or select below" as shown in Figure 4B.

[0075] Figure 4B shows some recommended song styles, such as pop, dance, rap, rock, folk, traditional Chinese style, ballad, electronic, jazz, etc., for users to choose from.

[0076] For example, on the interactive interface shown in Figure 4B, the user selects a recommended song style, such as folk, to change the song style from pop music, as shown in Figure 3, to folk music, thus obtaining a second piece of music. Next, the user returns to the second piece of music with adjusted rhythm on the user-agent interaction interface, as shown in Figure 4B.

[0077] Of course, if the first generated music is instrumental, you can also add lyrics and vocals to make adjustments first.

[0078] In some embodiments, responding to user input regarding adjustments to the first musical work, outputting a second musical work generated based on the adjustments includes: receiving adjustment information from the user regarding adding lyrics to the first music; and outputting first lyrics generated based on the first prompt information and the first music as part of the second musical work.

[0079] The following description, in conjunction with Figures 4C and 5A, further illustrates how to adjust musical works based on the user's adjustments made by adding lyrics to the music.

[0080] Figure 4C illustrates an interactive interface for adjusting a musical work according to some embodiments of the present disclosure. As shown in Figure 4C, in response to receiving a user input instruction to add lyrics, first lyrics corresponding to the first music are output. In some embodiments, a machine learning model is used to generate the first lyrics based on first prompt information and the first music.

[0081] As shown in Figure 4C, after generating the first set of lyrics, a guide message is displayed to help the user modify the lyrics, such as "Okay! Lyrics have been generated based on your description, and you can make changes now." The generated lyrics are displayed in a text box at the bottom of the interactive interface, where the user can further modify them.

[0082] Figure 5A illustrates a flowchart of generating a second musical work according to some embodiments of the present disclosure. As shown in Figure 5A, in response to user input regarding adjustments to the first musical work, outputting a second musical work generated based on the adjustments includes: step S31A, receiving adjustment information from the user regarding adding lyrics to the first music; step S32A, outputting first lyrics generated based on the first prompt information and the first music; step S33A, receiving adjustment information from the user regarding the first lyrics; and step S34A, displaying second lyrics obtained based on the adjustment information as part of the second musical work.

[0083] Adjustments to the initial lyrics can also be made using machine learning models, such as the AI-assisted control shown at the bottom of the text box in Figure 4C. Responding to the user triggering the AI-assisted control, a second set of lyrics matching the music can be generated based on the initial lyrics and the user's input, using a machine learning model. That is, by comprehensively considering the initial music, the initial lyrics, and the user's input, a new set of lyrics is generated using a machine learning model to better match the music. For vocal music, corresponding vocals are also required, including the singer's voice. The following sections, with reference to Figures 4D and 5B, further describe how to adjust the musical work based on the user's adjustments to the music by adding vocals.

[0084] Figure 5B illustrates a flowchart of generating a second musical work according to some other embodiments of the present disclosure. As shown in Figure 5B, in response to adjustments made to the first musical work by the user, outputting a second musical work generated according to the adjustments includes: step S31B, receiving adjustment information from the user regarding adding vocals to the first music; step S32B, providing the timbres of multiple singers according to the adjustment information; and step S33B, outputting vocals generated according to the timbres of one or more singers selected by the user as part of the second musical work.

[0085] Figure 4D illustrates an interactive interface for adjusting a musical work according to some embodiments of the present disclosure. Similar to adding lyrics in Figure 4C, as shown in Figure 4D, in response to receiving a user's instruction to add vocals, further guidance information can be returned on the user-agent interaction interface, such as "Okay! Please describe the vocals you want to me, or select below" as shown in Figure 4D.

[0086] Users can use a default vocal that matches the timbre of the music and its lyrics, based on the user's previous settings. Of course, users can also choose other vocals as needed. For example, users can choose their own voice or select the voice of another singer based on their preference.

[0087] When a user uses their own voice, the singer's voice can be generated using voice imitation technology. Alternatively, if the client or server has already stored the user's voice, it can be directly accessed with the user's authorization.

[0088] For users using other vocals, a machine learning model is used to generate the singer's voice based on the initial music and lyrics. Of course, if the vocal music does not include lyrics, the singer's voice can also be generated solely from the initial music. The singer's voice generated using the machine learning model can be of various types.

[0089] As shown in Figure 4D, the interactive interface for adjusting music tracks displays a variety of sounds for the user to choose from. Sound A in Figure 4D could be, for example, the voice of a poet; sound B could be, for example, the voice of an actor; sound C could be, for example, the voice of a character in a movie or TV series; and the default vocal could be, for example, the voice of a singer.

[0090] As shown in Figure 4D, users can select the singer's voice according to their actual needs, or they can input a description of the voice to obtain the corresponding voice.

[0091] The above, combined with Figures 3A-3C and 4A-4D, describes how a second musical piece is output based on the user's prompts, basic attributes, style type, added lyrics, and added vocal adjustments for the first piece of music. In addition to adjusting the generated musical piece through interaction with the intelligent agent, the user can also directly access the music editing interface.

[0092] The interfaces shown in Figures 3A-3C and 4A-4D display the generated music tracks, along with controls for editing, publishing, and sharing. The following section, using Figures 3C, 6A-6B, and 7A-7B, further describes how to output a second music track based on the user's adjustments to the first track.

[0093] Figure 6A shows a flowchart of generating a second musical work based on editing controls according to some embodiments of the present disclosure.

[0094] As shown in Figure 6A, the process of outputting a second musical work generated based on the user's input adjustment of the first musical work includes: step S31, displaying the editing control of the first musical work; step S32, entering the editing interface for the first musical work in response to the user's triggering of the editing control; and step S33, generating the second musical work based on the user's adjustment operation for the first musical work in the editing interface, wherein the adjustment operation includes at least one of the adjustment operation for the first prompt information and the adjustment operation for the first musical work.

[0095] Figure 3C shows the editing controls for the first musical piece. In response to the user's triggering of the editing controls in Figure 3C, a separate music creation interface is entered, as shown in Figure 7A.

[0096] Figure 7A illustrates an editing interface for a musical work according to some embodiments of the present disclosure, supporting fine-tuning of the musical work. As shown in Figure 7A, the input box for generating music prompt information displays previously entered prompt information, i.e., the first prompt information for generating the first piece of music. Users can modify the first prompt information based on this.

[0097] Figure 6B illustrates a flowchart of generating a second musical work based on an editing control according to other embodiments of the present disclosure. As shown in Figure 6B, generating the second musical work based on the user's adjustment operation on the first musical work in the editing interface includes: step S331, receiving the user's modification operation on the first prompt information; step S332, displaying a second prompt information obtained based on the modification operation; and step S333, generating the second musical work using a machine learning model based on the second prompt information.

[0098] In addition to displaying the initial prompts used to generate the first piece of music, further prompts can be recommended based on the generated first piece of music, i.e., suggestions can be given to help the user adjust the first piece of music. As shown in Figure 7A, the editing interface displays at least one recommended prompt based on the first piece of music. The recommended prompt includes at least one of the following: the theme of the music, the visuals of the music, the emotional type of the music, and the style of the music.

[0099] Figure 7A shows some recommended prompts, such as those related to music genres like pop, traditional Chinese style, and folk. Figure 7A also shows recommended prompts related to music themes, such as longing and nostalgia, as well as prompts related to music imagery, such as fishermen.

[0100] For example, after generating the first piece of music based on the initial prompt information such as "In the afterglow of the setting sun, the waves crash against the rocks, and the gentle singing of fishing boats can be heard in the distance, a peaceful scene," further prompt information can be recommended based on the visual imagery of the first piece of music ("In the afterglow of the setting sun, the waves crash against the rocks, and the gentle singing of fishing boats can be heard in the distance") and the emotional type of the music ("peaceful"). Examples include prompts related to themes such as "longing" and "nostalgia" (as shown in Figure 7A), styles such as "pop," "folk," and "traditional Chinese style," and visuals such as "fishermen." Users can select one or more of these recommended prompts, which can then be used as part of the second set of prompt information.

[0101] In steps S31 and S32, the user can modify the first prompt information to obtain the second prompt information. For example, if the first prompt information indicates the music genre is pop, the user can change it to folk to obtain the second prompt information. Similarly, if the first prompt information indicates the music's emotional type is melancholic, the user can change it to resignation to obtain the second prompt information. Likewise, the user can modify the music's theme, visuals, etc., in the first prompt information to obtain the second prompt information.

[0102] Of course, users can also delete or add descriptive information about the music to the first prompt message to obtain the second prompt message. For example, if the first prompt message includes the theme of celebrating success, the visual imagery of architecture, and the mood of cheerfulness, users can delete the visual description, resulting in a second prompt message with less descriptive information, reducing restrictions on the generated visuals. Similarly, in this case, users can add descriptive information about the music's genre, such as "pop," resulting in a second prompt message with more descriptive information, improving the match between the generated music and the user's needs.

[0103] In step S33, a second musical work that better meets the user's needs is generated by using a machine learning model and adjusting the first prompt information according to the aforementioned second prompt information.

[0104] It should be understood that user adjustments to the first piece of music can be made to the first prompt message, the first piece of music itself, or both. Adjustments to the first piece of music can include changes to the music's basic attributes.

[0105] In some embodiments, generating a second musical work based on the user's adjustment operation on the first musical work includes: generating a second musical work based on the user's adjustment operation on the basic attributes of the first musical work, wherein the basic attributes include the rhythm, melody, timbre, and harmony of the music.

[0106] If the user is not satisfied with the rhythm of the generated first track, they can adjust it individually. The rhythm of music generally includes beat, tempo, accents, and style. Users can adjust at least one of these factors to modify the rhythm of the first track.

[0107] Figure 7B illustrates an editing interface for a musical work according to other embodiments of the present disclosure. The editing interface shown in Figure 7B includes tempo controls. In response to the user triggering the tempo controls, a tempo editing interface pops up. As shown in Figure 7B, the user can move the tempo adjustment controls to adjust the tempo of the music.

[0108] Similarly, users can adjust the melody, timbre, harmony, etc. of the first piece of music based on their own understanding and preferences to obtain a second piece of music that better meets their needs.

[0109] Adjusting the basic attributes of the first piece of music allows for overall adjustments to the first piece of music. Since music can be divided into sections, adjustments can also be made to the first piece of music in sections.

[0110] For example, the first music includes multiple sections, and in response to the user's input of adjustments to the first music piece, outputting a second music piece generated according to the adjustments includes: generating the second music piece based on the user's adjustment operation on the basic attributes of at least one of the multiple sections of the first music piece, the basic attributes including the music's rhythm, melody, timbre, and harmony.

[0111] You can adjust the basic properties of only certain sections of the first music track, such as adjusting the rhythm of only the musical section corresponding to the intro of the first music track, without adjusting other parts of the first music track.

[0112] The rhythm adjustment of the first piece can be further refined. For example, each section of the music has multiple beats, and users can adjust the rhythm of the first piece by beat, which can more efficiently generate music works that match the user's needs.

[0113] For example, each of the multiple sections of the first music includes multiple beats, and the rhythm of the first music includes the beats of the music. In response to the user's input adjustment to the first music piece, outputting a second music piece generated according to the adjustment includes: displaying the audio track of the first music; generating the second music piece according to the user's adjustment operation on the first music based on the audio track, wherein the adjustment operation on the first music includes at least one of adjusting the time value of the beats and adjusting the combination rules between beats.

[0114] One audio track corresponds to one part of a piece of music. Each audio track can correspond to the performance of one instrument. The first piece of music may include multiple audio tracks. Displaying the audio tracks of the first piece of music allows users to adjust them based on the tracks. That is, based on the user's adjustment operation on at least one of the multiple audio tracks of the first piece of music, a second piece of music can be generated.

[0115] For example, in the multiple tracks of the first piece of music, the first track corresponds to the drum performance, giving the music its rhythmic backbone; the second track corresponds to the bass rhythm, complementing the drum performance; the third track corresponds to the guitar playing in sync with the beat; and the fourth track corresponds to the vocals accompanying the music. Users can adjust only some tracks instead of all of them. For instance, a user can adjust only the guitar playing in the third track. Specifically, they can adjust the duration of each beat (i.e., note time value) in the third track, as well as the combination of strong and weak beats, or both.

[0116] As described in detail above with reference to different embodiments, user adjustments to the first musical work include both overall adjustments, such as adjustments to the first prompt information used to generate the musical work, and partial adjustments, such as adjustments to the basic attributes of the generated music, such as rhythm and melody. Furthermore, the musical work includes both music and lyrics. Because the music and lyrics have a certain degree of independence, they can be adjusted separately to efficiently generate musical works that better meet the user's needs.

[0117] Several embodiments of the music generation method have been described in detail above with reference to Figures 1-7B. A specific application example of the music generation method is described below with reference to Figure 8. Figure 7 shows a flowchart of a music generation method according to other embodiments of this disclosure.

[0118] As shown in Figure 8, taking a musical work with vocals as an example, the music generation method includes: step S1', on the user-intelligent interaction interface, in response to the user's music generation request, displaying music generation guidance information; step S2', according to the first prompt information input by the user, outputting a first musical work, the first musical work including the first music and its title and cover; step S3', according to the user's adjustments to the first musical work, outputting a second musical work, the second musical work including the second music and its title and cover, lyrics, and vocals.

[0119] Step S1' is similar to step S1 in Figure 1. At least one prompt message suggestion can be provided on the interactive interface for the user to choose as part or all of the first prompt message.

[0120] The suggested prompts can be candidate prompt templates, as shown in Figure 2B, and users can only choose one. For example, a user could select the candidate prompt template "In the afterglow of the setting sun, waves crash against the rocks, and the gentle singing of fishing boats drifts from afar; tranquility reigns." The prompt information included in this template can then be directly applied to the prompt input box. Users can directly use this prompt information as the first prompt for generating the first piece of music. Of course, users can also adjust the prompt information based on this, such as adding descriptive information about the music's genre, such as "pop," to obtain a first prompt with more descriptive information about the music.

[0121] The suggested prompts can also be candidate prompts, as shown in Figure 2C. Users can select multiple candidate prompts and combine them. For example, users can select multiple candidate prompts such as "Sunset Glow," "Waves," "Reefs," "Fishing Boats," and "Tranquility." Similarly, these prompts can be directly applied to the prompt input box. Users can directly use these prompts as the first prompt for generating the first piece of music. Of course, users can also adjust the prompts based on this, for example, by adding descriptive information about the music's visuals, such as "Fisherman," to obtain a first prompt with more descriptive information about the music.

[0122] Next, in step S2', a machine learning model is used to generate the first piece of music based on the first prompt information. Considering the display requirements of the musical work, a title and cover art corresponding to the first piece of music can also be generated using the machine learning model based on the first prompt information. Alternatively, the cover art can be generated based on the controls describing the music visuals as shown in Figures 2B or 2C. Then, the first piece of music, its title, and cover art are combined to form the first musical work.

[0123] Next, in step S3', a second musical piece is generated using a machine learning model based on the user's adjustments to the first musical piece.

[0124] Adjustments to the first musical piece can include changes to the initial prompt information used to generate the first piece, adjustments to the basic attributes of the first piece, adjustments to adding lyrics or vocals, and possibly adjustments to the title or cover art.

[0125] Adjustments to the first musical piece can be made either based on user-inputted adjustment information or based on user adjustments made on the music piece's editing interface.

[0126] Regarding adjustments to the prompts, previously entered prompts can be displayed in the input box, allowing users to edit them accordingly. Additionally, as shown in Figure 3B, recommended prompts based on the first music track can be displayed on the interactive interface to help users adjust the first music track more effectively.

[0127] Of course, users can also access the music editing interface, as shown in Figure 7A, and adjust the prompts in the editing interface. The editing interface shown in Figure 7A will also display recommended prompts.

[0128] For example, the first piece of music is generated based on the first prompt message "Birds joyfully embrace the sky with their wings." In the interactive interface shown in Figure 3B or the editing interface shown in Figure 7A, prompt messages such as the music's emotional type ("joyful") and the music's theme type ("celebrating graduation") can be added. In this way, users can more efficiently input the second prompt message for generating the second piece of music, based on the first and recommended prompt messages.

[0129] Adjustments to the basic attributes of the music can be made as shown in Figure 3C. The interface with the AI ​​agent displays recommended adjustment options such as "Faster tempo," "Change song style," and "Add lyrics" for the user to choose from. Alternatively, users can directly input adjustment options such as "Slower tempo" to adjust the music's rhythm.

[0130] Adjustments to the basic attributes of music can also be made on the music editing interface, allowing users to adjust these attributes as a whole. As shown in Figure 7B, users can adjust the music's tempo using the speed control. Alternatively, the music can be divided into multiple sections, allowing users to adjust the basic attributes segment by segment.

[0131] Adjustments to the basic properties of music can also be made by displaying the music tracks, allowing users to adjust these properties track by track. Adjustments to each track can be made in conjunction with the beat. For example, one can adjust the duration of each beat within a track, or adjust the combination of strong and weak beats within that track.

[0132] For adjusting the addition of lyrics and vocals to the first track, the processes shown in Figures 5A and 5B can be used respectively, based on the adjustment information input by the user; alternatively, the user can enter the editing interface shown in Figures 7A and 7B and make adjustments based on the user's actions.

[0133] Lyrics can be generated using the first cue information used to generate the first piece of music; alternatively, lyrics can be generated based on other cue information different from the first cue information. Alternatively, based on the first lyrics generated from the first cue information, the lyrics can be adjusted according to at least one of the music's melody, rhythm, or style to obtain the second set of lyrics.

[0134] For vocal music, a matching voice—the singer's voice—must be generated. Similarly, users can use their own voice as the singer's. For example, voice imitation technology can be used to generate a matching voice. Alternatively, machine learning models can be used to generate one or more matching voices for the user to choose from. To further enhance the experience, titles and album art corresponding to the music can also be generated.

[0135] Finally, the generated second music track, along with its corresponding lyrics, vocals, title, and cover art, are combined into a second music work and output on the interactive interface. For example, the second music track with vocals can be played, and the title and cover art can be displayed. Users can perform operations such as publishing and sharing on the interactive interface shown in Figure 4C.

[0136] The above describes a music generation method provided by embodiments of this disclosure. The following description, with reference to Figures 9-11, describes a music generation apparatus according to embodiments of this disclosure, used to perform any of the above-described music generation methods.

[0137] Figure 9 shows a block diagram of a music generation apparatus according to some embodiments of the present disclosure.

[0138] As shown in Figure 9, the music generation device 9 includes: a display module 91 configured to display music generation guidance information on the user-smart agent interaction interface in response to the user's music generation request; a first output module 92 configured to output a first musical work generated according to the first prompt information input by the user based on the guidance information, wherein the first prompt information includes music description information and the first musical work includes first music; and a second output module 93 configured to output a second musical work generated according to the user's adjustment of the first musical work, wherein the adjustment includes at least one of adjustment to the first prompt information and adjustment to the first music.

[0139] The display module 91 of the music generating device 9 can be used to execute step S1 of FIG1 or step S1' of FIG8. The first output module 92 of the music generating device 9 can be used to execute step S2 of FIG1 or step S2' of FIG8. The second output module 93 of the music generating device 9 can be used to execute step S3 of FIG1 or step S3' of FIG8.

[0140] Figure 10 shows a block diagram of a music generation apparatus according to other embodiments of the present disclosure.

[0141] As shown in FIG10, the music generation device 10 includes: a memory 101; and a processor 102 coupled to the memory 101, the processor 102 being configured to execute the music generation method described in any of the foregoing embodiments based on instructions stored in the memory 101.

[0142] Memory 101 is used to store one or more computer-readable instructions. Memory 101 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 101 may, for example, store operating systems, application programs, boot loaders, databases, and other programs, as well as various application programs and various data.

[0143] The processor 102 is configured to execute computer-readable instructions to implement the music generation method described in any of the foregoing embodiments. Specific implementation details of each step of the music generation method can be found in the above embodiments; repeated details will not be elaborated upon here.

[0144] Processor 102 can be configured to execute steps S1-S3 of Figure 1 or steps S1'-S3' of Figure 8. Processor 102 can be various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) can be an x86 or ARM architecture, etc.

[0145] The processor 102 and the memory 101 can communicate with each other directly or indirectly. For example, the processor 102 and the memory 101 can communicate via a network. The network can include a wireless network, a wired network, and / or any combination of wireless and wired networks. The processor 102 and the memory 101 can also communicate with each other via a system bus, which is not limited in this disclosure.

[0146] It should be noted that the components of the music generation device 10 shown in Figure 10 are merely exemplary and not limiting. The music generation device 10 may also have other components depending on the specific application requirements. The processor 102 can control other components in the music generation device 10 to perform desired functions.

[0147] Music generation devices can be implemented by software, firmware, and / or hardware, and can be integrated into electronic devices with relevant applications installed.

[0148] Figure 11 shows a block diagram of an electronic device according to some embodiments of the present disclosure.

[0149] The electronic device 11 shown in Figure 11 can be a computer system with a dedicated hardware structure, which can perform corresponding functions when relevant applications are installed.

[0150] Electronic devices include, but are not limited to, mobile terminals such as smartphones, laptops, personal digital assistants (PDAs), tablet computers (PCs), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), wearable devices, and fixed terminals such as digital televisions and desktop computers.

[0151] As shown in Figure 11, the Central Processing Unit (CPU) 111 performs various processes based on a program stored in the Read-Only Memory (ROM) 112 or a program loaded from the Storage Section 118 into the Random Access Memory (RAM) 113. The RAM 113 stores data required as needed when the CPU 111 performs various processes. The CPU is merely exemplary and can also be other types of processors, such as the various processors described above. The ROM 112, RAM 113, and Storage Section 118 can be various forms of computer-readable storage media. It should be noted that although the ROM 112, RAM 113, and Storage Section 118 are shown separately in Figure 11, one or more of them can be combined or located in the same or different memories or storage modules.

[0152] CPU 111, ROM 112 and RAM 113 are interconnected via bus 114. Input / output interface 115 is also connected to bus 114.

[0153] The following components are connected to the input / output interface 115: input section 116, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output section 117, including displays such as cathode ray tube (CRT), liquid crystal display (LCD), speakers, vibrators, etc.; storage section 118, including hard disks, magnetic tapes, etc.; and communication section 119, including network interface cards such as LAN cards, modems, etc. The communication section 119 allows communication processing to be performed via a network such as the Internet. It is readily understood that although the various devices or modules in the electronic device 11 shown in Figure 11 communicate via bus 114, they can also communicate via a network or other means, wherein the network can include wireless networks, wired networks, and / or any combination of wireless and wired networks.

[0154] As needed, drive 1110 is also connected to input / output interface 115. Removable media 1111, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 1110 as needed, so that computer programs read from them can be installed into storage section 118 as needed.

[0155] When the above series of processes are implemented through software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as a removable medium 1111.

[0156] According to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product that, when run on a computer, causes the computer to implement the music generation method described in any of the foregoing embodiments. The computer program product includes a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 119, or installed from storage portion 118, or installed from ROM 112. When the computer program is executed by CPU 111, the music generation method of the embodiments of this disclosure is performed.

[0157] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0158] A computer-readable medium may be a computer-readable storage medium, a computer-readable signal medium, or any combination thereof.

[0159] Computer-readable storage media include, but are not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. A computer program is stored on the computer-readable storage medium that, when executed by a processor, implements the music generation method described in any of the foregoing embodiments.

[0160] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0161] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0162] In some embodiments, a computer program is also provided, comprising: instructions that, when executed by a processor, cause the processor to perform the music generation method of any of the above embodiments. For example, the instructions may be embodied in computer program code.

[0163] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0164] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0165] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0166] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.

Claims

1. A method for generating music, comprising: On the user-smart agent interaction interface, in response to the user's music generation request, guidance information for music generation is displayed; In response to the first prompt information input by the user based on the guidance information, the system outputs a first musical work generated according to the first prompt information, wherein the first prompt information includes description information of the music and the first musical work includes a first piece of music. In response to the user's adjustment to the first musical piece, a second musical piece generated according to the adjustment is output, the adjustment including at least one of the adjustment to the first prompt information and the adjustment to the first musical piece.

2. The music generation method according to claim 1, wherein, In response to the user input regarding adjustments to the first musical piece, the output of a second musical piece generated according to the adjustments includes: Receive adjustment information from the user regarding the basic attributes and / or style type of the first music, wherein the basic attributes include the music's rhythm, melody, timbre, and harmony; The output is a second musical work generated based on the adjustment information, the second musical work including a second piece of music.

3. The music generation method according to claim 1, wherein, In response to the user input regarding adjustments to the first musical piece, the output of a second musical piece generated according to the adjustments includes: Receive adjustment information from the user regarding the addition of lyrics to the first music; The first lyrics, generated based on the first prompt and the first music, are output as part of the second musical work.

4. The music generation method according to claim 1, wherein, In response to the user input regarding adjustments to the first musical piece, the output of a second musical piece generated according to the adjustments includes: Receive adjustment information from the user regarding the addition of lyrics to the first music; Output the first lyrics generated based on the first prompt information and the first music; Receive adjustment information from the user regarding the first lyrics; The second lyrics, obtained based on the adjustment information, are displayed as part of the second musical work.

5. The music generation method according to claim 1, wherein, In response to the user input regarding adjustments to the first musical piece, the output of a second musical piece generated according to the adjustments includes: Receive adjustment information from the user regarding adding vocals to the first music; Based on the aforementioned adjustment information, the timbres of multiple singers are provided; Output a human voice generated based on the timbre of one or more singers selected by the user, as part of the second musical work.

6. The music generation method according to claim 1, wherein, In response to the user input regarding adjustments to the first musical piece, the output of a second musical piece generated according to the adjustments includes: Displays the editing controls for the first musical piece; In response to the user's triggering of the editing control, the editing interface for the first musical work is entered; The second musical work is generated based on the user's adjustment operations on the first musical work in the editing interface, wherein the adjustment operations include at least one of the adjustment operations on the first prompt information and the adjustment operations on the first musical work.

7. The music generation method according to claim 6, wherein, Based on the user's adjustments to the first musical piece in the editing interface, the generation of the second musical piece includes: Receive the user's modification operation on the first prompt information; Display a second prompt message based on the modification operation; Using a machine learning model, the second musical piece is generated based on the second prompt information.

8. The music generation method according to claim 7, wherein, In response to the user input regarding adjustments to the first musical piece, the output of a second musical piece generated according to the adjustments includes: Display at least one recommendation message based on the first musical work, the recommendation message including at least one of the following: the theme of the music, the visuals of the music, the emotional type of the music, and the style of the music; The one or more recommended prompts selected by the user are used as part of the second prompt information.

9. The music generation method according to claim 8, wherein, The first piece of music includes multiple sections. In response to adjustments made by the user input to the first piece of music, a second piece of music generated based on the adjustments is output, including: The second musical work is generated based on the user's adjustment of the basic attributes of at least one of the multiple sections of the first music, the basic attributes including the rhythm, melody, timbre, and harmony of the music.

10. The music generation method according to claim 9, wherein, Each of the multiple segments includes multiple beats, and in response to the user input of adjustments to the first musical piece, outputs a second musical piece generated according to the adjustments, including: Display the audio track of the first piece of music; The second musical work is generated based on the user's adjustment operations on the first music track. The adjustment operations on the first music include at least one of adjusting the time value of the beat and adjusting the combination rules between beats.

11. The music generation method according to claim 1, wherein: The guidance information includes at least one candidate prompt template, each candidate prompt template corresponds to a complete music scene, and the information of each candidate prompt template includes at least one of the following: music theme, music visuals, music emotional type, and music style type. The prompt information in a candidate prompt template selected by the user is used as part or all of the first prompt information.

12. The music generation method according to claim 1, wherein: The guidance information includes at least one candidate prompt, and each candidate prompt includes at least one of the following: the theme of the music, the visuals of the music, the emotional type of the music, and the style of the music. The one or more candidate prompts selected by the user are used as part or all of the first prompt.

13. The music generation method according to any one of claims 1 to 12, wherein: The first musical work includes the title and cover art of the first musical piece; The second musical work includes the title and cover art of the second piece of music; The descriptive information of the music includes the theme, the imagery, the emotional type, and the style. At least one of the grid types.

14. A music generation device, comprising: The display module is configured to display music generation guidance information on the user's music generation request on the user-smart agent interaction interface. The first output module is configured to respond to a first prompt message input by the user based on guidance information, and output a first musical work generated according to the first prompt message, wherein the first prompt message includes description information of the music, and the first musical work includes a first piece of music; The second output module is configured to output a second musical work generated according to the user's adjustment to the first musical work, the adjustment including at least one of the adjustment to the first prompt information and the adjustment to the first music.

15. A music generation device, comprising: Memory; as well as A processor coupled to the memory, the processor being configured to execute the music generation method as described in any one of claims 1 to 13 based on instructions stored in the memory.

16. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the music generation method according to any one of claims 1 to 13.

17. A computer program product, when run on a computer, causes the computer to implement the music generation method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Lyric display method and device thereof, electronic equipment and computer readable storage medium

    CN112699269A

  • Music generation method and device, medium and computing equipment

    CN112785993A

  • Music work generation method and device thereof, music work synthesis method and device, equipment, medium and product

    CN113611268A

  • Method, device and equipment for generating arrangement, medium and computer program

    CN113838444A

  • Song creation method, song creation device, storage medium and electronic equipment

    CN115346503A