Music generation method, music generation apparatus, and computer readable storage medium
By receiving music description information input by users to generate preliminary works and allowing users to make adjustments, a final music work that meets the user's needs is generated. This solves the problem that music generated by machine learning models cannot meet user needs, and improves the quality of music generation and user experience.
Patent Information
- Application Number
- PCT/CN2024/089579
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-24
- Publication Date
- 2025-10-30
AI Technical Summary
Existing machine learning models cannot generate music that meets user needs, limited by model performance and the accuracy of users' textual descriptions of the music.
This paper provides a music generation method that receives music description information input by the user, generates a preliminary music piece, allows the user to adjust the piece, and generates a final music piece based on the adjustments using a machine learning model.
It improves the quality of music generation, enhances the user experience, and makes the generated music more in line with user expectations.
Smart Images

Figure CN2024089579_30102025_PF_FP_ABST
Abstract
Description
Music generation method, music generation device and computer-readable storage medium Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a music generation method, a music generation apparatus, a computer-readable storage medium, and a computer program product. Background Technology
[0002] Music generation technology is an important application in the field of computer technology. With the rapid development of artificial intelligence technology, natural language processing technology is widely used for music generation, utilizing machine learning models to automatically generate music data from text data.
[0003] However, due to limitations in the performance of machine learning models and the accuracy of users' textual descriptions of music, automatically generated music often fails to meet users' needs.
[0004] Summary of the Invention
[0005] In view of this, embodiments of this disclosure propose a music generation method, a music generation apparatus, a computer-readable storage medium, and a computer program product. After automatically generating music based on a text description, the user can make targeted adjustments to the generated music and regenerate the music based on the user's adjustments, thereby improving the quality of music generation and enhancing the user experience.
[0006] According to a first aspect of some embodiments of this disclosure, a music generation method is provided, comprising: receiving first prompt information input by a user, the first prompt information including description information of music; generating a first musical work based on the first prompt information, the first musical work including first music; and generating a second musical work based on an adjustment operation by the user on the first musical work, the adjustment operation including at least one of an adjustment operation on the first prompt information and an adjustment operation on the first music.
[0007] According to a second aspect of some embodiments of this disclosure, a music generation apparatus is provided, comprising: a receiving module configured to receive first prompt information input by a user, the first prompt information including description information of music; a first generation module configured to generate a first musical work based on the first prompt information, the first musical work including first music; and a second generation module configured to generate a second musical work using a machine learning model based on an adjustment operation by the user on the first musical work, the adjustment operation including at least one of an adjustment operation on the first prompt information and an adjustment operation on the first music.
[0008] According to a third aspect of some embodiments of the present disclosure, a music generation apparatus is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute any of the foregoing music generation methods based on instructions stored in the memory.
[0009] According to a fourth aspect of some embodiments of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements any of the aforementioned music generation methods.
[0010] According to a fifth aspect of some embodiments of the present disclosure, a computer program product is provided that, when the computer program product is run on a computer, causes the computer to implement any of the aforementioned music generation methods.
[0011] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0012] Other features and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0013] Preferred embodiments of the present disclosure are described below with reference to the accompanying drawings. The accompanying drawings, which are included to provide a further understanding of the present disclosure, and which, together with the following detailed description, are incorporated in and form a part of this specification and are used to explain the present disclosure. It should be understood that the drawings described below only relate to some embodiments of the present disclosure and are not intended to limit the present disclosure. In the drawings:
[0014] Figure 1 shows a flowchart of a music generation method according to some embodiments of the present disclosure;
[0015] Figure 2A shows an interactive interface for music generation according to some embodiments of the present disclosure;
[0016] Figure 2B shows an interactive interface for music generation according to other embodiments of the present disclosure;
[0017] Figure 2C shows an interactive interface for music generation according to some embodiments of the present disclosure;
[0018] Figure 2D shows an interactive interface for music generation according to some embodiments of the present disclosure;
[0019] Figure 3A shows the editing interface of a first musical work according to some embodiments of the present disclosure;
[0020] Figure 3B shows the editing interface of a first musical work according to other embodiments of the present disclosure;
[0021] Figure 3C shows the editing interface of a first musical work according to some embodiments of the present disclosure;
[0022] Figure 3D shows the editing interface of a first musical work according to some embodiments of the present disclosure;
[0023] Figure 4 illustrates a flowchart of generating a first musical work according to some embodiments of the present disclosure;
[0024] Figure 5A illustrates a flowchart of generating a second musical work according to some embodiments of the present disclosure;
[0025] Figure 5B illustrates a flowchart of generating a second musical work according to some other embodiments of the present disclosure;
[0026] Figure 6 shows an interactive interface for generating a second musical work according to some embodiments of the present disclosure;
[0027] Figure 7 shows a flowchart of a music generation method according to some other embodiments of the present disclosure;
[0028] Figure 8A shows the interactive interface after generating a second musical work according to some embodiments of the present disclosure;
[0029] Figure 8B shows the interactive interface after generating a second musical work according to some other embodiments of the present disclosure;
[0030] Figure 9 shows a block diagram of a music generation apparatus according to some embodiments of the present disclosure;
[0031] Figure 10 shows a block diagram of a music generation apparatus according to other embodiments of the present disclosure;
[0032] Figure 11 shows a block diagram of an electronic device according to some embodiments of the present disclosure.
[0033] It should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not necessarily drawn to actual scale. The same or similar reference numerals are used in the various drawings to denote the same or similar parts. Therefore, once an item is defined in one drawing, it may not be discussed further in subsequent drawings. Detailed Implementation
[0034] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. However, it is obvious that the described embodiments are only some embodiments of this disclosure, and not all embodiments. The following description of the embodiments is merely illustrative and is in no way intended to limit this disclosure or its application or use. It should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.
[0035] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of components and steps set forth in these embodiments should be interpreted as merely exemplary and do not limit the scope of this disclosure.
[0036] As used in this disclosure, the term "comprising" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". Furthermore, as used in this disclosure, the term "including" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". Therefore, "comprising" and "including" are synonymous. The term "based on" means "at least partially based on".
[0037] Throughout this specification, the terms "one embodiment," "some embodiments," or "embodiment" mean that a specific feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of the invention. For example, the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; and the term "some embodiments" means "at least some embodiments." Furthermore, the appearance of the phrases "in one embodiment," "in some embodiments," or "in an embodiment" in various places throughout the specification does not necessarily refer to the same embodiment, but may refer to the same embodiment.
[0038] It should be noted that the concepts of "first," "second," etc., used in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to define the order of functions performed by these devices, modules, or units or their interdependencies. Unless otherwise specified, the concepts of "first," "second," etc., are not intended to imply that the objects described herein must be in a given temporal, spatial, rank, or any other given order.
[0039] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0040] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0041] The embodiments of this disclosure are described in detail below with reference to the accompanying drawings; however, this disclosure is not limited to these specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. Furthermore, in one or more embodiments, specific features, structures, or characteristics can be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.
[0042] In existing music generation technologies, the automatically generated music often fails to meet user needs due to limitations in the performance of machine learning models and the accuracy of users' textual descriptions of the music. To improve the quality of music generation and enhance user experience, this disclosure provides a novel music generation method. After automatically generating music based on textual descriptions, users can make targeted adjustments to the generated music and regenerate the music according to their adjustments.
[0043] Figure 1 shows a schematic flowchart of a music generation method according to some embodiments of the present disclosure.
[0044] As shown in Figure 1, the music generation method includes: step S1, receiving first prompt information input by the user, the first prompt information including music description information; step S2, generating a first musical work based on the first prompt information, the first musical work including first music; step S3, generating a second musical work based on the user's adjustment operation on the first musical work, the adjustment operation including at least one of adjustment operation on the first prompt information and adjustment operation on the first music.
[0045] The music generation method in this embodiment can be executed on the client side or partially on the server side.
[0046] The themes of the music include, for example, longing for home, remembering friends, and celebrating success. The imagery in the music includes, for example, sunsets, waves, rocks, streets, and buildings. The emotional types of the music include, for example, sadness, joy, helplessness, and emotion. The musical styles include, for example, rock, jazz, pop, folk, and rap.
[0047] The first musical work can be a purely instrumental piece, without lyrics, consisting only of the first note, such as a piano piece or guitar piece. It can also be a vocal piece, including not only the first note but also the accompanying vocals. Vocal pieces usually also include lyrics, sung by the vocals.
[0048] The music generation method according to embodiments of the present disclosure will be further described below with reference to Figures 2A-2D. Figures 2A-2D respectively show the interactive interface for music generation according to different embodiments of the present disclosure.
[0049] As shown in Figure 2A, the interactive interface for music generation (also known as "music creation") provides input boxes, such as "Try describing your mood or imagery," guiding users to enter descriptive information about the music. The interface also displays some already generated musical works. For these works, the title and album art can be displayed.
[0050] Figure 2B illustrates another interactive interface for music creation. Figure 2B also provides an input box displaying "Describe the sound or image" to guide the user in entering descriptive information about the music. In response to the input box being activated, the user can enter descriptive information about the music using text or voice through, for example, a keyboard or voice input device, as a prompt to generate the musical work. The descriptive information can include at least one of the following: the theme of the music, the imagery of the music, the emotional type of the music, and the genre of the music. If the user wishes to generate vocal music using their own voice, they can select the "Use my voice" function shown in Figure 2B.
[0051] Of course, to help users input more efficient and accurate prompts, some suggestions can be provided on the interactive interface to facilitate user selection.
[0052] In some embodiments, receiving the first prompt information input by the user in step S1 includes: providing at least one candidate prompt template in the interactive interface, each candidate prompt template corresponding to a complete music scene, and the information of each candidate prompt template including at least one of the following: music theme, music visuals, music emotional type, and music style type; and using the prompt information in a candidate prompt template selected by the user as part or all of the first prompt information.
[0053] As shown in Figure 2C, the interactive interface displays multiple candidate prompt templates, such as templates A, B, C, D, and E, for the user to choose from. For example, one candidate prompt template might include information about a rainy day in Kyoto, reflecting the imagery of the music. Another candidate prompt template might include information about joyfully embracing the sky like birds, reflecting both the imagery of the music and its emotional tone.
[0054] Since each candidate prompt template corresponds to a complete music scene, users can only choose one. The prompt information from the selected candidate prompt template can be directly applied to the input box as the first prompt information. As shown in Figure 2C, if the user selects the candidate prompt template "Embrace the sky joyfully like a bird," the corresponding prompt information will be directly applied to the input box.
[0055] Of course, users can further adjust the prompts based on the prompts in the candidate prompt templates. For example, they can use the keyboard or voice input device shown at the bottom of Figure 2C to delete, add, or modify the prompts. In this case, the prompts in a candidate prompt template selected by the user are only a part of the first prompt.
[0056] In other embodiments, receiving the first prompt information input by the user includes: providing at least one candidate prompt information on the interactive interface, each candidate prompt information including at least one of the following: the theme of the music, the visuals of the music, the emotional type of the music, and the style type of the music; and using one or more candidate prompt information selected by the user as part or all of the first prompt information.
[0057] As shown in Figure 2D, multiple candidate prompts are displayed on the interactive interface for the user to choose from. For example, the candidate prompts may include prompts reflecting the theme, imagery, emotional type, and style of the music, such as longing, rainy day, Kyoto, birds, cheerful, and rock.
[0058] Users can select multiple candidate prompts and combine them. The selected prompts can be directly applied to the input box as the first prompt. For example, as shown in Figure 2D, users can select multiple candidate prompts such as "birds" and "joyful". Of course, users can also further adjust the prompts in the input box. For example, using the keyboard or voice input device shown at the bottom of Figure 2D, the prompt "birds, joyful" can be changed to "Birds joyfully embrace the sky with their wings". In this case, the multiple candidate prompts selected by the user are only a part of the first prompt.
[0059] The above description, with reference to Figures 2A-2D, illustrates how the first prompt information input by the user is received in step S1. The following description, with reference to Figures 3A-3D, describes how the first musical piece is generated based on the first prompt information in step S2.
[0060] In step S2, generating the first musical work based on the first prompt information includes: using a machine learning model to generate the first music based on the first prompt information.
[0061] Machine learning models include, for example, Large Language Models (LLMs) or other natural language processing (NLP) models. For instance, machine learning models such as Variational Autoencoders (VAEs), Generative Adversarial Networks (GANs), and Transformers can be used to achieve diverse and high-quality music generation.
[0062] Figure 3A shows an editing interface for a first musical work according to some embodiments of the present disclosure. After generating the first music, the user directly enters the interactive interface shown in Figure 3A, which allows editing of the first music, i.e., the editing interface for the first musical work.
[0063] As shown in Figure 3A, the editing interface for the first musical piece includes a control for generating lyrics. In response to the user's activation of the lyric generation control, first lyrics corresponding to the first musical piece are generated. In some embodiments, the machine learning model is used to generate the first lyrics based on the first prompt information and the first musical piece. As shown in Figure 3A, the generated lyrics are displayed in a text box at the bottom of the editing interface.
[0064] Of course, the first lyrics can also be generated based on other prompts different from the first prompts used to generate the first music, instead of the first prompts used to generate the first music. In other embodiments, other prompts input by the user are received, and the machine learning model is used to generate the first lyrics based on the other prompts, where the other prompts include descriptive information about the first lyrics. Alternatively, the lyrics can be generated by comprehensively considering the first music and the prompts input by the user to generate the lyrics, so that the generated lyrics better match the music. That is, the machine learning model is used to generate the first lyrics based on the other prompts and the first music.
[0065] In response to the user triggering the lyrics generation control in Figure 3A again, the lyrics editing interface is entered, as shown in Figure 3B. Figure 3B shows the editing interface of a first musical work according to other embodiments of the present disclosure.
[0066] Figure 3B shows the previously generated lyrics. Users can edit the previously generated lyrics using a keyboard or voice input device.
[0067] Of course, if the user wishes to use lyrics completely different from those previously generated, they can directly trigger the clear control to enter the editing interface shown in Figure 3C. Figure 3C shows the editing interface of a first musical work according to some embodiments of the present disclosure.
[0068] In Figure 3C, the user can re-enter prompts for generating lyrics via a keyboard or voice input device, such as other prompts different from the first prompt, as described above. In response to the user triggering the lyrics generation control, a machine learning model is used to generate first lyrics matching the first music based on the other prompts.
[0069] For vocal music, corresponding vocals are also required, including the singer's voice. As shown in Figures 3A-3C, the editing interface for the first music track also includes controls for adding vocals, such as controls with vocals. Responding to the user's triggering of the control with vocals on the editing interface of the first music track, the corresponding vocals for the first music track are generated. For example, the singer's voice can be generated based on the first music track and the first lyrics. Of course, if the vocal music track does not include lyrics, the singer's voice can also be generated based solely on the first music track.
[0070] For the singer's voice, the user can use a default vocal timbre that matches the music and its lyrics. In some embodiments, the singer's voice is generated using a machine learning model based on the first music and the first lyrics. Multiple singer voices can be generated using the machine learning model. The user can choose the singer's voice according to their actual needs.
[0071] Users can also use their own voices for the singer. In other embodiments, the singer's voice is generated using voice imitation technology based on the user's voice. Of course, if the client or server has already stored the user's voice, it can also be directly invoked based on the user's authorization.
[0072] Figure 3D illustrates an editing interface for a first musical work according to some embodiments of the present disclosure. As shown in Figure 3D, the editing interface for the first musical work displays a variety of sounds available for the user to select. Sound A in Figure 3D is, for example, the voice of a poet; sound B is, for example, the voice of an actor; sound C is, for example, the voice of a character in a movie or television series; and the default vocal is, for example, the voice of a singer.
[0073] Users can use the default voice, which can be based on their previous settings. Of course, users can also choose other voices according to their needs. For example, users can choose their own voice or choose the voice of other singers according to their preferences.
[0074] The following description, with reference to Figure 4, continues how to generate a first musical work with vocals. Figure 4 shows a flowchart of generating a first musical work according to some embodiments of this disclosure.
[0075] A first musical work with vocals includes the first melody and the singer's voice, and may or may not include the first lyrics. The following focuses on the generation of vocal music including the first lyrics.
[0076] As shown in Figure 4, generating the first musical work based on the first prompt information includes: Step S21, using the machine learning model to generate the first music based on the first prompt information; Step S22, receiving other prompt information input by the user, the other prompt information including the description information of the first lyrics; Step S23, using the machine learning model to generate the first lyrics based on the other prompt information; Step S24, generating the singer's voice based on the first music and the first lyrics; Step S25, synthesizing the first musical work based on the first music, the first lyrics, and the singer's voice.
[0077] In step S22, the other prompts can be the same as the first prompt, meaning the first music and first lyrics are generated based on the same prompts. Of course, the other prompts can also be different from the first prompt. Users can modify the first prompt by deleting, adding, or other means to obtain the other prompts. Alternatively, users can choose not to adjust the first prompt but instead input additional prompts solely from the perspective of the lyrics.
[0078] In step S24, if it is desired to generate a first piece with initial lyrics, a suitable vocal is selected by comprehensively considering the matching of sound, music, and lyrics. Of course, if there are no initial lyrics, the vocal can be selected by only considering the matching of sound and music.
[0079] In step S25, based on the basic attributes of music, the first music, the first lyrics, and the singer's voice are synchronized to synthesize a first musical work. Alternatively, the title and cover art of the first music can be further synthesized with the first music, the first lyrics, and the singer's voice to generate a more complete first musical work.
[0080] The steps of generating the first music in step S21, receiving user input prompts in step S22, generating the first lyrics in step S23, and generating the singer's voice in step S24 have been described previously and will not be repeated here.
[0081] The preceding sections, with reference to Figures 2A-3D, described in detail how to generate the first music, first lyrics, and first vocals to obtain the first musical piece. However, the first music generated in one go often fails to meet the user's quality requirements for the musical piece, thus requiring the user to make targeted adjustments to the music generation. The following sections, with reference to Figures 5A-5B and 6, will describe how, in step S3, a second musical piece is generated based on the user's adjustments to the first musical piece.
[0082] Figure 5A illustrates a flowchart of generating a second musical work according to some embodiments of the present disclosure. As shown in Figure 5A, generating a second musical work based on the user's adjustment operation on the first musical work includes: step S31, displaying the first prompt information on the interactive interface; step S32, receiving the user's modification operation on the first prompt information; step S33, displaying a second prompt information obtained based on the modification operation; and step S34, generating the second musical work using a machine learning model based on the second prompt information.
[0083] In step S31, the input box for the prompt message can be reactivated, and the previously entered prompt message, i.e., the first prompt message, can be displayed in the input box. For example, in response to the pen-shaped editing control shown in Figure 3A being triggered, an interactive interface similar to that shown in Figures 2B and 2C can be entered.
[0084] In addition to displaying the initial prompts used to generate the first music track, further prompts can be recommended based on the generated first music track, that is, suggestions can be given to help users adjust the first music track.
[0085] In some embodiments, the interactive interface displays at least one recommendation message based on the first musical work, the recommendation message including at least one of the following: the theme of the music, the visuals of the music, the emotional type of the music, and the style of the music.
[0086] Figure 3A shows some recommended prompts, such as music genres like pop, jazz, rap, rock, and folk. Regarding music genres, Figure 3A also shows a style control; responding to the user's activation of this control provides more music genre options. Figure 3A also shows recommended prompts related to the emotional tone of the music, such as sadness, emotion, helplessness, and joy.
[0087] For example, after generating the first piece of music based on the initial prompt message such as "Birds joyfully embrace the sky with their wings," if the user needs to modify the prompt message, additional prompt messages can be recommended based on the visual imagery of "birds embracing the sky with their wings" in the first piece of music, the emotional type of the music ("joyful"), etc. For example, prompt messages related to the theme could include "Celebrating Success" or "Celebrating Graduation," and prompt messages related to the genre could include "Pop" or "Folk." The user can select one or more recommended prompt messages, which can then be used as part of the second prompt message.
[0088] In steps S32 and S33, the user can modify the first prompt information to obtain the second prompt information. For example, if the first prompt information indicates the music's genre is pop, the user can change it to folk to obtain the second prompt information. Similarly, if the first prompt information indicates the music's emotional type is melancholic, the user can change it to resignation to obtain the second prompt information. Likewise, the user can modify the music's theme, visuals, etc., in the first prompt information to obtain the second prompt information.
[0089] Of course, users can also delete or add descriptive information about the music to the first prompt message to obtain the second prompt message. For example, if the first prompt message includes the theme of celebrating success, the visual imagery of architecture, and the mood of cheerfulness, users can delete the visual description, resulting in a second prompt message with less descriptive information, reducing restrictions on the generated visuals. Similarly, in this case, users can add descriptive information about the music's genre, such as "pop," resulting in a second prompt message with more descriptive information, improving the match between the generated music and the user's needs.
[0090] In step S34, a second musical piece that better meets the user's needs is generated using a machine learning model based on the second prompt information obtained by adjusting the first prompt information. For example, the musical piece can be regenerated in response to the triggering of the regenerated control shown in Figure 3A. Similar to the first musical piece, the second musical piece also includes a title and cover art.
[0091] It should be understood that user adjustments to the first piece of music can be made to the first prompt message, the first piece of music itself, or both. Adjustments to the first piece of music can include changes to the music's basic attributes.
[0092] In some embodiments, generating a second musical work based on the user's adjustment operation on the first musical work includes: generating a second musical work based on the user's adjustment operation on the basic attributes of the first musical work, wherein the basic attributes include the rhythm, melody, timbre, and harmony of the music.
[0093] If the user is not satisfied with the rhythm of the generated first track, they can adjust it individually. The rhythm of music generally includes beat, tempo, accents, and style. Users can adjust at least one of these factors to modify the rhythm of the first track.
[0094] The editing interface shown in Figure 3A illustrates a speed control. In response to the user triggering the speed control, the speed editing interface shown in Figure 6 is entered. Figure 6 illustrates an interactive interface for generating a second musical work according to some embodiments of this disclosure. As shown in Figure 6, the user can move the speed adjustment control to adjust the tempo of the music.
[0095] Similarly, users can adjust the melody, timbre, harmony, etc. of the first piece of music based on their own understanding and preferences to obtain a second piece of music that better meets their needs.
[0096] Adjusting the basic attributes of the first piece of music allows for overall adjustments to the first piece of music. Since music can be divided into sections, adjustments can also be made to the first piece of music section by section.
[0097] For example, if the first piece of music includes multiple sections, generating a second musical work based on the user's adjustment operations on the basic attributes of the first piece of music includes: generating the second musical work based on the user's adjustment operations on the basic attributes of at least one of the multiple sections of the first piece of music.
[0098] You can adjust the basic properties of only certain sections of the first music track, such as adjusting the rhythm of only the musical section corresponding to the intro of the first music track, without adjusting other parts of the first music track.
[0099] The rhythm adjustment of the first piece can be further refined. For example, each section of the music has multiple beats, and users can adjust the rhythm of the first piece by beat, which can more efficiently generate music works that match the user's needs.
[0100] For example, each of the multiple sections of the first music includes multiple beats, and the rhythm of the first music includes the beats of the music. Generating a second musical work based on the user's adjustment operations on the basic attributes of the first music includes: displaying the audio track of the first music; generating the second musical work based on the user's adjustment operations on the first music based on the audio track, wherein the adjustment operations on the first music include at least one of adjusting the duration of the beats and adjusting the combination rules between beats.
[0101] One audio track corresponds to one part of a piece of music. Each audio track can correspond to the performance of one instrument. The first piece of music may include multiple audio tracks. Displaying the audio tracks of the first piece of music allows users to adjust them based on the tracks. That is, based on the user's adjustment operation on at least one of the multiple audio tracks of the first piece of music, a second piece of music can be generated.
[0102] For example, in the multiple tracks of the first piece of music, the first track corresponds to the drum performance, giving the music its rhythmic backbone; the second track corresponds to the bass rhythm, complementing the drum performance; the third track corresponds to the guitar playing in sync with the beat; and the fourth track corresponds to the vocals accompanying the music. Users can adjust only some tracks instead of all of them. For instance, a user can adjust only the guitar playing in the third track. Specifically, they can adjust the duration of each beat (i.e., note time value) in the third track, as well as the combination of strong and weak beats, or both.
[0103] The foregoing, with reference to Figures 5A and 6, describes how a second musical work is generated based on user adjustments to a first musical work. As described in detail with reference to different embodiments, user adjustments to the first musical work include both overall adjustments, such as adjustments to the first prompt information used to generate the musical work, and partial adjustments, such as adjustments to the basic attributes of the generated music, such as rhythm and melody. Other embodiments for generating a second musical work based on user adjustments to the first musical work are further described below with reference to Figure 5B.
[0104] Figure 5B illustrates a flowchart of generating a second musical work according to some other embodiments of the present disclosure. As shown in Figure 5B, generating a second musical work based on the user's adjustment operation on the first musical work includes: step S31', generating the second music based on the user's adjustment operation on the first musical work; step S32', adjusting the first lyrics based on at least one of the melody, rhythm, and style type of the second music to obtain the second lyrics; and step S33', synthesizing the second musical work based on the second music and the second lyrics.
[0105] The first musical piece includes the first piece of music and the first set of lyrics. Because the music and lyrics have a certain degree of independence, the first piece of music and the first set of lyrics can be adjusted separately to efficiently generate a second musical piece that better suits the user's needs.
[0106] First, in step S31', a machine learning model can be used to generate a second piece of music based on the user's adjustments to the first piece of music. As mentioned earlier, the adjustments to the first piece of music can be adjustments to the prompt information used to generate the first piece of music, or adjustments to the basic attributes of the first piece of music.
[0107] Then, in step S32', the first lyrics can be adjusted according to the basic attributes of the second music, such as melody and rhythm, to obtain the second lyrics. For example, if the rhythm of the second music is faster than that of the first music, the first lyrics corresponding to the first music may not be suitable for the second music. Therefore, the first lyrics can be adjusted to obtain second lyrics that match the rhythm of the second music. Alternatively, the first lyrics can be adjusted according to the style of the second music to obtain the second lyrics. For example, if the style of the first music is rock and the style of the second music is pop, the first lyrics corresponding to the first music are usually not suitable for the second music. Therefore, the first lyrics can be adjusted to obtain second lyrics that match the style of the second music.
[0108] Of course, to accommodate the adjustments to the music and lyrics, the vocals were also adjusted accordingly. That is, the singer's voice was generated based on the second piece of music and the second set of lyrics.
[0109] Finally, in step S33', the second music and the second lyrics are combined into a second musical piece. Of course, if there are vocals, the singer's voice also needs to be combined into the second musical piece. The title and cover art of the second musical piece can also be combined into a single piece.
[0110] Several embodiments of the music generation method have been described in detail above with reference to Figures 1-6. A specific application example of the music generation method is described below with reference to Figure 7. Figure 7 shows a flowchart of a music generation method according to other embodiments of this disclosure.
[0111] As shown in Figure 7, the music generation method includes: step S1', receiving the first prompt information input by the user; step S2', generating the first music according to the first prompt information, and generating the first lyrics, first title, and first cover corresponding to the first music to synthesize the first music work; step S3', generating the second music according to the user's adjustments to the first music, and generating the second lyrics, second vocals, second title, and second cover corresponding to the second music to synthesize the second music work.
[0112] Step S1' is similar to step S1 in Figure 1. At least one prompt message suggestion can be provided on the interactive interface for the user to choose as part or all of the first prompt message.
[0113] The suggested prompts can be candidate prompt templates, as shown in Figure 2B, and users can only choose one. For example, a user can select the candidate prompt template "Kyoto on a Rainy Day," and the prompt information included in this template can be directly applied to the prompt information input box. Users can directly use these prompts as the first prompt information for generating the first piece of music. Of course, users can also adjust the prompt information based on this, such as adding descriptive information about the emotional type of the music, such as "sad," to obtain a first prompt information with more descriptive information about the music.
[0114] The suggested prompts can also be candidate prompts, as shown in Figure 2C. Users can select multiple candidate prompts and combine them. For example, users can select multiple candidate prompts such as "birds" and "cheerful," and similarly, these prompts can be directly applied to the prompt input box. Users can directly use these prompts as the first prompt for generating the first piece of music. Of course, users can also adjust the prompts based on this, such as adding descriptive information about the music's visuals, such as "wings" and "sky," to obtain a first prompt with more descriptive information about the music.
[0115] Next, in step S2', a machine learning model is used to generate the first piece of music based on the first prompt information. Alternatively, the machine learning model can also be used to generate the first lyrics corresponding to the first piece of music based on the first prompt information. Of course, considering the relative independence of music and lyrics, the first lyrics can also be generated based on other prompt information different from the first prompt information. Considering the display requirements of the musical work, a first title and a first cover art corresponding to the first piece of music can also be generated. Then, the first piece of music, the first lyrics, the first title, and the first cover art are combined to form the first musical work.
[0116] Since the music generated in one go is usually difficult to directly match the user's requirements, considering subsequent adjustments, in step S2', only the first piece of music can be generated first, without generating the first lyrics, first title, first cover, or corresponding vocals. The corresponding lyrics, vocals, title, and cover can be generated only after the generated music meets the user's needs.
[0117] Next, in step S3', a machine learning model is used to generate a second piece of music based on the user's adjustments to the first piece of music. These adjustments can be made to the initial prompt information used to generate the first piece of music, or to the basic attributes of the first piece of music.
[0118] Regarding adjustments to the prompts, the input box for the prompts can be reactivated, displaying the previously entered prompts for users to edit. Additionally, at least one recommended prompt based on the initial music can be displayed on the interface to help users adjust the initial music more effectively. For example, if the initial music is generated based on the prompt "Birds joyfully embrace the sky with their wings," the interface can recommend prompts such as the music's mood type ("joyful") and theme type ("celebrating graduation"). This allows users to more efficiently input secondary prompts for generating the secondary music, building upon the initial and recommended prompts.
[0119] Regarding adjustments to the basic attributes of the music, the system can display these attributes, allowing users to adjust the overall properties of the first track. As shown in Figure 6, users can adjust the tempo of the first track using the movement speed adjustment control. Alternatively, the music can be divided into multiple sections, allowing users to adjust the basic attributes of the first track segment by segment.
[0120] Adjustments to the basic properties of music can also be made by displaying the music tracks, allowing users to adjust the basic properties of the first track individually. Adjustments to each track can be made in conjunction with the beat. For example, the duration of each beat in a track can be adjusted, as can the combination pattern of strong and weak beats within that track.
[0121] After generating a second piece of music that matches the user's needs, a second set of lyrics corresponding to the second piece of music can be generated.
[0122] Similar to generating the first set of lyrics, the second set of lyrics can be generated based on the second cue information used to generate the second music; alternatively, it can be generated based on other cue information different from the second cue information. Alternatively, the first set of lyrics can be adjusted based on at least one of the following: melody, rhythm, or style of the second music, to obtain the second set of lyrics. This is particularly useful when only the basic attributes of the first music need to be adjusted.
[0123] After the second music track and second lyrics have been matched to the user's needs, for vocal music, a matching vocal track—the singer's voice—needs to be generated. Similarly, the user can use their own voice as the singer's voice. For example, voice imitation technology can be used to generate a second vocal track that matches the second music track and second lyrics. Alternatively, machine learning models can be used to generate one or more vocal tracks that match the second music track and second lyrics, allowing the user to choose. To demonstrate the user's needs, a second title and a second cover art corresponding to the second music track can also be generated.
[0124] Finally, the second music track, second lyrics, second vocals, second title, and second cover art are combined into a second musical work, which can be published and shared by the user. The save and publish controls can be triggered directly on the interactive interface shown in Figure 3A. Alternatively, publishing and sharing can be performed on the returned musical work's interactive interface.
[0125] Figures 8A and 8B show the interactive interface after generating a second musical work according to an embodiment of the present disclosure. As shown in Figure 8A, the generated second musical work is displayed in the interactive interface, including the title and cover art. For example, the title of the music is displayed as "Sunset Glow and Waves". The cover art can be displayed as an image including a sunset over the sea. The input box of the interactive interface in Figure 8A also displays the prompt message "In the sunset glow, waves crash against the rocks" for generating the second musical work. Users can publish the generated second musical work, for example, by triggering the publish control shown in Figure 8A. By triggering the publish control in Figure 8A in different ways, users can also perform operations such as sharing, viewing descriptions, or deleting the generated musical work, as shown in Figure 8.
[0126] The above describes a music generation method provided by embodiments of this disclosure. The following description, with reference to Figures 9-11, describes a music generation apparatus according to embodiments of this disclosure, used to perform any of the above-described music generation methods.
[0127] Figure 9 shows a block diagram of a music generation apparatus according to some embodiments of the present disclosure.
[0128] As shown in Figure 9, the music generation device 9 includes: a receiving module 91 configured to receive first prompt information input by a user, the first prompt information including music description information; a first generation module 92 configured to generate a first musical work based on the first prompt information, the first musical work including first music; and a second generation module 93 configured to generate a second musical work based on the user's adjustment operation on the first musical work using the machine learning model, the adjustment operation including at least one of adjustment operation on the first prompt information and adjustment operation on the first music.
[0129] The receiving module 91 of the music generating device 9 can be used to execute step S1 of FIG1 or step S1' of FIG7. The first generating module 92 of the music generating device 9 can be used to execute step S2 of FIG1 or step S2' of FIG7. The second generating module 93 of the music generating device 9 can be used to execute step S3 of FIG1 or step S3' of FIG7.
[0130] Figure 10 shows a block diagram of a music generation apparatus according to other embodiments of the present disclosure.
[0131] As shown in FIG10, the music generation device 10 includes: a memory 101; and a processor 102 coupled to the memory 101, the processor 102 being configured to execute the music generation method described in any of the foregoing embodiments based on instructions stored in the memory 101.
[0132] Memory 101 is used to store one or more computer-readable instructions. Memory 101 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 101 may, for example, store operating systems, application programs, boot loaders, databases, and other programs, as well as various application programs and various data.
[0133] The processor 102 is configured to execute computer-readable instructions to implement the music generation method described in any of the foregoing embodiments. Specific implementation details of each step of the music generation method can be found in the above embodiments; repeated details will not be elaborated upon here.
[0134] Processor 102 can be configured to execute steps S1-S3 of Figure 1 or steps S1'-S3' of Figure 7. Processor 102 can be various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) can be an x86 or ARM architecture, etc.
[0135] The processor 102 and the memory 101 can communicate with each other directly or indirectly. For example, the processor 102 and the memory 101 can communicate via a network. The network can include a wireless network, a wired network, and / or any combination of wireless and wired networks. The processor 102 and the memory 101 can also communicate with each other via a system bus, which is not limited in this disclosure.
[0136] It should be noted that the components of the music generation device 10 shown in Figure 10 are merely exemplary and not limiting. The music generation device 10 may also have other components depending on the specific application requirements. The processor 102 can control other components in the music generation device 10 to perform desired functions.
[0137] Music generation devices can be implemented by software, firmware, and / or hardware, and can be integrated into electronic devices with relevant applications installed.
[0138] Figure 11 shows a block diagram of an electronic device according to some embodiments of the present disclosure.
[0139] The electronic device 11 shown in Figure 11 can be a computer system with a dedicated hardware structure, which can perform corresponding functions when relevant applications are installed.
[0140] Electronic devices include, but are not limited to, mobile terminals such as smartphones, laptops, personal digital assistants (PDAs), tablet computers (PCs), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), wearable devices, and fixed terminals such as digital televisions and desktop computers.
[0141] As shown in Figure 11, the Central Processing Unit (CPU) 111 performs various processes based on a program stored in the Read-Only Memory (ROM) 112 or a program loaded from the Storage Section 118 into the Random Access Memory (RAM) 113. The RAM 113 stores data required as needed when the CPU 111 performs various processes. The CPU is merely exemplary and can also be other types of processors, such as the various processors described above. The ROM 112, RAM 113, and Storage Section 118 can be various forms of computer-readable storage media. It should be noted that although the ROM 112, RAM 113, and Storage Section 118 are shown separately in Figure 11, one or more of them can be combined or located in the same or different memories or storage modules.
[0142] CPU 111, ROM 112 and RAM 113 are interconnected via bus 114. Input / output interface 115 is also connected to bus 114.
[0143] The following components are connected to the input / output interface 115: input section 116, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output section 117, including displays such as cathode ray tube (CRT), liquid crystal display (LCD), speakers, vibrators, etc.; storage section 118, including hard disks, magnetic tapes, etc.; and communication section 119, including network interface cards such as LAN cards, modems, etc. The communication section 119 allows communication processing to be performed via a network such as the Internet. It is readily understood that although the various devices or modules in the electronic device 11 shown in Figure 11 communicate via bus 114, they can also communicate via a network or other means, wherein the network can include wireless networks, wired networks, and / or any combination of wireless and wired networks.
[0144] As needed, drive 1110 is also connected to input / output interface 115. Removable media 1111, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 1110 as needed, so that computer programs read from them can be installed into storage section 118 as needed.
[0145] When the above series of processes are implemented through software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as a removable medium 1111.
[0146] According to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product that, when run on a computer, causes the computer to implement the music generation method described in any of the foregoing embodiments. The computer program product includes a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 119, or installed from storage portion 118, or installed from ROM 112. When the computer program is executed by CPU 111, the music generation method of the embodiments of this disclosure is performed.
[0147] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0148] A computer-readable medium may be a computer-readable storage medium, a computer-readable signal medium, or any combination thereof.
[0149] Computer-readable storage media include, but are not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. A computer program is stored on the computer-readable storage medium that, when executed by a processor, implements the music generation method described in any of the foregoing embodiments.
[0150] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0151] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0152] In some embodiments, a computer program is also provided, comprising: instructions that, when executed by a processor, cause the processor to perform the music generation method of any of the above embodiments. For example, the instructions may be embodied in computer program code.
[0153] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0154] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0155] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0156] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.
Claims
1. A method for generating music, comprising: Receive first prompt information input by the user, the first prompt information including music description information; Based on the first prompt information, a first musical work is generated, the first musical work including first music; A second musical work is generated based on the user's adjustment operation on the first musical work, wherein the adjustment operation includes at least one of the adjustment operation on the first prompt information and the adjustment operation on the first musical work.
2. The music generation method according to claim 1, wherein, Based on the user's adjustment operations on the first musical piece, generating the second musical piece includes: The first prompt message is displayed on the interactive interface; Receive the user's modification operation on the first prompt information; Display a second prompt message based on the modification operation; Using a machine learning model, the second musical piece is generated based on the second prompt information.
3. The music generation method according to claim 2, wherein, Based on the user's adjustment operations on the first musical piece, generating the second musical piece includes: The interactive interface displays at least one recommendation prompt based on the first musical work, the recommendation prompt including at least one of the following: the theme of the music, the visuals of the music, the emotional type of the music, and the style of the music; The one or more recommended prompts selected by the user are used as part of the second prompt information.
4. The music generation method according to claim 1, wherein, Based on the user's adjustment operations on the first musical piece, generating the second musical piece includes: Based on the user's adjustment operations on the basic attributes of the first music, a second musical work is generated, wherein the basic attributes include the rhythm, melody, timbre, and harmony of the music.
5. The music generation method according to claim 4, wherein, The first piece of music includes multiple sections. Based on the user's adjustments to the basic attributes of the first piece of music, a second musical work is generated, including: The second musical work is generated based on the user's adjustment of the basic attributes of at least one of the multiple sections of the first music.
6. The music generation method according to claim 4, wherein, Based on the user's adjustments to the basic attributes of the first music, the generation of the second musical work includes: Display the audio tracks of the first music track, which includes multiple audio tracks; The second musical work is generated based on the user's adjustment operation on at least one of the multiple tracks of the first music.
7. The music generation method according to claim 5, wherein, Each of the multiple segments includes multiple beats, the rhythm of the first music includes the beats of the music, and the generation of the second musical work based on the user's adjustment operations on the basic attributes of the first music includes: Display the audio track of the first piece of music; The second musical work is generated based on the user's adjustment operations on the first music track. The adjustment operations on the first music include at least one of adjusting the time value of the beat and adjusting the combination rules between beats.
8. The music generation method according to claim 1, wherein, The first musical work also includes first lyrics. The generation of the first musical work based on the first prompt information includes: Using a machine learning model, the first music is generated based on the first prompt information; Using the machine learning model, the first lyrics are generated based on the first prompt information and the first music.
9. The music generation method according to claim 1, wherein, The first musical work is vocal music, including the first lyrics and the singer's voice. Based on the first prompt information, the first musical work is generated as follows: Using the machine learning model, the first music is generated based on the first prompt information; Receive a third prompt message input by the user, the third prompt message including the description information of the first lyrics; Using the machine learning model, the first lyrics are generated based on the third prompt information; The singer's voice is generated based on the first music and the first lyrics; The first musical work is synthesized based on the first music, the first lyrics, and the singer's voice.
10. The music generation method according to claim 9, wherein, Using the machine learning model, the lyrics for the first music are generated based on the third prompt information, including: Using the machine learning model, the first lyrics are generated based on the third prompt information and the first music.
11. The music generation method according to claim 9, wherein, Generating the singer's voice based on the first music and the first lyrics includes: Based on the user's voice, the singer's voice is generated using voice imitation technology; or Using the machine learning model, the singer's voice is generated based on the first music and the first lyrics.
12. The music generation method according to any one of claims 8 to 11, wherein, Based on the user's adjustment operations on the first musical piece, generating the second musical piece includes: The second music is generated based on the user's adjustment operations on the first music piece; The first lyrics are adjusted according to at least one of the melody, rhythm, and style of the second music to obtain the second lyrics; The second musical work is synthesized based on the second music and the second lyrics.
13. The music generation method according to any one of claims 12, wherein, The second musical work includes a second melody and second lyrics. The second melody comprises multiple sections, each section comprising multiple beats. The first lyrics are adjusted according to at least one of the rhythm and style of the second melody to obtain the second lyrics, which include: The first lyrics are adjusted based on at least one of the following: the time of each section and each beat of the second music, and the accent. This yields the second lyrics.
14. The music generation method according to claim 1, wherein, The first prompt message received upon receiving user input includes: The interactive interface provides at least one candidate prompt template. Each candidate prompt template corresponds to a complete music scene. The information of each candidate prompt template includes at least one of the following: music theme, music visuals, music emotional type, and music style. The prompt information in a candidate prompt template selected by the user is used as part or all of the first prompt information.
15. The music generation method according to claim 1, wherein, The first prompt message received upon receiving user input includes: The interactive interface provides at least one candidate prompt, each of which includes at least one of the following: the theme of the music, the visuals of the music, the emotional type of the music, and the style of the music. One or more candidate prompts selected by the user are used as part or all of the first prompt.
16. The music generation method according to any one of claims 1 to 11, wherein: The first musical work includes the title and cover art of the first musical piece; The second musical work includes the title and cover art of the second piece of music; The descriptive information of the music includes at least one of the following: the theme of the music, the imagery of the music, the emotional type of the music, and the style of the music.
17. A music generation device, comprising: The receiving module is configured to receive first prompt information input by the user, the first prompt information including music description information; The first generation module is configured to generate a first musical work based on the first prompt information, wherein the first musical work includes a first piece of music; The second generation module is configured to generate a second musical work based on the user's adjustment operation on the first musical work using the machine learning model, wherein the adjustment operation includes at least one of the adjustment operation on the first prompt information and the adjustment operation on the first musical work.
18. A music generation device, comprising: Memory; as well as A processor coupled to the memory, the processor being configured to execute the music generation method as described in any one of claims 1 to 16 based on instructions stored in the memory.
19. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the music generation method of any one of claims 1 to 16.
20. A computer program product, when run on a computer, causes the computer to implement the music generation method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Machines, systems and processes for automated music composition and generation employing linguistic and / or graphical icon based musical experience descriptors
CN108369799A
Song generation method and electronic equipment
CN110808019A
Audio synthesis method and device, electronic equipment and computer readable medium
CN111554267A
Lyric generation method and device, electronic equipment and computer readable storage medium
CN112632906A
Music generation method and device, medium and computing equipment
CN112785993A