Method for generating music data on the basis of text data, and related apparatus
By using an improved neural network model to process text data in a two-stage process, extracting musical symbol attributes and generating symbolic musical sequences, the problem of insufficient semantic information capture and controllability of generated music in existing technologies is solved, thus achieving higher quality and more personalized music generation.
Patent Information
- Application Number
- PCT/CN2025/087961
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-27
- Filing Date
- 2025-04-09
- Publication Date
- 2026-02-12
AI Technical Summary
Existing text-to-music generation technologies struggle to accurately capture the rich semantic information in text and integrate it into musical fragments. Furthermore, they lack the flexibility of human-computer collaboration and the controllability of generated music, resulting in insufficient quality and personalization capabilities of the generated music.
An improved neural network model is used to process text data in two stages: first, the attribute information of musical symbols is extracted, and then the musical symbol sequence is generated. Through a musical attribute information extractor and a data converter, the controllable generation of musical data is achieved.
It improves the interpretability and controllability of music data generation, enabling more precise control over music details and features, adapting to users' creative habits and processes, and generating music works that meet users' intentions.
Smart Images

Figure CN2025087961_12022026_PF_FP_ABST
Abstract
Description
Method and related apparatus for generating music data based on text data
[0001] The present application claims priority from the Chinese patent application No. 2024108531267 filed on June 27, 2024, and entitled "Method and apparatus for generating music data based on text data", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present disclosure relates to the field of artificial intelligence services, and more particularly, to generating music data based on text data. BACKGROUND
[0003] With the development of artificial intelligence technology, text-to-music generation has become an important technology. With this technology, users can create music through text descriptions. Such an interactive way is simple and easy to use, and it is convenient for users to participate in music creation. Therefore, the text-to-music generation technology can provide a friendly music creation channel for users who have no music theory or composition background.
[0004] In the industry and academia, a large number of researches have been devoted to the study of text input based music generation technology in the past. Some early text input based music generation systems consider using a regularized method to map text descriptions to music features, but due to the strictness of the rules, the scalability of the early systems is insufficient.
[0005] With the rapid development of deep learning technology in the fields of natural language processing and music signal processing, in recent years, some data-driven text-to-music generation models have appeared. These neural network models can capture the complex mapping relationship between text and music through learning a large number of "text-music" data pairs, thereby realizing end-to-end generation.
[0006] Although the text-to-music generation technology has made some progress, it still faces many challenges. For example, the current scheme still has difficulty in accurately capturing the rich semantic information in the text and fusing these semantic information into the music segments. At the same time, the current scheme still has difficulty in reasonably integrating the human creation process, and the flexibility is insufficient in the process of human-computer collaboration. In addition, it is also necessary to further improve the musicality, coherence and overall quality level of the generated music, and to enhance the controllability and individual customization ability of the generated results.
[0007] Therefore, the text-to-music generation technology still needs to be further optimized and improved. SUMMARY
[0008] The present disclosure provides a method for generating music data based on text data, a method for generating music data based on text data, an electronic device, a computer readable storage medium, and a computer program product.
[0009] The method for generating music data based on text data provided by the embodiments of the present disclosure comprises: determining description information of music symbols for describing music to be generated based on text data for describing the music to be generated, wherein the description information comprises at least one of the following two items: a music parameter set for describing music symbol attributes of the music to be generated, and a music symbol sequence for describing melody characteristics of the music to be generated; and determining music data in the form of symbolic music sequence of the music to be generated based on the description information.
[0010] The apparatus for generating music data based on text data provided by the embodiments of the present disclosure comprises: a music attribute information extraction model configured to: determine description information of music symbols for describing music to be generated based on text data for describing the music to be generated, wherein the description information comprises at least one of the following two items: a music parameter set for describing music symbol attributes of the music to be generated, and a music symbol sequence for describing melody characteristics of the music to be generated; and a music data converter configured to: determine music data in the form of symbolic music sequence of the music to be generated based on the description information.
[0011] In another aspect, the embodiments of the present disclosure provide a computer device, comprising:
[0012] a processor, a communication interface, a memory and a communication bus;
[0013] The processor, the communication interface and the memory complete mutual communication through the communication bus; the communication interface is an interface of a communication module;
[0014] The memory is configured to store a computer program and transmit the computer program to the processor; the processor is configured to call the computer program in the memory to execute the method in the above aspects.
[0015] In another aspect, the embodiments of the present disclosure provide a storage medium for storing a computer program, wherein the computer program is configured to execute the method in the above aspects.
[0016] In another aspect, the embodiments of the present disclosure provide a computer program product comprising a computer program, which, when running on a computer, causes the computer to execute the method in the above aspects.
[0017] The embodiments of the present disclosure better capture the description information of the input data at the granularity of music symbols, and use it as control information to control the generation of the final music data at a finer granularity. The embodiments of the present disclosure greatly improve the explainability and controllability of the music data generation process, and are beneficial to more finely control the detailed features of the generated music data. Attached Figure Description
[0018] Figure 1 is an example schematic diagram illustrating a scenario according to an embodiment of the present disclosure;
[0019] Figure 2 illustrates a schematic diagram of an application scenario according to an embodiment of the present disclosure;
[0020] Figure 3 shows a flowchart of a method for generating music data based on text data according to an embodiment of the present disclosure;
[0021] Figure 4 illustrates a schematic diagram of the operation "determining information of musical symbols for describing music data to be generated based on text data for describing music information or melody" according to an embodiment of the present disclosure.
[0022] Figure 5 illustrates a schematic diagram of text data for describing musical information or melody and information of musical symbols for describing musical data to be generated, according to an embodiment of the present disclosure.
[0023] Figure 6 shows another schematic diagram of a method for generating music data based on text data according to an embodiment of the present disclosure;
[0024] Figure 7 is a schematic diagram of obtaining a training dataset according to an embodiment of the present disclosure;
[0025] Figure 8 is a schematic diagram illustrating a training dataset for fine-tuning a music attribute extraction model according to an embodiment of the present disclosure;
[0026] Figure 9 shows a schematic diagram of an electronic device according to an embodiment of the present disclosure;
[0027] Figure 10 illustrates a schematic diagram of the architecture of an exemplary computing device according to an embodiment of the present disclosure;
[0028] Figure 11 shows a schematic diagram of a storage medium according to an embodiment of the present disclosure. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this disclosure more apparent, exemplary embodiments according to this disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments of this disclosure. It should be understood that this disclosure is not limited to the exemplary embodiments described herein.
[0030] In this specification and accompanying drawings, operations and elements that are substantially the same or similar are indicated by the same or similar reference numerals, and repeated descriptions of these operations and elements are omitted. Furthermore, in the description of this disclosure, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance or order.
[0031] To facilitate the description of the present disclosure, the following introduces concepts related to the present disclosure.
[0032] Symbolic music refers to a form of representing and recording musical works using musical symbols (such as musical symbols on the staff, rests, etc.). Unlike directly recording music as continuous digital audio signals, symbolic music adopts an abstract representation, decomposing musical works into a series of independent musical symbols with time duration, each of which optionally includes attribute information such as pitch, time value, and dynamics. The most common form of symbolic music is traditional staff notation, but it also includes other forms of musical notation, such as guitar tablature (Tablature), numbered musical notation (Numbered Musical Notation), etc.
[0033] The scheme provided by the embodiments of the present disclosure relates to technologies such as artificial intelligence and / or machine learning, which are specifically explained through the following embodiments.
[0034] First, the application scenario of the method of generating music data based on text data and the corresponding device and the like according to the embodiments of the present disclosure is described with reference to FIG. 1. FIG. 1 shows a schematic diagram of an application scenario 100 according to an embodiment of the present disclosure, in which a server 110 and a plurality of terminals 120 are schematically shown.
[0035] The neural network model of the embodiments of the present disclosure can be integrated in various computer devices, for example, any electronic device in the server 110 and the plurality of terminals 120 in FIG. 1. For example, the neural network model can be integrated in the terminal 120. The terminal 120 can be a mobile phone, a tablet computer, a notebook computer, a desktop computer, a personal computer (PC, Personal Computer), a smart speaker, or a smart watch, but is not limited thereto. For another example, the neural network model can also be integrated in the server 110. The server 110 can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN, Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in the present disclosure.
[0036] It can be understood that the device applying the neural network model of the embodiments of the present disclosure to perform inference can be a terminal, a server, or a system composed of a terminal and a server. The method for generating music data based on text data provided by the embodiments of the present disclosure can be executed on a terminal, a server, or a terminal and a server.
[0037] The artificial intelligence model provided by the embodiments of the present disclosure can also relate to artificial intelligence cloud services in the field of cloud technology. Among them, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and network in a wide area network or local area network to realize data calculation, storage, processing, and sharing. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. applied based on the cloud computing business model, which can form a resource pool, be used on demand, and be flexible and convenient. Cloud computing technology will become an important support. The background service of the technical network system needs a large amount of computing and storage resources, such as video websites, picture websites, drug research websites, and more portals. With the high development and application of the Internet industry, every item in the future may have its own identification mark and needs to be transmitted to the background system for logical processing. Different levels of data will be processed separately, and various industry data will need strong system support, which can only be realized through cloud computing.
[0038] It is worth noting that the terminal 110 and the server 120 according to the embodiments of the present disclosure both comply with the data protection principle, respect the data rights of users, and protect the data security and privacy of users. The terminal 110 and the server 120 according to the embodiments of the present disclosure will explicitly inform the user of the purpose, method, and scope of collecting, using, storing, transmitting, and deleting the user's data, and obtain the user's consent. The terminal 110 and the server 120 according to the embodiments of the present disclosure will take reasonable technical and management measures to prevent the user's data from being leaked, tampered with, damaged, or lost. The terminal 110 and the server 120 providers according to the embodiments of the present disclosure will regularly review and update the user's data, and timely delete expired or useless data. In addition, the cloud service provider using the embodiments of the present disclosure respects the user's rights of data access, correction, deletion, withdrawal of consent, complaint, and claim, and provides convenient channels and procedures for the user to effectively exercise these rights.
[0039] Further, the process of data analysis in terminal 120 or server 110 using artificial intelligence technology is carried out on the basis of the principles of legality, reasonableness and transparency. The data collected and processed by the artificial intelligence model according to the embodiments of the present disclosure is relevant, necessary and appropriate for the purpose of prediction, and does not contain any personal identity information or sensitive information. The neural network model according to the embodiments of the present disclosure adopts appropriate technical and organizational measures to protect the security and integrity of the data, and prevent unauthorized access, use or disclosure of the data.
[0040] The artificial intelligence-based neural network model according to the embodiments of the present disclosure will comply with relevant data protection regulations and ethical principles. The neural network model is trained based on a large amount of anonymized and de-identified data, without infringing the privacy rights of any individual or group. The artificial intelligence model has also been rigorously tested and evaluated to ensure that its output results are accurate and reliable, and will not cause any misleading or discrimination. The artificial intelligence model is designed only to improve service quality and customer satisfaction, and will not be used for any illegal or unethical purposes. In addition, the neural network model will be regularly reviewed and updated to adapt to changes in the data environment and legal norms.
[0041] Although traditional text-to-music generation systems / methods have made some progress, due to the presence of redundant information in the text, the output results of traditional text-to-music generation systems have inherent limitations in controllability. The industry and academia have proposed controlling music data generation through global music attributes or representations. The advantage of this method is that it can control the overall structure and style of the generated music to some extent. However, this method may overlook the element-level details of the generated music (such as the generation of specific musical symbols) due to the overemphasis on global music attributes or representations.
[0042] In recent years, content-based (such as melody or harmonic progression) audio music generation control methods have shown great potential. This method takes advantage of the similarities between symbolic music and natural language, and uses large language models (LLMs) to assist in the generation of symbolic music. However, this method ignores the significant structural differences between natural language and symbolic music. Specifically, in general, symbolic music revolves around a central melody (e.g., a refrain), and often has a higher degree of repetition, while natural language does not. Therefore, directly applying methods for simulating natural language to symbolic music generation can result in overly rigid generated music.
[0043] Secondly, the traditional text-to-music generation system / method does not take into account the importance of the creative process of human beings in music creation. One technique commonly adopted by many composers is to start with a set of simple musical symbols and first create a basic melody. As the creation progresses, more detailed musical symbols are gradually introduced to further develop and enrich the overall music work. However, the traditional text-to-music generation system / method cannot be integrated into such a creation process.
[0044] To this end, the embodiments of the present disclosure make improvements on the architecture of the neural network model. Compared with the single module structure of the traditional neural network model suitable for text-based music generation, the improved neural network model is composed of two decoupled modules: a music attribute information extractor and a music data converter. The music attribute information extractor is used to extract information describing musical symbols of music data to be generated from text data describing music information or melody, and the music data converter is used to determine music data in the form of symbolic music sequence based on the information describing musical symbols of music data to be generated. Of course, the present disclosure is not limited thereto.
[0045] Thus, the improved neural network model can perform two-stage processing on the input data (e.g., user input text data): a music information extraction stage and a controllable data generation stage.
[0046] Specifically, in the music information extraction stage, the improved neural network model can extract information describing musical symbols of music data to be generated from the user input text data describing music information or melody. Then, in the controllable data generation stage, the improved neural network model can generate the required music data in the form of symbolic music sequence through the music data converter based on the information describing musical symbols of music data to be generated.
[0047] Compared with the traditional single-stage scheme, i.e., directly generating the final output music data from the original user input data, the improved two-stage neural network model can better capture the information of the input data at the granularity of musical symbols and use it as control information to control the generation of the final music data at a finer granularity. The improved neural network model greatly improves the explainability and controllability of the music data generation process, which is conducive to more finely controlling the detailed features of the generated music data.
[0048] For the improved neural network model, the embodiment of the present disclosure provides a method for generating music data based on text data, so as to use the improved neural network model for inference. The method comprises: based on the text data for describing the music to be generated, determining the description information of the music symbol for describing the music to be generated, wherein the information of the music symbol for describing the music data to be generated comprises at least one of the following two items: a music parameter set for describing the music symbol attribute of the music to be generated, and a music symbol sequence for describing the melody characteristics of the music to be generated; and the description information, determining the music data in the form of the symbol music sequence of the music to be generated.
[0049] In the process of using the improved neural network for inference, by first determining the music attribute information and then predicting the music data based on the same, the controllability in the process of generating music data is significantly improved.
[0050] The method and device for generating music data based on text data according to the embodiment of the present disclosure are described below with reference to FIGS. 2 to 11.
[0051] FIG. 2 shows a schematic diagram of an application scenario according to an embodiment of the present disclosure.
[0052] Taking the user interaction interface 201 shown in FIG. 2 as an example, a text input control 202 can be generally provided for the user to input a descriptive text, such as “cheerful nocturne”, “melancholy waltz”, “like a sunny morning, children playing on the lawn”, etc. Once the user inputs the text description and clicks the generate button, the improved neural network model will extract the description information of the music symbol for describing the music to be generated based on the input text data. The generated description information will be displayed in the text display control 203 for the user to further confirm. Specifically, the description information is the music parameter set for describing the music symbol attribute of the music to be generated displayed in the text display control 203. Of course, the present disclosure is not limited thereto.
[0053] After the user clicks the confirmation button, the improved neural network model will controllably generate the corresponding music data based on the music parameters in the music parameter set, and display it in the music display control 204.
[0054] Optionally, the music data refers to the data for storing and representing the music information of the music to be generated in a structured form. When such music data is presented in the form of a symbolic music sequence, it will record and describe the music work (i.e. the music to be generated) in detail by mainly containing a series of symbolic elements. These symbolic elements include music symbols, beats, dynamic marks, expression marks, chord symbols and key signatures, etc.
[0055] Optionally, the music data includes, but is not limited to, at least one of the following: musical notation data, beat data, dynamic marking data, harmony data, tonality data, structure data. For example, the musical notation data includes pitch, duration and dynamics to reflect the specific information of the musical notation on the staff. The beat and rhythm data indicates the rhythmic structure of the music through time signature and rhythmic pattern. The dynamic marking data uses dynamic marking and hairpin marking to describe the dynamics changes of the music. The harmony and tonality data indicates the harmonic structure and tonality of the work through chord symbol and key signature. The structure data indicates the overall structure of the music through repeat sign and paragraph structure. Of course, the present disclosure is not limited thereto.
[0056] The data form of symbolic music can be represented in various ways, including traditional staff notation, guitar tablature, tablature, MIDI file and music XML, etc. Traditional staff notation is the most common form of symbolic music representation, suitable for human reading and playing; guitar tablature is used to represent the fingering and musical notation position of stringed instruments such as guitar; tablature uses numbers and symbols to represent pitch and rhythm, commonly used in music representation in some cultures; MIDI file is a digital music data format that transmits and plays music through MIDI protocol; music XML is an XML-based standardized music representation format for exchanging music data between different software and platforms. These symbolic music forms of music data are widely used in composition and arrangement, music education, music analysis, music performance and music software, etc.
[0057] In addition to static display of music data, a music playback engine can be integrated on the user interface to render the generated music data in real time as audio output. Users can not only visually view the generated music data, but also hear the generated music data, so as to determine whether the generated music data is what they want.
[0058] Through this text-to-music interaction mode, even ordinary users without professional music background can easily generate interesting music works quickly and controllably according to their own text description, greatly reducing the threshold and difficulty of music creation. In addition, as the creation progresses, users can manually adjust the descriptive information of the music data at the granularity of the musical notation to develop and enrich the overall music work.
[0059] In addition, as shown in FIG. 2, the user interaction interface 201 also optionally includes two input boxes: a melody input control 205 and a melody description input control 206. The melody input control 205 allows the user to directly input melody music symbols. The melody music symbols can be represented by standard music symbols such as C, D, E, etc., and include beat and time value marks such as quarter notes, eighth notes, etc. The melody description input control 206 provides a text input method, through which the user can describe the melody in words, for example, "C sharp, two eighth notes, and then an E note with a quarter note". These two input boxes are designed to provide flexible input options to meet the needs and habits of different users. Confirmation buttons are also designed on the user interaction interface 201 corresponding to these two input boxes to trigger the generation of description information, especially the generation of music symbol sequences for describing the melody characteristics of the music to be generated.
[0060] Thus, compared with traditional systems for generating music based on text, the user interaction interface 201 can better adapt to the habits of music creators. In addition, in order to further stimulate creative inspiration, the user interaction interface 201 also optionally provides an intelligent prompting function (not shown) to automatically generate melody suggestions or variation options based on the user's input, inspiring the user to expand on the original idea.
[0061] In addition, the user interaction interface 201 designs a visual feedback mechanism that will display the corresponding staff or tablature when the user inputs melody music symbols or word descriptions, helping the user to intuitively verify and adjust the input content. This heuristic creation method combines intelligent prompting and real-time feedback, not only improving the efficiency of music creation, but also providing a rich source of creative inspiration for users.
[0062] FIG. 3 shows a flowchart of a method 30 for generating music data based on text data according to an embodiment of the present disclosure.
[0063] The method 30 can be executed at a computer device, such as a terminal device or a server (such as the terminal 120 or the server 110 described in FIG. 1). The method 30 includes the following operations S301 to S302. Optionally, operation S301 is an operation in the music information extraction stage, and operation S302 can be performed in the conversion stage of the controllable data generation stage. Of course, the method 30 can also include more or fewer operations, and the present disclosure is not limited thereto.
[0064] In operation S301, based on the text data for describing the music to be generated, description information for describing the music symbols of the music to be generated is determined.
[0065] The description information includes at least one of a music parameter set for describing a music symbol attribute of the music to be generated and a music symbol sequence for describing a melody characteristic of the music to be generated.
[0066] Optionally, the music parameter set for describing the music symbol attribute of the music to be generated can be determined based on text data for describing the music to be generated. For example, the text data for describing the music to be generated refers to a set of strings capable of describing or explaining the (non-global) attribute of the music to be generated, including descriptive information of the music to be generated, etc. For example, an example of the text data for describing the music to be generated can be "like a sunny morning, children playing on the lawn" in FIG. 2.
[0067] Optionally, the music parameter set for describing the music symbol attribute of the music to be generated can also be referred to as "music attribute information" or "local music attribute", which refers to a set of elements capable of describing and defining the essential characteristics of the music to be generated at the granularity of the music symbol. For example, the music parameter set for describing the music symbol attribute of the music to be generated can include one or a combination of the following: instrument information, tonality information, speed information, number of notes information, style information, duration information, emotion information, rhythm information, genre information, intensity information, etc. Optionally, the "music attribute information" can be represented in the form of a set of key-value pairs. Of course, the present disclosure is not limited thereto.
[0068] Optionally, the music symbol sequence can be determined based on text data for describing the melody. For example, the text data for describing the melody refers to a set of strings capable of recording and analyzing the elements of the melody, such as the music symbol sequence or the text description of the melody, etc. For example, the text data for describing the melody can be "C sharp, two eighth notes, and one fourth note of E" or the melody music symbol "C, D, E" in FIG. 2. Of course, the present disclosure is not limited thereto.
[0069] Optionally, the music symbol sequence can also be referred to as "melody attribute information", which includes a plurality of music symbols arranged in time sequence, and each music symbol can be represented in the form of a triple. The mathematical representation of the triple can be (P, D, V), where P represents pitch, D represents time value, and V represents dynamics. Of course, the present disclosure is not limited thereto.
[0070] In one aspect of the present disclosure, the description information can be used to provide comprehensive and fine-grained control for the music to be generated. By manipulating these objective attributes, the embodiments of the present disclosure aim to shape the overall structure, emotional tone and dynamic characteristics of the music, thereby realizing a more customized and controllable music generation process.
[0071] Specifically, the description information can provide controllable adjustment space for the control of the subsequent operation S302. For example, changing the musical instrument can change the timbre; changing the tonality can affect the mood of the work; speeding up can enhance the dynamic; and increasing or decreasing the number of small sections can affect the overall length, etc. By reasonably adjusting, presenting and combining various music attribute parameters, the embodiments of the present disclosure can generate music data that is more in line with the user's intention, and provide more explicit control variables for the subsequent presentation of music data. In addition, as shown in FIG. 2, the user can also flexibly modify and adjust the extracted music attribute parameters, further optimize and customize the music attributes, and achieve the expected effect.
[0072] Optionally, S301 can be performed by any music attribute information extraction model. The music attribute information extraction model can extract description information of music symbols of music to be generated from text data based on neural network technology. For example, the music attribute information extraction model can be a large language model. Of course, the present disclosure is not limited thereto.
[0073] For example, the music attribute information extraction model can be LLaMA2 (Large Language Model Meta AI 2, second generation large language artificial intelligence model). LLaMA2 is good at capturing the deep semantics of text, so it is particularly suitable for the scene of describing music symbols of music to be generated in advance. In order to reduce the generation of redundant information, the traditional LLaMA2 can be fine-tuned to limit the output of LLaMA2 to only a set of music parameters for describing the music symbol attributes of the music to be generated or a music symbol sequence for describing the melody characteristics of the music to be generated, without generating irrelevant redundant information.
[0074] Those skilled in the art should understand that other similar models or algorithms (for example, BERT) can also be used to complete the operation of extracting information of music symbols for describing music data to be generated, depending on actual needs and application environments, and the present disclosure is not limited thereto. Some optional details of S301 will be described in detail with reference to FIGS. 4 and 5, and the present disclosure will not be repeated here.
[0075] In operation S302, based on the description information, the music data in the form of the symbolic music sequence of the music to be generated is determined.
[0076] The music data is represented in the form of symbolic music sequence, which precisely captures the core element information of the melody, rhythm, harmony, etc. of the music work to be generated. As described above, the symbolic music sequence adopts a computer-analyzable encoding format such as MIDI, ABC notation, etc., and represents the pitch, time value, dynamics and other attribute parameters of the music symbol in a series of numerical or text sequences. Compared with the traditional end-to-end processing of sound waveforms, symbolic music representation is more conducive to the interpretation of music attribute information. In addition, as shown in FIG. 2, the user can also edit the music symbols in the symbolic music sequence in the control 204 to achieve flexible control of different music attribute dimensions such as melody, harmony, rhythm, etc. At the same time, the symbolic music sequence also has a certain modularity and combinability, which provides more possibilities for music creation.
[0077] Optionally, the operation S302 includes: determining the word units corresponding to each music parameter in the music parameter set; generating a word sequence based on the word units corresponding to each music parameter in the music parameter set; and determining the music data in the form of symbolic music sequence of the music to be generated based on the word sequence. Thus, the embodiments of the present disclosure provide explicit control of the music data to be generated, so as to facilitate the interpretation of the influence of each music parameter on the determined music data.
[0078] Optionally, the operation S302 can be performed by any music data converter. The music data converter can utilize neural network technology to determine the music data in the form of symbolic music sequence based on the information of the music symbol used to describe the music data to be generated. For example, a transformer decoder can be used.
[0079] Optionally, in the process of generating music data, the embodiments of the present disclosure can also introduce a multi-head attention mechanism (MHA) in operation S302 to guide the generation of music data at a more fine-grained level. The multi-head attention mechanism can enable specific information (for example, any specific element in the music symbol sequence used to describe the melody characteristics of the music data to be generated) to be included in one or more time steps of music data generation, thereby achieving more accurate control of music data. Thus, it is ensured that the process of generating music data can effectively apply specific music elements and be more suitable for the user's creative purpose. Some optional details of operation S302 will be further described with reference to FIG. 6, and the present disclosure will not be repeated here.
[0080] The embodiments of the present disclosure can make the system understand how various attributes affect the music by using description information for describing the music symbols of the music to be generated. In this way, not only can more explicit guidance for generating music be provided, but also complex music expression and conversion can be realized in the creation process, providing creators with more extensive creation possibilities.
[0081] Thus, in the music information extraction stage, the embodiments of the present disclosure can extract relevant description information for describing the music symbols of the music to be generated from the natural language description input by the user through operation S301. Then, in the controllable data generation stage, the output of the required music data can be generated by the music data converter based on the extracted description information through operation S302.
[0082] Compared with the traditional single-stage scheme, i.e., directly generating the final output music data from the original user input data, the improved two-stage scheme can better capture the music symbol-specific information in the text data and use it as control information to control the generation of the final music data at a finer granularity. The improved neural network model greatly improves the explainability and controllability of the music data generation process, which is beneficial to more finely control the detailed features of the generated music.
[0083] For example, compared with the traditional scheme of using global music attributes to control music data generation, the embodiments of the present disclosure extract multiple music symbol granularity music parameters from text data, thereby realizing element-level control in the process of generating music data.
[0084] Compared with the scheme of directly applying the method for simulating natural language to symbolic music generation, the embodiments of the present disclosure generate music data in the form of symbolic music based on music symbol-specific description information, so that the music symbol-specific description information is always considered as a control variable in the process of generating music data, avoiding the generated music data deviating from the central melody.
[0085] Compared with the traditional scheme of directly generating the final output music data based on the original user input data, the embodiments of the present disclosure can also provide visual and editable controls in the process of generating music data, thereby integrating the artificial intelligence assisted composition technology into the creation process of the composer, realizing the technical effect of gradually developing and enriching the overall music work.
[0086] Next, some details in operation S301 will be further described with reference to FIG. 4 and FIG. 5.
[0087] FIG. 4 shows a schematic diagram of S301 according to an embodiment of the present disclosure. FIG. 5 shows a schematic diagram of text data for describing music information or melody and information for describing music symbols of music data to be generated according to an embodiment of the present disclosure.
[0088] Optionally, the text data for describing music information (shown as "text description of music information" in FIG. 4) can come from the text input control 202, in which a user can input a description of the music data to be generated, such as "like a sunny morning, children playing on the lawn".
[0089] Optionally, in mathematical expression, the task of determining the set of music parameters of attributes of music symbols for describing the music data to be generated based on the text data for describing music information can be represented as {I b , O}, where I b is the text data for describing the music to be generated input by the user (or a set of text data for describing the music to be generated), and O is the set of music parameters of attributes of music symbols for describing the music to be generated predicted by the music attribute information extraction model.
[0090] where, for each instance i b ∈ I b input by the user, it can be paired with a combination of k music parameters . Optionally, k = 10. Each instance i b , i.e., the text data related to music information input by the user by clicking the confirmation button once. As shown in FIG. 5, the output data for each instance i b can be represented in the form of key-value pairs. For example, the output data can be constrained as: "instrument information: {piano}; tonality information: {C major}; speed information: {fast}; number of measures information: {16}; style information: {classical}; length information: {3 minutes}; emotion information: {happy}; rhythm information: {4 / 4 beat}; genre information: {sonata}; intensity information: {weak}". This representation method ensures that the system can accurately identify and utilize various music attribute parameters when processing and generating music works, thereby achieving efficient and user-desired output.
[0091] Optionally, the text data for describing melody comes from the melody input control 205 or the melody description input control 206, in which a user can directly input melody music symbols (e.g., standard music symbols, etc.) or provide a text description of a simple melody (e.g., "C sharp").
[0092] For example, assume that the melody input control 205 is triggered. At this time, in response to the text data for describing the melody being a first music symbol sequence input, the first music symbol sequence is taken as a music symbol sequence for describing the melody feature of the music to be generated. Alternatively, the first music symbol sequence is extended into a second music symbol sequence of a specific length as a music symbol sequence for describing the melody feature of the music to be generated. As shown in FIG. 4, this process of “extension” can be performed by the music attribute information extraction model. Thus, such a music symbol sequence generated will be directly used for the generation of controllable music data in operation S302.
[0093] For example, the music symbol sequence input for describing the melody can be a plurality of triplets arranged in time sequence, such as {(60, 64, 70), (58, 64, 68), …, (55, 64, 66), (57, 64, 65), (54, 64, 63)}, where each triplet represents a feature of a music symbol. For example, the first number of the triplet can represent the pitch, ranging from 0 to 127; the second number can represent the time value, ranging from 0 to 127; and the third number can represent the dynamics, ranging from 0 to 31. Of course, the present disclosure is not limited thereto.
[0094] For the aforementioned “specific length”, the specific length is used to identify the number of music symbols in the second music symbol sequence, or in other words, at least a music symbol sequence of the specific length is required to accurately generate the music data. When the number of music symbols in the first music symbol sequence is less than the specific length, the computer device can perform the extension of the music symbols on the first music symbol sequence to obtain the second music symbol sequence.
[0095] For example, assume that the melody description input control 206 is triggered. At this time, in response to the text data for describing the melody being a natural language description of the melody, a music symbol sequence for describing the melody feature of the music to be generated is generated based on the natural language description of the melody.
[0096] Alternatively, as shown in FIG. 5, assume that an example of the text data for describing the melody input by the melody description input control 206 is: “The central melody starts in the middle register and gradually descends, occasionally jumping to high notes, forming a melody profile with a mix of high and middle notes. In terms of intensity variation, the music starts with soft intensity and gradually reaches a climax, creating a sense of tension and dynamics. There are also moments of alternating forte and piano, similar to the rhythm of a heartbeat, while the overall intensity remains relatively stable.” Of course, the present disclosure is not limited thereto.
[0097] Alternatively, in mathematical expression, the task of determining the music symbol sequence based on the text data for describing the melody can be represented as {I m , M}. Each instance im That is, the user inputs text data related to the melody to the system by clicking the confirmation button in the dashed box once. For each instance i m The output can also be predicted by the music attribute information extraction model, which can be represented in the form of a set of triples in the output data M. Optionally, the length of the set of triples is fixed. The set of triples is actually simple melody fragment data, which includes a combination of multiple melody musical symbols. Of course, the present disclosure is not limited thereto.
[0098] In an aspect of an embodiment of the present disclosure, the optional operation of "determining a sequence of musical symbols for describing the melody characteristics of the music data to be generated based on the text data for describing the melody" can also include: determining prompt input data based on the text data for describing the melody, the prompt input data including at least one of: an initial musical symbol, a format constraint of the output data, an expected distribution of the sequence of musical symbols in different tonal regions, a variation pattern of the sequence of musical symbols in time value, a variation pattern of the sequence of musical symbols in dynamics; and determining a sequence of musical symbols for describing the melody characteristics of the music data to be generated based on the prompt input data using a large language model.
[0099] This process draws on traditional composition techniques, considering the generation process of music data as a gradual process. In this process, a sequence of musical symbols for describing the melody characteristics of the music data to be generated is first obtained as a basic skeleton, and then embellishment notes are gradually added.
[0100] Specifically, assuming that the music attribute information extraction model is a large language model, the data input thereto is prompt input data for describing the melody after specific adjustment to guide the music attribute information extraction model to control the process in a finer granularity. A sequence of musical symbols describing the target melody characteristics can also be constructed as part of the prompt input data of the music attribute information extraction model.
[0101] For example, the prompt input data input to the music attribute information extraction model can contain a small number of initial musical symbols as the starting point of the melody generation. Then the music attribute information extraction model will infer a reasonable musical progress based on these musical symbols, and generate a sequence of musical symbols for describing the melody characteristics of the music data to be generated accordingly.
[0102] In addition, the prompt input data can also add the expected distribution of the sequence of musical symbols in different tonal regions, and the variation patterns of time value and dynamics as specific constraints to regulate the overall profile of the generated melody.
[0103] The music grammar rules such as mode, harmony, rhythm, and the like can be incorporated into the prompt input data. These parameters will ensure the structural and logical consistency of the generated content. In addition, style vocabulary and emotional descriptors can also be included in the prompt to guide the music attribute information extraction model to select appropriate musical symbols, rhythms, and dynamics.
[0104] In addition, in response to the text data describing the music to be generated being text data describing a melody, a format constraint of the output data is added to the prompt input data. That is, the format of the output data of the music attribute information extraction model can be limited by the input data input to the music attribute information extraction model. Specifically, the format of the output data will be limited to a sequence of triples, each triple containing musical symbol, time value, and dynamic information.
[0105] In actual application, as shown in FIG. 2, after the user inputs the initial musical symbol, the music attribute information extraction model will generate a sequence of triples based on the constraints in the prompt. Subsequently, the sequence can be presented in the form of standard musical notation or the key-value pair described above for further editing and confirmation by the user. In this way, by decomposing the composition task, the generation process of the music data is effectively regulated.
[0106] Next, some optional details in operation S302 are further described with reference to FIG. 6. FIG. 6 shows another schematic diagram of a method for generating music data based on text data according to an embodiment of the present disclosure.
[0107] Operation S301 is schematically described in the left box in FIG. 6, and the process is consistent with the process described with reference to FIGS. 2 to 5, which will not be described again in the present disclosure. Next, as shown in FIG. 6, in operation S302, a multi-head attention mechanism (MHA) can be introduced to guide the generation of music data at a more fine-grained level.
[0108] Optionally, when the description information includes the music parameter set and the musical symbol sequence, operation S302 can further include: determining, based on both the music parameter set and the musical symbol sequence, a probability distribution corresponding to each of a predetermined number (e.g., n) of time-sequentially arranged triples, wherein the probability distribution corresponding to each triple depends on the probability distribution of the previous triple; and determining, based on the probability distribution corresponding to each of the plurality of triples, the music data in the form of the symbolic music sequence of the music to be generated through a sampling operation.
[0109] More specifically, for each of the predetermined number of sequentially arranged triplets, a probability distribution of the triplet is determined based on the probability distribution of each of the preceding triplets of the triplet, the set of music parameters, and the sequence of music symbols. In a sequentially arranged order, for an i-th triplet of the plurality of triplets, the preceding triplets of the i-th triplet can be the first i-1 triplets of the plurality of triplets in the sequentially arranged order.
[0110] In particular, mathematically, the task of the controllable music generation stage corresponding to operation S302 can be defined as {A, M, Y}. Where A is the set of music parameters used to describe the attributes of the music symbols of the music to be generated, M is the sequence of music symbols used to describe the melody characteristics of the music to be generated, and Y is the dataset of symbolic music. That is, Y includes a plurality of triplets, which cover all individual music symbols and their characteristics within the human hearing threshold that are suitable for music generation. Each triplet, as described above, includes three numbers, which are pitch, duration, and dynamics, respectively.
[0111] For each instance y of a triplet in Y, the corresponding input data, that is, the set of music parameters used to describe the attributes of the music symbols of the music to be generated, can be represented as and the sequence of music symbols used to describe the melody characteristics of the music to be generated can be represented as m e M, the relationship of the three can be expressed mathematically as a probability distribution p(y|a, m) shown in formula (1).
[0112] That is, the probability of generating the entire instance y is equal to the product of the probability of generating each element y i in turn, where the generation of each y i is based on the probability distribution of all previously generated elements (i.e., the preceding triplets), the set of music parameters a used to describe the attributes of the music symbols of the music to be generated, and the sequence of music symbols m used to describe the melody characteristics of the music to be generated.
[0113] The task of the controllable music generation stage as shown in formula (1) can be implemented by a plurality of melody-guided transformer layers (Melo Transformer Layers) cascaded. The melody-guided transformer is a neural network layer designed for music data generation, which is improved on the basis of the standard transformer layer to better adapt to the characteristics of the music data generation task. Specifically, each melody-guided transformer layer can adjust the output of the previous melody-guided transformer layer with the information used to describe the music symbols of the music to be generated as attention information.
[0114] In each melody guide transformer layer, a special attention mechanism is optionally integrated to integrate the music parameter set a describing the attributes of the music symbol to be generated and the music symbol sequence m describing the melody characteristics of the music data to be generated, so as to guide the output of the melody guide transformer layer.
[0115] Taking FIG. 6 as an example, the melody guide transformer layer can include a first perceiver, a second perceiver and a fully connected layer; the first perceiver takes the music symbol sequence as a query vector, and takes the output of the previous melody guide transformer layer as a key vector and a value vector; the second perceiver takes the weighted sum of the output of the first perceiver and the output of the previous melody guide transformer layer as a query vector, a key vector and a value vector; and the fully connected layer takes the output of the second perceiver as an input, and the output of the fully connected layer is the output of the melody guide transformer layer.
[0116] Specifically, assuming that the input of the nth melody guide transformer layer is z n , and the output is z n+1 . Then the operation performed by the nth melody guide transformer can be mathematically expressed as formula (2).
[0117] Wherein, the first perceiver MHA1 and the second perceiver MHA2 are both cross-attention mechanism-based perceivers in the melody guide transformer, and Q, K and V are respectively a query vector, a key vector and a value vector. In each melody guide transformer layer, the music symbol sequence m describing the melody characteristics of the music data to be generated is added to the process of generating the music data through MHA1. The parameter η adjusts the influence degree of the music symbol sequence m describing the melody characteristics of the music data to be generated on the output z n+1 .
[0118] Specifically, the generation process of the nth melody guide transformer layer can be decomposed into the following processes. First, the input data z n of the nth layer is taken as a key vector (K) and a value vector (V), and the music symbol sequence m describing the melody characteristics of the music data to be generated is taken as a query vector. As shown in FIG. 5, the key vector K=z n , the value vector V=z n , and the query vector Q=m are jointly input into the first multi-head attention perceiver MHA1 to integrate the music symbol sequence m describing the melody characteristics of the music data to be generated into the current music data generation process. The output of the first multi-head attention perceiver MHA1 is MHA1 (Q=m, K=z n , V=z n ).
[0119] Next, based on the output of the first multi-head attention perceptron MHA1 and the nth layer input data z n Construct the input data for the second multi-head attention perceptron MHA2. Specifically, the query vector Q, key vector K, and value vector V of the second multi-head attention perceptron MHA2 can all be designed to be equal to ηz. n +(1-η)MHA1(Q=m,K=z n V=z n Where η is the value used to balance the output of the first multi-head attention perceptron MHA1 and the input data z of the nth layer. n The larger the coefficient of weight η, the weaker the influence of the music symbol sequence m used to describe the melodic characteristics of the music data to be generated; conversely, the stronger the influence of the music symbol sequence m used to describe the melodic characteristics of the music data to be generated.
[0120] The output data of the second multi-head attention perceptron, MHA2, will be further fused using a feedforward neural network (FFN). Optionally, the FFN can consist of two fully connected layers with a non-linear activation function (such as ReLU) in between. By introducing a non-linear transformation, the expressive power of the music data converter can be further enhanced.
[0121] In one optional embodiment, the music data converter may include more than eight cascaded melody-guided converter layers, wherein the η of the first four and last four melody-guided converter layers is greater than 0 and less than 1, while the η of the intermediate melody-guided converter layers is equal to 1. This design can further assist the music data converter in focusing more on the input of the previous layer and the music parameter set a of the music symbols used to describe the attributes of the music data to be generated, avoiding the reduction in the diversity of music data caused by focusing too much on the music symbol sequence m used to describe the melody characteristics of the music data to be generated. Of course, this disclosure is not limited thereto.
[0122] Next, the training process of the music attribute information extraction model and music data converter according to embodiments of the present disclosure will be further described with reference to Figures 7 and 8. Figure 7 is a schematic diagram of obtaining a training dataset according to an embodiment of the present disclosure; Figure 8 is a schematic diagram showing a training dataset for fine-tuning the music attribute extraction model according to an embodiment of the present disclosure.
[0123] In general, the training of the music data converter includes: extracting a music data sample from original training data; extracting a value of a preset music parameter corresponding to the music data sample from a description of the music data sample, taking the value of the preset music parameter as a true value of a sample music parameter set used to describe music symbol attributes of the corresponding music data sample; generating a first training sample set based on the music data sample and the true value of the corresponding sample music parameter set; determining a music symbol belonging to a central melody in the music data sample based on a melody tolerance and a chord tolerance, and taking the central melody as a true value of a sample music symbol sequence used to describe melody characteristics of the corresponding music data sample; generating a second training sample set based on the music data sample and the true value of the corresponding sample music symbol sequence; and training the music data converter by using the first training sample set and the second training sample set.
[0124] Optionally, the training of the music attribute information extraction model includes: determining text training data used to describe music information based on the true value of the sample music parameter set; generating a third training sample set based on the text training data used to describe music information and the true value of the corresponding sample music parameter set; determining text training data used to describe a melody based on the true value of the sample music symbol sequence; and generating a fourth training sample set based on the text training data used to describe the melody and the true value of the corresponding sample music symbol sequence; and fine-tuning the music attribute information extraction model by using the third training sample set and the fourth training sample set.
[0125] Specifically, the training data set collection process of the music attribute information extraction model and the music data converter mainly includes three links of MIDI collection, music attribute extraction and simple melody extraction. Since there are limited large-scale publicly available symbolic music data sets, the embodiments of the present disclosure integrate existing open source data sets to construct a large database as an original training data set. The sources of the training data set include EMOPIA, POP909, MetaMIDIDataset, SymphonyNet and LMD_full, etc. After deduplication processing, 16-bar fragments are randomly sampled from each MIDI file as samples. For files with a length less than 16 bars, the entire file is taken as a music data sample in the training process. Finally, about 1.5 million music data samples are collected.
[0126] Optionally, each music data sample is further processed into REMI-like representation. Specifically, the musical notes are represented as integers in the range of [0, 127], each integer corresponding to a specific pitch. The instruments are represented as integers in the range of [0, 128], each integer representing a specific instrument or timbre. Other musical elements are also digitally encoded in a similar manner. This data representation scheme facilitates the training and testing of the music attribute information extraction model and the music data converter.
[0127] As shown in FIG. 8, one of the labels for each music data sample - the music parameters describing the musical notation attributes of the music data to be generated, is collected by applying certain rules to the sample. These rules involve multiple music parameters such as tempo, tonality, harmony, instrument combination, etc., aiming to extract the core features of the music. This process provides the necessary high-level abstraction for controllable music generation, enabling the music attribute information extraction model to learn the overall style and structure of the music through the training data, and providing a clear control goal for the process of generating music data.
[0128] Correspondingly, the text training data for describing music information corresponding to the music parameters is designed, so as to take a pair of text training data for describing music information - music parameters describing the musical notation attributes of the music data to be generated as an instance of the training data. This instance will be used to fine-tune the music attribute information extraction model.
[0129] Optionally, the process is as follows: first, a template is created for each music parameter, in which the value of the music parameter is represented by a placeholder. This module can be represented in the form of xml as: "instrument information: {}; tonality information: {}; tempo information: {}; number of measures information: {}; style information: {}; duration information: {}; emotion information: {}; rhythm information: {}; genre information: {}; intensity information: {}". Subsequently, the values of the music parameters corresponding to the samples collected in FIG. 8 are filled into the placeholders {} of the prepared template. Then, the templates describing different combinations of music parameters are integrated together.
[0130] Specifically, in order to make the text training data for describing music information closer to real user input, additional modifications can be made to the music parameter descriptive prompt input data obtained during the integration process. Considering that user input in the real world may not contain all music parameters, the values of part of the music parameters are randomly hidden in half of the text training data for describing music information. For the other half of the text training data for describing music information, a large language model can be used to modify the structure and expression of these text training data for describing music information without changing the content delivered, so as to enhance the diversity of the training data.
[0131] To enrich the diversity of the text training data for describing music information-music parameters for describing the attributes of the music symbols of the music data to be generated and avoid the long tail problem, the embodiments of the present disclosure can randomly create instances based on the music parameters extracted in FIG. 8 and their values to ensure the balance of the data set, in which the occurrence probability of each value of each music parameter is close. This strategy helps to improve the performance of the model on rare music parameter values and enhance its generalization ability.
[0132] The label of the training sample of the music data converter not only includes the true value of the music parameter set for describing the attributes of the music symbols of the music data to be generated, but also includes the true value of the music symbol sequence for describing the melody characteristics of the music data to be generated. The collection of the true value is relatively difficult because there is no such label information in the existing training data set.
[0133] To this end, in one aspect of the embodiments of the present disclosure, the collection process of the true value of the music symbol sequence for describing the melody characteristics of the music data to be generated corresponding to the music data sample adopts a heuristic method, which takes pitch as reference information to determine whether a certain music symbol belongs to the central melody. This process can be summarized as follows: first, identify the continuous music symbol group with a duration of 0 (music symbols played simultaneously) and keep only the highest pitch in each group. Then, traverse the entire music symbol list.
[0134] Optionally, the embodiments of the present disclosure introduce two special thresholds: melody tolerance and chord tolerance. The preset values of the two thresholds are set to minor seventh and major sixth, respectively. For the first music symbol, if the pitch difference between the second music symbol and the first music symbol is greater than or equal to the chord tolerance, the first music symbol is marked as a chord, and the second music symbol becomes the first music symbol marked as the central melody. On the contrary, if the pitch difference is within the chord tolerance, the second music symbol is marked as a chord, and the first music symbol becomes the central melody music symbol.
[0135] Then, after determining the first music symbol belonging to the central melody, in the traversal process, the average pitch of the music symbols marked as the central melody is calculated with the length of the last measure as the window. If the total length corresponding to the currently marked central melody music symbol is less than one measure, the average pitch of all currently marked music symbols is calculated. If the pitch of the current first unmarked music symbol is higher than the average pitch, it continues to be marked as the central melody music symbol. If the pitch of the current first unmarked music symbol is lower but the pitch difference is within the melody tolerance, the current first unmarked music symbol can also be marked as the central melody music symbol.
[0136] In addition, if the pitch difference between the most recently marked center melody note and the first currently unmarked note is less than the chord tolerance, the first currently unmarked note is marked as a center melody note. This is because if the most recently marked center melody note needs to suddenly jump to the first currently unmarked note, even if the composer intends to make the first currently unmarked note part of the center melody, it is difficult for the listener to perceive the continuity of the melody. The essence of a melody is a series of notes with similar pitches, which is in contrast to chord tones (which usually have a large pitch difference from the melody). Based on this prior knowledge, the embodiments of the present disclosure take a more complex strategy when judging a note that is below the average pitch and exceeds the melody tolerance: only when the pitch difference between the note and the most recently marked center melody note, and the pitch difference between the note and the next note, are both less than the chord tolerance, can the note be marked as a center melody note. This rule attempts to capture the "jump in" phenomenon in melodies, i.e., a melody can temporarily deviate from its main range, but will quickly return thereafter. Thus, the embodiments of the present disclosure simulate, to some extent, the way a listener identifies a melody, i.e., a sequence of relatively high-pitched notes with smooth pitch changes is more likely to be considered a center melody.
[0137] At this point, the embodiments of the present disclosure have collected the identifiers, durations, and dynamics information of all notes identified as part of the center melody, and organized these information into a sequence of triples as the ground truth of the melody feature of the music data to be generated.
[0138] After obtaining the ground truth of the melody feature of the music data to be generated, text training data for describing the melody in the training sample also needs to be designed. This process can be simply described as: based on the ground truth of the melody feature of the music data to be generated, using a large language model to describe its characteristics from the pitch range and dynamics range, and then generating text data for describing the melody. This method converts abstract music note sequences into natural text language. The generated description is paired with the melody ground truth after screening to form a training sample of "text sample data for describing the melody - music note sequence sample for describing the melody feature of the music data to be generated".
[0139] At this point, the process of obtaining the training data set for training the music attribute information extraction model and the music data converter according to the embodiments of the present disclosure has been described in detail, and the process of training the music attribute information extraction model and the music data converter based on such a training sample set is introduced next.
[0140] For the training process of the music data converter, both the first sample dataset of paired “ground truth-music data sample of music parameter set for describing the attribute of music symbol to be generated” and the second sample dataset of paired “ground truth-music data sample of music symbol sequence for describing the melody feature of music data to be generated” can be used to train the music data converter. Of course, the present disclosure is not limited thereto.
[0141] For the fine-tuning process of the music attribute information extraction model, the first training sample set for fine-tuning the music attribute information extraction model can be “text sample data for describing music information-ground truth of music parameter set for describing the attribute of music symbol to be generated”. Meanwhile, the second training sample set for fine-tuning the music attribute information extraction model can also be “text sample data for describing melody-ground truth of music symbol sequence for describing the melody feature of music data to be generated”. The final fine-tuning goal is to minimize the cross-entropy loss on the two training sets. The cross-entropy loss measures the difference between the output of the fine-tuned music attribute information extraction model and the target, and its minimization means that the fine-tuned music attribute information extraction model’s prediction on the two training sets is increasingly close to the true label.
[0142] In addition, throughout the fine-tuning process, the embodiments of the present disclosure can also use the Low-Rank Adaptation (LoRA) technology to optimize efficiency and reduce memory occupation. The LoRA technology realizes parameter efficient fine-tuning by learning the low-rank update of the original weight matrix, significantly reduces the number of trainable parameters, and at the same time maintains good performance. This technology is particularly suitable for the domain adaptation task of fine-tuning large language models (e.g., fine-tuning the music attribute information extraction model).
[0143] In addition, the first training sample set and the second training sample set are guided by different prompt input data, and both training processes are used for fine-tuning of the same music attribute information extraction model. This learning strategy is conducive to the music attribute information extraction model learning more general feature representation, enhancing the generalization ability of the music attribute information extraction model. The design of different prompt input data enables the music attribute information extraction model to distinguish between task types and generate corresponding outputs accordingly. Of course, the present disclosure is not limited thereto.
[0144] Further, the embodiment of the present disclosure also provides a device for generating music data based on text data, the device comprising: a music attribute information extraction model configured to determine description information of a music symbol for describing to-be-generated music based on text data for describing the to-be-generated music, wherein the description information comprises at least one of the following two items: a music parameter set for describing a music symbol attribute of the to-be-generated music, and a music symbol sequence for describing a melody characteristic of the to-be-generated music; and a music data converter configured to determine music data in the form of a symbolic music sequence of the to-be-generated music based on the description information.
[0145] In a possible implementation, the music attribute information extraction model is further configured to:
[0146] determine the music parameter set based on the text data for describing the music information; or
[0147] determine the music symbol sequence based on the text data for describing the melody.
[0148] In a possible implementation, the music parameter set comprises a combination of one or more of the following items: instrument information, tonality information, speed information, number of notes information, style information, duration information, emotion information, rhythm information, genre information, intensity information.
[0149] In a possible implementation, the music attribute information extraction model is further configured to:
[0150] in response to the text data for describing the melody being a first music symbol sequence as input, take the first music symbol sequence as the music symbol sequence, or expand the first music symbol sequence into a second music symbol sequence of a specific length as the music symbol sequence; or
[0151] in response to the text data for describing the melody being a natural language description of the melody, generate the music symbol sequence based on the natural language description of the melody.
[0152] In a possible implementation, the music attribute information extraction model is further configured to:
[0153] determine prompt input data based on the text data for describing the melody, the prompt input data comprising at least one of the following items: an initial music symbol, a format constraint of output data, an expected distribution of the music symbol sequence in different sound regions, a variation pattern of the music symbol sequence in time value, and a variation pattern of the music symbol sequence in intensity; and
[0154] determine the music symbol sequence based on the prompt input data by using a large language model.
[0155] In a possible implementation, when the description information comprises the set of music parameters, the music data converter is further configured to:
[0156] determine a wordpiece corresponding to each music parameter in the set of music parameters;
[0157] generate a wordpiece sequence based on the wordpiece corresponding to each music parameter in the set of music parameters; and
[0158] determine music data in the form of a symbolic music sequence of the music to be generated based on the wordpiece sequence.
[0159] In a possible implementation, when the description information comprises the set of music parameters and the symbolic music sequence, the music data converter is further configured to:
[0160] determine, based on both the set of music parameters and the symbolic music sequence, a probability distribution corresponding to each of a predetermined number of time-sequentially arranged triplets, wherein the probability distribution corresponding to each of the triplets depends on a probability distribution of a preceding triplet; and
[0161] determine, based on the probability distribution corresponding to each of the triplets, the music data in the form of the symbolic music sequence of the music to be generated through a sampling operation;
[0162] wherein each of the triplets comprises a pitch, a duration, and a dynamics corresponding to a single music symbol.
[0163] In a possible implementation, for each of the triplets, the music data converter is further configured to:
[0164] determine the probability distribution of the triplet based on the probability distribution of each preceding triplet, the set of music parameters, and the symbolic music sequence.
[0165] In a possible implementation, the music data converter is further configured to:
[0166] determine, based on the description information, the music data in the form of the symbolic music sequence of the music to be generated by using a plurality of melody guide converter layers in cascade;
[0167] wherein each of the melody guide converter layers adjusts an output of a preceding melody guide converter layer by taking information of the music symbol used to describe the music data to be generated as attention information.
[0168] In a possible implementation, the melody guide converter layer comprises a first perceptron, a second perceptron, and a fully connected layer.
[0169] The first perceiver takes the music symbol sequence as a query vector, and takes the output of the previous melody guide converter layer as a key vector and a value vector.
[0170] The second perceiver takes a weighted sum of the output of the first perceiver and the output of the previous melody guide converter layer as a query vector, a key vector and a value vector.
[0171] The fully connected layer takes the output of the second perceiver as input, and the output of the fully connected layer is the output of the melody guide converter layer.
[0172] In a possible implementation manner,
[0173] The music attribute information extraction model is further configured to: based on the text data describing the to-be-generated music, determine, by using the music attribute information extraction model, description information of music symbols describing the to-be-generated music; and
[0174] The music data converter is further configured to: based on the description information, determine, by using the music data converter, music data in the form of a symbolic music sequence of the to-be-generated music.
[0175] In a possible implementation manner, the apparatus further includes a training unit, and the training unit is configured to:
[0176] extract a music data sample from original training data;
[0177] extract, from description of the music data sample, a value of a preset music parameter corresponding to the music data sample, and take the value of the preset music parameter as a real value of a sample music parameter set used to describe music symbol attributes of the corresponding music data sample;
[0178] generate a first training sample set based on the music data sample and the real value of the corresponding sample music parameter set;
[0179] determine, based on a melody tolerance and a chord tolerance, music symbols belonging to a central melody in the music data sample, and take the central melody as a real value of a sample music symbol sequence used to describe melody characteristics of the corresponding music data sample;
[0180] generate a second training sample set based on the music data sample and the real value of the corresponding sample music symbol sequence;
[0181] train the music data converter by using the first training sample set and the second training sample set.
[0182] In a possible implementation manner, the training unit is further configured to:
[0183] determine text training data for describing music information based on the ground truth of the sample music parameter set;
[0184] generate a third training sample set based on the text training data for describing music information and the ground truth of the corresponding sample music parameter set;
[0185] determine text training data for describing melody based on the ground truth of the sample music symbol sequence;
[0186] generate a fourth training sample set based on the text training data for describing melody and the ground truth of the corresponding sample music symbol sequence;
[0187] fine-tune the music attribute information extraction model by using the third training sample set and the fourth training sample set.
[0188] The present disclosure also provides an electronic device for implementing the method according to the embodiments of the present disclosure. FIG. 9 shows a schematic diagram of an electronic device 2000 according to an embodiment of the present disclosure.
[0189] As shown in FIG. 9, the electronic device 2000 can include one or more processors 2010 and one or more memories 2020. The memory 2020 stores computer readable code which, when executed by the one or more processors 2010, can perform the method described above.
[0190] The processor in the embodiments of the present disclosure can be an integrated circuit chip with a signal processing capability. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The methods, operations and logical block diagrams disclosed in the embodiments of the present disclosure can be implemented or executed by the processor. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor or the like, which can be of X86 architecture or ARM architecture.
[0191] In general, the various example embodiments of the present disclosure can be implemented in hardware or special-purpose circuitry, software, firmware, logic, or any combination thereof. Certain aspects can be implemented in hardware, while other aspects can be implemented in firmware or software which can be executed by a controller, microprocessor or other computing device. While aspects of the present disclosure have been illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that the blocks, devices, systems, techniques or methods described herein can be implemented in hardware, software, firmware, special purpose circuitry, general purpose hardware or controller or other computing devices which supports combinations of hardware and software, or it can be implemented in some other combination of hardware and software.
[0192] For example, the method or the apparatus according to the embodiments of the present disclosure can also be implemented by means of the architecture of the computing device 3000 shown in FIG. 10. As shown in FIG. 10, the computing device 3000 can include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port connected to a network 3050, an input / output component 3060, a hard disk 3070, etc. The storage devices in the computing device 3000, such as the ROM 3030 or the hard disk 3070, can store various data or files used in processing and / or communication of the method provided by the present disclosure and program instructions executed by the CPU. The computing device 3000 can also include a user interface 3080. Of course, the architecture shown in FIG. 10 is only exemplary, and when implementing different devices, one or more components in the computing device shown in FIG. 10 can be omitted according to actual needs.
[0193] The present disclosure also provides a computer-readable storage medium. FIG. 11 shows a schematic diagram of a storage medium 4000 according to the present disclosure.
[0194] As shown in FIG. 11, the computer storage medium 4020 stores computer readable instructions 4010. When the computer readable instructions 4010 are run by a processor, the method according to the embodiments of the present disclosure described with reference to the above figures can be executed. The computer readable storage medium in the embodiments of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of illustrative but not limiting illustration, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DRAM). It should be noted that the memory of the method described herein is intended to include, but not be limited to, these and any other suitable types of memory. It should be noted that the memory of the method described herein is intended to include, but not be limited to, these and any other suitable types of memory.
[0195] This disclosure also provides a computer program product or computer program comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform a method according to an embodiment of this disclosure.
[0196] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0197] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0198] The exemplary embodiments of this disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art will understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of this disclosure, and such modifications should fall within the scope of this disclosure.
Claims
1. A method for generating music data based on text data, the method being performed by a computer device, the method comprising: determining, based on text data for describing music to be generated, description information for describing music symbols of the music to be generated, wherein the description information comprises at least one of: a music parameter set for describing music symbol attributes of the music to be generated, and a music symbol sequence for describing melody characteristics of the music to be generated; and determining, based on the description information, music data in a form of symbolic music sequence of the music to be generated.
2. The method of claim 1, wherein the determining, based on text data for describing music to be generated, description information for describing music symbols of the music to be generated comprises at least one of: determining the music parameter set based on text data for describing music information; or determining the music symbol sequence based on text data for describing melody. The music parameter set comprises one or more combinations of: instrument information, tonality information, tempo information, number of notes information, style information, duration information, emotion information, rhythm information, genre information, intensity information. The determining, based on text data for describing melody, the music symbol sequence comprises at least one of: in response to the text data for describing melody being a first music symbol sequence inputted, taking the first music symbol sequence as the music symbol sequence, or expanding the first music symbol sequence into a second music symbol sequence of a specific length as the music symbol sequence; or in response to the text data for describing melody being a natural language description of melody, generating the music symbol sequence based on the natural language description of melody. The determining, based on text data for describing melody, the music symbol sequence comprises: determining, based on the text data for describing melody, prompt input data, the prompt input data comprising at least one of: initial music symbols, format constraints of output data, expected distribution of music symbol sequence in different tonal regions, variation pattern of music symbol sequence in time value, variation pattern of music symbol sequence in dynamics; and determining, based on the prompt input data, the music symbol sequence using a large language model. When the description information comprises the music parameter set, the determining, based on the description information, the music data in a form of symbolic music sequence of the music to be generated comprises: determining word pieces corresponding to each music parameter in the music parameter set; generating a word piece sequence based on the word pieces corresponding to each music parameter in the music parameter set; and determining the music data in a form of symbolic music sequence of the music to be generated based on the word piece sequence.
3. The method of claim 1, wherein, When the description information comprises the music parameter set and the music symbol sequence, the determining, based on the description information, the music data in a form of symbolic music sequence of the music to be generated comprises:
4. The method of claim 2, wherein, 5. The method of claim 2, wherein, 6. The method of any of claims 1-5, wherein, 7. The method of any of claims 1-5, wherein, determine, based on both the set of music parameters and the sequence of music symbols, a probability distribution corresponding to each of a predetermined number of sequentially-arranged triples, wherein the probability distribution corresponding to each triple depends on the probability distribution of a preceding triple; and determine, based on the probability distribution corresponding to each of the plurality of triples, music data in the form of a symbolic music sequence of the music to be generated through a sampling operation; wherein each triple comprises a pitch, a duration, and a dynamics corresponding to a single music symbol.
8. The method of claim 7, wherein, For each of the plurality of triples, the determining of the probability distribution corresponding to each of a predetermined number of sequentially-arranged triples comprises: determining the probability distribution of the triple based on the probability distribution of each preceding triple, the set of music parameters, and the sequence of music symbols.
9. The method of any of claims 1-8, wherein, The determining of the music data in the form of a symbolic music sequence based on the description information comprises: determining the music data in the form of a symbolic music sequence of the music to be generated based on the description information using a plurality of melody guide transducer layers in cascade; wherein each melody guide transducer layer adjusts an output of a preceding melody guide transducer layer using the information of the music symbol used to describe the music data to be generated as attention information.
10. The method of claim 9, wherein, The melody guide transducer layer comprises a first perceptron, a second perceptron, and a fully connected layer; The first perceptron uses the sequence of music symbols as a query vector and an output of a preceding melody guide transducer layer as a key vector and a value vector; The second perceptron uses a weighted sum of an output of the first perceptron and the output of the preceding melody guide transducer layer as a query vector, a key vector, and a value vector; The fully connected layer uses an output of the second perceptron as an input, and an output of the fully connected layer is the output of the melody guide transducer layer.
11. The method of any one of claims 1-10, wherein The determining of the description information of the music symbol used to describe the music to be generated based on the text data used to describe the music to be generated comprises: determining, based on the text data used to describe the music to be generated, the description information of the music symbol used to describe the music to be generated using a music attribute information extraction model; and The determining of the music data in the form of a symbolic music sequence of the music to be generated based on the description information comprises: determining, based on the description information, the music data in the form of a symbolic music sequence of the music to be generated using a music data transducer.
12. The method of claim 11, wherein, The training of the music data transducer comprises: extracting a music data sample from original training data; extracting, from a description of the music data sample, a value of a preset music parameter corresponding to the music data sample, and using the value of the preset music parameter as a true value of a sample music parameter set used to describe music symbol attributes of the corresponding music data sample; generating a first training sample set based on the music data sample and the true value of the corresponding sample music parameter set; determine music symbols belonging to a central melody in the music data sample based on melody tolerance and chord tolerance, and take the central melody as a ground truth of a sample music symbol sequence used to describe a melody feature of the corresponding music data sample; generate a second training sample set based on the music data sample and the ground truth of the corresponding sample music symbol sequence; train the music data converter using the first training sample set and the second training sample set.
13. The method of claim 12, wherein, The training of the music attribute information extraction model includes: determine text training data used to describe music information based on the ground truth of the sample music parameter set; generate a third training sample set based on the text training data used to describe music information and the ground truth of the corresponding sample music parameter set; determine text training data used to describe melody based on the ground truth of the sample music symbol sequence; generate a fourth training sample set based on the text training data used to describe melody and the ground truth of the corresponding sample music symbol sequence; fine-tune the music attribute information extraction model using the third training sample set and the fourth training sample set.
14. An apparatus for generating music data based on text data, the apparatus comprising: a music attribute information extraction model configured to determine description information of music symbols used to describe music to be generated based on text data used to describe the music to be generated, wherein the description information includes at least one of a music parameter set used to describe music symbol attributes of the music to be generated and a music symbol sequence used to describe melody features of the music to be generated; and a music data converter configured to determine music data in the form of symbolic music sequence of the music to be generated based on the description information.
15. An electronic device comprising: one or more processors; and one or more memories, wherein the memories store computer executable programs that, when executed by the processors, perform the method of any one of claims 1-13.
16. A computer readable storage medium having stored thereon a computer program, the computer program, when executed by a processor, implementing the method of any one of claims 1-13.
17. A computer program product, the computer program product comprising a computer program stored in a computer readable storage medium, the computer program being readable by a processor of a computer device from the computer readable medium and configured to cause the computer device to carry out the method of any one of claims 1-13.