Music Melody Arrangement Method, Device and System Based on Generative Adversarial Network

Through the improved generative adversarial network model, the generative adversarial network model uses two discriminators and relational memory units to solve the problem of poor generative melody effect in the prior art, and achieves more diverse and authentic musical melody generation, suitable for music creation of the general population.

CN115762452BActive Publication Date: 2025-07-25SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211403621.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-10
Publication Date
2025-07-25
Estimated Expiration
2042-11-10

AI Technical Summary

Technical Problem

In the prior art, algorithms based on lyrics to generate melody are mainly aimed at data practitioners, with poor generation results and cannot meet the music creation needs of the general population.

Method used

Using an improved generative adversarial network model, a machine learning model using two discriminators and relational memory units is used to generate music files based on the lyrics and instrument information entered by the user, and supports the playback and management of music files.

Benefits of technology

It improves the diversity and authenticity of music melody, expands the applicable population to the general population, and provides the functions of music file generation and playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115762452B_ABST
    Figure CN115762452B_ABST
Patent Text Reader

Abstract

The present invention discloses a music melody arrangement method, device and system based on a generative adversarial network. The method includes: receiving a form sent by a client, and synthesizing a music file according to the information in the form by using a machine learning model; wherein the machine learning model adopts an improved generative adversarial network model; the improved generative adversarial network model improves one discriminator into two discriminators, and at the same time improves the long short-term memory structure in the generator into a relational memory unit; returning the path of the music file to the client, so that the client uses the ref attribute to dynamically bind the path of the music file in the audio control on the page and automatically plays the music file. By adopting the generative adversarial network model, the present invention greatly increases the diversity and authenticity of music melody generation, and the generated music melody is more beautiful and pleasant; by synthesizing music files, the applicable population is extended to the general population.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence machine learning, and particularly relates to a music melody arrangement method, device and system based on a generative adversarial network. Background Art

[0002] Currently, most in this field use only an algorithm that generates melody data based on lyrics rather than the final audio. The disadvantage of such software is that it is only for data practitioners rather than those with real music arrangement business needs, and the music generation effect of such algorithms is also not satisfactory. Therefore, there is a need to provide a software system that can not only generate more suitable music melody data according to lyrics, but also convert it into a music file, and at the same time provide functions such as music playback, information management, and user permission management. Summary of the Invention

[0003] To solve the above-mentioned deficiencies of the prior art, the present invention provides a music melody arrangement method, device, system, server and storage medium based on a generative adversarial network. According to the lyrics input by the user and the selected musical instruments, a machine learning model is used to synthesize a music file. The machine learning model uses a generative adversarial network model and changes one discriminator to two discriminators, and at the same time replaces the long short-term memory structure in the generator with a relational memory unit, thereby greatly increasing the diversity and authenticity of music melody generation, and the generated music melody is more beautiful and pleasant. In addition, by synthesizing the music file, the applicable population is extended from data practitioners to the general population.

[0004] The first object of the present invention is to provide a music melody arrangement method based on a generative adversarial network.

[0005] The second object of the present invention is to provide a music melody arrangement device based on a generative adversarial network.

[0006] The third object of the present invention is to provide a music melody arrangement system based on a generative adversarial network.

[0007] The fourth object of the present invention is to provide a server.

[0008] The fifth object of the present invention is to provide a storage medium.

[0009] The first object of the present invention can be achieved by adopting the following technical solutions:

[0010] A music melody arrangement method based on a generative adversarial network, the method comprising:

[0011] Receive the form sent by the client, and synthesize a music file according to the information in the form by using a machine learning model; wherein the machine learning model adopts an improved generative adversarial network model; the improved generative adversarial network model improves one discriminator into two discriminators, and at the same time improves the long short-term memory structure in the generator into a relational memory unit;

[0012] Return the path of the music file to the client, so that the client can dynamically bind the path of the music file in the audio control on the page using the ref attribute and automatically play the music file.

[0013] Further, the information in the form includes lyrics and instrument information;

[0014] Synthesizing a music file according to the information in the form by using a machine learning model includes:

[0015] Use the dictionary module to convert the lyrics into vectors;

[0016] Generate a music melody according to the vectors by using a machine learning model;

[0017] Generate a music file according to the music melody and instrument information.

[0018] Further, generating a music melody according to the vectors by using a machine learning model includes:

[0019] Obtain a large number of existing songs on the network as a dataset, and use the songs in the dataset to train the machine learning model; during the training phase, adjust the parameters in the generator according to the output results of the two discriminators;

[0020] Input the vectors into the trained machine learning model, and output a music melody through the generator.

[0021] Further, the two discriminators are respectively a lyrics discriminator and a melody discriminator. Among them, the lyrics discriminator includes a linear layer, a bidirectional LSTM layer, and a sigmoid function, and the melody discriminator includes a linear layer, two cascaded LSTM layers, a linear layer, a Dropout module, and a sigmoid function;

[0022] Training the machine learning model by using the songs in the dataset includes:

[0023] Use the dictionary module to convert the lyrics in the song into vectors, splice the vectors with noise and input them into the generator, and the generator outputs a music melody; wherein, the noise is a random number vector with the same dimension as the vectors;

[0024] Input the music melody, vector, and the melody in the song into the lyric discriminator, and successively pass through a linear layer, a bidirectional LSTM layer, and a sigmoid function to obtain the output of the lyric discriminator;

[0025] Input the music melody and the melody in the song into the melody discriminator, and successively pass through a linear layer, two cascaded LSTM layers, a linear layer and a Dropout module, and a sigmoid function to obtain the output of the melody discriminator;

[0026] Generate a positive feedback effect on the generator through the outputs of the two discriminators, enabling the generator to generate more realistic melody data.

[0027] Further, generating a music file according to the music melody and instrument information includes:

[0028] Save the music melody in tensor format as an npy file;

[0029] Read the npy file as an ndarray file;

[0030] Convert the ndarray file and instrument information into a mid file;

[0031] Convert the mid file into an mp3 file or a wav file, and thus obtain the music file.

[0032] The second object of the present invention can be achieved by adopting the following technical solutions:

[0033] A music melody arrangement device based on a generative adversarial network, the device includes:

[0034] A music file synthesis unit, configured to receive a form sent by a client, and synthesize a music file according to the information in the form by using a machine learning model; wherein the machine learning model adopts an improved generative adversarial network model; the improved generative adversarial network model improves one discriminator into two discriminators, and at the same time improves the long short-term memory structure in the generator into a relational memory unit;

[0035] A music file playback unit, configured to return the path of the music file to the client, so that the client uses the ref attribute to dynamically bind the path of the music file in the audio control on the page, and automatically plays the music file.

[0036] The third object of the present invention can be achieved by adopting the following technical solutions:

[0037] A music melody arrangement system based on a generative adversarial network, including a client and a server, wherein:

[0038] The client is built with the Vue3 framework and is used for page display and user interaction with the system, including a music file generation module. The music file generation module is used to send form information to the server for processing via ajax based on the lyrics entered by the user on the form page and the selected musical instruments. After receiving the path of the music file returned by the server, it uses the ref attribute to dynamically bind the path of the music file in the audio control on the page and automatically plays the music file.

[0039] The main part of the server is built with the Django framework and includes a music synthesis module. The music synthesis module is used to synthesize a music file based on the form received from the client using a machine learning model and return the path of the music file to the client.

[0040] Furthermore, the client also includes a login module, a viewing module, and a commenting module. Among them, the login module is used to verify the user's identity information and provide an entry to log in to the system. The viewing module is used for logged-in users to view corresponding information according to their personal identity permissions. The commenting module is used to provide a communication channel between users and administrators. mainly by the user directly entering comments in the input box below the interface and sending an axios request to the server after clicking to submit the comment. The server parses it through django and sends the parsed information to the database for storage.

[0041] Furthermore, the system also provides a superuser for managing and viewing the historical records of all music generations, and the historical records are stored in the database.

[0042] The fourth object of the present invention can be achieved by adopting the following technical solution:

[0043] A server includes a processor and a memory for storing programs executable by the processor. When the processor executes the programs stored in the memory, it implements the above-mentioned music melody arrangement method.

[0044] The fifth object of the present invention can be achieved by adopting the following technical solution:

[0045] The fourth object of the present invention can be achieved by adopting the following technical solution:

[0046] A storage medium stores a program. When the program is executed by a processor, it implements the above-mentioned music melody arrangement method.

[0047] The present invention has the following beneficial effects compared with the prior art:

[0048] The method provided by the present invention synthesizes a music file according to the lyrics input by the user and the selected musical instrument. The machine learning model uses a generative adversarial network model and changes one discriminator into two discriminators. At the same time, the long short-term memory structure in the generator is replaced by a relational memory unit, thereby greatly increasing the diversity and authenticity of music melody generation. At the same time, the generated music melody is more beautiful and pleasant. In addition, by synthesizing the music file, the applicable population is extended from data practitioners to the general population. Description of the Drawings

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.

[0050] Figure 1 It is a flowchart of the music melody arrangement method based on the generative adversarial network in Embodiment 1 of the present invention.

[0051] Figure 2 It is a schematic flowchart of training the machine learning model in Embodiment 1 of the present invention.

[0052] Figure 3 It is a schematic diagram of the structure of the lyrics discriminator in Embodiment 1 of the present invention.

[0053] Figure 4 It is a schematic diagram of the structure of the melody discriminator in Embodiment 1 of the present invention.

[0054] Figure 5 It is to convert the music melody into a music file in Embodiment 2 of the present invention.

[0055] Figure 6 It is an architecture diagram of the music melody arrangement system based on the generative adversarial network in Embodiment 1 of the present invention.

[0056] Figure 7 It is a page for the system background to manage all music generation records in Embodiment 1 of the present invention.

[0057] Figure 8 It is a page of the comment module in the system in Embodiment 1 of the present invention.

[0058] Figure 9 It is a deployment diagram of the music melody arrangement system based on the generative adversarial network in Embodiment 1 of the present invention.

[0059] Figure 10It is a structural block diagram of a music melody arrangement device based on a generative adversarial network according to Embodiment 2 of the present invention.

[0060] Figure 11 It is a structural block diagram of a server according to Embodiment 3 of the present invention. Detailed implementation manners

[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. It should be understood that the described specific embodiments are only used to explain the present application and are not used to limit the present application.

[0062] Embodiment 1:

[0063] As Figure 1 shown, the music melody arrangement method based on a generative adversarial network provided in this embodiment includes the following steps:

[0064] (1) The client verifies the valid identity of the logged-in user.

[0065] If the user ID and the entered password match correctly, it enters the main page; otherwise, it prompts that the password is incorrect and requires re-entering the password. If the number of times of entering the password reaches 3 times and the password entered for the third time is still incorrect, it prompts the number of times entered and exits the login page.

[0066] (2) After the user logs in, enters the lyrics and selects an instrument on the page, and clicks the "Synthesis Button", the client collects the entered lyrics and the instrument used for the music melody selected from several alternative instruments through a form, and sends the form information to the server side by means of ajax.

[0067] After the user logs in, enters the lyrics and selects an instrument on the opened page, and clicks the "Synthesis Button", the UI control provided by element-plus is used to obtain the fields entered by the user as the lyrics input, and other configuration information of the music is transmitted to the backend server using the GET request in the request.

[0068] (3) After receiving the form, the server synthesizes a music file according to the information in the form using a machine learning model.

[0069] After the server receives the form, according to the information in the received form, mainly including lyrics information and music information, that is, the number of the instrument in the system. After parsing the lyrics information and converting it into a vector, a machine learning model is used to synthesize a music file, specifically including:

[0070] (3-1) Use the dictionary module to convert the lyrics in the form into a vector.

[0071] The dictionary module refers to dividing the common lyrics words into syllable forms and then using the machine learning word2vec method, that is, the method of converting words into vectors, to convert these common words into vectors, because only words in vector representation can be used for machine learning training.

[0072] Specifically, divide the sentence input by the user into syllable forms and then input them into the dictionary model in turn to convert them into vector representation forms.

[0073] (3-2) Generate a music melody by inputting the vector into a machine learning model.

[0074] (3-2-1) Machine learning model.

[0075] The machine learning model used in this embodiment is improved compared with the traditional machine learning model for generating melodies. The improvements are as follows: First, the machine learning model uses a generative adversarial network model; Second, one discriminator in the generative adversarial network model is changed to two discriminators; Third, the structure of the generator is improved to a relational memory cell model.

[0076] As Figure 2 shown, the generative adversarial network model includes a generator and two discriminators, and the two discriminators are a lyrics discriminator and a melody discriminator respectively. The advantage of using the generative adversarial network model is that it greatly increases the diversity and authenticity of music melody generation.

[0077] Generally, the generator in the generative adversarial network model uses a long short-term memory (LSTM) structure to generate data. Among them, LSTM is a machine learning network structure that includes a "forget gate", a "memory gate" and an "output gate", and it is used to process sequences. In this embodiment, the LSTM structure in the generator is replaced with a relational memory cell (RMC). The RMC unit also processes sequences. Compared with LSTM, it can better retain the relationship between words in a sequence. The machine learning model can better capture the relationship between each word in the lyrics input by the user through the generator, improve the quality of the generated music melody, make it more pleasant to listen to, and closer to real music.

[0078] As Figure 3As shown, the lyric discriminator mainly includes a linear layer, a bidirectional LSTM layer, and a sigmoid activation function layer.

[0079] As Figure 4 shown, the melody discriminator includes a linear layer, two cascaded LSTM layers, a linear layer and a Dropout module, and a sigmoid function. The input of the melody discriminator is only music triples, which may come from real data or may be fake data generated from the generator. By using the melody discriminator, the generated music melody effect is significantly improved.

[0080] (3-2-2) Train the machine learning model.

[0081] Obtain a large number of existing songs on the network as a dataset, such as Figure 2 shown, use the songs in the dataset to train the machine learning model.

[0082] Use the dictionary module to convert the lyrics in the song into vectors. After splicing the vectors with noise, input the spliced vectors into the linear layer and the relational memory unit layer in the generator. The generator outputs music triples, where the noise is a random number vector of the same dimension as the above vectors.

[0083] Input the music triples, vectors, and the melody in the song into the lyric discriminator, and pass through the linear layer, bidirectional LSTM layer, and sigmoid layer in sequence to obtain the output of the lyric discriminator. The output of the lyric discriminator mainly includes two parts. One part is the vector generated by the lyrics through the dictionary module. The other part is the melody, that is, the vector in the form of music triples. The music triples mentioned here may come from real data or may be fake data generated from the generator. If the output is true, it means that the lyric discriminator has been successfully "deceived" by the generator during the generation process of this music melody, and the generator achieves the expected effect of simulating artificial synthesized music melody. If the output is false, it means that the lyric discriminator recognizes that this music melody is generated by the generator rather than real data.

[0084] Input the music triples and the melody in the song into the melody discriminator, and pass through the linear layer, two cascaded LSTM layers, a linear layer and a Dropout module, and a sigmoid function in sequence to obtain the output of the melody discriminator.

[0085] In the training stage of the machine learning model, adjust the parameters in the generator according to the output results of the two discriminators. Whether the judgment results of the two discriminators are true or false, they will have a positive feedback effect on the generator, enabling the generator to generate more realistic melody data.

[0086] During the training process, if the output results of the two discriminators remain unchanged after several consecutive training sessions, the training ends, that is, the training of the machine learning model ends, and a trained machine learning model is obtained.

[0087] (3-2-3) Input the vector into the trained machine learning model to generate a music melody.

[0088] Concatenate the vector with noise and input it into the generator in the trained machine learning model. The generator outputs a music triple, that is, a music melody, and the music melody is melody data in tensor form.

[0089] Specifically, the generator generates a music melody based on the vector converted from the lyrics input by the user.

[0090] (3-3) According to the music melody and the instrument information in the form, use a tool function to generate a music file.

[0091] Such as Figure 5 As shown, the music melody generated by the generator is melody data in tensor form. It can be persistently saved only by saving it as an npy format file. If you want to generate a music file, you need to use the numpy library in Python to read the npy file into numpy's ndarray format data, and then use the instrument number parsed from the request sent by the front end and the just-mentioned ndarray format data to convert it into a mid file through pretty_midi. Among them, pretty_midi is an external Python library used to convert ndarray format data into midi format music files. Since the browser does not support playing midi format music files natively, to play a music file in the browser, it needs to be converted from midi format to mp3 format through midi2audio to be playable. Among them, midi2audio is an external Python library used to convert midi format files into mp3 format music files. Converting the music melody into a common mp3 format file can be directly played in the browser.

[0092] (4) If the music file is successfully synthesized, display a message indicating successful synthesis. Otherwise, display a message indicating failed synthesis and return to step (2) to re-enter the lyrics, and execute the subsequent steps.

[0093] After the music file is successfully synthesized, the server returns the path of the music file to the client. The client uses the ref attribute to dynamically bind the path of the music file in the audio control on the page and automatically plays the music file. At this time, the user can hear the music melody synthesized by the browser, and the music arrangement process ends.

[0094] Such as Figure 6As shown in the figure, this embodiment also provides a music melody arrangement system based on a generative adversarial network. Its architecture includes a client, a gateway, a server, storage, and continuous integration. Among them, the client contains a series of pages for display built by the front-end framework Vue3; the gateway includes a server with nginx load balancing; the main part of the server is built by the Django framework, and the Django framework includes a generator, a communication interface, and a data persistence layer. Among them, the generator can generate music melodies and is a machine learning module built by the pytorch framework; the MySQL database is used for data storage, and docker is used for deployment in continuous integration.

[0095] The music melody arrangement system based on the generative adversarial network provided by this embodiment mainly includes two parts: the front end (user side) and the back end (server side) of the system. The front end mainly bears page display and user interaction with the system, including four modules: a login module, a music file generation module, a viewing module, and a comment module. Among them, the login module is used to verify the user's identity information and provide an entry to log in to the system; the music file generation is used to obtain the generated music file after clicking the "synthesis button" according to the lyrics input by the user and the selected instrument. At the same time, a melody waveform diagram is displayed on the page; the viewing module is used for logged-in users to view their personal information; the comment module is used for logged-in users to comment on this system. Users can leave comments on this system and also view the comments left by other users on the system. The back end mainly includes two modules: a music synthesis module, which is used to receive requests from the front end and synthesize music files using a machine learning model; a storage module, which is used for data storage and stores user information and user comments and other data in the database.

[0096] Specifically, both the comment module and the viewing module in the system use the MVC architecture. The front-end form page obtains the user's input data and also passes it to the Django framework on the back end for data processing in the form of Ajax through axios, and then the data is rendered to the page through the reactive update of the vue framework. The internationalization adaptation of the front-end page is achieved by using vue-i18n to replace strings in different languages and render them on the page by configuring different json format data.

[0097] Specifically, the music synthesis module in the system collects the music lyrics information input by the user and the instrument information selected by the user through a form, and sends the form information to the backend interface through the ajax method. The specific process is to use the UI controls provided by element-plus to obtain the fields input by the user as the lyrics input, and use the GET request in the request to transmit other configuration information of the music to the backend. The backend first parses the lyrics information from the frontend, converts the lyrics into vectors through the loaded dictionary module, then performs machine learning derivation on these vectors to synthesize a data in the npy format, and then uses a tool function to convert it into a music file in the mp3 format or wav format supported by the browser for storage, and returns the path of this music file on the server as the content of the response. After the frontend obtains the Promise data wrapped with the above response, it asynchronously parses to obtain the path of the music file, dynamically binds the path of the music file in the audio control on the page, and finally plays the music file. The backend deployment uses a deep learning algorithm based on a generative adversarial network. This machine learning algorithm separately takes out the generator for machine learning derivation after completion of training.

[0098] As Figure 7 shown, the system provided in this embodiment also provides a background for super users to manage all music generation records. Specifically, it uses MySQL to store all music generation histories. After the administrator logs in, they can query the data in MySQL and send it to the foreground for display. The function of adding this function is to provide a more intuitive way to manage the data of music synthesis. After logging in with the super user, the history of all music melody syntheses in the system can be seen.

[0099] As Figure 8 shown, the comment module in the system provided in this embodiment mainly allows users to directly enter comments in the input box below the interface, and after clicking to submit the comment, the frontend sends an axios request to the backend. The django on the backend parses and stores the information in MySQL. At the same time, all information can also be queried in MySQL for display. Adding the evaluation function provides a communication channel between website users and website administrators.

[0100] As Figure 9 shown, the system provided in this embodiment uses the web distribution and proxy functions of nginx, achieves data exchange by using a socket to connect to the backend information, and the data persistence server uses MySQL.

[0101] Since the above system has a multi - language adaptable interface for users of different languages. The above system provides functions such as information management and user permission settings, enabling users to view and operate relevant information. At the same time, the super - user logged in as the system administrator can view all music synthesis records in the system background. Compared with other music synthesis systems, the above system provides user comments, and users can view their own music synthesis history and select different musical instruments, greatly enriching the functionality for users.

[0102] Those skilled in the art can understand that all or part of the steps in the method of implementing the above - mentioned embodiments can be completed by instructing relevant hardware through a program, and the corresponding program can be stored in a computer - readable storage medium.

[0103] It should be noted that although the method operations of the above - mentioned embodiments are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the shown operations must be performed to achieve the desired result. On the contrary, the described steps can be changed in the execution order. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.

[0104] Embodiment 2:

[0105] As Figure 10 shown, this embodiment provides a music melody arrangement device based on a generative adversarial network. The device includes a music file synthesis unit 1001 and a music file playback unit 1002, where:

[0106] The music file synthesis unit 1001 is configured to receive a form sent by a client, and synthesize a music file according to the information in the form by using a machine learning model; where the machine learning model uses an improved generative adversarial network model; the improved generative adversarial network model improves one discriminator into two discriminators, and at the same time improves the long - short - term memory structure in the generator into a relational memory unit;

[0107] The music file playback unit 1002 is configured to return the path of the music file to the client, so that the client uses the ref attribute to dynamically bind the path of the music file in the audio control on the page and automatically play the music file.

[0108] The specific implementation of each unit in this embodiment can be referred to the above-mentioned Embodiment 1, which will not be elaborated here one by one. It should be noted that the device provided in this embodiment is only exemplified by the above division of each functional unit. In practical applications, the above functions can be allocated to different functional units as needed, that is, the internal structure is divided into different functional units to complete all or part of the functions described above.

[0109] Embodiment 3:

[0110] This embodiment provides a server, which can be a computer, such as Figure 11 shown, which is connected to a processor 1102, a memory, an input device 1103, a display 1104, and a network interface 1105 through a system bus 1101. The processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium 1106 and an internal memory 1107. The non-volatile storage medium 1106 stores an operating system, a computer program, and a database. The internal memory 1107 provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the processor 1102 executes the computer program stored in the memory, the music melody arrangement method of the above-mentioned Embodiment 1 is implemented as follows:

[0111] Receive a form sent by a client, and synthesize a music file according to the information in the form by using a machine learning model; wherein the machine learning model uses an improved generative adversarial network model; the improved generative adversarial network model improves one discriminator into two discriminators, and at the same time improves the long short-term memory structure in the generator into a relational memory unit;

[0112] Return the path of the music file to the client, so that the client can use the ref attribute to dynamically bind the path of the music file in the audio control on the page and automatically play the music file.

[0113] Embodiment 4:

[0114] This embodiment provides a storage medium, which is a computer-readable storage medium and stores a computer program. When the computer program is executed by a processor, the music melody arrangement method of the above-mentioned Embodiment 1 is implemented as follows:

[0115] Receive a form sent by a client, and synthesize a music file according to the information in the form by using a machine learning model; wherein the machine learning model uses an improved generative adversarial network model; the improved generative adversarial network model improves one discriminator into two discriminators, and at the same time improves the long short-term memory structure in the generator into a relational memory unit;

[0116] Return the path of the music file to the client, so that the client can use the ref attribute to dynamically bind the path of the music file in the audio control on the page and automatically play the music file.

[0117] It should be noted that the computer-readable storage medium in this embodiment can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0118] As mentioned above, only the preferred embodiments of the present invention for patents are described, but the protection scope of the present invention for patents is not limited thereto. Any person skilled in the art within the scope disclosed by the present invention for patents, according to the technical solution and inventive concept of the present invention for patents, makes equivalent substitutions or changes, and all belong to the protection scope of the present invention for patents.

Claims

1. A music melody arrangement method based on a generative adversarial network, characterized in that, The method includes: Receiving a form sent by a client, and synthesizing a music file according to the information in the form by using a machine learning model; wherein, the information in the form includes lyrics and instrument information; the machine learning model adopts an improved generative adversarial network model; the improved generative adversarial network model improves one discriminator into two discriminators, and at the same time improves the long short-term memory structure in the generator into a relational memory unit; the two discriminators are a lyrics discriminator and a melody discriminator respectively, the lyrics discriminator includes a linear layer, a bidirectional LSTM layer and a sigmoid function, and the melody discriminator includes a linear layer, two cascaded LSTM layers, a linear layer, a Dropout module and a sigmoid function; Returning the path of the music file to the client, so that the client can dynamically bind the path of the music file in the audio control on the page by using the ref attribute and automatically play the music file.

2. The music melody arrangement method according to claim 1, characterized in that The synthesizing the music file according to the information in the form by using a machine learning model includes: Using a dictionary module to convert the lyrics into vectors; Generating a music melody according to the vectors by using a machine learning model; Generating a music file according to the music melody and instrument information.

3. The music melody arrangement method according to claim 2, characterized in that, The generating the music melody according to the vectors by using a machine learning model includes: Obtaining a large number of existing songs on the network as a data set, and training the machine learning model by using the songs in the data set; during the training phase, adjusting the parameters in the generator according to the output results of the two discriminators; Inputting the vectors into the trained machine learning model, and outputting a music melody through the generator.

4. The music melody arrangement method according to claim 3, characterized in that, The training the machine learning model by using the songs in the data set includes: Using a dictionary module to convert the lyrics in the song into vectors, splicing the vectors with noise and inputting them into the generator, and the generator outputs a music melody; wherein, the noise is a random number vector with the same dimension as the vectors; Inputting the music melody, the vectors and the melody in the song into the lyrics discriminator, and passing through a linear layer, a bidirectional LSTM layer and a sigmoid function in sequence to obtain the output of the lyrics discriminator; Inputting the music melody and the melody in the song into the melody discriminator, and passing through a linear layer, two cascaded LSTM layers, a linear layer, a Dropout module and a sigmoid function in sequence to obtain the output of the melody discriminator; Generating a positive feedback effect on the generator through the outputs of the two discriminators, so that the generator can generate more realistic melody data.

5. The music melody arranging method according to claim 2, wherein The generating the music file according to the music melody and instrument information includes: Saving the music melody in tensor format as an npy file; Reading the npy file as an ndarray file; Converting the ndarray file and the instrument information into a mid file; Converting the mid file into an mp3 file or a wav file, and thus obtaining the music file.

6. A music melody arrangement device based on a generative adversarial network, characterized in that, The device includes: A music file synthesis unit, which is used to receive the form sent by the client, and synthesize a music file according to the information in the form by using a machine learning model; wherein, the information in the form includes lyrics and instrument information; the machine learning model adopts an improved generative adversarial network model; the improved generative adversarial network model improves one discriminator into two discriminators, and at the same time improves the long short-term memory structure in the generator into a relational memory unit; the two discriminators are respectively a lyrics discriminator and a melody discriminator, the lyrics discriminator includes a linear layer, a bidirectional LSTM layer and a sigmoid function, and the melody discriminator includes a linear layer, two cascaded LSTM layers, a linear layer, a Dropout module and a sigmoid function; A music file playback unit, which is used to return the path of the music file to the client, so that the client can dynamically bind the path of the music file in the audio control in the page by using the ref attribute and automatically play the music file.

7. A music melody arrangement system based on a generative adversarial network, characterized in that, It includes a client and a server, where: The client is built by the Vue3 framework and is used for page display and user interaction with the system, including a music file generation module; the music file generation module is used to send the form information to the server for processing by ajax according to the lyrics input by the user on the form page and the selected instrument. After receiving the path of the music file returned by the server, use the ref attribute to dynamically bind the path of the music file in the audio control in the page and automatically play the music file; The main part of the server is built by the Django framework and includes a music synthesis module; the music synthesis module is used to synthesize a music file according to the received form sent by the client by using a machine learning model and return the path of the music file to the client; the machine learning model adopts an improved generative adversarial network model; the improved generative adversarial network model improves one discriminator into two discriminators, and at the same time improves the long short-term memory structure in the generator into a relational memory unit; the two discriminators are respectively a lyrics discriminator and a melody discriminator, the lyrics discriminator includes a linear layer, a bidirectional LSTM layer and a sigmoid function, and the melody discriminator includes a linear layer, two cascaded LSTM layers, a linear layer, a Dropout module and a sigmoid function.

8. The music melody arrangement system according to claim 7, characterized in that, The client also includes a login module, a viewing module and a comment module; wherein, the login module is used to verify the user's identity information and provide an entrance to log in to the system; the viewing module is used for the logged-in user to view corresponding information according to the permissions of the personal identity; the comment module is used to provide a channel for communication between the user and the administrator. mainly by the user directly entering a comment in the input box below the interface, clicking the submit comment and then sending an axios request to the server, and the server parses it through django and sends the parsed information to the database for storage.

9. The music melody arrangement system according to claim 7, wherein The system also provides a superuser for managing and viewing the historical records of all music generated, and the historical records are stored in a database.

10. A server, comprising a processor and a memory for storing processor-executable programs, characterized in that, When the processor executes the program stored in the memory, it implements the music melody arrangement method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Automatic song creation method via computer

    CN106652984A

  • Melody generation method based on generative adversarial network

    CN109584846A