A method, system, device and storage medium for automatic composition of guzheng music
Through GRU-based and reinforcement learning methods, the problem of tone simulation and skill expression in the generation of Guzheng music in the prior art is solved. The generated music is closer to the real Guzheng folk music, and the tone and skill expression is richer and more accurate.
Patent Information
- Application Number
- CN202210162198.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-22
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-02-22
AI Technical Summary
The prior art is difficult to effectively generate high-quality guzheng music, especially in terms of tone simulation and technical expression.
Using GRU and reinforcement learning methods, by obtaining the simplified score data of the Guzheng into MIDI format, the GRU network is trained to generate preliminary melody data, and the Guzheng performance skills are added through reinforcement learning training, and finally the tone conversion is performed to generate the music data of the Guzheng tone.
The generated music is closer to the real Guzheng folk music, and the tone and technique expression is richer and more accurate, solving the problems of tone simulation and technique expression in the existing technology.
Smart Images

Figure CN114664275B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automatic music composition, and particularly to a method, system, device and storage medium for automatic Guzheng music composition. Background Art
[0002] Computer music generation refers to the computer as an aid to generate a specific sequence of notes, which includes musical elements such as pitch, beat, melody, etc., so as to assist composers in music creation and minimize the degree of intervention of composers in the music creation process. However, the task of using a computer to generate Guzheng music is very challenging. First, the timbres of Guzheng and similar Chinese folk instruments cannot be found in the standard MIDI timbre library, so a post-processing link is needed to simulate the timbres of Guzheng or related folk instruments. Second, there are many unique fingering techniques for Guzheng, and it is difficult to express them in MIDI format. Therefore, the generation of technical music is the key to algorithm generation. Third, the generation of music is a discrete sequence, and the pleasant melody and generation speed are problems that the algorithm needs to train and solve. Summary of the Invention
[0003] To solve at least one of the technical problems existing in the prior art to a certain extent, the purpose of the present invention is to provide a method, system, device and storage medium for automatic Guzheng music composition based on GRU and reinforcement learning.
[0004] The technical solution adopted by the present invention is as follows:
[0005] An automatic Guzheng music composition method includes the following steps:
[0006] Obtain the numbered musical notation data of Guzheng, convert the numbered musical notation data of Guzheng into staff notation data, obtain MIDI format data according to the staff notation data, and construct a Guzheng dataset;
[0007] Obtain the chorus part of each song in the Guzheng dataset, and use the chorus part and the full song as a training set;
[0008] Train the GRU network with the training set, and obtain preliminary melody data after training;
[0009] Perform reinforcement learning training on the preliminary melody data to obtain music data with Guzheng performance skills;
[0010] Perform timbre conversion on the music data with Guzheng performance skills to generate music data with Guzheng timbre.
[0011] Further, the conversion of the numbered musical notation data of Guzheng into staff notation data and the obtaining of MIDI format data according to the staff notation data include:
[0012] Convert each note in the numbered musical notation data according to the staff standard to obtain staff data;
[0013] Use Musescore software to convert the staff data into MIDI format data; among them, in the MIDI format, the piano timbre is used instead of the guzheng timbre.
[0014] Further, the obtaining of the chorus part of each song in the guzheng dataset includes:
[0015] Segment the MIDI audio of the song, calculate the similarity between any two segments of audio using a similarity function, and obtain the chorus part of the song according to the similarity.
[0016] Further, the expression of the similarity function is:
[0017]
[0018] In the formula, X i and X i+1 are the note vectors of any two pieces of music, and the whole song consists of a set of frames of 12 basic notes.
[0019] Further, the process of training the GRU network includes:
[0020] Represent a piece of music data in MIDI format as a matrix: x = {x t,j , t = 1, 2, …, N, j = 1, 2, …, 128}; where N represents the input length of the music segment;
[0021] Use One-Hot encoding to encode the music data to obtain the input matrix: X = {X t,j , t = 1, 2, …, N, j = 1, 2, …, 128};
[0022] Input the input matrix into the GRU network to obtain the output matrix: Y = {Y t,j , t = 1, 2, …, M, j = 1, 2, …, 128}; where M represents the output length of the music segment (the input and output lengths are set according to the code);
[0023] Perform One-Hot decoding on the output matrix to obtain the melody matrix: y = {y t,j , t = 1, 2, …, M, j = 1, 2, …, 128};
[0024] Use the built-in MIDI library in python to convert the melody matrix y into MIDI format;
[0025] Convert the MIDI format into the notes of the music.
[0026] The loss function of the GRU network is the cross-entropy function, and the expression of the cross-entropy function is as follows:
[0027]
[0028] In the formula, P represents the number of training samples, and Y i represents the actual output of the i-th training sample, Y j represents the expected output of the i-th training sample.
[0029] Furthermore, the step of performing reinforcement learning training on the preliminary melody data to obtain music data with guzheng performance skills includes:
[0030] Setting the reward and punishment rules for skills and setting the parameters of the reinforcement learning network;
[0031] Inputting the preliminary melody data into the reinforcement learning network and processing it according to the reward and punishment rules to obtain music data with guzheng performance skills.
[0032] Furthermore, the expression of the loss function of the reinforcement learning network is as follows:
[0033]
[0034] In the formula, E D represents the expectation of the agent random variable, R t is the reward and punishment reward, γ is the attenuation factor, y t is the current action state, that is, the input note state; Q(s t ,y t ; θ t - ) represents the output of the target value state, Q(s t ,y t ; θ t ) represents the output of the current value state, θ t and θ t - are the iterative parameters of the network at time t;
[0035] The steps of processing in the reinforcement learning DQN network are as follows:
[0036] Initializing the experience pool D, initializing the DNN network parameters θ t of the DQN in the DQN, setting the value of the action value Q to a random number, and setting θ t - = θ t ;
[0037] Taking the preliminary melody data y output by the GRU network as the initial value of the reinforcement learning DQN;
[0038] Calculate the reward and punishment reward R t , and calculate the average return rate according to the reward and punishment reward R t Calculate the average return rate
[0039] Calculate the action value Q(s t , y t ; θ t ), and calculate the target Q value according to the action value Q
[0040] Solve the loss function of the reinforcement learning network to obtain (s t , R t ; θ t ), and put (s t , R t ; θ t ) into the experience pool D;
[0041] According to the average return rate and the loss L, determine whether to stop training; if not, update θ t - = θ t , t = t + 1, and recalculate the reward and punishment reward R t .
[0042] Another technical solution adopted by the present invention is:
[0043] An automatic guzheng composition system, comprising:
[0044] A data acquisition module, configured to acquire the numbered musical notation data of the guzheng, convert the numbered musical notation data of the guzheng into staff notation data, acquire MIDI format data according to the staff notation data, and construct a guzheng data set;
[0045] A data preprocessing module, configured to acquire the chorus part of each song in the guzheng data set, and use the chorus part and the full song as a training set;
[0046] A preliminary composition module, configured to train a GRU network using the training set, and obtain preliminary melody data after training;
[0047] A technique generation module, configured to perform reinforcement learning training on the preliminary melody data to obtain music data with guzheng performance techniques;
[0048] A timbre conversion module, configured to perform timbre conversion on the music data with guzheng performance techniques to generate music data with guzheng timbre.
[0049] Another technical solution adopted by the present invention is:
[0050] An automatic guzheng composition device, comprising:
[0051] At least one processor;
[0052] At least one memory for storing at least one program;
[0053] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned method.
[0054] Another technical solution adopted by the present invention is:
[0055] A computer-readable storage medium stores a program executable by a processor, and the program executable by the processor is used to execute the method as described above when executed by the processor.
[0056] The beneficial effects of the present invention are: The present invention uses the gated recurrent unit (GRU) algorithm to generate the melody of zheng string music, and then uses reinforcement learning to train in terms of traditional Chinese zither techniques of traditional Chinese music, making the generated music closer to real traditional Chinese zither music. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the accompanying drawings related to the technical solutions in the embodiments of the present invention or the prior art. It should be understood that the accompanying drawings in the following introduction are only for conveniently and clearly presenting some embodiments of the technical solutions in the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0058] Figure 1 It is a schematic flowchart of a method for automatically composing zheng music in an embodiment of the present invention;
[0059] Figure 2 It is a schematic diagram of the GRU network structure in an embodiment of the present invention;
[0060] Figure 3 It is a flowchart of the composition part of the GRU network in an embodiment of the present invention;
[0061] Figure 4 It is a schematic diagram of the deep Q-network (DQN) for reinforcement learning in an embodiment of the present invention;
[0062] Figure 5 It is a schematic diagram of the timbre conversion software in an embodiment of the present invention;
[0063] Figure 6 It is a schematic diagram of the final result of DQN training for reinforcement learning in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as a limitation to the present invention. For the step numbers in the following embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0065] In the description of the present invention, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc., is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention.
[0066] In the description of the present invention, the meaning of several is one or more, the meaning of multiple is two or more, greater than, less than, exceeding, etc. are understood as not including the present number, and above, below, within, etc. are understood as including the present number. If the first and second are described only for the purpose of distinguishing technical features, they should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.
[0067] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense. Those skilled in the art can reasonably determine the specific meaning of the above terms in the present invention in combination with the specific content of the technical solution.
[0068] Term Explanation:
[0069] Gated Recurrent Unit: The Gated Recurrent Unit (GRU) is an improved network of the Recurrent Neural Network (RNN). It is most suitable for processing and predicting important events at intervals in time series, and is simpler than the Long Short-Term Memory (LSTM) network. Therefore, it can solve problems more quickly.
[0070] Reinforcement Learning: Reinforcement Learning (RL) is one of the methods of machine learning, which is used to describe and solve the problem that an agent (agent) maximizes the reward or achieves a specific goal by learning strategies during the interaction with the environment.
[0071] DQN: DQN, Deep Q-learning is one of the implementation methods of reinforcement learning.
[0072] MIDI: It is one of the ways to represent music in a computer.
[0073] Existing machine learning methods have achieved certain performance in the automatic composition of Western musical instruments with beautiful melodies. However, in current domestic research, due to the small amount of folk music data, unique timbres, and the special nature of its techniques that cannot be correctly expressed in MIDI format, this may greatly reduce the performance of machine learning, making it impossible to produce good melodies and approach the characteristics of Chinese folk music. Therefore, there is no research on the automatic composition of folk music. Therefore, the purpose of the embodiments of the present invention is to use the GRU and reinforcement learning methods. Through the preprocessing link of chorus extraction, after enriching the data volume, use the GRU network to train to generate good melodies and use reinforcement learning to generate a certain amount of folk music techniques. Then use a guzheng player to play, so as to be able to generate beautiful and similar guzheng melodies in the case of small data volume and the absence of a standard timbre format.
[0074] As Figure 1 shown, this embodiment provides a method for automatic guzheng composition, including the following steps:
[0075] S1. Obtain the numbered musical notation data of the guzheng, convert the numbered musical notation data of the guzheng into staff notation data, obtain MIDI format data according to the staff notation data, and construct a guzheng dataset.
[0076] For the collection of the guzheng dataset, adopt the method of manual score input. By converting the numbered musical notation of the guzheng into staff notation and then into MIDI format that can be recognized by a computer, construct a unique guzheng dataset.
[0077] S2. Obtain the chorus part of each song in the guzheng dataset, and use the chorus part and the full song as the training set.
[0078] For the generation of the guzheng music melody, a preprocessing link and a music generation link are proposed. For the preprocessing link, this article will use a similarity function for calculating the similarity between fragments as the preprocessing of music, and extract the main melody (i.e., the chorus part) of each guzheng piece as part of the dataset.
[0079] S3. Use the training set to train the GRU network, and obtain preliminary melody data after training.
[0080] For the automatic composition link, this embodiment will use a GRU network, merge and train the main melody and all music data, establish a model, and generate corresponding music results.
[0081] S4. Perform reinforcement learning training on the preliminary melody data to obtain music data with zheng performance techniques.
[0082] For the special technique processing link that is different between the zheng and Western musical instruments, the DQN method of reinforcement learning is proposed. Three techniques are selected as examples to formulate rules and train the reinforcement learning training parameters, so that the notes generated by the GRU are re-trained by reinforcement learning to generate zheng performance techniques.
[0083] S5. Perform timbre conversion on the music data with zheng performance techniques to generate music data with zheng timbre.
[0084] For the generation of the zheng music timbre, a post-processing link is proposed. For the zheng timbre, the method of removing the timbre in the training part and reusing the zheng timbre through the post-processing link is adopted. In the final performance link, the generated zheng music is played by using the MIDI player that records the zheng timbre.
[0085] The above method will be explained in detail with the accompanying drawings in combination with specific embodiments.
[0086] Step 1: Obtain data and preprocess the data.
[0087] The original zheng music data all comes from the zheng grading examination repertoire. Select representative pieces from levels 1 - 7. Each zheng music is in the form of numbered musical notation. Convert each note of the numbered musical notation into the standard of staff notation, and after changing it to the staff notation format, import it into the Musescore software for the visualization of the staff notation. Finally, use the built-in options of the Musescore software to export it in MIDI format and save the data. Since the MIDI format does not have the timbre data of the zheng, the timbre in the MIDI format will be replaced by the piano timbre (the timbre conversion will be performed in the post-processing link).
[0088] To more intuitively see the structural hierarchy of the music, the self-similarity matrix will be used for representation, and the prelude, interlude, and coda of the song will be visually edited through mathematical methods. Scan the MIDI audio 100 times per second, with each frame being 0.1 seconds, and each frame being 128-dimensional. Generate a two-dimensional matrix, with the two axes being the scanned frame numbers. For each point of the self-similarity matrix S, calculate the distance function using the Euclidean distance formula. The Euclidean distance formula is as follows:
[0089]
[0090] Traverse the entire song composed of a set of frames of 12 basic notes, compare parts of a selected length with other parts and calculate the similarity to check for repetitions. When the similarity score is relatively high, indicating that two segments of music are very similar, extract the segment with the high score as the chorus part of the song. Among them, a similarity function is used to calculate the similarity, and the formula is as follows:
[0091]
[0092] In the formula, X i and X i+1 are the note vectors of any two segments of music. A very high S score indicates that the two segments of music are very similar.
[0093] Step 2: Compose music based on the gated recurrent unit (GRU) network.
[0094] In this stage, in the part of initially generating the track, this embodiment adopts a multi-input multi-output GRU network structure, and the network structure of one GRU is as Figure 2 shown.
[0095] The training process is as Figure 3 shown, and the training is based on the minimization of the cross-entropy function, that is, the minimization of Formula 3:
[0096]
[0097] The training process includes the following 6 steps:
[0098] A1: Convert the sheet music into MIDI format. Through the MIDI format, a piece of music can be represented as a matrix: x = {x t,j , t = 1, 2, …, N, j = 1, 2, …, 128};
[0099] A2: Encode the data using One-Hot encoding to obtain the input matrix: X = {X t,j , t = 1, 2, …, N, j = 1, 2, …, 128};
[0100] A3: Input the data into the LSTM network as Figure 2 shown to obtain the output matrix: Y = {Y t,j , t = 1, 2, …, M, j = 1, 2, …, 128};
[0101] A4: Perform One-Hot decoding to make the output matrix Y become y = {y t,j , t = 1, 2, …, M, j = 1, 2, …, 128};
[0102] A5: Use the built-in MIDI library in Python to convert y into MIDI format;
[0103] A6: Convert the MIDI format into musical notes.
[0104] Step 3: The generation part of the deep Q - network (DQN) technique of reinforcement learning.
[0105] In the third stage, the data generated by the GRU is sent to the processing stage of the reinforcement learning DQN. Set the technique reward and punishment mechanism, set the DQN network parameters, and after training the DQN network, high - quality guzheng technique data is generated, effectively solving the problem that there are no guzheng techniques in the data generated by the GRU. Figure 4 This is the basic logic of the reinforcement learning DQN.
[0106] First, set the reward and punishment rules for the techniques. The design of the reward policy is to capture the characteristics of guzheng music. Taking arpeggio as an example, this is a technique usually composed of seven notes. Many even - numbered beats in guzheng music end with arpeggios. Therefore, assume that at time t, the input state is where represents the 7 notes generated at time t, then the reward policy is defined as shown in Equation 4, and the reward and punishment mechanisms for other techniques are shown in Equations 5 and 6.
[0107]
[0108]
[0109]
[0110] To sum up and organize, the total reward at time t is obtained, as shown in Equation 7:
[0111]
[0112] Define the total return as Equation 8, and calculate the average return rate formula 9, which is used as one of the bases for the subsequent network to stop calculating.
[0113] R total = R t + γR t-1 + γ 2 R t-2 +…+ γ n-t R1 (8)
[0114]
[0115] The selection of each action is based on the current maximum action selection mechanism, as shown in Equation 10. The direction of the next action is calculated based on the previous action, approaching the target Q - value infinitely, that is, as shown in Equation 11.
[0116]
[0117]
[0118] The loss calculation of the entire network is shown in Equation 12, which is one of the bases for subsequent network termination calculation.
[0119]
[0120] The complete steps for the reinforcement learning DQN to generate music include:
[0121] B1. Initialize the experience pool D and the DNN network parameters θ in DQN t , set the value of action Q as a random number, and set θ t - = θ t ;
[0122] B2. Input y generated by GRU as the initial value of DQN;
[0123] B3. Calculate the reward R using Equation 7 t , and calculate the average return rate using Equation 9
[0124] B4. Calculate the action value Q(s t , y t ; θ t ) using Equation 10, and calculate the target Q value using Equation 11
[0125] B5. Solve the loss function defined in Equation 12 to obtain (s t , R t ; θ t ), and put it into the experience pool D;
[0126] B6. If the target ( and L → 0) is reached, stop the iteration; otherwise, update θ t - = θ t , t = t + 1, and go to step B3.
[0127] Step 4: Conversion of the timbre of traditional Chinese zither.
[0128] MIDI has many built-in timbre settings, with different instruments and sounds, but no traditional Chinese zither. To convert the piano timbre of the generated music into the timbre of a traditional Chinese zither, we use software for conversion, and embed the MIDI timbre of each string of the zither into the software as shown in Figure 5 , and play the final generated track through this player.
[0129] For the DQN of reinforcement learning, the average return defined in Equation 9 and the loss function defined in Equation 12 are used to evaluate the output. As Figure 6 shown, after covering 1400 generations, the final audio track is obtained.
[0130] As can be seen from the above, a method for automatic composition of traditional Chinese music guzheng provided in this embodiment realizes the basic function of automatic composition by using a gated recurrent unit network (GRU) to generate multiple pieces of guzheng music data, and enriches the guzheng techniques through a reinforcement learning DQN network to realize the perfect function of automatic composition of traditional Chinese music. The elements of guzheng music research will be divided into three parts: melody, technique, and timbre. Different treatments will be carried out on the three parts separately to achieve the effect that the finally generated music conforms to guzheng music. In the processing of the melody part of guzheng music, the other two elements, timbre and technique, are adjusted to fixed values for processing. The timbre will be specified as the timbre of a certain piece of music, and the beat will be processed as 4 / 4 beats for the whole piece for training.
[0131] At the same time, a preprocessing link will be added to the melody part. The preprocessing will extract the chorus part of guzheng music from all guzheng data sets through the algorithm of the similarity function, which will be used as part of the data for subsequent model establishment. The algorithm part of melody generation will adopt the GRU model. After generating the optimal model, music generation will be carried out.
[0132] After generating the music, the technique processing of reinforcement learning will be carried out. Sort out the techniques of the guzheng, determine the reward mechanism for each technique, and add it to the generated music after DQN network calculation. At the same time, there are important differences in timbre between the generation of guzheng music and the generation of traditional ordinary music. Therefore, the post-processing of guzheng music will add a timbre part. The timbre of each string of the guzheng will be collected and stored in the MIDI player, and the guzheng timbre will be broadcast in MIDI format.
[0133] Therefore, the method of this embodiment can well solve the problem of automatic composition of various technical music and improve the accuracy of automatic composition.
[0134] In summary, the method of this embodiment has the following advantages and beneficial effects compared with the prior art:
[0135] (1) After obtaining the traditional Chinese music data in this embodiment, in the preprocessing stage of traditional Chinese music data, the similarity algorithm is used to extract the chorus part. Combining the traditional Chinese music songs with the chorus part extracted by the similarity algorithm to create a unique database, enrich the original data set, and effectively solve the problem of small data volume samples before the training of the gated recurrent unit (GRU).
[0136] (2) This embodiment uses the DQN algorithm of reinforcement learning and sets up a reward and punishment mechanism for training in traditional Chinese zither techniques, making the generated music closer to real traditional Chinese zither music. The combination of the gated recurrent unit (GRU) and the DQN (Deep Q-learning) of reinforcement learning can effectively extract the complex technique features in traditional Chinese music techniques, improving the accuracy and richness of the generated music.
[0137] (3) This embodiment uses a traditional Chinese music timbre correction player to process the timbre of traditional Chinese music, making a breakthrough in automatic composition jointly from the aspects of timbre, melody, and technique, showing great potential in practical applications.
[0138] (4) This embodiment uses the GRU algorithm of the gated recurrent unit to generate the melody of traditional Chinese zither string music. Compared with other machine learning methods in existing automatic composition technologies, our gating system is simpler than the LSTM gating system, achieving the same excellent effect while having a faster training speed.
[0139] This embodiment also provides a traditional Chinese zither automatic composition system, including:
[0140] A data acquisition module, used to acquire the numbered musical notation data of the traditional Chinese zither, convert the numbered musical notation data of the traditional Chinese zither into staff notation data, acquire MIDI format data according to the staff notation data, and construct a traditional Chinese zither dataset;
[0141] A data preprocessing module, used to acquire the chorus part of each song in the traditional Chinese zither dataset and use the chorus part and the full song as the training set;
[0142] A preliminary composition module, used to train the GRU network with the training set and obtain preliminary melody data after training;
[0143] A technique generation module, used to perform reinforcement learning training on the preliminary melody data to obtain music data with traditional Chinese zither performance techniques;
[0144] A timbre conversion module, used to perform timbre conversion on the music data with traditional Chinese zither performance techniques to generate music data with traditional Chinese zither timbre.
[0145] A traditional Chinese zither automatic composition system of this embodiment can execute a traditional Chinese zither automatic composition method provided by the method embodiment of the present invention, can execute any combination of implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.
[0146] This embodiment also provides a traditional Chinese zither automatic composition device, including:
[0147] At least one processor;
[0148] At least one memory, used to store at least one program;
[0149] When the at least one program is executed by the at least one processor, the at least one processor implements Figure 1 the method shown.
[0150] An automatic guzheng composition device according to this embodiment can execute an automatic guzheng composition method provided in an embodiment of the method of the present invention, can execute any combination of implementation steps of the method embodiment, and has corresponding functions and beneficial effects of the method.
[0151] An embodiment of the present application also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 the method shown.
[0152] This embodiment also provides a storage medium storing instructions or a program that can execute an automatic guzheng composition method provided in an embodiment of the method of the present invention. When the instructions or the program are run, any combination of implementation steps of the method embodiment can be executed, and corresponding functions and beneficial effects of the method are provided.
[0153] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order mentioned in the operation diagrams. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated, in which the order of various operations is changed and sub-operations described as part of a larger operation are executed independently.
[0154] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the described functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. Rather, considering the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Thus, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0155] If the described functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0156] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a predefined sequence of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0157] More specific examples (nonexhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0158] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0159] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0160] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
[0161] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A method for automatically composing Guzheng music, characterized in that, It includes the following steps: Obtain the numbered musical notation data of the guzheng, convert the numbered musical notation data of the guzheng into staff notation data, obtain MIDI format data according to the staff notation data, and construct a guzheng dataset; Obtain the chorus part of each song in the guzheng dataset, and use the chorus part and the full song as the training set; Train the GRU network using the training set, and obtain preliminary melody data after training; Conduct reinforcement learning training on the preliminary melody data to obtain music data with guzheng performance techniques; Perform timbre conversion on the music data with guzheng performance techniques to generate music data with guzheng timbre; The conducting reinforcement learning training on the preliminary melody data to obtain music data with guzheng performance techniques includes: Set the reward and punishment rules for the techniques and set the parameters of the reinforcement learning network; Input the preliminary melody data into the reinforcement learning network, and process it according to the reward and punishment rules to obtain music data with guzheng performance techniques; The expression of the loss function of the reinforcement learning network is as follows: where, E D represents the expectation of the agent random variable, R t is the reward and punishment, γ is the attenuation factor, and y t is the current action state; Q(s t ,y t ; θ t - ) represents the output of the target value state, Q(s t ,y t ; θ t ) represents the output of the current value state, θ t and θ t - are the iterative parameters of the network at time t; The processing steps in the reinforcement learning DQN network are as follows: Initialize the experience pool D and initialize the DNN network parameters θ in the DQN t , set the value of the action value Q to a random number, and set θ t - = θ t ; Use the preliminary melody data y output by the GRU network as the initial value of the reinforcement learning DQN; Calculate the reward and punishment compensation R t , according to the reward and punishment compensation R t Calculate the average rate of return Calculate the action value Q(s t , y t ; θ t ), and calculate the target Q value according to the action value Q Solve the loss function of the reinforcement learning network to obtain (s t , R t ; θ t ), and put (s t , R t ; θ t ) into the experience pool D; According to the average return rate and the loss L, determine whether to stop training; if not, update θ t - = θ t , t = t + 1, and recalculate the reward and punishment reward R t .
2. The method for automatically composing Guzheng music according to claim 1, characterized in that, The converting the numbered musical notation data of the guzheng into staff notation data and obtaining MIDI format data according to the staff notation data includes: Convert each note in the numbered musical notation data according to the staff notation standard to obtain staff notation data; Use Musescore software to convert the staff notation data into MIDI format data; among them, in the MIDI format, the piano timbre is used instead of the guzheng timbre.
3. The method for automatically composing Guzheng music according to claim 1, characterized in that, The obtaining the chorus part of each song in the guzheng dataset includes: Segment the MIDI audio of the song, calculate the similarity between any two segments of audio using a similarity function, and obtain the chorus part of the song according to the similarity.
4. The method for automatically composing Guzheng music according to claim 3, characterized in that, The expression of the similarity function is: where X i and X i+1 are the note vectors of any two segments of music, and the whole song consists of a set of frames of 12 basic notes; Use the segment with a high similarity score as the chorus part of the song.
5. The method for automatically composing Guzheng music according to claim 1, characterized in that, The process of training the GRU network includes: Represent a piece of MIDI - formatted music data as a matrix: \(x = \{x t,j , t = 1,2,\cdots,N, j = 1,2,\cdots,128\}\); Encode the music data using One-Hot encoding to obtain the input matrix: X = {X t,j , t = 1, 2, …, N, j = 1, 2, …, 128}; Input the input matrix into the GRU network to obtain the output matrix: Y = {Y t,j , t = 1, 2, …, M, j = 1, 2, …, 128}; Perform One-Hot decoding on the output matrix to obtain the melody matrix: y = {y t,j , t = 1, 2, …, M, j = 1, 2, …, 128}; The loss function of the GRU network is the cross-entropy function, and the expression of the cross-entropy function is as follows: where P represents the number of training samples, and Y i represents the actual output of the i-th training sample, and Y j represents the expected output of the i-th training sample.
6. An automatic guzheng composition system, applied to an automatic guzheng composition method described in any one of claims 1-5, characterized in that, It includes: A data acquisition module, which is used to obtain the numbered musical notation data of the guzheng, convert the numbered musical notation data of the guzheng into staff notation data, obtain MIDI format data according to the staff notation data, and construct a guzheng dataset; A data preprocessing module, which is used to obtain the chorus part of each song in the guzheng dataset, and use the chorus part and the full song as the training set; A preliminary composition module, which is used to train the GRU network using the training set and obtain preliminary melody data after training; A technique generation module, which is used to conduct reinforcement learning training on the preliminary melody data to obtain music data with guzheng performance techniques; A timbre conversion module, which is used to perform timbre conversion on the music data with guzheng performance techniques to generate music data with guzheng timbre.
7. An automatic guzheng composition device, characterized in that, It includes: At least one processor; At least one memory, which is used to store at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method described in any one of claims 1-5.
8. A computer-readable storage medium, in which a program executable by a processor is stored, characterized in that, The program executable by the processor, when executed by the processor, is used to execute the method described in any one of claims 1-5.
Citation Information
Patent Citations
A taxpayer industry two-level classification method based on an MIMO recurrent neural network
CN109710768A
Audio processing method, audio processing device, electronic equipment and storage medium
CN109979418A
Automatic composition system and method for note vectors based on contextual information
CN111583891A
Music composition method based on deep reinforcement learning
KR1020170128073A