Music generation device, music generation method, music generation program, model generation device, model generation method, and model generation program
The music generation device uses a trained generative model to adjust music difficulty levels, addressing the monotony of conventional methods and enabling diverse and high-quality song generation suitable for commercial use.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- YAMAHA CORP
- Filing Date
- 2026-02-26
- Publication Date
- 2026-05-01
AI Technical Summary
Conventional methods for automating the adjustment of music difficulty levels result in monotonous arrangements and fail to accommodate diverse difficulty levels, making it difficult to generate high-quality arranged songs suitable for commercial use.
A music generation device and method that utilizes a trained generative model to adjust music difficulty levels based on a difficulty parameter, allowing for diverse control over the difficulty of generated songs by specifying parameters such as playable pitch range and maximum simultaneous controls.
Enables the easy generation of music with varying difficulty levels, accommodating different performer skills and preferences, thereby enhancing the quality and versatility of arranged songs for commercial applications.
Smart Images

Figure 2026074295000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a music generation device, a music generation method, a music generation program, a model generation device, a model generation method, and a model generation program.
Background Art
[0002] Conventionally, the difficulty level of music has mainly been changed manually by humans. However, if all the work of changing the difficulty level of music and generating new music is done manually, the cost of such work becomes high. Therefore, the development of a method for automating at least a part of the work on music using computer technology has been underway.
[0003] For example, Patent Document 1 proposes a technique for automatically generating simplified score information based on rules. According to this method, at least a part of the work of changing the difficulty level can be automated. Therefore, it is possible to reduce the cost of the work of changing the difficulty level.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] The inventors of this case have found the following problems with the conventional difficulty level modification method described above. Specifically, they have found that even when attempting to generate new songs with modified difficulty levels based on rules, the arrangements tend to become uniform. For example, the conventional method generates arranged songs using simple rules such as omitting notes. However, this method tends to result in monotonous arrangements, and the quality of the arranged songs is low. For example, the appropriate difficulty level of a song can vary depending on the performer (for example, certain controls may be difficult to operate depending on the size of the performer's hands). The conventional method has the problem that it is difficult to accommodate the specification of diverse difficulty levels, and therefore difficult to apply to commercial uses such as the automatic generation of arranged songs and the creation of corresponding sheet music.
[0006] In one respect, this invention was made in consideration of these circumstances, and its purpose is to provide a technology that can easily generate musical pieces of varying difficulty levels. [Means for solving the problem]
[0007] To solve the above-mentioned problems, the present invention employs the following configuration.
[0008] In other words, a music generation device according to one aspect of the present invention comprises: a data acquisition unit that acquires target music data representing at least a part of a music; a parameter acquisition unit that acquires a value of a difficulty parameter for specifying the modified difficulty of the music; a generation unit that uses a trained generation model to generate new music data representing at least a part of a new music obtained by changing the difficulty of the music to the difficulty specified by the difficulty parameter, from the acquired target music data and the value of the difficulty parameter; and an output unit that outputs the generated new music data.
[0009] In this configuration, a difficulty parameter is used in the inference process for generating new songs. The difficulty parameter allows for diverse control over the difficulty of the generated new songs. For example, by setting the difficulty level specified by the difficulty parameter to multiple levels, new songs corresponding to each specified level can be generated. Furthermore, it's possible to generate not only songs with reduced difficulty, but also songs with increased difficulty. In other words, by specifying the difficulty level using the difficulty parameter, it's possible to adjust the level of new songs, for example, by increasing or decreasing the difficulty. Therefore, this configuration makes it easy to generate songs of various difficulty levels.
[0010] In the music generation device relating to the above aspect, the value of the difficulty parameter may be configured to indicate the range of the playable pitch of the new music after changing the difficulty level. With this configuration, the difficulty level can be specified by the range of the playable pitch. This makes it possible to easily generate music with various ranges of playable pitch.
[0011] In the music generation device relating to the above aspect, the value of the difficulty parameter may be configured to indicate the maximum number of controls that can be operated simultaneously on the instrument used for the new music after the difficulty has been changed. With this configuration, the difficulty can be specified by the maximum number of controls that can be operated simultaneously. This makes it easy to generate music with various maximum number of controls.
[0012] In the music generation device relating to the above aspect, the target music data may consist of an input token sequence arranged to represent at least a part of the music, and the new music data may consist of an output token sequence output from the trained generation model and arranged to represent at least a part of the new music. With this configuration, music of various difficulty levels can be accurately generated using a machine learning model (for example, a Transformer described later) configured to handle token sequences.
[0013] Embodiments of the present invention are not limited to music generation devices configured to generate new music using a trained generative model. One aspect of the present invention may be a model generation device configured to generate the trained generative model used in any of the above embodiments by machine learning.
[0014] For example, a model generation apparatus according to one aspect of the present invention includes: a learning data acquisition unit that acquires a plurality of learning datasets, each composed of a combination of training data and ground truth data, wherein the training data includes learning song data representing at least a part of a song and a difficulty parameter for learning, and the ground truth data includes new learning song data representing at least a part of a new song generated by changing the difficulty of the song in the learning song data to the difficulty level specified by the difficulty parameter; and a learning processing unit that performs machine learning of a generative model using the acquired plurality of learning datasets, wherein the machine learning is performed by training the generative model so that, for each learning dataset, the song data generated by the generative model from the values of the learning song data and the difficulty parameter included in the training data fits the learning song data included in the ground truth data. According to this configuration, it is possible to generate a trained generative model that has acquired the ability to generate songs of various difficulty levels based on the difficulty parameter.
[0015] In the model generation device relating to the above aspect, the value of the difficulty parameter included in the training data may be configured to indicate the range of the performance pitch of the new song after changing the difficulty level. With this configuration, it is possible to generate a trained generation model that has acquired the ability to generate songs with various performance pitch ranges.
[0016] In the model generation device relating to the above aspect, the value of the difficulty parameter included in the training data may be configured to indicate the maximum number of controls to be operated simultaneously on the instrument used for the new song after the difficulty has been changed. With this configuration, it is possible to generate a trained generation model that has acquired the ability to generate songs with various maximum numbers of controls.
[0017] As another form of the music generation apparatus and model generation apparatus according to each of the above forms, one aspect of the present invention may be an information processing method, an information processing system, a program, or a storage medium readable by a computer or other device or machine that stores such a program. Here, a storage medium readable by a computer or the like is a medium that stores information such as a program by electrical, magnetic, optical, mechanical, or chemical action.
[0018] For example, a music generation method according to one aspect of the present invention is an information processing method in which a computer performs the steps of: acquiring target music data that represents at least a part of a music; acquiring a value of a difficulty parameter; using a trained generation model, generating new music data that represents at least a part of a new music obtained by changing the difficulty of the music to the difficulty specified by the difficulty parameter from the acquired target music data and the value of the difficulty parameter; and outputting the generated new music data.
[0019] Furthermore, for example, a music generation program according to one aspect of the present invention is a program for a computer to perform the following steps: acquiring target music data representing at least a part of a music; acquiring a value for a difficulty parameter; using a trained generation model, generating new music data representing at least a part of a new music obtained by changing the difficulty of the music to the difficulty specified by the difficulty parameter from the acquired target music data and the value for the difficulty parameter; and outputting the generated new music data.
[0020] Furthermore, for example, a model generation method according to one aspect of the present invention is an information processing method that includes the steps of: a computer acquiring a plurality of learning datasets, each composed of a combination of training data and ground truth data, wherein the training data includes learning music data representing at least a portion of a musical piece and a difficulty parameter for learning, and the ground truth data includes new learning music data representing at least a portion of a new musical piece generated by changing the difficulty level of the musical piece in the learning music data to the difficulty level specified by the difficulty parameter; and performing machine learning of a generative model using the acquired plurality of learning datasets, wherein the machine learning consists of training the generative model for each learning dataset so that the musical piece data generated by the generative model from the values of the learning music data and the difficulty parameter included in the training data fits the learning music data included in the ground truth data.
[0021] Further, for example, a model generation program according to one aspect of the present invention causes a computer to perform steps of acquiring a plurality of learning data sets each constituted by a combination of training data and correct answer data, wherein the training data includes learning music data indicating at least a part of a music piece and a difficulty parameter for learning, and the correct answer data includes new learning music data indicating at least a part of a new music piece generated by changing the difficulty of the music piece of the learning music data to a difficulty specified by the difficulty parameter; and performing machine learning of a generation model using the acquired plurality of learning data sets, the machine learning being configured by training the generation model so that music data generated by the generation model from the values of the learning music data and the difficulty parameter included in the training data conforms to the learning music data included in the correct answer data for each of the learning data sets. The program is for causing the above steps to be executed.
Advantages of the Invention
[0022] According to the present invention, it is possible to provide a technique capable of easily generating music pieces of various difficulties.
Brief Description of the Drawings
[0023] [Figure 1] FIG. 1 schematically shows an example of a scene to which the present invention is applied. [Figure 2] FIG. 2 schematically shows an example of the hardware configuration of a model generation device according to an embodiment. [Figure 3] FIG. 3 schematically shows an example of the hardware configuration of a music piece generation device according to an embodiment. [Figure 4] FIG. 4 schematically shows an example of the software configuration of a model generation device according to an embodiment. [Figure 5] FIG. 5 is a musical score showing an example of a music piece. [Figure 6A] FIG. 6A shows an example of an input token sequence generated from the music piece of FIG. 5. [Figure 6B]Figure 6B shows an example of an input token sequence generated from the music in Figure 5. [Figure 7] Figure 7 shows an example of a musical score (generated result) where the difficulty level has been changed. [Figure 8A] Figure 8A shows an example of the true value of the output token sequence corresponding to the song in Figure 7. [Figure 8B] Figure 8B shows an example of the true value of the output token sequence corresponding to the song in Figure 7. [Figure 9] Figure 9 schematically shows an example of the configuration of the generation model according to the embodiment. [Figure 10] Figure 10 schematically shows an example of the software configuration of a music generation device according to an embodiment. [Figure 11] Figure 11 is a flowchart showing an example of the processing procedure of the model generation device according to the embodiment. [Figure 12] Figure 12 is a flowchart showing an example of the processing procedure of a music generation device according to an embodiment. [Modes for carrying out the invention]
[0024] Hereinafter, an embodiment relating to one aspect of the present invention (hereinafter also referred to as "this embodiment") will be described based on the drawings. However, this embodiment described below is merely illustrative in all respects of the present invention. Needless to say, various improvements and modifications can be made without departing from the scope of the present invention. In other words, in carrying out the present invention, specific configurations according to the embodiment may be adopted as appropriate. Although the data appearing in this embodiment is described in natural language, more specifically, it is specified in pseudo-language, commands, parameters, machine code, etc., that can be recognized by a computer.
[0025] §1 Examples of Application Figure 1 schematically illustrates an example of a scenario in which the present invention is applied. As shown in Figure 1, the generation system 100 according to this embodiment comprises a model generation device 1 and a music generation device 2.
[0026] The model generation device 1 according to this embodiment is a computer configured to generate a trained generation model 5 by machine learning for generating new songs with altered difficulty levels from original songs. First, the model generation device 1 acquires a plurality of training datasets 3. Each training dataset 3 consists of a combination of training data 31 and ground truth data 32. The training data 31 is configured to include training song data 311 that represents at least a portion of a song and a difficulty parameter 313 for learning. The ground truth data 32 is configured to include new training song data 321 that represents at least a portion of a new song generated by changing the difficulty level of the song in the training song data 311 to the difficulty level specified by the difficulty parameter 313.
[0027] Next, the model generation device 1 performs machine learning on the generative model 5 using the acquired training datasets 3. The machine learning consists of training the generative model 5 so that, for each training dataset 3, the music data generated by the generative model 5 from the training music data 311 and the difficulty parameter 313 values contained in the training data 31 fits the new training music data 321 contained in the ground truth data 32. Through this machine learning process, a trained generative model 5 can be produced that has acquired the ability to generate new music of a difficulty level specified by the difficulty parameter from the original music.
[0028] On the other hand, the music generation device 2 according to this embodiment is a computer configured to generate a new music from an original music using a trained generation model 5. First, the music generation device 2 acquires target music data 221 that represents at least a part of the music. The music generation device 2 also acquires a value for a difficulty parameter 223 to specify the difficulty level of the music after modification. Next, using the trained generation model 5, the music generation device 2 generates new music data 225 that represents at least a part of the new music obtained by changing the difficulty level of the music to the difficulty level specified by the difficulty parameter 223, from the acquired target music data 221 and the value of the difficulty parameter 223. The music generation device 2 outputs the generated new music data 225.
[0029] The difficulty parameters (223, 313) are configured to indicate the difficulty level of the new song generated by the generative model 5 from the original song. In the learning phase, the difficulty parameter 313 is configured to indicate the difficulty level of the song shown by the learning song data 321 that constitutes the ground truth data 32, relative to the learning song data 311 that constitutes the training data 31. In the inference phase, the difficulty parameter 223 is configured to indicate the difficulty level of the new song generated by the trained generative model 5 from the target song data 221.
[0030] The difficulty level can be specified in any format. For example, the difficulty level may be set to multiple levels (e.g., "Beginner," "Lower Intermediate," "Intermediate," "Upper Intermediate," "Advanced," etc.), and the difficulty parameters (223, 313) may be configured to indicate one of these multiple levels.
[0031] In another example, difficulty may be specified by the range of the playable notes. In this case, the difficulty parameters (223, 313) may be configured to indicate the range of the playable notes of the new piece after the difficulty has been changed.
[0032] In another example, the difficulty level may be specified by the maximum number of controls that can be operated simultaneously on the instrument used for performance. In this case, the difficulty parameter (223, 313) may be configured to indicate the maximum number of controls that can be operated simultaneously on the instrument used for the new piece of music after the difficulty level has been changed. The instrument used for the piece of music may be selected as appropriate depending on the embodiment. The instrument may be, for example, a piano. The controls may be, for example, piano keys. In the difficulty parameter (223, 313), the difficulty level may be specified relatively or absolutely.
[0033] As described above, in this embodiment, the generation model 5 is configured to generate a new song from an input original song based on the difficulty level specified by the difficulty parameter. This difficulty parameter allows for diverse control over the difficulty level of the newly generated song. Therefore, the model generation device 1 can generate a trained generation model 5 that has acquired the ability to generate songs of various difficulty levels based on the difficulty parameter. The song generation device 2 can easily generate songs of various difficulty levels by using such a trained generation model 5.
[0034] In the example shown in Figure 1, the model generation device 1 and the music generation device 2 are connected to each other via a network. The type of network may be appropriately selected from, for example, the Internet, a wireless communication network, a mobile communication network, a telephone network, a dedicated network, etc. However, the method of exchanging data between the model generation device 1 and the music generation device 2 is not limited to this example and may be appropriately selected depending on the embodiment. For example, data may be exchanged between the model generation device 1 and the music generation device 2 using a storage medium.
[0035] Furthermore, in the example shown in Figure 1, the model generation device 1 and the music generation device 2 are each configured by separate computers. However, the configuration of the generation system 100 according to this embodiment is not limited to this example and may be determined as appropriate depending on the embodiment. For example, the model generation device 1 and the music generation device 2 may be a single computer. Also, for example, at least one of the model generation device 1 and the music generation device 2 may be configured by multiple computers. When configured by multiple computers, the distribution of information processing may be determined as appropriate depending on the embodiment.
[0036] §2 Example Configuration [Hardware configuration] <Model Generator> Figure 2 schematically illustrates an example of the hardware configuration of the model generation device 1 according to this embodiment. As shown in Figure 2, the model generation device 1 according to this embodiment is a computer in which a control unit 11, a storage unit 12, a communication interface 13, an external interface 14, an input device 15, an output device 16, and a drive 17 are electrically connected. In Figure 2, the communication interface and the external interface are referred to as "communication I / F" and "external I / F," respectively.
[0037] The control unit 11 includes an example of a hardware processor (processor resource), such as a CPU (Central Processing Unit), RAM (Random Access Memory), and ROM (Read Only Memory), and is configured to perform information processing based on a program and various data. The storage unit 12 is an example of memory, and is composed of, for example, a hard disk drive or a solid-state drive. In this embodiment, the storage unit 12 stores various information such as a model generation program 81, multiple training datasets 3, and training result data 125.
[0038] The model generation program 81 is a program that causes the model generation device 1 to execute the machine learning information processing (Figure 11) described later, which generates a trained generative model 5. The model generation program 81 includes a series of instructions for said information processing. Multiple training datasets 3 are used to generate the trained generative model 5. The training result data 125 shows information about the generated trained generative model 5. In this embodiment, the training result data 125 is generated as a result of executing the model generation program 81. Details will be described later.
[0039] The communication interface 13 is, for example, a wired LAN (Local Area Network) module, a wireless LAN module, etc., and is an interface for wired or wireless communication over a network. The model generation device 1 can use the communication interface 13 to perform data communication over a network with other information processing devices. The external interface 14 is, for example, a USB (Universal Serial Bus) port, a dedicated port, etc., and is an interface for connecting to external devices. The type and number of external interfaces 14 may be arbitrarily selected.
[0040] The model generation device 1 may be connected to a device for obtaining each training dataset 3 via at least one of the communication interface 13 and the external interface 14. For example, the training music data 311 that constitutes the training data 31 may be obtained by an electronic musical instrument. In this case, the model generation device 1 may be connected to the electronic musical instrument via at least one of the communication interface 13 and the external interface 14, and the training music data 311 may be collected by the electronic musical instrument.
[0041] The input device 15 is a device for inputting data, such as a mouse or keyboard. The output device 16 is a device for outputting data, such as a display or speaker. Users or other operators can operate the model generation device 1 by using the input device 15 and the output device 16.
[0042] Drive 17 is, for example, a CD drive, DVD drive, etc., and is a drive device for reading various information such as programs stored in the storage medium 91. The storage medium 91 is a medium that stores information such as programs by electrical, magnetic, optical, mechanical, or chemical means so that computers and other devices, machines, etc., can read the stored information such as programs. At least one of the above model generation program 81 and the multiple training datasets 3 may be stored in the storage medium 91. The model generation device 1 may obtain at least one of the above model generation program 81 and the multiple training datasets 3 from this storage medium 91. In Figure 2, a disk-type storage medium such as a CD or DVD is shown as an example of the storage medium 91. However, the type of storage medium 91 is not limited to disk type and may be other types. Examples of storage media other than disk type include semiconductor memory such as flash memory. The type of drive 17 may be arbitrarily selected according to the type of storage medium 91.
[0043] Regarding the specific hardware configuration of the model generation device 1, components can be omitted, replaced, and added as appropriate depending on the embodiment. For example, the control unit 11 may include multiple hardware processors. The hardware processors may consist of microprocessors, FPGAs (field-programmable gate arrays), etc. The storage unit 12 may consist of RAM and ROM included in the control unit 11. At least one of the communication interface 13, external interface 14, input device 15, output device 16, and drive 17 may be omitted. The model generation device 1 may consist of multiple computers. In this case, the hardware configuration of each computer may or may not be the same. Furthermore, the model generation device 1 may be an information processing device designed specifically for the services provided, as well as a general-purpose server device, a PC (Personal Computer), etc.
[0044] <Music Generation Device> Figure 3 schematically illustrates an example of the hardware configuration of the music generation device 2 according to this embodiment. As shown in Figure 3, the music generation device 2 according to this embodiment is a computer in which a control unit 21, a storage unit 22, a communication interface 23, an external interface 24, an input device 25, an output device 26, and a drive 27 are electrically connected.
[0045] The control units 21 to 27 and the storage medium 92 of the music generation device 2 may be configured in the same way as the control units 11 to 17 and the storage medium 91 of the model generation device 1. The control unit 21 includes a CPU, RAM, ROM, etc., which are examples of hardware processors, and is configured to perform various information processing based on programs and data. The storage unit 22 is composed of, for example, a hard disk drive, a solid-state drive, etc. In this embodiment, the storage unit 22 stores various information such as the music generation program 82 and the learning result data 125.
[0046] The music generation program 82 is a program that causes the music generation device 2 to execute the information processing described later (Figure 12) which generates a new music by changing the difficulty level from the original music using the trained generation model 5. The music generation program 82 includes a series of instructions for said information processing. At least one of the music generation program 82 and the learning result data 125 may be stored in the storage medium 92. The music generation device 2 may also retrieve at least one of the music generation program 82 and the learning result data 125 from the storage medium 92.
[0047] The music generation device 2 may be connected to a device for obtaining target music data 221 via at least one of the communication interface 23 and the external interface 24. For example, the target music data 221 may be obtained from an electronic musical instrument. In this case, the music generation device 2 may be connected to the electronic musical instrument via at least one of the communication interface 23 and the external interface 24. The music generation device 2 may also accept operations and inputs from an operator such as a user by utilizing the input device 25 and the output device 26.
[0048] Regarding the specific hardware configuration of the music generation device 2, components can be omitted, replaced, and added as appropriate depending on the embodiment. For example, the control unit 21 may include multiple hardware processors. The hardware processors may consist of microprocessors, FPGAs, etc. The storage unit 22 may consist of RAM and ROM included in the control unit 21. At least one of the communication interface 23, external interface 24, input device 25, output device 26, and drive 27 may be omitted. The music generation device 2 may consist of multiple computers. In this case, the hardware configuration of each computer may or may not be the same. Furthermore, the music generation device 2 may be an information processing device designed specifically for the service provided, as well as a general-purpose server device, a general-purpose PC, etc.
[0049] [Software Configuration] <Model Generator> Figure 4 schematically illustrates an example of the software configuration of the model generation device 1 according to this embodiment. The control unit 11 of the model generation device 1 interprets the instructions contained in the model generation program 81 stored in the storage unit 12 and executes control processing according to the interpreted instructions. Thus, the model generation device 1 according to this embodiment is configured to include a learning data acquisition unit 111, a learning processing unit 112, and a storage processing unit 113 as software modules. In other words, in this embodiment, the software modules of the model generation device 1 are realized by the control unit 11 (CPU).
[0050] The learning data acquisition unit 111 is configured to acquire multiple learning datasets 3. Each learning dataset 3 consists of a combination of training data 31 and ground truth data 32. The training data 31 is configured to include learning song data 311 that represents at least a portion of a song and a difficulty parameter 313 for learning. The ground truth data 32 is configured to include new learning song data 321 that represents at least a portion of a new song generated by changing the difficulty level of the song in the learning song data 311 to the difficulty level specified by the difficulty parameter 313.
[0051] The learning processing unit 112 is configured to perform machine learning on the generative model 5 using the acquired multiple learning datasets 3. The machine learning consists of training the generative model 5 so that, for each learning dataset 3, the music data generated by the generative model 5 from the learning music data 311 and the difficulty parameter 313 values contained in the training data 31 fits the learning music data 321 contained in the ground truth data 32. Upon completion of this machine learning process, a trained generative model 5 is produced that has acquired the ability to generate new music (i.e., music with a changed difficulty level from the original music) by changing the difficulty level of the music indicated by the given music data to the difficulty level specified by the difficulty parameter.
[0052] The storage processing unit 113 is configured to generate information about the trained generative model 5 generated by machine learning as training result data 125, and to store the generated training result data 125 in a predetermined memory area. The training result data 125 may be configured as appropriate to include information for reconstructing the trained generative model 5.
[0053] (Example of song data) Music can be obtained in any format, such as encoded data (MIDI, etc.) or musical notation. Furthermore, during processing by the generative model 5, the music can be obtained in a format compatible with the generative model 5's processing. That is, the format of the target music data 221 and the training music data 311 is not particularly limited and can be determined appropriately depending on the embodiment, as long as it is processable by the generative model 5. The format of the new music data 225 and the new training music data 321 (ground truth data 32) is not particularly limited and can be determined appropriately depending on the embodiment, as long as it is obtainable by the generative model 5. As an example, the generative model 5 may be configured to accept at least a portion of the music before the difficulty level change as input in the form of a token sequence, and to output the result (inference result) of generating at least a portion of the new music with the difficulty level changed as a token sequence. Accordingly, the music can be represented by a token sequence.
[0054] If the generative model 5 is configured to handle token sequences, during the training phase, the training music data 311 may consist of input token sequences arranged to represent at least a portion of the music before the difficulty is changed (the original music). The training music data 321 may consist of true values of output token sequences arranged to represent at least a portion of the new music obtained by changing the difficulty (new music generated by changing the difficulty of the music in the training music data 311 to the difficulty specified by the difficulty parameter 313).
[0055] Similarly, in the generation (inference) stage, the target song data 221 may consist of an input token sequence arranged to represent at least a portion of the songs (original songs) whose difficulty level is to be changed. The new song data 225 generated by the trained generative model 5 may consist of an output token sequence arranged to represent at least a portion of the new songs obtained by changing the difficulty level of the songs in the target song data 221 to the difficulty level specified by the difficulty parameter 223.
[0056] The tokens constituting the input and output token sequences may include any symbols, such as numbers, letters, or figures. The symbols (token representations) and data formats used for the tokens are not particularly limited, as long as they are recognizable by a computer, and may be appropriately selected depending on the embodiment. Below, two tokenization methods, action-based and note-based, are given as examples of tokenization methods.
[0057] Figure 5 shows a musical score that is an example of at least a portion of the song before the difficulty level is changed. Figure 6A shows an example of an input token sequence generated from the song in Figure 5 using an action-based tokenization scheme. Figure 6B shows an example of an input token sequence generated from the song in Figure 5 using a note-based tokenization scheme. Figure 7 shows a musical score that is an example of at least a portion of the song after the difficulty level has been changed. Figure 8A shows an example of the true value of the output token sequence obtained for the song in Figure 7 using an action-based tokenization scheme. Figure 8B shows an example of the true value of the output token sequence obtained for the song in Figure 7 using a note-based tokenization scheme.
[0058] Action-based tokenization is a method of tokenizing to represent actions corresponding to notes or elements of a musical piece. Table 1 shows examples of token types and their representations in action-based tokenization. On the other hand, note-based tokenization is a method of tokenizing to represent notes of a musical piece directly. Table 2 shows examples of token types and their representations in note-based tokenization. Note that the following token types and representations are examples and may be modified as appropriate depending on the embodiment. [Table 1] [Table 2]
[0059] The tokenization method and token representation of the input token sequence and output token sequence may be one of the two methods described above. As an example of how to acquire each training dataset 3, performance information representing at least a part of the musical piece exemplified in Figure 5 may be appropriately acquired. The data format of the performance information may be appropriately selected depending on the embodiment. For example, the performance information may be obtained in the form of encoded data (MIDI, etc.), musical score, etc. The training musical piece data 311 in the training data 31 of each training dataset 3 may be appropriately generated from the acquired performance information so as to include the input token sequence shown in Figure 6A or Figure 6B. In addition, corresponding to at least a part of the musical piece exemplified in Figure 5, ground truth performance information showing the true value of the musical piece after changing the difficulty level, as exemplified in Figure 7, may be appropriately acquired. The new training musical piece data 321 constituting the ground truth data 32 of each training dataset 3 may be appropriately generated from the acquired ground truth performance information so as to include the true value of the output token sequence shown in Figure 8A or Figure 8B. Any conversion process, such as natural language processing, may be employed to generate the true values of the input token sequence and output token sequence. The true values of the input token column and the output token column may be generated manually.
[0060] Furthermore, the input token sequence may include multiple beat tokens, each positioned to indicate the beat location of the music. This allows the input token sequence to be configured in a way that identifies the beat structure of the music. Among the tokens included in the input token sequences illustrated in Figures 6A and 6B, "bar" and "beat" are examples of beat tokens. "bar" is an example of a token indicating a bar line, and "beat" is an example of a token indicating a beat (time signature).
[0061] Each beat token may be positioned in the input token sequence to indicate at least one of the locations of a bar line or a beat in the musical score. A bar line indicates the division of a measure. A measure is a division of appropriate length to make the musical score easier to read. A beat is a unit that divides the temporal continuity of music.
[0062] In one example, each beat token may be positioned to indicate either a bar line or a beat, or only one of the two. This allows the rhythmic structure of the piece to be understood using each beat token as a clue. However, rhythmic structures vary from piece to piece. Some pieces even have changes in time signature midway through. It is difficult to fully understand the rhythmic structure of various types of pieces using only either bar lines or beats. Therefore, it is preferable that each beat token be positioned at both the bar line and beat locations in the input token sequence.
[0063] As illustrated in Figures 8A and 8B, the output token sequence may also include multiple beat tokens, each positioned to indicate the beat positions in the newly generated song (i.e., the song after the difficulty change). The generation model 5 may be configured to output an output token sequence containing beat tokens as a result of generating a new song after the difficulty change. Similar to the input token sequence, in one example, each beat token may be positioned in the output token sequence to indicate at least one of a bar line and a beat. Preferably, each beat token is positioned at the bar line and the beat, respectively, in the output token sequence.
[0064] Here, the representations of the beat tokens ("bar", "beat") in Figures 6A, 6B, 8A, and 8B are examples. The representations of the beat tokens are not limited to these examples and may be determined as appropriate depending on the embodiment. Any symbols such as numbers, letters, or shapes may be used as beat tokens. The symbols and data formats used for beat tokens are not particularly limited as long as they are computer-recognizable and may be selected as appropriate depending on the embodiment.
[0065] Note that the same tokenization scheme and token representation may be used for both the input token sequence and the output token sequence. In the example above, both the input token sequence and the output token sequence may employ either an action-based or note-based tokenization scheme. However, the format of the input token sequence and the output token sequence is not limited to this example. The input token sequence and the output token sequence do not necessarily have to use the same tokenization scheme and token representation. Different tokenization schemes and different token representations may be used for the input token sequence and the output token sequence.
[0066] If the computer can recognize at least a portion of the song whose difficulty level is to be changed, the format of the tokens used in the input token sequence is not particularly limited and may be determined appropriately depending on the embodiment. If the computer can recognize the result of generating a new song after changing the difficulty level, the format of the tokens used in the output token sequence is not particularly limited and may be determined appropriately depending on the embodiment. Furthermore, if the computer can recognize the rhythmic structure, the format of the rhythmic tokens is not particularly limited and may be determined appropriately depending on the embodiment. Difficulty parameters may also be represented by tokens in a similar manner.
[0067] (An example of a generative model) Figure 9 schematically illustrates an example of the configuration of the generation model 5 according to this embodiment. The generation model 5 is composed of a machine learning model having parameters that are adjusted by machine learning. The type of machine learning model is not particularly limited and may be appropriately selected depending on the embodiment. The structure of the machine learning model is not particularly limited and may be appropriately determined depending on the embodiment, as long as it is configured to accept input of the original song to be changed in difficulty and difficulty parameters, and to output the result of generating a new song obtained by changing the difficulty of the original song to the difficulty level specified by the difficulty parameters. For example, the machine learning model may be configured to accept input of the original song in the form of a token sequence and output the result of generating a new song in the form of a token sequence. As an example of the structure of a machine learning model in this case, as shown in Figure 9, generative model 5 may have a configuration based on the Transformer proposed in the reference "Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, 2017." The Transformer is a machine learning model that processes sequential data (such as natural language) and has an attention-based configuration.
[0068] In the example shown in Figure 9, the generative model 5 comprises an encoder 50 and a decoder 55. The encoder 50 has a structure composed of stacking multiple blocks, each having a multi-head attention layer for self-attention and a feed-forward layer. The decoder 55, on the other hand, has a structure composed of stacking multiple blocks, each having a masked multi-head attention layer for self-attention, a multi-head attention layer for source-target attention, and a feed-forward layer. As shown in Figure 9, each layer of the encoder 50 and decoder 55 may be provided with an addition and normalization layer. Each layer may contain one or more nodes, and each node may have a threshold value set. The threshold value may be expressed by an activation function. In addition, weights (connection weights) may be set for the connections between nodes in adjacent layers. The weights and threshold values of the connections between nodes are examples of parameters for the generative model 5.
[0069] In the example shown in Figure 9, the generation model 5 is configured to accept tokens in the input token sequence in order from the beginning. The tokens input to the generation model 5 are each converted into vectors with a predetermined number of dimensions by the input embedding process, and after being assigned a value that identifies their position within the music (within a phrase) by the position encoding process, they are input to the encoder 50. The encoder 50 repeatedly performs processing by the multi-head attention layer and the feedforward layer on the input for the number of blocks to obtain feature representations, and supplies the obtained feature representations to the next stage decoder 55 (multi-head attention layer). The difficulty parameter may be input to the generation model 5 at any time. In one example, the value of the difficulty parameter may be input before the input token sequence.
[0070] The decoder 55 (masked multi-head attention layer) is supplied with input from the encoder 50, as well as known (past) outputs from the decoder 55. That is, the generative model 5 illustrated in Figure 9 is configured to have a recursive structure. The decoder 55 takes the above input and repeatedly performs processing by the masked multi-head attention layer, the multi-head attention layer, and the feedforward layer for the number of blocks to obtain and output a feature representation. The output from the decoder 55 is transformed in the linear layer and the softmax layer and obtained as a token indicating the result of generating a new song.
[0071] The learning processing unit 112 is configured to perform machine learning on the generative model 5 for each learning dataset 3, using the training data 31 as input data and the corresponding ground truth data 32 (true value of the output token sequence) as a training signal. Specifically, for each learning dataset 3, the learning processing unit 112 is configured to input the learning music data 311 (input token sequence) and difficulty parameter 313 contained in the training data 31 into the generative model 5, and train the generative model 5 so that the output token sequence obtained by executing the calculation process of the generative model 5 fits the new learning music data 321 (true value of the output token sequence) of the corresponding ground truth data 32. In other words, for each learning dataset 3, the learning processing unit 112 is configured to adjust the parameter values of the generative model 5 so that the error between the output token sequence generated from the learning music data 311 and difficulty parameter 313 and the true value indicated by the corresponding ground truth data 32 is reduced. Any method, such as backpropagation, may be used to adjust the parameters. Furthermore, multiple normalization techniques (e.g., label smoothing, residual dropout, attention dropout) may be applied to the machine learning processing of generative model 5.
[0072] <Music Generation Device> Figure 10 schematically illustrates an example of the software configuration of the music generation device 2 according to this embodiment. The control unit 21 of the music generation device 2 interprets the instructions contained in the music generation program 82 stored in the storage unit 22 and executes control processing according to the interpreted instructions. Thus, the music generation device 2 according to this embodiment is configured to include a data acquisition unit 211, a parameter acquisition unit 212, a generation unit 213, and an output unit 214 as software modules. That is, in this embodiment, similar to the model generation device 1, the software modules of the music generation device 2 are realized by the control unit 21 (CPU).
[0073] The data acquisition unit 211 is configured to acquire target song data 221 that indicates at least a portion of the songs whose difficulty level is to be changed. If the generation model 5 is configured to handle token sequences, the target song data 221 may be configured to include an input token sequence arranged to indicate at least a portion of a song. In this case, the input token sequence included in the target song data 221 may be obtained in a similar format to the learning song data 311 of the training data 31 exemplified in Figures 6A and 6B. The input token sequence included in the target song data 221 may include a plurality of beat tokens, each positioned to indicate the beat position of the song. In one example, each beat token may be positioned in the input token sequence to indicate at least one of a bar line and a beat. Preferably, each beat token is positioned in the input token sequence at the bar line and the beat, respectively.
[0074] The parameter acquisition unit 212 acquires the value of the difficulty parameter 223 for specifying the difficulty level after the song has been modified. The value of the difficulty parameter 223 may be specified manually or by computer processing. The value of the difficulty parameter 223 may be acquired by any method.
[0075] The generation unit 213 has a trained generative model 5, which holds the learning result data 125. The generation unit 213 is configured to use the trained generative model 5 to generate new song data 225 that represents at least a portion of a new song obtained by changing the difficulty level of a song to the difficulty level specified by the difficulty parameter 223, from the acquired target song data 221 and the value of the difficulty parameter 223.
[0076] In the example shown in Figure 9, the generation unit 213 sequentially inputs the input token sequence and the difficulty parameter 223 values contained in the target song data 221 to the encoder 50 of the trained generative model 5 (specifically, the multi-head attention layer, which is placed first after passing through the input embedding layer), and performs calculations on the encoder 50 and decoder 55. As a result of these calculations, the generation unit 213 sequentially acquires the tokens output from the trained generative model 5 (the softmax layer, which is placed last in the example shown in Figure 9), thereby generating an output token sequence that constitutes new song data 225.
[0077] During this process, the output token sequence may be generated using a search method such as beam search. More specifically, the generation unit 213 may generate the output token sequence by holding n candidate tokens in descending order of score from the probability distribution of the values output from the generation model 5, and selecting candidate tokens such that the combined score of m consecutive tokens is the highest (n and m are integers greater than or equal to 2).
[0078] The generated output token sequence may be obtained in the same format as the new training music data 321 of the ground truth data 32 exemplified in Figures 8A and 8B. The output token sequence included in the new music data 225 may include multiple beat tokens, each positioned to indicate the beat position of the music. In one example, each beat token may be positioned in the output token sequence to indicate at least one of a bar line and a beat. Preferably, each beat token is positioned at the bar line and the beat, respectively, in the output token sequence.
[0079] The output unit 214 is configured to output the newly generated music data 225. The output format of the new music data 225 is not particularly limited and may be determined as appropriate depending on the embodiment. For example, if the new music data 225 is composed of an output token sequence, the output token sequence may be output as is. As another example, the output token sequence may be converted into an appropriate format, and the information obtained from the conversion may be output.
[0080] <Other> The software modules of the model generation device 1 and the music generation device 2 according to this embodiment will be described in detail in the operation examples described later. In this embodiment, an example is described in which the software modules of the model generation device 1 and the music generation device 2 are all implemented by a general-purpose CPU. However, some or all of the above software modules may be implemented by one or more dedicated processors (for example, application-specific integrated circuits (ASICs)). Each of the above modules may also be implemented as a hardware module. Furthermore, regarding the software configuration of the model generation device 1 and the music generation device 2, software modules may be omitted, replaced, and added as appropriate, depending on the embodiment.
[0081] §3 Example of Operation <Model Generator> Figure 11 is a flowchart showing an example of the processing procedure of the model generation apparatus 1 according to this embodiment. The processing procedure of the model generation apparatus 1 described below is an example of a model generation method. However, the processing procedure of the model generation apparatus 1 described below is merely an example, and each step may be modified as much as possible. In addition, steps may be omitted, replaced, or added as appropriate in the following processing procedure, depending on the embodiment.
[0082] (Step S101) In step S101, the control unit 11 acquires multiple training datasets 3, each composed of a combination of training data 31 and ground truth data 32.
[0083] Each training dataset 3 may be generated as appropriate. The training music data 311 of the training data 31 may be generated as appropriate so as to show at least a portion of the music before the difficulty level was changed. The new training music data 321 of the ground truth data 32 may be generated as appropriate so as to show at least a portion of the music after the difficulty level was changed. For example, the training music data 311 may be generated to include the input token sequence shown in Figure 6A or Figure 6B, and the new training music data 321 may be generated to include the true values of the output token sequence shown in Figure 8A or Figure 8B. Each token sequence may be generated from other forms of performance information, such as encoded data or musical scores. Alternatively, each token sequence may be generated directly.
[0084] The difficulty parameter 313 for training the training data 31 may be generated as appropriate to indicate the modified difficulty (i.e., the difficulty of the piece indicated by the corresponding ground truth data 32). In one example, the difficulty parameter 313 may be configured to indicate one of several levels. In another example, the value of the difficulty parameter 313 may be determined according to one or more factors related to difficulty. Specifically, the difficulty parameter 313 may be configured to indicate at least one of the following: the range of notes that can be played in the new piece after the difficulty has been changed, and the maximum number of controls that can be operated simultaneously on the instrument used in the new piece. Each training dataset 3 can be generated by associating the corresponding training data 31 and ground truth data 32 with each other.
[0085] The process of generating the training dataset 3 may be performed on any computer. In one example, the process of generating each training dataset 3 may be performed by the model generation device 1 (control unit 11). In another example, at least a portion of the multiple training datasets 3 may be generated by other computers. In this case, the model generation device 1 (control unit 11) may acquire the training datasets 3 generated by other computers via a network, storage medium 91, etc. The number of training datasets 3 to acquire may be determined appropriately so as to be sufficient for machine learning. Once multiple training datasets 3 have been acquired, the control unit 11 proceeds to the next step S102.
[0086] (Step S102) In step S102, the control unit 11 operates as a learning processing unit 112 and performs machine learning on the generative model 5 using the acquired multiple learning datasets 3.
[0087] As an example of a specific machine learning process, the control unit 11 sequentially inputs the values of the learning song data 311 and difficulty parameter 313 included in the training data 31 of each learning dataset 3 into the generation model 5, repeatedly executes the calculation process of the generation model 5, and sequentially obtains the output tokens. Through this calculation process, the control unit 11 can obtain an output token sequence that shows the generation result of a new song after the difficulty level has been changed, corresponding to the training data 31. Subsequently, the control unit 11 calculates the error between the obtained output token sequence and the true value indicated by the new learning song data 321 included in the corresponding ground truth data 32, and further calculates the gradient of the calculated error. The control unit 11 calculates the error in the parameter values of the generation model 5 by backpropagating the calculated error gradient using the backpropagation method. Based on the calculated error, the control unit 11 adjusts the parameter values of the generation model 5. The control unit 11 may repeat the adjustment of the parameter values of the generation model 5 by the above series of processes until predetermined conditions (for example, executed a specified number of times, or the sum of the calculated errors being less than or equal to a threshold) are met.
[0088] Through this machine learning, the generative model 5 is trained so that, for each training dataset 3, the music data generated from the training music data 311 and the difficulty parameter 313 contained in the training data 31 fits the training music data 321 contained in the ground truth data 32. Therefore, as a result of the machine learning, a trained generative model 5 can be generated that has acquired the ability to generate new music by changing the difficulty level of a given music to the difficulty level specified by the difficulty parameter. Once the machine learning process is complete, the control unit 11 proceeds to the next step S103.
[0089] (Step S103) In step S103, the control unit 11 operates as a storage processing unit 113 and generates learning result data 125 containing information about the trained generative model 5 generated by machine learning. The learning result data 125 contains information for reproducing the trained generative model 5. For example, the learning result data 125 may include information indicating the values of each parameter of the generative model 5 obtained by the adjustment of the machine learning described above. In some cases, the learning result data 125 may include information indicating the structure of the generative model 5. The structure may be specified, for example, by the number of layers, the type of each layer, the number of nodes included in each layer, the connection relationships between nodes in adjacent layers, etc. The control unit 11 stores the generated learning result data 125 in a predetermined memory area.
[0090] The predetermined memory area may be, for example, RAM in the control unit 11, memory unit 12, external storage device, storage media, or a combination thereof. The storage media may be, for example, a CD or DVD, and the control unit 11 may store the learning result data 125 in the storage media via the drive 17. The external storage device may be, for example, a data server such as a NAS. In this case, the control unit 11 may use the communication interface 13 to store the learning result data 125 in the data server via the network. Alternatively, the external storage device may be, for example, an external storage device connected to the model generation device 1.
[0091] Once the saving of the learning result data 125 is complete, the control unit 11 terminates the processing procedure of the model generation device 1 related to this operation example.
[0092] The generated learning result data 125 may be provided to the music generation device 2 at any time. For example, the control unit 11 may transfer the learning result data 125 to the music generation device 2 as part of the processing in step S103 or separately from the processing in step S103. The music generation device 2 may acquire the learning result data 125 by receiving this transfer. Alternatively, for example, the music generation device 2 may acquire the learning result data 125 by accessing the model generation device 1 or data server via a network using the communication interface 23. Alternatively, for example, the music generation device 2 may acquire the learning result data 125 via the storage medium 92. Alternatively, for example, the learning result data 125 may be pre-loaded into the music generation device 2.
[0093] Furthermore, the control unit 11 may update or generate new learning result data 125 by periodically or irregularly repeating the processes described in steps S101 to S103. During this repetition, at least a portion of the multiple training datasets 3 used for machine learning may be modified, corrected, added, deleted, etc., as appropriate. As a result, the control unit 11 may update or generate new trained generative models 5. The control unit 11 may then update the learning result data 125 held by the music generation device 2 by providing the updated or newly generated learning result data 125 to the music generation device 2 in any way.
[0094] <Music Generation Device> Figure 12 is a flowchart showing an example of the processing procedure of the music generation device 2 according to this embodiment. The processing procedure of the music generation device 2 described below is an example of a music generation method. However, the processing procedure of the music generation device 2 described below is merely an example, and each step may be modified as much as possible. In addition, steps may be omitted, replaced, and added as appropriate in the following processing procedure, depending on the embodiment.
[0095] (Step S201) In step S201, the control unit 21 operates as a data acquisition unit 211 and acquires target song data 221 that represents at least a portion of the song.
[0096] The target song data 221 may be generated as appropriate. For example, the target song data 221 may be configured to include an input token sequence arranged to represent at least a portion of the song. In this case, the input token sequence may be generated from other forms of performance information, such as encoded data or musical scores. Alternatively, the input token sequence may be generated directly.
[0097] Furthermore, the target song data 221 may be acquired through any route. In one example, the target song data 221 may be generated in the song generation device 2. In this case, the control unit 21 may acquire the target song data 221 as a result of executing the generation process. In another example, the generation of the target song data 221 may be performed by a computer other than the song generation device 2. In this case, the control unit 21 may acquire the target song data 221 via, for example, a network, a storage medium 92, etc. Once the target song data 221 is acquired, the control unit 21 proceeds to the next step S202.
[0098] (Step S202) In step S202, the control unit 21 operates as a parameter acquisition unit 212 and acquires the value of the difficulty parameter 223.
[0099] The value of difficulty parameter 223 may be obtained by any method. For example, the value of difficulty parameter 223 may be obtained by manual input by an operator, selection from a list of available difficulty levels, random selection, determination according to any rule, etc.
[0100] Similar to the training data 31, in one example, the value of the difficulty parameter 223 may be configured to represent one of several levels. In another example, the value of the difficulty parameter 223 may be determined according to one or more factors related to difficulty. Specifically, the value of the difficulty parameter 223 may be configured to represent at least one of the following: the range of notes that can be played in the new piece after the difficulty has been changed, and the maximum number of controls that can be operated simultaneously on the instrument used in the new piece.
[0101] Once the value of the difficulty parameter 223 is obtained, the control unit 21 proceeds to the next step S203. Note that the timing of executing steps S201 and S202 is not limited to this example. In another example, step S202 may be executed before step S201. In yet another example, steps S201 and S202 may be executed in parallel.
[0102] (Step S203) In step S203, the control unit 21 operates as a generation unit 213 and, referring to the learning result data 125, sets up the generative model 5 that has been trained by machine learning. Using the trained generative model 5, the control unit 21 generates new song data 225 that represents at least a portion of the new song obtained by changing the difficulty level of the song to the difficulty level specified by the difficulty parameter 223, based on the acquired target song data 221 and the value of the difficulty parameter 223.
[0103] In one example, the control unit 21 inputs the input token sequence and the difficulty parameter 223 values contained in the target song data 221 into the trained generative model 5 and executes the arithmetic processing of the trained generative model 5. In the example shown in Figure 9 above, the control unit 21 sequentially inputs the input token sequence and the difficulty parameter 223 values into the trained generative model 5 and repeatedly executes the forward propagation arithmetic processing of the trained generative model 5 to sequentially generate tokens that make up the output token sequence. As a result of this arithmetic processing, the control unit 21 obtains new song data 225 composed of the output token sequence from the trained generative model 5. Once the generation of a new song due to the change in difficulty (arithmetic processing of the trained generative model 5) is complete, the control unit 21 proceeds to the next step S204.
[0104] (Step S204) In step S204, the control unit 21 operates as an output unit 214 and outputs the new music data 225 generated by the processing in step S203. The output destination and output format are not particularly limited and may be appropriately selected depending on the embodiment. In one example, the output destination may be, for example, RAM, storage unit 22, storage medium, external storage device, other computer, other device, etc. In another example, the control unit 21 may output the obtained new music data 225 (for example, an output token sequence) as is. In yet another example, the control unit 21 may convert the new music data 225 into an appropriate format and output the information obtained by the conversion. As a specific example, if the new music data 225 is obtained in the format of an output token sequence, the control unit 21 may convert the obtained output token sequence into a musical score and output the obtained musical score. In this case, the control unit 21 may, for example, output a command to a printing device (not shown) to print the musical score on paper.
[0105] Once the output of the newly generated song data 225 is complete, the control unit 21 terminates the processing procedure of the song generation device 2 according to this example of operation. The control unit 21 may, for example, repeatedly execute the processing of steps S201 to S204 periodically or irregularly in response to a request from an operator. During this repetition, at least a portion of the target song data 221 obtained in step S201 and at least one of the values of the difficulty parameter 223 obtained in step S202 may be changed, modified, added, deleted, etc., as appropriate. This allows the control unit 21 to generate songs with varying difficulty levels in various variations using the trained generation model 5.
[0106] <Features> As described above, in this embodiment, the training data 31 of each training dataset 3 used in the machine learning in step S102 is configured to include a training difficulty parameter 313 that indicates the modified difficulty level. As a result, the generative model 5 is trained to generate a new song (new training song data 321) from the original input song (training song data 311) based on the difficulty level specified by the difficulty parameter 313. The difficulty level of the newly generated song can be controlled in various ways according to the difficulty parameter. Therefore, the machine learning process in step S102 can generate a trained generative model 5 that has acquired the ability to generate songs of various difficulty levels based on the difficulty parameter.
[0107] Furthermore, in step S202, the value of the difficulty parameter 223 is obtained in order to specify the new difficulty level. Then, in step S203, the obtained value of the difficulty parameter 223 is used by the trained generative model 5 to generate a new song. This allows the trained generative model 5 to perform a process that changes the difficulty level of the song indicated by the target song data 221 according to the value of the difficulty parameter 223. As a result, songs of various difficulty levels can be easily generated.
[0108] Furthermore, in this embodiment, the values of the difficulty parameters (223, 313) may be configured to indicate at least one of the range of playable notes in the new song after the difficulty has been changed, and the maximum number of controls that can be operated simultaneously on the instrument used in the new song. This allows the machine learning process in step S102 to generate a trained generative model 5 that has acquired the ability to control at least one of the range of playable notes and the maximum number of controls in the generated new song using the difficulty parameters. In the generation process in step S203, it is possible to easily generate at least one of songs with various ranges of playable notes and songs with various maximum numbers of controls.
[0109] Furthermore, in this embodiment, the generation model 5 may be configured to accept musical input in the form of a token sequence. The input token sequence (target musical data 221 and training musical data 311) may be configured to include multiple beat tokens arranged to indicate the beat positions of the musical piece. This allows the generation model 5 to grasp the beat structure of the musical piece using the beat tokens and then perform the process of generating a new musical piece. Therefore, in the machine learning process of step S102, a trained generation model 5 that is less prone to temporal errors caused by the beat structure can be generated. In the generation process of step S203, the probability of temporal errors occurring in the newly generated musical piece (new musical data 225) can be reduced.
[0110] Furthermore, in this embodiment, the generative model 5 may be configured to output the newly generated song in token format. The output token sequence (new song data 225 and new training song data 321) may be configured to include multiple beat tokens arranged to indicate the beat positions of the song, similar to the input token sequence. This allows for easy identification of the location of the time error, even if a time error occurs during the generation process in step S203, based on the position of the beat tokens included in the output token sequence. As a result, the resulting new song can be easily corrected. The machine learning process in step S102 can generate a trained generative model 5 with such capabilities.
[0111] §4 Variant Although embodiments of the present invention have been described in detail above, the above description is merely illustrative in all respects of the present invention. Needless to say, various improvements or modifications can be made without departing from the scope of the present invention.
[0112] For example, in the above embodiment, a machine learning model having a recursive structure with a Transformer configuration (Figure 9) was given as the generative model 5. However, the recursive structure is not limited to the example shown in Figure 9. A recursive structure is a structure configured to perform processing on the target (current) input by referring to past inputs from the target. As long as such operations are possible, the recursive structure is not particularly limited and may be determined as appropriate depending on the embodiment. In another example, the recursive structure may be composed of known structures such as RNN (Recurrent Neural Network) or LSTM (Long short-term memory).
[0113] In the above embodiment, the generative model 5 is configured to have a recursive structure. However, the configuration of the generative model 5 is not limited to this example. The recursive structure may be omitted. The generative model 5 may be composed of a neural network having a known structure, such as a fully connected neural network or a convolutional neural network. Furthermore, the form in which the input token sequence is input to the generative model 5 is not limited to the example in the above embodiment. In another example, the generative model 5 may be configured to accept multiple tokens included in the input token sequence at once.
[0114] In the above embodiment, the input and output data formats of the generative model 5 are not limited to token sequences, and any data format may be used. The type of machine learning model constituting the generative model 5 is not particularly limited and may be appropriately selected depending on the embodiment, as long as it can accept input of song and difficulty parameters and output a new song after changing the difficulty. In the above embodiment, when the generative model 5 is composed of a machine learning model having multiple layers, the type of each layer may be appropriately selected depending on the embodiment. Each layer may include, for example, a convolutional layer, a pooling layer, a dropout layer, a normalization layer, a fully connected layer, etc. With respect to the structure of the generative model 5, components can be omitted, replaced, and added as appropriate. Furthermore, the generative model 5 may be configured to accept input of information other than the song and difficulty parameters. The generative model 5 may be configured to output further information other than the new song.
[0115] §5 Reference examples To verify the effectiveness of the beat token, we generated trained generative models for the following first and second reference examples, and evaluated the accuracy of the generated trained generative models.
[0116] Specifically, 261,396 samples of original music were prepared, and using the action-based tokenization method shown in Table 1, Figure 6A, and Figure 8A above, input token sequences constituting training data were generated from each sample of music. In addition, 261,396 samples of music with varying difficulty levels corresponding to each of the original songs were prepared. Then, similar to the input token sequences, the action-based tokenization method was used to generate true values of output token sequences constituting the ground truth data from each sample of the difficulty-changed music. By associating the generated input token sequences (training data) with the true values of the output token sequences (ground truth data), a training dataset of 261,396 samples was generated. In the first example, as shown in Figures 6A and 8A, the training dataset was obtained by placing beat tokens at the bar lines and beat positions in the true values of both the input and output token sequences. On the other hand, in the second example, the training dataset was obtained without including beat tokens in the true values of both the input and output token sequences (otherwise the same as in the first example).
[0117] The generative models for the first and second reference examples employ the Transformer structure illustrated in Figure 9. Using the same method as in the above embodiment, machine learning was performed using the prepared training dataset of 261,396 samples to generate the trained generative models for the first and second reference examples.
[0118] In addition to the training data, 1000 samples of music (each sample having a duration of 4 measures) were prepared, and 1000 input token sequences (target music data) were obtained from these prepared songs. Similar to the training dataset, the input token sequence for the first reference example had each beat token placed at the bar line and beat position. On the other hand, the input token sequence for the second reference example did not have beat tokens placed (otherwise it was the same as the first reference example).
[0119] Next, using the trained generative models for the first and second reference examples, we obtained output token sequences representing new song data from the target song data of each sample. Then, we evaluated whether a discrepancy in the number of beats occurred in the difficulty-modified song, as indicated by the output token sequence, compared to the original song (i.e., the song indicated by the target song data). As a result, in the second reference example, a discrepancy in the number of beats occurred with a probability of 17.4%. On the other hand, in the first reference example, a discrepancy in the number of beats occurred with a probability of 4.1%. From these results, it was found that the probability of temporal errors occurring can be significantly reduced by including beat tokens that indicate the beat structure. [Explanation of Symbols]
[0120] 1...Model generation device, 11...Control unit, 12...Storage unit, 13...Communication interface, 14…External interface, 15...Input device, 16...Output device, 17...Drive, 81...Model generation program, 91...Storage medium, 111...Learning data acquisition unit, 112...Learning processing unit, 113...Storage processing unit, 125...Training result data, 2... Music generation device, 21...Control unit, 22...Storage unit, 23...Communication interface, 24…External interface, 25...Input device, 26...Output device, 27...Drive, 82... Music generation program, 92... Storage medium, 211...Data acquisition unit, 212...Parameter acquisition unit, 213...Generation unit, 214...Output unit, 221...Target song data, 223...Difficulty parameter, 225... New song data, 3…Training dataset, 31... Training data, 311...Learning song data, 313...Difficulty parameter, 32... Correct answer data, 321... New learning song data, 5…Generative Model
Claims
1. A data acquisition unit that acquires target song data representing at least a portion of the song, A parameter acquisition unit that acquires a value for a difficulty parameter that relatively specifies a new difficulty level relative to the difficulty level of the aforementioned song, A generation unit that uses a trained generation model to generate new song data representing at least a portion of a new song obtained by changing the difficulty level of the song to the new difficulty level specified by the difficulty parameter, based on the acquired target song data and the value of the difficulty parameter, An output unit that outputs the newly generated music data, Equipped with, Music generation device.
2. The value of the difficulty parameter is configured to indicate the range of the playable notes of the new song after the difficulty level has been changed. The music generation apparatus according to claim 1.
3. The value of the difficulty parameter is configured to indicate the maximum number of controls to be operated simultaneously on the instrument used for the new song after the difficulty has been changed. A music generation device according to claim 1 or 2.
4. The aforementioned target song data consists of a sequence of input tokens arranged to represent at least a portion of the song, The new music data consists of a sequence of output tokens output from the trained generative model and arranged to represent at least a portion of the new music. A music generation device according to any one of claims 1 to 3.
5. The input token sequence includes a plurality of beat tokens, each arranged to indicate the beat position of the music; Each of the aforementioned multiple beat tokens is positioned at the bar line and beat in the input token sequence. The output token sequence includes a plurality of beat tokens, each positioned to indicate the beat position of the new song, Each of the aforementioned plurality of beat tokens is positioned at the bar line and beat in the output token sequence. The music generation apparatus according to claim 4.
6. Computers A step of obtaining target song data that represents at least a portion of the song, The steps include obtaining a value for a difficulty parameter that relatively specifies a new difficulty level relative to the difficulty level of the aforementioned song, The steps include: using a trained generative model to generate new song data that represents at least a portion of a new song obtained by changing the difficulty of the song to the new difficulty specified by the difficulty parameter, from the acquired target song data and the value of the difficulty parameter; The steps include outputting the newly generated music data, Execute Music generation method.
7. On the computer, A step of obtaining target song data that represents at least a portion of the song, The steps include obtaining a value for a difficulty parameter that relatively specifies a new difficulty level relative to the difficulty level of the aforementioned song, The steps include: using a trained generative model to generate new song data that represents at least a portion of a new song obtained by changing the difficulty of the song to the new difficulty specified by the difficulty parameter, from the acquired target song data and the value of the difficulty parameter; The steps include outputting the newly generated music data, To execute A music generation program.
8. A learning data acquisition unit that acquires multiple learning datasets, each composed of a combination of training data and ground truth data, The training data includes training music data representing at least a portion of a musical piece and difficulty parameters for learning, which include difficulty parameters that relatively specify a new difficulty level relative to the difficulty level of the musical piece. The aforementioned correct answer data includes new learning song data that shows at least a portion of the new songs generated by changing the difficulty level of the songs in the learning song data to the new difficulty level specified by the difficulty parameter. The training data acquisition unit, A learning processing unit that performs machine learning of a generative model using the acquired plurality of training datasets, wherein the machine learning is comprised of training the generative model so that, for each of the training datasets, the music data generated by the generative model from the training music data and the difficulty parameter values included in the training data fits the training music data included in the ground truth data, Equipped with, Model generation device.
9. The value of the difficulty parameter included in the training data is configured to indicate the range of the playable notes of the new song after the difficulty level has been changed. The model generation apparatus according to claim 8.
10. The value of the difficulty parameter included in the training data is configured to indicate the maximum number of controls to be operated simultaneously on the instrument used for the new song after the difficulty has been changed. The model generation apparatus according to claim 8 or 9.
11. Computers A step of obtaining multiple training datasets, each composed of a combination of training data and ground truth data, The training data includes training music data representing at least a portion of a musical piece and difficulty parameters for learning, which include difficulty parameters that relatively specify a new difficulty level relative to the difficulty level of the musical piece. The aforementioned correct answer data includes new learning song data that shows at least a portion of the new songs generated by changing the difficulty level of the songs in the learning song data to the new difficulty level specified by the difficulty parameter. Steps and A step of performing machine learning on a generative model using the acquired plurality of training datasets, wherein the machine learning consists of training the generative model for each training dataset so that the music data generated by the generative model from the training music data and the difficulty parameter values included in the training data fits the training music data included in the ground truth data. Execute Model generation method.
12. On the computer, A step of obtaining multiple training datasets, each composed of a combination of training data and ground truth data, The training data includes training music data representing at least a portion of a musical piece and difficulty parameters for learning, which include difficulty parameters that relatively specify a new difficulty level relative to the difficulty level of the musical piece. The aforementioned correct answer data includes new learning song data that shows at least a portion of the new songs generated by changing the difficulty level of the songs in the learning song data to the new difficulty level specified by the difficulty parameter. Steps and A step of performing machine learning on a generative model using the acquired plurality of training datasets, wherein the machine learning consists of training the generative model for each training dataset so that the music data generated by the generative model from the training music data and the difficulty parameter values included in the training data fits the training music data included in the ground truth data. To execute Model generation program.
Citation Information
Patent Citations
Simple musical score creating device and simple musical score creating program
JP2007241026A