Arrangement generation method, arrangement generation device and computer program product
The generative model trained through machine learning solves the problems of high cost and lack of diversity in generating arrangement data, and realizes the automatic generation of a variety of arrangement data.
Patent Information
- Application Number
- CN202180009202.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-17
- Filing Date
- 2021-02-09
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2041-02-09
AI Technical Summary
In the prior art, the cost of generating arrangement data is high and it is difficult to generate a variety of arrangement data. In particular, when the performance information is not suitable for the specified algorithm, the arrangement may deviate from the original song and appropriate arrangement data cannot be generated.
A generative model trained through machine learning is used to generate arrangement data based on the target music data, and metadata is used to control the generation conditions of the arrangement data to achieve automation of the arrangement data.
By generating models through machine learning, the cost of generating arrangement data can be reduced, and a variety of appropriate arrangement data can be generated.
Smart Images

Figure CN115004294B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an arrangement generation method, an arrangement generation device, and a computer program product for generating an arrangement of music using a trained generative model generated by machine learning. Background Art
[0002] Generating a musical score requires a variety of steps. Generally, a musical score is created through the following steps: creating the basic structure of the piece (melody, rhythm, and harmony); creating an arrangement based on the basic structure; laying out the elements, such as notes and instrumentation, corresponding to the created piece (or arrangement) to create the score data; and finally, outputting the score data to a paper medium. Traditionally, these steps have primarily been performed by humans (e.g., manually operating computer software).
[0003] However, if the process of generating the music score is entirely performed manually, the cost of generating the music score becomes high. Therefore, in recent years, the development of technology that automates at least part of the process of generating the music score is being promoted. For example, in Patent Document 1, a method for automatically generating an accompaniment based on an arrangement is proposed.
[0004] According to this technology, a part of the process of generating an arrangement can be automated, thereby reducing the cost of generating the arrangement.
[0005] Prior art literature
[0006] Patent Literature
[0007] Patent Document 1: Japanese Patent Application Laid-Open No. 2017-58594 Summary of the Invention
[0008] Problems to be solved by the invention
[0009] The inventors of the present invention have discovered that the following problems exist in the previous arrangement generation method proposed in Patent Document 1 and the like. That is, in the previous technology, accompaniment data is generated based on the performance information according to a prescribed algorithm. However, the music that serves as the basis for automatic arrangement is diverse, so the prescribed algorithm may not always be suitable for the performance information (music). In the case where the original performance information is not suitable for the prescribed algorithm, it is possible to perform an arrangement that deviates from the original song, and it is impossible to generate appropriate arrangement data. In addition, in the previous method, only consistent arrangement data based on the prescribed algorithm can be generated, and it is difficult to automatically generate a variety of arrangement data. Therefore, in the previous method, it is difficult to appropriately generate a variety of arrangement data.
[0010] In one aspect, the present invention has been made in view of the above circumstances, and an object of the present invention is to provide a technology for reducing the cost of generating arrangement data and appropriately generating a variety of arrangement data.
[0011] Means for solving problems
[0012] To address the aforementioned issues, the present invention employs the following configuration. Specifically, one aspect of the present invention relates to an arrangement generation method in which a computer executes the following steps: acquiring target music data, the target music data including performance information representing the melody and harmony of at least a portion of the music, and meta-information representing characteristics related to at least a portion of the music; generating arrangement data based on the acquired target music data using a generative model trained through machine learning, wherein the arrangement data is obtained by arranging the performance information in accordance with the meta-information; and outputting the generated arrangement data.
[0013] In the above structure, a trained generative model generated by machine learning is used to generate arrangement data based on target music data containing original performance information. By appropriately implementing machine learning using sufficient learning data, the trained generative model can acquire the ability to appropriately generate arrangement data based on a variety of original performance information. Therefore, by using a trained generative model that has acquired such an ability, arrangement data can be appropriately generated. Moreover, in this structure, meta-information is included in the input of the generative model. Based on the meta-information, the generation conditions of the arrangement data can be controlled. Thus, according to this structure, a variety of arrangement data can be generated. Furthermore, according to this structure, the process of generating arrangement data can be automated, thereby reducing the cost of generating arrangement data. Thus, according to the above structure, it is possible to reduce the cost of generating arrangement data and appropriately generate a variety of arrangement data.
[0014] Effects of the Invention
[0015] According to the present invention, it is possible to provide a technique for reducing the cost of generating arrangement data and appropriately generating a variety of arrangement data. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 An example of a scenario to which the present invention is applied is schematically illustrated.
[0017] Figure 2 An example of the hardware configuration of the arrangement generating device according to the embodiment is schematically illustrated.
[0018] Figure 3 An example of the software configuration of the arrangement generating device according to the embodiment is schematically illustrated.
[0019] Figure 4 This is a music score that shows an example of melody and harmony of the performance information according to the embodiment.
[0020] Figure 5 Is based on Figure 4 An example of a music score generated by the melody and harmony shown.
[0021] Figure 6 An example of the structure of the generation model according to the embodiment is schematically illustrated.
[0022] Figure 7 This is a diagram for explaining an example of a token input to a generation model according to an embodiment.
[0023] Figure 8 This is a diagram for explaining an example of a token output from a generation model according to an embodiment.
[0024] Figure 9 This is a flowchart showing an example of a processing procedure of machine learning of a generation model performed by the arrangement generation device according to the embodiment.
[0025] Figure 10 This is a flowchart showing an example of a procedure of a music arrangement data generation process (inference process performed by a generation model) performed by the music arrangement generation device according to the embodiment.
[0026] Figure 11 This is a diagram for explaining an example of a token input to a generation model according to a modification.
[0027] Figure 12 This is a diagram for explaining an example of a token output from a generation model according to a modification.
[0028] Figure 13 Another example of a scenario to which the present invention is applied is schematically illustrated. DETAILED DESCRIPTION
[0029] Hereinafter, an embodiment of one aspect of the present invention (hereinafter also referred to as "this embodiment") is described based on the accompanying drawings. The present embodiment described below is merely an illustration of the present invention. Various improvements or modifications can obviously be made without departing from the scope of the present invention. In the implementation of the present invention, a specific structure corresponding to the embodiment can also be appropriately adopted. In addition, the data appearing in this embodiment is described by natural language, but more specifically, it is specified by pseudo-language, commands, parameters, machine language, etc. that can be recognized by a computer.
[0030] <1. Application Examples>
[0031] Figure 1An example of a scenario in which the present invention is applied is schematically shown. The arrangement generation device 1 according to the present embodiment is a computer configured to generate arrangement data 25 of a musical piece using a trained generation model 5 .
[0032] First, the arrangement generation device 1 according to this embodiment acquires target music data 20. This target music data 20 includes performance information 21 representing the melody and chord of at least a portion of the music, and meta-information 23 representing characteristics related to at least a portion of the music. Next, the arrangement generation device 1 generates arrangement data 25 based on the acquired target music data 20 using a generative model 5 trained through machine learning. The arrangement data 25 is obtained by arranging the performance information 21 according to the meta-information 23. In other words, the meta-information 23 corresponds to the conditions for generating the arrangement. The arrangement generation device 1 then outputs the generated arrangement data 25.
[0033] As described above, in this embodiment, a trained generative model 5 generated through machine learning is used to generate arrangement data 25 based on target music data 20 including original performance information 21. By appropriately implementing machine learning using sufficient learning data, the trained generative model 5 acquires the ability to appropriately generate arrangement data based on a wide variety of original performance information. Therefore, by using the trained generative model 5 that has acquired such an ability, it is possible to appropriately generate arrangement data 25. Furthermore, the conditions for generating arrangement data 25 can be controlled using the meta-information 23. Furthermore, by using the trained generative model 5, at least a portion of the process for generating arrangement data 25 can be automated. Thus, according to this embodiment, it is possible to reduce the cost of generating arrangement data 25 and appropriately generate a wide variety of arrangement data 25.
[0034] <2. Example of structure>
[0035] <2.1 Hardware Structure>
[0036] Figure 2 An example of the hardware structure of the arrangement generating device 1 according to this embodiment is schematically illustrated. Figure 2 As shown, the arrangement generating device 1 according to this embodiment is a computer formed by electrically connecting a control unit 11, a storage unit 12, a communication interface 13, an input device 14, an output device 15, and a driver 16. Figure 2 In the example, the communication interface is referred to as "communication I / F".
[0037] The control unit 11 includes a CPU (Central Processing Unit), RAM (Random Access Memory), and ROM (Read Only Memory), which are examples of hardware processors (processor resources), and is configured to execute information processing based on programs and various data. The storage unit 12 is an example of a memory, and is comprised of, for example, a hard disk drive or solid-state drive. In this embodiment, the storage unit 12 stores various information, such as the generation program 81, the learning data 3, and the learning result data 125.
[0038] The generation program 81 is used to cause the arrangement generation device 1 to perform the information processing (described later) related to the machine learning of the generation model 5 and the generation of the arrangement data 25 using the trained generation model 5. Figure 9 as well as Figure 10 ). The generation program 81 includes a series of instructions for processing this information. The learning data 3 is used for machine learning of the generation model 5. The learning result data 125 represents information related to the trained generation model 5. In this embodiment, the learning result data 125 is generated as a result of executing machine learning processing on the generation model 5. Details are described later.
[0039] The communication interface 13 is an interface for performing wired or wireless communication via a network, such as a wired LAN (Local Area Network) module or a wireless LAN module. The arrangement generating device 1 can perform data communication with other information processing devices via a network using the communication interface 13 .
[0040] The input device 14 is, for example, a device for input, such as a mouse or keyboard. Furthermore, the output device 15 is, for example, a device for output, such as a display or speaker. In one example, the input device 14 and the output device 15 may be separate devices. In another example, the input device 14 and the output device 15 are integrally formed, such as a touch panel display. An operator, such as a user, can operate the arrangement generation device 1 by utilizing the input device 14 and the output device 15.
[0041] The drive 16 is, for example, a CD drive, a DVD drive, etc., and is a drive device for reading various information such as programs stored in the storage medium 91. The storage medium 91 is a medium that stores various information such as stored programs by means of electrical, magnetic, optical, mechanical or chemical effects in a manner that enables a computer or other device or machine to read the stored programs. At least one of the above-mentioned generation program 81 and the learning data 3 may also be stored in the storage medium 91. The arrangement generation device 1 may also obtain at least one of the above-mentioned generation program 81 and the learning data 3 from the storage medium 91. In addition, in Figure 2 In the embodiment, a disk-type storage medium such as a CD or DVD is shown as an example of storage medium 91. However, the type of storage medium 91 is not limited to a disk-type storage medium and may be a storage medium other than a disk-type storage medium. Examples of storage mediums other than a disk-type storage medium include semiconductor memories such as flash memories. The type of drive 16 can be arbitrarily selected according to the type of storage medium 91.
[0042] In addition, regarding the specific hardware structure of the arrangement generation device 1, structural elements can be omitted, replaced, and added as appropriate according to the implementation method. For example, the control unit 11 may also include multiple hardware processors. The type of hardware processor is not limited to the CPU. The hardware processor can be composed of, for example, a microprocessor, an FPGA (field-programmable gate array), a GPU (graphics processing unit), etc. The storage unit 12 can also be composed of the RAM and ROM contained in the control unit 11. At least one of the communication interface 13, the input device 14, the output device 15, and the driver 16 can also be omitted. The arrangement generation device 1 can have an external interface for connecting to an external device. The external interface can be composed of, for example, a USB (Universal Serial Bus) port, a dedicated port, etc. The arrangement generation device 1 can also be composed of multiple computers. In this case, the hardware structure of each computer can be consistent or inconsistent. Furthermore, the arrangement generating device 1 may be a general-purpose server device, a general-purpose PC (Personal Computer), a portable terminal (eg, a smartphone, a tablet PC), or the like, in addition to an information processing device specifically designed for providing services.
[0043] <2.2 Software Structure>
[0044] Figure 3An example of the software structure of the arrangement generation device 1 according to this embodiment is schematically illustrated. The control unit 11 of the arrangement generation device 1 controls each component by interpreting and executing instructions contained in the generation program 81 stored in the storage unit 12 using the CPU. Thus, the arrangement generation device 1 according to this embodiment includes a learning data acquisition unit 111, a learning processing unit 112, a storage processing unit 113, an object data acquisition unit 114, an arrangement generation unit 115, a musical score generation unit 116, and an output unit 117 as software modules. Specifically, in this embodiment, each software module of the arrangement generation device 1 is implemented by the control unit 11 (CPU).
[0045] The learning data acquisition unit 111 is configured to acquire learning data 3. Learning data 3 is composed of a plurality of learning data sets 300. Each learning data set 300 is composed of a combination of training music data 30 and known arrangement data 35. The training music data 30 is music data used as training data for machine learning in the generative model 5. The training music data 30 includes performance information 31 representing the melody and harmony of at least a portion of the music, and meta-information 33 representing characteristics related to at least a portion of the music. The meta-information 33 indicates the conditions for generating the corresponding known arrangement data 35 based on the performance information 31.
[0046] The learning processing unit 112 is configured to perform machine learning on the generative model 5 using the acquired plurality of learning data sets 300. The storage processing unit 113 is configured to generate information related to the trained generative model 5 generated through machine learning as learning result data 125, and to store the generated learning result data 125 in a predetermined storage area. The learning result data 125 can be appropriately configured to include information for reproducing the trained generative model 5.
[0047] The target data acquisition unit 114 is configured to acquire target music data 20, which includes performance information 21 representing the melody and harmony of at least a portion of the music, and meta-information 23 representing characteristics related to at least a portion of the music. The target music data 20 is music data that, when input into the trained generative model 5, becomes the target of arrangement (i.e., the basis for the arrangement). The arrangement generation unit 115 stores the trained generative model 5 through learning result data 125, thereby possessing the trained generative model 5. The arrangement generation unit 115 uses the generative model 5 trained through machine learning to generate arrangement data 25 based on the acquired target music data 20. The arrangement data 25 is obtained by arranging the performance information 21 according to the meta-information 23. The musical score generation unit 116 is configured to generate musical score data 27 using the generated arrangement data 25. The output unit 117 is configured to output the generated arrangement data 25. In this embodiment, outputting the arrangement data 25 may also be configured to output the generated musical score data 27.
[0048] (Various data)
[0049] The performance information (21, 31) can be appropriately configured to represent the melody and chords of at least a portion of a piece of music. At least a portion of the piece of music can be defined by a length such as four measures. As one example, the performance information (21, 31) can be provided directly. In another example, the performance information (21, 31) can be obtained from data in other formats such as musical notation. As a specific example, the performance information (21, 31) can be obtained from various types of raw data representing a performance of a piece of music containing melody and chords. The raw data can be, for example, MIDI data, audio waveform data, etc. In one example, the raw data can also be read from a memory resource of the device, such as the storage unit 12 or the storage medium 91. In another example, the raw data can also be obtained from an external device, such as another smartphone, a music providing server, or a NAS (Network Attached Storage). The raw data can also include data other than melody and harmony. The harmony in the performance information (21, 31) can be determined by performing harmony estimation processing on the raw data. A well-known method can be used for the harmony estimation process.
[0050] The meta-information (23, 33) is preferably configured to indicate a condition for generating an arrangement. In the present embodiment, the meta-information (23, 33) may be configured to include at least one of difficulty information, style information, composition information, and speed information. The difficulty information is configured to indicate the difficulty of playing as a condition for the arrangement. In one example, the difficulty information may be configured to include a value indicating a category of difficulty (e.g., one of "beginner," "beginner-intermediate," "intermediate," "upper-intermediate," and "advanced"). The style information is configured to indicate a style of music to be arranged as a condition for the arrangement. In one example, the style information may be configured to include at least one of arranger information (e.g., arranger ID) for identifying an arranger and artist information (e.g., artist ID) for identifying an artist.
[0051] The composition information is configured to represent the instrument composition of the music as a condition for the arrangement. In one example, the composition information is configured to include values representing the categories of the instruments used in the arrangement. The categories of the instruments can be given, for example, in accordance with the GM (General MIDI) standard. The tempo information is configured to represent the tempo of the music. In one example, the tempo information can be configured to include values representing the tempo range to which the music belongs, among multiple tempo ranges (e.g., BPM = less than 60, greater than 60 and less than 84, greater than 84 and less than 108, greater than 108 and less than 144, greater than 144 and less than 192, and greater than 192).
[0052] In the machine learning scenario, the meta-information 33 can be pre-associated with the corresponding known arrangement data 35. In this case, the meta-information 33 can be obtained from the known arrangement data 35. The meta-information 33 can be obtained by analyzing the corresponding known arrangement data 35. The meta-information 33 can also be obtained by an operator who specifies the performance information 31 (for example, inputs the original data) and inputs it via the input device 14. On the other hand, in the inference processing (arrangement generation) scenario, the meta-information 23 can be appropriately determined to specify the conditions for the arrangement to be generated. In one example, the meta-information 23 is automatically selected by the arrangement generation device 1 or other computer, for example, by random selection or determination according to prescribed rules. In another example, the meta-information 23 can also be obtained by a user who wishes to generate arrangement data and inputs it via the input device 14.
[0053] The arrangement data (25, 35) is configured to include accompaniment sounds (arrangement sounds) corresponding to the melody and harmony of at least a portion of the music. The arrangement data (25, 35) can be obtained, for example, based on a standard MIDI file (SMF) or other format. In the context of machine learning, the known arrangement data 35 can be appropriately obtained based on the performance information 31 and the meta-information 33 so that it can be used as correct answer data. The known arrangement data 35 can be automatically generated based on the performance information 31 according to a predetermined algorithm, or can be generated at least partially manually. The known arrangement data 35 can be generated, for example, based on existing musical score data.
[0054] Figure 4 An example of a music score showing the melody and harmony of the performance information (21, 31) according to this embodiment is shown. Figure 4 As illustrated, the performance information (21, 31) can be configured to include a melody (simple melody) composed of a sequence of single notes (including rests) and harmony (harmony information such as Am, F, etc.) that progresses over time.
[0055] Figure 5 Example representation based on Figure 4 The music score of an example of the arrangement of the melody and harmony shown in FIG. Figure 5 As illustrated, the arrangement data (25, 35) may include a plurality of performance parts (in one example, a right-hand part and a left-hand part of a piano). The arrangement data (25, 35) may be configured to include not only the melody tones constituting the melody included in the performance information (21, 31) but also accompaniment tones (arrangement tones) corresponding to the melody and harmony.
[0056] exist Figure 4 as well as Figure 5In the example, at the beginning of measure 1, the melody included in the performance information (21, 31) is the note A (dotted quarter note), and the harmony is the key of A minor (chord VI of C major, which is the key of this example). Correspondingly, the arrangement data (25, 35) includes not only the melody included in the right hand part, but also the notes A (eighth note on the forward beat) and E (dotted quarter note on the forward beat and eighth note on the reverse beat) as the constituent notes of the key of A minor as accompaniment tones based on the harmony method.
[0057] Furthermore, as shown in the figure, the accompaniment tones included in the arrangement data (25, 35) are not limited to tones obtained by simply extending the tones that constitute the harmony. The arrangement data (25, 35) may include tones that correspond not only to the harmony but also to the pitch and rhythm of the melody (for example, tones composed in counterpoint).
[0058] (Structure example of generation model)
[0059] Figure 6 An example of the structure of the generation model 5 involved in this embodiment is schematically illustrated. The generation model 5 is composed of a machine learning model having parameters adjusted by machine learning. The type of machine learning model is not particularly limited and can be appropriately selected according to the embodiment. In one example, Figure 6 As shown, the generative model 5 may have a structure based on the transformer proposed in the reference “Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, 2017.” The transformer is a machine learning model that processes sequence data (natural language, etc.) and has a structure based on attention.
[0060] exist Figure 6In the example, the generative model 5 includes an encoder 50 and a decoder 55. The encoder 50 has a structure composed of a plurality of blocks as a stack, each of which has a multi-head attention layer (Multi-HeadAttention Layer) seeking self-attention and a feedforward layer (Feed Forward Layer). On the other hand, the decoder 55 has a structure composed of a plurality of blocks as a stack, each of which has a masked multi-head attention layer (MaskedMulti-HeadAttention Layer) seeking self-attention, a multi-head attention layer seeking source / target attention, and a feedforward layer. Figure 6 As shown, each layer of the encoder 50 and decoder 55 can include an addition / normalization layer. Each layer can contain one or more nodes, and each node can be assigned a threshold. The threshold can be expressed using an activation function. In addition, the connections between nodes in adjacent layers can be assigned weights (connection loads). The weights and thresholds of the connections between nodes are examples of parameters of the generative model 5.
[0061] Furthermore, use Figure 7 as well as Figure 8 An example of the input format and output format of the generation model 5 will be described. Figure 7 This is a diagram for explaining an example of an input format (token) of music data to be input to the generation model 5 according to the present embodiment. Figure 8 1 is a diagram for explaining an example of the output format (token) of arrangement data output from the generation model 5 according to the present embodiment. Figure 7 As shown, in the context of machine learning and inference processing, music data (20, 30) is converted into an input token sequence containing a plurality of tokens T. The input token sequence can be appropriately generated in a manner corresponding to the music data (20, 30).
[0062] During the machine learning phase, the learning processing unit 112 is configured to input the tokens included in the input token sequence corresponding to the training music data 30 into the generation model 5, and execute calculations by the generation model 5 to generate an output token sequence corresponding to the arrangement data (inference result). Meanwhile, during the inference phase, the arrangement generation unit 115 is configured to input the tokens included in the input token sequence corresponding to the target music data 20 to the trained generation model 5, and execute calculations by the trained generation model 5 to generate an output token sequence corresponding to the arrangement data 25.
[0063] like Figure 7As shown in the example, each token T included in the input token sequence is an information element representing performance information (21, 31) or meta-information (23, 33). The difficulty token (e.g., level_400) represents the difficulty information (e.g., intermediate piano level) included in the meta-information (23, 33). The style token (e.g., arr_1) represents the style information (e.g., arranger A) included in the meta-information (23, 33). The tempo token (e.g., tempo_72) represents the tempo information (e.g., a tempo range around quarter note = 72) included in the meta-information (23, 33).
[0064] A harmony token (e.g., chord_0root_0) indicates the harmony included in the performance information (21, 31) (e.g., C major with the root note being C). A note-on token (e.g., on_67), a hold token (e.g., wait_4), and a note-off token (e.g., off_67) indicate the notes (e.g., a quarter note with a pitch of G4) that constitute the melody included in the performance information (21, 31). Furthermore, a note-on token indicates the pitch of a new note to be sounded, a note-off token indicates the pitch of a note to be stopped, and a hold token indicates the length of time the sounding (or silent) state is maintained. Thus, a note-on token is used to sound a predetermined note, a hold token is used to maintain the sounding state of the aforementioned note, and a note-off token is used to stop the aforementioned note.
[0065] In this embodiment, the input token sequence is configured such that after the token T corresponding to the meta-information (23, 33) is configured, the token T corresponding to the performance information (21, 31) is configured in time series. Figure 7 In the example, the tokens T corresponding to the various information included in the meta-information (23, 33) are arranged in the order of difficulty token, style token, and speed token in the input token sequence. However, when the meta-information (23, 33) includes multiple types of information, the order of arranging the tokens T corresponding to the various types of information in the meta-information (23, 33) in the input token sequence is not limited to this example and may be appropriately determined according to the embodiment.
[0066] like Figure 6 As shown, the generation model 5 involved in this embodiment is configured to accept the input of the tokens T included in the input token sequence in sequence from the beginning. The tokens T input to the generation model 5 are respectively transformed into vectors with a predetermined dimension through input embedding processing, and are assigned values for determining their positions in the music (inside the phrase) through position encoding processing, and are then input to the encoder 50. The encoder 50 repeatedly performs the processing of the multi-head attention layer and the feedforward layer for this input in an amount corresponding to the number of blocks to obtain feature representation, and supplies the obtained feature representation to the decoder 55 (multi-head attention layer) of the next stage.
[0067] The decoder 55 (masked multi-head attention layer) is supplied with not only the input from the encoder 50 but also the known (past) output from the decoder 55. That is, the generative model 5 involved in this embodiment is configured to have a regression structure. The decoder 55 repeatedly performs the processing of the masked multi-head attention layer, the multi-head attention layer, and the feedforward layer on the above input in an amount corresponding to the number of blocks to obtain feature representation and output it. The output from the decoder 55 is transformed in the linear layer and the SOFTMAX layer and output as a token T after being given information equivalent to the arrangement.
[0068] like Figure 8 As shown in the example, each token T output from the generation model 5 is an information element representing performance information or meta-information, and constitutes arrangement data. The output token sequence corresponding to the arrangement data is formed by the multiple tokens T obtained in sequence from the generation model 5. The token T corresponding to the meta-information is different from the input token sequence ( Figure 7 )Similarly, the description is omitted.
[0069] The tokens T (note sound tokens, note stop tokens) representing the performance information contained in the arrangement data can correspond to the notes of multiple performance parts (the right hand part and the left hand part of the piano). Figure 5 As shown, the multiple tokens T (output token column) output from the generation model 5 can be constructed to represent not only the melody sound constituting the melody indicated by the token T corresponding to the input performance information (21, 31), but also the accompaniment sound (arrangement sound) corresponding to the melody and harmony.
[0070] The output token sequence, like the input token sequence, is structured such that, after the tokens T corresponding to the meta-information are arranged, the tokens T corresponding to the performance information are arranged in a time-sequential manner. The order in which the tokens T corresponding to the various pieces of meta-information are arranged in the output token sequence is not particularly limited and can be determined as appropriate depending on the implementation.
[0071] During the machine learning phase, the learning processing unit 112 performs machine learning on the generative model 5 for each learning dataset 300, using a plurality of tokens T (input token sequence) representing the training music data 30 as training data (input data) and a plurality of tokens T (output token sequence) representing the corresponding arrangement data 35 as correct answer data (teacher signal). Specifically, the learning processing unit 112 is configured to input the input token sequence corresponding to the training music data 30 into the generative model 5 for each learning dataset 300, and train the generative model 5 so that the output token sequence (the inference result of the arrangement data) obtained by executing the calculations of the generative model 5 is appropriate for the corresponding correct answer data (the known arrangement data 35). In other words, the learning processing unit 112 is configured to adjust the parameter values of the generative model 5 for each learning dataset 300 so that the error between the arrangement data represented by the output token sequence generated by the generative model 5 based on the input token sequence corresponding to the training music data 30 and the corresponding known arrangement data 35 is minimized. In the machine learning process of generating model 5, various normalizing methods (such as label smoothing, residual dropout, attention dropout) can be applied.
[0072] In the inference (arrangement generation) stage, the arrangement generation unit 115 sends a plurality of tokens T (input token sequence) representing the target music data 20 of the arrangement to the encoder 50 (in the case of the trained generation model 5) Figure 6 In the example of , after passing through the input embedding layer, it is sequentially input to the initially configured multi-head attention layer) from the beginning, and the operation processing of the encoder 50 is performed. As a result of the operation processing, the arrangement generation unit 115 sequentially obtains the generated model 5 (in the trained Figure 6 In the example, the token T output by the last configured SOFTMAX layer) is used to generate the arrangement data 25 (output token column). In this process, the arrangement data 25 can be generated using a search method such as a beam search. More specifically, the arrangement generation unit 115 can maintain n candidate tokens in descending order of scores based on the probability distribution of the values output from the generation model 5, and select candidate tokens so that the total score of the consecutive m tokens is the highest, thereby generating the arrangement data 25 (n, m are integers greater than 2). This process can also be applied to the process of obtaining inference results in machine learning.
[0073] (other)
[0074] The software modules of the arrangement generation device 1 will be described in detail in the operational examples described below. Furthermore, in this embodiment, the software modules of the arrangement generation device 1 are implemented using a general-purpose CPU. However, some or all of the software modules may be implemented using one or more dedicated processors (e.g., application-specific integrated circuits (ASICs)). Each of the modules may also be implemented as hardware modules. Furthermore, the software structure of the arrangement generation device 1 may be modified to include, replace, or add software modules as appropriate, depending on the implementation.
[0075] <3. Action Example>
[0076] <3.1 Machine Learning Process>
[0077] Figure 9 This is a flowchart illustrating an example of a process related to machine learning for generating model 5, performed by the arrangement generation device 1 according to this embodiment. The process related to machine learning described below is an example of a model generation method. The process of the model generation method described below is merely an example, and each step can be modified within the scope of the present invention. Furthermore, steps in the following process may be omitted, replaced, or added as appropriate, depending on the implementation.
[0078] In step S801, the control unit 11 functions as the learning data acquisition unit 111 to acquire the performance information 31 constituting each learning data set 300. In one example, the performance information 31 may be directly provided. In another example, the performance information 31 may be obtained from other forms of data, such as musical notation. As a specific example, the performance information 31 may be generated by analyzing the melody and harmony of known original data.
[0079] In step S802, the control unit 11 operates as the learning data acquisition unit 111 to acquire meta-information 33 corresponding to each piece of performance information 31. The meta-information 33 can be appropriately configured to represent characteristics related to the arranged music. In this embodiment, the meta-information 33 can be configured to include at least one of difficulty information, style information, composition information, and tempo information. The meta-information 33 can also be obtained by an operator specifying the performance information 31 (for example, by inputting raw data) via the input device 14. Through the processing of steps S801 and S802, the training music data 30 of each learning data set 300 can be acquired.
[0080] In step S803, the control unit 11 acts as a learning data acquisition unit 111 to acquire known arrangement data 35 corresponding to each piece of training music data 30. The known arrangement data 35 can be appropriately generated so that it can be used as correct answer data. That is, the known arrangement data 35 can be appropriately generated based on the conditions indicated by the corresponding meta-information 33 so as to represent the music obtained by arranging the music indicated by the corresponding performance information 31. In one example, the known arrangement data 35 can be generated corresponding to the known original data used to obtain the performance information 31. The above-mentioned meta-information 33 can also be acquired from the corresponding known arrangement data 35. The acquired known arrangement data 35 can be appropriately associated with the corresponding training music data 30. Through the processing of steps S801 to S803, a plurality of learning data sets 300 can be acquired.
[0081] In step S804, the control unit 11 operates as the learning processing unit 112 and converts the training music data 30 (performance information 31 and meta-information 33) of each learning data set 300 into a plurality of tokens T. Thus, the control unit 11 generates an input token sequence corresponding to the training music data 30 of each learning data set 300. As described above, in this embodiment, the input token sequence is configured such that, after the tokens T corresponding to the meta-information 33 are arranged, the tokens T corresponding to the performance information 31 are arranged in time series.
[0082] Furthermore, as long as the processing of steps S801 and S802 is performed before step S804, the order of the processing of steps S801 through S804 is not limited to the above example and can be determined appropriately depending on the implementation. In another example, the processing of step S802 can be performed before step S801. Alternatively, the processing of steps S801 and S802 can be performed in parallel. In another example, the processing of step S804 can be performed in correspondence with each of steps S801 and S802. That is, the control unit 11 can generate a token T for a portion of the performance information 31 in response to obtaining the performance information 31, and can obtain a token T for a portion of the meta-information 33 in response to obtaining the meta-information 33. In another example, the processing of step S804 can be performed before at least one of steps S801 through S803. In another example, the processing of steps S803 and S804 can also be performed in parallel.
[0083] In addition, at least a part of the processing of steps S801 to S804 can be performed by other computers. In this case, the control unit 11 can also obtain the calculation results from other computers via the network, storage medium 91, other external storage devices (such as NAS, external storage media, etc.), etc., so as to achieve at least a part of the processing of steps S801 to S804. In one example, each learning data set 300 can be generated by another computer. In this case, the control unit 11 can also obtain each learning data set 300 from another computer as the processing of steps S801 to S803. It is also possible that at least a part of the multiple learning data sets 300 are generated by other computers, and the rest are generated by the arrangement generation device 1.
[0084] In step S805, the control unit 11 functions as the learning processing unit 112, performing machine learning on the generative model 5 using multiple learning data sets 300 (learning data 3). In this embodiment, the control unit 11 sequentially inputs the tokens T included in the input token sequence obtained through the processing of step S804 into the generative model 5 for each learning data set 300 as a forward propagation operation, and repeatedly performs the operations of the generative model 5, thereby sequentially generating the tokens T that constitute the output token sequence. Through this operation, the control unit 11 can obtain the arrangement data (output token sequence) corresponding to each piece of training music data 30 as an inference result. Next, the control unit 11 calculates the error between the obtained arrangement data and the corresponding known arrangement data 35 (correct answer data), and further calculates the gradient of the calculated error. The control unit 11 backpropagates the calculated error gradient using the error backpropagation method to calculate the error in the parameter values of the generative model 5. Based on the calculated error, the control unit 11 adjusts the parameter values of the generative model 5. The control unit 11 may repeatedly adjust the parameter values of the generation model 5 through the above series of processes until a predetermined condition is satisfied (for example, the process is performed a predetermined number of times and the sum of the calculated errors becomes equal to or smaller than a threshold value).
[0085] Through this machine learning, the generative model 5 is trained for each learning dataset 300 so that the arrangement data generated from the training music data 30 is adapted to the corresponding known arrangement data 35. Thus, as a result of the machine learning, a trained generative model 5 can be generated that has learned the correspondence between the input token sequence (training music data 30) and the output token sequence (known arrangement data 35) provided by each learning dataset 300. In other words, a trained generative model 5 can be generated that has acquired the ability to arrange the melody and harmony of the performance information 31 (original) according to the conditions indicated by the meta-information 33 so that it is adapted to the known arrangement data 35 (correct answer data).
[0086] In step S806, the control unit 11 acts as the storage processing unit 113 and generates information related to the trained generative model 5 generated by machine learning as learning result data 125. The learning result data 125 stores information for reproducing the trained generative model 5. As an example, the learning result data 125 may include information indicating the values of the various parameters of the generative model 5 obtained through the adjustment of the above-mentioned machine learning. Depending on the circumstances, the learning result data 125 may include information indicating the structure of the generative model 5. The structure can be determined, for example, based on the number of layers, the type of each layer, the number of nodes included in each layer, the connection relationship between the nodes of adjacent layers, etc. The control unit 11 saves the generated learning result data 125 to a predetermined storage area.
[0087] The predetermined storage area may be, for example, the RAM within the control unit 11, the storage unit 12, an external storage device, a storage medium, or a combination thereof. The storage medium may be, for example, a CD or DVD, and the control unit 11 may store the learning result data 125 on the storage medium via the drive 16. The external storage device may be, for example, a data server such as a NAS. In this case, the control unit 11 may also use the communication interface 13 to store the learning result data 125 on the data server via a network. Furthermore, the external storage device may be, for example, an external storage device connected to the arrangement generation device 1.
[0088] If the saving of the learning result data 125 is completed, the control unit 11 ends the processing of the machine learning of the generation model 5 involved in this action example. In addition, the control unit 11 can also repeat the processing of the above steps S801 to S806 regularly or irregularly to update or newly generate the learning result data 125. During this repetition, at least a part of the learning data 3 used for machine learning can be appropriately changed, corrected, added, deleted, etc. In this way, the control unit 11 can also update or newly generate the trained generation model 5. In addition, if there is no need to save the results of machine learning, the processing of step S806 can be omitted.
[0089] <3.2 Processing of Arrangement Generation>
[0090] Figure 10 This is a flowchart illustrating an example of a process related to arrangement generation performed by the arrangement generation device 1 according to this embodiment. The process related to arrangement generation described below is an example of an arrangement generation method. However, steps in the process described below may be omitted, replaced, or added as appropriate depending on the embodiment.
[0091] In step S901, the control unit 11 operates as the target data acquisition unit 114 to acquire performance information 21 representing the melody and harmony of at least a portion of the musical composition. In one example, the performance information 21 may be directly provided. In another example, the performance information 21 may be obtained from other forms of data, such as musical notation. As a specific example, the performance information 21 may be obtained by analyzing raw data serving as the target of the arrangement.
[0092] In step S902, the control unit 11 functions as the object data acquisition unit 114 to acquire meta-information 23 representing characteristics related to at least a portion of the music. In this embodiment, the meta-information 23 may include at least one of difficulty information, style information, composition information, and tempo information. In one example, the meta-information 23 may be automatically selected by the arrangement generation device 1 or another computer, for example, randomly or according to a predetermined rule. In another example, the meta-information 23 may be obtained by a user inputting it via the input device 14. In this case, the user can specify desired arrangement conditions. Through the processing of steps S901 and S902, the control unit 11 can acquire the target music data 20 including the performance information 21 and the meta-information 23.
[0093] In step S903, the control unit 11 operates as the arrangement generation unit 115, converting the performance information 21 and meta-information 23 included in the target music data 20 into a plurality of tokens T. Thus, the control unit 11 generates an input token sequence corresponding to the arranged target music data 20. As described above, in this embodiment, the input token sequence is configured such that, after the tokens T corresponding to the meta-information 23 are arranged, the tokens T corresponding to the performance information 21 are arranged in time sequence.
[0094] In addition, as long as the processing of step S901 and step S902 is performed before step S903, the order of the processing of step S901 to step S903 may not be limited to the above example, and may be appropriately determined according to the implementation method. In another example, the processing of step S902 may be performed before step S901. Alternatively, the processing of step S901 and step S902 may be performed in parallel. In another example, the processing of step S903 may be performed corresponding to step S901 and step S902, respectively. That is, the control unit 11 may also generate a token T of a portion of the performance information 21 corresponding to the acquisition of the performance information 21, and generate a token T of a portion of the meta-information 23 corresponding to the acquisition of the meta-information 23.
[0095] In step S904, the control unit 11 acts as the arrangement generation unit 115, and refers to the learning result data 125 to set the generation model 5 trained by machine learning. In the case where the setting of the trained generation model 5 has been completed, this process can be omitted. The control unit 11 uses the generation model 5 trained by machine learning to generate the arrangement data 25 based on the acquired object music data 20. In this embodiment, the control unit 11 inputs the token T contained in the generated input token column to the trained generation model 5, and performs the operation of the trained generation model 5, thereby generating an output token column corresponding to the arrangement data 25. Furthermore, in this embodiment, the trained generation model 5 is constructed to have a regression structure. In the above-mentioned step of generating the output token column, the control unit 11 inputs the token T contained in the input token column to the trained generation model 5 in sequence from the beginning, and repeatedly performs the operation of the trained generation model 5 (the above-mentioned forward propagation operation), thereby sequentially generating the tokens constituting the output token column.
[0096] As a result of this calculation, arrangement data 25 can be generated by arranging the performance information 21 according to the meta-information 23. That is, even if the performance information 21 is the same, different arrangement data 25 can be generated by changing the meta-information 23. If the meta-information 23 includes difficulty information, in step S904, the control unit 11 can use the trained generation model 5 to generate arrangement data 25 corresponding to the difficulty level indicated by the difficulty information from the target music data 20. If the meta-information 23 includes style information, in step S904, the control unit 11 can use the trained generation model 5 to generate arrangement data 25 corresponding to the style (arranger, artist) indicated by the style information from the target music data 20. If the meta-information 23 includes composition information, in step S904, the control unit 11 can use the trained generation model 5 to generate arrangement data 25 corresponding to the instrumental composition indicated by the composition information from the target music data 20. When the meta-information 23 includes tempo information, in step S904 , the control unit 11 can generate arrangement data 25 corresponding to the tempo indicated by the tempo information from the target music data 20 using the trained generation model 5 .
[0097] In step S905, the control unit 11 operates as the score generating unit 116 and generates the score data 27 using the generated arrangement data 25. In one example, the control unit 11 generates the score data 27 by laying out elements such as notes and instrumentation marks using the arrangement data 25.
[0098] In step S906, the control unit 11 operates as the output unit 117 and outputs the generated arrangement data 25. The output destination and output format are not particularly limited and can be determined appropriately according to the implementation method. In one example, the control unit 11 outputs the arrangement data 25 as is to an output destination such as RAM, storage unit 12, storage medium, external storage device, or other information processing device. In another example, the output arrangement data 25 can also be configured to output musical score data 27. In this case, the control unit 11 can also output the musical score data 27 to an output destination such as RAM, storage unit 12, storage medium, external storage device, or other information processing device. In addition to this, the control unit 11 can also output a command to a printing device (not shown) to print the musical score data 27 onto a medium such as paper. In this way, a printed musical score can also be output.
[0099] Once the output of the arrangement data 25 is complete, the control unit 11 terminates the arrangement generation process involved in this example operation. Furthermore, the control unit 11 may also periodically or irregularly repeat the above-described processes of steps S901 to S906, for example, in response to a user request. During this repetition, at least a portion of the performance information 21 and meta-information 23 input to the trained generation model 5 may be appropriately modified, corrected, added to, or deleted. This allows the control unit 11 to generate different arrangement data 25 using the trained generation model 5.
[0100] Features
[0101] As described above, in this embodiment, in step S904, the trained generative model 5 generated through machine learning is used to generate arrangement data 25 based on the target musical piece data 20 including the original performance information 21. By appropriately performing machine learning in step S805 using sufficient learning data 3, the trained generative model 5 acquires the ability to appropriately generate arrangement data based on a wide variety of original performance information. Therefore, in step S904, by using the trained generative model 5 that has acquired this ability, the arrangement data 25 can be appropriately generated. Furthermore, the conditions for generating the arrangement data 25 can be controlled using the meta-information 23, enabling the generation of a wide variety of arrangement data 25 based on the same performance information 21. Furthermore, by using the trained generative model 5, at least a portion of the process for generating the arrangement data 25 can be automated, thereby reducing manual labor. Consequently, according to this embodiment, the cost of generating the arrangement data 25 can be reduced while also enabling the appropriate generation of a wide variety of arrangement data 25.
[0102] Furthermore, in this embodiment, the musical score data 27 can be automatically generated based on the generated arrangement data 25 through step S905. Furthermore, the musical score data 27 can be automatically output to various media (e.g., storage media, paper media, etc.) through step S906. Thus, according to this embodiment, the generation and output of musical scores can be automated, further reducing manual work hours.
[0103] Furthermore, in this embodiment, the meta-information (23, 33) can be configured to include at least one of difficulty information, style information, composition information, and tempo information. Thus, in step S904, a variety of arrangement data 25 can be generated that is suitable for at least one of the difficulty, style, instrument composition, and tempo indicated by the meta-information 23. Thus, according to this embodiment, the cost of generating multiple variations (arrangement patterns) of the arrangement data 25 based on the same performance information 21 can be reduced. Similarly, the performance information (21, 23) includes not only melody information but also chord information. Therefore, according to this embodiment, the harmony in the generated arrangement data 25 can also be controlled.
[0104] Furthermore, in this embodiment, the music data (20, 30) is converted into an input token sequence, which is configured such that, after the token T corresponding to the meta-information (23, 33) is arranged, the token T corresponding to the performance information (21, 31) is arranged in a time series. Furthermore, the generation model 5 is configured to have a regression structure, and the tokens T included in the input token sequence are input to the generation model 5 sequentially from the beginning. Thus, in the generation model 5, the calculation results corresponding to the portion preceding the target of the meta-information (23, 33) and the performance information (21, 31) can be reflected in the calculation corresponding to the portion of the target of the performance information (21, 31). Thus, according to this embodiment, the context of the meta-information and performance information can be appropriately reflected in the inference processing, so that the generation model 5 can generate appropriate arrangement data. In the machine learning stage, a trained generation model 5 capable of generating such appropriate arrangement data can be generated. In the stage of arrangement generation, in step S905 , by using the trained generation model 5 that has acquired such capabilities, appropriate arrangement data 25 can be generated.
[0105] <4. Modifications>
[0106] The embodiments of the present invention have been described above in detail. The descriptions thus far are merely illustrative of the present invention. It is apparent that various modifications or variations are possible without departing from the scope of the present invention. For example, the following modifications are possible. Furthermore, in the following, the same reference numerals are used for the same structural elements as in the above-described embodiments, and descriptions of aspects common to the above-described embodiments are omitted as appropriate. The following modifications may be combined as appropriate.
[0107] <4.1>
[0108] In the above example, the generative model 5 is configured to generate arrangement data for the right and left hands of a piano based on the single melody and harmony included in the performance information. However, the arrangement is not limited to this example. In the above embodiment, the meta-information (23, 33) can also be configured to include composition information, and the instrument configuration indicated by the composition information can be appropriately controlled (e.g., by user specification) to generate arrangement data for any desired portion in the generative model 5. Examples of instrument configurations include an orchestra configuration including vocals / guitar / bass / drums / keyboards, a chorus configuration including soprano / alto / tenor / bass, and a wind ensemble configuration including multiple woodwind instruments / multiple brass instruments / double bass / percussion instruments. With this configuration, in step S904, arrangement data 25 for portions having different instrument configurations can be generated based on the same performance information 21. In the above machine learning stage, a trained generative model 5 with such capabilities can be generated.
[0109] use Figure 11 as well as Figure 12 , an example of the input form and output form of the generation model 5 involved in this modification example is described. Figure 11 This is a diagram for explaining an example of the input format (token) of music data to the generation model 5 according to this modification. Figure 12 This is a diagram for explaining an example of the output format (token) of arrangement data output from the generation model 5 according to this modification.
[0110] like Figure 11 As shown in the example, the input token sequence involved in this modification includes the above Figure 7 The token T exemplified in FIG, and the instrument composition token (eg <inst> elg bas apf< / inst> The instrument composition token includes: a plurality of instrument specific tokens each representing an instrument (for example, elg for guitar, bas for bass guitar, apf for piano), a start tag token ( <inst> ), and the end tag token (< / inst> ).
[0111] Therefore, if Figure 12 As illustrated, the generation model 5 can determine the instrument composition based on the instrument composition token and generate arrangement data (output token sequence) corresponding to the determined instrument composition. Figure 12 In the example of , the output token sequence output from the generation model 5 includes tokens T representing sounds (performance information) corresponding to a plurality of musical instruments (eg, guitar, bass guitar, piano) respectively identified by the musical instrument composition tokens.
[0112] <4.2>
[0113] Furthermore, in the above embodiment, the information included in the performance information (21, 31) is not limited to information indicating the melody and harmony included in the music. The performance information (21, 31) may also include information other than melody and harmony.
[0114] As an example, Figure 11 As illustrated, the performance information (21, 31) may include not only melody and harmony information but also beat information indicating the rhythm of at least a portion of the music. Figure 11 In the example, the input token list contains beat tokens representing beat information (e.g., Figure 11 According to this structure, in the above step S904, the arrangement data 25 that more appropriately reflects the structure (rhythm) of the music can be generated. In the above machine learning stage, a trained generation model 5 that has acquired such capabilities can be generated.
[0115] <4.3>
[0116] In steps S901 and S902 involved in the above-mentioned embodiment, the arrangement generation device 1 (control unit 11) can also obtain a plurality of object music data 20 corresponding to a plurality of parts obtained by dividing a music piece (for example, dividing it into a predetermined length such as every 4 bars). Correspondingly, the control unit 11 can also perform the steps of generating arrangement data 25 for each of the obtained plurality of object music data 20 (steps S903 and S904), thereby generating a plurality of arrangement data 25. In addition, the control unit 11 can also act as the arrangement generation unit 115 to integrate the generated plurality of arrangement data 25 to generate arrangement data corresponding to a music piece. According to this structure, the amount of calculation of the generation model 5 executed once can be suppressed, and the data size of the reference object of the attention layer can also be suppressed. As a result, the computational load in the generation process can be reduced, and arrangement data can be generated for the entire music piece.
[0117] <4.4>
[0118] Furthermore, in the above embodiment, the arrangement generation device 1 is configured to perform both machine learning and arrangement generation (inference) operations. However, the structure of the arrangement generation device 1 is not limited to this example. If the arrangement generation device 1 is composed of multiple computers, the operations for each step can be decentralized by executing each step on at least one of the multiple computers. Data can be exchanged between the computers via a network, storage media, external storage devices, etc. In one example, the machine learning and arrangement generation processes can be performed by separate computers.
[0119] Figure 13 This schematically illustrates another example of a scenario in which the present invention is applied. The model generation device 101 is one or more computers configured to generate a trained generative model 5 by performing machine learning. The arrangement generation device 102 is one or more computers configured to use the trained generative model 5 to generate arrangement data 25 based on target music data 20.
[0120] The hardware structure of the model generation device 101 and the arrangement generation device 102 can be the same as that of the arrangement generation device 1 described above. As a specific example, the model generation device 101 can be a general-purpose server device, and the arrangement generation device 1 can be, for example, a general-purpose PC, tablet PC, smartphone, or other user terminal. The model generation device 101 and the arrangement generation device 102 can be connected directly or via a network. In the case where the model generation device 101 and the arrangement generation device 102 are connected via a network, the type of network is not particularly limited, and for example, it can be appropriately selected from the Internet, a wireless communication network, a mobile communication network, a telephone network, a dedicated network, etc. The method for exchanging data between the model generation device 101 and the arrangement generation device 102 is not limited to this example and can be appropriately selected according to the implementation method. For example, data can be exchanged between the model generation device 101 and the arrangement generation device 102 using a storage medium.
[0121] In this variation, the generation program 81 can be divided into a first program containing instructions for information processing related to machine learning of the generation model 5, and a second program containing instructions for information processing related to generating arrangement data 25 using the trained generation model 5. In this case, the first program can be referred to as the model generation program, and the second program can be referred to as the arrangement generation program. The arrangement generation program is an example of the generation program of the present invention.
[0122] The model generation device 101 operates as a computer having a learning data acquisition unit 111, a learning processing unit 112, and a storage processing unit 113 as software modules by executing the portion (first program) related to the machine learning processing of the generation program 81. Meanwhile, the arrangement generation device 102 operates as a computer having an object data acquisition unit 114, an arrangement generation unit 115, a musical score generation unit 116, and an output unit 117 as software modules by executing the portion (second program) related to the arrangement generation processing of the generation program 81.
[0123] In this variant, the model generation device 101 generates a trained generation model 5 by executing the above-mentioned processing of steps S801 to S806. The generated trained generation model 5 is generated. The generated trained generation model 5 can be provided to the arrangement generation device 102 at any timing. The generated trained generation model 5 (learning result data 125) can be provided to the arrangement generation device 102 via a network, a storage medium, an external storage device, etc. Alternatively, the generated trained generation model 5 (learning result data 125) can be pre-incorporated into the arrangement generation device 102. On the other hand, the arrangement generation device 102 generates arrangement data 25 based on the target music data 20 using the trained generation model 5 by executing the above-mentioned processing of steps S901 to S906.
[0124] <4.5>
[0125] In the above embodiment, the generation model 5 has Figure 6 The regression structure based on the transformer structure shown in FIG. However, the regression structure is not limited to Figure 6 The example shown. The regression structure refers to a structure that is configured to perform processing corresponding to the input of the object (current) with reference to the input closer to the past than the object. As long as such an operation can be performed, the regression structure is not particularly limited and can be appropriately determined according to the implementation method. In another example, the regression structure can be composed of well-known structures such as RNN (Recurrent Neural Network) and LSTM (Long Short-Term Memory).
[0126] Furthermore, in the above embodiment, the generation model 5 is configured to have a regression structure. However, the structure of the generation model 5 is not limited to this example. The regression structure can be omitted. The generation model 5 can be configured, for example, by a neural network having a well-known structure such as a fully connected neural network or a convolutional neural network. Furthermore, the method of inputting the input token sequence to the generation model 5 is not limited to the example in the above embodiment. In another example, the generation model 5 can also be configured to accept multiple tokens T contained in the input token sequence at a time.
[0127] Furthermore, in the above embodiment, the generation model 5 is configured to receive an input token sequence corresponding to music data and output an output token sequence corresponding to arrangement data. However, the input and output formats of the generation model 5 are not limited to this example. In another example, the generation model 5 can be configured to directly acquire music data. Furthermore, the generation model 5 can also be configured to directly output arrangement data.
[0128] Furthermore, in the above-described embodiment, as long as arrangement data can be generated from music data, the type of machine learning model that constitutes the generative model 5 is not particularly limited and can be appropriately selected depending on the embodiment. Furthermore, in the above-described embodiment, when the generative model 5 is composed of multiple layers, the type of each layer can be appropriately selected depending on the embodiment. For example, each layer can be a convolutional layer, a pooling layer, a dropout layer, a normalization layer, a fully connected layer, etc. Regarding the structure of the generative model 5, structural elements can be omitted, replaced, or added as appropriate.
[0129] <4.6>
[0130] In the above embodiment, the generation of the musical score data 27 can be omitted. Accordingly, the musical score generation unit 116 can be omitted in the software structure of the arrangement generation device 1. In the above arrangement generation related processing, the processing of step S905 can be omitted.
[0131] Label Description
[0132] 1...arrangement generation device, 11...control unit, 12...storage unit, 111...learning data acquisition unit, 112...learning processing unit, 113...save processing unit, 114...object data acquisition unit, 115...arrangement generation unit, 116...music score generation unit, 117...output unit, 5...generation model.
Claims
1. A method for generating an arrangement, comprising executing the following steps by a computer: acquiring target music data, the target music data including performance information indicating melody and harmony of at least a portion of the music, and meta-information indicating characteristics related to at least a portion of the music; generating arrangement data based on the acquired target music data using a generative model trained by machine learning, wherein the arrangement data is obtained by arranging the performance information based on the meta-information; and outputting the generated arrangement data, The meta information includes difficulty information indicating the difficulty of playing the music as a condition for arranging the music. In the step of generating the arrangement data, the computer generates the arrangement data corresponding to the difficulty level indicated by the difficulty level information based on the acquired target music data using the trained generation model.
2. The arrangement generating method according to claim 1, The meta information includes style information indicating the musical style of the music as a condition for arrangement, In the step of generating the arrangement data, the computer generates the arrangement data corresponding to the style indicated by the style information based on the acquired target music data using the trained generation model.
3. The arrangement generation method according to claim 2, The style information includes arranger information for identifying an arranger.
4. The arrangement generation method according to claim 1, The meta information includes composition information indicating the composition of the musical instruments in the music as a condition for arrangement. In the step of generating the arrangement data, the computer generates the arrangement data corresponding to the instrument composition indicated by the composition information based on the acquired target music data using the trained generation model.
5. The arrangement generating method according to claim 1, The performance information includes beat information indicating a rhythm in at least a portion of the music.
6. The arrangement generating method according to claim 1, The step of generating the arrangement data comprises the following steps: The computer generates an input token sequence corresponding to the target music data; and The computer inputs the tokens included in the generated input token sequence into the trained generation model and performs operations on the trained generation model, thereby generating an output token sequence corresponding to the arrangement data.
7. The arrangement generating method according to claim 6, The input token sequence is configured such that, after arranging the tokens corresponding to the meta information, the tokens corresponding to the performance information are arranged in time sequence. The trained generative model is constructed to have a regression structure. In the step of generating the output token column, the computer inputs the tokens contained in the input token column into the trained generation model in sequence from the beginning, and repeatedly performs operations of the trained generation model, thereby generating tokens constituting the output token column in sequence.
8. The arrangement generating method according to claim 1, In the acquisition step, the computer acquires a plurality of target music data corresponding to each of a plurality of parts obtained by dividing a music piece. The computer generates a plurality of arrangement data by executing the step of generating the arrangement data for each of the plurality of target music data obtained, thereby generating the plurality of arrangement data. The computer integrates the generated plurality of arrangement data to generate arrangement data corresponding to the one piece of music.
9. The arrangement generating method according to claim 1, The computer further executes the step of generating musical score data using the generated arrangement data.
10. A music arrangement generating device comprising: a target data acquisition unit configured to acquire target music data including performance information indicating melody and harmony of at least a portion of the music and meta-information indicating characteristics related to at least a portion of the music; an arrangement generating unit configured to generate arrangement data based on the acquired target music data using a generation model trained by machine learning, wherein the arrangement data is obtained by arranging the performance information based on the meta-information; and an output unit configured to output the generated arrangement data, The meta information includes difficulty information indicating the difficulty of playing the music as a condition for arranging the music. The arrangement generation unit generates the arrangement data corresponding to the difficulty level indicated by the difficulty level information based on the acquired target music data using the trained generation model.
11. The arrangement generating device according to claim 10, The arrangement generating device further includes: a music score generating unit configured to generate music score data using the generated arrangement data; Outputting the arrangement data is configured to output the generated musical score data.
12. A computer program product storing a generating program for causing a computer to execute the following steps: acquiring target music data, the target music data including performance information indicating melody and harmony of at least a portion of the music, and meta-information indicating characteristics related to at least a portion of the music; generating arrangement data based on the acquired target music data using a generative model trained by machine learning, wherein the arrangement data is obtained by arranging the performance information based on the meta-information; and outputting the generated arrangement data, The meta information includes difficulty information indicating the difficulty of playing the music as a condition for arranging the music. In the step of generating the arrangement data, the computer generates the arrangement data corresponding to the difficulty level indicated by the difficulty level information based on the acquired target music data using the trained generation model.
Citation Information
Patent Citations
Automatic arrangement device and program
JP2017058594A
Automatic rhythm creation device
JP1992157499A
Automatic music arranging device
JP1994274171A