Music inference device, music inference method, and model generation device
Patent Information
- Application Number
- CN202211440226.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-11-24
- Filing Date
- 2022-11-17
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2042-11-17
AI Technical Summary
然而,若通过手工完成所有针对乐曲的推断作业,则导致该作业所花费的成本变高
Smart Images

Figure CN116168666B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a music inference device, a music inference method, a model generation device, a model generation method, and a storage medium. Background Technology
[0002] For example, in the past, inferences about music, such as generating the arranged music, generating the score, and estimating the music's attributes, were mainly performed manually. However, performing all inferences manually increases the cost of this process. Therefore, efforts are underway to develop methods that use computer technology to automate at least part of the inference process for music.
[0003] For example, Patent Document 1 proposes a technique for automatically generating backing data based on musical arrangements. Furthermore, in recent years, methods for automating inference tasks for musical pieces have utilized AI (artificial intelligence) technology. For instance, Non-Patent Document 1 proposes a method for automatically generating musical pieces using a model trained through machine learning. Based on these techniques, it is possible to reduce the cost of inference tasks for musical pieces.
[0004] Existing technical documents
[0005] Patent documents
[0006] Patent Document 1: Japanese Patent Application Publication No. 2017-58594
[0007] Non-patent literature
[0008] Non-Patent Document 1: Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Noam Shazeer, Ian Simon, Curtis Hawthorne, Andrew M. Dai, Matthew D. Hoffman, Monica Dinculescu, Douglas Eck, "Music Transformer", [online], [accessed September 24, 2008], Internet <URL: https: / / arxiv.org / abs / 1809.04281>
[0009] Non-Patent Document 2: Sageev Oore, Ian Simon, Sander Dieleman, Douglas Eck, Karen Simonyan, "This Time with Feeling: Learning Expressive Musical Performance", [online], [accessed September 24, 2008], Internet <URL: https: / / arxiv.org / abs / 1808.03715>
[0010] Non-Patent Document 3: Yu-Siang Huang, Yi-Hsuan Yang, "Pop Music Transformer: Beat-based Modeling and Generation of Expressive Pop Piano Compositions", [online], [Searched on September 24, 2002], Internet <URL: https: / / arxiv.org / abs / 2002.00212> Summary of the Invention
[0011] The inventors of this invention discovered the following problem in a music inference method using existing AI technology. Specifically, in methods using existing AI technology, for example, information representing the music, such as musical notes, is tokenized, and the resulting token list is input into a trained model to perform computational processing (e.g., non-patent documents 2 and 3). Through this computational processing, an output token list representing the inference result is obtained from the trained model. However, timing errors sometimes occur in the obtained inference result. For example, consider a scenario where an accompaniment is automatically generated from a piece of music using a trained model. In this case, an error may occur where the performance time of the generated accompaniment is inconsistent with the performance time of the original piece of music. When such timing errors occur, it is difficult to determine the location and cause of the error, thus making it difficult to correct the obtained inference result (for example, correcting the duration of the obtained accompaniment data in the above case). Furthermore, the scenario of timing errors is not limited to the case of automatically generated accompaniment; the same problem can occur in all scenarios where inference processing for a piece of music is performed using a trained model.
[0012] In one aspect, the present invention was made in view of such a situation, and its object is to provide a technique for reducing the probability of time-related errors in inferences about musical pieces.
[0013] To address the aforementioned issues, the present invention employs the following structure.
[0014] That is, a music inference apparatus according to one aspect of the present invention includes: a data acquisition unit that acquires object data, the object data including an input token column arranged in a manner representing at least a portion of a music piece, the input token column including a plurality of beat tokens arranged in a manner representing the position of the beats of the music piece; an inference unit that generates an output token column representing the result of an inference for the music piece from the input token column included in the object data using a trained inference model; and an output unit that outputs the result of the inference. Alternatively, the output token column may also be configured to include beat tokens.
[0015] In the music inference device mentioned above, the plurality of beat tokens may also be respectively configured in the bar lines and beat positions of the music in the input token column.
[0016] In the music inference device involved in the above aspect, the input token column may also be generated corresponding to at least a portion of the note column of the music, and the output token column may also be generated in such a way as to show at least a portion of the note column of the arranged music as a result of the inference for the music.
[0017] In the music inference apparatus of the aforementioned aspect, the input token column may also be generated corresponding to the note column of at least a portion of the music, and the output token column may also be generated in a manner that shows the result of inferring local attributes of at least a portion of the music as a result of inference for the music.
[0018] In the music inference device involved in the above aspect, the input token column may also be generated corresponding to at least a portion of the note column of the music, and the output token column may also be generated in such a way as to show at least a portion of the score of the music as a result of the inference for the music.
[0019] In the music inference device involved in the above aspect, the input token column may also be generated corresponding to the element column of at least a portion of the music, and the output token column may also be generated in such a way as to show the note column of at least a portion of the arranged music as a result of the inference for the music.
[0020] The invention is not limited to a music inference apparatus configured using a trained inference model. One aspect of the invention is also a model generation apparatus configured to generate the trained inference model used in any of the above-described forms.
[0021] For example, a model generation apparatus according to one aspect of the present invention includes: a learning data acquisition unit that acquires a plurality of learning datasets each composed of a combination of training data and correct answer labels, wherein the training data includes an input token column arranged in a manner that represents at least a portion of a piece of music for learning, the input token column includes a plurality of beat tokens arranged in a manner that represents the position of the beats of the music, and the correct answer labels are configured to represent the truth value of an output token column corresponding to the inference result for the music; and a learning processing unit that performs machine learning of an inference model using the acquired plurality of learning datasets, wherein the machine learning is configured by training the inference model on each of the learning datasets, so that the output token column generated by the inference model from the input token column included in the training data is adapted to the truth value represented by the correct answer labels.
[0022] In the model generation apparatus described in the above aspect, the plurality of beat tokens can also be respectively configured in the bar lines and beat positions of the music in the input token column.
[0023] In the model generation apparatus of the aforementioned aspect, the input token column contained in the training data of each of the aforementioned learning datasets may also be generated corresponding to at least a portion of the note column of the aforementioned musical piece, and the output token column of the correct answer label of each of the aforementioned learning datasets may also be configured to show the truth value of at least a portion of the note column of the arranged musical piece as the truth value of the inference result for the aforementioned musical piece.
[0024] In the model generation apparatus of the aforementioned aspect, the input token column contained in the training data of each of the aforementioned learning datasets may also be generated corresponding to the note column of at least a portion of the aforementioned musical piece, and the output token column of the correct answer label of each of the aforementioned learning datasets may also be configured to show the truth value of the result of inferring local attributes in at least a portion of the aforementioned musical piece as the truth value of the result of inference for the aforementioned musical piece.
[0025] In the model generation apparatus of the aforementioned aspect, the input token column contained in the training data of each of the aforementioned learning datasets may also be generated corresponding to at least a portion of the note column of the aforementioned musical piece, and the output token column of the correct answer label of each of the aforementioned learning datasets may also be configured to show the truth value of at least a portion of the score of the aforementioned musical piece as the truth value of the result of the inference for the aforementioned musical piece.
[0026] In the model generation apparatus of the aforementioned aspect, the input token column contained in the training data of each of the aforementioned learning datasets may also be generated corresponding to the element column of at least a portion of the aforementioned musical piece, and the output token column of the correct answer label of each of the aforementioned learning datasets may also be configured to show the truth value of the note column of at least a portion of the arranged musical piece as the truth value of the inference result for the aforementioned musical piece.
[0027] As other forms of the music deduction device and model generation device involved in the above-described forms, one aspect of the present invention may also be an information processing method that implements all or part of the above structures, an information processing system, a program, or a storage medium that can be read by a device, machine, or other means storing such a program. Here, a storage medium that can be read by a computer or other means is a medium that stores information such as programs using electrical, magnetic, optical, mechanical, or chemical actions.
[0028] For example, one aspect of the present invention relates to a music inference method, which is an information processing method in which a computer performs the following steps: obtaining object data, the object data including an input token column arranged in a manner that represents at least a portion of a piece of music, and the input token column including a plurality of beat tokens configured in a manner that represents the position of the beats of the music; generating an output token column representing the result of the inference for the music from the input token column included in the object data by using a trained inference model; and outputting the result of the inference.
[0029] Furthermore, for example, one aspect of the present invention relates to a music inference program that is a program for causing a computer to perform the following steps: obtaining object data, the object data comprising an input token column arranged in a manner representing at least a portion of a music piece, and the input token column comprising a plurality of beat tokens configured in a manner representing the position of the beats of the music piece; generating an output token column representing the result of an inference for the music piece from the input token column contained in the object data using a trained inference model; and outputting the result of the inference.
[0030] Furthermore, for example, one aspect of the model generation method according to the present invention is an information processing method in which a computer performs the following steps: obtaining a plurality of learning datasets consisting of combinations of training data and positive solution labels, wherein the training data includes an input token column arranged in a manner that represents at least a portion of a piece of music for learning, the input token column includes a plurality of beat tokens configured in a manner that represents the position of the beats of the music, and the positive solution labels are configured to represent the truth values of an output token column corresponding to the inference result for the music; and performing machine learning of an inference model using the obtained plurality of learning datasets, wherein the machine learning is configured by training the inference model on each of the learning datasets such that the output token column generated by the inference model from the input token column included in the training data is adapted to the truth values represented by the positive solution labels.
[0031] Furthermore, for example, one aspect of the present invention relates to a model generation program that enables a computer to perform the following steps: obtaining multiple learning datasets, each consisting of a combination of training data and positive response labels, wherein the training data includes an input token column arranged to represent at least a portion of a piece of music for learning, the input token column including multiple beat tokens configured to represent the positions of the beats in the music, and the positive response labels being configured to represent the truth values of an output token column corresponding to the inference result for the music; and performing machine learning on an inference model using the obtained multiple learning datasets, wherein the machine learning is configured by training the inference model on each of the learning datasets such that the output token column generated by the inference model from the input token columns included in the training data is adapted to the truth values represented by the positive response labels.
[0032] Invention Effects
[0033] According to the present invention, a technique is provided for reducing the probability of time-related errors in inferences about musical pieces. Attached Figure Description
[0034] Figure 1 This is an illustrative example of a scenario in which the present invention is applied.
[0035] Figure 2 This is an example of the hardware structure of the model generation apparatus involved in the implementation method.
[0036] Figure 3 This is an example of the hardware structure of the music inference device involved in the implementation.
[0037] Figure 4 This is an example of the software structure of the model generation apparatus involved in the implementation method.
[0038] Figure 5 This is a musical score representing one example of a musical piece.
[0039] Figure 6A Indicates from Figure 5 An example of an input token list generated from a piece of music.
[0040] Figure 6B Indicates from Figure 5 An example of an input token list generated from a piece of music.
[0041] Figure 7 It is a musical score representing an example of an arranged musical piece (inferred result).
[0042] Figure 8A Indicates and Figure 7 An example of the truth value of the output token column corresponding to the music.
[0043] Figure 8B Indicates and Figure 7 An example of the truth value of the output token column corresponding to the music.
[0044] Figure 9 This schematically illustrates an example of the structure of the inference model involved in the implementation method.
[0045] Figure 10 This schematically illustrates an example of the software structure of the music deduction device involved in the implementation.
[0046] Figure 11 This is a flowchart illustrating an example of the processing procedure of the model generation apparatus involved in the implementation method.
[0047] Figure 12 This is a flowchart illustrating an example of the processing procedure of the music deduction device involved in the implementation.
[0048] Explanation of reference numerals in the attached figures
[0049] 1...Model generation device; 11...Control unit; 12...Storage unit; 13...Communication interface; 14...External interface; 15...Input device; 16...Output device; 17...Driver; 81...Model generation program; 91...Storage medium; 111...Learning data acquisition unit; 112...Learning processing unit; 113...Storage processing unit; 125...Learning result data; 2...Music inference device; 21...Control unit; 22...Storage unit; 23...Communication interface; 24...External interface; 25...Input device; 26...Output device; 27...Driver; 82...Music inference program; 92...Storage medium; 211...Data acquisition unit; 212...Inference unit; 213...Output unit; 221...Object data; 3...Learning dataset; 31...Training data; 32...Positive solution label; 5...Inference model. Detailed Implementation
[0050] Hereinafter, an embodiment of one aspect of the present invention (hereinafter also referred to as "this embodiment") will be described based on the accompanying drawings. However, the embodiments described below are merely illustrative examples of the present invention. Various modifications and variations can obviously be made without departing from the scope of the present invention. In other words, specific structures corresponding to the embodiments may be appropriately adopted in the implementation of the present invention. Furthermore, the data appearing in this embodiment is described using natural language, but more specifically, it is specified using pseudo-language, commands, parameters, machine language, etc., that can be recognized by a computer.
[0051] §1 Application Examples
[0052] Figure 1 An example of a scenario in which the present invention is applied is illustrated. For example... Figure 1 As shown, the inference system 100 according to this embodiment includes a model generation device 1 and a music inference device 2.
[0053] The model generation apparatus 1 according to this embodiment is a computer configured to generate a trained inference model 5 for performing an inference task for a piece of music through machine learning. First, the model generation apparatus 1 acquires multiple learning datasets 3. Each learning dataset 3 consists of a combination of training data 31 and correct answer labels 32. The training data 31 is configured to include an input token column arranged to represent at least a portion of the music to be learned. The input token column includes multiple beat tokens arranged to represent the positions of the beats in the music. The correct answer labels 32 are configured to represent the truth values of the output token column corresponding to the inference result for the music.
[0054] Next, the model generation device 1 uses the acquired multiple learning datasets 3 to perform machine learning on the inference model 5. Machine learning is configured by training the inference model 5 on each learning dataset 3, so that the output token column generated by the inference model 5 from the input token column contained in the training data 31 is suitable for the true value represented by the corresponding positive solution label 32. Through this machine learning process, a trained inference model 5 capable of performing inference tasks for musical pieces can be generated.
[0055] On the other hand, the music inference device 2 according to this embodiment is a computer configured to perform a music inference task using a trained inference model 5. First, the music inference device 2 acquires object data 221, which includes an input token column arranged to represent at least a portion of the music. The input token column is configured to include multiple beat tokens arranged to represent the positions of the beats in the music. Next, the music inference device 2 generates inference result data from the input token column included in the object data 221 using the trained inference model 5. This inference result data includes an output token column representing the result of the inference for the music. The music inference device 2 outputs the obtained inference result.
[0056] The input token columns for training data 31 and object data 221 can also be appropriately obtained according to the implementation method. For example, the music can be obtained through encoded data (MIDI, etc.), sheet music, or other forms of performance information. For example, the input token columns for training data 31 and object data 221 can be generated from the obtained performance information through transformation processing such as natural language processing. The transformation processing can also be performed by a computer other than the respective devices (1, 2). Furthermore, the transformation processing can be performed at any time. Each device (1, 2) can directly obtain the input token column, or it can obtain other forms of performance information and generate the input token column from the obtained performance information.
[0057] The inference task that enables inference model 5 to complete can also include all inferences for at least a portion of the music. The input token column and the output token column can also be appropriately constructed according to the inference task.
[0058] As an example, the inference task could also be to generate the note sequence of the arranged music from the note sequence of the music. A note sequence is a list of notes that constitute the music. Arrangement could also be, for example, a change in the difficulty of the music, a reduction in scale (a transformation from a multi-instrument note sequence to a single-instrument note sequence, etc.). In this case, the input token sequences contained in the training data 31 and the object data 221 could also be configured to correspond to at least a portion of the note sequence of the music. The output token sequence in the inference result data could also be generated in such a way that it shows at least a portion of the note sequence of the arranged music as the result of the inference for the music. The output token sequence of the correct answer label 32 is configured to show the truth value of at least a portion of the note sequence of the arranged music, corresponding to the associated training data 31, as the truth value of the inference result.
[0059] As another example, the inference task can also be to infer local attributes of a piece of music from its note sequence. Local attributes could be, for example, chords, pitch, rhythm, and the timing of their changes. In this case, the input token sequences contained in the training data 31 and the object data 221 can also be configured to correspond to at least a portion of the note sequence of the music. The output token sequence in the inference result data can also be generated in a manner that shows the result of inferring local attributes in at least a portion of the music as a result of the inference for the music. The output token sequence of the correct answer label 32 can also be configured to show the truth value of the result of inferring local attributes in at least a portion of the music, corresponding to the associated training data 31, as the truth value of its inference result.
[0060] As another example, the inference task could also be to generate a musical score from the notes of a piece of music. In this case, the input token columns contained in the training data 31 and the object data 221 could also be configured to correspond to at least a portion of the notes of the music. The output token columns in the inference result data could also be generated in such a way that they represent at least a portion of the musical score as a result of the inference for the music. The output token column of the correct answer label 32 could also be configured to represent the truth value of at least a portion of the musical score of the music as a truth value of its inference result, corresponding to the associated training data 31.
[0061] As another example, the inference task could also be to generate a sequence of notes for the arranged music from a sequence of elements. A sequence of elements is a list of musical elements (materials) that constitute the music. Elements could be, for example, melody, chords, rhythm, etc. In this case, the input token sequences contained in training data 31 and object data 221 could also be configured to correspond to at least a portion of the sequence of elements of the music. The output token sequence in the inference result data could also be generated in such a way that it represents at least a portion of the sequence of notes for the arranged music as a result of the inference for the music. The output token sequence of the correct answer label 32 could also be configured to represent the truth value of at least a portion of the sequence of notes for the arranged music, corresponding to the associated training data 31, as the truth value of its inference result.
[0062] As another example, the inference task could also be to generate a note sequence from an element sequence representing a theme in a piece of music. The generated note sequence could also be a melody or an arranged piece of music. In this case, the input token sequences contained in the training data 31 and the object data 221 could also be configured to correspond to at least a portion of the element sequence of the music. The output token sequence in the inference result data could also be generated in such a way that it represents at least a portion of the note sequence of the music as a result of the inference for the music. The output token sequence of the correct answer label 32 could also be configured to represent the truth value of at least a portion of the note sequence of the music, corresponding to the associated training data 31, as the truth value of its inference result.
[0063] Multiple beat tokens are appropriately configured in the input token column to indicate the beat structure of the music. Specifically, each beat token is configured in the input token column to indicate the position of at least one of the bar lines and beats of the music. A bar line represents a section of a measure. A measure is a section of appropriate length that is divided for easy reading of the score. A beat is a unit that divides the duration of music. In one example, each beat token can also be configured to indicate only either the bar line or the beat. Thus, the beat structure of the music can be grasped by using each beat token as a clue. However, the beat structure varies from piece to piece. There are also pieces where the beat changes midway. It is difficult to fully grasp the beat structure of various types of music by only indicating either the bar line or the beat. Therefore, it is preferable that each beat token is configured in the input token column of both training data 31 and object data 221 at the respective positions of the bar line and the beat.
[0064] The tokens in the input and output token columns can also be appropriately constructed using symbols such as numbers, letters, and graphics. Similarly, the tokens for each beat can also be appropriately constructed using symbols such as numbers, letters, and graphics. If the computer can distinguish them, the symbols and data format used for the tokens are not particularly limited and can be appropriately selected according to the implementation method. (The following will be discussed...) Figure 6A , Figure 6B , Figure 8A as well as Figure 8B This represents an example of each token.
[0065] In existing methods, the token column input into the trained model does not contain information representing the beat structure of the music. Therefore, while it is possible to infer the beat structure of music with a certain level of accuracy, it is difficult to properly infer the beat structure of various types of music, such as music with altered beats or music with a beat structure different from the training data. This is presumed to be one of the main reasons for the aforementioned timing errors.
[0066] In contrast, in this embodiment, the input token column used for inference is configured as described above to include multiple beat tokens representing the positions of the beats in the music. Therefore, the inference model 5 can complete the inference process for the music after determining its beat structure. Thus, the model generation device 1 can generate a trained inference model 5 that is less prone to timing errors caused by the beat structure. In the music inference device 2, the trained inference model 5 is used to perform the inference task with respect to object data 221 containing multiple beat tokens. Therefore, the probability of timing errors can be reduced in the music inference task.
[0067] In one example, model generation device 1 can generate a trained inference model 5 that is less prone to timing errors caused by beat construction. This inference model 5 is a trained inference model 5 that has acquired the ability to perform inference processes such as generating the note sequence of the arranged piece from the note sequence of the piece being inferred, inferring local attributes of the piece from the note sequence of the piece being inferred, generating a musical score from the note sequence of the piece being inferred, and generating the note sequence of the arranged piece from the element sequence of the piece being inferred. In the music inference device 2, when these inference processes are performed, the probability of timing errors can be reduced.
[0068] In addition, Figure 1 In this example, the model generation device 1 and the music inference device 2 are interconnected via a network. The type of network can be appropriately selected from, for example, the Internet, wireless communication networks, mobile communication networks, telephone networks, private networks, etc. However, the method for exchanging data between the model generation device 1 and the music inference device 2 is not limited to this example and can be appropriately selected according to the implementation method. For example, data can also be exchanged between the model generation device 1 and the music inference device 2 using a storage medium.
[0069] In addition, Figure 1In the example described, the model generation device 1 and the music inference device 2 are each composed of a separate computer. However, the structure of the inference system 100 according to this embodiment is not limited to such an example and can be appropriately determined according to the embodiment. For example, the model generation device 1 and the music inference device 2 can also be an integrated computer. Furthermore, for example, at least one of the model generation device 1 and the music inference device 2 can be composed of multiple computers. In the case of multiple computers, the allocation of information processing can also be appropriately determined according to the embodiment.
[0070] §2 Structural Examples
[0071] [Hardware Structure]
[0072] <Model Generation Device>
[0073] Figure 2 An example of the hardware structure of the model generation apparatus 1 according to this embodiment is illustrated schematically. For example... Figure 2 As shown, the model generation apparatus 1 according to this embodiment is a computer that electrically connects a control unit 11, a storage unit 12, a communication interface 13, an external interface 14, an input device 15, an output device 16, and a driver 17. Furthermore, Figure 2 In this context, the communication interface and the external interface are recorded as "Communication I / F" and "External I / F".
[0074] The control unit 11 includes, for example, a CPU (Central Processing Unit), RAM (Random Access Memory), and ROM (Read Only Memory), and is configured to perform information processing based on programs and various data. The storage unit 12 is an example of a memory, such as a hard disk drive or a solid-state drive. In this embodiment, the storage unit 12 stores various information, including the model generation program 81, multiple learning datasets 3, and learning result data 125.
[0075] Model generation program 81 is used to enable model generation device 1 to perform machine learning information processing to generate the trained inference model 5 (described later). Figure 11 The model generation program 81 contains a series of commands for this information processing. Multiple learning datasets 3 are used to generate the trained inference model 5. The learning result data 125 represents information related to the generated trained inference model 5. In this embodiment, the learning result data 125 is generated as a result of executing the model generation program 81. Details will be described later.
[0076] Communication interface 13, such as a wired LAN (Local Area Network) module or a wireless LAN module, is used for wired or wireless communication via a network. The model generation device 1 can utilize communication interface 13 to perform data communication via a network with other information processing devices. External interface 14, such as a USB (Universal Serial Bus) port or a dedicated port, is used for connecting to external devices. The type and number of external interfaces 14 can be arbitrarily selected.
[0077] The model generation device 1 can also be connected to a device for obtaining each training dataset 3 via at least one of the communication interface 13 and the external interface 14. For example, the input token column for the training data 31 can also be generated from performance information obtained from an electronic musical instrument. When the model generation device 1 performs the generation of the input token column based on this performance information, the model generation device 1 can also be connected to the electronic musical instrument via at least one of the communication interface 13 and the external interface 14, and can also collect performance information for generating the training data 31 through the electronic musical instrument.
[0078] Input device 15 is, for example, a mouse, keyboard, or other device for input. Output device 16 is, for example, a display, speaker, or other device for output. Users or other operators can operate the model generation device 1 using input device 15 and output device 16.
[0079] The drive 17 is, for example, a CD drive, a DVD drive, etc., and is a drive device for reading various information such as programs stored in the storage medium 91. The storage medium 91 is a medium that stores various information such as programs using electrical, magnetic, optical, mechanical, or chemical actions in a manner that allows devices or machines such as computers to read the stored information. The aforementioned model generation program 81 and at least one of the multiple learning datasets 3 can also be stored in the storage medium 91. The model generation device 1 can also obtain the aforementioned model generation program 81 and at least one of the multiple learning datasets 3 from the storage medium 91. Furthermore, in Figure 2 In this example, storage medium 91 is exemplified by disc-type storage media such as CDs and DVDs. However, the type of storage medium 91 is not limited to disc type and can be other types besides disc type. Examples of storage media other than disc type include semiconductor memory such as flash memory. The type of drive 17 can also be arbitrarily selected depending on the type of storage medium 91.
[0080] Furthermore, the specific hardware structure of the model generation apparatus 1 can be appropriately omitted, replaced, or added according to the implementation method. For example, the control unit 11 may include multiple hardware processors. The hardware processors may be composed of microprocessors, FPGAs (Field-Programmable Gate Arrays), etc. The storage unit 12 may be composed of RAM and ROM included in the control unit 11. At least one of the communication interface 13, external interface 14, input device 15, output device 16, and driver 17 may be omitted. The model generation apparatus 1 may also be composed of multiple computers. In this case, the hardware structure of each computer may be the same or different. In addition, the model generation apparatus 1 may be a general-purpose server device, a PC (Personal Computer), etc., in addition to information processing devices specifically designed for the services provided.
[0081] <Music Inference Device>
[0082] Figure 3 An example of the hardware structure of the music deduction device 2 according to this embodiment is illustrated schematically. For example... Figure 3 As shown, the music deduction device 2 involved in this embodiment is a computer that electrically connects the control unit 21, the storage unit 22, the communication interface 23, the external interface 24, the input device 25, the output device 26, and the driver 27.
[0083] The control unit 21, driver 27, and storage medium 92 of the music inference device 2 can be configured identically to the control unit 11, driver 17, and storage medium 91 of the model generation device 1 described above. The control unit 21 includes, for example, a CPU, RAM, ROM, etc., which are hardware processors, and is configured to perform various information processing based on programs and data. The storage unit 22 is configured, for example, a hard disk drive, a solid-state drive, etc. In this embodiment, the storage unit 22 stores various information such as the music inference program 82 and the learning result data 125.
[0084] The music inference procedure 82 is used to enable the music inference device 2 to perform the inference task for the music using the trained inference model 5, as described later. Figure 12 The music inference program 82 contains a series of commands for information processing. At least one of the music inference program 82 and the learning result data 125 can also be stored in the storage medium 92. Furthermore, the music inference device 2 can also retrieve at least one of the music inference program 82 and the learning result data 125 from the storage medium 92.
[0085] The music inference device 2 can also be connected to a device for obtaining object data 221 via at least one of the communication interface 23 and the external interface 24. For example, the input token string for object data 221 can also be generated from performance information obtained from an electronic musical instrument. When generating the input token string based on this performance information is performed in the music inference device 2, the music inference device 2 can also be connected to the electronic musical instrument via at least one of the communication interface 23 and the external interface 24. Furthermore, the music inference device 2 can also accept operations and inputs from users or other operators through the use of the input device 25 and the output device 26.
[0086] Furthermore, the specific hardware structure of the music deduction device 2 can be appropriately omitted, replaced, or added depending on the implementation method. For example, the control unit 21 may include multiple hardware processors. The hardware processors may also be composed of microprocessors, FPGAs, etc. The storage unit 22 may also be composed of RAM and ROM included in the control unit 21. At least one of the communication interface 23, external interface 24, input device 25, output device 26, and driver 27 may also be omitted. The music deduction device 2 may also be composed of multiple computers. In this case, the hardware structure of each computer may be the same or different. In addition, the music deduction device 2 may be a general-purpose server device, a general-purpose PC, etc., in addition to an information processing device specifically designed for the service provided.
[0087] [Software Structure]
[0088] <Model Generation Device>
[0089] Figure 4 An example of the software structure of the model generation apparatus 1 according to this embodiment is illustrated schematically. The control unit 11 of the model generation apparatus 1 interprets the commands contained in the model generation program 81 stored in the storage unit 12 and executes control processing corresponding to the interpreted commands. Thus, the model generation apparatus 1 according to this embodiment is configured to have a learning data acquisition unit 111, a learning processing unit 112, and a storage processing unit 113 as software modules. That is, in this embodiment, the software modules of the model generation apparatus 1 are implemented by the control unit 11 (CPU).
[0090] The learning data acquisition unit 111 is configured to acquire multiple learning datasets 3. Each learning dataset 3 consists of a combination of training data 31 and correct answer labels 32. The training data 31 contains an input token column arranged in a manner representing at least a portion of the music to be learned. The input token column contains multiple beat tokens configured in a manner representing the position of the beats in the music. At least a portion of the music can also be defined by a predetermined length, such as 4 measures. The correct answer labels 32 are configured to represent the truth values of the output token column corresponding to the inference result for the music.
[0091] The learning processing unit 112 is configured to implement machine learning of the inference model 5 using multiple acquired learning datasets 3. Machine learning is configured by training the inference model 5 on each learning dataset 3, such that the output token column generated by the inference model 5 from the input token column contained in the training data 31 is suitable for the true value represented by the corresponding positive solution label 32. This machine learning process ends, thereby generating a trained inference model 5 that has acquired the ability to perform the desired inference task.
[0092] The storage processing unit 113 is configured to generate information related to the trained inference model 5 generated through machine learning as learning result data 125, and to store the generated learning result data 125 in a designated storage area. The learning result data 125 may also be appropriately configured to include information for reproducing the trained inference model 5.
[0093] (An example of a token)
[0094] The tokens constituting the input and output token strings can appropriately use any symbols such as numbers, letters, or graphics. The symbols used for the tokens (token representations) and the data format are not particularly limited and can be appropriately selected depending on the implementation method, provided the computer can recognize them. The same applies to beat tokens. Below, as examples of tokenization methods, two tokenization methods are illustrated: action-based and note-based.
[0095] Figure 5 It is a musical score that represents at least a part of a musical piece. Figure 6A This indicates that tokenization is performed using an action-based method. Figure 5 An example of an input token list generated from a piece of music. Figure 6B This indicates that the tokenization method is used to... Figure 5 An example of an input token list generated from a piece of music. Figure 7 This is shown as an example of an inference task. Figure 5 The musical score is an example of the arranged music (inferred result) obtained from the music. Figure 8A This indicates that the tokenization method is used to connect with... Figure 7This is an example of the truth value of the output token column corresponding to the music. Figure 8B This indicates that the tokenization method is used to connect with... Figure 7 This is an example of the truth value of the output token column corresponding to the music.
[0096] Action-based tokenization is a method of tokenizing music by representing actions corresponding to notes or elements of the music. Table 1 shows an example of the types of tokens and their representations in action-based tokenization. On the other hand, note-based tokenization is a method of tokenizing music by displaying the notes of the music as is. Table 2 shows an example of the types of tokens and their representations in note-based tokenization. Furthermore, the following token types and representations are examples and can be appropriately changed depending on the implementation method.
[0097] [Table 1]
[0098] type content An example of a token Press (Note on) Press the keyboard on_R72, on_R64, on_L48,… Release (Note off) Free up the keyboard off_L48, off_R72, off_R64,… The amount of time difference (Delta time) To make time pass (wait) wait_12, ...
[0099] [Table 2]
[0100] type content An example of a token Note Pitch pitch note_R72, note_R64, note_L48,… Note Value Length of sound len_24, len_12, ...
[0101] The tokenization method and token representation of the input and output token columns can also adopt either of the two methods mentioned above. As an example of a method for obtaining each training dataset, the representation can also be appropriately obtained. Figure 5 The example above contains at least a portion of the music data. The format of this music data can also be appropriately chosen depending on the implementation method. For example, the music data can also be obtained in the form of encoded data (MIDI, etc.), sheet music, etc. The training data 31 for each learning dataset 3 can also be derived from the obtained music data, to include... Figure 6A or Figure 6B The input token column shown is generated appropriately. Furthermore, it is obtained in accordance with... Figure 5 At least a portion of the illustrated musical piece corresponds to [the specific musical piece]. Figure 7 The inference result of the illustrated piece of music ( Figure 7 In the example, the true value of the arrangement result of the music is the positive solution data. The positive solution labels 32 of each learning dataset 3 can also be obtained from the positive solution data to contain Figure 8A or Figure 8B The truth values of the output token column shown are generated appropriately. The generation of the truth values of the input and output token columns can also employ arbitrary transformation processes such as natural language processing. The truth values of the input and output token columns can also be generated through manual operations.
[0102] in addition, Figure 7 , Figure 8A as well as Figure 8BThis example illustrates a scenario where generating the arranged music is used as an inference task. However, the inference task performed by inference model 5 is not limited to this, as described above. Similarly, the inference results and the truth values of the output token columns can be appropriately obtained for other inference tasks.
[0103] Figure 6A , Figure 6B , Figure 8A as well as Figure 8B The input and output token columns shown illustrate this, with the tokens "bar" and "beat" being examples of beat tokens. "Bar" is an example of a token representing a bar line, and "beat" is an example of a token representing a beat (time of action). Figure 6A as well as Figure 6B As illustrated, the input token column is configured to contain multiple beat tokens. Thus, the input token column is configured to determine the beat structure of the music. Furthermore, as... Figure 8A as well as Figure 8B As illustrated, the output token column can also be configured to include beat tokens. Furthermore, the behavior of these beat tokens is just one example. The behavior of the beat tokens is not limited to these examples and can be appropriately determined depending on the implementation method.
[0104] Furthermore, the input and output token strings can use the same tokenization method and the same token representation. In the example above, both the input and output token strings could use either action-based or note-based tokenization. However, the form of the input and output token strings is not limited to this example. The input and output token strings do not necessarily have to use the same tokenization method and the same token representation. They can also use any of the different tokenization methods and different token representations.
[0105] If the computer can identify at least a portion of the music being inferred, the form of the tokens used in the input token column is not particularly limited and can be appropriately determined depending on the implementation method. If the computer can identify the inference result, the form of the tokens used in the output token column is not particularly limited and can be appropriately determined depending on the implementation method. Furthermore, if the computer can identify the beat structure, the form of the beat tokens is not particularly limited and can be appropriately determined depending on the implementation method.
[0106] (An example of an inference model)
[0107] Figure 9An example of the structure of the inference model 5 according to this embodiment is illustrated schematically. The inference model 5 is constructed from a machine learning model having parameters adjusted through machine learning. The type of machine learning model is not particularly limited and can be appropriately selected according to the embodiment. If it is configured to accept an input token column as input and output an output token column representing the inference result, the construction of the machine learning model is not particularly limited and can be appropriately determined according to the embodiment. As an example, such as Figure 9 As shown, inference model 5 can also have a structure based on the transformer proposed in the reference "Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, 2017." A transformer is a machine learning model that processes sequential data (natural language, etc.) and has an attention-based structure.
[0108] exist Figure 9 In one example, inference model 5 includes an encoder 50 and a decoder 55. The encoder 50 has a stack-like structure consisting of multiple blocks, each containing a multi-head attention layer seeking self-attention and a feed-forward layer. Conversely, the decoder 55 has a stack-like structure consisting of multiple blocks, each containing a masked multi-head attention layer seeking self-attention, a multi-head attention layer seeking source / target attention, and a feed-forward layer. Figure 9 As shown, addition and normalization layers can also be set in each layer of encoder 50 and decoder 55. Each layer can also contain more than one node, and a threshold can be set for each node. The threshold can also be represented by an activation function. In addition, the connections between nodes in adjacent layers can also be weighted (connection load). The weights and thresholds of the connections between nodes are an example of the parameters of inference model 5.
[0109] exist Figure 9In one example, inference model 5 is configured to accept tokens from a sequence of input tokens sequentially from the beginning. The tokens input to inference model 5 are transformed into vectors of a specified dimension through input embedding processing, and then assigned values to determine their positions within the music (phrases) through positional encoding processing before being input to encoder 50. Encoder 50 repeatedly performs processing based on multi-head attention layers and feedforward layers on this input in an amount corresponding to the number of blocks to obtain feature representations, and then feeds the obtained feature representations to the next-level decoder 55 (multi-head attention layer).
[0110] The decoder 55 (masked multi-head attention layer) is supplied not only with input from the encoder 50, but also with known (past) outputs from the decoder 55. That is, Figure 9 The illustrated inference model 5 is constructed with a regression structure. For the input described above, decoder 55 repeatedly performs processing based on a masked multi-head attention layer, a multi-head attention layer, and a feedforward layer in an amount corresponding to the number of blocks to obtain feature representations and output them. The output from decoder 55 is transformed in a linear layer and a SOFTMAX layer and obtained as a token representing the inference result.
[0111] The learning processing unit 112 is configured to perform machine learning on the inference model 5 by using the input token columns (multiple tokens) contained in the training data 31 as input data for each learning dataset 3, and using the true values of the output token columns represented by the corresponding positive solution labels 32 as teacher signals. Specifically, the learning processing unit 112 is configured to input the input token columns contained in the training data 31 into the inference model 5 for each learning dataset 3, and train the inference model 5 so that the output token columns obtained by performing the computational processing of the inference model 5 are suitable for the true values represented by the corresponding positive solution labels 32. In other words, the learning processing unit 112 is configured to adjust the values of the parameters of the inference model 5 for each learning dataset 3 so that the error between the output token columns generated by the inference model 5 from the input token columns contained in the training data 31 and the true values represented by the corresponding positive solution labels 32 is reduced. Parameter adjustment can also use any method such as error backpropagation. Furthermore, various standardization methods (e.g., label smoothing, residual discarding, attention discarding) can be applied to the machine learning processing of the inference model 5.
[0112] <Music Inference Device>
[0113] Figure 10An example of the software structure of the music deduction device 2 according to this embodiment is illustrated schematically. The control unit 21 of the music deduction device 2 interprets the commands contained in the music deduction program 82 stored in the storage unit 22 and executes control processing corresponding to the interpreted commands. Thus, the music deduction device 2 according to this embodiment is configured to have a data acquisition unit 211, a deduction unit 212, and an output unit 213 as software modules. That is, in this embodiment, similar to the model generation device 1, the software modules of the music deduction device 2 are implemented by the control unit 21 (CPU).
[0114] The data acquisition unit 211 is configured to acquire object data 221, which includes an input token column arranged to represent at least a portion of a musical piece. The input token column is configured to include multiple beat tokens arranged to represent the positions of beats in the musical piece. The input token column included in the object data 221 is connected to... Figure 6A as well as Figure 6B The training data 31 shown is generated in the same form as the input token column.
[0115] The inference unit 212 possesses the trained inference model 5 while maintaining the learning result data 125. The inference unit 212 is configured to generate an output token column representing the inference result for the music from the input token column included in the object data 221 using the trained inference model 5. Figure 9 In one example, the inference unit 212 sequentially inputs the input token sequence contained in the object data 221 into the encoder 50 of the trained inference model 5 (specifically, the multi-head attention layer initially configured after the input embedding layer), and performs computational processing on the encoder 50 and decoder 55. As a result of this computational processing, the inference unit 212 sequentially obtains the input tokens from the trained inference model 5 ( Figure 9 In one example, the tokens output by the last configured SOFTMAX layer are used to generate an output token column representing the inference result. During this process, the output token column can also be generated using a search method such as beam search. More specifically, the inference unit 212 can also maintain n candidate tokens in descending order of score based on the probability distribution of the values output from the inference model 5, and select candidate tokens that result in the highest total score among m consecutive tokens, thereby generating an output token column (n and m are integers greater than 2). The output token column generated by the inference unit 212 can also be used in conjunction with... Figure 8A as well as Figure 8B The output tokens of the example correct solution label 32 are constructed in the same form.
[0116] The output unit 213 is configured to output the inference result obtained through the processing of the inference unit 212. The output format of the inference result is not particularly limited and can be appropriately determined depending on the implementation method. For example, the output token list may be output as is. Alternatively, the output unit 213 may transform the output token list into an appropriate form. For instance, in the case where the inference task generates an arranged musical piece, the output token list may be transformed into information representing the musical piece in the form of a note list, score, etc. Furthermore, the output unit 213 may output the information obtained through the transformation as the inference result.
[0117] <Other>
[0118] The software modules of the model generation apparatus 1 and the music inference apparatus 2 involved in this embodiment will be described in detail in the operation examples described later. Furthermore, in this embodiment, an example will be described where each software module of the model generation apparatus 1 and the music inference apparatus 2 is implemented by a general-purpose CPU. However, some or all of the above-mentioned software modules may also be implemented by one or more dedicated processors (e.g., application-specific integrated circuits (ASICs)). The above-mentioned modules may also be implemented as hardware modules. Moreover, regarding the software structure of the model generation apparatus 1 and the music inference apparatus 2, the software modules may be appropriately omitted, replaced, or added depending on the embodiment.
[0119] §3 Action Examples
[0120] <Model Generation Device>
[0121] Figure 11 This is a flowchart illustrating an example of the processing procedure of the model generation apparatus 1 according to this embodiment. The processing procedure of the model generation apparatus 1 described below is an example of a model generation method. However, the processing procedure of the model generation apparatus 1 described below is merely an example, and each step can be changed within the possible range. Furthermore, for the following processing procedure, steps can be appropriately omitted, substituted, or added according to the embodiment.
[0122] (Step S101)
[0123] In step S101, the control unit 11 operates as the learning data acquisition unit 111 to acquire multiple learning datasets 3.
[0124] Each learning dataset 3 can also be appropriately generated. Musical data representing the music in other forms such as encoded data or musical scores can also be obtained, and the input token column constituting the training data 31 can also be appropriately generated from the obtained musical data. The correct answer label 32 can also be appropriately generated as an output token column representing the true value of the inference result for the music.
[0125] The process of generating the learning dataset 3 can also be performed by any computer. In one example, the process of generating each learning dataset 3 can also be performed by the model generation device 1 (control unit 11). In another example, at least a portion of the multiple learning datasets 3 can also be generated by other computers. In this case, the model generation device 1 (control unit 11) can also obtain the learning datasets 3 generated by other computers via a network, storage medium 91, etc. The number of obtained learning datasets 3 can also be appropriately determined to be sufficient for machine learning. If multiple learning datasets 3 are obtained, the control unit 11 advances the process to the next step S102.
[0126] (Step S102)
[0127] In step S102, the control unit 11 operates as the learning processing unit 112, and uses the acquired multiple learning datasets 3 to perform machine learning of the inference model 5.
[0128] As an example of specific machine learning processing, the control unit 11 inputs the input token columns contained in the training data 31 sequentially into the inference model 5 for each learning dataset 3, and repeatedly executes the computational processing of the inference model 5 to sequentially generate tokens constituting the output token column. Through this computational processing, the control unit 11 can obtain the output token column representing the inference result corresponding to the training data 31 of each learning dataset 3. Next, the control unit 11 calculates the error between the obtained output token column and the true value represented by the corresponding positive solution label 32, and also calculates the angle of the calculated error. The control unit 11 backpropagates the calculated error angle using the error backpropagation method, thereby calculating the error of the parameter values of the inference model 5. Based on the calculated error, the control unit 11 adjusts the parameter values of the inference model 5. The control unit 11 can also repeatedly adjust the parameter values of the inference model 5 through the above series of processes until a predetermined condition is met (e.g., a predetermined number of executions, the sum of the calculated errors being below a threshold).
[0129] Through this machine learning, the inference model 5 is trained on each learning dataset 3 so that the output token column generated from the input token column contained in the training data 31 is adapted to the true value represented by the corresponding positive solution label 32. Therefore, as a result of the machine learning, a trained inference model 5 is generated that has acquired the ability to perform the inference task in a manner adapted to the true value given by the positive solution label 32. If the machine learning process ends, the control unit 11 causes the process to proceed to the next step S103.
[0130] (Step S103)
[0131] In step S103, the control unit 11 operates as a storage processing unit 113, generating learning result data 125 containing information related to the trained inference model 5 generated by machine learning. The learning result data 125 contains information for reproducing the trained inference model 5. For example, the learning result data 125 may also include information representing the values of each parameter of the inference model 5 obtained through the aforementioned machine learning adjustments. Depending on the situation, the learning result data 125 may also include information representing the structure of the inference model 5. The structure may be determined, for example, by the number of layers, the types of layers, the number of nodes in each layer, and the connection relationships between nodes in adjacent layers. The control unit 11 stores the generated learning result data 125 in a designated storage area.
[0132] The designated storage area may be, for example, RAM within the control unit 11, storage unit 12, external storage device, storage medium, or a combination thereof. The storage medium may be, for example, a CD, DVD, etc., and the control unit 11 may store the learning result data 125 on the storage medium via drive 17. The external storage device may be, for example, a data server such as a NAS. In this case, the control unit 11 may also use communication interface 13 to store the learning result data 125 on the data server via a network. Furthermore, the external storage device may be, for example, the storage device of a peripheral connected to the model generation device 1.
[0133] If the saving of the learning result data 125 is completed, the control unit 11 ends the processing of the model generation device 1 involved in this action example.
[0134] Furthermore, the generated learning result data 125 can be provided to the music inference device 2 at any time. For example, the control unit 11 can transfer the learning result data 125 to the music inference device 2 as part of step S103, or it can transfer the learning result data 125 to the music inference device 2 separately from the processing of step S103. The music inference device 2 can also obtain the learning result data 125 by receiving this transfer. In addition, for example, the music inference device 2 can also obtain the learning result data 125 by accessing the model generation device 1 or the data server via a network using the communication interface 23. In addition, for example, the music inference device 2 can also obtain the learning result data 125 via the storage medium 92. In addition, for example, the learning result data 125 can also be pre-loaded into the music inference device 2.
[0135] Furthermore, the control unit 11 can update or generate new learning result data 125 by periodically or irregularly repeating the processes of steps S101 to S103. During this repetition, changes, corrections, additions, deletions, etc., of at least a portion of the multiple learning datasets 3 used for machine learning can also be appropriately performed. Thus, the control unit 11 can also update or generate new inference models 5 after training. Moreover, the control unit 11 can also provide the updated or newly generated learning result data 125 to the music inference device 2 using any method, thereby updating the learning result data 125 held by the music inference device 2.
[0136] <Music Inference Device>
[0137] Figure 12 This is a flowchart illustrating an example of the processing procedure of the music inference device 2 according to this embodiment. The processing procedure of the music inference device 2 described below is an example of a music inference method. However, the processing procedure of the music inference device 2 described below is merely an example, and each step can be changed within the possible range. Furthermore, for the following processing procedure, steps can be appropriately omitted, substituted, or added according to the embodiment.
[0138] (Step S201)
[0139] In step S201, the control unit 21 operates as a data acquisition unit 211 to acquire object data 221, which includes a column of input tokens arranged in a manner representing at least a portion of a musical piece.
[0140] The input token column contained in object data 221 can also be generated by any method. In one example, the input token column can also be generated from other forms of data such as coded data, musical scores, etc. In another example, the input token column can also be generated directly by any method (e.g., by manual input).
[0141] Furthermore, object data 221 can be obtained via any path. In one example, the input token column can also be generated in the music inference device 2. In this case, the control unit 21 can also obtain object data 221 as a result of performing the generation process. In another example, the generation of the input token column can also be performed by a computer other than the music inference device 2. In this case, the control unit 21 can obtain object data 221, for example, via a network, storage medium 92, etc. If object data 221 is obtained, the control unit 21 advances the process to the next step S202.
[0142] (Step S202)
[0143] In step S202, the control unit 21 operates as the inference unit 212, referring to the learning result data 125, and sets the trained inference model 5 through machine learning. The control unit 21 uses the trained inference model 5 to generate an output token column representing the inference result for the music from the input token column contained in the object data 221. Specifically, the control unit 21 inputs the obtained input token column contained in the object data 221 into the trained inference model 5 and performs the computational processing of the trained inference model 5. In the above... Figure 9 In one example, the control unit 21 inputs the tokens contained in the input token column sequentially from the beginning into the trained inference model 5, and repeatedly performs the forward propagation operation of the trained inference model 5, thereby generating tokens that constitute the output token column in sequence. As a result of this operation, the control unit 21 obtains an output token column from the trained inference model 5, representing the result of completing the inference task relative to at least a part of the music represented by the object data 221. If the completion of the inference task (the operation of the trained inference model 5) ends, the control unit 21 advances the process to the next step S203.
[0144] (Step S203)
[0145] In step S203, the control unit 21 operates as the output unit 213, outputting the inference result obtained through the processing in step S202. The output destination and output format are not particularly limited and can be appropriately selected depending on the implementation method. For example, the output destination could be RAM, storage unit 22, storage medium, external storage device, other computer, other device, etc. Furthermore, in one example, the control unit 21 may output the output token list as is. In another example, the control unit 21 may transform the output token list into an appropriate form and output the information obtained through the transformation. As a specific example, when the inference task generates an arranged piece of music, the control unit 21 may also generate information in the form of note sequences, musical scores, etc., of the arranged music from the obtained output token list and output the generated information. When the inference result is obtained in the form of a musical score, the control unit 21 may, for example, cause a printing device (not shown) to output an instruction to print the musical score data on a paper medium.
[0146] If the output of the inference result ends, the control unit 21 terminates the processing of the music inference device 2 involved in this action example. Alternatively, the control unit 21 may, for example, periodically or irregularly repeat the processing of steps S201 to S203 as requested by the operator. During this repetition, at least a portion of the object data 221 (input token column) obtained in step S201 may be appropriately modified, corrected, added to, or deleted. Thus, the control unit 21 can use the trained inference model 5 to generate inference results for new music.
[0147] <Characteristics>
[0148] As described above, in this embodiment, the training data 31 of each learning dataset 3 used in the machine learning of step S102, i.e., the input token column, is composed of multiple beat tokens containing the positions of the beats of the music. Thus, the inference model 5 is trained to perform inference processing for the music after mastering the beat structure of the music through the beat tokens. As a result, through the machine learning processing of step S102, a trained inference model 5 that is less prone to timing errors caused by the beat structure can be generated.
[0149] Furthermore, in step S201 above, an input token column containing multiple beat tokens is obtained as object data 221. Moreover, in step S202, the input token column containing multiple beat tokens is used for inference of the piece of music by the trained inference model 5. Thus, the trained inference model 5 can perform inference processing for the piece of music after mastering its beat structure. As a result, the probability of timing errors can be reduced in the inference task for the piece of music in step S202.
[0150] Furthermore, in this embodiment, each beat token can also be configured to represent the position of the bar line and the beat in the input token column. That is, multiple beat tokens can also be configured to grasp both the bar line and the beat. Thus, the inference model 5 can completely determine the beat structure of the music represented by the input token column based on each beat token. Therefore, in the processing of step S102 above, a trained inference model 5 that is less prone to timing errors can be generated. In the music inference task of step S202 above, the probability of timing errors can be further reduced.
[0151] In this embodiment, the output token column can also be configured to include beat tokens. Therefore, even if a timing error occurs during the inference process in step S202, the location of the timing error can be easily determined based on the position of the beat tokens included in the output token column. As a result, the obtained inference result can be easily corrected.
[0152] §4 Variations
[0153] The embodiments of the present invention have been described in detail above, and the description herein is merely illustrative of the invention. It is obvious that various modifications or variations can be made without departing from the scope of the present invention.
[0154] For example, in the above embodiment, as inference model 5, a machine learning model with a regression construction based on a transformer structure is illustrated. Figure 9 However, regression construction is not limited to Figure 9 The example shown illustrates a regression construct, which is configured to perform processing on the (current) input of an object by referring to inputs from the past. The regression construct is not particularly limited as long as such operations can be performed, and can be appropriately determined depending on the implementation method. In another example, the regression construct can also be constructed using known constructs such as RNNs (Recurrent Neural Networks) and LSTMs (Long Short-Term Memory).
[0155] Furthermore, in the above embodiment, the inference model 5 is configured to have a regression structure. However, the structure of the inference model 5 is not limited to such an example. The regression structure may also be omitted. The inference model 5 may also be configured using a neural network with a known structure such as a fully connected neural network or a convolutional neural network. Moreover, the form in which the input token column is input into the inference model 5 is not limited to the example of the above embodiment. In another example, the inference model 5 may also be configured to accept multiple tokens contained in the input token column at once.
[0156] Furthermore, in the above embodiments, as long as an output token column representing the inference result can be generated from the input token column corresponding to the music, the type of machine learning model constituting the inference model 5 is not particularly limited and can be appropriately selected according to the implementation method. Moreover, in the above embodiments, when the inference model 5 is constituted by a machine learning model with multiple layers, the type of each layer can also be appropriately selected according to the implementation method. Each layer can be, for example, a convolutional layer, a pooling layer, a dropout layer, a normalization layer, a fully connected layer, etc. Regarding the construction of the inference model 5, structural elements can be appropriately omitted, replaced, or added.
[0157] In the above embodiments, the inference model 5 can also be configured to further accept input information other than the aforementioned input token column. Furthermore, the inference model 5 can also be configured to further output information other than the aforementioned output token column.
[0158] §5 Examples
[0159] To verify the effectiveness of the present invention, the trained inference models involved in the following embodiments and comparative examples were generated, and the inference accuracy of the generated trained inference models was evaluated.
[0160] Specifically, 261,396 original musical pieces were prepared, using the methods described in Table 1 above. Figure 6A as well as Figure 8A The action-based tokenization method shown generates an input token column constituting the training data from the prepared samples of music. Furthermore, as an inference task, an arrangement of the music is conceived, and 261,396 samples of the arranged music are prepared, each corresponding to an original piece. Moreover, similar to the input token column, an action-based tokenization method is used to generate ground values for the output token column constituting the correct answer label from each sample of the arranged music. By matching the generated input token column (training data) with the ground values (correct answer labels) of the output token column, a training dataset of 261,396 samples is generated. In the embodiment, as... Figure 6A as well as Figure 8A As shown, beat tokens are configured at the positions of bar lines and beats, respectively, for the true values of the input token column and the output token column, thereby obtaining a learning dataset. On the other hand, in the comparative example, the true values of both the input token column and the output token column are not included in the beat tokens (otherwise, it is the same as the embodiment), and a learning dataset is obtained.
[0161] The inference models involved in the embodiments and comparative examples are constructed using... Figure 9 The illustrated Transformer. Using the same method as the embodiments described above, machine learning is performed on a prepared learning dataset of 261,396 samples to generate the trained inference model involved in the embodiments and comparative examples.
[0162] A set of 1000 musical pieces (each sample being 4 measures in length) was prepared separately from the training data. From the prepared musical pieces, a set of 1000 input tokens (object data) was obtained. Similar to the training dataset, the input tokens in this embodiment were configured with beat tokens at the bar lines and the respective positions of the beats. On the other hand, beat tokens were not configured in the input tokens of the comparative example (otherwise, it was the same as the embodiment).
[0163] Next, using the trained inference models from both the embodiment and the comparative example, an inference task was performed relative to the object data of each sample, resulting in an output token column representing the inference result. Furthermore, it was evaluated whether a beat count deviation occurred in the arrangement represented by the output token column compared to the original composition (i.e., the composition represented by the object data). As a result, in the comparative example, a beat count deviation occurred with a probability of 17.4%. On the other hand, in the embodiment, a beat count deviation occurred with a probability of 4.1%. Based on these results, it can be seen that including beat tokens representing beat construction can significantly reduce the probability of timing errors.
Claims
1. A music inference device, comprising: The data acquisition unit acquires object data, which includes an input token column arranged in a manner that represents at least a portion of a musical piece, the input token column containing a plurality of beat tokens configured in a manner that represents the position of the beats in the musical piece; The inference unit generates an output token column representing the inference result for the music by using a trained inference model from the input token column contained in the object data. as well as The output section outputs the result of the inference. Multiple beat tokens are respectively configured in the bar lines of the music and the position of each beat in the input token column.
2. The music deduction device as described in claim 1, wherein, The input token column is generated in correspondence with at least a portion of the note column of the music. The output token column is generated in such a way that it represents at least a portion of the note sequence of the arranged music as a result of inferences made for the music.
3. The music deduction device as described in claim 1, wherein, The input token column is generated in correspondence with at least a portion of the note column of the music. The output token column is generated in such a way that it shows the result of inferences about local attributes in at least a part of the music.
4. The music deduction device as described in claim 1, wherein, The input token column is generated in correspondence with at least a portion of the note column of the music. The output token column is generated in such a way that it represents at least a portion of the score of the music as a result of inferences made about the music.
5. The music deduction device as described in claim 1, wherein, The input token column is generated in correspondence with at least a portion of the element columns of the music. The output token column is generated in such a way that it represents a sequence of notes of at least a portion of the arranged music as a result of inferences made about the music.
6. A method for inferring musical composition, comprising the following steps performed by a computer: The step of obtaining object data includes an input token column arranged in a manner that represents at least a portion of a musical piece, the input token column containing a plurality of beat tokens configured in a manner that represents the position of beats in the musical piece; The step of generating an output token column representing the result of the inference for the music from the input token column contained in the object data by using a trained inference model; as well as The step of outputting the inference result, Multiple beat tokens are respectively configured in the bar lines of the music and the position of each beat in the input token column.
7. A storage medium storing a program for causing a computer to perform the following steps: The step of obtaining object data includes an input token column arranged in a manner that represents at least a portion of a musical piece, the input token column containing a plurality of beat tokens configured in a manner that represents the position of beats in the musical piece; The step of generating an output token column representing the result of the inference for the music from the input token column contained in the object data by using a trained inference model; as well as The step of outputting the inference result, Multiple beat tokens are respectively configured in the bar lines of the music and the position of each beat in the input token column.
8. A model generation apparatus, comprising: The learning data acquisition unit acquires multiple learning datasets, each consisting of a combination of training data and correct answer labels. The training data includes an input token column arranged to represent at least a portion of a piece of music for learning. The input token column includes multiple beat tokens configured to represent the positions of beats in the music. The correct answer labels are configured to represent the truth values of an output token column corresponding to the inference result for the music. The learning processing unit performs machine learning on an inference model using multiple acquired learning datasets. The machine learning is configured by training the inference model on each of the learning datasets, such that an output token column generated by the inference model from the input token column contained in the training data is adapted to the truth value represented by the positive solution label. Multiple beat tokens are respectively configured in the bar lines of the music and the position of each beat in the input token column.
9. The model generation apparatus as described in claim 8, wherein, The input token column included in the training data is generated in correspondence with at least a portion of the note column of the music. The output token column of the correct answer label is configured to show the truth value of at least a portion of the notes of the arranged music as the truth value of the inference result for the music.
10. The model generation apparatus as claimed in claim 8, wherein, The input token column included in the training data is generated in correspondence with at least a portion of the note column of the music. The output token column of the correct answer label is configured to show the truth value of the result of inferring local attributes in at least a part of the music, as the truth value of the result of the inference for the music.
11. The model generation apparatus as claimed in claim 8, wherein, The input token column included in the training data is generated in correspondence with at least a portion of the note column of the music. The output token column of the correct answer label is configured to show the truth value of at least a portion of the score of the music as the truth value of the result of the inference for the music.
12. The model generation apparatus as claimed in claim 8, wherein, The input token column contained in the training data is generated in correspondence with at least a portion of the element columns of the music. The output token column of the correct answer label is configured to show the truth value of at least a portion of the note column of the arranged music as the truth value of the result of the inference for the music.
13. A model generation method, wherein a computer performs the following steps: The steps include obtaining multiple learning datasets composed of combinations of training data and correct answer labels, wherein the training data includes an input token column arranged to represent at least a portion of a piece of music for learning, the input token column including multiple beat tokens configured to represent the positions of beats in the music, and the correct answer labels being configured to represent the truth values of an output token column corresponding to the inference result for the music; and The machine learning steps of the inference model are implemented using multiple acquired learning datasets. The machine learning is configured by training the inference model on each learning dataset so that the output token column generated by the inference model from the input token column contained in the training data is adapted to the truth value represented by the positive solution label. Multiple beat tokens are respectively configured in the bar lines of the music and the position of each beat in the input token column.
14. A storage medium storing a program for causing a computer to perform the following steps: The steps include obtaining multiple learning datasets composed of combinations of training data and correct answer labels, wherein the training data includes an input token column arranged to represent at least a portion of a piece of music for learning, the input token column including multiple beat tokens configured to represent the positions of beats in the music, and the correct answer labels being configured to represent the truth values of an output token column corresponding to the inference result for the music; and The machine learning steps of the inference model are implemented using multiple acquired learning datasets. The machine learning is configured by training the inference model on each learning dataset so that the output token column generated by the inference model from the input token column contained in the training data is adapted to the truth value represented by the positive solution label. Multiple beat tokens are respectively configured in the bar lines of the music and the position of each beat in the input token column.
Citation Information
Patent Citations
Automatic arrangement device and program
JP2017058594A
Beat extraction device and beat extraction method
CN101375327A
Method and device for correcting time delay between accompaniment and dry sound, and storage medium
CN108711415A