Data processing methods, programs, and systems
The use of AI models to convert performance data into tone control data and generate musical scores addresses the lack of automated sound control data processing, enabling enhanced musical experiences and instructional services.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-10-18
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies lack the ability to automatically generate sound control data based on performance and process it according to specific user purposes, such as generating musical scores or adding desired arrangements, using AI for enhanced musical performance and transcription.
A data processing method utilizing trained AI models to convert performance data into tone control data, which can be further processed to generate musical scores or include user-specified arrangements, and generate performance control signals, leveraging machine learning models like CNNs and RNNs to analyze and transform performance data into sound control data and musical scores.
Enables automatic generation and processing of sound control data to create musical scores and perform desired arrangements, allowing for enhanced musical experiences and instructional services, including sheet music provision and automatic performance.
Smart Images

Figure 0007845491000001 
Figure 0007845491000002 
Figure 0007845491000003
Abstract
Description
[Technical Field]
[0001] This invention relates to a technology for processing data. [Background technology]
[0002] A technology is known that displays musical scores based on a performer's performance. For example, Patent Document 1 discloses a device that determines and displays the performance location in the musical score data of a corresponding song based on pitch data of the sound. [Prior art documents] [Patent Documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2001-337675 [Overview of the project] Problems that the invention aims to solve
[0004] One of the objectives of this invention is to use AI to automatically generate sound control data based on performance, and to perform desired processing on the generated sound control data according to the purpose. [Means for solving the problem]
[0005] According to one embodiment of the present invention, a data processing method is provided which includes: obtaining first tone control data, including pitch information, note value information, and sound production timing, from a first trained model into which performance data has been input; inputting parameters corresponding to the first tone control data and first user-specified information into a second trained model; and obtaining second tone control data from the second trained model.
[0006] One embodiment of the present invention provides a data processing method that includes obtaining first tone control data, including pitch information, note value information, and sound production timing, from a first trained model into which performance data has been input; inputting the first tone control data into a third trained model; and obtaining musical score data from the third trained model.
[0007] According to one embodiment of the present invention, a data processing method is provided which includes obtaining first tone control data, including pitch information, note value information, and sound production timing, from a first trained model into which performance data has been input; obtaining desired tempo information; and generating a performance control signal based on the first tone control data and the tempo information.
[0008] According to one embodiment of the present invention, a data processing method is provided which includes: obtaining first tone control data, including pitch information, note value information, and sound production timing, from a first trained model into which performance data has been input; inputting parameters corresponding to the first tone control data and second user-specified information into a fourth trained model; and obtaining image control data corresponding to the sound production timing from the fourth trained model.
[0009] According to one embodiment of the present invention, a program is provided that causes a computer to perform any of the data processing methods described above. [Effects of the Invention]
[0010] According to the present invention, AI can be used to automatically generate sound control data based on performance, and the generated sound control data can be processed according to the purpose. [Brief explanation of the drawing]
[0011] [Figure 1] This is a diagram illustrating the system configuration in one embodiment. [Figure 2] This is a block diagram illustrating the configuration of an electronic musical instrument according to one embodiment. [Figure 3] This is a block diagram illustrating the configuration of a data processing device 10 according to one embodiment. [Figure 4] This block diagram shows the configuration of an automatic music transcription function according to one embodiment. [Figure 5] This block diagram shows the configuration of an automatic music transcription function according to one embodiment. [Figure 6]A diagram for explaining a data processing method according to an embodiment. [Figure 7] A diagram for explaining a data processing method according to an embodiment. [Figure 8] A diagram for explaining a data processing method according to an embodiment. [Figure 9] A block diagram for explaining the configuration of a data processing apparatus according to another embodiment. [Figure 10] A block diagram showing the configuration of an automatic score-taking function in another embodiment. [Figure 11] A diagram for explaining a data processing method according to another embodiment. [Figure 12] A diagram for explaining a data processing method according to another embodiment. [Figure 13] A diagram for explaining a model generation function according to an embodiment. [Figure 14] A diagram for explaining a model generation function according to an embodiment. [Figure 15] A diagram for explaining a model generation function according to an embodiment. [Figure 16] A diagram for explaining a model generation function according to an embodiment. [Figure 17] A diagram for explaining the generation process of first sound control data. [Figure 18] A diagram for explaining the generation process of first sound control data. [Figure 19] A diagram for explaining the generation process of first sound control data. [Figure 20] A diagram for explaining the generation process of first sound control data. [Figure 21] A diagram for explaining the generation process of first sound control data. [Figure 22] A diagram for explaining the generation process of first sound control data. [Figure 23] A diagram for explaining the generation process of first sound control data. [Figure 24] A diagram for explaining the generation process of first sound control data. [Figure 25]This is a diagram illustrating the generation process of the first tone control data. [Figure 26] This is a diagram illustrating the generation process of the first tone control data. [Figure 27] This is a diagram illustrating the generation process of the first tone control data. [Figure 28] This is a diagram illustrating the generation process of the first tone control data. [Figure 29] This is a diagram illustrating the generation process of the first tone control data. [Figure 30] This is a diagram illustrating the generation process of the first tone control data. [Figure 31] This is a diagram illustrating the generation process of the first tone control data. [Figure 32] This is a diagram illustrating the generation process of the first tone control data. [Figure 33] This is a diagram illustrating the generation process of the first tone control data. [Figure 34] This is a diagram illustrating the generation process of the first tone control data. [Figure 35] This is a diagram illustrating the generation process of the first tone control data. [Figure 36] This is a diagram illustrating the generation process of the first tone control data. [Figure 37] This is a diagram illustrating the generation process of the first tone control data. [Figure 38] This is a diagram illustrating the generation process of the first tone control data. [Figure 39] This is a diagram illustrating the generation process of the first tone control data. [Figure 40] This is a diagram illustrating the generation process of the first tone control data. [Figure 41] This is a diagram illustrating the generation process of the first tone control data. [Figure 42] This is a diagram illustrating the generation process of the first tone control data. [Modes for carrying out the invention]
[0012] Hereinafter, one embodiment of the present invention will be described in detail with reference to the drawings. The embodiments shown below are examples, and the present invention is not limited to these embodiments. In the drawings referenced in this embodiment, the same parts or parts having similar functions are denoted by the same or similar reference numerals (simply a number followed by A, B, etc.), and repeated descriptions may be omitted.
[0013] <First Embodiment> [System Configuration] Figure 1 is a diagram illustrating the system configuration in the first embodiment. The system 1 shown in Figure 1 includes a data processing device 10 and an electronic musical instrument 20. In the system 1 shown in Figure 1, the data processing device 10 and the electronic musical instrument 20 are connected to each other, but the data processing device 10 and the electronic musical instrument 20 may be connected via a network such as the Internet.
[0014] The data processing device 10 is, for example, a computer such as a smartphone, tablet PC, laptop PC, or desktop PC. Alternatively, the data processing device 10 may be a server connected to the electronic musical instrument 20 via a network. In this example, the electronic musical instrument 20 is an electronic keyboard device such as an electronic piano.
[0015] When a user performs a predetermined performance operation on the electronic instrument 20, the data processing device 10 generates sound control data based on the performance data output in response to the performance operation. The data processing device 10 processes this sound control data for purposes such as generating a musical score, automatic performance on the electronic instrument 20, and adding arrangements desired by the user. A detailed explanation of the data processing device 10 will be given later.
[0016] [Electronic musical instruments] Figure 2 is a block diagram illustrating the configuration of the electronic instrument 20 according to this embodiment. In this example, the electronic instrument 20 is an electronic keyboard device such as an electronic piano. The electronic instrument 20 includes a performance control unit 201, a sound source 203, a speaker 205, a drive control unit 207, a drive unit 209, and an interface 211. The performance control unit 201 includes multiple keys and outputs a performance signal to the sound source 203 or interface 211 according to the operation of each key. The performance signal is sound generation control information and is output sequentially in real time.
[0017] The sound source 203 includes a DSP (Digital Signal Processor). When performance is performed in the electronic instrument 20 based on operations on the performance control 201, the sound source 203 generates playback sound data based on the performance signal and outputs it to the speaker 205. Also, when performance is performed in the electronic instrument 20 based on sound control data processed by the data processing device 10, the sound source 203 generates playback sound data based on the sound control data provided from the data processing device 10 via the interface 211 and outputs it to the speaker 205. Here, the playback sound data is sound waveform data.
[0018] Speaker 205 converts the playback sound data provided by sound source 203 into air vibrations and delivers them to the user.
[0019] The drive control unit 207 generates a drive signal based on sound control data provided from the data processing device 10 via the interface 211, and outputs the generated drive signal to the drive unit 209. The drive unit 209 is a drive mechanism that operates the performance control element 201, and is, for example, a solenoid.
[0020] Interface 211 includes a module for sending and receiving data to and from external devices wirelessly or via a wired connection. In this example, interface 211 connects to the data processing device 10 wirelessly or via a wired connection and sequentially transmits performance signals output in response to operations on the performance control unit 201 to the data processing device 10. Interface 211 also receives performance control signals generated by the data processing device 10 and outputs the received performance control signals to the sound source 203 and / or drive control unit 207.
[0021] [Data Processing Device] Figure 3 is a block diagram illustrating the configuration of the data processing device 10 according to this embodiment. The data output device 10 includes a control unit 101, a storage unit 103, an operation unit 105, and an interface 107.
[0022] The control unit 101 is an example of a computer equipped with a processor such as a CPU and a memory device such as RAM. The control unit 101 executes the program 131 stored in the memory unit 103 using the CPU (processor) and implements functions for performing various processes in the data processing device 10. The functions implemented in the data processing device 10 include the automatic music transcription function described later. Furthermore, the functions implemented in the data processing device 10 may also include a model training function.
[0023] The memory unit 103 is a storage device such as RAM, ROM, non-volatile memory, or a hard disk drive. The memory unit 103 stores the program 131 executed by the control unit 101 and various data necessary for executing this program 131. The memory unit 103 stores multiple trained models obtained by machine learning. The trained models stored in the memory unit 103 include a first trained model 132, a second trained model 133, and a third trained model 134. Furthermore, the memory unit 103 includes a storage area 135 for temporarily storing data output from each trained model.
[0024] Program 131 may be downloaded from an external server via a network and stored in the storage unit 103, thereby being installed in the data processing device 10. Alternatively, Program 131 may be provided recorded on a non-transient, computer-readable recording medium (e.g., magnetic recording medium, optical recording medium, magneto-optical recording medium, semiconductor memory, etc.). In this case, the data processing device 10 only needs to be equipped with a device to read this recording medium. The storage unit 103 can be considered an example of a recording medium. Details of the first trained model 132, the second trained model 133, and the third trained model 134 will be described later.
[0025] The memory area 135 is composed of a storage device such as RAM and temporarily stores data used in processing performed by the data processing device 10 and data generated by said processing. In this example, the memory area 135 includes a performance data storage area 135a, a sound control data storage area 135b, a musical score data storage area 135c, and so on.
[0026] The performance data storage area 135a is an area that stores performance signals sequentially provided from the electronic instrument 20 via the interface 107 as a single data file (performance data Pd). The performance signals are converted into sequence data in a predetermined format by the control unit 101 and stored in the performance data storage area 135a in association with time information. The predetermined format is, for example, MIDI format. In other words, the performance signals are converted into data that includes sound control information, such as note-on, note-off, and note number, which define the content of the sound produced by the performer's performance, and stored in the performance data storage area 135a in association with time information.
[0027] The sound control data storage area 135b is a region that temporarily stores the first sound control data SC1 output from the first trained model 132 (described later) and the second sound control data SC2 output from the second trained model 133. Although not shown in the diagram, the sound control data storage area 135b includes the first sound control data storage area and the second sound control data storage area. The first sound control data storage area temporarily stores the first sound control data SC1, and the second sound control data storage area temporarily stores the second sound control data SC2.
[0028] The first tone control data SC1 is output from the first trained model 132 in response to the input of the performance data Pd. Details of the first trained model 132 will be described later. The first tone control data SC1 is data that includes pitch information, note value information, and sound timing. Pitch information, note value information, and sound timing are defined for each note and are related to each other. Pitch information is information corresponding to the note number. Note value information is information indicating the length of the note on the musical score. In this specification, sound timing corresponds to the performance timing on the musical score and is not timing in absolute time. For example, sound timing is information indicating relative time, which is determined by the number of measures, beats, etc. Sound timing can be converted to timing in absolute time by determining the tempo (performance speed) of the piece. Pitch information, note value information, and sound timing are defined for each note that makes up the piece.
[0029] The second tone control data SC2 is output from the second trained model 133 in response to the input of the first tone control data SC1. Details of the second trained model 133 will be described later. The second tone control data SC2, like the first tone control data SC1, is data that includes pitch information, note value information, and pronunciation timing.
[0030] The musical score data storage area 135c is an area for temporarily storing the musical score data Sd output from the third trained model 134, which will be described later. The musical score data Sd is output from the third trained model 134 in response to the input of sound control data (first sound control data SC1 or second sound control data SC2). Details of the third trained model 134 will be described later. The musical score data Sd is data for displaying the musical score generated based on the sound control data on a display device.
[0031] The operation unit 105 is an operating device that outputs signals to the control unit 101 in response to user operations. In this embodiment, the signals in response to user operations include user instruction information UI, tempo information, first user-specified information UD1, and second user-specified information UD2. The user instruction information UI is information indicating the processing to be executed in the data processing device 10. The tempo information is information indicating the performance speed (tempo) of the musical piece. The first user-specified information UD1 and the second user-specified information UD2 will be described later. The interface 107 includes a module for communicating with an external device by wireless or wired communication. In this example, the external device includes an electronic musical instrument 20.
[0032] [Pre-trained model] The first trained model 132 is a computational model used to convert the input performance data Pd of a predetermined musical piece into first tone control data SC1. In this embodiment, the first trained model 132 has two computational models. The two computational models correspond to the first encoder and the first decoder. A known machine learning model is applied to each computational model. The two computational models may be different models from each other. A known machine learning model is, for example, a neural network model using CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), etc. The first tone control data SC1 is data for reconstructing the musical score that the performer is assumed to have seen and played, by removing the performer's performance habits from the input performance data Pd. The performer's performance habits include the tempo (speed) of the performance and the dynamics of the piece.
[0033] In other words, the first trained model 132 is a trained model obtained by machine learning the correlation between performance data Pd and first tone control data SC1. The correlation between performance data Pd and first tone control data SC1 indicates the correspondence between the sound production control information of the performance data and the first tone control data SC1. The first trained model 132 is a trained model obtained by learning the performance content when various performers play the song. When performance data Pd is input to the first trained model 132, it outputs the first tone control data SC1 corresponding to the input data.
[0034] The second trained model 133 is a computational model used when adding user-desired arrangements to the first sound control data SC1. In this embodiment, the second trained model 133 has three computational models. The three computational models correspond to the second encoder, the second decoder, and the third encoder. A known machine learning model is applied to each computational model. Known machine learning models include, for example, models using neural networks such as CNNs and RNNs.
[0035] The second trained model 133 is a trained model obtained by machine learning the correlation between the first sound control data SC1 and the first user-specified information UD1 and the second sound control data SC2. When the first sound control data SC1 and the first user-specified information UD1 are input to the second trained model 133, it outputs the second sound control data SC2 corresponding to the input data.
[0036] The first user-specified information UD1 is input by the user via the operation unit 105. The first user-specified information UD1 is information indicating the performance genre desired by the user, specifically, the genre of performance the user wishes to reproduce. The performance genre includes, for example, the performer desired by the user, and the genre of music desired by the user, including pop, jazz, rock, Latin, etc. For example, if the user wishes to reproduce a performance by a specific performer (for example, a specific pianist), the first user-specified information UD1 will include information indicating the performer desired by the user. Alternatively, if the user wishes to perform a performance based on a specific genre, the first user-specified information UD1 will include information indicating the genre desired by the user.
[0037] The second tone control data SC2 output by the second trained model 133 is data obtained by processing the first tone control data SC1 based on the first user-specified information UD1. For example, if the first user-specified information UD1 includes information indicating a specific pianist, the second tone control data SC2 output by the second trained model 133 is arranged so that the playing habits of that specific pianist are added to the first tone control data SC1. Also, for example, if the first user-specified information UD1 includes information indicating jazz, the second tone control data SC2 output by the second trained model 133 is arranged so that the characteristics of jazz are added to the first tone control data SC1.
[0038] The third pre-trained model 134 is a computational model used when generating musical score data based on sound control data. In this embodiment, the third pre-trained model 134 has two computational models. The two computational models correspond to the third encoder and the third decoder. A known machine learning model is applied to each computational model. A known machine learning model is, for example, a model using a neural network such as a CNN or RNN.
[0039] In other words, the third pre-trained model 134 is a pre-trained model obtained by machine learning the correlation between sound control data and musical score data. When sound control data is input to the third pre-trained model 134, it outputs musical score data Sd corresponding to the input data. The sound control data input to the third pre-trained model 134 is either the first sound control data SC1 output from the first pre-trained model 132 or the second sound control data SC2 output from the second pre-trained model 133. The musical score data Sd output from the third pre-trained model 134 is data for displaying the musical score.
[0040] [Automatic music transcription function] The automatic music transcription function, which is realized by the control unit 101 executing program 131, will be described below. At least some of the functions of the automatic music transcription function described below may be realized by other devices connected to the data processing device 10 via a network. The sound conversion function may be realized by multiple devices connected via a network working together.
[0041] Figures 4 and 5 are block diagrams showing the configuration of the automatic music transcription function 40 in this embodiment. The automatic music transcription function 40 includes a first note conversion unit 401, a second note conversion unit 403, a music score generation unit 405, and a performance control signal generation unit 407.
[0042] The performance signal input from the electronic instrument 20 is stored in the performance data storage area 135a as a single data file called performance data Pd, associated with time information. The generated performance data Pd is input to the first tone conversion unit 401.
[0043] The first sound conversion unit 401 includes a first trained model 132 and a feature extraction unit 323. The first trained model 132 includes a first encoder 321 and a first decoder 322. The first sound conversion unit 401 is supplied with performance data Pd from the performance data storage area 135a. The performance data Pd is input to the first trained model 132. As described above, the first trained model 132 generates and outputs first sound control data SC1 corresponding to the input performance data Pd. The performance data Pd provided from the performance data storage area 135a is also input to the feature extraction unit 323. The feature extraction unit 323 extracts sound features contained in the provided performance data Pd and provides feature information indicating these features to the first decoder 322 of the first trained model 132. The feature information is used to generate the first sound control data SC1. Details of the first encoder 321 and first decoder 322 and feature extraction unit 323 of the first trained model 132 will be described later. Although not shown in the figures, the first sound control data SC1 output from the first trained model 132 is temporarily stored in the first sound control data storage area of the sound control data storage area 135b and output to the second sound conversion unit 403 or the music score generation unit 405.
[0044] Figure 4 shows a configuration in which the first sound control data SC1 is output to the second sound conversion unit 403 based on user instruction information UI. In this case, the user instruction information UI includes information instructing the unit to execute a process to generate the second sound control data SC2. The second sound conversion unit 403 includes a second trained model 133. The second trained model 133 has a second encoder 331, a second decoder 332, and a third encoder 333. The third encoder 333 receives first user-specified information UD1 from the user via the operation unit 105. The third encoder 333 outputs parameters corresponding to the input first user-specified information UD1. These parameters are vector values. The parameters output from the second encoder 333 are input to the second decoder 332.
[0045] The first sound control data SC1 is input to the second encoder 331. The second encoder 331 converts the input first sound control data SC1 into a vector value and outputs it. The vector value output from the second encoder 331 is input to the second decoder 332.
[0046] The second decoder 332 receives the vector value output from the second encoder 331 and the parameters corresponding to the first user-specified information UD1 output from the third encoder 333. Based on the input first sound control data SC1 and parameters, the second decoder 332 generates and outputs the second sound control data SC2. Although not shown in the diagram, the second sound control data SC2 output from the second trained model 133 is temporarily stored in the second sound control data storage area of the sound control data storage area 135b and output to the score generation unit 405 and / or the performance control signal generation unit 407.
[0047] The music score generation unit 405 includes a third trained model 134. The third trained model 134 includes a fourth encoder 341 and a fourth decoder 342. When the second tone control data SC2 output from the second decoder 332 of the second trained model 133 is input to the music score generation unit 405, the second tone control data SC2 is input to the fourth encoder 341. The fourth encoder 341 converts the input second tone control data SC2 into a vector value and outputs it. The vector value output from the fourth encoder 341 is input to the fourth decoder 342.
[0048] The fourth decoder 342 receives a vector value output from the fourth encoder 341. Based on the input vector value, the fourth decoder 342 generates and outputs musical score data Sd. Although not shown in the diagram, the musical score data Sd output from the third trained model 134 is temporarily stored in the musical score data storage area 135c and provided via the interface 107 to a display device capable of displaying a musical score based on the musical score data Sd. The display device may be included in the electronic instrument 20. In this case, the musical score data Sd is provided to the electronic instrument 20. Alternatively, the display device may be an external device different from the electronic instrument 20. In this case, the musical score data Sd is provided to the external device. The external device, including the display device, is a device capable of sending and receiving data with the data processing device 10. The external device may be a device capable of sending and receiving data from the data processing device via a network.
[0049] The second sound control data SC2 output from the second decoder 332 of the second trained model 133 can be input to the performance control signal generation unit 407 according to the user instruction information UI. The performance control signal generation unit 407 generates a performance control signal based on the second sound control data SC2 and tempo information and outputs it to the electronic instrument 20.
[0050] As described above, the performance control signal is a MIDI-format sound generation control signal generated based on the second tone control data SC2. The tempo information is input to the user via the operation unit 105 and indicates the performance speed (tempo) desired by the user. The performance control signal generated by the performance control signal generation unit 407 is sequentially transmitted via the interface 107 to the interface 211 of the electronic instrument 20 shown in Figure 2, and provided to the sound source 203 and / or drive control unit 207.
[0051] The sound source 203 generates playback sound data based on the provided performance control signal. The playback sound data is a sound waveform signal based on the performance control signal. The sound source 203 outputs the playback sound data to the speaker 205. The speaker 205 converts the provided sound waveform signal into air vibrations and provides them to the user.
[0052] The drive control unit 207 generates a drive control signal based on the performance control signal. The drive control unit 207 outputs the generated drive signal to the drive unit 209. The drive unit 209 drives the performance control element 201 based on the drive signal.
[0053] As a result, the electronic instrument 20 can reproduce or automatically play a performance sound based on the second tone control data SC2. As described above, the second tone control data SC2 is arranged so that the user's desired features are added to the first tone control data SC1 based on the first user-specified information UD1. Therefore, the electronic instrument 20 can reproduce or automatically play a performance sound with the user's desired features added for a given musical piece.
[0054] Figure 4 shows the configuration in which the first tone control data SC1 is output to the second tone conversion unit 403. However, the first tone control data SC1 output from the first trained model 132 may be output to the score generation unit 405 without passing through the second tone conversion unit 403.
[0055] Figure 5 shows the configuration in which the first tone control data SC1 is output to the score generation unit 405. In this case, the user instruction information UI does not contain information instructing the user to execute the process of generating the second tone control data SC2. In this example, the first tone control data SC1 is output to the score generation unit 405 without passing through the second tone conversion unit 403. The first tone control data SC1 is input to the fourth encoder 341 of the score generation unit 405. The fourth encoder 341 converts the input first tone control data SC1 into a vector value and outputs it. The vector value output from the fourth encoder 341 is input to the fourth decoder 342.
[0056] The fourth decoder 342 receives the vector value output from the fourth encoder 341. Based on the input vector value, the fourth decoder 342 generates and outputs musical score data Sd. The musical score data Sd output from the third trained model 134 is temporarily stored in the musical score data storage area 135c and provided to a display device capable of displaying a musical score based on the musical score data Sd. In this way, it is possible to remove the performer's performance habits from the input performance data Pd and generate a musical score based on the first tone control data SC1 to reproduce the musical score that the performer was assumed to have seen and played.
[0057] Furthermore, although not shown in Figure 5, the first tone control data SC1 may also be output to the performance control signal generation unit 407 according to the user instruction information UI. When the first tone control data SC1 is provided to the performance control signal generation unit 407, the performance control signal generation unit 407 generates a performance control signal based on the first tone control data SC1 and tempo information and provides it to the electronic instrument 20. The electronic instrument 20 can then reproduce or automatically play a performance sound that recreates the musical score that is assumed to have been viewed and played by the performer for a given piece of music.
[0058] The sound control data (first sound control data SC1, second sound control data SC2) and the musical score data Sd, output from the first trained model 132, the second trained model 133, and the third trained model 134, are associated with each other and stored in the sound control data storage area 135b or the musical score data storage area 135c, respectively. Therefore, for example, if the user instruction information UI includes both the generation of musical score data Sd and the generation of performance control signals, the display of the musical score based on the musical score data Sd and the driving of the performance controls based on the performance control signals can be synchronized with each other.
[0059] [Data processing method] The data processing method performed in the automatic music transcription function 40 will now be described. The data processing method described here begins when program 131 is executed by the control unit 101.
[0060] Figures 6 to 8 illustrate the data processing method according to this embodiment. As shown in Figure 6, the control unit 101 generates performance data Pd based on performance signals sequentially provided from the electronic musical instrument 20 (S101). The control unit 101 inputs the generated performance data Pd into the first trained model 132 (S102) to obtain the first tone control data SC1 (S103). The control unit 101 determines whether or not to generate a musical score and / or perform based on the first tone control data SC1 based on user instruction information UI input in advance via the operation unit 105 (S104). If the user instruction information UI contains information instructing to execute the musical score generation process and / or performance control signal generation process based on the first tone control data SC1 (S104; Yes), the control unit 101 inputs the first tone control data SC1 and parameters based on the first user specified information UD1 into the second trained model 134 (S105) to obtain the second tone control data SC2 (S106).
[0061] If the user instruction information UI does not contain information instructing the execution of a musical score generation process and / or a performance control signal generation process based on the first tone control data SC1 (S104; No), as shown in Figure 7, the control unit 101 inputs the first tone control data SC1 to the third trained model 134 (S107) to obtain musical score data Sd (S108). The control unit 101 also outputs the first tone control data SC1 to the performance control signal generation unit 407 to generate a performance control signal based on the first tone control data SC1 and tempo information (S109). Either the processes in S107-S108 or S109 may be executed, or both may be executed.
[0062] On the other hand, when the second tone control data SC2 is obtained by the processing in S106, as shown in Figure 8, the control unit 101 inputs the second tone control data SC2 to the third learned model 134 (S110) and obtains the musical score data Sd (S111). The control unit 101 also outputs the second tone control data SC2 to the performance control signal generation unit 407 and generates a performance control signal based on the second tone control data SC2 and tempo information (S112). The processing in S110 to S111 and the processing in S112 may be performed individually or both.
[0063] Through the above process, sound control data for generating musical scores and / or automatic performance is automatically generated based on the performance signal, and further processing can be performed according to the user's purpose. The generated sound control data can be used for various purposes.
[0064] The first tone control data SC1 is used for sheet music provision services, instructional services, etc. When generating sheet music based on the first tone control data SC1, it is possible to generate sheet music that is assumed to be what the user was looking at and playing, with the user's individual habits removed. By generating such sheet music, a sheet music provision service can be realized. Furthermore, when performing automatic performance based on the first tone control data SC1, automatic performance based on the sheet music assumed to be what the user was playing can be realized on the electronic instrument 20. Such automatic performance can realize an instructional service. As described above, when performance data Pd based on the performance signal is input to the first trained model 132, the first tone control data SC1 is output. This allows for the automatic acquisition of the first tone control data SC1 corresponding to the sheet music assumed to be what the user was looking at and playing, and enables the provision of customer experiences such as sheet music generation and / or automatic performance based on this first tone control data SC1.
[0065] The second tone control data SC2 is used for sheet music provision services, instructional services, and music generation services. When generating sheet music based on the second tone control data SC2, it is possible to generate sheet music with arrangements desired by the user added to the sheet music assumed to have been played by the user. This type of sheet music generation enables the realization of sheet music provision services. Furthermore, when performing automatic performance based on the second tone control data SC2, the electronic instrument 20 can realize automatic performance based on sheet music with arrangements desired by the user added to the sheet music assumed to have been played by the user. Such automatic performance enables instructional services and music generation services. As described above, when the first tone control data SC1 and the first user-specified information UD1 input by the user are input to the second trained model 133, the second tone control data SC2 is output. As a result, the second tone control data SC2, which has been processed to add arrangements based on the user's desired performance genre, is automatically obtained from the first tone control data SC1, and customer experiences such as sheet music generation, automatic performance, and music generation based on this second tone control data SC2 can be provided.
[0066] Furthermore, when sound control data is used in a training service, for example, the performance data Pd based on the user's performance of the electronic instrument 20 may be compared with the first sound control data SC1 or the second sound control data SC2, and comparison information showing the comparison result may be generated by the data processing device 10. In this case, an image based on the comparison information may be displayed on the display device so that the user can visually confirm the comparison result. This provides a customer experience in which the user can visually confirm the difference between their own performance and the performance based on the first sound control data SC1 or the second sound control data SC2.
[0067] <Second Embodiment> A second embodiment of the present invention will now be described. In the second embodiment, the storage unit of the data processing device includes a fourth trained model. The fourth trained model provides image control data based on sound control data and a desired genre of operation specified by the user. Based on this image control data, the data processing device can generate image data for displaying an XR image. The configuration of the system 1 and the electronic instrument 20 in the second embodiment is the same as in the first embodiment described with reference to Figures 1 and 2, so redundant descriptions will be omitted.
[0068] [Data Processing Device] Figure 9 is a block diagram illustrating the configuration of the data processing device 10A according to the second embodiment. The data output device 10A includes a control unit 101, a storage unit 103A, an operation unit 105, and an interface 107. The other components, excluding the storage unit 103A, are the same as those of the data processing device 10A according to the first embodiment. Therefore, this section will mainly describe the differences between the storage unit 103A and the storage unit 103 in the first embodiment.
[0069] The memory unit 103A stores multiple trained models obtained by machine learning. The trained models stored in the memory unit 103A include a first trained model 132, a second trained model 133, a third trained model 134, and a fourth trained model 136. The fourth trained model 136 is a computational model used when generating motion image data for XR image display based on sound control data and second user-specified information UD2. In this embodiment, the fourth trained model 136 has three computational models. A known machine learning model is applied to each computational model. Known machine learning models include, for example, models using neural networks such as CNNs and RNNs.
[0070] In other words, the fourth trained model 136 is a trained model obtained by machine learning the correlation between sound control data and second user-specified information UD2 and image control data. Here, the sound control data is the first sound control data SC1 output from the first trained model 132 or the second sound control data SC2 output from the second trained model 133. The image control data is data for generating motion image data for XR image display. The image control data is information indicating the motion corresponding to the sound timing included in the sound control data. The second user-specified information UD2 is information indicating the genre of motion desired by the user, specifically, the genre of motion that the user wishes to display in an image. Here, the genre of motion includes, for example, the performer, conductor, or body part desired by the user. For example, if the user desires to display fingerings by a specific performer (for example, a specific pianist), the second user-specified information UD2 includes information indicating the performer desired by the user and information indicating "both hands" as the body part to be displayed. Furthermore, if the user requests to display the posture of a specific performer (for example, a specific pianist) while they are performing, the second user-specified information UD2 will include information indicating the performer the user requests and information indicating "the whole body" as the body part to be displayed.
[0071] [Automatic music transcription function] Figure 10 is a block diagram showing the configuration of the automatic music transcription function 40A in this embodiment. The automatic music transcription function 40A includes a first tone conversion unit 401, a second tone conversion unit 403, a musical score generation unit 405, a performance control signal generation unit 407, and an action image generation unit 409. The first tone conversion unit 401, second tone conversion unit 403, musical score generation unit 405, and performance control signal generation unit 407 in the automatic music transcription function 40A are substantially the same as those described with reference to Figures 4 and 5, so a detailed explanation is omitted. Although not shown, the first tone conversion unit 401 includes a first trained model 132 and a feature extraction unit 323. Also, although not shown, the second tone conversion unit 403 includes a second trained model 133. Although not shown in the diagram, the score generation unit 405 also includes a third trained model 134.
[0072] The motion image generation unit 409 includes a fourth trained model 136 and a motion image data generation unit 364. The fourth trained model 136 includes a fifth encoder 361, a fifth decoder 362, and a sixth encoder 363. Sound control data is provided to the fourth trained model 136. The sound control data input to the fourth trained model 136 is either the first sound control data SC1 output from the first sound conversion unit 401 or the second sound control data SC2 output from the second sound conversion unit 403. Here, the case where the sound control data input to the fourth trained model 136 is the second sound control data SC2 will be explained as an example.
[0073] The sixth encoder 363 of the fourth trained model 136 receives second user-specified information UD2 from the user via the operation unit 105. If the user wishes to display the actions of a predetermined performer, the predetermined performer indicated by the information contained in the first user-specified information UD1 input to the second trained model 133 of the second sound conversion unit 403 may be the same as the predetermined performer indicated by the information contained in the second user-specified information UD2. The sixth encoder 363 outputs a parameter corresponding to the input second user-specified information UD2. This parameter is a vector value. The parameter output from the sixth encoder 363 is input to the fifth decoder 362.
[0074] The second sound control data SC2 is input to the fifth encoder 361. The fifth encoder 361 converts the input second sound control data SC2 into a vector value and outputs it. The vector value output from the fifth encoder 361 is input to the fifth decoder 362.
[0075] The fifth decoder 362 receives the vector value output from the fifth encoder 361 and the parameters corresponding to the second user-specified information UD2 output from the sixth encoder 363. Based on the input second sound control data SC2 and parameters, the sixth decoder 362 generates and outputs image control data. The image control data is information indicating the operation corresponding to the sound timing included in the second sound control data SC2. The image control data output from the fourth trained model 136 is output to the operation image data generation unit 364.
[0076] The motion image data generation unit 364 generates motion image data based on the image control data output from the fourth trained model 136. The motion image data is display data for displaying XR images. XR images are, for example, VR images, MR images, etc. The motion image data generation unit 364 outputs the generated motion image data to a display device or an external goggles, head-mounted display, or other terminal that displays XR images. The terminal is a device capable of sending and receiving data with the data processing device 10, and may be capable of sending and receiving data with the data processing device 10 via a network.
[0077] As described above, the image control data is information indicating the operation corresponding to the sound generation timing included in the second sound control data SC2. Therefore, in this embodiment, when a performance control signal is generated in the performance control signal generation unit 407 based on the second sound control data SC2 and provided to the electronic instrument 20, the sound generation timing and / or the driving timing of the performance operator 201 based on the performance control signal can be synchronized with the display timing of the XR image based on the operation image data.
[0078] The above explanation described the case where the sound control data input to the fourth trained model 136 is the second sound control data SC2 as an example. However, the sound control data input to the fourth trained model 136 may also be the first sound control data SC1. In this case, the image control data is information indicating the operation corresponding to the pronunciation timing included in the first sound control data SC1.
[0079] [Data processing method] The data processing method performed in the automatic music transcription function 40A will now be described. The data processing method described here begins when program 131 is executed by the control unit 101.
[0080] Figures 11 and 12 illustrate the data processing method according to this embodiment. As shown in Figure 11, the control unit 101 generates performance data Pd based on performance signals sequentially provided from the electronic musical instrument 20 (S201). The control unit 101 inputs the generated performance data Pd into the first trained model 132 (S202) to obtain the first sound control data SC1 (S203). The control unit 101 determines whether or not to generate an action image based on the first sound control data SC1 based on user instruction information UI input in advance via the operation unit 105 (S204). If the user instruction information UI includes information instructing to execute the action image generation process based on the first sound control data SC1 (S204; Yes), the control unit 101 inputs the first sound control data SC1 and parameters based on the second user-specified information UD2 into the fourth trained model 136 (S205) to obtain image control data (S206). The control unit 101 inputs image control data to the motion image data generation unit 364 and generates motion image data based on the image control data (S207).
[0081] If the user instruction information UI does not contain information instructing the execution of motion image generation processing based on the first sound control data SC1 (S204; No), as shown in Figure 12, the control unit 101 inputs the first sound control data SC1 and parameters based on the first user-specified information UD1 to the second trained model 133 (S208) to obtain the second sound control data SC2 (S209). Next, the control unit 101 inputs the second sound control data SC1 and parameters based on the second user-specified information UD2 to the fourth trained model 136 (S210) to obtain image control data (S211). The control unit 101 inputs the image control data to the motion image data generation unit 364 to generate motion image data based on the image control data (S207).
[0082] This enables the automatic generation of image control data based on the first control data SC1 or the second sound control data SC2, and the display of an action image based on the image control data.
[0083] The processes S201-S203 shown in Figure 11 correspond to the processes S101-S103 shown in Figure 6. Although not shown in Figures 11 and 12, the processes S205-S206 and S209-S211 shown in Figures 11 and 12 can be executed in parallel with the processes S107-S109 shown in Figure 7, or the processes S110-S112 shown in Figure 8.
[0084] Through the above processing, sound control data for generating musical scores and / or automatic performance can be automatically generated, and further, motion images desired by the user can be generated based on the sound control data. The generated motion images can be used, for example, in a training data service. The motion images may be displayed on a display device or external terminal in synchronization with the musical score or automatic performance based on the first sound control data SC1 or the second sound control data SC2. This allows the user to visually observe the fingering and body movements of a desired performer along with the musical score or automatic performance and use them as a model. As described above, when the first sound control data SC1 or the second sound control data SC2 and the second user-specified information UD2 input from the user are input to the fourth trained model 136, image control data is output. This generates motion images based on the image control data, providing a customer experience in which the user can correct their fingering and posture when performing.
[0085] <Third Embodiment> The following describes the model generation function for generating the first trained model 132, the second trained model 133, the third trained model 134, and the fourth trained model 136 in the first and second embodiments. As described above, the first trained model 132, the second trained model 133, the third trained model 134, and the fourth trained model 136 are obtained by machine learning. Here, the model generation function is realized by the control unit 101 in the data processing devices 10 and 10A implementing a predetermined program. The model generation function is implemented for each of the first trained model 132, the second trained model 133, the third trained model 134, and the fourth trained model 136. The term "teaching data" described below may be replaced with the term "training data." The expression "to train a model" may be replaced with the expression "to train a model." For example, the expression "the computer trains the learning model using teaching data" may be replaced with the expression "the computer trains the learning model using training data." Furthermore, the model generation function may be performed on an external device such as a server that can communicate with the data processing devices 10 and 10A via a network.
[0086] Figure 13 is a diagram illustrating the model generation function 50 for generating the first trained model 132. The model generation function 50 includes a machine learning unit (first trained model learning unit) 501. The machine learning unit 501 is provided with performance data 503 and first tone control data 505. Performance data 503 corresponds to the performance data Pd described above, and first tone control data 505 corresponds to first tone control data SC1. Performance data 503 and first tone control data 505 correspond to training data in machine learning. The machine learning unit 501 uses this training data to perform machine learning and generate the first trained model 132. In other words, the first trained model 132 can also be said to be generated by having the computer train a learning model using training data.
[0087] Figure 14 is a diagram illustrating the model generation function 51 for generating the second trained model 133. The model generation function 51 includes a machine learning unit (second trained model learning unit) 511. The machine learning unit 511 is provided with first sound control data 513, first user-specified information 515, and second sound control data 517. The first sound control data 513 corresponds to the first sound control data SC1 described above, the first user-specified information 515 corresponds to the first user-specified information UD1, and the second sound control data 517 corresponds to the second sound control data SC2. The first sound control data 513, the first user-specified information 515, and the second sound control data 517 correspond to training data in machine learning. The machine learning unit 511 uses this training data to perform machine learning and generate the second trained model 133. In other words, the second trained model 133 can also be said to be generated by having the computer train a learning model using training data.
[0088] Figure 15 is a diagram illustrating the model generation function 52 for generating the third pre-trained model 134. The model generation function 52 includes a machine learning unit (third pre-trained model learning unit) 521. The machine learning unit 521 is provided with sound control data 523 and musical score data 525. The sound control data 523 corresponds to the first sound control data SC1 and second sound control data SC2 mentioned above, and the musical score data 525 corresponds to the musical score data Sd. The sound control data 523 and musical score data 525 correspond to training data in machine learning. The machine learning unit 521 uses this training data to perform machine learning and generate the third pre-trained model 134. In other words, the third pre-trained model 134 can also be said to be generated by having the computer train a learning model using training data.
[0089] Figure 16 is a diagram illustrating the model generation function 53 for generating the fourth pre-trained model 136. The model generation function 53 includes a machine learning unit (fourth pre-trained model learning unit) 531. The machine learning unit 531 is provided with sound control data 533, second user-specified information 535, and image control data 537. The sound control data 533 corresponds to the first sound control data SC1 and second sound control data SC2 described above, the second user-specified information 535 corresponds to the second user-specified information UD2, and the image control data 517 corresponds to information indicating actions corresponding to the pronunciation timing included in the sound control data. The sound control data 533, second user-specified information 535, and image control data 537 correspond to training data in machine learning. The machine learning unit 531 uses this training data to perform machine learning and generate the fourth pre-trained model 136. In other words, the fourth pre-trained model 136 can also be said to be generated by having the computer train a learning model using training data.
[0090] [Generation process for the first tone control data] The following describes the general process for generating the first tone control data SC1 by the first tone conversion unit 401. Figures 17 to 42 are diagrams illustrating the process for generating the first tone control data. Figures 17 to 42 and their descriptions include the content described in U.S. Provisional Application No. 63 / 416,941.
[0091] As described above, the first tone control data SC1 is output in response to the input of the performance data Pd to the first trained model 132. The first tone control data SC1 includes data in which pitch information, note value information, and sound production timing are defined for each note through quantization, and these data are related to each other.
[0092] As shown in Figure 17, the first tone data SC1 obtained by quantizing the performance data (MIDI performance) can be used for various purposes. For example, the first tone data SC1 can be used for processes such as generating musical scores, generating motion image data for displaying motion images, and adding arrangements desired by the user, as described in the first and second embodiments above. The first tone data SC1 may also be used for composition processing using a composition AI, time signature analysis, and so on.
[0093] Figure 18 illustrates quantization. As shown in Figure 14, beat tracking and quantization are similar but different tasks; beat tracking outputs beat positions, while quantization outputs detailed information about the beat (time signature) and each note. There is more research on beat tracking than on quantization, and more research on speech-based beat tracking than on symbolic beat tracking. However, quantization is a fundamental task that directly deals with the basic meaning (meter structure) in symbolic music.
[0094] The process for generating the first tone control data SC1 by the first tone conversion unit 401 will now be described. As described above, the first trained model 132 of the first tone conversion unit 401 includes a first encoder 321 and a first decoder 322. Performance data Pd is input to the first encoder 321 and converted into a vector value.
[0095] As shown in Figure 19, the proposed method addresses the problems of "beat tracking" and "quantization" using a single unified model. Latent chords are learned from the performance and the score. Quantization is predicted from the latent chords, and beat tracking is scored by a contrasting loss. In Figure 19, "Performance" corresponds to the performance data Pd in the embodiment described above, "Perf enc" corresponds to the first encoder 321, "Scoer dec" corresponds to the first decoder 322, and "Score quantization" corresponds to the first tone control data SC1.
[0096] As shown in Figure 20, the first encoder 321 converts the performance data based on the musical score that the performer is presumed to have seen and played into a vector value, Zs + Convert the performance data Pd to a vector value Zp so that it is close to the desired value.
[0097] Figure 21 shows that the implementation features a non-autoregressive and conditional design for the score decoder (MuseBERT) and a multimodal view of symbolic data: an image view and a note sequence view.
[0098] In detail, as shown in Figures 22 and 23, the duration (duration of sound production) and velocity (volume) of each note are converted into vector values based on the performance data. This conversion uses two layers included in the first encoder 321: CNN*2 and Bi-GRU.
[0099] As shown in Figures 24 to 26, the first encoder 321 converts the data based on the musical score that the performer is presumed to have seen and played into a vector value, Zs + , and the value Zs when data different from the sheet music that the performer was presumably looking at and playing is converted into a vector value. - Referencing multiple Zs + The performance data Pd is converted into a vector value Zp so that it is close to the desired value. The formula shown in Figure 27 is used for this purpose. The value of the vector value Zp is adjusted so that the formula shown in Figure 27 equals "=1".
[0100] As shown in Figure 28, the vector value Zp output from the first encoder 321 is input to the first decoder 322. The first decoder 322 is also input with feature information based on the performance data Pd. For example, MuseBERT is used for the first decoder 322.
[0101] Feature information is generated by the feature extraction unit 323 of the first sound conversion unit 401 in the above-described embodiment and is information indicating the sound features contained in the performance data Pd. Feature information includes information indicating pitch (hereinafter, pitch information), information indicating velocity (hereinafter, velocity information), and information indicating the relative relationship between sounds. Pitch information and velocity information are each associated with time information. The information indicating the relative relationship between sounds is, as shown in Figure 29, the onset (R) between a predetermined sound and another sound. o ), duration (R d ), pitch (R p This shows relative relationships such as ).
[0102] As shown in Figure 30, the first decoder 322 outputs first tone control data SC1 in response to the input. The first tone control data SC1 includes pitch information, duration information, and onset (onset; beat and position within that beat, or position within a measure). The pitch information, duration information, and onset are related to each other. The first tone control data SC1 may also include velocity information.
[0103] As shown in Figure 31, the quantization and beat tracking results from the ablation test, the beat tracking interference algorithm, and the self-updating training strategy are described below.
[0104] Figure 32 shows the accuracy of quantization using the model proposed here.
[0105] As shown in Figure 33, the performance of the computational model can be improved by sending the quantization results to a new version of the computational model and providing multiple negative samples until convergence occurs. Figures 34 and 35 show the accuracy of quantization by a model that uses the quantization results as negative samples.
[0106] According to FIG. 36, this model can only apply quantization to music segments of 6 to 10 beats. It is relatively easy to estimate the time range that satisfies this condition, and this model is robust to the selection of the time range.
[0107] According to FIG. 37, if the prediction is incorrect, there may be a large error in the beat estimation. Therefore, it is considered that only the subdivided estimation should be used to perform better beat tracking. It is possible to infer beats from the subdivided ones when the detailed estimation accuracy is medium and prior knowledge such as a smooth tempo curve and a certain tempo range is used. There is no existing method that applies a "sawtooth" waveform as shown in FIG. 37.
[0108] According to FIG. 38, the tempo curve can be estimated by pairwise subdivision. Given a limited tempo range (for example, 30 - 300 bpm), only a finite number of tempos are possible.
[0109] According to FIG. 39, first, the curve of the whole piece is estimated. This is to estimate local subdivision pairs (within 3 seconds) to estimate the BPM candidates. Specifically, p(BPM t / subdiv [t-δ,t+δ] ) is estimated, and a smooth BPM transition, that is, p(BPM t / BPM t-1This involves assuming the curve is approximately diagonal and applying the Viterbi algorithm to estimate the most likely tempo curve. Furthermore, it proposes a reliable beat position. This involves double-checking whether pairs of local subdivisions match the tempo curve, giving high confidence scores to consistent pairs, using these to propose possible beat positions, and merging close beat proposals using a clustering algorithm. Furthermore, it traces the beats from left to right. This involves assuming the previous beat is known, estimating the time range of the next beat based on the tempo curve, deciding on that beat if a beat proposal is strictly within the time range, if the beat proposal is approximately inside the time range and that beat is found consecutively, that beat is the next beat, and otherwise choosing the average of the time range as the next beat position.
[0110] As shown in Figure 40, the beat tracking results can be visualized using a sawtooth waveform based on the confidence score and proposed beat position.
[0111] As shown in Figure 41, future challenges include detailed and complete research on the utility of contrasting losses, detailed and complete research on self-renewal performance, and the possibility of generalizing current methods to downbeat tracking and quantitative hierarchical analysis.
[0112] According to Figure 42, the signal processing method described with reference to Figures 17 to 41 can be understood as consisting of a performance encoder that outputs a performance expression from a performance signal, a first model that estimates the pitch from the performance signal, and a second model that takes the performance expression, pitch, and timing information as input and outputs a quantized pitch. [Explanation of Symbols]
[0113] 1: System, 10, 10A: Data processing unit, 20: Electronic musical instrument, 101: Control unit, 50, 51, 52, 53: Model generation function, 103, 103A: Memory unit, 105: Operation unit, 107: Interface, 131: Program, 132: First trained model, 133: Second trained model, 134: Third trained model, 135: Memory area, 136: Fourth trained model, 201: Performance control unit, 203: Sound source, 205: Speaker, 207: Drive control unit 207, 209: Drive unit, 21 1: Interface, 321: First encoder, 322: First decoder, 323: Feature extraction unit, 331: Second encoder, 332: Second encoder, 333: Third encoder, 341: Fourth encoder, 342: Fourth encoder, 361: Fifth encoder, 362: Fifth decoder, 363: Sixth encoder, 364: Motion image data generation unit, 401: First sound conversion unit, 403: Second sound conversion unit, 405: Score generation unit, 407: Performance control signal generation unit, 409: Motion image generation unit
Claims
1. To obtain first tone control data, including pitch information, note value information, and sound production timing, from a first trained model into which performance data has been input. Inputting the parameters corresponding to the first sound control data and the first user-specified information into the second trained model, To obtain the second sound control data from the second trained model, Includes, The first trained model is a model trained to estimate the first control data using the performance data as input. A data processing method wherein the second trained model is a model trained to estimate the second sound control data by taking the first sound control data and parameters corresponding to the first user-specified information as input.
2. The data processing method according to claim 1, wherein the first user-specified information is information indicating the genre of performance desired by the user.
3. The data processing method according to claim 1, wherein the second trained model is a trained model obtained by machine learning the correlation between the first sound control data and the first user-specified information and the second sound control data.
4. The first trained model is a trained model obtained by machine learning the correlation between the performance data and the first sound control data. The data processing method according to claim 1, wherein the machine learning uses first sound control data containing a correct beat and first sound control data containing an incorrect beat.
5. This further includes extracting features from the performance data and inputting the feature information representing the features into the first trained model, The data processing method according to claim 1, wherein the feature information includes pitch information, velocity information, and information indicating the relative relationship between sounds.
6. To obtain first tone control data, including pitch information, note value information, and sound production timing, from a first trained model into which performance data has been input. Inputting the first sound control data into the third trained model, and To obtain musical score data from the aforementioned third trained model, Includes, The first trained model is a model trained to estimate the first control data using the performance data as input. A data processing method wherein the third trained model is a model trained to estimate the musical score data using the first control data as input.
7. Inputting the aforementioned second sound control data into the third trained model, and To obtain musical score data from the aforementioned third trained model, It further includes, The data processing method according to claim 1, wherein the third trained model is a model trained to estimate the musical score data using the second control data as input.
8. To obtain first tone control data, including pitch information, note value information, and sound production timing, from a first trained model into which performance data has been input. To obtain the desired tempo information, and Based on the first sound control data and the tempo information, a performance control signal is generated. Includes, A data processing method wherein the first trained model is a model trained to estimate the first control data using the performance data as input.
9. To obtain the desired tempo information, and Based on the second sound control data and the tempo information, a performance control signal is generated. The data processing method according to claim 1, further comprising:
10. To obtain first tone control data, including pitch information, note value information, and sound production timing, from a first trained model into which performance data has been input. Inputting parameters corresponding to the first sound control data and the second user-specified information into the fourth trained model, and To obtain image control data corresponding to the pronunciation timing from the fourth trained model, Includes, A data processing method wherein the fourth trained model is a model trained to estimate the image control data using the first sound control data and parameters corresponding to the second user-specified information as input.
11. The data processing method according to claim 10, wherein the second user-specified information is information indicating the genre of operation desired by the user.
12. The data processing method according to claim 10, wherein the fourth trained model is a trained model obtained by machine learning the correlation between the first sound control data and the second user-specified information and the image control data.
13. Inputting the parameters corresponding to the second sound control data and the second user-specified information into the fourth trained model, and To obtain image control data corresponding to the pronunciation timing from the fourth trained model, It further includes, The data processing method according to claim 1, wherein the fourth trained model is a model trained to estimate the image control data using the second sound control data and parameters corresponding to the second user-specified information as input.
14. The data processing method according to claim 1, further comprising comparing the performance data with the second sound control data and generating comparison information indicating the comparison result.
15. The data processing method according to claim 6, further comprising comparing the performance data with the first sound control data and generating comparison information indicating the comparison result.
16. On the computer, A program for causing the data processing method described in any one of claims 1 to 15 to be executed.
17. A system for performing the data processing method described in any one of Claims 1 to 15.
Citation Information
Patent Citations
Device and method for analyzing musical information and recording medium with musical information analyzing program
JP2001290474A
Playing support device and playing support method
JP2001337675A
Automatic transcription of musical content and real-time musical accompaniment
JP2016136251A
Control system and control method
JP2021043258A
Information processing device for musical-score data
WO2020031544A1