Information processing methods, information processing systems, and programs

The information processing system uses machine learning to generate control data for finger movements, addressing the challenge of detailed finger representation in musical performances, achieving realistic and varied finger behaviors.

JP2026048321APending Publication Date: 2026-03-17YAMAHA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-05
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing techniques for controlling the behavior of objects, such as robots performing music, struggle to accurately reproduce the detailed movements of fingers, lacking the ability to vary and realistically represent finger behaviors in response to musical performances.

Method used

An information processing system that utilizes machine learning to generate control data for the movement of fingers by analyzing musical notes, employing a trained generation model to produce detailed finger movements synchronized with music, allowing for adjustments and corrections to ensure natural and varied finger behaviors.

Benefits of technology

The system effectively generates control data that accurately represents the movements of fingers playing musical notes, enabling realistic and varied finger behaviors, enhancing the visual representation of musical performances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026048321000001_ABST
    Figure 2026048321000001_ABST
Patent Text Reader

Abstract

This generates control data to vary the movement of the fingers in accordance with the performance of music. [Solution] The information processing system 100 acquires music data M representing the time series of musical notes, and processes the music data M using a machine learning-trained generative model G to generate control data Z representing the behavior of multiple fingers playing the time series of musical notes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005]

[0001] The present disclosure relates to a technique for controlling the movement of an object representing the behavior of a performer's fingers.

Background Art

[0002] Techniques for controlling the behavior of an object performing a performance such as dancing or playing in accordance with the performance of music have been proposed conventionally. For example, Patent Document 1 discloses a technique in which a robot that has learned the association between music and movement patterns dances in accordance with the music.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the technique of Patent Document 1, control data for controlling the overall behavior of an object is generated by using music data and time-series joint angle parameters. However, there is room for improvement in terms of reproducing partial and detailed behaviors such as the behavior of fingers of an object. In view of the above circumstances, one aspect of the present disclosure aims to generate control data for variously changing the behavior of fingers in accordance with the performance of music.

Means for Solving the Problems

[0005] In order to solve the above problems, an information processing method according to one aspect of the present disclosure is realized by a computer system that acquires music data representing a time series of musical notes and processes the music data by a machine learning-trained generation model to generate control data representing the behavior of a plurality of fingers playing the time series of musical notes. <\

[0006] To solve the above problems, an information processing system according to one aspect of this disclosure comprises an acquisition unit that acquires musical data representing a time series of musical notes, and a generation unit that processes the musical data using a machine learning-based generation model to generate control data representing the behavior of multiple fingers playing a time series of musical notes.

[0007] To solve the above problems, a program according to one aspect of this disclosure causes a computer system to function as an acquisition unit that acquires musical data representing a time series of musical notes, and a generation unit that processes the musical data using a machine learning-prepared generation model to generate control data representing the behavior of the fingers playing the time series of musical notes. [Brief explanation of the drawing]

[0008] [Figure 1] This is a block diagram illustrating the configuration of an information processing system according to the first embodiment. [Figure 2] This is an explanatory diagram of the music data and analysis data. [Figure 3] This is a schematic diagram of a display image by a display device according to the first embodiment. [Figure 4] This is a block diagram illustrating the functional configuration of the control device according to the first embodiment. [Figure 5] This is an explanatory diagram of the control data. [Figure 6] This is a block diagram illustrating the functional configuration of the generation unit according to the first embodiment. [Figure 7] This is a performance matrix representing individual data points. [Figure 8] This is a block diagram illustrating the configuration of the second model. [Figure 9] This is a flowchart illustrating the information processing according to the first embodiment. [Figure 10] This is a diagram illustrating the training process. [Figure 11] This is a flowchart of the training process. [Figure 12] This is a block diagram illustrating the functional configuration of the control device according to the second embodiment. [Figure 13]It is a schematic diagram of associated data according to the second embodiment. [Figure 14] It is an explanatory diagram of a process for correcting control data according to the second embodiment. [Figure 15] It is a flowchart illustrating information processing according to the second embodiment. [Figure 16] It is a block diagram illustrating a functional configuration of a control device according to the third embodiment. [Figure 17] It is a schematic diagram of a display image by a display device according to the third embodiment. [Figure 18] It is a flowchart illustrating information processing according to the third embodiment. [Figure 19] It is a block diagram illustrating a configuration of an information processing system according to the fourth embodiment. [Figure 20] It is a block diagram illustrating a functional configuration of a control device according to the fourth embodiment. [Figure 21] It is a flowchart illustrating information processing according to the fourth embodiment.

Embodiments for Carrying Out the Invention

[0009] A: First Embodiment FIG. 1 is a block diagram illustrating a configuration of an information processing system 100 according to the first embodiment. The information processing system 100 is a computer system that analyzes music. Specifically, the information processing system 100 analyzes the fingering of a music piece and generates a three-dimensional image of the fingers during performance by a virtual performer in a virtual space according to the result of the analysis. Fingering is a method (i.e., finger technique) by which a performer operates each key of a keyboard instrument with each finger of the left and right hands in the performance of a keyboard instrument. That is, information on which finger a performer operates each key of the keyboard instrument with is analyzed as the fingering of the performer and reflected in the generated three-dimensional image of the fingers.

[0010] The information processing system 100 includes an operation device 10, a control device 11, a storage device 12, a display device 13, a sound source device 14, and a sound playback device 15. The information processing system 100 is realized by, for example, a portable information device such as a smartphone or a tablet terminal, or a portable or stationary information device such as a personal computer. The information processing system 100 can be realized not only as a single device but also as a plurality of devices separately configured from each other.

[0011] The operation device 10 is an input device that receives instructions from the user U. The operation device 10 is, for example, an operator operated by the user U or a touch panel that detects contact by the user U. An operation device 10 (for example, a mouse or a keyboard) separate from the information processing system 100 may be connected to the information processing system 100 by wire or wirelessly.

[0012] The control device 11 is composed of one or more processors that control each element of the information processing system 100. For example, the control device 11 is composed of one or more types of processors such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an SPU (Sound Processing Unit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or an ASIC (Application Specific Integrated Circuit).

[0013] The storage device 12 is one or more memories that store programs executed by the control device 11 and various data used by the control device 11. The storage device 12 is composed of, for example, a known recording medium such as a magnetic recording medium or a semiconductor recording medium. The storage device 12 may be composed of a combination of multiple types of recording media. Also, a portable recording medium detachable from the information processing system 100 or a recording medium (for example, cloud storage) on which the control device 11 can perform writing or reading via a communication network may be used as the storage device 12.

[0014] The storage device 12 stores, for example, music data M representing the performance of a musical piece. Figure 2 is an explanatory diagram of the music data M and the performance matrix C, which will be described later. The music data M is time-series data that specifies the pitch and duration of each of the multiple notes that make up the musical piece. Specifically, the music data M is time-series data in which note data that specifies a note and instructs the performance, and time data that specifies the time when each note data is read, are arranged. The note data specifies, for example, the pitch and intensity of the note. The time data specifies, for example, the timing of reading consecutive note data. The music data M is, for example, data in a format compliant with the MIDI (Musical Instrument Digital Interface) standard.

[0015] The display device 13 shown in Figure 1 displays a three-dimensional image in a virtual space under the control of the control device 11. Various types of display panels are envisioned for the display device 13, such as liquid crystal display panels or organic EL (electroluminescence) panels. The display device 13 may be a separate unit from the information processing system 100 and connected to the information processing system 100 by wire or wireless connection.

[0016] The display device 13 of the first embodiment displays a three-dimensional image (hereinafter referred to as "finger object Oa") representing the movement of the fingers in response to playing. Figure 3 shows an example of the display of the finger object Oa. The finger object Oa is a three-dimensional image of both hands in a virtual space. An image representing a virtual keyboard instrument played by the finger object Oa (hereinafter referred to as "instrument object Ob") is also displayed on the display device 13 together with the finger object Oa.

[0017] The sound source device 14 shown in Figure 1 generates an acoustic signal that represents the waveform of the note data specified by the music data M. Alternatively, the control device 11 may execute a program to realize the functions of the sound source device 14.

[0018] The sound-emitting device 15 reproduces sound waves under the control of the control device 11. The sound-emitting device 15 is, for example, an output device such as a speaker or headphones. Specifically, the sound-emitting device 15 reproduces the musical tones of the target piece of music represented by the acoustic signal generated by the sound source device 14. Note that the sound source device 14 or the sound-emitting device 15, which are separate from the information processing system 100, may be connected to the information processing system 100 by wire or wireless connection.

[0019] Figure 4 is a block diagram illustrating the functional configuration of the control device 11. The control device 11 executes a program stored in the storage device 12 in response to instructions from the user U to the operating device 10, thereby realizing multiple functions (acquisition unit 21, generation unit 22, display control unit 23) for generating a finger object Oa. The functions of the control device 11 may be realized by a collection of multiple devices (i.e., a system), or some or all of the functions of the control device 11 may be realized by a dedicated electronic circuit (e.g., a signal processing circuit).

[0020] The acquisition unit 21 acquires music data M from the storage device 12. Specifically, the acquisition unit 21 sequentially reads and outputs each note data constituting the music data M at timings specified by the time data. The acquisition unit 21 outputs each note data of the music data M to the generation unit 22 and the sound source device 14.

[0021] The generation unit 22 generates control data Z from music data M. The control data Z represents the actions of multiple fingers playing the time series of notes represented by the music data M. The generation unit 22 sequentially generates control data Z from the corresponding music data M within the analysis period Q. The analysis period Q is a period that divides the time series of notes represented by the music data M into specific time lengths. That is, the control data Z is generated sequentially for each analysis period Q.

[0022] Figure 5 is an explanatory diagram of control data Z. Control data Z represents the skeleton of the finger object Oa using a plurality of control points 41 and a plurality of connecting parts 42. Each control point 41 is a point that can move in the virtual space, and the connecting parts 42 are straight lines that connect each control point 41 to each other. Each control point 41 corresponds to the position of each joint or fingertip of the finger. The movement of the finger object Oa is controlled by moving each control point 41. Control data Z represents the behavior of multiple fingers playing a time series of musical notes. Therefore, it has the advantage that the behavior of the fingers can be controlled by control data Z regardless of the size of the fingers. Note that the positions and number of control points 41 and connecting parts 42 are arbitrary and are not limited to the above examples.

[0023] The control data Z generated by the generation unit 22 is a vector representing the position of each of the multiple control points 41 in coordinate space. The control data Z represents the coordinates of each control point 41 in a three-dimensional coordinate space where mutually orthogonal Ax and Ay axes, and an Az axis orthogonal to the Ax-Ay plane, are set. In other words, the control data Z indicates the position of each of the multiple control points 41 corresponding to the fingers. For each of the multiple control points 41, a vector is used as the control data Z, which is an array of coordinates on the Ax axis, Ay axis, and Az axis. However, the format of the control data Z is arbitrary. The time series of the control data Z exemplified above represents the movement of the finger object Oa (i.e., the movement of each control point 41 and each connecting part 42 over time).

[0024] The generation unit 22 uses a trained generative model G to generate the control data Z. The generative model G is a statistical model that has learned the relationship between training music data Mt and training control data Zt through machine learning. The generation unit 22 generates the control data Z by processing the music data M with the trained generative model G.

[0025] Figure 6 is a block diagram illustrating the functional configuration of the generation unit 22. The generation unit 22 includes a fingering data generation unit 31, an analysis data generation unit 32, and a control data generation unit 33. The generation model G includes a first model G1 and a second model G2.

[0026] The fingering data generation unit 31 performs a first process to generate fingering data F by processing music data M. Fingering data F is data that specifies a finger number for each note represented by music data M. Therefore, a time series of multiple fingering data F corresponding to different note data specified by music data M is generated. A finger number is information used to identify one of several fingers. For example, finger number "1" is the right thumb, finger number "2" is the right index finger, finger number "3" is the right middle finger, finger number "4" is the right ring finger, finger number "5" is the right little finger, finger number "6" is the left thumb, finger number "7" is the right index finger, finger number "8" is the left middle finger, finger number "9" is the left ring finger, and finger number "10" is the left little finger. The numbers assigned to each finger are called finger numbers. Therefore, for example, in the music data M, the first note "C" is assigned finger number "1", the second note "F" is assigned finger number "4", and so on, with each note being assigned a finger number. However, the fingering data generation unit 31 may not be able to uniquely estimate the finger number corresponding to each note. Notes for which the fingering data generation unit 31 cannot uniquely estimate the finger number are treated as notes with unknown finger numbers.

[0027] The first processing uses the first model G1. The first model G1 is a statistical model that has learned the relationship between music data M and fingering data F through prior machine learning. In other words, the first model G1 generates statistically valid fingering data F for the music data M.

[0028] The first model G1 is implemented by a combination of a program that causes the control device 11 to perform a calculation to generate fingering data F from music data M, and multiple variables (weights and biases) applied to the calculation. The multiple variables are set by machine learning (especially deep learning) using multiple training data T and stored in the storage device 12. The fingering data generation unit 31 generates fingering data F by processing the music data M with the trained first model G1.

[0029] For example, a neural network such as a Transformer, which is an encoder-decoder model that includes a self-attention mechanism (specifically, a multi-head attention mechanism), can be used as the first model G1. The Transformer is disclosed, for example, in Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin in "Attention Is All You Need," 31st Conference on Neural Information Processing Systems (NIPS 2017). However, the configuration of the first model G1 is arbitrary and not limited to the examples above.

[0030] The analysis data generation unit 32 performs analysis data generation processing to generate analysis data P according to the music data M and fingering data F. The analysis data P represents the relationship between the time series of notes represented by the music data M and the fingerings represented by the fingering data F. Specifically, the analysis data generation unit 32 sequentially acquires the music data M and the fingering data F and generates analysis data P corresponding to the portion of the music data M within the analysis period Q. That is, analysis data P is generated sequentially for each analysis period Q. Note that the analysis data generation processing is an example of the "third processing".

[0031] The first process of generating fingering data F from music data M and the analysis data generation process of generating analysis data P from music data M and fingering data F are separate processes. That is, fingering data F is created while the generation unit 22 is generating control data Z from music data M. Therefore, compared to a configuration in which control data Z is generated from music data M in a series of processes, the user U can easily correct the fingering tendency in the analysis data P. The fingering tendency refers to the bias in the finger numbers assigned to each of the notes of the same pitch among multiple notes. Specifically, the user U can correct finger movements that are impossible in actual performance by correcting the fingering tendency through operations on the operating device 10. In addition, the user U can also change the finger numbers when the finger movements differ due to differences in playing techniques.

[0032] The analysis data P includes one overall data point Pa and multiple individual data points Pb[n] (where n is a natural number).

[0033] Each of the multiple individual data points Pb[n] corresponds to a different finger. Specifically, each of the multiple individual data points Pb[n] represents the time series of notes in the music data M corresponding to the portion within the analysis period Q, where the finger number "n" corresponding to that individual data point Pb[n] is specified by the fingering data F. Notes for which the finger number is unknown are not included in any of the multiple individual data points Pb[n]. Therefore, the fingering tendencies of each finger can be reflected in the control data Z from the multiple individual data points Pb[n]. Furthermore, when user U modifies the fingering tendencies of each finger, they can do so by replacing the individual data point Pb[n] to be modified with another individual data point Pb[n]. In other words, user U can partially modify the analysis data P.

[0034] The comprehensive data Pa represents the time series of all notes corresponding to the portion of the music data M within the analysis period Q. That is, notes with unknown finger numbers are also included in the comprehensive data Pa. The comprehensive data Pa accurately reflects the musical piece represented by the music data M. Therefore, it can reflect not only the fingering tendencies of individual fingers but also the overall tendencies of the piece in the control data Z.

[0035] Each of the single overall data point Pa and the multiple individual data points Pb[n] is represented by a performance matrix C.

[0036] The performance matrix C shown in Figure 2 represents an I-row, J-column matrix (where I and J are natural numbers). The performance matrix C is a binary matrix representing the time series of one overall data Pa or multiple individual data Pb[n] sequentially output by the analysis data generation unit 32. The horizontal axis of the performance matrix C corresponds to the time axis. Any column of the performance matrix C corresponds to one of the J (e.g., 60) unit periods included in the analysis period Q. The vertical axis of the performance matrix C corresponds to the pitch axis. Any row of the performance matrix C corresponds to one of the I (e.g., 88) pitches. One element in the i-th row and j-th column (i=1 to I, j=1 to J) of the performance matrix C indicates whether or not the pitch corresponding to the i-th row is pronounced during the unit period corresponding to the j-th column. Specifically, among the J elements in the i-th row corresponding to any given pitch, the element corresponding to each unit period in which that pitch is played is set to "1", and the element corresponding to each unit period in which that pitch is not played is set to "0". The performance matrix C shown in Figure 2 represents the time series of all notes in the analysis period Q, and therefore represents the overall data Pa from the analysis data P.

[0037] Figure 7 shows the performance matrix C representing individual data Pb[n]. Individual data Pb[n] corresponding to finger number "n" represents the performance by the finger with finger number "n". For example, individual data Pb[1] for finger number "1" is the performance matrix C representing the time series of notes played by the right thumb. The same applies to other finger numbers.

[0038] The control data generation unit 33 in Figure 6 performs a second process to generate control data Z for controlling the movement of the finger object Oa from the analysis data P. The control data Z is generated sequentially for each analysis period Q. Specifically, the control data Z for any one analysis period Q is generated from the analysis data P of that analysis period Q.

[0039] The analysis data generation process, which generates analysis data P from music data M and fingering data F, and the second process, which generates control data Z from analysis data P, are separate processes. That is, analysis data P is created while the generation unit 22 is generating control data Z from music data M. Therefore, compared to a configuration in which control data Z is generated from music data M in a series of processes, it is easy to pre-create analysis data P or replace the analysis data P that is the target of the second process.

[0040] The second processing step utilizes the second model G2. The second model G2 is a statistical model that has learned the relationship between the analysis data P and the control data Z through prior machine learning. In other words, the second model G2 generates statistically valid control data Z for the analysis data P.

[0041] The second model G2 is implemented by a combination of a program that causes the control device 11 to perform an operation to generate control data Z from analysis data P, and multiple variables (weights and biases) applied to the operation. The multiple variables are set by machine learning (especially deep learning) using multiple training data T and stored in the memory device 12. Specifically, multiple variables defining the first neural network G2a and multiple variables defining the second neural network G2b are set collectively by machine learning using multiple training data T. The control data generation unit 33 generates control data Z by processing the analysis data P with the trained second model G2. Figure 8 is a block diagram illustrating the configuration of the second model G2. The second model G2 is configured by connecting the first neural network G2a and the second neural network G2b in series.

[0042] The first neural network G2a generates a feature vector V that represents the features of the analysis data P. For example, a temporal convolutional neural network (TCN), which is suitable for feature extraction, is used as the first neural network G2a.

[0043] The second neural network G2b generates control data Z corresponding to the feature vector V. For example, a recurrent neural network (RNN) containing a Long Short-Term Memory (LSTM) unit suitable for processing time-series data can be used as the second neural network G2b. As illustrated above, the combination of the first neural network G2a and the second neural network G2b can generate appropriate control data Z corresponding to the time series of the analysis data P. However, the configuration of the second model G2 is arbitrary and is not limited to the examples above.

[0044] The display control unit 23 shown in Figure 4 generates a finger object Oa from the control data Z and displays the finger object Oa and the instrument object Ob shown in Figure 3 on the display device 13. The display control unit 23 shown in Figure 4 dynamically changes the finger object Oa in parallel with the playback of the performance by the sound emission device 15. Specifically, the display device 13 moves each control point 41 to the coordinates specified by the control data Z over time. That is, the display device 13 displays the behavior of the finger object Oa within the analysis period Q. Through this control, the finger object Oa performs the performance actions within the analysis period Q. Therefore, users U (e.g., spectators) who view the displayed image on the display device 13 can visually and intuitively grasp the behavior of the fingers playing the time series of musical notes represented by the music data M.

[0045] Figure 9 is a flowchart illustrating the specific steps of the performance display processing performed by the information processing system 100. The performance display processing is executed for each analysis period Q on the time axis.

[0046] When the performance display process is started, the control device 11 (acquisition unit 21) acquires music data M from the storage device 12 (S1). Specifically, the control device 11 (acquisition unit 21) outputs the note data corresponding to the time series of notes from the music data M to the control device 11 (generation unit 22) and the sound source device 14.

[0047] When music data M is acquired, the control device 11 (generation unit 22) processes the music data M using the trained first model G1 to generate fingering data F (S2). The control device 11 (generation unit 22) performs analysis data generation processing on the music data M and fingering data F to generate analysis data P corresponding to the portion of the music data M within the analysis period Q (S3). As described above, the analysis data P includes multiple individual data Pb[n] and a combined data Pa. The control device 11 (generation unit 22) processes the analysis data P using the trained second model G2 to generate control data Z (S4). As described above, the second model G2 includes the first neural network G2a and the second neural network G2b.

[0048] The control device 11 (display control unit 23) generates a finger object Oa according to the control data Z (S5). The control device 11 (display control unit 23) updates the finger object Oa displayed on the display device 13 (S6). The updated finger object Oa performs the performance operation within the analysis period Q. In parallel with the performance operation of the finger object Oa, the sound source device 14 causes the sound output device 15 to play the music represented by the music data M. In other words, the operation of the finger object Oa and the playback of the music are synchronized.

[0049] The control device 11 determines whether the termination condition is met (S7). The termination condition is, for example, that the user U has instructed the user to terminate the process by operating the control device 10, or that processing of the entire music data M has been completed. If it is determined that the termination condition is not met (S7: No), the playback display process moves to step S1, and for the subsequent analysis period Q, the processing from the acquisition of the music data M onward (S1 to S7) is repeated. If it is determined that the termination condition is met (S7: Yes), the playback display process terminates.

[0050] Machine learning will be explained using the second model G2 as an example. Figure 10 is an explanatory diagram of the process of establishing the second model G2 using machine learning (hereinafter referred to as the "training process"). The control device 11, by executing the program stored in the storage device 12, functions not only as the elements exemplified in Figure 4 (acquisition unit 21, generation unit 22, display control unit 23), but also as the training processing unit 50 in Figure 10. The training processing unit 50 establishes the second model G2 using machine learning with multiple training data T.

[0051] Multiple training data sets T are stored in the memory device 12. Each of the multiple training data sets T consists of a combination of training analysis data Pt and training control data Zt. Each training data set T is training data in which the training analysis data Pt and training control data Zt are correlated with each other. The training processing unit 50 establishes a second model G2 by machine learning using the multiple training data sets T.

[0052] Figure 11 is a flowchart of the training processing unit 50. For example, the training process is started in response to instructions from the user U to the operating device 10.

[0053] When the training process begins, the control device 11 (training processing unit 50) selects one of the multiple training data T stored in the memory device 12 (hereinafter referred to as "selected training data T") (Sb1). The control device 11 (training processing unit 50) iteratively updates multiple variables of the initial or provisional second model G2 (hereinafter referred to as "provisional second model G2p") using the selected training data T (Sb2~Sb4).

[0054] The control device 11 generates control data Z by processing the analysis data P of the selected training data T using the provisional second model G2p (Sb2). The control device 11 calculates a loss function that represents the error between the control data Z generated by the provisional second model G2p and the control data Z of the selected training data T (Sb3). The control device 11 updates several variables of the provisional second model G2p so that the loss function is reduced (ideally minimized) (Sb4). For example, backpropagation is used to update each variable according to the loss function.

[0055] The control device 11 determines whether a predetermined termination condition has been met (Sb5). The termination condition is, for example, that the loss function falls below a predetermined threshold, or that the amount of change in the loss function falls below a predetermined threshold. If the termination condition is not met (Sb5: NO), the control device 11 selects the unselected training data T stored in the memory device 12 as the new selected training data T (Sb1). That is, the process of updating multiple variables of the provisional second model G2p (Sb2~Sb4) is repeated until the termination condition is met (Sb5: YES). If the termination condition is met (Sb5: YES), the control device 11 terminates the training process. The provisional second model G2p at the time the termination condition is met is finalized as the trained second model G2.

[0056] As can be understood from the above explanation, the second model G2 learns the latent relationship between the analysis data Pt and the control data Zt in multiple training data sets T. Therefore, the trained second model G2 outputs statistically valid control data Z for unknown analysis data P under these relationships.

[0057] As described above, in the first embodiment, the trained second model G2 is used to generate the control data Z. Therefore, it is possible to generate control data Z that reflects the trends present in the multiple training data T used to train the second model G2.

[0058] Although the above explanation focuses on training the second model G2, the first model G1 is also trained using the same procedure as shown in the example in Figure 11. Multiple training data T, including training music data Mt and training fingering data Ft, are used to train the first model G1. Specifically, the control device 11 (training processing unit 50) updates multiple variables of the first model G1 so that the error between the fingering data F generated by the initial or provisional first model G1p from the music data M of each training data T and the training fingering data F included in the training data T is reduced (ideally minimized).

[0059] As explained above, the information processing system 100 generates control data Z by inputting music data M into a machine learning-trained generative model G. Under the latent relationship between music data M and control data Z in the multiple training datasets T used for machine learning, a variety of control data Z representing multiple reasonable finger movements for unknown music data M is generated. In other words, it can generate control data Z that allows for diverse changes in finger movements in response to music performance.

[0060] B: Second Embodiment A second embodiment of this disclosure will now be described. For elements whose function is the same as in the first embodiment in each of the embodiments described below, the same reference numerals as in the first embodiment will be used, and detailed descriptions of each will be omitted as appropriate.

[0061] Figure 12 is a block diagram illustrating the functional configuration of the control device 11 of the second embodiment. The control device 11 of the second embodiment functions as a modification unit 24 in addition to having elements similar to those of the first embodiment.

[0062] The modification unit 24 modifies the control data Z. Specifically, the modification unit 24 performs processing to make the behavior of multiple fingers represented by the control data Z closer to natural behavior. A specific example of the processing by which the modification unit 24 modifies the control data Z is shown below. Two or more processes arbitrarily selected from the following examples may be merged as appropriate, within a range that does not contradict each other.

[0063] (1) One of the processes by which the modification unit 24 modifies the control data Z is to adjust the position of the fingertip of the finger corresponding to the note that the music data M instructs to play, so that it contacts the key of the keyboard instrument that corresponds to that note. Specifically, the position of the control point 41 corresponding to the fingertip of the finger number specified for the note that instructs to play is fixed to the surface of the key of the keyboard instrument that corresponds to that note, for at least a portion of the duration of the note's sound production. That is, in the finger object Oa and instrument object Ob displayed on the display device 13, the fingertip of the finger with the finger number specified for the note being played and the key corresponding to that note come into contact for at least a portion of the duration of the note's sound production. Therefore, the behavior of the fingers represented by the control data Z can be made to more closely resemble the behavior of fingers actually playing a keyboard instrument.

[0064] (2) One of the processes by which the modification unit 24 modifies the control data Z is to modify the length between multiple control points 41 (i.e., the total length of each connecting part 42) to be constant. "The length between multiple control points 41 is constant" means that the movement of each control point 41 is restricted so that the distance between two control points 41 connected to each other in the connecting part 42 is maintained within a predetermined range. The predetermined range is a range in which a user U (e.g., an audience member) observing the finger object Oa cannot perceive the expansion and contraction of the fingers. In other words, the total length of each connecting part 42 does not expand or contract unnaturally due to the movement of the fingers. Also, the total length of each connecting part may differ for each combination of two control points 41. Therefore, the length of the fingers in the finger object Oa displayed on the display device 13 does not change unnaturally due to the movement of the fingers during performance. As explained above, there is an advantage in that the shape of the multiple fingers represented by the control data Z remains stable regardless of the behavior.

[0065] (3) One of the processes by which the modification unit 24 modifies the control data Z is to link the finger corresponding to the musical note that the music data M instructs to be played, with the fingers adjacent to that finger. In human behavior, it is difficult to move only one finger, and even if one does not intend to move one finger, the fingers adjacent to the moved finger may move in conjunction. The ease with which adjacent fingers can move in conjunction differs for each combination of fingers.

[0066] The process of modifying control data Z so that adjacent fingers move in conjunction uses linkage data B. Linkage data B is stored in the memory device 12. Specifically, linkage data B is stored for both the right hand and the left hand. Figure 13 is a schematic diagram of linkage data B. Linkage data B defines an index (hereinafter referred to as "linkage index R(r,s)") that represents the degree to which the movement of the fingertip r of one finger affects the fingertip s of the other finger, for all possible combinations of selecting two adjacent fingers from the five fingers.

[0067] The modification unit 24 uses the linkage data B to link adjacent fingers. Specifically, the modification unit 24 identifies a linkage index R(r,s) from the linkage data B that corresponds to the combination of the fingertip r of the finger specified by the fingering data F (hereinafter referred to as the "playing finger") and the fingertip s of another finger adjacent to the playing finger. The modification unit 24 calculates the movement distance d·R(r,s) of fingertip s by multiplying the linkage index R(r,s) identified from the linkage data B by the movement distance d of the playing finger. The movement distance d of the fingertip r of the playing finger is identified from the control data Z before modification. Then, the modification unit 24 modifies the control data Z so that fingertip s moves by a distance d·R(r,s) in the same direction as fingertip r of the playing finger. In addition, one linkage data B may be shared for the movement of mutually corresponding fingers between the right hand and the left hand (for example, the index finger of the right hand and the index finger of the left hand).

[0068] Figure 14 is an explanatory diagram of the process for modifying the control data Z. When the fingertip r of a playing finger moves d [cm] among multiple fingers, the modification unit 24 modifies the fingertip s of the other fingers adjacent to the playing finger by d·R(r,s) [cm]. Therefore, the natural behavior of fingers during actual performance, where other fingers adjacent to a specific finger move in conjunction with that finger, can be faithfully reproduced by modifying the control data Z. Note that "adjacent fingers" can be multiple fingers as long as they are adjacent to the finger corresponding to a musical note.

[0069] (4) One of the processes by which the modification unit 24 modifies the control data Z is to modify the multiple fingers represented in the control data Z to match the physical information, according to the physical information that represents the requirements of the virtual performer's body. "Physical information" refers to physical characteristics such as the size of the body parts of the virtual performer or the range of motion of each joint. The user U inputs the physical information using the operating device 10, and the modification unit 24 reflects the physical information in the control data Z.

[0070] The movement range of each control point 41 and the total length of each connecting part 42 represented by the control data Z before modification are set to standard values ​​corresponding to the standard physical characteristics or physique of a virtual performer. As physical information, the user U specifies one of several stages relative to the standard physical characteristics or physique. The modification unit 24 increases or decreases the standard values ​​represented by the control data Z according to the physical information. For example, if the physical information is modified to a larger physique, the modification unit 24 increases the total length of each connecting part 42 from the standard value. As a result of the increase in the total length of the connecting part 42, the fingers represented by the finger object Oa become longer. Also, if the physical information is modified to a stiffer physique, the modification unit 24 decreases the movement range of each control point 41 from the standard value. As a result of the decrease in the movement range of the control points 41, the range of motion of the fingers represented by the finger object Oa becomes narrower.

[0071] As explained above, the process of modifying the control data Z according to the physical information reflects the physical characteristics represented by the physical information in the finger object Oa by modifying the range of movement of each control point 41 or the total length of each connection. Therefore, when user U uses their own physical information to modify the control data Z, user U can observe the behavior of the finger object Oa in performance that reflects their own physical information, which has the advantage of making it easier to use as a reference for their own performance.

[0072] (5) Various data processing methods other than those exemplified above may be employed as the processing by the modification unit 24 to modify the control data Z. One example of data processing performed by the modification unit 24 is a filtering process that mitigates the discontinuous behavior of the multiple fingers represented by the control data Z. Filtering to mitigate discontinuous behavior smooths the movement of each control point 41 and each connection part 42 over time. An example of filtering to mitigate discontinuous behavior is a One-euro filter. As a result, the awkward behavior of the finger object Oa becomes smooth, and the behavior of the finger object Oa approaches natural behavior. However, the processing to modify the control data Z is not limited to the examples above.

[0073] Figure 15 is a flowchart illustrating the specific procedure for the performance display processing in the second embodiment. In the performance display processing of the second embodiment, the modification of control data Z (S8) is added to the processing similar to that of the first embodiment. Specifically, the control device 11 (modification unit 24) modifies the control data Z generated by the control device 11 (generation unit 22) as described above (S8). The operation of elements other than the modification unit 24 is the same as in the first embodiment.

[0074] The same effects as in the first embodiment are achieved in the second embodiment. Furthermore, the control data Z modified by the modification unit 24 represents more natural finger movements compared to a configuration in which the control data Z is not modified.

[0075] C: Third Embodiment Figure 16 is a block diagram illustrating the functional configuration of the control device 11 according to the third embodiment. The control device 11 of the third embodiment has the same elements as the second embodiment, plus an additional whole-body generation unit 25. However, in the third embodiment, modification of the control data Z is not required.

[0076] The whole-body generation unit 25 generates performer data W that represents the behavior of the performer's body (whole body or upper body). Specifically, the whole-body generation unit 25 generates performer data W using analysis data P. For generating performer data W, any known performance analysis method, such as that described in Japanese Patent Application Publication No. 2019-139295, can be arbitrarily employed.

[0077] Figure 17 is a three-dimensional image (hereinafter referred to as "performer object Oc") representing the upper body of a virtual performer in a virtual space, including multiple fingers, arms, chest, and head. The display control unit 23 displays the performer object Oc and the instrument object Ob on the display device 13. The display control unit 23 generates the performer object Oc by linking the arms represented by performer data W and the hands represented by control data Z. Specifically, the display control unit 23 links the hands represented by control data Z so that the direction of the hands is aligned with the direction of the elbows of the arms represented by performer data W. In addition, the right and left hands move in conjunction with the swaying of the body represented by performer data W. Note that known data processing can be arbitrarily employed to link performer data W and control data Z. For example, it is assumed that IK (Inverse Kinematics) will be used to link performer data W and control data Z. However, the processing for linking performer data W and control data Z is not limited to the above examples.

[0078] Figure 18 is a flowchart illustrating the specific procedure of the performance display processing in the third embodiment. In the performance display processing of the third embodiment, the generation of performer data W (S9) and the synthesis of control data Z and performer data W (S10) are added to the processing similar to that of the second embodiment. Specifically, the control device 11 (whole body generation unit 25) performs the process of generating performer data W using analysis data P (S9). Then, the control device 11 (display control unit 23) synthesizes the control data Z and performer data W (S10). The operation of elements other than the whole body generation unit 25 and the display control unit 23 is the same as in the second embodiment. Note that step S9 may be performed in parallel with step S4 or step S8, or the order of execution may be changed.

[0079] As explained above, by combining control data Z and performer data W, the performer's body movements can be reproduced in a three-dimensional image. Therefore, the user U can visually and intuitively grasp both the overall or general movements of the performer's body and the partial or detailed movements of the performer's fingers.

[0080] D: Fourth Embodiment Figure 19 is a block diagram illustrating the configuration of the information processing system 100 according to the fourth embodiment. The information processing system 100 of the fourth embodiment has the same elements as the third embodiment, with the addition of a sound collection device 16. However, in the fourth embodiment, modification of the control data Z or generation of performer data W is not required.

[0081] The sound pickup device 16 is a device that picks up sounds (e.g., instrument sounds or singing sounds) produced by user U's performance. Specifically, the sound pickup device 16 is a microphone that picks up the sounds of the keyboard instrument 200 played by user U. The sound pickup device 16 generates an acoustic signal E representing the waveform of the sound from the sounds produced by user U's performance. Note that user U playing the keyboard instrument 200 and user U operating the control device 10 may be different people. Furthermore, although the example illustrates a configuration in which the acoustic signal E is generated by picking up the performance sounds produced by the keyboard instrument 200, a configuration in which the acoustic signal E is generated by playing an electric instrument that generates an acoustic signal from performance is also possible. Therefore, the sound pickup device 16 may be omitted.

[0082] Figure 20 is a block diagram illustrating the functional configuration of the control device 11 according to the fourth embodiment. The acquisition unit 21 estimates the time when user U is actually performing within the music (hereinafter referred to as "performance time") by analyzing the acoustic signal E. The estimation of the performance time is performed sequentially in parallel with the actual performance by user U. For the estimation of the performance time, known acoustic analysis techniques (score alignment), such as those described in Japanese Patent Application Publication No. 2015-79183, can be arbitrarily employed.

[0083] The acquisition unit 21 controls the automatic performance by the sound emission device 15 and the behavior of the performer object Oc on the display device 13 to synchronize with the progress of the performance time. Specifically, each time the performance time reaches a point specified by each time data of the music data M, the acquisition unit 21 outputs the note data corresponding to that time data to the sound source device 14 and the generation unit 22. Therefore, the progress of the automatic performance by the sound emission device 15 and the progress of the behavior of the performer object Oc displayed on the display device 13 are synchronized with the actual performance by the user U.

[0084] Figure 21 is a flowchart illustrating the specific procedure for the performance display processing in the fourth embodiment. In the performance display processing of the fourth embodiment, for example, acquisition of an acoustic signal E (S11) and analysis of the acoustic signal E (S12) are added to the processing similar to that of the third embodiment. Specifically, when the user U starts playing, the control device 11 (acquisition unit 21) acquires an acoustic signal E from the sound pickup device 16 (S11). The control device 11 (acquisition unit 21) analyzes the acoustic signal E and estimates the performance time (S12). Specifically, the control device 11 (acquisition unit 21) outputs the note data of the notes corresponding to the performance time from the music data M to the sound source device 14 and the control device 11 (generation unit 22). The operation of elements other than the acquisition unit 21 is the same as in the third embodiment.

[0085] As described above, the automatic performance by the sound emission device 15 and the progress of the performer object Oc's behavior as shown by the display device 13 are synchronized with the actual performance by the user U. Therefore, an atmosphere is created as if the sound emission device 15, the display device 13, and the user U are coordinating with each other to perform together.

[0086] E: Variation The following are examples of specific modifications that may be added to each of the embodiments exemplified above. Two or more embodiments may be arbitrarily selected from the following examples and merged as appropriate, provided they do not contradict each other.

[0087] (1) In the second embodiment, a configuration is shown in which the modification unit 24 performs the process of modifying the control data Z according to the physical information. However, in each embodiment, the processing is not limited to the modification unit 24, as long as the physical information can be reflected in the finger object Oa (or performer object Oc). For example, a configuration in which the generation unit 22 in each embodiment performs the process of generating the control data Z according to the physical information, or a configuration in the third embodiment in which the display control unit 23 performs the process of adjusting the control data Z (or performer data W) according to the physical information when combining the control data Z and performer data W is conceivable.

[0088] (2) In the third embodiment, a configuration in which analysis data P is used to generate performer data W is shown as an example, but a configuration in which only the overall data Pa from the analysis data P is used to generate performer data W is also possible. However, the configuration in which analysis data P is used to generate performer data W has the advantage that the positions of the right hand and left hand can be reflected in the performer data W by using multiple individual data Pb[n] from the analysis data P.

[0089] (3) In the fourth embodiment, a configuration is shown in which the performance time is estimated by analyzing the acoustic signal E of the sound picked up by the sound pickup device 16, but the performance time may also be estimated by analyzing the performance data transmitted from a MIDI instrument (e.g., an electronic keyboard instrument).

[0090] (4) In each embodiment, a binary matrix representing the time series of notes in each analysis period Q of one comprehensive data Pa or multiple individual data Pb[n] is exemplified as the performance matrix C, but the performance matrix C is not limited to the above examples. For example, a performance matrix C representing the performance intensity (volume) of notes in each analysis period Q of one comprehensive data Pa or multiple individual data Pb[n] may be generated. Specifically, one element in the i-th row and j-th column of the performance matrix C represents the intensity in which the pitch corresponding to the i-th row is played in the unit period corresponding to the j-th column. With the above configuration, since the performance intensity of each note is reflected in the control data Z, it is possible to impose on the movement of the finger object Oa (or performer object Oc) a tendency for the performer's movements to differ depending on the strength of the performance intensity.

[0091] (5) In each embodiment, a configuration in which control data Z is generated at each predetermined length of analysis period Q is illustrated, but the temporal unit of the control data Z is arbitrary. For example, a configuration in which control data Z is generated in advance for the entire music data M, and then the generated control data Z is used to display an object on the display device 13 is conceivable.

[0092] (6) In each embodiment, an example was given in which the control data Z is stored in the storage device 12, but configurations in which the control data Z is transmitted to another device, or recorded on a portable recording medium, etc., are also conceivable.

[0093] (7) In each embodiment, the instrument displayed on the display device 13 is exemplified as a keyboard instrument, but the instrument displayed on the display device 13 is not limited to a keyboard instrument. For example, wind instruments such as saxophones and flutes, string instruments such as guitars, etc. are also conceivable. Furthermore, this disclosure applies to various instruments that include performance controls operated by the performer's fingers. Performance controls are parts of an instrument that are operated by the performer's fingers, such as the keys on a keyboard instrument. The shape or configuration of performance controls differs depending on the type of instrument. For example, keys on wind instruments and strings on string instruments are exemplified as performance controls. The same applies to the keyboard instrument 200 played by user U in the fourth embodiment.

[0094] (8) In each embodiment, a finger number was given as an example of information (finger information) that identifies any of the multiple fingers in the fingering data F, but the finger information is not limited to a number as long as it can identify multiple fingers. For example, the finger information may consist of a character corresponding to each finger, or a combination of a character and a number.

[0095] (9) For example, the information processing system 100 may be implemented by a server device that communicates with an information device such as a smartphone or tablet terminal. For example, the information processing system 100 generates control data Z using music data M received from the information device and transmits the control data Z to the information device.

[0096] In the configuration in which analysis data P is transmitted from the information device to the information processing system 100 (i.e., in the configuration in which the analysis data generation unit 32 is mounted on the information device), the analysis data generation unit 32 may be omitted from the information processing system 100. Also, in the configuration in which control data Z is transmitted from the information device to the information processing system 100 (i.e., in the configuration in which the generation unit 22 is mounted on the information device), the generation unit 22 may be omitted from the information processing system 100.

[0097] (10) The functions of the information processing system 100 exemplified above are realized through the cooperation of one or more processors constituting the control device 11 and a program stored in the storage device 12, as described above. The program according to this disclosure can be provided in a form stored on a computer-readable recording medium and installed on a computer. The recording medium is, for example, a non-transitory recording medium, such as an optical recording medium (optical disc) like a CD-ROM, but also includes any known form of recording medium such as a semiconductor recording medium or a magnetic recording medium. A non-transitory recording medium includes any recording medium except for transient propagation signals (transitory, propagating signals), and volatile recording media are not excluded. Furthermore, in a configuration in which a distribution device distributes a program via a communication network, the storage medium in which the distribution device stores the program corresponds to the non-transitory recording medium described above.

[0098] (11) The notation "nth" (where n is a natural number) in this application is used solely as a formal and convenient label to distinguish each element in notation and has no substantive meaning whatsoever. Therefore, there is no room for restrictive interpretation of the position or manufacturing order of each element based on the notation "nth".

[0099] F: Note From the forms exemplified above, the following configuration can be understood, for example.

[0100] An information processing method according to one aspect of this disclosure (Aspect 1) is realized by a computer system that acquires musical data representing a time series of musical notes and processes the musical data using a machine learning-trained generative model to generate control data representing the behavior of multiple fingers playing the time series of musical notes. In this aspect, since control data is generated by inputting musical data into a machine learning-trained generative model, a variety of control data representing appropriate behavior of multiple fingers for unknown musical data is generated under the latent relationship between musical data and control data in the multiple training data used for machine learning. In other words, it is possible to generate control data that allows for diverse changes in finger behavior in response to musical performance.

[0101] In a specific example of Embodiment 1 (Embodiment 2), the generation model includes a first model and a second model, and the generation of the control data includes a first process of generating fingering data that specifies finger information for each note by processing the music data with the first model, and a second process of generating the control data by processing analysis data representing the relationship between the time series of notes represented by the music data and the fingerings represented by the fingering data with the second model. In the above embodiment, the first process of generating fingering data from music data and the second process of generating control data from analysis data are performed separately. Therefore, the first model and the second model can be trained individually by machine learning. In addition, either the first model or the second model can be selectively modified (adjusted or replaced, etc.). Furthermore, it is possible to perform processes such as partial modification of the fingering data generated by the first process, or replacement with different fingering data.

[0102] In a specific example of Embodiment 1 or Embodiment 2 (Embodiment 3), the generation of the analysis data includes a third process that generates the analysis data according to the music data and the fingering data generated by the first process, and in the second process, the analysis data generated by the third process is processed by the second model. In the above embodiments, the generation of analysis data is performed by the third process. Therefore, before performing the generation of analysis data by the third process, it is possible to perform processes such as partial modification of the fingering data or replacement with different fingering data.

[0103] In any specific example of Embodiments 2 to 3 (Embodiment 4), the analysis data includes a plurality of individual data corresponding to different fingers, and each of the plurality of individual data represents the time series of notes in the time series of notes represented by the music data, where the finger information of the finger corresponding to the individual data is specified by the fingering data. In the above embodiments, control data that reflects the fingering tendencies of each finger can be generated. Furthermore, when correcting the fingering tendencies of individual data, the correction can be made by replacing the individual data to be corrected with another individual data from among the plurality of individual data. Note that "finger information" refers to information for identifying any of the multiple fingers (e.g., finger number).

[0104] In any specific example of Embodiments 2 to 4 (Embodiment 5), the analysis data further includes comprehensive data representing the time series of notes represented by the music data. In the above embodiments, control data can be generated that reflects not only the fingering tendencies of each different finger but also the overall tendencies of the piece. Furthermore, since the music data includes notes for which fingering could not be estimated, control data that more accurately reflects the piece can be generated compared to using only individual data. Note that "more accurately reflects the piece" means that there are no omissions in the notes represented by the music data.

[0105] In any specific example of Embodiments 2 to 5 (Embodiment 6), the second model includes a first neural network and a second neural network, and in the second process, The first neural network processes the analysis data to generate feature vectors representing the characteristics of the analysis data, and the second neural network processes the feature vectors to generate control data corresponding to the feature vectors. In this embodiment, since the trained model includes a combination of the first neural network and the second neural network, it is possible to generate appropriate control data corresponding to the music data.

[0106] In any specific example of Embodiments 1 to 6 (Embodiment 7), the control data indicates the position of each of the multiple control points corresponding to the fingers. According to the above embodiments, there is an advantage that the behavior of the fingers can be controlled by the control data regardless of the size of the fingers.

[0107] In any specific example of Embodiments 1 to 7 (Embodiment 8), the control data is further modified. In the above embodiments, the movement of the fingers becomes closer to natural movement compared to a configuration in which the control data is not modified.

[0108] In a specific example of Embodiment 8 (Embodiment 9), the modification of the control data includes a process of adjusting the fingertip of the finger corresponding to the note that the music data instructs to be played, among the plurality of fingers, to a position that contacts the performance control corresponding to that note among the plurality of performance controls of the instrument. In the above embodiments, the control data is modified so that the fingertip of the finger corresponding to the note that is instructed to be played contacts the performance control corresponding to that note. Therefore, the behavior of the plurality of fingers represented by the control data can be made closer to the behavior of fingers actually playing the instrument. Note that "performance control" refers to the part of an instrument that is operated by the fingers when playing. For example, this could include keys in a keyboard instrument such as a piano, keys in a wind instrument such as a saxophone or flute, or strings in a string instrument such as a guitar. However, performance control is not limited to the above examples.

[0109] In a specific example of Embodiment 8 or Embodiment 9 (Embodiment 10), the modification of the control data includes a process of linking the finger corresponding to the note that the music data instructs to be played with the fingers adjacent to that finger. In the embodiments described above, the natural behavior of multiple fingers during actual performance, where other fingers adjacent to a particular finger move in conjunction with that finger, can be faithfully reproduced by the control data. Note that "adjacent fingers" may refer to multiple fingers as long as they are adjacent to the finger corresponding to the note.

[0110] In any specific example of embodiments 8 to 10 (embodiment 11), the modification of the control data includes a process of modifying the length between a plurality of control points corresponding to the fingers to be constant. In the above embodiments, the control data is modified so that the length of the fingers does not change depending on the behavior. Therefore, there is an advantage that the shape of the fingers represented by the control data is stable.

[0111] In a specific example of Embodiment 8 or Embodiment 11 (Embodiment 12), the modification of the control data includes a process of modifying the multiple fingers represented by the control data to match the physical information representing the physical requirements of a virtual performer. In the embodiments described above, a user can generate control data that reflects their own or another person's physical information. Therefore, a user with a body similar to the physical information used for modification has the advantage of being able to easily use the behavior of the multiple fingers represented by the control data as a reference for playing. "Physical information" refers to physical characteristics such as the size of body parts or the range of motion of each joint.

[0112] In any specific example of Embodiments 1 to 12 (Embodiment 13), an image representing the performer's fingers is further generated according to the control data. In the above embodiments, an image representing the movement of the fingers can be generated using the control data. Therefore, a user who views the image can visually and intuitively grasp the movement of the fingers playing the time series of notes represented by the music data.

[0113] In a specific example of Embodiment 13 (Embodiment 14), performer data representing the performer's physical movements is further generated, and the image generation includes generating an image including the performer's fingers and body by combining the performer data and the control data. In the above embodiments, the generated physical movements of the performer and the generated wrist movements of the performer can be combined to reproduce the performer's entire body movements in an image. Therefore, users can visually and intuitively grasp both the overall or general movements of the performer's body and the partial or detailed movements of the performer's fingers.

[0114] In a specific example of Embodiment 13 or Embodiment 14 (Embodiment 15), the control data is generated by processing analysis data that represents the relationship between the time series of notes represented by the music data and the fingerings corresponding to the time series of notes, and the performer data is generated using the analysis data. In the above embodiments, the analysis data is reused for both the generation of control data and the generation of performer data. Therefore, compared to forms in which separate data is used for the generation of control data and performer data, the processing load for generating control data and performer data is reduced. Furthermore, since common analysis data is used for both the generation of control data and the generation of performer data, a sense of unity can be maintained between the behavior of multiple fingers represented by the control data and the behavior of the performer's body represented by the performer data.

[0115] In any specific example of embodiments 13 to 15 (embodiment 16), the analysis data includes, for each of the multiple fingers, individual data specifying the time series of notes to be played by that finger from the time series of notes represented by the music data, and comprehensive data representing the time series of notes represented by the music data. In generating the performer data, the comprehensive data from the analysis data is used to generate the performer data. In the above embodiments, the comprehensive data is reused for generating the control data and the performer data. Therefore, compared to the form in which individual data is used for generating the control data and the performer data, the processing load for generating the control data and the performer data is reduced. Furthermore, since a common comprehensive data is used for generating the control data and the performer data, a sense of unity can be maintained between the behavior of the multiple fingers represented by the control data and the behavior of the performer's body represented by the performer data.

[0116] An information processing system according to one aspect of this disclosure (Aspect 17) comprises an acquisition unit that acquires musical data representing a time series of musical notes, and a generation unit that processes the musical data using a machine learning-trained generation model to generate control data representing the behavior of multiple fingers playing the time series of musical notes. In this aspect, since control data is generated by inputting musical data into a machine learning-trained generation model, a variety of control data representing appropriate behavior of multiple fingers for unknown musical data is generated under the latent relationship between musical data and control data in the multiple training data used for machine learning. In other words, it is possible to generate control data that allows for diverse changes in finger behavior in response to musical performance.

[0117] A program according to one aspect of this disclosure (Aspect 18) causes a computer system to function as an acquisition unit that acquires musical data representing a time series of musical notes, and a generation unit that processes the musical data using a machine learning-trained generation model to generate control data representing the behavior of the fingers playing the time series of musical notes. In this aspect, since control data is generated by inputting musical data into a machine learning-trained generation model, a variety of control data representing appropriate finger behaviors for unknown musical data is generated under the latent relationship between musical data and control data in the multiple training datasets used for machine learning. In other words, it is possible to generate control data that allows for diverse changes in finger behavior according to the performance of music. [Explanation of Symbols]

[0118] 10...Memory device, 11...Control device, 12...Display device, 13...Sound source device, 14...Sound emission device, 15...Operation device, 16...Sound collection device, 21...Acquisition unit, 22...Generation unit, 23...Display control unit, 24...Modification unit, 25...Whole body generation unit, 31...Fingering data generation unit, 32...Analysis data generation unit, 33...Control data generation unit, 41...Control point, 42...Connection unit, 50...Training processing unit, 100...Information processing system, 200...Keyboard instrument, C...Performance matrix, E...Acoustic signal, F...Fingering data, Ft...Fingering data for training, G...Generated model, G1...First model G1p…Provisional Model 1, G2…Model 2, G2a…First Neural Network, G2b…Second Neural Network, G2p…Provisional Model 2, M…Music Data, Mt…Training Music Data, P…Analysis Data, P1…Overall Data, P2…Individual Data, Pt…Training Analysis Data, Q…Analysis Period, R(r,s)…Coordination Index, r…Fingertip of Playing Finger, s…Other Fingertips Adjacent to Playing Finger, T…Training Data, U…User, V…Feature Vector, W…Performer Data, Z…Control Data, Zt…Training Control Data.

Claims

1. Obtain music data representing the time series of musical notes, By processing the music data using a machine learning-based generative model, control data representing the behavior of multiple fingers playing the time series of notes is generated. Information processing methods implemented by computer systems.

2. The aforementioned generation model includes a first model and a second model, The generation of the aforementioned control data is The first process involves processing the music data using the first model to generate fingering data that specifies finger information for each note, The process includes a second process which generates the control data by processing analysis data representing the relationship between the time series of notes represented by the music data and the fingerings represented by the fingering data using the second model. The information processing method of claim 1.

3. The generation of the aforementioned analysis data is The process includes a third process that generates the analysis data in accordance with the music data and the fingering data generated by the first process, In the second process, the analysis data generated by the third process is processed by the second model. The information processing method of claim 2.

4. The aforementioned analysis data includes multiple individual data corresponding to different fingers, Each of the aforementioned individual data represents the time series of notes in the musical data, where the finger information of the hand corresponding to that individual data represents the time series of notes specified by the fingering data. The information processing method of claim 2 or claim 3.

5. The aforementioned analysis data further includes comprehensive data representing the time series of notes represented by the music data. The information processing method of claim 4.

6. The second model includes a first neural network and a second neural network, In the second process described above, By processing the analysis data with the first neural network, a feature vector representing the characteristics of the analysis data is generated. The second neural network processes the feature vectors to generate the control data corresponding to the feature vectors. The information processing method of claim 2.

7. The control data indicates the position of each of the multiple control points corresponding to the fingers. The information processing method of claim 1.

8. Furthermore, the control data is modified. The information processing method of claim 1.

9. The modification of the control data includes a process of adjusting the fingertip of the finger corresponding to the note that the music data instructs to produce sound among the plurality of fingers to a position that contacts the performance control of the instrument that corresponds to that note among the plurality of performance controls of the instrument. The information processing method of claim 8.

10. The modification of the control data includes a process of linking the finger corresponding to the note that the music data instructs to produce sound among the multiple fingers with the adjacent fingers. The information processing method of claim 8 or claim 9.

11. The modification of the control data includes a process of correcting the length between multiple control points corresponding to the fingers to a constant value. The information processing method of claim 8.

12. The modification of the control data includes a process of modifying a plurality of fingers represented by the control data to match the physical information that represents the physical requirements of a virtual performer. The information processing method of claim 8.

13. Furthermore, an image representing the performer's fingers is generated according to the control data. The information processing method of claim 1.

14. Furthermore, performer data representing the physical movements of the performers is generated, The generation of the aforementioned image includes generating an image that includes the performer's fingers and body by combining the performer data and the control data. The information processing method of claim 13.

15. In generating the aforementioned control data, The control data is generated by processing analysis data that represents the relationship between the time series of notes represented by the aforementioned music data and the fingerings corresponding to the time series of notes. In generating the performer data, the performer data is generated using the analysis data. The information processing method of claim 14.

16. The aforementioned analysis data is, For each of the multiple fingers, there is individual data specifying the time series of notes to be played by that finger from the time series of notes represented by the aforementioned musical data, This includes comprehensive data representing the time series of musical notes represented by the aforementioned music data, In generating the performer data, the performer data is generated using the comprehensive data from the analysis data. The information processing method of claim 15.

17. An acquisition unit that acquires music data representing the time series of musical notes, A generation unit processes the music data using a machine learning-based generative model to generate control data that represents the actions of multiple fingers playing the time series of notes. An information processing system equipped with the following features.

18. An acquisition unit that acquires music data representing the time series of musical notes, and A generation unit processes the music data using a machine learning-based generative model to generate control data that represents the movement of the fingers playing the time series of the musical notes. A program that makes a computer system function.

Citation Information

Patent Citations

  • System and method for teaching movement to leg type robot

    JP2002086378A