Information processing method, information processing system, and program

The information processing system uses a machine-learned generative model to generate control data for finger movements in response to music, addressing the challenge of detailed finger behavior reproduction by employing neural networks and deep learning, achieving realistic and natural finger movements.

WO2026053772A1PCT designated stage Publication Date: 2026-03-12YAMAHA CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately control the detailed behavior of an object's fingers in response to music performance, particularly in reproducing the nuanced movements of fingers playing musical instruments.

Method used

An information processing system that utilizes a machine-learned generative model to generate control data for the behavior of multiple fingers based on music data, employing a combination of neural networks and deep learning to analyze and generate fingering data and control data, allowing for the creation of realistic finger movements.

Benefits of technology

The system effectively generates control data that accurately represents the behavior of fingers playing musical notes, enabling a three-dimensional display of finger movements that closely mimic human performance, with the ability to correct and adjust for natural finger interplay and physical characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025029651_12032026_PF_FP_ABST
    Figure JP2025029651_12032026_PF_FP_ABST
Patent Text Reader

Abstract

An information processing system 100 acquires music data M representing a time series of notes, and processes the music data M by using a machine-trained generation model G, thereby generating control data Z representing the behavior of a plurality of fingers that will play the time series of notes.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing method, information processing system, and program

[0001] The present disclosure relates to a technique for controlling the movement of an object that represents the behavior of a player's fingers.

[0002] Techniques have been proposed for controlling the behavior of an object performing a dance, musical instrument, or other performance in accordance with the performance of music. For example, Patent Literature 1 discloses a technique in which a robot that has learned the correspondence between music and movement patterns dances to the music.

[0003] JP 2002-86378

[0004] The technology of Patent Document 1 generates control data for the behavior of the entire body of an object by using music data and time-series joint angle parameters. However, there is room for improvement in terms of reproducing detailed partial behavior, such as the behavior of fingers of an object. In consideration of the above circumstances, one aspect of the present disclosure aims to generate control data for changing the behavior of fingers in a variety of ways in response to the performance of music.

[0005] In order to solve the above problems, an information processing method according to one aspect of the present disclosure is realized by a computer system that acquires music data representing a time sequence of musical notes, processes the music data using a machine-learned generative model, and generates control data representing the behavior of multiple fingers playing the time sequence of musical notes.

[0006] In order to solve the above problems, an information processing system according to one aspect of the present disclosure includes an acquisition unit that acquires music data representing a time sequence of musical notes, and a generation unit that processes the music data using a machine-learned generative model to generate control data representing the behavior of multiple fingers playing the time sequence of musical notes.

[0007] In order to solve the above problems, a program according to one aspect of the present disclosure causes a computer system to function as an acquisition unit that acquires music data representing a time sequence of musical notes, and a generation unit that processes the music data using a machine-learned generative model to generate control data representing the behavior of fingers playing the time sequence of musical notes.

[0008] 1 is a block diagram illustrating the configuration of an information processing system according to a first embodiment; FIG. 2 is an explanatory diagram of music data and analysis data; FIG. 3 is a schematic diagram of a display image by a display device according to the first embodiment; FIG. 4 is a block diagram illustrating the functional configuration of a control device according to the first embodiment; FIG. 5 is an explanatory diagram of control data; FIG. 6 is a block diagram illustrating the functional configuration of a generation unit according to the first embodiment; FIG. 7 is a performance matrix representing individual data; FIG. 8 is a block diagram illustrating the configuration of a second model; FIG. 9 is a flowchart illustrating information processing according to the first embodiment; FIG. 10 is an explanatory diagram of training processing; FIG. 11 is a flowchart of training processing; FIG. 11 is a block diagram illustrating the functional configuration of a control device according to a second embodiment; FIG. 12 is a schematic diagram of linkage data according to the second embodiment; FIG. 13 is an explanatory diagram of processing for correcting control data according to the second embodiment; FIG. 14 is a flowchart illustrating information processing according to the second embodiment; FIG. 15 is a block diagram illustrating the functional configuration of a control device according to a third embodiment; FIG. 16 is a schematic diagram of a display image by a display device according to the third embodiment; FIG. 17 is a flowchart illustrating information processing according to the third embodiment; FIG. 18 is a block diagram illustrating the configuration of an information processing system according to a fourth embodiment; FIG. 19 is a block diagram illustrating the functional configuration of a control device according to the fourth embodiment;

[0009] A: First Embodiment FIG. 1 is a block diagram illustrating the configuration of an information processing system 100 according to a first embodiment. The information processing system 100 is a computer system that analyzes music. Specifically, the information processing system 100 analyzes the fingering of a piece of music and generates, in a virtual space, a three-dimensional image of the fingers of a virtual performer performing the piece of music based on the analysis results. Fingering is the method by which a performer operates each key of a keyboard instrument using the fingers of their left and right hands. In other words, information about which fingers the performer uses to operate each key of the keyboard instrument is analyzed as the performer's fingering and is reflected in the generated three-dimensional image of the fingers.

[0010] The information processing system 100 includes an operation device 10, a control device 11, a storage device 12, a display device 13, a sound source device 14, and a sound emission device 15. The information processing system 100 is realized by a portable information device such as a smartphone or a tablet terminal, or a portable or stationary information device such as a personal computer. The information processing system 100 may be realized as a single device, or may be realized by multiple devices configured separately from each other.

[0011] The operation device 10 is an input device that receives instructions from the user U. The operation device 10 is, for example, an operator operated by the user U, or a touch panel that detects contact by the user U. Note that an operation device 10 (for example, a mouse or a keyboard) separate from the information processing system 100 may be connected to the information processing system 100 by wire or wirelessly.

[0012] The control device 11 is configured with one or more processors that control each element of the information processing system 100. For example, the control device 11 is configured with one or more types of processors, such as a central processing unit (CPU), a graphics processing unit (GPU), a sound processing unit (SPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or an application specific integrated circuit (ASIC).

[0013] The storage device 12 is one or more memories that store programs executed by the control device 11 and various data used by the control device 11. The storage device 12 is configured with a known storage medium such as a magnetic storage medium or a semiconductor storage medium. The storage device 12 may also be configured with a combination of multiple types of storage media. Furthermore, the storage device 12 may be a portable storage medium that is detachable from the information processing system 100, or a storage medium (e.g., cloud storage) that the control device 11 can write to or read from via a communication network.

[0014] The storage device 12 stores, for example, music data M representing the performance of a piece of music. FIG. 2 is an explanatory diagram of the music data M and a performance matrix C, which will be described later. The music data M is time-series data that specifies the pitch and sound duration of each of the multiple notes that make up a piece of music. Specifically, the music data M is time-series data in which note data that specifies a note and instructs its performance, and time data that specifies the time point at which each note data is read, are arranged. The note data specifies, for example, the pitch and intensity of the note. The time data specifies, for example, the timing at which consecutive note data is read. The music data M is data in a format that complies with the MIDI (Musical Instrument Digital Interface) standard.

[0015] 1 displays a three-dimensional image in a virtual space under the control of the control device 11. As the type of the display device 13, various display panels such as a liquid crystal display panel or an organic EL (Electroluminescence) panel are assumed. Note that the display device 13, which is separate from the information processing system 100, may be connected to the information processing system 100 by wire or wirelessly.

[0016] The display device 13 of the first embodiment displays a three-dimensional image (hereinafter referred to as a "finger object Oa") representing the behavior of fingers in response to a performance. Fig. 3 shows an example of the display of the finger object Oa. The finger object Oa is a three-dimensional image of both hands in a virtual space. An image (hereinafter referred to as a "musical instrument object Ob") representing a virtual keyboard instrument played by the finger object Oa is also displayed on the display device 13 together with the finger object Oa.

[0017] 1 generates an acoustic signal representing the waveform of note data specified by the music data M. Note that the function of the sound source device 14 may be realized by the control device 11 executing a program.

[0018] The sound emitting device 15 reproduces sound waves under the control of the control device 11. The sound emitting device 15 is an output device such as a speaker or headphones. Specifically, the sound emitting device 15 reproduces the musical tones of the target music piece represented by the acoustic signal generated by the sound source device 14. Note that the sound source device 14 or the sound emitting device 15, which are separate from the information processing system 100, may be connected to the information processing system 100 via a wired or wireless connection.

[0019] 4 is a block diagram illustrating an example of the functional configuration of the control device 11. The control device 11 executes a program stored in the storage device 12 in response to an instruction from the user U to the operation device 10, thereby realizing multiple functions (an acquisition unit 21, a generation unit 22, and a display control unit 23) for generating the hand object Oa. Note that the functions of the control device 11 may be realized by a set of multiple devices (i.e., a system), or some or all of the functions of the control device 11 may be realized by a dedicated electronic circuit (e.g., a signal processing circuit).

[0020] The acquisition unit 21 acquires the music data M from the storage device 12. Specifically, the acquisition unit 21 sequentially reads and outputs each piece of note data constituting the music data M at a timing specified by the time data. The acquisition unit 21 outputs each piece of note data of the music data M to the generation unit 22 and the sound source device 14.

[0021] The generation unit 22 generates control data Z from music data M. The control data Z represents the behavior of multiple fingers playing the time series of notes represented by the music data M. The generation unit 22 sequentially generates the control data Z from the corresponding music data M within an analysis period Q. The analysis period Q is a period that divides the time series of notes represented by the music data M into specific time lengths. In other words, the control data Z is sequentially generated for each analysis period Q.

[0022] FIG. 5 is an explanatory diagram of the control data Z. The control data Z represents the skeleton of the finger object Oa using a plurality of control points 41 and a plurality of connectors 42. Each control point 41 is a point that can be moved in virtual space, and the connectors 42 are lines connecting the control points 41 to each other. Each control point 41 corresponds to the position of each joint or fingertip of a finger. The movement of the finger object Oa is controlled by moving each control point 41. The control data Z represents the behavior of multiple fingers that play a time sequence of musical notes. Therefore, there is an advantage that the behavior of the fingers can be controlled by the control data Z regardless of the size of the fingers. Note that the positions and number of the control points 41 and connectors 42 are arbitrary and are not limited to the above example.

[0023] The control data Z generated by the generation unit 22 is a vector representing the position of each of the multiple control points 41 in the coordinate space. The control data Z represents the coordinates of each control point 41 in a three-dimensional coordinate space in which the Ax-axis and Ay-axis, which are orthogonal to each other, and the Az-axis, which is orthogonal to the Ax-Ay plane, are set. That is, the control data Z represents the position of each of the multiple control points 41 corresponding to the fingers. A vector in which the coordinate on the Ax-axis, the coordinate on the Ay-axis, and the coordinate on the Az-axis are arranged for each of the multiple control points 41 is used as the control data Z. However, the format of the control data Z is arbitrary. The time series of the control data Z exemplified above represents the movement of the finger object Oa (i.e., the movement of each control point 41 and each connection portion 42 over time).

[0024] A trained generative model G is used by the generation unit 22 to generate the control data Z. The generative model G is a statistical model that has learned the relationship between the training music data Mt and the training control data Zt through machine learning. The generation unit 22 generates the control data Z by processing the music data M using the trained generative model G.

[0025] 6 is a block diagram illustrating an example of the functional configuration of the generation unit 22. The generation unit 22 includes a fingering data generation unit 31, an analysis data generation unit 32, and a control data generation unit 33. The generation model G includes a first model G1 and a second model G2.

[0026] The fingering data generating unit 31 performs a first process of generating fingering data F by processing the music data M. The fingering data F is data that specifies a finger number for each note represented by the music data M. Therefore, a time series of multiple fingering data F corresponding to different note data specified by the music data M is generated. The finger number is information for identifying one of multiple fingers. For example, finger number "1" is the thumb of the right hand, finger number "2" is the index finger of the right hand, finger number "3" is the middle finger of the right hand, finger number "4" is the ring finger of the right hand, finger number "5" is the little finger of the right hand, finger number "6" is the thumb of the left hand, finger number "7" is the index finger of the right hand, finger number "8" is the middle finger of the left hand, finger number "9" is the ring finger of the left hand, and finger number "10" is the little finger of the left hand. The finger numbers are assigned to each finger. Therefore, for example, a fingering number is assigned to each note among the multiple notes in the music data M, such that the first note "C" is assigned fingering number "1," the second note "Fa" is assigned fingering number "4," and so on. However, the fingering data generation unit 31 may not be able to uniquely estimate the fingering number corresponding to each note during processing. Notes for which the fingering data generation unit 31 cannot uniquely estimate the fingering number are treated as notes with unknown fingering numbers.

[0027] The first process utilizes a first model G1. The first model G1 is a statistical model that has learned the relationship between the music data M and the fingering data F through prior machine learning. In other words, the first model G1 generates fingering data F that is statistically valid for the music data M.

[0028] The first model G1 is realized by a combination of a program that causes the control device 11 to execute a calculation to generate fingering data F from music data M, and multiple variables (weights and biases) that are applied to the calculation. The multiple variables are set by machine learning (particularly deep learning) using multiple training data T and are stored in the storage device 12. The fingering data generation unit 31 generates the fingering data F by processing the music data M with the trained first model G1.

[0029] For example, a neural network such as a Transformer, which is an encoder-decoder model including a self-attention mechanism (specifically, a multi-head attention mechanism), is used as the first model G1. The Transformer is disclosed, for example, in "Attention Is All You Need," by Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, 31st Conference on Neural Information Processing Systems (NIPS 2017). However, the configuration of the first model G1 is arbitrary and is not limited to the above examples.

[0030] The analysis data generation unit 32 performs an analysis data generation process that generates analysis data P in accordance with the music data M and the fingering data F. The analysis data P represents the relationship between the time series of notes represented by the music data M and the fingering represented by the fingering data F. Specifically, the analysis data generation unit 32 sequentially acquires the music data M and the fingering data F, and generates analysis data P that corresponds to the portion of the music data M within the analysis period Q. In other words, the analysis data P is generated sequentially for each analysis period Q. The analysis data generation process is an example of a "third process."

[0031] The first process of generating fingering data F from music data M and the analysis data generation process of generating analysis data P from music data M and fingering data F are separate processes. That is, the fingering data F is created while the generation unit 22 is generating control data Z from music data M. Therefore, compared to a configuration in which control data Z is generated from music data M through a series of processes, the user U can easily correct the fingering tendency in the analysis data P. The fingering tendency refers to the bias in the finger numbers assigned to notes of the same pitch among multiple notes. Specifically, by correcting the fingering tendency through operations on the operation device 10, the user U can correct fingering behavior that is impossible in actual performance. The user U can also change the fingering numbers to match fingering behaviors that differ depending on the playing technique.

[0032] The analysis data P includes one piece of comprehensive data Pa and a plurality of pieces of individual data Pb[n] (n is a natural number).

[0033] Each of the multiple individual data Pb[n] corresponds to a different finger. Specifically, each of the multiple individual data Pb[n] represents a time series of notes, among the time series of notes corresponding to the portion of the analysis period Q in the music data M, for which the finger number "n" corresponding to the individual data Pb[n] is specified by the fingering data F. It should be noted that notes with unknown finger numbers are not included in any of the multiple individual data Pb[n]. Therefore, the fingering tendency of each finger can be reflected in the control data Z from the multiple individual data Pb[n]. Furthermore, when user U modifies the fingering tendency of each finger, he or she can do so by replacing the individual data Pb[n] to be modified with another individual data Pb[n]. In other words, user U can partially change the analysis data P.

[0034] The comprehensive data Pa represents the time series of all notes in the music data M that correspond to the portion of the analysis period Q. In other words, the comprehensive data Pa also includes notes with unknown fingering numbers. The comprehensive data Pa accurately reflects the music represented by the music data M. Therefore, the control data Z can reflect not only the fingering tendencies of each of the different fingers, but also the overall tendency of the music.

[0035] The one piece of general data Pa and each of the plurality of individual data Pb[n] are represented by a performance matrix C.

[0036] The performance matrix C shown in FIG. 2 is a matrix with I rows and J columns (I and J are natural numbers). The performance matrix C is a binary matrix representing the time series of one comprehensive data Pa or each of multiple individual data Pb[n] sequentially output by the analysis data generation unit 32. The horizontal direction of the performance matrix C corresponds to the time axis. Any one column of the performance matrix C corresponds to one of the J (e.g., 60) unit periods included in the analysis period Q. The vertical direction of the performance matrix C corresponds to the pitch axis. Any one row of the performance matrix C corresponds to one of the I (e.g., 88) pitches. An element in the ith row and jth column (i = 1 to I, j = 1 to J) of the performance matrix C indicates whether the pitch corresponding to the ith row is sounded in the unit period corresponding to the jth column. Specifically, of the J elements in the i-th row corresponding to a given pitch, the element corresponding to each unit period in which that pitch is sounded is set to "1," and the element corresponding to each unit period in which that pitch is not sounded is set to "0." The performance matrix C shown in Figure 2 represents the time series of all notes in the analysis period Q, and therefore represents the overall data Pa of the analysis data P.

[0037] 7 shows a performance matrix C representing the individual data Pb[n]. The individual data Pb[n] corresponding to finger number "n" is data representing a performance by the finger with finger number "n." For example, the individual data Pb[1] for finger number "1" is a performance matrix C representing the time sequence of notes played by the thumb of the right hand. The same applies to the other finger numbers.

[0038] 6 performs a second process of generating control data Z for controlling the movement of the hand object Oa from the analysis data P. The control data Z is generated sequentially for each analysis period Q. Specifically, the control data Z for any one analysis period Q is generated from the analysis data P for that analysis period Q.

[0039] The analysis data generation process for generating analysis data P from music data M and fingering data F and the second process for generating control data Z from analysis data P are separate processes. That is, the analysis data P is created while the generation unit 22 is generating control data Z from music data M. Therefore, compared to a configuration in which control data Z is generated from music data M through a series of processes, processes such as creating analysis data P in advance and replacing the analysis data P that is the target of the second process can be easily performed.

[0040] The second process utilizes a second model G2. The second model G2 is a statistical model that has learned the relationship between the analysis data P and the control data Z through prior machine learning. In other words, the second model G2 generates control data Z that is statistically appropriate for the analysis data P.

[0041] The second model G2 is realized by a combination of a program that causes the control device 11 to execute a calculation to generate control data Z from the analysis data P and multiple variables (weights and biases) applied to the calculation. The multiple variables are set by machine learning (particularly deep learning) using multiple training data T and stored in the storage device 12. Specifically, multiple variables that define the first neural network G2a and multiple variables that define the second neural network G2b are set collectively by machine learning using the multiple training data T. The control data generation unit 33 generates the control data Z by processing the analysis data P using the trained second model G2. FIG. 8 is a block diagram illustrating an example of the configuration of the second model G2. The second model G2 has a configuration in which the first neural network G2a and the second neural network G2b are connected in series.

[0042] The first neural network G2a generates a feature vector V that represents the features of the analysis data P. For example, a temporal convolutional neural network (TCN), which is suitable for extracting features, is used as the first neural network G2a.

[0043] The second neural network G2b generates control data Z according to the feature vector V. For example, a recurrent neural network (RNN) including a long short-term memory (LSTM) unit suitable for processing time-series data is used as the second neural network G2b. As illustrated above, by combining the first neural network G2a and the second neural network G2b, appropriate control data Z according to the time series of the analysis data P can be generated. However, the configuration of the second model G2 is arbitrary and is not limited to the above example.

[0044] The display control unit 23 shown in FIG. 4 generates a finger object Oa from the control data Z and displays the finger object Oa and musical instrument object Ob shown in FIG. 3 on the display device 13. The display control unit 23 shown in FIG. 4 dynamically changes the finger object Oa in parallel with the reproduction of the performance by the sound emitting device 15. Specifically, the display device 13 moves each control point 41 over time to coordinates specified by the control data Z. That is, the display device 13 displays the behavior of the finger object Oa within the analysis period Q. Through the above control, the finger object Oa executes a performance action within the analysis period Q. Therefore, a user U (e.g., an audience member) viewing the image displayed by the display device 13 can visually and intuitively grasp the behavior of the fingers playing the time series of musical notes represented by the music data M.

[0045] 9 is a flowchart illustrating a specific procedure of the performance display process performed by the information processing system 100. The performance display process is executed for each analysis period Q on the time axis.

[0046] When the performance display process is started, the control device 11 (acquisition unit 21) acquires music data M from the storage device 12 (S1). Specifically, the control device 11 (acquisition unit 21) outputs note data corresponding to the time series of notes in the music data M to the control device 11 (generation unit 22) and the sound source device 14.

[0047] Upon acquiring the music data M, the control device 11 (generation unit 22) processes the music data M using a trained first model G1 to generate fingering data F (S2). The control device 11 (generation unit 22) performs analysis data generation processing on the music data M and fingering data F to generate analysis data P corresponding to the portion of the music data M within the analysis period Q (S3). As described above, the analysis data P includes multiple individual data Pb[n] and comprehensive data Pa. The control device 11 (generation unit 22) processes the analysis data P using a trained second model G2 to generate control data Z (S4). As described above, the second model G2 includes a first neural network G2a and a second neural network G2b.

[0048] The control device 11 (display control unit 23) generates a hand object Oa in accordance with the control data Z (S5). The control device 11 (display control unit 23) updates the hand object Oa displayed on the display device 13 (S6). The updated hand object Oa performs a performance action within the analysis period Q. In parallel with the performance action of the hand object Oa, the sound source device 14 causes the sound output device 15 to play a piece of music represented by the music data M. In other words, the action of the hand object Oa and the playback of the music are synchronized.

[0049] The control device 11 determines whether a termination condition is met (S7). The termination condition may be, for example, when the user U issues an instruction to terminate the process via the operation device 10, or when processing of all of the music data M has been completed. If it is determined that the termination condition is not met (S7: No), the performance display process returns to step S1, and the process (S1 to S7) following acquisition of the music data M is repeated for the immediately following analysis period Q. If it is determined that the termination condition is met (S7: Yes), the performance display process ends.

[0050] Machine learning will be described using the second model G2 as an example. Fig. 10 is an explanatory diagram of a process (hereinafter referred to as "training process") for establishing the second model G2 through machine learning. By executing a program stored in the storage device 12, the control device 11 functions as the training processing unit 50 in Fig. 10 in addition to the elements (acquisition unit 21, generation unit 22, display control unit 23) illustrated in Fig. 4. The training processing unit 50 establishes the second model G2 through machine learning using multiple pieces of training data T.

[0051] A plurality of pieces of training data T are stored in the storage device 12. Each of the plurality of pieces of training data T is composed of a combination of training analysis data Pt and training control data Zt. Each piece of training data T is teacher data in which the training analysis data Pt and the training control data Zt are associated with each other. The training processing unit 50 establishes a second model G2 through machine learning using the plurality of pieces of training data T.

[0052] 11 is a flowchart of the training processing unit 50. For example, the training processing is started in response to an instruction from the user U to the operation device 10.

[0053] When the training process is started, the control device 11 (training processing unit 50) selects one of the multiple training data T (hereinafter referred to as "selected training data T") stored in the storage device 12 (Sb1). The control device 11 (training processing unit 50) iteratively updates multiple variables of an initial or provisional second model G2 (hereinafter referred to as "provisional second model G2p") using the selected training data T (Sb2 to Sb4).

[0054] The control device 11 generates control data Z by processing the analysis data P of the selected training data T using the provisional second model G2p (Sb2). The control device 11 calculates a loss function that represents the error between the control data Z generated by the provisional second model G2p and the control data Z of the selected training data T (Sb3). The control device 11 updates multiple variables of the provisional second model G2p so that the loss function is reduced (ideally minimized) (Sb4). For example, backpropagation is used to update each variable in accordance with the loss function.

[0055] The control device 11 determines whether a predetermined termination condition is met (Sb5). The termination condition may be, for example, that the loss function falls below a predetermined threshold, or that the change in the loss function falls below a predetermined threshold. If the termination condition is not met (Sb5: NO), the control device 11 selects unselected training data T stored in the storage device 12 as new selected training data T (Sb1). That is, the process of updating multiple variables of the provisional second model G2p (Sb2 to Sb4) is repeated until the termination condition is met (Sb5: YES). If the termination condition is met (Sb5: YES), the control device 11 terminates the training process. The provisional second model G2p at the time the termination condition is met is confirmed as the trained second model G2.

[0056] As can be understood from the above explanation, the second model G2 learns the latent relationship between the analytical data Pt and the control data Zt in the multiple training data T. Therefore, the trained second model G2 outputs control data Z that is statistically valid for the unknown analytical data P based on the above relationship.

[0057] As described above, in the first embodiment, the trained second model G2 is used to generate the control data Z. Therefore, it is possible to generate the control data Z that reflects the trends present in the multiple pieces of training data T used to train the second model G2.

[0058] While the above description has focused on training the second model G2, the first model G1 is also trained using the same procedure as the example shown in Fig. 11. Training the first model G1 uses a plurality of training data T including training music data Mt and training fingering data Ft. Specifically, the control device 11 (training processing unit 50) updates a plurality of variables of the first model G1 so as to reduce (ideally minimize) the error between the fingering data F generated by the initial or provisional first model G1p from the music data M of each training data T and the training fingering data F included in the training data T.

[0059] As explained above, the information processing system 100 generates control data Z by inputting music data M into the machine-learned generative model G, and therefore generates a variety of control data Z that represent appropriate behaviors of a plurality of fingers for unknown music data M, based on the underlying relationship between the music data M and the control data Z in the plurality of training data T used in the machine learning. In other words, it is possible to generate control data Z that causes a variety of changes in finger behavior depending on the musical performance.

[0060] B: Second Embodiment A second embodiment of the present disclosure will be described. Note that, for elements in the following exemplary aspects that have the same functions as those in the first embodiment, the same reference numerals as those in the first embodiment will be used, and detailed descriptions of each will be omitted as appropriate.

[0061] 12 is a block diagram illustrating the functional configuration of the control device 11 of the second embodiment. The control device 11 of the second embodiment functions as a correction unit 24 in addition to the same elements as those of the first embodiment.

[0062] The correction unit 24 corrects the control data Z. Specifically, the correction unit 24 performs processing to make the behavior of the multiple fingers represented by the control data Z closer to natural behavior. Specific examples of processing by the correction unit 24 to correct the control data Z are shown below. Two or more processes arbitrarily selected from the following examples may be combined as appropriate within a range that does not contradict each other.

[0063] (1) One of the processes by which the correction unit 24 corrects the control data Z is to correct the fingertip of a finger corresponding to a note instructed by the music data M to a position that contacts a key corresponding to the note among the keys of the keyboard instrument. Specifically, the position of the control point 41 corresponding to the fingertip of the finger number specified for the instructed note is fixed to the surface of the key corresponding to the note among the keys in the image displayed by the display device 13, for at least a portion of the note's sounding period. That is, in the finger object Oa and the musical instrument object Ob displayed on the display device 13, the fingertip of the finger of the finger number specified for the currently sounding note contacts the key corresponding to the note for at least a portion of the note's sounding period. Therefore, the behavior of the fingers represented by the control data Z can be made closer to the behavior of fingers actually playing a keyboard instrument.

[0064] (2) One of the processes by which the correction unit 24 corrects the control data Z is to correct the distance between the multiple control points 41 (i.e., the total length of each connecting portion 42) to be constant. "Maintaining a constant distance between the multiple control points 41" means restricting the movement of each control point 41 so that the distance between two control points 41 connected to each connecting portion 42 is maintained within a predetermined range. The predetermined range may be a range within which a user U (e.g., an audience member) observing the finger object Oa cannot recognize the extension or shortening of the fingers. In other words, the total length of each connecting portion 42 does not unnaturally extend or shorten depending on the behavior of the fingers. Furthermore, the total length of each connecting portion may differ for each combination of two control points 41. Therefore, the length of the fingers of the finger object Oa displayed on the display device 13 does not unnaturally change depending on the behavior of the fingers during performance. As explained above, this has the advantage that the shape of the multiple fingers represented by the control data Z is stable regardless of the behavior of the fingers.

[0065] (3) One of the processes by which the correction unit 24 corrects the control data Z is to link a finger of a plurality of fingers that corresponds to a note that the music data M instructs to produce with the fingers adjacent to that finger. In human behavior, it is difficult to move only one finger, and even if an unintended movement occurs, adjacent fingers may move in conjunction with the moved finger. The ease with which adjacent fingers can move in conjunction with each other varies depending on the combination of fingers.

[0066] Interlocking data B is used in the process of correcting control data Z so that adjacent fingers are interlocked. Interlocking data B is stored in storage device 12. Specifically, interlocking data B is stored for each of the right hand and the left hand. FIG. 13 is a schematic diagram of interlocking data B. Interlocking data B defines an index (hereinafter referred to as "interlocking index R(r, s)") that represents the degree to which movement of a fingertip r of one hand affects a fingertip s of the other hand for all combinations of two adjacent fingers selected from five fingers.

[0067] The correction unit 24 uses the linkage data B to link adjacent fingers. Specifically, the correction unit 24 identifies, from the linkage data B, a linkage index R(r, s) corresponding to the combination of the fingertip r of a finger (hereinafter referred to as a "performing finger") specified by the fingering data F and the fingertip s of another finger adjacent to the performing finger. The correction unit 24 calculates the movement distance d·R(r, s) of the fingertip s by multiplying the linkage index R(r, s) identified from the linkage data B by the movement distance d of the performing finger. The movement distance d of the fingertip r of the performing finger is identified from the control data Z before correction. The correction unit 24 then corrects the control data Z so that the fingertip s moves by the movement distance d·R(r, s) in the same direction as the fingertip r of the performing finger. Furthermore, one linkage data B may be shared between the right and left hands for the movements of corresponding fingers (e.g., the index finger of the right hand and the index finger of the left hand).

[0068] FIG. 14 is an explanatory diagram of the process for correcting the control data Z. When the fingertip r of a performing finger among multiple fingers moves by d [cm], the correction unit 24 corrects the fingertip s of other fingers adjacent to the performing finger by d·R(r, s) [cm]. Therefore, by correcting the control data Z, it is possible to faithfully reproduce the natural behavior of fingers during actual performance, in which when a specific finger moves for performance, the other fingers adjacent to that finger also move in conjunction with that finger. Note that "adjacent fingers" may refer to multiple fingers adjacent to the finger corresponding to the note.

[0069] (4) One of the processes by which the correction unit 24 corrects the control data Z is to correct the multiple fingers represented by the control data Z to match the physical information representing the requirements of the virtual performer's body. "Physical information" refers to physical characteristics of the virtual performer, such as the size of the body parts or the range of motion of each joint. The user U inputs the physical information using the operation device 10, and the correction unit 24 reflects the physical information in the control data Z.

[0070] The movement range of each control point 41 and the total length of each linking portion 42 represented by the pre-modification control data Z are set to standard values ​​corresponding to the standard constitution or physique of the virtual player. The user U specifies one of a plurality of stages for the standard constitution or physique as the physical information. The modification unit 24 increases or decreases the standard value represented by the control data Z according to the physical information. For example, when modifying the physical information to a larger physique, the modification unit 24 increases the total length of each linking portion 42 from the standard value. As a result of increasing the total length of the linking portion 42, the fingers represented by the finger object Oa become longer. Furthermore, when modifying the physical information to a stiffer body constitution, the modification unit 24 decreases the movement range of each control point 41 from the standard value. As a result of decreasing the movement range of the control point 41, the movable range of the fingers represented by the finger object Oa becomes narrower.

[0071] As explained above, the process of correcting the control data Z in accordance with the physical information reflects the physical characteristics represented by the physical information in the finger object Oa by correcting the range of movement of each control point 41 or the overall length of each connecting part. Therefore, when the user U uses his / her own physical information to correct the control data Z, there is an advantage in that the user U can observe the behavior of the finger object Oa in performance that reflects his / her own physical information, and can easily use this as reference for his / her own performance.

[0072] (5) Various data processes other than the processes exemplified above may be employed as the process by the correction unit 24 to correct the control data Z. One of the data processes executed by the correction unit 24 is assumed to be a filter process that alleviates discontinuous behavior of the multiple fingers represented by the control data Z. The filter process that alleviates discontinuous behavior smooths the movement of each control point 41 and each connecting portion 42 over time. An example of a filter process that alleviates discontinuous behavior is a One-euro filter. Therefore, the awkward behavior of the finger object Oa becomes smooth, and the behavior of the finger object Oa approaches natural behavior. However, the process of correcting the control data Z is not limited to the above examples.

[0073] 15 is a flowchart illustrating the specific steps of the performance display process in the second embodiment. In the performance display process in the second embodiment, the correction of the control data Z (S8) is added to the same process as in the first embodiment. Specifically, the control device 11 (correction unit 24) performs the process of correcting the control data Z as described above for the control data Z generated by the control device 11 (generation unit 22) (S8). The operations of the elements other than the correction unit 24 are the same as in the first embodiment.

[0074] The second embodiment also achieves the same effects as the first embodiment. Furthermore, the control data Z corrected by the corrector 24 represents more natural finger behavior than a configuration in which the control data Z is not corrected.

[0075] C: Third Embodiment Fig. 16 is a block diagram illustrating the functional configuration of a control device 11 according to a third embodiment. The control device 11 of the third embodiment has a configuration in which a whole-body generation unit 25 is added to the same elements as those of the second embodiment. However, in the third embodiment, it is not necessary to modify the control data Z.

[0076] The whole-body generation unit 25 generates performer data W that represents the behavior of the performer's body (whole body or upper body). Specifically, the whole-body generation unit 25 generates the performer data W by using the analysis data P. To generate the performer data W, any known performance analysis method, such as that disclosed in Japanese Patent Application Laid-Open No. 2019-139295, may be adopted.

[0077] FIG. 17 shows a three-dimensional image (hereinafter referred to as a "player object Oc") representing the upper body of a virtual player in a virtual space, including multiple fingers, arms, a chest, and a head. The display control unit 23 displays the player object Oc and the musical instrument object Ob on the display device 13. The display control unit 23 generates the player object Oc by connecting the arm represented by the player data W with the hand represented by the control data Z. Specifically, the display control unit 23 connects the arm represented by the player data W so that the direction of the hand represented by the control data Z matches the direction of the elbow of the arm. Furthermore, the right and left hands move in conjunction with the swinging of the body represented by the player data W. Note that any known data processing may be used to connect the player data W and the control data Z. For example, IK (Inverse Kinematics) may be used to connect the player data W and the control data Z. However, the process for connecting the player data W and the control data Z is not limited to the above example.

[0078] FIG. 18 is a flowchart illustrating the specific steps of the performance display process in the third embodiment. The performance display process in the third embodiment adds the steps of generating performer data W (S9) and combining control data Z with the performer data W (S10) to the same process as in the second embodiment. Specifically, the control device 11 (whole-body generation unit 25) generates performer data W using analysis data P (S9). Then, the control device 11 (display control unit 23) combines the control data Z with the performer data W (S10). The operations of the elements other than the whole-body generation unit 25 and the display control unit 23 are the same as in the second embodiment. Note that step S9 may be performed in parallel with step S4 or step S8, or the order of execution may be changed.

[0079] As explained above, the movement of the performer's body can be reproduced in a three-dimensional image by combining the control data Z and the performer data W. Therefore, the user U can visually and intuitively grasp both the overall or general movement of the performer's body and the partial or detailed movement of the performer's fingers.

[0080] D: Fourth Embodiment Fig. 19 is a block diagram illustrating the configuration of an information processing system 100 according to a fourth embodiment. The information processing system 100 according to the fourth embodiment has the same elements as those of the third embodiment, but with the addition of a sound pickup device 16. However, in the fourth embodiment, it is not necessary to modify the control data Z or generate the performer data W.

[0081] The sound collection device 16 is a device that collects sounds (e.g., instrument sounds or singing sounds) produced by the performance of the user U. Specifically, the sound collection device 16 is a microphone that collects sounds from the keyboard instrument 200 played by the user U. The sound collection device 16 generates an audio signal E that represents an audio waveform from the sounds produced by the performance of the user U. Note that the user U who plays the keyboard instrument 200 may be a different person from the user U who operates the operation device 10. Furthermore, although the configuration in which the audio signal E is generated by collecting performance sounds produced by the keyboard instrument 200 has been exemplified, the audio signal E may also be generated by playing an electric instrument that generates an audio signal from performance. Therefore, the sound collection device 16 may be omitted.

[0082] 20 is a block diagram illustrating the functional configuration of the control device 11 according to the fourth embodiment. The acquisition unit 21 estimates the time in the song at which the user U is currently playing (hereinafter referred to as the "performance time") by analyzing the audio signal E. The estimation of the performance time is performed sequentially in parallel with the actual performance by the user U. For the estimation of the performance time, any known audio analysis technology (score alignment) such as that disclosed in Japanese Patent Laid-Open No. 2015-79183 can be adopted.

[0083] The acquisition unit 21 controls the automatic performance by the sound emitting device 15 and the behavior of the player object Oc on the display device 13 so as to be synchronized with the progress of the performance time. Specifically, each time the performance time reaches a time point designated by each time data of the music data M, the acquisition unit 21 outputs the note data corresponding to the time data to the sound source device 14 and the generation unit 22. Therefore, the progress of the automatic performance by the sound emitting device 15 and the progress of the behavior of the player object Oc displayed on the display device 13 are synchronized with the actual performance by the user U.

[0084] FIG. 21 is a flowchart illustrating the specific steps of the performance display process in the fourth embodiment. In the performance display process of the fourth embodiment, for example, the steps of acquiring an audio signal E (S11) and analyzing the audio signal E (S12) are added to the same process as in the third embodiment. Specifically, when the user U starts playing, the control device 11 (acquisition unit 21) acquires the audio signal E from the sound collection device 16 (S11). The control device 11 (acquisition unit 21) analyzes the audio signal E and estimates the performance time (S12). Specifically, the control device 11 (acquisition unit 21) outputs note data of the note in the music data M corresponding to the performance time to the sound source device 14 and the control device 11 (generation unit 22). The operation of the elements other than the acquisition unit 21 is the same as in the third embodiment.

[0085] As explained above, the automatic performance by the sound emitting device 15 and the progress of the behavior of the player object Oc by the display device 13 are synchronized with the live performance by the user U. Therefore, an atmosphere is created in which the sound emitting device 15, the display device 13, and the user U are playing together in cooperation with each other.

[0086] E: Modifications Specific modifications that can be added to the above-mentioned embodiments are exemplified below. Two or more embodiments arbitrarily selected from the following examples may be combined as appropriate within a range that does not contradict each other.

[0087] (1) In the second embodiment, the modification unit 24 modifies the control data Z in accordance with the physical information. However, in each embodiment, the modification unit 24 is not limited to performing the processing as long as the physical information can be reflected in the finger object Oa (or the performer object Oc). For example, a configuration in which the generation unit 22 in each embodiment performs processing to generate the control data Z in accordance with the physical information, or a configuration in which the display control unit 23 in the third embodiment adjusts the control data Z (or the performer data W) in accordance with the physical information when combining the control data Z and the performer data W, can be envisioned.

[0088] (2) In the third embodiment, the analysis data P is used to generate the performer data W, but only the comprehensive data Pa of the analysis data P may be used to generate the performer data W. However, the use of the analysis data P to generate the performer data W has the advantage that the positions of the right and left hands can be reflected in the performer data W by using a plurality of individual data Pb[n] of the analysis data P.

[0089] (3) In the fourth embodiment, the performance time is estimated by analyzing the sound signal E of the sound picked up by the sound pickup device 16. However, the performance time may also be estimated by analyzing performance data transmitted from a MIDI instrument (e.g., an electronic keyboard instrument).

[0090] (4) In each embodiment, the performance matrix C is exemplified as a binary matrix representing the time series of notes in the analysis period Q of one piece of comprehensive data Pa or each of multiple individual data Pb[n]. However, the performance matrix C is not limited to the above example. For example, a performance matrix C may be generated that represents the performance intensity (volume) of notes in the analysis period Q of one piece of comprehensive data Pa or each of multiple individual data Pb[n]. Specifically, one element in the i-th row and j-th column of the performance matrix C represents the intensity with which the pitch corresponding to the i-th row is played in the unit period corresponding to the j-th column. With the above configuration, the performance intensity of each note is reflected in the control data Z, so that the tendency for the player's movements to differ depending on the strength of the performance intensity can be imparted to the movement of the finger object Oa (or player object Oc).

[0091] (5) In each embodiment, the control data Z is generated for each analysis period Q of a predetermined length, but the time unit for the control data Z may be any time unit. For example, the control data Z may be generated in advance for the entire music data M, and then the generated control data Z may be used to display an object on the display device 13.

[0092] (6) In each embodiment, a configuration in which the control data Z is stored in the storage device 12 is illustrated, but a configuration in which the control data Z is transmitted to another device or recorded on a portable recording medium is also envisioned.

[0093] (7) In each embodiment, the instrument displayed on the display device 13 is exemplified as a keyboard instrument, but the instruments displayed on the display device 13 are not limited to keyboard instruments. Examples include wind instruments such as a saxophone and a flute, and string instruments such as a guitar. The present disclosure also applies to various instruments that include performance controls operated by the performer's fingers. A performance control is a part of an instrument that is operated by the performer's fingers, such as a key on a keyboard instrument. The shape or configuration of the performance control varies depending on the type of instrument. For example, keys on a wind instrument and strings on a string instrument are exemplified as performance controls. The same applies to the keyboard instrument 200 played by the user U in the fourth embodiment.

[0094] (8) In each embodiment, finger numbers are used as an example of information (finger information) for identifying one of multiple fingers in the fingering data F. However, finger information is not limited to numbers as long as multiple fingers can be identified. For example, the finger information may be a letter corresponding to each finger, or a combination of a letter and a number.

[0095] (9) The information processing system 100 may be realized by a server device that communicates with an information device such as a smartphone or a tablet terminal. For example, the information processing system 100 generates control data Z using music data M received from the information device and transmits the control data Z to the information device.

[0096] In a configuration in which the analysis data P is transmitted from the information device to the information processing system 100 (i.e., the analysis data generation unit 32 is mounted on the information device), the analysis data generation unit 32 may be omitted from the information processing system 100. In addition, in a configuration in which the control data Z is transmitted from the information device to the information processing system 100 (i.e., the generation unit 22 is mounted on the information device), the generation unit 22 may be omitted from the information processing system 100.

[0097] (10) As described above, the functions of the information processing system 100 illustrated above are realized through cooperation between one or more processors constituting the control device 11 and a program stored in the storage device 12. The program according to the present disclosure may be provided in a form stored on a computer-readable recording medium and installed on a computer. The recording medium may be, for example, a non-transitory recording medium, such as an optical recording medium (optical disk) such as a CD-ROM, but may also include any known type of recording medium, such as a semiconductor recording medium or a magnetic recording medium. Note that a non-transitory recording medium includes any recording medium other than a transient, propagating signal, and does not exclude volatile recording media. Furthermore, in a configuration in which a distribution device distributes a program via a communication network, the storage medium that stores the program in the distribution device corresponds to the non-transitory recording medium described above.

[0098] (11) The term "nth" (n is a natural number) in this application is used only as a formal and convenient label to distinguish each element in the description and does not have any substantive meaning. Therefore, there is no room for restrictive interpretation of the position of each element or the order of production, etc., based on the term "nth."

[0099] F: Supplementary Note From the above-described exemplary embodiments, the following configurations can be understood, for example.

[0100] An information processing method according to one aspect (aspect 1) of the present disclosure is realized by a computer system that acquires music data representing a time series of musical notes and processes the music data using a machine-learned generative model to generate control data representing the behaviors of multiple fingers playing the time series of musical notes. In this aspect, the control data is generated by inputting the music data into the machine-learned generative model. Therefore, a variety of control data representing appropriate behaviors of multiple fingers for unknown music data is generated based on the underlying relationship between the music data and the control data in the multiple training data used in the machine learning. In other words, it is possible to generate control data for changing the behavior of the fingers in a variety of ways depending on the musical performance.

[0101] In a specific example (Aspect 2) of Aspect 1, the generative model includes a first model and a second model, and the generation of the control data includes a first process of generating fingering data specifying fingering information for each note by processing the music data using the first model, and a second process of generating the control data by processing analysis data representing a relationship between a time series of notes represented by the music data and the fingering represented by the fingering data using the second model. In the above aspect, the first process of generating fingering data from the music data and the second process of generating control data from the analysis data are performed separately. Therefore, the first model and the second model can be trained separately by machine learning. Furthermore, either the first model or the second model can be selectively changed (adjusted, replaced, etc.). Furthermore, processing such as partial change of the fingering data generated by the first process or replacement with different fingering data can be performed.

[0102] In a specific example (Aspect 3) of Aspect 1 or Aspect 2, the generation of the analysis data includes a third process of generating the analysis data based on the music data and the fingering data generated by the first process, and the second process processes the analysis data generated by the third process using the second model. In the above aspects, the analysis data is generated by the third process. Therefore, before generating the analysis data by the third process, it is possible to perform processes such as partially changing the fingering data or replacing it with different fingering data.

[0103] In a specific example (Aspect 4) of any of Aspects 2 to 3, the analysis data includes a plurality of individual data corresponding to different fingers, and each of the plurality of individual data represents a time series of notes, among the time series of notes represented by the music data, for which fingering information for the finger corresponding to the individual data is specified by the fingering data. In the above aspects, control data reflecting the fingering tendencies of each finger can be generated. Furthermore, when correcting the fingering tendencies of the individual data, the correction can be made by replacing the individual data to be corrected among the plurality of individual data with another individual data. Note that "finger information" refers to information (e.g., a finger number) for identifying one of the plurality of fingers.

[0104] In a specific example (Aspect 5) of any of Aspects 2 to 4, the analysis data further includes comprehensive data representing a time sequence of notes represented by the music data. In these aspects, control data can be generated that reflects not only the fingering tendencies of different fingers but also the overall tendencies of the music. Furthermore, since the music data includes notes for which fingering could not be estimated, control data can be generated that more accurately reflects the music than when using only individual data. Note that "accurately reflecting the music" means that no notes represented by the music data are omitted.

[0105] In a specific example (Aspect 6) of any of Aspects 2 to 5, the second model includes a first neural network and a second neural network, and in the second processing, the first neural network processes the analysis data to generate a feature vector representing characteristics of the analysis data, and the second neural network processes the feature vector to generate the control data corresponding to the feature vector. In the above aspect, since the trained model includes a combination of the first neural network and the second neural network, appropriate control data corresponding to music data can be generated.

[0106] In a specific example (Aspect 7) of any of Aspects 1 to 6, the control data indicates the positions of each of a plurality of control points corresponding to the fingers. According to the above aspects, there is an advantage that the behavior of the fingers can be controlled by the control data regardless of the size of the fingers.

[0107] In a specific example (Aspect 8) of any one of Aspects 1 to 7, the control data is further modified. In the above aspect, the behavior of the fingers becomes more natural compared to a configuration in which the control data is not modified.

[0108] In a specific example (aspect 9) of aspect 8, modifying the control data includes modifying the fingertips of the fingers of the plurality of fingers corresponding to the note instructed to be sounded by the music data to positions that contact the performance controller of the musical instrument corresponding to that note. In the above aspect, the control data is modified so that the fingertips of the fingers corresponding to the note instructed to be sounded contact the performance controller corresponding to that note. Therefore, the behavior of the plurality of fingers represented by the control data can be made closer to the behavior of fingers actually performing the musical instrument. Note that a "performance controller" refers to a part of a musical instrument that is operated by fingers when playing the instrument. Examples include keys on a keyboard instrument such as a piano, keys on a wind instrument such as a saxophone or flute, and strings on a stringed instrument such as a guitar. However, performance controllers are not limited to the above examples.

[0109] In a specific example (Aspect 10) of Aspect 8 or Aspect 9, the modification of the control data includes processing for linking a finger of the plurality of fingers corresponding to a note instructed by the music data to be produced with adjacent fingers of the finger. In the above aspect, the control data can faithfully reproduce the natural behavior of multiple fingers during actual performance, in which when a specific finger moves to perform a performance, the other fingers adjacent to the finger also move in conjunction with the finger. Note that the "adjacent fingers" may refer to multiple fingers adjacent to the finger corresponding to the note.

[0110] In a specific example (Aspect 11) of any of Aspects 8 to 10, the modification of the control data includes a process of modifying the lengths between the plurality of control points corresponding to the fingers to be constant. In the above aspects, the control data is modified so that the lengths of the fingers do not change depending on the behavior. This has the advantage of stabilizing the shapes of the fingers represented by the control data.

[0111] In a specific example (Aspect 12) of Aspect 8 or Aspect 11, modifying the control data includes modifying the multiple fingers represented by the control data to match physical information representing the physical requirements of a virtual performer. The above aspects allow a user to generate control data that reflects their own or another user's physical information. This has the advantage that a user with a body similar to the physical information used for modification can easily use the behavior of the multiple fingers represented by the control data as a reference for their performance. Note that "physical information" refers to physical characteristics such as the size of body parts or the range of motion of each joint.

[0112] In a specific example (Aspect 13) of any of Aspects 1 to 12, an image representing the performer's fingers is further generated in accordance with the control data. In the above aspects, an image representing the behavior of the fingers can be generated using the control data. Therefore, a user who views the image can visually and intuitively grasp the behavior of the fingers playing the time sequence of notes represented by the music data.

[0113] In a specific example (Aspect 14) of Aspect 13, player data representing the player's physical behavior is further generated, and the generation of the image includes generating an image including the player's fingers and body by combining the player data and the control data. In the above aspect, the generated player's physical behavior and the generated player's wrist behavior are combined to reproduce the player's entire body behavior in an image. Therefore, a user can visually and intuitively grasp both the overall or general behavior of the player's body and the partial or detailed behavior of the player's fingers.

[0114] In a specific example (Aspect 15) of Aspect 13 or Aspect 14, the control data is generated by processing analysis data that represents the relationship between the time series of notes represented by the music data and the fingerings corresponding to the time series of notes, and the performer data is generated by using the analysis data. In the above aspects, the analysis data is used to generate both the control data and the performer data. Therefore, the processing load for generating the control data and the performer data is reduced compared to an embodiment in which separate data is used to generate the control data and the performer data. Furthermore, because the same analysis data is used to generate the control data and the performer data, a sense of unity can be maintained between the behavior of the multiple fingers represented by the control data and the behavior of the performer's body represented by the performer data.

[0115] In a specific example (Aspect 16) of any of Aspects 13 to 15, the analysis data includes, for each of a plurality of fingers, individual data specifying the time series of notes to be played by that finger among the time series of notes represented by the music data, and comprehensive data representing the time series of notes represented by the music data, and the performer data is generated using the comprehensive data from the analysis data. In the above aspects, the comprehensive data is used to generate the control data and the performer data. Therefore, the processing load for generating the control data and the performer data is reduced compared to an embodiment in which separate data is used to generate the control data and the performer data. Furthermore, because the same comprehensive data is used to generate the control data and the performer data, a sense of unity can be maintained between the behavior of the plurality of fingers represented by the control data and the behavior of the performer's body represented by the performer data.

[0116] An information processing system according to one aspect (aspect 17) of the present disclosure includes an acquisition unit that acquires music data representing a time series of musical notes, and a generation unit that processes the music data using a machine-learned generative model to generate control data representing the behaviors of multiple fingers playing the time series of musical notes. In this aspect, the control data is generated by inputting the music data into the machine-learned generative model. Therefore, based on the underlying relationship between the music data and the control data in the multiple training data used in the machine learning, a variety of control data representing appropriate behaviors of multiple fingers for unknown music data is generated. In other words, it is possible to generate control data for changing the finger behaviors in a variety of ways depending on the musical performance.

[0117] A program according to one aspect (aspect 18) of the present disclosure causes a computer system to function as an acquisition unit that acquires music data representing a time series of notes, and a generation unit that processes the music data using a machine-learned generative model to generate control data representing the behavior of fingers playing the time series of notes. In this aspect, the control data is generated by inputting music data into the machine-learned generative model. Therefore, based on the underlying relationship between the music data and the control data in the multiple training data used in the machine learning, a variety of control data representing appropriate behaviors of multiple fingers for unknown music data is generated. In other words, it is possible to generate control data for changing the behavior of fingers in a variety of ways depending on the musical performance.

[0118] 10...storage device, 11...control device, 12...display device, 13...sound source device, 14...sound emission device, 15...operation device, 16...sound collection device, 21...acquisition unit, 22...generation unit, 23...display control unit, 24...correction unit, 25...whole body generation unit, 31...fingering data generation unit, 32...analysis data generation unit, 33...control data generation unit, 41...control point, 42...connection unit, 50...training processing unit, 100...information processing system, 200...keyboard instrument, C...performance matrix, E...acoustic signal, F...fingering data, Ft...training fingering data, G...generative model, G1...first model G1p...tentative first model, G2...second model, G2a...first neural network, G2b...second neural network, G2p...tentative second model, M...music data, Mt...music data for training, P...analysis data, P1...overall data, P2...individual data, Pt...analysis data for training, Q...analysis period, R(r, s)...linkage index, r...fingertip of performing finger, s...other fingertip adjacent to the performing finger, T...training data, U...user, V...feature vector, W...performer data, Z...control data, Zt...control data for training.

Claims

1. An information processing method implemented by a computer system that acquires music data representing a time sequence of musical notes, processes the music data using a machine-learned generative model, and generates control data representing the behavior of multiple fingers playing the time sequence of musical notes.

2. The information processing method of claim 1, wherein the generative model includes a first model and a second model, and the generation of the control data includes: a first process of generating fingering data specifying fingering information for each of the notes by processing the music data using the first model; and a second process of generating the control data by processing analysis data representing the relationship between the time series of notes represented by the music data and the fingering represented by the fingering data using the second model.

3. The information processing method of claim 2, wherein the generation of the analysis data includes a third process of generating the analysis data in accordance with the music data and the fingering data generated by the first process, and in the second process, the analysis data generated by the third process is processed using the second model.

4. An information processing method according to claim 2 or 3, wherein the analysis data includes a plurality of individual data corresponding to different fingers, and each of the plurality of individual data represents a time series of notes, among the time series of notes represented by the music data, for which fingering information for the finger corresponding to that individual data is specified by the fingering data.

5. The information processing method according to claim 4, wherein said analysis data further includes comprehensive data representing a time sequence of notes represented by said music data.

6. The information processing method of claim 2, wherein the second model includes a first neural network and a second neural network, and the second processing involves processing the analysis data using the first neural network to generate a feature vector representing the characteristics of the analysis data, and processing the feature vector using the second neural network to generate the control data corresponding to the feature vector.

7. The information processing method according to claim 1, wherein the control data indicates the positions of a plurality of control points corresponding to the fingers.

8. The information processing method of claim 1, further comprising modifying said control data.

9. An information processing method according to claim 8, wherein the modification of the control data includes modifying the fingertip of a finger among the plurality of fingers corresponding to a note that the music data instructs to sound to a position that contacts a performance operator among a plurality of performance operators of a musical instrument that corresponds to that note.

10. An information processing method according to claim 8 or 9, wherein the modification of the control data includes a process of linking a finger among the plurality of fingers that corresponds to the musical note that the music data instructs to produce with the fingers adjacent to that finger.

11. The information processing method according to claim 8, wherein said control data correction includes a process for correcting the lengths between a plurality of control points corresponding to said fingers to be constant.

12. An information processing method according to claim 8, wherein said control data modification includes a process of modifying the plurality of fingers represented by said control data to match physical information representing requirements relating to the body of a virtual player.

13. The information processing method of claim 1, further comprising generating an image representing the performer's fingers in response to said control data.

14. The information processing method of claim 13, further comprising generating performer data representing the movement of the performer's body, and generating said image by combining said performer data with said control data to generate an image including the performer's fingers and body.

15. An information processing method according to claim 14, wherein in generating the control data, the control data is generated by processing analysis data that represents the relationship between the time series of notes represented by the music data and the fingerings corresponding to the time series of notes, and in generating the performer data, the performer data is generated using the analysis data.

16. An information processing method according to claim 15, wherein the analysis data includes, for each of a plurality of fingers, individual data specifying the time series of notes to be played by that finger from the time series of notes represented by the music data, and comprehensive data representing the time series of notes represented by the music data, and wherein in generating the performer data, the comprehensive data from the analysis data is used to generate the performer data.

17. An information processing system comprising: an acquisition unit that acquires music data representing a time sequence of musical notes; and a generation unit that processes the music data using a machine-learned generative model to generate control data representing the behavior of multiple fingers playing the time sequence of musical notes.

18. A program that causes a computer system to function as an acquisition unit that acquires music data representing a time sequence of musical notes, and a generation unit that processes the music data using a machine-learned generative model to generate control data representing the behavior of fingers playing the time sequence of musical notes.

Citation Information

Patent Citations

  • Information processing method

    WO2019156092A1

  • Information processing method, information processing system, and program

    WO2023181570A1