Information processing system, electronic musical instrument, information processing method, program, and machine learning system

The information processing system addresses the challenge of personalized musical practice by using a trained model to analyze performance data and provide tailored practice phrases, enhancing practice effectiveness and efficiency.

JP7823706B2Active Publication Date: 2026-03-04YAMAHA CORP
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-10-03
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Existing systems fail to effectively assist users in practicing musical instruments by considering individual performance tendencies, making it difficult to improve performance based on personal mistakes and weaknesses.

Method used

An information processing system that utilizes a trained model to analyze a user's performance data, identify performance tendencies, and provide tailored practice phrases to address these tendencies, incorporating elements like deep neural networks and machine learning to generate personalized practice content.

Benefits of technology

Enables effective practice by providing personalized practice phrases that align with a user's performance tendencies, improving practice efficiency and reducing the load of identifying suitable practice material.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007823706000001
    Figure 0007823706000001
  • Figure 0007823706000002
    Figure 0007823706000002
  • Figure 0007823706000003
    Figure 0007823706000003
Patent Text Reader

Abstract

To achieve effective performance practice according to a user's performance tendency.SOLUTION: An information processing system includes a receiving unit that receives performance data representing a performance of a piece of music by a practitioner, and a generating unit that uses the performance data to generate tendency data representing a performance tendency designated by an instructor for the performance.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a technique for assisting playing musical instruments such as electronic musical instruments. [Background technology]

[0002] Various technologies have been proposed to support the performance of musical instruments such as electronic musical instruments. For example, Patent Document 1 discloses a technology that calculates statistical values ​​such as standard deviations from the differences between parameters of pre-prepared music data and parameters of performance data representing a performance by a user, and then tallying the statistical values ​​using a method appropriate for the type of the parameters. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2005-55635 Summary of the Invention [Problem to be solved by the invention]

[0004] However, simply presenting a user with an evaluation value that is the result of evaluating the performance makes it difficult to effectively practice a performance that takes into account the performance tendencies of each individual user (for example, tendencies in performance mistakes, etc.) In consideration of the above circumstances, one aspect of the present disclosure aims to realize effective performance practice that is in line with the user's performance tendencies. [Means for solving the problem]

[0005] In order to solve the above problems, an information processing system according to one aspect of the present disclosure includes a performance data acquisition unit that acquires performance data representing a musical piece played by a user, a tendency identification unit that generates tendency data representing the performance tendency of the user by inputting the performance data acquired by the performance data acquisition unit into a first trained model that has learned the relationship between training performance data representing the musical piece performance and training tendency data that represents the performance tendency represented by the training performance data, and a practice phrase identification unit that identifies practice phrases according to the tendency data generated by the tendency identification unit.

[0006] An electronic musical instrument according to one aspect of the present disclosure includes a performance acceptance unit that accepts a musical piece performance by a user, a performance data acquisition unit that acquires performance data representing the performance accepted by the performance acceptance unit, a tendency identification unit that inputs the performance data acquired by the performance data acquisition unit into a first learned model that has learned the relationship between training performance data representing the musical piece performance and training tendency data that represents the performance tendency represented by the training performance data, and outputs tendency data that represents the performance tendency of the user from the first learned model, a practice phrase identification unit that uses the tendency data output by the tendency identification unit to identify practice phrases that correspond to the performance tendency of the user, and a presentation processing unit that presents the practice phrases to the user.

[0007] An information processing method according to one aspect of the present disclosure acquires performance data representing a musical piece played by a user, and inputs the acquired performance data into a first trained model that has learned the relationship between training performance data representing the musical piece performance and training trend data representing the performance trend represented by the training performance data, thereby generating trend data representing the performance trend of the user and identifying practice phrases corresponding to the trend data.

[0008] A machine learning system according to one aspect of the present disclosure includes a first learning data acquisition unit that acquires performance data representing a performance of a piece of music by a user and instruction data representing a time point within the piece of music and a performance tendency at that time point; and a first learning processing unit that establishes a first trained model that learns the relationship between the training performance data and the training tendency data through machine learning using first learning data that represents a combination of training performance data representing a performance within a section of the performance data that includes a time point represented by the instruction data and training tendency data representing the performance tendency represented by the instruction data. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a block diagram illustrating the configuration of a performance system according to a first embodiment. [Figure 2] FIG. 1 is a block diagram illustrating the configuration of an electronic musical instrument. [Figure 3] FIG. 1 is a block diagram illustrating a configuration of an information processing system. [Figure 4] FIG. 2 is a block diagram illustrating an example of a functional configuration of an information processing system. [Figure 5] 10 is a flowchart illustrating a specific procedure of a specification process. [Figure 6] FIG. 1 is a block diagram illustrating a configuration of a machine learning system. [Figure 7] FIG. 1 is a block diagram illustrating an example of the functional configuration of a machine learning system. [Figure 8] FIG. 10 is a block diagram illustrating the configuration of an information device used by an instructor. [Figure 9] FIG. 10 is a schematic diagram of indicated data. [Figure 10] 10 is a flowchart illustrating a specific procedure of a preparation process. [Figure 11] 10 is a flowchart illustrating a specific procedure of a learning process. [Figure 12] FIG. 10 is a block diagram illustrating an example of the functional configuration of an information processing system according to a second embodiment. [Figure 13]10 is a flowchart illustrating a procedure of a specification process in the second embodiment. [Figure 14] FIG. 10 is a block diagram illustrating an example of the functional configuration of an information processing system according to a third embodiment. [Figure 15] 10 is a flowchart illustrating a procedure of a specification process in the third embodiment. [Figure 16] FIG. 10 is a block diagram illustrating an example of the functional configuration of a machine learning system according to a third embodiment. [Figure 17] 10 is a flowchart illustrating a procedure of a learning process in the third embodiment. [Figure 18] FIG. 10 is a block diagram illustrating the functional configuration of an electronic musical instrument according to a fourth embodiment. [Figure 19] FIG. 13 is a block diagram illustrating an example of the functional configuration of an information device according to a fifth embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] A: First embodiment FIG. 1 is a block diagram illustrating the configuration of a performance system 100 according to the first embodiment. The performance system 100 is a computer system that allows a user U of an electronic musical instrument 10 to practice playing the electronic musical instrument 10, and includes the electronic musical instrument 10, an information processing system 20, and a machine learning system 30. The elements that make up the performance system 100 communicate with each other via a communication network 200, such as the Internet. Note that while the performance system 100 actually includes multiple electronic musical instruments 10, the following description will focus on one arbitrary electronic musical instrument 10 for convenience.

[0011] FIG. 2 is a block diagram illustrating the configuration of an electronic musical instrument 10. The electronic musical instrument 10 is a performance device used by a user U to play music. The electronic musical instrument 10 of the first embodiment is an electronic keyboard instrument equipped with multiple keys operated by the user U. The electronic musical instrument 10 is realized by a computer system including a control device 11, a storage device 12, a communication device 13, a performance device 14, a display device 15, a sound source device 16, and a sound output device 17. The electronic musical instrument 10 may be realized as a single device, or may be realized as multiple devices configured separately from each other.

[0012] The control device 11 is composed of one or more processors that control each element of the electronic musical instrument 10. For example, the control device 11 is composed of one or more types of processors, such as a CPU (Central Processing Unit), an SPU (Sound Processing Unit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or an ASIC (Application Specific Integrated Circuit).

[0013] The storage device 12 is one or more memories that store programs executed by the control device 11 and various data used by the control device 11. The storage device 12 is configured from a known storage medium such as a magnetic storage medium or a semiconductor storage medium, or a combination of multiple types of storage media. Note that the storage device 12 may be a portable storage medium that is detachable from the electronic musical instrument 10, or a storage medium (e.g., cloud storage) that the control device 11 can write to or read from via the communication network 200, for example.

[0014] The storage device 12 of the first embodiment stores multiple pieces of music data X representing different pieces of music. The music data X for each piece of music specifies the time sequence of multiple notes that make up part or all of the piece of music. Specifically, the music data X specifies the pitch and duration of each note in the piece of music. The music data X is data in a format that complies with, for example, the MIDI (Musical Instrument Digital Interface) standard.

[0015] The communication device 13 communicates with the information processing system 20 via a communication network 200. Note that communication between the communication device 13 and the communication network 200 may be either wired or wireless. Alternatively, a communication device 13 separate from the electronic musical instrument 10 may be connected to the electronic musical instrument 10 by wire or wirelessly. Examples of the communication device 13 separate from the electronic musical instrument 10 include information terminals such as smartphones and tablet terminals.

[0016] The display device 15 displays an image under the control of the control device 11. For example, various display panels such as a liquid crystal display panel or an organic EL (Electroluminescence) panel are used as the display device 15. The display device 15 displays the musical score of a piece of music performed by the user U, for example, using music data X of the piece of music.

[0017] The performance device 14 is an input device that accepts a performance by the user U. Specifically, the performance device 14 has a keyboard on which a plurality of keys corresponding to different pitches are arranged. The user U plays a piece of music by sequentially operating the desired keys on the performance device 14. The performance device 14 is an example of a "performance acceptance unit."

[0018] The control device 11 generates performance data Y representing a musical piece performed by the user U. Specifically, the performance data Y specifies the pitch and duration of each of a plurality of notes designated by the user U through operations on the performance device 14. Like the musical piece data X, the performance data Y is time-series data in a format conforming to, for example, the MIDI standard. The communication device 13 transmits the performance data Y representing the musical piece performed by the user U and the musical piece data X of the musical piece to the information processing system 20. The musical piece data X is data representing an exemplary or standard performance of a musical piece, and the performance data Y is data representing the actual performance of the musical piece by the user U. Therefore, although the musical notes designated by the musical piece data X and the musical piece data Y are mutually correlated, they do not completely match. The difference between the musical piece data X and the performance data Y is particularly noticeable in parts of the musical piece where the user U is likely to make performance mistakes or where the user U is not good at playing.

[0019] The sound source device 16 generates an audio signal A corresponding to the performance on the performance device 14. The audio signal A is a signal representing the waveform of a musical tone instructed by the performance on the performance device 14. Specifically, the sound source device 16 is a MIDI sound source that generates an audio signal A representing the musical tone of each note specified in time series by the performance data Y. That is, the sound source device 16 generates an audio signal A representing a musical tone of a pitch corresponding to a key pressed by the user U among multiple keys on the performance device 14. Note that the control device 11 may also realize the function of the sound source device 16 by executing a program stored in the storage device 12. That is, the sound source device 16 dedicated to generating the audio signal A may be omitted.

[0020] The sound emitting device 17 emits the performance sound represented by the acoustic signal A. For example, a speaker or headphones are used as the sound emitting device 17. As can be understood from the above explanation, the sound source device 16 and the sound emitting device 17 in the first embodiment function as a playback system 18 that plays back musical sounds corresponding to the performance by the user U.

[0021] 3 is a block diagram illustrating the configuration of an information processing system 20. The information processing system 20 provides a user U with musical phrases Z (hereinafter referred to as "practice phrases") suitable for performance practice by the user U. The information processing system 20 is realized by a computer system including a control device 21, a storage device 22, and a communication device 23. The information processing system 20 may be realized as a single device, or may be realized as multiple devices configured separately from each other.

[0022] The control device 21 is composed of one or more processors that control each element of the information processing system 20. For example, the control device 21 is composed of one or more types of processors, such as a CPU, SPU, DSP, FPGA, or ASIC. The communication device 23 communicates with each of the electronic musical instrument 10 and the machine learning system 30 via the communication network 200. Note that communication between the communication device 23 and the communication network 200 may be either wired communication or wireless communication.

[0023] The storage device 22 is one or more memories that store programs executed by the control device 21 and various data used by the control device 21. The storage device 22 is configured from a known recording medium such as a magnetic recording medium or a semiconductor recording medium, or a combination of multiple types of recording media. Note that the storage device 22 may be a portable recording medium that is detachable from the information processing system 20, or a recording medium (e.g., cloud storage) that the control device 21 can write to or read from via the communication network 200, for example.

[0024] 4 is a block diagram illustrating an example of the functional configuration of the information processing system 20. The storage device 22 stores a plurality of practice phrases Z corresponding to different trend data D. In other words, the storage device 22 stores a table in which each of the plurality of trend data D and each of the plurality of practice phrases Z are associated with each other.

[0025] The tendency data D is data in any format that represents the tendencies of a performer's performance (hereinafter referred to as "performance tendencies"). Performance tendencies may be, for example, the tendency of a performer to make performance errors or tendencies in a performing technique that the performer is not good at. For example, the tendency data D may specify one of a number of performance tendencies, such as "playing keys at the wrong time," "playing keys adjacent to the intended key," "playing the wrong pitch," "poor performance of leap progressions," "poor performance of chords," or "poor performance of finger crossing." A leap progression is a section in which two notes are played in succession, with a pitch difference greater than a predetermined value (e.g., a third). Finger crossing is a performing technique in which a finger of one hand moves under a finger that is pressing a key corresponding to a note, thereby playing a higher note.

[0026] A practice phrase Z is time-series data representing a piece of music made up of multiple notes, and more specifically, a melody suitable for practicing the electronic musical instrument 10 (for example, part or all of a practice piece). A practice phrase Z is composed of a time series of single notes or chords. A practice phrase Z corresponding to each piece of tendency data D represents a piece of music suitable for improving the performance tendency specified by the tendency data D. For example, for the tendency data D for a performance tendency of "not good at jump progressions," a practice phrase Z rich in jump progressions is registered. Similarly, for the tendency data D for a performance tendency of "not good at playing chords," a practice phrase Z rich in chords is registered. A practice phrase Z is, for example, MIDI-format data that specifies the pitch and sound duration for each of multiple notes.

[0027] The control device 21 of the information processing system 20 executes a program stored in the storage device 22 to realize multiple elements (a performance data acquisition unit 71, a tendency identification unit 72, and a practice phrase identification unit 73) for identifying a practice phrase Z from music data X and performance data Y.

[0028] The performance data acquisition unit 71 acquires performance data Y representing a musical piece performed by the user U. Specifically, the performance data acquisition unit 71 receives the musical piece data X and the performance data Y transmitted from the electronic musical instrument 10 via the communication device 23. The performance data acquisition unit 71 generates control data C including the musical piece data X and the performance data Y.

[0029] The tendency identification unit 72 generates tendency data D representing the performance tendency of the user U in accordance with the control data C. The tendency identification unit 72 generates the tendency data D using the trained model Ma. The trained model Ma is an example of a "first trained model."

[0030] There is a correlation between the difference between the score of a piece of music played by a performer (music data X) and the performer's actual performance (performance data Y), and the performer's performance tendency (tendency data D). For example, if the timing of each note differs between the music data X and the performance data Y, a performance tendency of "shifting key pressing times" is estimated. Also, if the performance data Y specifies a note close to the note represented by the music data X, a performance tendency of "pressing a key adjacent to the target key" is estimated. Also, if the difference between the music data X and the performance data Y is significant in a section of the music that contains a jump progression, a performance tendency of "not being good at jump progressions" is estimated. The trained model Ma is a statistical estimation model that has learned the above tendencies. In other words, the trained model Ma is a statistical estimation model that has learned the relationship between the combination of the music data X and the performance data Y (i.e., the control data C) and the tendency data D that represents the performer's performance tendency. The tendency identification unit 72 inputs control data C including music piece data X and performance data Y into the trained model Ma, and outputs tendency data D representing the performance tendency of the user U from the trained model Ma.

[0031] The trained model Ma is configured, for example, by a deep neural network (DNN). For example, any type of neural network, such as a recurrent neural network (RNN) or a convolutional neural network (CNN), may be used as the trained model Ma. The trained model Ma may also be configured by combining multiple types of deep neural networks. Additionally, additional elements, such as a long short-term memory (LSTM), may be incorporated into the trained model Ma.

[0032] The trained model Ma is realized by a combination of a program that causes the control device 21 to execute a calculation to generate trend data D from control data C and a plurality of variables (specifically, weights and biases) that are applied to the calculation. The program and the plurality of variables that realize the trained model Ma are stored in the storage device 22. The numerical values ​​of each of the plurality of variables that define the trained model Ma are set in advance by machine learning.

[0033] The practice phrase identification unit 73 uses the tendency data D identified by the tendency identification unit 72 to identify a practice phrase Z that corresponds to the performance tendency of the user U. Specifically, the practice phrase identification unit 73 searches the storage device 22 for a practice phrase Z that corresponds to the tendency data D identified by the tendency identification unit 72, among the multiple practice phrases Z stored in the storage device 22. In other words, a practice phrase Z that is suitable for improving the performance tendency of the user U represented by the tendency data D is identified.

[0034] The practice phrase Z identified by the practice phrase identification unit 73 is transmitted from the communication device 23 to the electronic musical instrument 10. The communication device 13 of the electronic musical instrument 10 receives the practice phrase Z transmitted from the information processing system 20. The control device 11 displays the musical score of the practice phrase Z on the display device 15. The user U plays the practice phrase Z while checking the musical score displayed on the display device 15.

[0035] FIG. 5 is a flowchart illustrating a specific procedure of a process Sa executed by the control device 21 of the information processing system 20 (hereinafter referred to as a "specific process").

[0036] When the identification process Sa is started, the performance data acquisition unit 71 waits until the music piece data X and performance data Y transmitted from the electronic musical instrument 10 are received by the communication device 23 (Sa1: NO). When the performance data acquisition unit 71 acquires the music piece data X and performance data Y (Sa1: YES), the tendency identification unit 72 inputs control data C including the music piece data X and performance data Y into the trained model Ma, thereby outputting tendency data D from the trained model Ma (Sa2). The practice phrase identification unit 73 identifies a practice phrase Z corresponding to the tendency data D from among the multiple practice phrases Z stored in the storage device 22 (Sa3). The practice phrase identification unit 73 transmits the practice phrase Z from the communication device 23 to the electronic musical instrument 10 (Sa4).

[0037] As described above, in the first embodiment, performance data Y representing a musical piece played by a user U is input into a trained model Ma to generate tendency data D representing the performance tendency of the user U, and a practice phrase Z is identified according to the tendency data D. Therefore, when the user U plays the practice phrase Z, effective practice according to the performance tendency of the user U is realized.

[0038] In the first embodiment, the practice phrase Z corresponding to the performance tendency of the user U is identified from among a plurality of practice phrases Z corresponding to different performance tendencies (tendency data D). Therefore, the processing load for identifying the practice phrase Z according to the performance tendency of the user U is reduced.

[0039] The machine learning system 30 in FIG. 1 generates the trained model Ma exemplified above. FIG. 6 is a block diagram illustrating the configuration of the machine learning system 30. The machine learning system 30 includes a control device 31, a storage device 32, and a communication device 33. The machine learning system 30 can be realized as a single device, or as multiple devices configured separately from each other.

[0040] The control device 31 is composed of one or more processors that control each element of the machine learning system 30. For example, the control device 31 is composed of one or more types of processors, such as a CPU, SPU, DSP, FPGA, or ASIC. The communication device 33 communicates with the information processing system 20 via the communication network 200. Note that communication between the communication device 33 and the communication network 200 may be either wired communication or wireless communication.

[0041] The storage device 32 is one or more memories that store programs executed by the control device 31 and various data used by the control device 31. The storage device 32 is configured from a known recording medium such as a magnetic recording medium or a semiconductor recording medium, or a combination of multiple types of recording media. The storage device 32 may also be a portable recording medium that is detachable from the machine learning system 30, or a recording medium (e.g., cloud storage) that the control device 31 can write to or read from via the communication network 200.

[0042] 7 is a block diagram illustrating an example of the functional configuration of the machine learning system 30. The control device 31 executes a program stored in the storage device 32 to function as multiple elements (a learning data acquisition unit 81a and a learning processing unit 82a) for establishing a trained model Ma through machine learning.

[0043] The learning processing unit 82a establishes a learned model Ma through supervised machine learning (learning process Sc, described below) using multiple pieces of learning data Ta. The learning data acquisition unit 81a acquires multiple pieces of learning data Ta. The multiple pieces of learning data Ta acquired by the learning data acquisition unit 81a are stored in the storage device 32. Each of the multiple pieces of learning data Ta is composed of a combination of learning control data Ct and learning trend data Dt. The control data Ct includes learning music data Xt and learning performance data Yt. The music data Xt is an example of "learning music data," the performance data Yt is an example of "learning performance data," and the trend data Dt is an example of "learning trend data." Furthermore, the music represented by the music data Xt is an example of a "reference music." The learning data acquisition unit 81a is an example of a "first learning data acquisition unit," and the learning processing unit 82a is an example of a "first learning processing unit." Furthermore, the learning data Ta is an example of "first learning data."

[0044] As illustrated in FIG. 7, the learning data Ta is generated using the results of a performance of a musical piece by a learner U1 and instruction on that performance by an instructor U2. The learner U1 plays the musical piece using an electronic musical instrument 10. The instructor U2 evaluates and provides instruction on the performance of the learner U1 using an information device 40. The information device 40 is, for example, an information terminal such as a smartphone or a tablet terminal. The learner U1 and the instructor U2 are located, for example, in remote locations. However, the learner U1 and the instructor U2 may also be located in the same place.

[0045] The electronic musical instrument 10 transmits music piece data X0 representing a piece of music and performance data Y0 representing a performance of the piece of music by a learner U1 to the information device 40 and the machine learning system 30. The music piece data X0, like the aforementioned music piece data X, specifies the time series of multiple notes that make up the piece of music. The performance data Y0, like the aforementioned performance data Y, specifies the time series of multiple notes instructed by the learner U1 through operations on the performance device 14.

[0046] 8 is a block diagram illustrating the configuration of an information device 40. The information device 40 is a computer system that allows an instructor U2 to evaluate and provide instruction to a learner U1's performance of the electronic musical instrument 10, and includes a control device 41, a storage device 42, a communication device 43, an operation device 44, a display device 45, and a playback system 46. The information device 40 may be realized as a single device, or may be realized as multiple devices configured separately from each other.

[0047] The control device 41 is configured with one or more processors that control each element of the information device 40. For example, the control device 41 is configured with one or more types of processors such as a CPU, an SPU, a DSP, an FPGA, or an ASIC.

[0048] The storage device 42 is one or more memories that store programs executed by the control device 41 and various data used by the control device 41. The storage device 42 is configured from a known recording medium such as a magnetic recording medium or a semiconductor recording medium, or a combination of multiple types of recording media. Note that the storage device 42 may be a portable recording medium that is detachable from the information device 40, or a recording medium (e.g., cloud storage) that the control device 41 can write to or read from via the communication network 200, for example.

[0049] The communication device 43 communicates with each of the electronic musical instrument 10 and the machine learning system 30 via the communication network 200. Note that communication between the communication device 43 and the communication network 200 may be either wired or wireless. The communication device 43 receives, for example, music piece data X0 and performance data Y0 transmitted from the electronic musical instrument 10.

[0050] The operating device 44 is an input device that receives instructions from the instructor U2. The operating device 44 may be, for example, a plurality of controls operated by the instructor U2, or a touch panel that detects contact by the instructor U2. The display device 45 displays images under the control of the control device 41. Specifically, the display device 45 displays a time series of notes specified by the performance data Y received by the communication device 43. That is, an image representing the performance by the learner U1 is displayed on the display device 45. Note that the time series of notes specified by the music data X may be displayed in parallel with the notes in the performance data Y. The playback system 46, like the playback system 18 of the electronic musical instrument 10, plays back the musical tones of each note specified by the performance data Y. That is, the musical tones played by the learner U1 are played back by the playback system 46.

[0051] The instructor U2 can check the performance of the piece by the learner U1 by viewing the image displayed on the display device 45 and listening to the sound reproduced by the reproduction system 46. The instructor U2 operates the operation device 44 to input performance tendencies to be pointed out in the performance of the piece by the learner U1. The instructor U2 specifies the performance tendencies of the piece by the learner U1 and the time points within the piece at which the performance tendencies are observed. The instructor U2 selects the performance tendencies from a plurality of options by operating the operation device 44. For example, one of a plurality of performance tendencies such as "off-time key pressing," "pressing a key adjacent to the intended key," "mistaking pitch," "poor performance of leap progressions," "poor performance of chords," and "poor performance of quick short notes such as sixteenth notes" is selected as a point of criticism regarding the performance of the learner U1.

[0052] The control device 41 generates instruction data P in response to instructions from the instructor U2. FIG. 9 is a schematic diagram of the instruction data P. The instruction data P includes tendency data Dt and time data τ for each instruction from the instructor U2. The tendency data Dt is data representing the performance tendency indicated by the instructor U2. The time data τ is data representing the time in the music piece at which the performance tendency is observed. As can be understood from the above explanation, the instruction data P is data representing the time in the music piece and the performance tendency at that time.

[0053] The communication device 43 transmits the instruction data P generated by the control device 41 to the electronic musical instrument 10 and the machine learning system 30. The communication device 13 of the electronic musical instrument 10 receives the instruction data P transmitted from the information device 40. The control device 11 displays the performance tendency represented by the instruction data P on the display device 15. The learner U1 can confirm the instructions (performance tendency) from the instructor U2 by visually checking the image on the display device 15.

[0054] 7, the learning data acquisition unit 81a in the machine learning system 30 receives, via the communication device 33, the music piece data X0 and performance data Y0 transmitted from the electronic musical instrument 10, and the instruction data P transmitted from the information device 40. The learning data acquisition unit 81a generates learning data Ta using the music piece data X0, performance data Y0, and instruction data P. Note that the electronic musical instrument 10 is an example of a "first device," and the information device 40 is an example of a "second device."

[0055] 10 is a flowchart illustrating the specific steps of a process (hereinafter referred to as "preparatory process") Sb in which the learning data acquisition unit 81a generates learning data Ta. For example, the preparation process Sb is started when the communication device 33 receives music piece data X0, performance data Y0, and instruction data P. When the preparation process Sb is started, the learning data acquisition unit 81a acquires the music piece data X0, performance data Y0, and instruction data P from the communication device 33 (Sb1).

[0056] The learning data acquiring unit 81a extracts, as the music piece data Xt, a portion of the music piece data X0 within a section (hereinafter referred to as a "specific section") that includes the time point specified by the time data τ of the instruction data P (Sb2). The specific section is, for example, a section of a predetermined length with the time point specified by the time data τ as its midpoint. The learning data acquiring unit 81a also extracts, as the performance data Yt, a portion of the performance data Y0 within the specific section that includes the time point specified by the time data τ of the instruction data P (Sb3). That is, for each of the music piece data X0 and the performance data Y0, a specific section that includes the time point at which the instructor U2 pointed out the performance tendency is extracted.

[0057] The learning data acquisition unit 81a generates learning control data Ct including the music piece data Xt and performance data Yt generated by the above procedure (Sb4).The learning data acquisition unit 81a then generates learning data Ta by associating the learning control data Ct with the tendency data Dt included in the indication data P (Sb5).

[0058] By repeating the preparation process Sb illustrated above, a large number of learning data Ta are generated for the performance of various pieces of music by a large number of learners U1, including music data Xt and performance data Yt corresponding to specific sections, and tendency data Dt of the performance tendencies pointed out by instructor U2 for the specific sections.

[0059] 11 is a flowchart illustrating specific steps of the learning process Sc in which the control device 31 of the machine learning system 30 establishes a trained model Ma. The learning process Sc can also be expressed as a method for generating a trained model Ma through machine learning (a trained model generation method).

[0060] When the learning process Sc is started, the learning processing unit 82a selects one of the plurality of learning data Ta (hereinafter referred to as "selected learning data Ta") stored in the storage device 32 (Sc1). As illustrated in Fig. 7, the learning processing unit 82a inputs the control data Ct of the selected learning data Ta to an initial or provisional model (hereinafter referred to as "provisional model Ma0") (Sc2), and obtains the tendency data D output by the provisional model Ma0 in response to the input (Sc3).

[0061] The learning processing unit 82a calculates a loss function that represents the error between the trend data D generated by the provisional model Ma0 and the trend data Dt of the selected learning data Ta (Sc4). The learning processing unit 82a updates multiple variables of the provisional model Ma0 so that the loss function is reduced (ideally minimized) (Sc5). The multiple variables are updated according to the loss function using, for example, backpropagation.

[0062] The learning processing unit 82a determines whether a predetermined termination condition is met (Sc6). The termination condition may be, for example, that the loss function falls below a predetermined threshold, or that the amount of change in the loss function falls below a predetermined threshold. If the termination condition is not met (Sc6: NO), the learning processing unit 82a selects the unselected training data Ta as new selected training data Ta (Sc1). That is, the process of updating multiple variables of the provisional model Ma0 (Sc2-Sc5) is repeated until the termination condition is met (Sc6: YES). If the termination condition is met (Sc6: YES), the learning processing unit 82a terminates the update of multiple variables defining the provisional model Ma0 (Sc2-Sc5). The provisional model Ma0 at the time the termination condition is met is determined as the trained model Ma. That is, the multiple variables of the trained model Ma are determined to be the values ​​at the time the learning process Sc is terminated.

[0063] As can be understood from the above explanation, the trained model Ma outputs statistically valid trend data D for unknown control data C based on the underlying relationship between the control data Ct and trend data Dt in multiple training data Ta. In other words, as described above, the trained model Ma is a statistical learning model that has learned the relationship between a performer's performance of a piece of music (control data C) and the performer's performance tendency (trend data D).

[0064] The learning processing unit 82a transmits the trained model Ma established by the above procedure to the information processing system 20 from the communication device 33 (Sc7). Specifically, the learning processing unit 82a transmits multiple variables of the trained model Ma from the communication device 33 to the information processing system 20. The control device 21 of the information processing system 20 stores the trained model Ma received from the machine learning system 30 in the storage device 22. Specifically, multiple variables that define the trained model Ma are stored in the storage device 22.

[0065] B: Second embodiment A second embodiment will be described. Note that, for elements in the following exemplary aspects that have the same functions as those in the first embodiment, the same reference numerals as those in the first embodiment will be used, and detailed descriptions of each will be omitted as appropriate.

[0066] 12 is a block diagram illustrating the functional configuration of an information processing system 20 according to the second embodiment. In the first embodiment, a plurality of practice phrases Z are stored in the storage device 22. In the second embodiment, one reference phrase Zref is stored in the storage device 22 instead of the plurality of practice phrases Z of the first embodiment.

[0067] Like the practice phrase Z in the first embodiment, the reference phrase Zref is time-series data representing a piece of music made up of a plurality of notes. Specifically, the reference phrase Zref is a melody suitable for practicing the electronic musical instrument 10 (for example, part or all of a practice piece). The practice phrase identification unit 73 in the second embodiment generates the practice phrase Z by editing the reference phrase Zref in accordance with the tendency data D generated by the tendency identification unit 72. Specifically, the practice phrase identification unit 73 edits the reference phrase Zref so as to reduce the difficulty of performance for the portion of the reference phrase Zref that is related to the performance tendency specified by the tendency data D.

[0068] 13 is a flowchart illustrating a specific procedure of the specific process Sa in the second embodiment. The specific process Sa in the second embodiment is a process in which step Sa3 in the specific process Sa in the first embodiment is replaced with step Sa13.

[0069] The acquisition of music piece data X and performance data Y by the performance data acquisition unit 71 (Sa1), and the generation of trend data D by the trend identification unit 72 (Sa2) are the same as in the first embodiment. The practice phrase identification unit 73 of the second embodiment generates practice phrase Z by editing the reference phrase Zref stored in the storage device 22 in accordance with the trend data D (Sa13). The process of transmitting practice phrase Z to the electronic musical instrument 10 by the practice phrase identification unit 73 (Sa4) is the same as in the first embodiment. A specific example of editing the reference phrase Zref (Sa13) will be described below.

[0070] For example, if the tendency data D indicates a performance tendency of "difficulty playing chords," the practice phrase identification unit 73 generates the practice phrase Z by changing one or more chords included in the reference phrase Zref. For example, for a chord that includes more than a predetermined number of constituent notes, the practice phrase identification unit 73 omits, for example, one or more constituent notes other than the root note from among the multiple constituent notes. Also, for a chord in which the pitch difference between the lowest note and the highest note exceeds a predetermined value, the practice phrase identification unit 73 omits a predetermined number of constituent notes including the highest note. Omitting constituent notes reduces the difficulty of playing the chord. As shown in the above examples, the editing of the reference phrase Zref by the practice phrase identification unit 73 includes changing the chords.

[0071] Furthermore, if the tendency data D indicates a performance tendency of "not good at jump progressions," the practice phrase identification unit 73 generates the practice phrase Z by omitting or changing the jump progression included in the reference phrase Zref. For example, the practice phrase identification unit 73 omits the latter note of the two notes involved in the jump progression. Furthermore, the practice phrase identification unit 73 changes the latter note of the two notes involved in the jump progression to another note on the lower pitch side. As shown in the above examples, the editing of the reference phrase Zref by the practice phrase identification unit 73 includes omitting or changing the jump progression.

[0072] The reference phrase Zref includes, for example, a specification of a performance technique such as fingering. Specifically, the practice phrase Z includes, for each of a plurality of notes, a specification of the finger number that should play that note. If the tendency data D indicates a performance tendency of "poor at finger crossing," the practice phrase identification unit 73 generates the practice phrase Z by changing the fingering for the reference phrase Zref. For example, assuming that playing keys with the little finger is difficult for beginners, the practice phrase identification unit 73 changes the numbers of notes in the reference phrase Zref that are designated with the little finger to the numbers of fingers other than the little finger. In the electronic musical instrument 10 that receives the edited practice phrase Z, the fingering (finger number for each note) changed by the practice phrase identification unit 73 is displayed on the display device 15 together with the score of the practice phrase Z. As shown in the above example, the editing of the reference phrase Zref by the practice phrase identification unit 73 includes a change in the performance technique of the instrument.

[0073] The second embodiment also achieves the same effects as the first embodiment. Furthermore, in the second embodiment, the practice phrase Z is generated by editing the reference phrase Zref, so that the user U can be provided with an appropriate practice phrase Z according to the level of the user U's performance technique.

[0074] C: Third embodiment FIG. 14 is a block diagram illustrating the functional configuration of the information processing system 20 in the third embodiment. In the first embodiment, a configuration was exemplified in which the practice phrase identification unit 73 identifies the practice phrase Z corresponding to the tendency data D of the user U from among multiple practice phrases Z stored in the storage device 22. The practice phrase identification unit 73 in the third embodiment uses a trained model Mb to identify the practice phrase Z corresponding to the tendency data D. The trained model Mb is an example of a "second trained model."

[0075] As can be understood from the description of the first embodiment, there is a correlation between a performer's performance tendency (tendency data D) and a practice phrase Z that is suitable for that performance tendency. For example, a practice phrase Z corresponding to each tendency data D is a piece of music that is suitable for improving the performance tendency specified by the tendency data D. The trained model Mb is a statistical estimation model that learns the relationship between the tendency data D and the practice phrase Z. The practice phrase identification unit 73 of the third embodiment inputs the tendency data D generated by the tendency identification unit 72 into the trained model Mb to identify a practice phrase Z that corresponds to the performance tendency represented by the tendency data D. For example, the trained model Mb outputs an index of appropriateness for the tendency data D for each of multiple different practice phrases Z (i.e., the degree to which each practice phrase Z is appropriate for the user U's performance tendency). The practice phrase identification unit 73 identifies the practice phrase Z with the largest index from among the multiple practice phrases Z stored in the storage device 22.

[0076] The trained model Mb is configured, for example, by a deep neural network. For example, any type of neural network, such as a recurrent neural network or a convolutional neural network, can be used as the trained model Mb. The trained model Mb may be configured by combining multiple types of deep neural networks. In addition, additional elements, such as a long short-term memory (LSTM), may be incorporated into the trained model Mb.

[0077] The trained model Mb is realized by a combination of a program that causes the control device 21 to execute a calculation to estimate the practice phrase Z from the tendency data D, and a plurality of variables (specifically, weights and biases) that are applied to the calculation. The program and the plurality of variables that realize the trained model Mb are stored in the storage device 22. The numerical values ​​of each of the plurality of variables that define the trained model Mb are set in advance by machine learning.

[0078] 15 is a flowchart illustrating a specific procedure of the specific process Sa in the third embodiment. The specific process Sa in the third embodiment is a process in which step Sa3 in the specific process Sa in the first embodiment is replaced with step Sa23.

[0079] The acquisition of music piece data X and performance data Y by the performance data acquisition unit 71 (Sa1), and the generation of trend data D by the trend identification unit 72 (Sa2) are the same as in the first embodiment. The practice phrase identification unit 73 of the third embodiment inputs the trend data D into the trained model Mb to identify a practice phrase Z (Sa23). The process of transmitting practice phrase Z to the electronic musical instrument 10 by the practice phrase identification unit 73 (Sa4) is the same as in the first embodiment.

[0080] The trained model Mb illustrated above is generated by the machine learning system 30. Figure 16 is a block diagram illustrating a functional configuration of the machine learning system 30 related to the generation of the trained model Mb. The control device 31 executes a program stored in the storage device 32, thereby functioning as multiple elements (a training data acquisition unit 81b and a training processing unit 82b) for establishing the trained model Mb through machine learning.

[0081] The learning processing unit 82b establishes a learned model Mb through supervised machine learning (learning processing Sd described below) using a plurality of pieces of learning data Tb. The learning data acquisition unit 81b acquires a plurality of pieces of learning data Tb. Specifically, the learning data acquisition unit 81b acquires a plurality of pieces of learning data Tb stored in the storage device 32 from the storage device 32. The learning data acquisition unit 81b is an example of a "second learning data acquisition unit," and the learning processing unit 82b is an example of a "second learning processing unit." Furthermore, the learning data Tb is an example of "second learning data."

[0082] Each of the multiple learning data Tb is composed of a combination of learning tendency data Dt and learning practice phrases Zt. The practice phrases Zt of each learning data Tb are songs that are suitable for the performance tendency indicated by the tendency data Dt of that learning data Tb. The combination of the tendency data Dt and the practice phrases Zt is selected, for example, by the creator of the learning data T. The tendency data Dt is an example of "learning tendency data," and the practice phrases Zt are an example of "learning practice phrases."

[0083] 17 is a flowchart illustrating a specific procedure of the learning process Sd in which the control device 31 establishes the learned model Mb. The learning process Sd can also be expressed as a method for generating the learned model Mb by machine learning (a method for generating a learned model).

[0084] When the learning process Sd starts, the learning data acquisition unit 81b selects one of the multiple learning data Tb (hereinafter referred to as "selected learning data Tb") stored in the storage device 32 (Sd1). As illustrated in Fig. 16, the learning processing unit 82b inputs the tendency data Dt of the selected learning data Tb into an initial or provisional model (hereinafter referred to as "provisional model Mb0") (Sd2), and acquires the training phrase Z estimated by the provisional model Mb0 in response to the input (Sd3).

[0085] The learning processing unit 82b calculates a loss function that represents the error between the training phrase Z estimated by the provisional model Mb0 and the training phrase Zt in the selected training data Tb (Sd4). The learning processing unit 82b updates multiple variables of the provisional model Mb0 so that the loss function is reduced (ideally minimized) (Sd5). The multiple variables are updated in accordance with the loss function using, for example, backpropagation.

[0086] The learning processing unit 82b determines whether a predetermined termination condition is met (Sd6). If the termination condition is not met (Sd6: NO), the learning processing unit 82b selects unselected learning data Tb as new selected learning data Tb (Sd1). That is, the process of updating multiple variables of the provisional model Mb0 (Sd2-Sd5) is repeated until the termination condition is met (Sd6: YES). The provisional model Mb0 at the time the termination condition is met (Sd6: YES) is determined to be the trained model Mb.

[0087] As can be understood from the above explanation, the trained model Mb estimates a training phrase Z that is statistically valid for unknown trend data D based on the underlying relationship between trend data Dt and training phrases Zt in multiple training data Tb. In other words, the trained model Mb is a statistical estimation model that has learned the relationship between trend data D and training phrases Z. The training phrase identification unit 73 of the third embodiment identifies training phrases Z by inputting trend data D to the trained model Mb that has learned the relationship between trend data Dt and training phrases Zt.

[0088] The learning processing unit 82b transmits the trained model Mb established by the above procedure to the information processing system 20 from the communication device 33 (Sd7). The control device 21 of the information processing system 20 stores the trained model Mb received from the machine learning system 30 in the storage device 22.

[0089] The third embodiment also achieves the same effects as the first embodiment. Furthermore, in the third embodiment, the practice phrase Z is identified by inputting the trend data D output by the trend identification unit 72 into the trained model Mb. Therefore, a statistically valid practice phrase Z can be identified based on the underlying relationship between the training trend data Dt and the training practice phrase Zt.

[0090] D: Fourth embodiment FIG. 18 is a block diagram illustrating the functional configuration of an electronic musical instrument 10 according to a fourth embodiment. In each of the above-described embodiments, the information processing system 20 includes a performance data acquisition unit 71, a tendency identification unit 72, and a practice phrase identification unit 73. In the fourth embodiment, the electronic musical instrument 10 includes the performance data acquisition unit 71, the tendency identification unit 72, and the practice phrase identification unit 73. The above elements are realized by the control device 11 executing a program stored in the storage device 12. The control device 11 also functions as a presentation processing unit 74.

[0091] The storage device 12 of the electronic musical instrument 10 stores a plurality of pieces of music data X similar to those in the first embodiment, as well as a trained model Ma and a plurality of practice phrases Z. The trained model Ma established by the machine learning system 30 is transferred to the electronic musical instrument 10, and the trained model Ma is stored in the storage device 12. Furthermore, each of the plurality of practice phrases Z corresponds to a different set of tendency data D.

[0092] As in the first embodiment, the performance data acquisition unit 71 acquires performance data Y representing a performance of a piece of music by the user U, and music data X of the piece of music. Specifically, the performance data acquisition unit 71 generates the performance data Y in response to an operation from the user U on the performance device 14. The performance data acquisition unit 71 also acquires music data X of the piece of music to be performed by the user U from the storage device 12. The performance data acquisition unit 71 generates control data C including the music data X and the performance data Y.

[0093] As in the first embodiment, the tendency identification unit 72 generates tendency data D representing the performance tendency of the user U in accordance with the control data C. Specifically, the tendency identification unit 72 identifies the tendency data D by inputting the control data C including the music piece data X and the performance data Y into the trained model Ma.

[0094] As in the first embodiment, the practice phrase identification unit 73 uses the tendency data D identified by the tendency identification unit 72 to identify a practice phrase Z that corresponds to the performance tendency of the user U. Specifically, the practice phrase identification unit 73 searches the storage device 12 for a practice phrase Z that corresponds to the tendency data D identified by the tendency identification unit 72 from among the multiple practice phrases Z stored in the storage device 12.

[0095] The presentation processing unit 74 presents the practice phrase Z identified by the practice phrase identification unit 73 to the user U. Specifically, the presentation processing unit 74 causes the display device 15 to display the musical score of the practice phrase Z. The presentation processing unit 74 may also cause the playback system 18 to play back the performance sound of the practice phrase Z.

[0096] As can be understood from the above explanation, the fourth embodiment achieves the same effects as the first embodiment. Note that the configuration of the second embodiment in which the practice phrase identification unit 73 generates the practice phrase Z by editing the reference phrase Zref, and the configuration in which the practice phrase identification unit 73 identifies the practice phrase Z using the trained model Mb, are similarly applied to the fourth embodiment in which the practice phrase identification unit 73 is built into the electronic musical instrument 10.

[0097] E: Fifth embodiment 19 is a block diagram illustrating the configuration of a performance system 100 according to a fifth embodiment. The performance system 100 includes an electronic musical instrument 10 and an information device 50. The information device 50 is, for example, a smartphone or a tablet terminal. The information device 50 is connected to the electronic musical instrument 10, for example, by wire or wirelessly.

[0098] The information device 50 is realized as a computer system including a control device 51 and a storage device 52. The control device 51 is composed of one or more processors that control each element of the information device 50. For example, the control device 51 is composed of one or more types of processors, such as a CPU, SPU, DSP, FPGA, or ASIC. The storage device 52 is one or more memories that store programs executed by the control device 51 and various data used by the control device 51. The storage device 52 is composed of a known storage medium, such as a magnetic storage medium or a semiconductor storage medium, or a combination of multiple types of storage media. Note that the storage device 52 may be a portable storage medium that is detachable from the information device 50, or a storage medium (e.g., cloud storage) to which the control device 51 can write or read data via, for example, a communication network 200.

[0099] The control device 51 executes the programs stored in the storage device 52 to implement a performance data acquisition unit 71, a tendency identification unit 72, and a practice phrase identification unit 73. The configurations and operations of the performance data acquisition unit 71, the tendency identification unit 72, and the practice phrase identification unit 73 are the same as those illustrated in the first to fourth embodiments. The practice phrase Z identified by the practice phrase identification unit 73 is transmitted to the electronic musical instrument 10. The control device 11 of the electronic musical instrument 10 causes the display device 15 to display the musical score of the practice phrase Z.

[0100] As can be understood from the above explanation, the fifth embodiment also achieves the same effects as the first to fourth embodiments. The information processing system 20 of the first to third embodiments, the electronic musical instrument 10 of the fourth embodiment, and the information device 50 of the fifth embodiment are all examples of the "information processing system 20."

[0101] F: Variation Specific modified embodiments that can be added to each of the embodiments exemplified above are exemplified below. Multiple embodiments arbitrarily selected from the following examples may be combined as appropriate within the scope of not mutually contradicting each other.

[0102] (1) In each of the above-described embodiments, the trend data D is generated using one trained model Ma. However, the trend data D may be generated by selectively using multiple trained models Ma. For example, multiple trained models Ma corresponding to different instruments are prepared. The trend identification unit 72 selects the trained model Ma corresponding to the instrument played by the user U from the multiple trained models Ma and generates the trend data D by inputting the control data C to the trained model Ma. The relationship between the content of the performance by the user U (performance data Y) and the performance tendency of the user U (trend data D) differs for each instrument. By selectively using multiple trained models Ma corresponding to different instruments, it is possible to generate trend data D that appropriately represents the performance tendency of the instrument actually played by the user U.

[0103] (2) In the third embodiment, practice phrase Z is generated using one trained model Mb. However, practice phrase Z may be generated by selectively using multiple trained models Mb. For example, multiple trained models Mb corresponding to different instruments are prepared. The practice phrase identification unit 73 selects the trained model Mb corresponding to the instrument played by the user U from the multiple trained models Mb, and generates practice phrase Z by inputting tendency data D into the trained model Mb.

[0104] (3) Any of a plurality of trained models Ma established by the machine learning system 30 may be selectively transferred to the electronic musical instrument 10 of the fourth embodiment. For example, of a plurality of trained models Ma corresponding to different musical instruments, the trained model Ma corresponding to an instrument specified by a user U of the electronic musical instrument 10 is transferred from the machine learning system 30 to the electronic musical instrument 10. Similarly, any of a plurality of trained models Ma established by the machine learning system 30 may be selectively transferred to the information device 50 of the fifth embodiment. In the third embodiment, any of a plurality of trained models Mb established by the machine learning system 30 may be selectively transferred to the information processing system 20.

[0105] (4) In each of the above-described embodiments, the instruction data P is generated in response to instructions from the instructor U2. However, the control device 11 of the electronic musical instrument 10 may generate the instruction data P in response to instructions from the learner U1. For example, the learner U1 indicates a performance tendency (e.g., a playing style that the learner is not good at) and the time point at which the performance tendency is observed. The control device 11 generates the instruction data P in response to instructions from the user U and transmits the instruction data P to the machine learning system 30 from the communication device 13.

[0106] (5) In the above-described embodiments, the control data C includes music data X and performance data Y. However, the content of the control data C is not limited to these examples. For example, the control data C may include image data of an image captured of the user U playing the electronic musical instrument 10. For example, the control data C includes image data of both hands of the user U during performance. Similarly, the learning control data Ct includes image data of an image of the performer. With the above configuration, an appropriate practice phrase Z that also reflects the performance of the user U can be identified. It is also possible for the control data C to not include music data X. As can be understood from the above explanation, the control data C that includes at least performance data Y is input to the trained model Ma. That is, the tendency identification unit 72 generates tendency data D by inputting the performance data Y to the trained model Ma.

[0107] (6) In the first embodiment, a piece of music suitable for improving the performance tendencies of user U was given as an example of practice phrase Z. However, as in the second embodiment, the practice phrase identification unit 73 may identify a practice phrase Z that is easy to play in parts related to user U's performance tendencies.

[0108] (7) The configuration of the first embodiment in which one of a plurality of practice phrases Z is selected according to the trend data D may be combined with the configuration of the second embodiment in which the reference phrase Zref is edited according to the trend data D. For example, the practice phrase identification unit 73 selects one practice phrase Z according to the trend data D from the plurality of practice phrases Z stored in the storage device 22 as the reference phrase Zref (Sa3), and generates the practice phrase Z by editing the reference phrase Zref according to the trend data D (Sa13). In other words, the trend data D is used both in selecting the practice phrase Z (Sa3) and in editing the reference phrase Zref (Sa13).

[0109] (8) In the second embodiment, the practice phrase identification unit 73 generated the practice phrase Z by editing one reference phrase Zref stored in the storage device 22, but the practice phrase Z may also be generated by selectively using multiple reference phrases Zref stored in the storage device 22. For example, the practice phrase generation unit may generate the practice phrase Z by using the reference phrase Zref of a piece of music selected by the user U of the electronic musical instrument 10 from the multiple reference phrases Zref stored in the storage device 22.

[0110] (9) In the above-described embodiments, an electronic keyboard instrument is exemplified as the electronic musical instrument 10, but the type of instrument played by the user U is arbitrary. For example, the user U may play an electric string instrument such as an electric guitar. The performance data Y may be an audio signal (audio data) representing the vibration of the strings of the electric string instrument, or MIDI-format data generated by analyzing the musical notes produced by the electric string instrument. Performance tendencies for electric string instruments include, for example, "insufficient muting of notes that should be muted" and "strings other than those corresponding to the intended notes are sounding." For example, if the user U plays a wind instrument such as a trumpet or saxophone, performance tendencies represented by the tendency data D may include "unstable volume of musical notes" and "inaccurate pitch." For example, if the user U plays a percussion instrument such as a drum, performance tendencies represented by the tendency data D may include "offset of striking timing" and "difficulty in striking repeatedly at short intervals."

[0111] (10) In the above-described embodiments, a deep neural network is exemplified as the trained model Ma, but the trained model Ma is not limited to a deep neural network. For example, a statistical estimation model such as an HMM (Hidden Markov Model) or an SVM (Support Vector Machine) may be used as the trained model Ma. A trained model Ma using an SVM is described in detail below.

[0112] For example, an SVM is prepared for each of all possible combinations of two performance tendencies selected from multiple performance tendencies. For the SVM corresponding to the combination of two performance tendencies, a hyperplane in multidimensional space is established by machine learning (learning process Sc). The hyperplane is a boundary surface that separates the space where control data C corresponding to one of the two performance tendencies is distributed from the space where control data C corresponding to the other performance tendency is distributed. The trained model Ma is composed of multiple SVMs corresponding to different combinations of performance tendencies (multi-class SVM).

[0113] The tendency identification unit 72 inputs the control data C to each of the multiple SVMs of the trained model Ma. The SVM corresponding to each combination selects one of two types of performance tendencies associated with that combination depending on which of the two spaces separated by the hyperplane the control data C is located in. The selection of a performance tendency is performed in the same manner in each of the multiple SVMs corresponding to different combinations. The tendency identification unit 72 generates tendency data D that represents the performance tendency that is selected the most frequently by the multiple SVMs among the multiple types of performance tendencies.

[0114] As can be understood from the above examples, regardless of the type of trained model Ma, the tendency identification unit 72 functions as an element that generates tendency data D representing the performance tendency of the user U by inputting control data C into the trained model Ma. Note that while the above description has focused on the trained model Ma, a statistical estimation model such as an HMM or an SVM is also used for the trained model Mb of the third embodiment in a similar manner.

[0115] (11) In the above-described embodiments, supervised machine learning using multiple pieces of training data T has been exemplified as the training process Sc. However, the trained model Ma may be established by unsupervised machine learning that does not require training data T, or by reinforcement learning that maximizes rewards. An example of unsupervised machine learning is machine learning that uses well-known clustering. Similarly, the trained model Mb of the third embodiment may be established by unsupervised machine learning or reinforcement learning.

[0116] (12) In each of the above-described embodiments, the machine learning system 30 established the trained model Ma. However, the functions (training data acquisition unit 81a and learning processing unit 82a) by which the machine learning system 30 establishes the trained model Ma may be installed in the information processing system 20 of the first to third embodiments, the electronic musical instrument 10 of the fourth embodiment, or the information device 50 of the fifth embodiment. The same applies to the trained model Mb of the third embodiment. That is, the functions (training data acquisition unit 81b and learning processing unit 82b) by which the machine learning system 30 establishes the trained model Mb may be installed in the information processing system 20 of the third embodiment, the electronic musical instrument 10 of the fourth embodiment, or the information device 50 of the fifth embodiment.

[0117] (13) In the above-described embodiments, the trained model Ma is used to generate the tendency data D corresponding to the control data C. However, the use of the trained model Ma may be omitted. For example, a table in which each of the plurality of control data C and each of the plurality of tendency data D are associated with each other may be used to generate the tendency data D. The table in which the association between the control data C and the tendency data D is registered is stored, for example, in the storage device 22 of the first embodiment, the storage device 12 of the fourth embodiment, or the storage device 52 of the fifth embodiment. The tendency identification unit 72 searches the table for the tendency data D corresponding to the control data C generated by the performance data acquisition unit 71.

[0118] (14) In the above-described embodiments, a trained model Ma is used that has trained the relationship between the control data C, which includes music piece data X and performance data Y, and the tendency data D. However, the configuration and method for generating the tendency data D from the control data C are not limited to the above examples. For example, a lookup table in which the tendency data D is associated with each of a plurality of different control data C may be used by the tendency identification unit 72 to generate the tendency data D. The lookup table is a data table in which the correspondence between the control data C and the tendency data D is registered, and is stored in, for example, the storage device 22 (storage device 12 in the fourth embodiment). The tendency identification unit 72 searches the lookup table for the control data C corresponding to the combination of music piece data X and performance data Y, and obtains the tendency data D associated with the control data C from the lookup table among the plurality of tendency data D.

[0119] (15) In the third embodiment, a trained model Mb that has learned the relationship between trend data D and practice phrases Z is used. However, the configuration and method for generating practice phrases Z from trend data D are not limited to the above example. For example, a reference table in which practice phrases Z are associated with multiple different sets of trend data D may be used by the practice phrase identification unit 73 to generate practice phrases Z. The reference table is a data table in which the correspondence between trend data D and practice phrases Z is registered, and is stored in, for example, the storage device 22 (storage device 12 in the fourth embodiment). The practice phrase identification unit 73 searches the reference table for the practice phrase Z that corresponds to the trend data D, and obtains the practice phrase Z that is associated with the trend data D from the reference table among the multiple practice phrases Z.

[0120] (16) In each of the above-described embodiments, the performance data acquisition unit 71 acquires the performance data Y representing the performance of the user U from the electronic musical instrument 10. However, the method by which the performance data acquisition unit 71 acquires the performance data Y is not limited to the above examples. For example, the performance data acquisition unit 71 does not need to acquire the performance data Y in real time in parallel with the performance on the performance device 14. For example, the performance data acquisition unit 71 may acquire performance data Y that records past performances by the user U from the electronic musical instrument 10. In other words, in the present disclosure, it is irrelevant whether the performance data acquisition unit 71 acquires the performance data Y in real time in response to a performance by the user U.

[0121] Furthermore, for example, the performance data acquisition unit 71 does not need to receive the performance data Y representing the sequence of notes played by the user U from the electronic musical instrument 10. For example, the performance data acquisition unit 71 may receive video data of the user U's performance via the communication device 23 and generate the performance data Y by analyzing the video data. In other words, the "acquisition" of the performance data Y by the performance data acquisition unit 71 includes not only the process of receiving the performance data Y from an external device such as the electronic musical instrument 10, but also the process of generating the performance data Y from information such as the video data.

[0122] (17) In the above-described embodiments, the learning data acquisition unit 81a acquires performance data Y0 representing a performance of a piece of music by the learner U1 and instruction data P representing instructions from the instructor U2. However, the method by which the learning data acquisition unit 81a acquires the learning data Ta is not limited to the above examples. For example, the learning data acquisition unit 81a does not need to acquire the performance data Y0 and instruction data P (and further the learning data Ta) in parallel with the performance by the learner U1 and the instruction by the instructor U2. For example, the learning data acquisition unit 81a may acquire performance data Y0 recording a past performance by the learner U1 and instruction data P recording a past instruction by the instructor U2. In other words, the present disclosure does not consider whether the learning data acquisition unit 81a acquires the performance data Y0 and instruction data P in real time in response to the performance by the learner U1 and the instruction by the instructor U2.

[0123] Furthermore, for example, the learning data acquisition unit 81a does not need to receive performance data Y0 representing a sequence of notes played by learner U1 from the electronic musical instrument 10. For example, the learning data acquisition unit 81a may receive video data of learner U1's performance via the communication device 23 and generate performance data Y0 by analyzing the video data. In other words, the "acquisition" of performance data Y0 by the learning data acquisition unit 81a includes not only the process of receiving performance data Y0 from an external device such as the electronic musical instrument 10, but also the process of generating performance data Y0 from information such as the video data.

[0124] Similarly, the learning data acquiring unit 81a does not need to receive the instruction data P representing the instruction by the instructor U2 from the information device 40. For example, the learning data acquiring unit 81a may receive video data of the instructor U2's instruction via the communication device 23 and analyze the video data to generate the instruction data P. In other words, the "acquisition" of the instruction data P by the learning data acquiring unit 81a includes not only the process of receiving the instruction data P from an external device such as the information device 40, but also the process of generating the instruction data P from information such as the video data.

[0125] (18) In the above-described embodiments, the learning data acquisition unit 81a extracts, as performance data Yt, a portion of the performance data Y0 transmitted from the electronic musical instrument 10 within a specific interval including a time point designated by the time data τ of the instruction data P. However, the learning performance data Yt may be transmitted from the electronic musical instrument 10 to the machine learning system 30. For example, the control device 11 of the electronic musical instrument 10 receives the instruction data P from the information device 40, and transmits, as performance data Yt, a portion of the performance data Y0 within a specific interval corresponding to the time data τ of the instruction data P from the communication device 13 to the machine learning system 30. The learning data acquisition unit 81a receives the performance data Yt transmitted from the electronic musical instrument 10 via the communication device 33. With the above configuration, the machine learning system 30 does not need to acquire the time data τ from the information device 40. That is, the time data τ may be omitted from the instruction data P transmitted from the information device 40 to the machine learning system 30.

[0126] While the above description focuses on the performance data Yt, learning music data Xt may also be transmitted from the electronic musical instrument 10 to the machine learning system 30. For example, the control device 11 of the electronic musical instrument 10 transmits a portion of the music data X0 within a specific section corresponding to the time data τ of the pointed-out data P as music data Xt from the communication device 13 to the machine learning system 30. The learning data acquisition unit 81a receives the music data Xt transmitted from the electronic musical instrument 10 via the communication device 33.

[0127] (19) As described above, the functions exemplified in each of the above embodiments (the performance data acquisition unit 71, the tendency identification unit 72, and the practice phrase identification unit 73) are realized by the cooperation of one or more processors constituting the control device and a program stored in a storage device. The above programs can be provided in a form stored on a computer-readable recording medium and installed on a computer. The recording medium is, for example, a non-transitory recording medium. A good example is an optical recording medium (optical disk) such as a CD-ROM, but it also includes any known type of recording medium, such as a semiconductor recording medium or a magnetic recording medium. Note that a non-transitory recording medium includes any recording medium other than a transient, propagating signal, and does not exclude volatile recording media. Furthermore, in a configuration in which a distribution device distributes a program via a communication network 200, the recording medium storing the program in the distribution device corresponds to the non-transitory recording medium described above.

[0128] G: Notes From the above-described exemplary embodiments, the following configurations can be understood, for example.

[0129] An information processing system according to one aspect (Aspect 1) includes a performance data acquisition unit that acquires performance data representing a user's performance of a musical piece; a tendency identification unit that generates tendency data representing the user's performance by inputting the performance data acquired by the performance data acquisition unit into a first trained model that has learned the relationship between learning performance data representing a performance of a reference musical piece and learning tendency data representing the performance tendency represented by the learning performance data; and a practice phrase identification unit that identifies practice phrases corresponding to the tendency data generated by the tendency identification unit. According to the above aspect, inputting the performance data representing the user's performance of a musical piece into the first trained model generates tendency data representing the user's performance tendency, and practice phrases corresponding to the user's performance tendency are identified according to the tendency data. Therefore, by playing the practice phrases, effective practice corresponding to the user's performance tendency is realized.

[0130] "Performance data" is data in any format that represents a performance by a user. Examples of performance data include music data (e.g., MIDI data) that represents a time series of notes played by the user, and audio data that represents the sounds produced by an instrument played by the user. Performance data may also include video data that captures the user's performance.

[0131] "Tendency data" is data in any format that represents the tendencies of a user's performance. "Performance tendencies" may be, for example, the tendency of a user to make performance mistakes or tendencies in playing techniques that the user finds difficult. For example, the tendency data may specify one of multiple types of tendencies related to performance mistakes or playing techniques.

[0132] A "practice phrase" is a sequence of notes (melody) for a user to practice playing. A "practice phrase according to the user's playing tendencies" is, for example, a sequence of notes suitable for overcoming playing mistakes that tend to occur in the user's playing or a playing style that the user is not good at. A practice phrase may be an entire piece of music or a part of that piece of music.

[0133] In a specific example (Aspect 2) of Aspect 1, the first trained model is a model that has learned the relationship between the training control data, which includes training music piece data representing the score of the reference music piece and the training performance data, and the training tendency data, and the tendency identification unit generates the tendency data by inputting control data, which includes the performance data and music piece data representing the score of the music piece, into the first trained model. According to the above aspect, since the control data includes music piece data in addition to the performance data, it is possible to generate appropriate tendency data that reflects the relationship (e.g., similarity or difference) between the performance data and the music piece data.

[0134] In a specific example (Aspect 3) of Aspect 1 or Aspect 2, the practice phrase identification unit selects a practice phrase that corresponds to the tendency represented by the tendency data from among a plurality of practice phrases corresponding to different performance tendencies. According to the above aspect, a practice phrase that corresponds to the user's performance tendency is selected from a plurality of practice phrases, thereby reducing the processing load on the practice phrase identification unit in identifying the practice phrase.

[0135] In a specific example (Aspect 4) of Aspect 1 or Aspect 2, the practice phrase identification unit generates the practice phrase by editing the reference phrase in accordance with the tendency represented by the tendency data. According to the above aspect, since the practice phrase is generated by editing the reference phrase, it is possible to provide the user with a practice phrase appropriate to the user's level of performance technique.

[0136] "Editing the reference phrase" refers to the process of changing the reference phrase so that the difficulty of playing it changes according to the tendency indicated by the tendency data. Examples of "editing" include simplifying the chords in the reference phrase (for example, omitting the constituent notes of the chord), omitting a leap progression (a section in which two notes with a large pitch difference are played one after the other), or simplifying the fingering when playing.

[0137] In a specific example of Aspect 4 (Aspect 5), the reference phrase includes a time sequence of chords, and editing the reference phrase includes changing the chords. In another specific example of Aspect 4 (Aspect 6), the reference phrase includes a leap progression with a pitch difference exceeding a predetermined value, and editing the reference phrase includes omitting or changing the leap progression. In another specific example of Aspect 4 (Aspect 7), the reference phrase includes a specification of a playing technique for an instrument, and editing the reference phrase includes changing the playing technique. "Playing technique" refers to the way an instrument is played. Examples of "playing techniques" include fingering for keyboard instruments or string instruments, and special playing techniques such as hammering, pulling, or cutting for string instruments such as guitar or bass.

[0138] In a specific example (Aspect 8) of any of Aspects 1 to 4, the practice phrase identification unit identifies the practice phrase by inputting the trend data output by the trend identification unit into a second trained model that has learned the relationship between learning trend data that represents performance trends and learning practice phrases that correspond to the trends represented by the learning trend data. According to the above aspect, the practice phrase identification unit identifies the practice phrase by inputting the trend data output by the trend identification unit into the second trained model. Therefore, it is possible to identify practice phrases that are statistically valid based on the underlying relationship between the learning trend data and the learning practice phrases.

[0139] In a specific example (aspect 9) of aspect 8, the practice phrase identification unit identifies the practice phrase by selectively using one of a plurality of second trained models corresponding to different instruments. According to the above aspect, compared to a configuration using only one second trained model, it is possible to identify practice phrases appropriate for the instrument that the user actually plays.

[0140] In a specific example (Aspect 10) of any of Aspects 1 to 9, the tendency identification unit selectively uses one of a plurality of first trained models corresponding to different musical instruments to generate the tendency data. According to the above aspect, since a plurality of first trained models corresponding to different musical instruments are selectively used to generate the tendency data, it is possible to generate tendency data that appropriately represents the playing tendency of the instrument actually played by the user, compared to a configuration that uses only one first trained model.

[0141] An electronic musical instrument according to one aspect (aspect 11) of the present disclosure comprises a performance acceptance unit that accepts a musical piece performance by a user; a performance data acquisition unit that acquires performance data representing the performance accepted by the performance acceptance unit; a tendency identification unit that inputs the performance data acquired by the performance data acquisition unit into a first learned model that has learned the relationship between learning performance data representing the musical piece performance and learning tendency data that represents the performance tendency represented by the learning performance data, and outputs tendency data that represents the performance tendency of the user from the first learned model; a practice phrase identification unit that uses the tendency data output by the tendency identification unit to identify practice phrases that correspond to the performance tendency of the user; and a presentation processing unit that presents the practice phrases to the user.

[0142] The presentation processing unit presents the practice phrase to the user in a manner that the user can perceive visually or audibly. For example, the presentation processing unit may be an element that displays the musical score of the practice phrase on a display device, or an element that causes a sound output device to output the sound of the practice phrase being played.

[0143] An information processing method according to one aspect (aspect 12) of the present disclosure acquires performance data representing a musical piece played by a user, and inputs the acquired performance data into a first trained model that has learned the relationship between training performance data representing the musical piece performance and training trend data representing the performance trend represented by the training performance data, thereby generating trend data representing the performance trend of the user and identifying practice phrases according to the trend data.

[0144] In a specific example (Aspect 13) of Aspect 12, the practice phrase identification comprises selecting a practice phrase that corresponds to the tendency represented by the tendency data from among a plurality of practice phrases corresponding to different performance tendencies. Also, in a specific example (Aspect 14) of Aspect 12, the practice phrase identification comprises generating the practice phrase by editing a reference phrase in accordance with the tendency represented by the tendency data. In another specific example (Aspect 15) of Aspect 12, the practice phrase identification comprises inputting the tendency data into a second trained model that has learned the relationship between training tendency data representing performance tendencies and training practice phrases corresponding to the tendencies represented by the training tendency data.

[0145] A machine learning system according to one aspect (aspect 16) of the present disclosure includes a first learning data acquisition unit that acquires first learning data including learning performance data representing a user's performance of a piece of music and learning tendency data representing the performance tendency represented by the instruction data, and a first learning processing unit that establishes a first trained model that learns the relationship between the learning performance data and the learning tendency data through machine learning using the first learning data. According to the above aspect, it is possible to generate trend data that is statistically valid for performance data based on the underlying relationship between the learning performance data and the learning tendency data using the first trained model.

[0146] In a specific example (aspect 17) of aspect 16, the first learning data acquisition unit acquires performance data representing a performance of the musical piece by the user and instruction data representing a time point in the musical piece and a tendency of the performance at that time point, and generates the first learning data including the learning performance data representing the performance within a section of the performance data that includes the time point represented by the instruction data, and the learning tendency data representing the tendency of the performance represented by the instruction data. According to the above aspect, it is not necessary for a source of the performance data (e.g., the first device) to extract a section of the user's performance that corresponds to the time point represented by the instruction data.

[0147] In a specific example (Aspect 18) of Aspect 17, the first learning data acquisition unit acquires the performance data from a first device and acquires the instruction data from a second device separate from the first device. According to the above aspect, data for machine learning can be prepared using data (performance data and instruction data) acquired from, for example, the first device and the second device, which are located remotely from each other. The first device is, for example, a terminal device used by a learner who practices playing a musical instrument, and the second device is, for example, a terminal device used by an instructor who evaluates and instructs the learner's performance.

[0148] In a specific example (Aspect 19) of any of Aspects 16 to 18, the first trained model is a model that has learned the relationship between the training control data, which includes training music piece data representing the musical score of the reference music piece and the training performance data, and the training trend data. In the above aspects, since the training control data includes the training music piece data in addition to the training performance data, it is possible to establish a first trained model that can generate appropriate trend data that reflects the relationship (e.g., similarity or difference) between the training performance data and the training music piece data.

[0149] In a specific example (Aspect 20) of any of Aspects 16 to 19, the system further includes a second learning data acquisition unit that acquires a plurality of second learning data including learning tendency data that represents performance tendencies and learning practice phrases that correspond to the tendencies represented by the learning tendency data, and a second learning processing unit that establishes a second trained model that learns the relationship between the learning tendency data and the learning practice phrases in the second learning data through machine learning using the plurality of second learning data.

[0150] A machine learning method according to one aspect (aspect 21) of the present disclosure acquires performance data representing a performance of a piece of music by a user and comment data representing a time point within the piece of music and the performance tendency at that time point, and establishes a first trained model that learns the relationship between the training performance data and the training tendency data through machine learning using first training data that includes training performance data representing the performance within a section of the performance data that includes the time point represented by the comment data, and training tendency data representing the performance tendency represented by the comment data. [Explanation of symbols]

[0151] 100...performance system, 10...electronic musical instrument, 11,21,31,41,51...control device, 12,22,32,42,52...storage device, 13,23,33,43...communication device, 14...performance device, 15,45...display device, 16...sound source device, 17...sound emission device, 18,46...playback system, 20...information processing system, 30...machine learning system, 40...information device, 44...operation device, 50...information device, 71...performance data acquisition unit, 72...trend identification unit, 73...practice phrase identification unit, 74...presentation processing unit, 81a,81b...learning data acquisition unit, 82a,82b...learning processing unit.

Claims

1. a receiving unit for receiving performance data representing a performance of a piece of music by a learner; a generation unit that uses the performance data to generate tendency data representing a performance tendency selected by an instructor from a plurality of options for the performance; An information processing system comprising:

2. a display unit that displays the performance tendency represented by the tendency data; The information processing system of claim 1 further comprising:

3. a learning data acquisition unit that acquires a plurality of learning data each including the performance data and the tendency data; a learning processing unit that establishes a learned model that learns a relationship between the performance data and the tendency data in the plurality of learning data by machine learning using the plurality of learning data; The information processing system of claim 1 further comprising:

4. each of the plurality of learning data includes control data including music piece data representing a score of the music piece and the performance data, and the tendency data; The trained model is a model that has learned the relationship between the control data and the trend data. The information processing system of claim 3.

5. a transmitting unit that transmits performance data representing a performance of a piece of music by a learner; a receiving unit that receives tendency data representing a performance tendency selected by an instructor from a plurality of options for the performance using the performance data; An electronic musical instrument comprising:

6. receiving performance data representing a performance of a piece of music by a learner; Using the performance data, tendency data is generated that represents a performance tendency selected by the instructor from a plurality of options regarding the performance. An information processing method implemented by a computer system.

7. A receiving unit that receives performance data representing a performance of a piece of music by a learner; a generation unit that uses the performance data to generate tendency data representing a performance tendency selected by an instructor from a plurality of options for the performance; A program that makes a computer system function as a

8. a learning data acquisition unit that acquires a plurality of learning data, each of which includes performance data representing a performance of a piece of music by a learner and tendency data representing a performance tendency selected by an instructor from a plurality of options for the performance, using the performance data; a learning processing unit that establishes a learned model that learns a relationship between the performance data and the tendency data in the plurality of learning data by machine learning using the plurality of learning data; A machine learning system comprising:

Citation Information

Patent Citations

  • Performance evaluation system of electronic musical instrument

    JP2005055635A

  • Timbre and / or effect setting device and program

    JP2007140308A

  • Automatic accompaniment for voice melodies

    JP2010538335A

  • Electronic musical instrument and electronic musical instrument system

    JP2018109690A

  • Information processing device for musical-score data

    WO2020031544A1