Fingering display device, training device, fingering display method, and training method

A machine learning-based fingering suggestion device addresses the challenge of determining optimal fingerings by using trained models to estimate and suggest appropriate fingerings for musical instrument players, considering individual performer characteristics and style, thereby improving performance.

JP7841641B2Active Publication Date: 2026-04-07YAMAHA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies struggle to provide appropriate finger movements for musical instrument players, especially when determining optimal fingerings for musical pieces, as there are numerous combinations and no single optimal solution.

Method used

A fingering suggestion device using a machine learning model that estimates appropriate fingerings based on time-series data and performer identifiers, constructed through training with reference performer data, to suggest optimal fingerings for musical instrument players.

Benefits of technology

The device provides accurate and personalized fingering suggestions for musical instrument players, considering their physical characteristics and playing style, enhancing their performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007841641000001
    Figure 0007841641000001
  • Figure 0007841641000002
    Figure 0007841641000002
  • Figure 0007841641000003
    Figure 0007841641000003
Patent Text Reader

Abstract

To provide a fingering presentation device, training device, fingering presentation method, and training method for presenting fingering positions when playing a musical instrument.SOLUTION: In a presented fingering presentation device 100 comprising a training device 10 and a fingering indication device 20, the fingering indication device includes a reception unit and an estimation unit. The reception unit receives time-series data including a sequence of multiple musical notes and a performer identifier indicating a performer who plays the sequence of musical notes. The estimation unit estimates musical note information using a trained model. The note information indicates the notes that are objects of fingering to be assigned from the note sequence based on the performer identifier. The trained model is a machine learning model that has learned an input-output relationship between input time-series data including a reference note sequence consisting of multiple notes and a reference performer identifier indicating a reference performer who plays the reference note sequence, and output note information indicating the notes that are objects of the fingering by the reference performer to be assigned from the reference note sequence.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a finger movement presentation device, a training device, a finger movement presentation method, and a training method for presenting finger movements when playing a musical instrument.

Background Art

[0002] Devices for assisting in the practice of playing a musical instrument are known. For example, in the information processing device described in Patent Document 1, the playing technique level of a player is calculated, and based on the calculated playing technique level, musical pieces that the player can play are presented. However, when the player is inexperienced, it is not easy to appropriately determine the fingerings (hereinafter referred to as finger movements) when playing each note on the musical instrument. On the other hand, Patent Document 2 describes a finger movement determination method for determining finger movements at each note of a note sequence based on a probability model.

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0003] According to Patent Document 2, a player can recognize finger movements in musical instrument playing based on a probability model. However, in reality, there are innumerable combinations of finger movements, and there is not just one optimal finger movement for playing a musical piece. Therefore, it is desired that more appropriate finger movements be presented.

[0004] An object of the present invention is to provide a finger movement presentation device, a training device, a finger movement presentation method, and a training method capable of presenting appropriate finger movements when playing a musical instrument.

Means for Solving the Problems

[0005] A fingering suggestion device according to the first aspect of the present invention comprises a receiving unit that receives time-series data including a sequence of multiple notes and a performer identifier indicating a performer who plays the sequence of notes, and an estimation unit that uses a trained model to estimate note information indicating the notes to which fingering is to be assigned from the sequence of notes based on the performer identifier, wherein the trained model is a machine learning model that has learned the input-output relationship between input time-series data including a reference sequence of multiple notes and a reference performer identifier indicating a reference performer who plays the reference sequence of notes, and output note information indicating the notes to which fingering is to be assigned from the reference sequence by the reference performer.

[0006] A training device according to the second aspect of the present invention includes: a first acquisition unit that acquires input time-series data including a reference note sequence consisting of multiple notes and a reference performer identifier indicating a reference performer who plays the reference note sequence; a second acquisition unit that acquires output note information indicating the notes to which the reference performer's fingerings are to be assigned from the reference note sequence; and a construction unit that constructs a trained model that has learned the input-output relationship between the input time-series data and the output note information.

[0007] A fingering suggestion method according to the third aspect of the present invention receives time-series data including a sequence of multiple notes and a performer identifier indicating the performer playing the sequence of notes, and uses a trained model to estimate note information indicating the note to which fingering is to be assigned from the sequence of notes based on the performer identifier. The trained model is a machine learning model that has learned the input-output relationship between input time-series data including a reference sequence of multiple notes and a reference performer identifier indicating the reference performer playing the reference sequence, and output note information indicating the note to which fingering is to be assigned from the reference sequence by the reference performer, and is executed by a computer.

[0008] A training method according to the fourth aspect of the present invention involves acquiring input time-series data including a reference note sequence consisting of multiple notes and a reference performer identifier indicating a reference performer who plays the reference note sequence, acquiring output note information indicating the notes to which the reference performer's fingering is to be assigned from the reference note sequence, constructing a trained model that has learned the input-output relationship between the input time-series data and the output note information, and executing it by a computer. [Effects of the Invention]

[0009] According to the present invention, it is possible to provide appropriate fingerings for playing a musical instrument. [Brief explanation of the drawing]

[0010] [Figure 1] Figure 1 is a block diagram showing the configuration of a processing system including a finger placement suggestion device and a training device according to a first embodiment of the present invention. [Figure 2] Figure 2 shows an example of each training data set. [Figure 3] Figure 3 is a block diagram showing the configuration of the training device and the finger placement suggestion device. [Figure 4] Figure 4 shows an example of an auxiliary musical score displayed on the display unit. [Figure 5] Figure 5 is a flowchart showing an example of the training process using the training device shown in Figure 3. [Figure 6] Figure 6 is a flowchart showing an example of the finger placement suggestion process using the finger placement suggestion device shown in Figure 3. [Figure 7] Figure 7 shows another example of input time series data. [Figure 8] Figure 8 shows an example of input time series data in a modified example. [Figure 9] Figure 9 shows an example of output finger information in a modified example. [Figure 10] Figure 10 shows an example of input time-series data in the second embodiment. [Figure 11] Figure 11 shows an example of output signal information in the third embodiment. [Figure 12] Figure 12 is a flowchart showing an example of finger placement guidance in a modified example. [Figure 13] Figure 13 shows an example of finger information estimated in step S24 of the finger placement suggestion process. [Modes for carrying out the invention]

[0011] [1] First Embodiment (1) Configuration of the Processing System Hereinafter, the operation instruction presentation device, training device, operation instruction presentation method, and training method according to the embodiments of the present invention will be described in detail with reference to the drawings. FIG. 1 is a block diagram showing the configuration of a processing system including an operation instruction presentation device and a training device according to the first embodiment of the present invention. As shown in FIG. 1, the processing system 100 includes a RAM (Random Access Memory) 110, a ROM (Read Only Memory) 120, a CPU (Central Processing Unit) 130, a storage unit 140, an operation unit 150, and a display unit 160.

[0012] The processing system 100 is realized by a computer such as a personal computer, a tablet terminal, or a smartphone. Alternatively, the processing system 100 may be realized by the cooperative operation of a plurality of computers connected by a communication path such as Ethernet, or may be realized by an electronic musical instrument having a performance function such as an electronic piano.

[0013] The RAM 110, ROM 120, CPU 130, storage unit 140, operation unit 150, and display unit 160 are connected to a bus 170. The training device 10 and the operation instruction presentation device 20 are constituted by the RAM 110, ROM 120, and CPU 130. In the present embodiment, the training device 10 and the operation instruction presentation device 20 are constituted by a common processing system 100, but may be constituted by separate processing systems.

[0014] The RAM 110 is, for example, composed of a volatile memory and is used as a working area for the CPU 130. The ROM 120 is, for example, composed of a non-volatile memory and stores a training program and an operation instruction presentation program. The CPU 130 performs training processing by executing the training program stored in the ROM 120 on the RAM 110. In addition, the CPU 130 performs operation instruction presentation processing by executing the operation instruction presentation program stored in the ROM 120 on the RAM 110. Details of the training processing and the operation instruction presentation processing will be described later.

[0015] The training program or the fingering guidance program may be stored in the storage unit 140 instead of the ROM 120. Alternatively, the training program or the fingering guidance program may be provided in a form stored in a computer-readable storage medium and installed in the ROM 120 or the storage unit 140. Alternatively, when the processing system 100 is connected to a network such as the Internet, the training program or the fingering guidance program distributed from a server on the network (including a cloud server) may be installed in the ROM 120 or the storage unit 140.

[0016] The storage unit 140 includes a storage medium such as a hard disk, an optical disk, a magnetic disk, or a memory card, and stores the trained model M and a plurality of training data D. The trained model M or each training data D may not be stored in the storage unit 140 but may be stored in a computer-readable storage medium. Alternatively, when the processing system 100 is connected to a network, the trained model M or each training data D may be stored in a server on the network.

[0017] (2) Training data The trained model M is a machine learning model trained to present fingering when a user of the fingering guidance device 20 (hereinafter referred to as a performer) plays a music piece on an instrument, and is constructed using a plurality of training data D. The user of the training device 10 can generate the training data D by operating the operation unit 150. The training data D is data created based on the performance knowledge or performance style of a reference performer. The reference performer has relatively high skills in playing a music piece. The reference performer may be an instructor or a teacher of a performer in playing a music piece.

[0018] Training data D represents a pair of input time-series data and output finger information. The input time-series data represents a reference note sequence consisting of multiple notes. The input time-series data may also be image data representing an image of a musical score. The output finger information indicates the fingers of the reference performer used when playing each note in the reference note sequence on an instrument, and can be used to present the fingering when playing the reference note sequence. The output finger information may also be a unique number assigned to each finger. In this example, the thumb, index finger, middle finger, ring finger, and little finger are assigned the numbers "1" through "5", respectively.

[0019] Here, the optimal fingering for playing a piece of music varies depending on the performer's physical characteristics or their playing style. Therefore, in this embodiment, the input time-series data further includes a reference performer identifier indicating the classification (category) of the reference performer playing the reference note sequence. The reference performer identifier is determined to be different for at least one of the reference performer's physical characteristics or their playing style. The reference performer's physical characteristics include, for example, the size of the performer's hands (finger length), age, gender, or whether they are an adult or a child.

[0020] Figure 2 shows an example of each training data D. The example in Figure 2 shows a portion of the input time-series data and output finger information when a reference performer plays the piano. As shown in Figure 2, the input time-series data A includes elements A0 to A16. Element A0 corresponds to the reference performer identifier and is represented by a different string for at least one of the reference performer's physical characteristics and the style of playing by the reference performer. Elements A1 to A16 correspond to the reference note sequence. In this example, element A0 is placed at the beginning of the input time-series data A, i.e., before the reference note sequence (elements A1 to A16), but it may be placed at any position in the input time-series data A.

[0021] In elements A1, A3, A5, ..., A15, "L" means left hand, the numbers mean the numbers assigned to the keys, and "on" and "off" mean pressing and releasing the keys, respectively. In elements A2, A4, A6, ..., A16, "wait" means waiting, and the numbers mean the length of time. Therefore, elements A1 to A4 mean pressing the key numbered "66" and holding it for 13 units of time, then releasing the key numbered "66" and holding it for 2 units of time.

[0022] Output finger information B contains elements B0 to B16, which correspond to elements A0 to A16 of input time-series data A. Element B0 indicates the reference performer identifier and is represented by the same string as element A0. In elements B1, B3, B5, ..., B15, "L" means left hand, the numbers mean the numbers assigned to the fingers, and "down" and "up" mean pushing up and pushing down, respectively. In elements B2, B4, B6, ..., B16, "wait" means waiting, and the numbers mean the length of time. Therefore, elements B1 to B4 mean pressing down the middle finger of the left hand and waiting for 13 units of time, then pushing up the middle finger of the left hand and holding it for 2 units of time.

[0023] The training data D in Figure 2 is generated to show the fingering of the left hand, but the embodiment is not limited to this. The training data D may be generated to show the fingering of the right hand, or to show the fingering of the left hand and the right hand separately. In the input time-series data A and the output finger information B elements for showing the fingering of the right hand, the letter "R" may be used instead of the letter "L".

[0024] (3) Training device and finger placement display device Figure 3 is a block diagram showing the configuration of the training device 10 and the finger placement suggestion device 20. As shown in Figure 3, the training device 10 includes a first acquisition unit 11, a second acquisition unit 12, and a construction unit 13 as functional units. The functional units of the training device 10 are realized when the CPU 130 in Figure 1 executes the training program. At least a portion of the functional units of the training device 10 may be realized by hardware such as electronic circuits.

[0025] The first acquisition unit 11 acquires input time series data A from each training data D stored in the memory unit 140, etc. The second acquisition unit 12 acquires output finger information B from each training data D. The construction unit 13 performs machine learning on each training data D, using the input time series data A acquired by the first acquisition unit 11 as the input element and the output finger information B acquired by the second acquisition unit 12 as the output element. By repeating machine learning on multiple training data D, the construction unit 13 constructs a trained model M that shows the input / output relationship between the input time series data A and the output finger information B.

[0026] In this example, the construction unit 13 constructs a trained model M by training a Transformer, but the embodiment is not limited to this. The construction unit 13 may also construct a trained model M by training a machine learning model of another type that handles time series. The trained model M constructed by the construction unit 13 is stored, for example, in the storage unit 140. The trained model M constructed by the construction unit 13 may also be stored on a server on a network or the like.

[0027] The finger placement suggestion device 20 includes a reception unit 21, an estimation unit 22, and a generation unit 23 as its functional units. The CPU 130 in Figure 1 executes the finger placement suggestion program, thereby realizing the functional units of the finger placement suggestion device 20. At least a portion of the functional units of the finger placement suggestion device 20 may be realized by hardware such as electronic circuits.

[0028] In this embodiment, the reception unit 21 receives time-series data including a sequence of musical notes consisting of multiple notes. The performer can provide the reception unit 21 with image data showing an image of a musical score as time-series data. Alternatively, the performer can generate time-series data by operating the operation unit 150 and provide it to the reception unit 21.

[0029] In this example, the time-series data has a similar structure to input time-series data A in Figure 2, and further includes performer identifiers indicating the classification (category) of the performers playing the note sequences. The performer identifiers are determined to differ for at least one of the performer's physical characteristics and the style of performance by the performer. The performer's physical characteristics include, for example, the size of the performer's hands, age, gender, or whether they are an adult or a child.

[0030] The estimation unit 22 estimates finger information using a trained model M stored in the memory unit 140, etc. The finger information indicates the fingers of the performer used when playing each note in the note sequence received by the reception unit 21, and is estimated based on the note sequence and the performer identifier. The finger information may also be a unique number assigned to each finger. The generation unit 23 generates musical score information based on the note sequence of time-series data received by the reception unit 21 and the finger information estimated by the estimation unit 22.

[0031] The display unit 160 displays an auxiliary score based on the musical score information generated by the generation unit 23. Figure 4 shows an example of an auxiliary score displayed on the display unit 160. As shown in Figure 4, the auxiliary score shows finger information estimated by the estimation unit 22 corresponding to each note in the note sequence received by the reception unit 21. In the example in Figure 4, the finger information shows the finger numbers of one hand.

[0032] To distinguish between the finger numbers of the left and right hands, a predetermined letter such as "L" may be placed near the finger numbers of the left hand, and another predetermined letter such as "R" may be placed near the finger numbers of the right hand. Alternatively, a predetermined color such as red may be placed on the finger numbers of the left hand or the corresponding musical notes, and another predetermined color such as blue may be placed on the finger numbers of the right hand or the corresponding musical notes.

[0033] (4) Training process and finger placement suggestion process Figure 5 is a flowchart illustrating an example of the training process performed by the training device 10 in Figure 3. The training process in Figure 5 is performed by the CPU 130 in Figure 1 executing the training program. First, the first acquisition unit 11 acquires input time-series data A from each training data D (step S1). The second acquisition unit 12 acquires output signal information B from each training data D (step S2). Steps S1 and S2 may be executed either first or simultaneously.

[0034] Next, the construction unit 13 performs machine learning on each training data D, using the input time series data A obtained in step S1 as the input element and the output signal information B obtained in step S2 as the output element (step S3). Subsequently, the construction unit 13 determines whether sufficient machine learning has been performed (step S4). If the machine learning is insufficient, the construction unit 13 returns to step S3. Steps S3 and S4 are repeated with changing parameters until sufficient machine learning is performed. The number of machine learning iterations varies according to the quality conditions that the constructed trained model M must satisfy.

[0035] If sufficient machine learning has been performed, the construction unit 13 saves the input-output relationship between the input time series data A and the output signal information B, which was acquired by machine learning in step S3, as a trained model M (step S5). This completes the training process.

[0036] Figure 6 is a flowchart showing an example of the finger placement suggestion process by the finger placement suggestion device 20 in Figure 3. The finger placement suggestion process in Figure 6 is performed by the CPU 130 in Figure 1 executing the finger placement suggestion program. First, the reception unit 21 receives time-series data (step S11). Next, the estimation unit 22 uses the trained model M saved in step S5 of the training process to estimate finger information from the time-series data received in step S11 (step S12).

[0037] Subsequently, the generation unit 23 generates musical score information based on the note sequence of time-series data received in step S11 and the finger information estimated in step S12 (step S13). Based on the generated musical score information, an auxiliary musical score may be displayed on the display unit 160. This completes the fingering suggestion process.

[0038] (5) Effects of the embodiment As described above, the fingering suggestion device 20 according to this embodiment includes a receiving unit 21 that receives time-series data including a sequence of multiple notes, and an estimation unit 22 that uses a trained model M to estimate fingering information indicating the fingers to be used when playing each note in the sequence on an instrument. With this configuration, appropriate fingering information is estimated from the temporal flow of multiple notes in the time-series data using the trained model M. This makes it possible to suggest appropriate fingering when playing an instrument.

[0039] The trained model M may be a machine learning model that has learned the input-output relationship between input time-series data A, which includes a reference note sequence consisting of multiple notes, and output finger information B, which indicates which finger is used to play each note in the reference note sequence on an instrument. In this case, finger information can be easily estimated from the time-series data.

[0040] The time-series data further includes a performer identifier indicating the performer playing the note sequence, and the estimation unit 22 may estimate finger information based on the performer identifier. In this case, appropriate finger information can be estimated according to the performer.

[0041] The performer identifier may be determined to correspond to the performer's physical characteristics. In this case, appropriate finger information can be estimated according to the performer's physical characteristics.

[0042] The performer identifier may be determined to correspond to the performer's playing style. In this case, appropriate fingering information can be estimated according to the performer's playing style.

[0043] The fingering display device 20 may further include a generation unit 23 that generates musical score information, which shows an auxiliary musical score with finger information attached to each note in the musical score sequence. In this case, the performer can easily recognize the finger corresponding to each note in the musical score sequence by looking at the auxiliary musical score.

[0044] The training device 10 according to this embodiment includes a first acquisition unit 11 that acquires input time-series data A including a reference note sequence consisting of multiple notes, a second acquisition unit 12 that acquires output finger information B indicating the fingers used when playing each note in the reference note sequence on an instrument, and a construction unit 13 that constructs a trained model M that has learned the input-output relationship between the input time-series data A and the output finger information B. With this configuration, a trained model M that has learned the input-output relationship between the input time-series data A and the output finger information B can be easily constructed.

[0045] (6) Other examples of training data In this embodiment, the input time-series data A includes a reference performer identifier, and the time-series data includes a performer identifier; however, the embodiment is not limited to this. The input time-series data A may include a reference note sequence but may not include a reference performer identifier. Similarly, the time-series data may include a note sequence but may not include a performer identifier.

[0046] Furthermore, in this embodiment, the input time-series data A and output finger information B are described in a so-called action-based manner, indicating key presses or releases in the MIDI (Musical Instrument Digital Interface) standard, but the embodiment is not limited to this. The input time-series data A and output finger information B may be described in other ways. For example, the input time-series data A and output finger information B may be described in a so-called note-based manner, indicating the starting position or length of a note in the MIDI standard. The same applies to the time-series data and finger information.

[0047] Figure 7 shows another example of input time series data A. The top of Figure 7 shows input time series data A(Ax) described in motion-based terms. The middle of Figure 7 shows input time series data A(Ay) described in note-based terms. Input time series data Ax and input time series data Ay contain the same reference note sequence (the reference note sequence in the musical score shown in the bottom of Figure 7). In input time series data Ax and Ay, "bar" and "beat" are elements that indicate the meter structure of the reference note sequence.

[0048] As shown in Figure 7, describing the input time-series data A in terms of musical notes shortens the length of the input time-series data A. This makes it easier to process longer input time-series data A. The output finger information B corresponding to the input time-series data A can be described by inserting an element indicating the finger number immediately after the element indicating the pitch number ("note_○○") in the input time-series data A.

[0049] Alternatively, the input time-series data A and output finger information B may be described in a manner that represents musical notation. Details of the input time-series data A and output finger information B described in a manner that represents musical notation will be explained in the following modified examples.

[0050] (7) Variant Figure 8 shows an example of input time-series data A in a modified example. The upper part of Figure 8 shows the input time-series data A (Az) described in a musical notation style. The lower part of Figure 8 shows the musical notation represented by the input time-series data Az. As shown in the upper part of Figure 8, the input time-series data Az contains multiple elements A0 to A24. Some elements have attributes. The attributes of an element are described after the element (after the underscore).

[0051] Element A0 indicates the percentage of notes in the reference note sequence that require fingering. Element A0 is placed at the beginning of the input time series data Az, but it may be placed at any position in the input time series data Az. The percentage is specified by the "fingerrate" attribute of element A0. In this example, the attribute "5" means a percentage of 100%. The percentage may have a range, such as 20-40% or 40-60%, or it may be divided into multiple ranges.

[0052] Element A1 indicates a part. Element A1 is placed immediately after element A0, but may be placed at any position in the input time series data Az. Element A1, "R" and "L", indicate the right-hand and left-hand parts, respectively. In this example, the element corresponding to the right hand is placed after "R". "L" is then placed after "L", and the element corresponding to the left hand is placed after "L". The elements corresponding to "R" and the right hand may be placed after the elements corresponding to the left hand. If there is no distinction between parts, the input time series data Az does not include element A1.

[0053] Elements A2, A15, and A24 represent bar lines in a musical score. Therefore, in the example in Figure 8, the range separated by the "bar" in element A2 and the "bar" in element A15 corresponds to the first measure. The range separated by the "bar" in element A15 and the "bar" in element A24 corresponds to the second measure.

[0054] Element A3 indicates the clef of the musical score. The type of clef is specified by the "clef" attribute of element A3. In the example in Figure 8, the attribute is "treble," so element A3 specifies a treble clef as the clef. If the attribute were "bass," element A3 would specify a bass clef as the clef.

[0055] Element A4 represents the time signature of the musical score. The type of time signature is specified by the "time" attribute of element A4. In the example in Figure 8, since the attribute is "4 / 4", element A4 specifies "4 / 4" as the time signature.

[0056] In the reference note sequence, notes are represented by a pair of pitch and duration. The pitch is specified by the "note" attribute in elements A5, A9, A11, A13, A16, A18, and A20. The duration is specified by the "len" attribute in elements A6, A10, A12, A14, A17, A19, and A21. In this example, "len_1" corresponds to one beat.

[0057] The direction of the note stem in a musical score is specified by other attributes of “len” in elements A6, A10, A12, A14, A17, A19, and A21. If the other attribute is “down”, the stem extends downward from the note head. If the other attribute is “up”, the stem extends upward from the note head. When multiple notes, such as eighth notes or sixteenth notes, are connected by a beam, the start, intermediate, and end positions of the beam are specified by further attributes of “len” in elements A10, A12, and A14, namely “start”, “continue”, and “stop”.

[0058] Rests in the reference note sequence are specified by the "rest" attribute in elements A7 and A22. The note value of a rest is described by the "len" attribute in elements A8 and A23.

[0059] In the example in Figure 8, elements A5 and A6 represent note N1, and elements A7 and A8 represent rest R1. Elements A9 and A10 represent note N2, elements A11 and A12 represent note N3, and elements A13 and A14 represent note N4. Elements A16 and A17 represent note N5, and elements A18 and A19 represent note N6. Elements A20 and A21 represent note N7, and elements A22 and A23 represent rest R2.

[0060] Figure 9 shows an example of output finger information B in a modified example. The upper part of Figure 9 shows the output finger information B (Bz) described in a musical notation format. Output finger information Bz corresponds to the input time-series data Az in Figure 8. The lower part of Figure 9 shows the musical notation represented by the output finger information Bz.

[0061] As shown in the upper part of Figure 9, the output signal information Bz includes multiple elements B0 to B24. Furthermore, the output signal information Bz also includes elements B5f, B9f, B11f, B13f, B16f, B18f, and B20f, which are placed immediately after elements B5, B9, B11, B13, B16, B18, and B20, respectively. Elements B0 to B24 are the same as elements A0 to A24 of the input time series data Az in Figure 8. Therefore, the first acquisition unit 11 in Figure 3 can acquire the input time series data Az by deleting elements B5f, B9f, B11f, B13f, B16f, B18f, and B20f from the output signal information Bz.

[0062] Elements B5f, B9f, B11f, B13f, B16f, B18f, and B20f indicate the finger numbers used when playing the notes corresponding to the preceding elements B5, B9, B11, B13, B16, B18, and B20 on an instrument. The finger numbers are specified by the "finger" attribute in elements B5f, B9f, B11f, B13f, B16f, B18f, and B20f. Therefore, as shown in the lower part of Figure 9, elements B5f, B9f, B11f, B13f, B16f, B18f, and B20f indicate the finger numbers "1", "1", "2", "1", "3", "3", and "2" used when playing notes N1 to N7, respectively, in the musical score.

[0063] (8) Effects of the modified form In a modified version of the first embodiment, the attribute of element A0 allows for the arbitrary specification of the proportion of notes in the reference note sequence to which fingerings should be assigned. When the proportion is 100%, the estimation unit 22 estimates fingering information for all notes in the note sequence. In this case, appropriate fingerings for an entry-level player can be presented.

[0064] Furthermore, when the percentage is 100%, the generation unit 23 may generate a video file that shows the finger movements using animation or the like, based on the finger information estimated by the estimation unit 22. This makes it possible to visualize the finger movements. The generation of such a video file may be performed before or after step S13 in the finger movement presentation process in Figure 6, in parallel with step S13, or in place of step S13.

[0065] On the other hand, when the percentage is less than 100%, the estimation unit 22 estimates some of the notes in the note sequence that are to be assigned fingerings, and fingering information for those notes. In this case, it is possible to suggest appropriate fingerings for beginner or intermediate level players who are playing the instrument, which is higher than the introductory level. In this configuration, the output fingering information Bz does not include some of the elements B5f, B9f, B11f, B13f, B16f, B18f, and B20f.

[0066] Furthermore, when the percentage is less than 100%, the estimation unit 22 may estimate note information indicating the note to which fingering should be assigned from the note sequence, without estimating finger information. Details will be explained in the third embodiment described later.

[0067] [2] Second embodiment (1) Processing system The processing system 100 in the second embodiment will now be described in terms of its differences from the processing system 100 in the first embodiment. In the training device 10 shown in Figure 3, the first acquisition unit 11 and the second acquisition unit 12 acquire the input time-series data A and output instruction information B of the training data D, respectively.

[0068] Figure 10 shows an example of input time-series data A in the second embodiment. The upper part of Figure 10 shows the input time-series data Az described in a musical notation style. The lower part of Figure 10 shows the musical notation represented by the input time-series data Az.

[0069] As shown in the upper part of Figure 10, the input time series data Az includes multiple elements A0 to A24. Elements A0 to A24 in Figure 10 are the same as elements A0 to A24 in the modified example in the first embodiment (Figure 8). The input time series data Az also includes additional elements placed immediately after some of the elements A5, A9, A11, A13, A16, A18, and A20 that correspond to musical notes. In the example in Figure 10, the input time series data Az further includes elements A5f, A11f, A16f, and A20f, which are placed immediately after elements A5, A11, A16, and A20, respectively.

[0070] Elements A5f, A11f, A16f, and A20f are finger information (hereinafter referred to as basic finger information) that indicate the finger numbers used when playing the notes corresponding to the preceding elements A5, A11, A16, and A20 on an instrument. The finger numbers are specified by the "finger" attribute in elements A5f, A11f, A16f, and A20f. Therefore, as shown in the lower part of Figure 10, elements A5f, A11f, A16f, and A20f indicate the finger numbers "1", "2", "3", and "2" used when playing notes N1, N3, N5, and N7, respectively, in the musical score.

[0071] The output finger information Bz in this embodiment is the same as the output finger information Bz in the modified example of the first embodiment (Figure 9). Therefore, the first acquisition unit 11 can acquire the input time series data Az by randomly deleting some of the elements B5f, B9f, B11f, B13f, B16f, B18f, and B20f from the output finger information Bz. The proportion of the elements B5f, B9f, B11f, B13f, B16f, B18f, and B20f to be deleted can be specified by the user of the training device 10 by operating the operation unit 150 in Figure 1.

[0072] In this example, the input time series data Az is obtained by deleting elements B9f, B13f, and B18f from the output time series data Bz. The elements B5f, B11f, B16f, and B20f that are not deleted remain as the basic time series data elements A5f, A11f, A16f, and A20f.

[0073] The construction unit 13 in Figure 3 performs machine learning using the input time series data Az as the input element and the output signal information Bz as the output element. By repeating the machine learning process with multiple training data D, a trained model M is constructed that shows the input-output relationship between the input time series data Az and the output signal information Bz.

[0074] In the fingering suggestion device 20, the reception unit 21 receives time-series data. The time-series data further includes basic finger information indicating which fingers are used when playing some of the notes in the note sequence on an instrument. The estimation unit 22 estimates the finger information indicating which fingers are used when playing the notes in the note sequence on an instrument, based on the constructed trained model M and the basic finger information. The generation unit 23 generates musical score information based on the note sequence and finger information of the time-series data.

[0075] (2) Effects of the embodiment According to this embodiment, even if only some of the finger information (basic finger information) for some of the notes in the time-series data of notes is known, and finger information for the remaining notes is not provided, the finger information for the remaining notes is supplemented. This makes it possible to present appropriate fingering for beginner-level players when playing an instrument. The generation unit 23 may generate a video file that shows the finger movements using animation or the like, based on the finger information estimated by the estimation unit 22. In this case, the finger movements can be visualized.

[0076] (3) Variant In this embodiment, the estimation unit 22 estimates finger information for all notes included in the time-series data of notes, but the embodiment is not limited to this. If finger information is provided for a first proportion of the notes included in the note sequence, the estimation unit 22 may estimate finger information for a second proportion of notes, which is greater than the first proportion, included in the note sequence. In this case, it is possible to suggest appropriate fingering for beginner or intermediate level players when playing an instrument.

[0077] In the modified example, the output signal information B of the training data D does not need to include parts of elements B5f, B9f, B11f, B13f, B16f, B18f, and B20f. For example, if the input time series data Az includes elements A5f, A11f, A16f, and A20f, then the output signal information B will include elements B5f, B11f, B16f, and B20f. On the other hand, the output signal information B does not need to include parts of elements B9f, B13f, and B18f.

[0078] [3] Third embodiment (1) Processing system The processing system 100 in the third embodiment differs from the processing system 100 in the first embodiment. In this embodiment, the training data D represents a pair of input time-series data A and output note information. In the training device 10 of Figure 3, the first acquisition unit 11 and the second acquisition unit 12 acquire the input time-series data A and output note information of the training data D, respectively. The acquisition of output note information is performed instead of step S2 in the sound learning process of Figure 5.

[0079] The input time series data Az in this embodiment is the same as the input time series data Az in the modified example of the first embodiment (Figure 8). The first acquisition unit 11 can acquire the input time series data Az by deleting elements C9f, C11f, and C16f from the output note information Cz in Figure 11, which will be described later.

[0080] Figure 11 shows an example of output note information C in the third embodiment. The upper part of Figure 11 shows the output note information C(Cz) described in a musical notation format. The lower part of Figure 11 shows the musical notation represented by the output note information Cz.

[0081] As shown in the upper part of Figure 11, the output note information Cz includes multiple elements C0 to C24. The elements C0 to C24 in Figure 11 are the same as the elements B0 to B24 of the output finger information Bz in the modified example in the first embodiment (Figure 9). In addition, the output note information Cz includes additional elements placed immediately after some of the elements C5, C9, C11, C13, C16, C18, and C20 that correspond to the notes.

[0082] In this example, the "fingerrate" attribute of element C0 is "2," and an attribute of "2" represents a proportion of 40%. Therefore, the output note information Cz further includes elements C9f, C11f, and C16f, which are placed immediately after elements C9, C11, and C16, respectively, which represent approximately 40% of the elements C5, C9, C11, C13, C16, C18, and C20.

[0083] Elements C9f, C11f, and C16f each indicate the notes corresponding to the preceding elements C9, C11, and C16, respectively, as the notes to which fingering is to be assigned from the reference note sequence. As shown in the lower part of Figure 11, elements C9f, C11f, and C16f allow notes N2, N3, and N5 corresponding to elements C9, C11, and C16 to be clearly written in the musical score.

[0084] The construction unit 13 in Figure 3 performs machine learning using the input time series data Az as the input element and the output note information Cz as the output element. By repeating the machine learning process with multiple training data D, a trained model M is constructed that shows the input-output relationship between the input time series data Az and the output note information Cz.

[0085] In the fingering suggestion device 20, the reception unit 21 receives time-series data. The estimation unit 22 estimates note information indicating the notes to which fingerings are to be assigned from the note sequence, based on the trained model M constructed by the training device 10 and the time-series data received by the reception unit 21. The estimation of note information is performed in place of step S12 in the fingering suggestion process shown in Figure 6. The generation unit 23 generates musical score information indicating an auxiliary musical score in which the notes indicated by the note information are displayed in an identifiable manner.

[0086] (2) Effects of the embodiment According to this embodiment, it is possible to present notes from a sequence of notes to which fingerings should be assigned. This allows beginner or intermediate level players to recognize the key notes when playing the instrument.

[0087] (3) Variant The estimation unit 22 may use the first trained model M constructed in the first embodiment and the second trained model M constructed in this embodiment to estimate finger information indicating which fingers to use when playing some of the notes in the note sequence on an instrument. Figure 12 is a flowchart showing an example of finger placement suggestion processing in a modified example.

[0088] First, the reception unit 21 receives time-series data (step S21). Next, the estimation unit 22 uses the first trained model M constructed in the first embodiment to estimate middle finger information from the time-series data received in step S11 (step S22). The middle finger information indicates which finger is used when playing each note in the note sequence on the instrument.

[0089] Furthermore, the estimation unit 22 uses the second trained model M constructed in this embodiment to estimate note information from the time-series data received in step S21 (step S23). Steps S22 and S23 may be executed in any order, or they may be executed simultaneously.

[0090] Next, the estimation unit 22 estimates finger information for the notes in the note sequence that are not indicated by the note information estimated in step S23, based on the intermediate finger information estimated in step S22 (step S24). Subsequently, the generation unit 23 generates musical score information based on the note sequence of time-series data received in step S21 and the finger information estimated in step S24 (step S25). This completes the fingering suggestion process.

[0091] In this fingering suggestion process, the intermediate finger information estimated in step S22 has a similar structure to, for example, the output finger information Bz in the modified example in the first embodiment (Figure 9). Also, the note information estimated in step S23 has a similar structure to the output note information Cz in Figure 11. Figure 13 shows an example of finger information estimated in step S24 of the fingering suggestion process.

[0092] The upper part of Figure 13 shows finger information F(Fz) described using a musical notation method. The lower part of Figure 13 shows an auxiliary musical score represented by finger information Fz. Finger information Fz is estimated by deleting elements B9f, B11f, and B16f, which correspond to elements C9f, C11f, and C16f that indicate the notes to which fingering is to be assigned in the note information (see Figure 11), from the intermediate finger information (see Figure 9).

[0093] Specifically, as shown in the upper part of Figure 13, the finger information Fz includes multiple elements F1 to F24. Elements F1 to F24 in Figure 13 are the same as elements B1 to B24 of the output finger information Bz in the modified example in the first embodiment (Figure 9). Furthermore, the finger information Fz includes additional elements placed immediately after some of the elements F5, F9, F11, F13, F16, F18, and F20 that correspond to the notes. In this example, the finger information Fz further includes elements F5f, F13f, F18f, and F20f, which are placed immediately after elements F5, F13, F18, and F20, respectively.

[0094] Elements F5f, F13f, F18f, and F20f indicate the finger numbers used when playing the notes corresponding to the preceding elements F5, F13, F18, and F20 on an instrument. The finger numbers are specified by the "finger" attribute in elements F5f, F13f, F18f, and F20f. Therefore, as shown in the lower part of Figure 13, elements F5f, F13f, F18f, and F20f indicate the finger numbers "1", "1", "3", and "2" used when playing notes N1, N4, N6, and N7, respectively, in the auxiliary score.

[0095] In one variation, fingering information for some notes is filtered from the total fingering information for all notes in a time-series data sequence. In this case, it is possible to present appropriate fingering for beginner or intermediate level players when they play an instrument. For example, since fingering information for key notes is filtered out, beginner or intermediate level players can develop the ability to judge appropriate fingering when practicing an instrument.

[0096] [4] Other embodiments In the above embodiment, the fingering suggestion device 20 includes a generation unit 23, but the embodiment is not limited thereto. The performer can create an auxiliary score by transcribing the fingering information estimated by the estimation unit 22 onto a desired score. Therefore, the fingering suggestion device 20 does not necessarily have to include a generation unit 23.

[0097] In the above embodiment, training data D is trained to estimate finger information when playing a piano, but the embodiment is not limited to this. Training data D may also be trained to estimate finger information when playing other instruments such as drums.

[0098] In the above embodiment, the case where the user of the fingering suggestion device 20 is a performer was described as an example, but the user of the fingering suggestion device 20 may be, for example, a staff member of a music production company. Also, machine learning by the training device 10 may be performed in advance by the staff member of the music production company.

Claims

1. A receiving unit that receives time-series data including a sequence of notes consisting of multiple notes and a performer identifier indicating the performer who plays the sequence of notes, The system includes an estimation unit that uses a trained model to estimate note information indicating the note to which fingerings are to be assigned from the note sequence based on the performer identifier, The aforementioned trained model is a machine learning model that has learned the input-output relationship between input time-series data including a reference note sequence consisting of multiple notes and a reference performer identifier indicating a reference performer who plays the reference note sequence, and output note information indicating the notes to which the reference performer's fingering is to be assigned from the reference note sequence, and is a fingering suggestion device.

2. The fingering display device according to claim 1, wherein the performer identifier is determined to correspond to the physical characteristics of the performer.

3. The fingering suggestion device according to claim 1 or 2, wherein the performer identifier is determined to correspond to the style of performance by the performer.

4. The fingering suggestion device according to any one of claims 1 to 3, further comprising a generation unit that generates musical score information showing an auxiliary musical score in which the musical notes indicated by the aforementioned musical note information are displayed in an identifiable manner.

5. A first acquisition unit acquires input time-series data including a reference note sequence consisting of multiple notes and a reference performer identifier indicating the reference performer who plays the reference note sequence. A second acquisition unit that acquires output note information indicating the note to which the fingering of the reference performer is to be assigned from the aforementioned reference note sequence, A training device comprising a construction unit that constructs a trained model that has acquired the input-output relationship between the input time-series data and the output musical note information.

6. It accepts time-series data including a sequence of notes consisting of multiple notes and a performer identifier indicating the performer who plays the sequence of notes. Using a trained model, note information indicating the note to which fingering should be assigned is estimated from the note sequence based on the performer identifier. The trained model is a machine learning model that has learned the input-output relationship between input time-series data including a reference note sequence consisting of multiple notes and a reference performer identifier indicating a reference performer who plays the reference note sequence, and output note information indicating the notes to which the fingering of the reference performer is to be assigned from the reference note sequence. A method of displaying fingerings, executed by a computer.

7. The system obtains input time-series data including a reference note sequence consisting of multiple notes and a reference performer identifier indicating the reference performer who plays the reference note sequence. From the aforementioned reference note sequence, output note information is obtained that indicates the note to which the fingering of the reference performer will be assigned. A trained model is constructed that has learned the input-output relationship between the input time-series data and the output note information. A training method performed by a computer.

Citation Information

Patent Citations

  • Fingering display device and program

    JP2007178695A

  • Fingering determination method and system for musical instrument performance

    JP2007241034A

  • Method of automated musical instrument finger finding

    WO2008147368A1