Information processing method, information processing system, and program
The information processing system addresses the challenge of intuitively confirming fingering by displaying a virtual performer and instrument, leveraging trained models to provide accurate fingering information for stringed instruments.
Patent Information
- Application Number
- JP2024118729
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2025-08-20
- Estimated Expiration
- 2042-03-25
AI Technical Summary
Existing techniques do not effectively enable users to visually and intuitively confirm fingering for stringed instruments.
An information processing system that acquires fingering information and displays a reference image of a virtual performer and instrument corresponding to the fingering, utilizing a generative model trained on performer data to generate fingering information based on sound and image analysis.
Enables users to visually and intuitively confirm correct fingering through a virtual representation, enhancing their playing experience.
Smart Images

Figure 0007726343000001 
Figure 0007726343000002 
Figure 0007726343000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to techniques for analyzing the performance of stringed instruments. [Background technology]
[0002] Various techniques for supporting the playing of stringed instruments have been proposed. For example, Patent Document 1 discloses a technique for displaying, on a display device, a fingering image that shows the fingering for playing chords on a stringed instrument. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2005-241877 Summary of the Invention [Problem to be solved by the invention]
[0004] One aspect of the present disclosure aims to enable a user to visually and intuitively confirm fingering for a stringed instrument. [Means for solving the problem]
[0005] In order to solve the above problems, an information processing method according to one aspect of the present disclosure acquires fingering information representing the fingering of a stringed instrument, and displays, on a display device, a reference image representing a virtual performer corresponding to the fingering represented by the fingering information and a virtual stringed instrument played by the performer.
[0006] An information processing system according to one aspect of the present disclosure includes an information processing unit that acquires fingering information representing the fingering of a stringed instrument, and a presentation processing unit that displays, on a display device, a reference image representing a virtual performer corresponding to the fingering represented by the fingering information and a virtual stringed instrument played by the performer.
[0007] A program according to one aspect of the present disclosure causes a computer system to function as an information processing unit that acquires fingering information representing the fingering of a stringed instrument, and a presentation processing unit that displays, on a display device, a reference image representing a virtual performer corresponding to the fingering represented by the fingering information and a virtual stringed instrument played by the performer. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a block diagram illustrating a configuration of an information processing system. [Figure 2] FIG. 2 is a schematic diagram of a performance image. [Figure 3] FIG. 2 is a block diagram illustrating an example of a functional configuration of an information processing system. [Figure 4] 10 is a flowchart of an image analysis process. [Figure 5] FIG. 2 is a schematic diagram of a reference image. [Figure 6] 10 is a flowchart of a performance analysis process. [Figure 7] FIG. 1 is a block diagram illustrating a configuration of a machine learning system. [Figure 8] FIG. 1 is a block diagram illustrating an example of the functional configuration of a machine learning system. [Figure 9] 1 is a flowchart of a machine learning process. [Figure 10] FIG. 10 is a block diagram illustrating an example of the functional configuration of an information processing system according to a third embodiment. [Figure 11] FIG. 10 is a block diagram illustrating an example of the functional configuration of an information processing system according to a fourth embodiment. [Figure 12] FIG. 10 is a block diagram illustrating an example of the functional configuration of a machine learning system according to a fourth embodiment. [Figure 13] FIG. 10 is a schematic diagram of a reference image in a modified example. [Figure 14] FIG. 10 is a block diagram illustrating a functional configuration of an information processing system according to a modified example. [Figure 15] FIG. 10 is a block diagram illustrating a functional configuration of an information processing system according to a modified example. DETAILED DESCRIPTION OF THE INVENTION
[0009] A: First embodiment FIG. 1 is a block diagram illustrating the configuration of an information processing system 100 according to a first embodiment. The information processing system 100 is a computer system (performance analysis system) for analyzing a performance of a stringed instrument 200 by a user U. The stringed instrument 200 is, for example, a natural musical instrument such as an acoustic guitar that includes a fingerboard and multiple strings. The information processing system 100 according to the first embodiment analyzes fingering when the user U plays the stringed instrument 200. Fingering is a method in which the user U uses his or her fingers to play the stringed instrument 200. Specifically, the fingers with which the user U presses each string against the fingerboard (hereinafter referred to as "string pressing") and the position of the pressed string on the fingerboard (combination of the string and the fret) are analyzed as the fingering of the stringed instrument 200.
[0010] The information processing system 100 includes a control device 11, a storage device 12, an operation device 13, a display device 14, a sound collection device 15, and an imaging device 16. The information processing system 100 is realized by, for example, a portable information device such as a smartphone or a tablet terminal, or a portable or stationary information device such as a personal computer. The information processing system 100 may be realized as a single device, or may be realized by multiple devices configured separately from each other.
[0011] The control device 11 is one or more processors that control the operation of the information processing system 100. Specifically, the control device 11 is configured by one or more types of processors, such as a central processing unit (CPU), a graphics processing unit (GPU), a sound processing unit (SPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or an application specific integrated circuit (ASIC).
[0012] The storage device 12 is one or more memories that store programs executed by the control device 11 and various data used by the control device 11. For example, a known recording medium such as a semiconductor recording medium or a magnetic recording medium, or a combination of multiple types of recording media, is used as the storage device 12. Note that, for example, a portable recording medium that is detachable from the information processing system 100, or a recording medium that the control device 11 can access via a communication network (e.g., cloud storage) may also be used as the storage device 12.
[0013] The operation device 13 is an input device that accepts operations by the user U. For example, an operator operated by the user U or a touch panel that detects contact by the user U is used as the operation device 13. The display device 14 displays various images under the control of the control device 11. For example, various display panels such as a liquid crystal display panel or an organic EL panel are used as the display device 14. Note that the operation device 13 or the display device 14, which are separate from the information processing system 100, may be connected to the information processing system 100 by wire or wirelessly.
[0014] The sound collection device 15 is a microphone that generates an audio signal Qx by collecting musical sounds produced from the stringed instrument 200 when played by the user U. The audio signal Qx is a signal that represents the waveform of the musical sounds produced by the stringed instrument 200. Note that the sound collection device 15, which is separate from the information processing system 100, may be connected to the information processing system 100 by wire or wirelessly. For convenience, an A / D converter that converts the audio signal Qx from analog to digital is not shown in the illustration.
[0015] The imaging device 16 generates an image signal Qy by capturing an image of the user U playing the stringed instrument 200. The image signal Qy is a signal representing a video of the user U playing the stringed instrument 200. Specifically, the imaging device 16 includes an optical system such as a photographing lens, an imaging element that receives incident light from the optical system, and a processing circuit that generates an image signal Qy according to the amount of light received by the imaging element. Note that the imaging device 16, which is separate from the information processing system 100, may be connected to the information processing system 100 via a wired or wireless connection.
[0016] FIG. 2 is an explanatory diagram of an image captured by the imaging device 16. The image represented by the image signal Qy (hereinafter referred to as the "performance image") G includes a player image Ga and a musical instrument image Gb. The player image Ga is an image of the user U playing the stringed instrument 200. The musical instrument image Gb is an image of the stringed instrument 200 played by the user U. The player image Ga includes an image of the user U's left hand (hereinafter referred to as the "left hand image") Ga1 and an image of the user U's right hand (hereinafter referred to as the "right hand image") Ga2. In the following description, it is assumed that the user U presses the strings with their left hand and plucks them with their right hand. However, the user U may also pluck the strings with their left hand and press them with their right hand. The musical instrument image Gb includes an image of the stringed instrument's fingerboard (hereinafter referred to as the "fingerboard image") Gb1.
[0017] 3 is a block diagram illustrating an example of the functional configuration of the information processing system 100. The control device 11 executes a program stored in the storage device 12 to realize a plurality of functions (an information acquisition unit 21, an information generation unit 22, and a presentation processing unit 23) for analyzing a performance of the string instrument 200 by the user U.
[0018] The information acquisition unit 21 acquires input information C. The input information C is control data including sound information X and finger information Y. The sound information X is data relating to musical sounds played by the user U on the stringed instrument 200. The finger information Y is data relating to a performance image G of the user U playing the stringed instrument 200. The generation of the input information C by the information acquisition unit 21 is repeated sequentially in parallel with the performance of the stringed instrument 200 by the user U. The information acquisition unit 21 of the first embodiment includes an acoustic analysis unit 211 and an image analysis unit 212.
[0019] The acoustic analysis unit 211 generates sound information X by analyzing the sound signal Qx. The sound information X in the first embodiment specifies the pitch of the sound played by the user U on the stringed instrument 200. That is, the acoustic analysis unit 211 estimates the pitch of the sound represented by the sound signal Qx and generates sound information X that specifies the pitch. Note that any known analysis technique may be employed to estimate the pitch of the sound signal Qx.
[0020] The acoustic analysis unit 211 also sequentially detects sound generation points by analyzing the acoustic signal Qx. A sound generation point is the time point at which the stringed instrument 200 starts to generate sound (i.e., onset). Specifically, the acoustic analysis unit 211 sequentially determines the volume of the acoustic signal Qx at a predetermined cycle, and detects the time point at which the volume exceeds a predetermined threshold as the sound generation point. The stringed instrument 200 generates sound when the user U plucks the string. Therefore, the sound generation point of the stringed instrument 200 can also be said to be the time point at which the user U plucks the stringed instrument 200.
[0021] The acoustic analysis unit 211 generates sound information X upon detecting a sound generation point. That is, sound information X is generated for each sound generation point of the stringed instrument 200. For example, the acoustic analysis unit 211 generates sound information X by analyzing a sample of the sound signal Qx taken a predetermined time (e.g., 150 milliseconds) after each sound generation point. The sound information X corresponding to each sound generation point is information representing the pitch of the musical tone generated at that sound generation point.
[0022] The image analysis unit 212 generates finger information Y by analyzing the image signal Qy. The finger information Y in the first embodiment represents a left hand image Ga1 of the user U and a fingerboard image Gb1 of the stringed instrument 200. The image analysis unit 212 generates finger information Y in response to detection of a sound generation point by the acoustic analysis unit 211. That is, finger information Y is generated for each sound generation point of the stringed instrument 200. For example, the image analysis unit 212 generates finger information Y by analyzing the performance image G of the image signal Qy at a point when a predetermined time (for example, 150 milliseconds) has elapsed since each sound generation point. The finger information Y corresponding to each sound generation point represents the left hand image Ga1 and fingerboard image Gb1 at that sound generation point.
[0023] FIG. 4 is a flowchart of the process Sa3 (hereinafter referred to as "image analysis process") in which the image analysis unit 212 generates finger information Y. The image analysis process Sa3 is started in response to the detection of the sound producing point. When the image analysis process Sa3 is started, the image analysis unit 212 executes image detection process (Sa31). The image detection process is a process for extracting an image Ga1 of the left hand of the user U and an image Gb1 of the fingerboard of the stringed instrument 200 from the performance image G represented by the image signal Qy. The image detection process utilizes, for example, object detection process using a statistical model such as a deep neural network.
[0024] The image analysis unit 212 executes image conversion processing (Sa32). As illustrated in FIG. 2, the image conversion processing is image processing that converts the performance image G so that the fingerboard image Gb1 is converted into an image obtained by observing the fingerboard from a predetermined direction and distance. For example, the image analysis unit 212 converts the performance image G so that the fingerboard image Gb1 approximates a rectangular reference image Gref arranged in a predetermined direction. The left hand image Ga1 of the user U is also converted along with the fingerboard image Gb1. The image conversion processing utilizes well-known image processing such as projective transformation, which applies a transformation matrix generated from the fingerboard image Gb1 and the reference image Gref to the performance image G. The image analysis unit 212 generates fingering information Y that represents the performance image G after the image conversion processing.
[0025] As explained above, the sound information X and fingering information Y are generated for each sound generation point. That is, the information acquiring unit 21 generates the input information C for each sound generation point of the stringed instrument 200. A time series of multiple pieces of input information C corresponding to different sound generation points is generated.
[0026] The information generating unit 22 in FIG. 3 generates fingering information Z using input information C. The fingering information Z is data in any format that represents the fingering of the stringed instrument 200. Specifically, the fingering information Z specifies the finger numbers of one or more fingers used to press the strings of the stringed instrument 200 and the fingering position of the fingers. The fingering position is specified, for example, by combining one of the multiple strings of the stringed instrument 200 with one of the multiple frets provided on the fingerboard.
[0027] As described above, input information C is generated for each sound generation point. Therefore, the information generation unit 22 generates fingering information Z for each sound generation point. In other words, a time series of multiple pieces of fingering information Z corresponding to different sound generation points is generated. The fingering information Z corresponding to each sound generation point is information that represents the fingering at that sound generation point. As can be understood from the above explanation, in the first embodiment, acquisition of input information C and generation of fingering information Z are executed for each sound generation point of the stringed instrument 200. Therefore, it is possible to prevent fingering information from being generated unnecessarily when the user U is pressing but not plucking the strings. However, acquisition of input information C and generation of fingering information Z may be repeated at a predetermined cycle that is unrelated to the sound generation point.
[0028] The information generating unit 22 uses a generative model M to generate the fingering information Z. Specifically, the information generating unit 22 processes the input information C using the generative model M to generate the fingering information Z. The generative model M is a trained model that has learned the relationship between the input information C and the fingering information Z through machine learning. In other words, the generative model M outputs fingering information Z that is statistically appropriate for the input information C.
[0029] The generative model M is realized by a combination of a program that causes the control device 11 to execute a calculation to generate fingering information Z from input information C, and a plurality of variables (e.g., weights and biases) that are applied to the calculation. The program and the plurality of variables that realize the generative model M are stored in the storage device 12. The plurality of variables of the generative model M are set in advance by machine learning.
[0030] The generative model M is configured, for example, by a deep neural network. For example, any type of deep neural network, such as a recurrent neural network (RNN) or a convolutional neural network (CNN), can be used as the generative model M. The generative model M may also be configured by combining multiple types of deep neural networks. In addition, the generative model M may be equipped with additional elements such as a long short-term memory (LSTM) or attention.
[0031] The presentation processing unit 23 presents the fingering information Z to the user U. Specifically, the presentation processing unit 23 displays a reference image R1 illustrated in FIG. 5 on the display device 14. The reference image R1 includes a musical score B (B1, B2) corresponding to the performance of the stringed instrument 200 by the user U. The musical score B1 is a staff corresponding to the fingering represented by the fingering information Z. The musical score B2 is a tablature corresponding to the fingering represented by the fingering information Z. In other words, the musical score B2 is an image including multiple (six) horizontal lines corresponding to different strings of the stringed instrument 200. In the musical score B2, the fret numbers corresponding to the fingering positions are displayed in chronological order for each string. The presentation processing unit 23 generates musical score information P using the time series of the fingering information Z. The musical score information P is data in any format representing the musical score B of FIG. 5. The presentation processing unit 23 displays the musical score B represented by the musical score information P on the display device 14.
[0032] 6 is a flowchart of a process (hereinafter referred to as a "performance analysis process") Sa executed by the control device 11. For example, an instruction from the user U via the operation device 13 is a trigger for starting the performance analysis process Sa.
[0033] When the performance analysis process Sa is started, the control device 11 (acoustic analysis unit 211) waits until a sound generation point is detected by analyzing the sound signal Qx (Sa1: NO). If a sound generation point is detected (Sa1: YES), the control device 11 (acoustic analysis unit 211) generates sound information X by analyzing the sound signal Qx (Sa2). Furthermore, the control device 11 (image analysis unit 212) generates fingering information Y by image analysis process Sa3 of FIG. 4. Note that the order of generating sound information X (Sa2) and fingering information Y (Sa3) may be reversed. As explained above, input information C is generated for each sound generation point of the stringed instrument 200. Note that the input information C may be generated at a predetermined cycle.
[0034] The control device 11 (information generation unit 22) processes the input information C using the generation model M to generate fingering information Z (Sa4). Furthermore, the control device 11 (presentation processing unit 23) presents the fingering information Z to the user U (Sa5, Sa6). Specifically, the control device 11 generates score information P representing score B from the fingering information Z (Sa5), and displays score B represented by the score information P on the display device 14 (Sa6).
[0035] The control device 11 determines whether a predetermined termination condition has been met (Sa7). The termination condition may be, for example, when the user U issues an instruction to terminate the performance analysis process Sa via the operation device 13, or when a predetermined time has passed since the most recent sound generation point of the stringed instrument 200. If the termination condition has not been met (Sa7: NO), the control device 11 proceeds to step Sa1. That is, the acquisition of input information C (Sa2, Sa3), the generation of fingering information Z (Sa4), and the presentation of fingering information Z (Sa5, Sa6) are repeated for each sound generation point of the stringed instrument 200. On the other hand, if the termination condition has been met (Sa7: YES), the performance analysis process Sa ends.
[0036] As can be understood from the above explanation, in the first embodiment, fingering information Z is generated by processing input information C including sound information X and fingering information Y using a generative model M. Therefore, fingering information Z can be generated that corresponds to musical sounds (audio signals Qx) produced by the stringed instrument 200 when played by the user U, and an image (image signal Qy) of the user U playing the stringed instrument 200. In other words, fingering information Z that corresponds to the performance of the stringed instrument 200 by the user U can be provided. In the first embodiment, in particular, music score information P is generated using the fingering information Z. Therefore, the user U can effectively use the fingering information Z by displaying the music score B.
[0037] 7 is a block diagram illustrating the configuration of a machine learning system 400 according to the first embodiment. The machine learning system 400 is a computer system that establishes, through machine learning, a generative model M used by the information processing system 100. The machine learning system 400 includes a control device 41 and a storage device 42.
[0038] The control device 41 is composed of one or more processors that control each element of the machine learning system 400. For example, the control device 41 is composed of one or more types of processors such as a CPU, a GPU, an SPU, a DSP, an FPGA, or an ASIC.
[0039] The storage device 42 is one or more memories that store programs executed by the control device 41 and various data used by the control device 41. The storage device 42 is configured with a known recording medium, such as a magnetic recording medium or a semiconductor recording medium. The storage device 42 may also be configured with a combination of multiple types of recording media. Note that the storage device 42 may be a portable recording medium that is detachable from the machine learning system 400, or a recording medium that the control device 41 can access via a communication network (e.g., cloud storage).
[0040] 8 is a block diagram illustrating an example of the functional configuration of a machine learning system 400. A storage device 42 stores multiple pieces of training data T. Each of the multiple pieces of training data T is teacher data that includes training input information Ct and training fingering information Zt.
[0041] The training input information Ct includes sound information Xt and fingering information Yt. The sound information Xt is data related to musical tones played by a number of performers (hereinafter referred to as "reference performers") on the stringed instrument 201. Specifically, the sound information Xt specifies the pitches played by the reference performers on the stringed instrument 201. The fingering information Yt is data related to images of the reference performer's left hand and the fingerboard of the stringed instrument 201. Specifically, the fingering information Yt represents an image of the reference performer's left hand and an image of the fingerboard of the stringed instrument 201.
[0042] The fingering information Zt of the training data T is data representing the fingering of the stringed instrument 201 by the reference performer. That is, the fingering information Zt of each training data T is the correct label that the generative model M should generate for the input information Ct of the training data T.
[0043] Specifically, the fingering information Zt specifies the finger numbers of the left hand used by the reference performer to press the strings of the stringed instrument 201, as well as the fingering positions. The fingering positions in the fingering information Zt are positions detected by a detection device 250 installed on the stringed instrument 201. The detection device 250 is, for example, an optical or mechanical sensor installed on the fingerboard of the stringed instrument 201. Note that the detection of the fingering positions in the fingering information Zt can employ any known technology, such as the technology described in U.S. Pat. No. 9,646,591. As can be understood from the above explanation, the fingering information Zt for training is generated using the results of the detection device 250 installed on the stringed instrument 201 detecting the performance of the reference performer. This reduces the burden of preparing training data T used in machine learning of the generative model M.
[0044] The control device 41 of the machine learning system 400 executes a program stored in the storage device 42 to realize multiple functions (training data acquisition unit 51, learning processing unit 52) for generating a generative model M. The training data acquisition unit 51 acquires multiple pieces of training data T. The learning processing unit 52 establishes the generative model M through machine learning using the multiple pieces of training data T.
[0045] 9 is a flowchart of a process Sb in which the control device 41 establishes a generative model M through machine learning (hereinafter referred to as the “machine learning process”). For example, the machine learning process Sb is started in response to an instruction from the operator of the machine learning system 400.
[0046] When the machine learning process Sb is started, the control device 41 (training data acquisition unit 51) selects one of the multiple training data T (hereinafter referred to as "selected training data T") (Sb1). The control device 41 (learning processing unit 52) iteratively updates multiple coefficients of an initial or provisional generative model M (hereinafter referred to as "provisional model M0") using the selected training data T (Sb2 to Sb4).
[0047] The control device 41 generates fingering information Z by processing input information Ct of the selected training data T using the provisional model M0 (Sb2). The control device 41 calculates a loss function that represents the error between the fingering information Z generated by the provisional model M0 and the fingering information Zt of the selected training data T (Sb3). The control device 41 updates multiple variables of the provisional model M0 so that the loss function is reduced (ideally minimized) (Sb4). For example, backpropagation is used to update each variable according to the loss function.
[0048] The control device 41 determines whether a predetermined termination condition is met (Sb5). The termination condition is that the loss function falls below a predetermined threshold, or that the amount of change in the loss function falls below a predetermined threshold. If the termination condition is not met (Sb5: NO), the control device 41 selects unselected training data T as new selected training data T (Sb1). That is, the process of updating multiple variables of the provisional model M0 (Sb1 to Sb4) is repeated until the termination condition is met (Sb5: YES). If the termination condition is met (Sb5: YES), the control device 41 ends the machine learning process Sb. The provisional model M0 at the time the termination condition is met is confirmed as the trained generative model M.
[0049] As can be understood from the above explanation, the generative model M learns the latent relationship between the input information Ct and the fingering information Zt in multiple training data T. Therefore, the trained generative model M outputs fingering information Z that is statistically valid for unknown input information C based on the above relationship.
[0050] The control device 41 transmits the generative model M established by the machine learning process Sb to the information processing system 100. Specifically, multiple variables that define the generative model M are transmitted to the information processing system 100. The control device 11 of the information processing system 100 receives the generative model M transmitted from the machine learning system 400 and stores the generative model M in the storage device 12.
[0051] B: Second embodiment A second embodiment will be described. Note that, for elements in the following exemplary aspects that have the same functions as those in the first embodiment, the same reference numerals as those in the first embodiment will be used, and detailed descriptions of each will be omitted as appropriate.
[0052] The configuration and operation of the information processing system 100 in the second embodiment are the same as those in the first embodiment. Therefore, the second embodiment also achieves the same effects as those in the first embodiment. In the second embodiment, the fingering information Zt of the training data T applied to the machine learning process Sb is different from that in the first embodiment.
[0053] In the first embodiment, training data T including input information Ct (sound information Xt and fingering information Yt) corresponding to performances by each of a plurality of reference performers and fingering information Zt corresponding to performances by each reference performer is used in the machine learning process Sb of the generative model M. In other words, the input information Ct and fingering information Zt in the training data T correspond to performances by a common reference performer.
[0054] In the second embodiment, the input information Ct of each piece of training data T is information (sound information Xt and fingering information Yt) corresponding to performances by a number of reference performers, as in the first embodiment. On the other hand, the fingering information Zt of each piece of training data T in the second embodiment represents the fingering used in performance by a specific performer (hereinafter referred to as the "target performer"). The target performer may be, for example, a musical artist who plays the stringed instrument 200 with distinctive fingering, or a musical instructor who plays the stringed instrument 200 with exemplary fingering. In other words, the input information Ct and fingering information Zt in the training data T in the second embodiment correspond to performances by different performers (reference performers / target performers).
[0055] The fingering information Zt of the target player in the training data T is prepared by analyzing images of the target player playing a stringed instrument. For example, the fingering information Zt is generated from images of a live music performance or a music video in which the target player appears. Therefore, the fingering information Zt reflects the fingering that is unique to the target player. For example, the fingering information Zt may reflect a tendency to frequently press strings within a specific range on the fingerboard of a stringed instrument, or a tendency to frequently press strings with a specific finger of the left hand.
[0056] As can be understood from the above explanation, the generative model M of the second embodiment generates fingering information Z that corresponds to the performance by the user U (sound information Xt and fingering information Yt) and reflects the fingering tendencies of the target player. For example, the fingering information Z represents fingering that is likely to be adopted by the target player if the target player were to play a piece of music similar to that played by the user U. Therefore, by checking the score B displayed according to the fingering information Z, the user U can confirm what fingering the target player would use to play the piece of music played by the user U.
[0057] According to the second embodiment, a target performer, such as a music artist or a music instructor, can enjoy the customer experience of being able to easily provide his / her fingering information Z to a large number of users U. Furthermore, the users U can enjoy the customer experience of practicing a stringed instrument while referring to the fingering information Z of a desired target performer.
[0058] C: Third embodiment FIG. 10 is a block diagram illustrating the functional configuration of an information processing system 100 according to the third embodiment. In the third embodiment, multiple generative models M corresponding to different target players are selectively used. Each of the multiple generative models M corresponds to one generative model M in the second embodiment. One generative model M corresponding to each target player is a model that has learned the relationship between training input information Ct and training fingering information Zt representing the fingering performed by that target player.
[0059] Specifically, in the third embodiment, multiple sets of training data T are prepared for each target player. A generative model M for each target player is established by a machine learning process Sb that uses the multiple sets of training data T for that target player. Therefore, the generative model M for each target player generates fingering information Z that corresponds to the performance (sound information Xt and fingering information Yt) by the user U and reflects the fingering tendencies of the target player.
[0060] The user U can select one of a plurality of target players by operating the operation device 13. The information generation unit 22 accepts the user U's selection of the target player. The information generation unit 22 processes the input information C using a generative model M corresponding to the target player selected by the user U from among a plurality of generative models M, thereby generating fingering information Z (Sa4). Therefore, the fingering information Z generated by the generative model M represents fingering that is likely to be adopted by the target player selected by the user U, assuming that the target player plays a similar piece of music to that of the user U.
[0061] The third embodiment also achieves the same effects as the second embodiment. In particular, the third embodiment selectively uses one of a plurality of generative models M corresponding to different target players. Therefore, fingering information Z can be generated that reflects the fingering tendencies specific to each target player.
[0062] D: Fourth embodiment 11 is a block diagram illustrating the functional configuration of an information processing system 100 according to the fourth embodiment. The input information C according to the fourth embodiment includes identification information D in addition to the sound information X and fingering information Y similar to those in the first embodiment. The identification information D is a code string for identifying one of a plurality of target players.
[0063] As in the third embodiment, the user U can select one of a plurality of target performers by operating the operation device 13. The information acquisition unit 21 generates identification information D of the target performer selected by the user U. That is, the information acquisition unit 21 generates input information C including sound information X, finger information Y, and the identification information D.
[0064] FIG. 12 is a block diagram illustrating the functional configuration of a machine learning system 400 according to the fourth embodiment. In the fourth embodiment, as in the third embodiment, multiple sets of training data T are prepared for each target player. The training data T corresponding to each target player includes learning identification information Dt in addition to the sound information Xt and fingering information Yt similar to those in the first embodiment. The identification information Dt is a code string for identifying one of multiple target players. Furthermore, the fingering information Zt in the training data T corresponding to each target player represents the fingering of the stringed instrument 200 by that target player. In other words, the fingering information Zt of each target player reflects the performance tendencies of that target player on the stringed instrument 200.
[0065] In the third embodiment, a generative model M is individually generated for each target player through machine learning processing Sb using multiple pieces of training data T for each target player. In the fourth embodiment, a single generative model M is generated through machine learning processing Sb using multiple pieces of training data T corresponding to different target players. That is, the generative model M in the fourth embodiment is a model that learns the relationship between training input information Ct, including identification information D of each target player, and training fingering information Zt representing the fingering of the target player. Therefore, the generative model M generates fingering information Z that corresponds to the performance by a user U (sound information Xt and fingering information Yt) and reflects the fingering tendency of the target player selected by the user U.
[0066] As explained above, the fourth embodiment also achieves the same effects as the second embodiment. In particular, in the fourth embodiment, the input information C includes the target player's identification information D. Therefore, as in the third embodiment, fingering information Z can be generated that reflects the fingering tendencies unique to each target player.
[0067] E: Fifth embodiment The presentation processing unit 23 of the fifth embodiment uses fingering information Z to display the reference image R2 of Fig. 13 on the display device 14. Note that the configuration and operation other than that of the presentation processing unit 23 are the same as those of the first to fourth embodiments. Therefore, the fifth embodiment also achieves the same effects as those of the first to fourth embodiments.
[0068] The reference image R2 includes a virtual object (hereinafter referred to as "virtual object") O that exists in a virtual space. The virtual object O is a three-dimensional image that represents a virtual performer Oa playing a virtual stringed instrument Ob. The virtual performer Oa includes a left hand Oa1 that presses the stringed instrument Ob and a right hand Oa2 that plucks the stringed instrument Ob. The state of the virtual object O (particularly the state of the left hand Oa1) changes over time according to the fingering information Z that is sequentially generated by the information generation unit 22. As described above, the presentation processing unit 23 of the fifth embodiment displays the reference image R2 representing the virtual performer Oa (Oa1, Oa2) and the virtual stringed instrument Ob on the display device 14.
[0069] The fifth embodiment also achieves the same effects as the first to fourth embodiments. In particular, in the fifth embodiment, a virtual performer Oa corresponding to the fingering represented by the fingering information Z is displayed on the display device 14 together with a virtual stringed instrument Ob. Therefore, the user U can visually and intuitively confirm the fingering represented by the fingering information Z.
[0070] The display device 14 may be mounted on a head-mounted display (HMD) worn on the head of the user U. The presentation processor 23 displays a virtual object O (a performer Oa and a string instrument Ob) captured by a virtual camera in the virtual space as a reference image R2 on the display device 14. The presentation processor 23 dynamically controls the position and direction of the virtual camera in the virtual space according to the behavior (e.g., position and direction) of the user U's head. Therefore, the user U can view the virtual object O from any position and direction in the virtual space by appropriately moving his or her head. The HMD equipped with the display device 14 may be either a transparent type that allows the user U to view the real space as the background of the virtual object O, or a non-transparent type that displays the virtual object O together with a background image of the virtual space. A transparent HMD displays the virtual object O using, for example, augmented reality (AR) or mixed reality (MR), while a non-transparent HMD displays the virtual object O using, for example, virtual reality (VR).
[0071] The display device 14 may also be mounted on a terminal device capable of communicating with the information processing system 100 via a communication network such as the Internet. The presentation processing unit 23 transmits image data representing the reference image R2 to the terminal device, thereby displaying the reference image R2 on the display device 14 of the terminal device. The display device 14 of the terminal device may or may not be worn on the head of the user U.
[0072] F: Variation Specific modified embodiments that can be added to each of the above-described embodiments are exemplified below. Multiple embodiments arbitrarily selected from the above-described embodiments and the modified embodiments exemplified below may be combined as appropriate within the scope of not mutually contradicting each other.
[0073] (1) In the above-described embodiments, the musical score B corresponding to the fingering information Z is displayed on the display device 14, but the use of the fingering information Z is not limited to the above examples. For example, as illustrated in FIG. 14, the presentation processing unit 23 may generate content N according to the fingering information Z and the sound information X. The content N includes the above-described musical score B generated from the time series of the fingering information Z and the time series of pitches specified by the sound information X for each sound generation point. When the content is played back by the playback device, musical tones corresponding to the pitches of each sound information X are played back in parallel with the display of the musical score B. Therefore, viewers of the content can listen to the musical notes being played while viewing the musical score B of the musical piece. The above-described content is useful as teaching materials for practicing or teaching performance of the stringed instrument 200, for example.
[0074] (2) In the above-described embodiments, the sound information X specifies the pitch, but the information specified by the sound information X is not limited to the pitch. For example, the frequency characteristics of the sound signal Qx may be used as the sound information X. The frequency characteristics of the sound signal Qx may be, for example, information such as an intensity spectrum (amplitude spectrum or power spectrum) or MFCC (Mel-Frequency Cepstrum Coefficients). Furthermore, the time series of samples constituting the sound signal Qx may be used as the sound information X. As can be understood from the above examples, the sound information X is comprehensively expressed as information related to the sound played by the user U on the stringed instrument 200.
[0075] (3) In the above-described embodiments, the sound information X is generated by analyzing the audio signal Qx. However, the method of generating the sound information X is not limited to the above. For example, as illustrated in FIG. 15 , the acoustic analysis unit 211 may generate the sound information X from performance information E sequentially supplied from the electronic stringed instrument 202. The electronic stringed instrument 202 is a MIDI (Musical Instrument Digital Interface) instrument that outputs performance information E representing a performance by a user U. The performance information E is event data specifying the pitch and intensity of the note played by the user U, and is output from the electronic stringed instrument 202 each time the user U plucks a string. The acoustic analysis unit 211 generates, for example, the pitch included in the performance information E as the sound information X. The acoustic analysis unit 211 may also detect a sound generation point from the performance information E. For example, the time when the performance information E representing a sound generation is supplied from the electronic stringed instrument 202 is detected as the sound generation point.
[0076] (4) In the above-described embodiments, the sound source of the stringed instrument 200 is detected by analyzing the sound signal Qx, but the method for detecting the sound source is not limited to the above examples. For example, the image analysis unit 212 may detect the sound source of the stringed instrument 200 by analyzing the image signal Qy. As described above, the player image Ga represented by the image signal Qy includes a right hand image Ga2 of the right hand used by the user U to pluck the strings. The image analysis unit 212 extracts the right hand image Ga2 from the performance image G and detects the plucked strings by analyzing changes in the right hand image Ga2. The time when the user U plucks the strings is detected as the sound source.
[0077] (5) Techniques for playing a stringed instrument 200 such as a guitar include an arpeggio style in which multiple musical notes are played in sequence, and a stroke style in which multiple musical notes constituting a chord are played substantially simultaneously. When analyzing the performance (particularly the onset point) of the stringed instrument 200, a distinction may be made between arpeggio style and stroke style. For example, for multiple musical notes played sequentially at intervals greater than a predetermined threshold, an onset point is detected for each musical note (arpeggio style). On the other hand, for multiple musical notes played at intervals less than the predetermined threshold, a single onset point is detected for the multiple musical notes (stroke style). As described above, the playing style of the stringed instrument 200 may be reflected in the detection of the onset point. Furthermore, the onset point may be discretized on the time axis. In a mode in which the onset point is discretized, a single onset point is identified for multiple musical notes played at intervals less than the predetermined threshold.
[0078] (6) In the above-described embodiments, the fingering information Y includes a left hand image Ga1 and a fingerboard image Gb1. However, the fingering information Y may also include a right hand image Ga2 in addition to the left hand image Ga1 and fingerboard image Gb1. With the above configuration, the fingering information Z is generated based on the user U's right hand plucking in addition to the left hand pressing. Similarly, the fingering information Yt in the input information Ct of each training data T may also include an image of the right hand used by the reference performer for plucking.
[0079] (7) In the above-described embodiments, the finger information Y includes the player image Ga (left hand image Ga1 and right hand image Ga2) and the musical instrument image Gb (fingerboard image Gb1). However, the format of the finger information Y is arbitrary. The image analysis unit 212 may generate the finger information Y from the coordinates of feature points extracted from the performance image G. The finger information Y, for example, specifies the coordinates of each node (e.g., joint or tip) in the left hand image Ga1 of the user U, or the coordinates of the points where each string intersects with each fret in the fingerboard image Gb1 of the stringed instrument 200. In an embodiment in which the right hand image Ga2 is reflected in the finger information Y, the finger information Y specifies the coordinates of each node (e.g., joint or tip) in the right hand image Ga2 of the user U. As can be understood from the above examples, the finger information Y is comprehensively expressed as information relating to the player image Ga and the musical instrument image Gb.
[0080] (8) In the third embodiment, one of the multiple generative models M was selected in response to an instruction from the user U, but the method for selecting the generative model M is not limited to the above example. That is, any method for selecting one of the multiple target performers may be used. For example, the information generating unit 22 may select one of the multiple generative models M in response to an instruction from an external device or the result of a predetermined calculation process. Similarly, in the fourth embodiment, any method for selecting one of the multiple target performers may be used. For example, the information acquiring unit 21 may generate identification information D of one of the multiple target performers in response to an instruction from an external device or the result of a predetermined calculation process.
[0081] (9) In the above-described embodiments, a deep neural network is used as an example of the generative model M for generating the fingering information Z. However, the form of the generative model M is not limited to the above examples. For example, a statistical model such as an HMM (Hidden Markov Model) or an SVM (Support Vector Machine) may be used as the generative model M.
[0082] (10) In each of the above-described embodiments, a generation model M that has learned the relationship between input information C and fingering information Z is used. However, the configuration and method for generating fingering information Z from input information C are not limited to the above examples. For example, a reference table in which fingering information Z is associated with each of a plurality of different pieces of input information C may be used by the information generating unit 22 to generate fingering information Z. The reference table is a data table in which the correspondence between input information C and fingering information Z is registered, and is stored in, for example, the storage device 12. The information generating unit 22 searches the reference table for fingering information Z that corresponds to the input information C acquired by the information acquiring unit 21.
[0083] (11) In each of the above-mentioned embodiments, the machine learning system 400 established the generative model M, but the function of establishing the generative model M (the training data acquisition unit 51 and the learning processing unit 52) may be installed in the information processing system 100.
[0084] (12) In the above-described embodiments, fingering information Z specifying finger numbers and fingering positions has been exemplified, but the form of fingering information Z is not limited to these examples. For example, in addition to normal fingering defined by finger numbers and fingering positions, various performance techniques for musical expression may be specified by fingering information Z. Examples of performance techniques specified by fingering information Z include vibrato, slide, glissando, pulling, hammering, and choking. A known facial expression estimation model is used to estimate the performance technique.
[0085] (13) The string instrument 200 may be of any type. The string instrument 200 is collectively referred to as an instrument that produces sound by vibrating strings, and includes, for example, plucked string instruments and bowed string instruments. A plucked string instrument is a string instrument 200 that produces sound by plucking a string. Examples of plucked string instruments include an acoustic guitar, an electric guitar, an acoustic bass, an electric bass, a ukulele, a banjo, a mandolin, a koto, or a shamisen. A bowed string instrument is a string instrument that produces sound by bowing a string. Examples of bowed string instruments include a violin, a viola, a cello, or a double bass. The present disclosure can be applied to any of the above-listed types of string instruments for performance analysis.
[0086] (14) The information processing system 100 may be realized by a server device that communicates with a terminal device such as a smartphone or a tablet terminal. For example, the information acquisition unit 21 of the information processing system 100 receives an audio signal Qx (or performance information E) and an image signal Qy from the terminal device, and generates sound information X corresponding to the audio signal Qx and fingering information Y corresponding to the image signal Qy. The information generation unit 22 generates fingering information Z from input information C including the sound information X and the fingering information Y. The presentation processing unit 23 generates music score information P from the fingering information Z and transmits the music score information P to the terminal device. The display device of the terminal device displays a music score B represented by the music score information P.
[0087] In a configuration in which the acoustic analysis unit 211 and the image analysis unit 212 are installed in a terminal device, the information acquisition unit 21 receives the sound information X and the finger information Y from the terminal device. As can be understood from the above explanation, the information acquisition unit 21 is an element that generates the sound information X and the finger information Y, or an element that receives the sound information X and the finger information Y from another device such as a terminal device. In other words, the "acquisition" of the sound information X and the finger information Y includes both generation and reception.
[0088] Furthermore, in a configuration in which the presentation processing unit 23 is installed in a terminal device, the fingering information Z generated by the information generating unit 22 is transmitted from the information processing system 100 to the terminal device. The presentation processing unit 23 generates music score information P from the fingering information Z and displays it on a display device. As can be understood from the above explanation, the presentation processing unit 23 may be omitted from the information processing system 100.
[0089] (15) As described above, the functions of the information processing system 100 according to each of the above-described embodiments are realized through cooperation between one or more processors constituting the control device 11 and a program stored in the storage device 12. The programs exemplified above can be provided in a form stored on a computer-readable recording medium and installed on a computer. The recording medium is, for example, a non-transitory recording medium, such as an optical recording medium (optical disk) such as a CD-ROM, but also includes any known type of recording medium, such as a semiconductor recording medium or a magnetic recording medium. Note that a non-transitory recording medium includes any recording medium other than a transitory, propagating signal, and does not exclude volatile recording media. Furthermore, in a configuration in which a distribution device distributes a program via a communication network, the recording medium storing the program in the distribution device corresponds to the non-transitory recording medium described above.
[0090] G: Notes A particular pitch of a stringed instrument can be played using a plurality of different fingerings. When a user practices playing a stringed instrument, they may wish to check fingerings other than their own, such as exemplary fingerings or the fingerings of a particular performer. Furthermore, a user who plays a stringed instrument may wish to check their own fingerings when playing. In consideration of the above circumstances, one aspect of the present disclosure aims to provide fingering information regarding fingerings for users when playing a stringed instrument.
[0091] An information processing method according to one aspect (aspect 1) of the present disclosure acquires input information including finger information relating to the fingers of a user playing a stringed instrument and an image of the fingerboard of the stringed instrument, and sound information relating to the sound played by the user on the stringed instrument, and generates fingering information representing fingering by processing the acquired input information using a generative model that has learned the relationship between training input information and training fingering information. In the above aspect, fingering information is generated by processing the input information including finger information and sound information using a machine-learned generative model. In other words, it is possible to provide fingering information relating to the fingering when a user plays a stringed instrument.
[0092] "Finger information" is data in any format relating to an image of a user's fingers and an image of a stringed instrument fingerboard. For example, image information representing an image of a user's fingers and an image of a stringed instrument fingerboard, or analysis information generated by analyzing the image information, is used as finger information. The analysis information is, for example, information representing the coordinates of each node (joint or tip) of the user's fingers, information representing the line segments between the nodes, information representing the fingerboard, and information representing the frets on the fingerboard.
[0093] "Sound information" is data in any format related to a sound played by a user on a stringed instrument. For example, the sound information represents a feature of the sound played by the user. The feature may be, for example, a pitch or frequency characteristic, and may be determined by analyzing an audio signal representing the vibration of the strings of a stringed instrument. In addition, for example, in a stringed instrument that outputs performance information in MIDI format, sound information specifying the pitch of the performance information is generated. A time series of samples of the audio signal may also be used as sound information.
[0094] "Fingering information" is data in any format that represents the fingering of a stringed instrument. For example, finger numbers that indicate which fingers press the strings and the positions of the fingers (combinations of frets and strings) are used as fingering information.
[0095] A "generative model" is a trained model that has learned the relationship between input information and fingering information through machine learning. Multiple training data are used for machine learning of the generative model. Each training data set contains input information for learning and fingering information for learning (correct answer labels). Examples of generative models include various statistical models such as a deep neural network (DNN), a hidden Markov model (HMM), or a support vector machine (SVM).
[0096] In a specific example (Aspect 2) of Aspect 1, the sound generation points of the stringed instrument are further detected, and the input information is acquired and the fingering information is generated for each sound generation point. In the above aspect, the input information is acquired and the fingering information is generated for each sound generation point of the stringed instrument. This prevents fingering information from being generated unnecessarily when the user is pressing a string but not performing a sound generation operation. A "sound generation operation" is a user action that causes the stringed instrument to produce a sound corresponding to a pressing operation. Specifically, a sound generation operation is, for example, a plucking action for a plucked string instrument, or a bowing action for a bowed string instrument.
[0097] In a specific example (Aspect 3) of Aspect 1 or Aspect 2, the fingering information is further used to generate score information representing a score corresponding to the performance of the stringed instrument by the user. In the above aspects, the fingering information is used to generate the score information. The user can effectively use the fingering information by outputting (e.g., displaying or printing) the score. The "score" represented by the "score information" is, for example, a tablature notation indicating the fingering position for each string of the stringed instrument. However, it is also conceivable that the score information represents a staff notation in which the finger numbers to be used in playing each pitch are specified.
[0098] In a specific example (Aspect 4) of any of Aspects 1 to 3, a reference image representing a virtual performer corresponding to the fingering represented by the fingering information and a virtual stringed instrument played by the fingers is further displayed on a display device. In the above aspects, the virtual fingers corresponding to the fingering represented by the fingering information are displayed on a display device together with the virtual stringed instrument, allowing the user to visually and intuitively confirm the fingering represented by the fingering information.
[0099] In a specific example (Aspect 5) of Aspect 4, the display device is worn on the user's head, and in displaying the reference image, an image of the virtual performer and the virtual stringed instrument in the virtual space is captured by a virtual camera whose position and direction in the virtual space are controlled in accordance with the behavior of the user's head, and is displayed on the display device as the reference image. According to the above aspect, the user can view the virtual performer and the virtual stringed instrument from a desired position and direction.
[0100] In a specific example (aspect 6) of aspect 4 or aspect 5, the reference image is displayed by transmitting image data representing the reference image to a terminal device via a communication network, and displaying the reference image on the display device of the terminal device. According to the above aspect, even if the terminal device does not have a function for generating fingering information, the user of the terminal device can visually recognize a virtual performer and string instrument corresponding to the fingering information.
[0101] In a specific example (Aspect 7) of any of Aspects 1 to 6, content is further generated according to the sound information and the fingering information. According to the above aspects, content can be generated that allows the correspondence between sound information and fingering information to be confirmed. The above content is useful for practicing or teaching string instrument performance.
[0102] In a specific example (Aspect 8) of any of Aspects 1 to 7, the input information includes identification information of one of multiple players, and the generative model is a model that learns the relationship between the training input information including the identification information of each player and the training fingering information representing the fingerings performed by each player, for each of the multiple players. In the above aspects, the input information includes the identification information of the player. Therefore, it is possible to generate fingering information that reflects the fingering tendencies unique to each player.
[0103] In a specific example (Aspect 9) of any of Aspects 1 to 7, the fingering information is generated by processing the acquired input information using one of a plurality of generative models corresponding to different players, and each of the plurality of generative models is a model that has learned the relationship between the training input information and the training fingering information representing the fingering by the player corresponding to that generative model. In the above aspects, one of a plurality of unit models corresponding to different players is selectively used. Therefore, it is possible to generate fingering information that reflects the fingering tendencies unique to each player.
[0104] In a specific example (Aspect 10) of any of Aspects 1 to 9, the fingering information for learning is generated using the results of a detection device installed on the stringed instrument detecting the performance of the performer. In the above aspects, the fingering information for learning is generated using the detection results of the detection device installed on the stringed instrument. This reduces the burden of preparing training data to be used in machine learning of the generative model.
[0105] An information processing system according to one aspect (aspect 11) of the present disclosure includes an information acquisition unit that acquires input information including finger information relating to the fingers of a user playing a stringed instrument and an image of the fingerboard of the stringed instrument, and sound information relating to the sound played by the user on the stringed instrument, and an information generation unit that processes the acquired input information using a generation model that has learned the relationship between learning input information and learning fingering information, thereby generating fingering information representing fingering.
[0106] A program according to one aspect (aspect 12) of the present disclosure causes a computer system to function as an information acquisition unit that acquires input information including finger information relating to the fingers of a user playing a stringed instrument and an image of the fingerboard of the stringed instrument, and sound information relating to the sound played by the user on the stringed instrument, and an information generation unit that processes the acquired input information using a generation model that has learned the relationship between learning input information and learning fingering information, thereby generating fingering information representing fingering. [Explanation of symbols]
[0107] 100...information processing system, 200, 201...stringed instrument, 202...electronic stringed instrument, 250...detection device, 11, 41...control device, 12, 42...storage device, 13...operation device, 14...display device, 15...sound collection device, 16...imaging device, 21...information acquisition unit, 211...acoustic analysis unit, 212...image analysis unit, 22...information generation unit, 23...presentation processing unit, 400...machine learning system, 51...training data acquisition unit, 52...learning processing unit.
Claims
1. A method for acquiring input information including finger information relating to an image of the fingers of a user playing a stringed instrument and the fingerboard of the stringed instrument, generating fingering information representing fingering for the stringed instrument by processing the acquired input information using a generative model that has learned the relationship between the learning input information and the learning fingering information; a reference image representing a virtual player corresponding to the fingering represented by the fingering information and a virtual stringed instrument played by the player is displayed on a display device; An information processing method implemented by a computer system.
2. Furthermore, a score image representing a score corresponding to the fingering indicated by the fingering information is displayed on the display device. The information processing method according to claim 1.
3. The fingering information for learning represents the fingering used when a specific player plays the instrument. The information processing method according to claim 1.
4. the input information includes identification information of any one of a plurality of performers; The generative model is a model that learns the relationship between the learning input information, which includes identification information of each of the multiple players, and the learning fingering information, which represents fingering by each of the players. The information processing method according to claim 1.
5. generating the fingering information by processing the acquired input information using one of a plurality of generation models corresponding to different players; Each of the plurality of generative models is a model that has learned the relationship between the learning input information and the learning fingering information that represents the fingering of a player corresponding to the generative model. The information processing method according to claim 1.
6. The display device is worn on the user's head, In displaying the reference image, an image of the virtual performer and the virtual stringed instrument in the virtual space is captured by a virtual camera whose position and direction in the virtual space are controlled in accordance with the behavior of the user's head, and the image is displayed on the display device as the reference image. The information processing method according to claim 1.
7. An information acquisition unit that acquires input information including finger information regarding an image of the fingers of a user playing a stringed instrument and the fingerboard of the stringed instrument; an information generation unit that processes the acquired input information using a generation model that has learned the relationship between learning input information and learning fingering information, thereby generating fingering information that represents fingering for the stringed instrument; a presentation processing unit that displays, on a display device, a reference image representing a virtual player corresponding to the fingering represented by the fingering information and a virtual stringed instrument played by the player; An information processing system comprising:
8. An information acquisition unit that acquires input information including finger information regarding an image of the fingers of a user playing a stringed instrument and the fingerboard of the stringed instrument; an information generation unit that processes the acquired input information using a generation model that has learned the relationship between learning input information and learning fingering information, thereby generating fingering information that represents fingering for the stringed instrument; and a presentation processing unit that displays, on a display device, a reference image representing a virtual player corresponding to the fingering represented by the fingering information and a virtual stringed instrument played by the player; A program that makes a computer system function as a
9. an information acquisition unit that acquires input information including finger information relating to an image of the fingers of a user playing a stringed instrument and the fingerboard of the stringed instrument; an information generation unit that processes the acquired input information using a generation model that has learned the relationship between learning input information and learning fingering information, thereby generating fingering information that represents the fingering of the stringed instrument; A program that makes a computer system function as a
10. The input information further includes sound information regarding the sound that the user plays on the stringed instrument. The program of claim 9.
11. Furthermore, the computer system is caused to function as a presentation processing unit that displays, on a display device, a score image that represents a score corresponding to the fingering indicated by the fingering information. The program of claim 9.
Citation Information
Patent Citations
Musical performance practicing device and program for practicing musical performance
JP2003058157A
Fingering instruction apparatus and program
JP2005241877A
Musical chord chart generation device
JP2014153523A
Playing apparatus, motion capture method for fingering, and performance support system
JP2017181850A