Data processing method, data processing system and program

JP2025170154A5Pending Publication Date: 2026-01-08YAMAHA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025153584
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-03-25
Filing Date
2025-09-16
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

The accuracy of automatic performance following a user's performance is affected by the accuracy of identifying the performance position on a musical score, which can be reduced due to factors such as the sequence of notes in a piece of music.

Method used

A data output method that includes sequentially acquiring input data, providing it to multiple estimation models to generate estimated information, and identifying a musical score performance position, using models that indicate the relationship between performance data and musical score positions, thereby improving accuracy.

Benefits of technology

Enhances the accuracy of identifying performance positions on a musical score based on user input, allowing for precise reproduction of musical data and providing a realistic customer experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To improve accuracy when identifying performance positions on a musical score on the basis of performance of a user.SOLUTION: A data output method includes steps of: sequentially acquiring input data about performance operation; acquiring a plurality of pieces of estimation information including first estimation information and second estimation information, by providing the input data for a plurality of estimation models including a first estimation model and a second estimation model; identifying musical score performance positions on the basis of the plurality of pieces of estimation information; and reproducing and outputting predetermined data on the basis of the musical score performance positions. The first estimation model is a model indicating relation between performance data about the performance operation and musical score positions in a predetermined musical score. The second estimation model is the model obtained by learning the relation between the performance data and a position within a measure.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technique for outputting data. Background technology

[0002] A technology has been proposed that identifies the playing position on the score of a given piece of music by analyzing sound data obtained when the user plays the piece. A technology has also been proposed that applies this technology to automatic performance, realizing automatic performance that follows the user's performance (for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-207615 Summary of the Invention [Problem to be solved by the invention]

[0004] The accuracy with which an automatic performance follows a user's performance is affected by the accuracy of the performance position that is identified, which can sometimes be reduced due to factors such as the sequence of notes that make up a piece of music.

[0005] One object of the present invention is to improve the accuracy of identifying a performance position on a musical score based on a user's performance. Means to solve the problem

[0006] According to one embodiment, a data output method is provided, which includes: sequentially acquiring input data related to performance operations; providing the input data to a plurality of estimation models, including a first estimation model and a second estimation model, thereby acquiring a plurality of pieces of estimated information, including first estimated information and second estimated information; identifying a musical score performance position corresponding to the input data based on the plurality of pieces of estimated information; and reproducing and outputting predetermined data based on the musical score performance position. The first estimation model is a model that indicates the relationship between performance data related to performance operations and a musical score position in a predetermined musical score, and when the input data is provided, outputs the first estimated information related to the musical score position corresponding to the input data. The second estimation model is a model that indicates the relationship between the performance data and a position in a measure, and when the input data is provided, outputs the second estimated information related to the position in a measure corresponding to the input data. Effect of the invention

[0007] According to the present invention, it is possible to improve the accuracy when identifying a performance position on a musical score based on a user's performance. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram illustrating a system configuration according to a first embodiment. [Figure 2] 1 is a diagram illustrating the configuration of an electronic musical instrument according to a first embodiment. [Figure 3] FIG. 2 is a diagram illustrating the configuration of a data output device in the first embodiment. [Figure 4] FIG. 2 is a diagram illustrating a performance follow-up function in the first embodiment. [Figure 5] FIG. 3 is a diagram illustrating a data output method in the first embodiment. [Figure 6] FIG. 10 is a diagram illustrating a musical score position model in the second embodiment. [Figure 7] FIG. 11 is a diagram illustrating a data generation function in the third embodiment. [Figure 8]FIG. 13 is a diagram illustrating a model generation function for generating a musical score position model in the fourth embodiment. [Figure 9] FIG. 13 is a diagram illustrating a model generation function for generating an intra-bar position model in the fourth embodiment. [Figure 10] 10 is a diagram illustrating a model generation function for generating a beat position model in the fourth embodiment. MODES FOR CARRYING OUT THE INVENTION

[0009] An embodiment of the present invention will be described in detail below with reference to the drawings. The embodiments described below are merely examples, and the present invention should not be construed as being limited to these embodiments. In the drawings referred to in the following embodiments, identical parts or parts having similar functions are designated by the same or similar symbols (symbols consisting of a number followed by A, B, etc.), and repeated explanations may be omitted. For clarity of explanation, the drawings may be described schematically, with some components omitted from the drawings.

[0010] First Embodiment [overview] A data output device according to one embodiment of the present invention realizes automatic performance of a predetermined piece of music by tracking a user's performance on an electronic musical instrument. In this example, the electronic musical instrument is an electronic piano, and the instrument being automatically performed is a vocalist. The data output device provides the user with singing sounds obtained by the automatic performance and a video including an image of the singer. This data output device can accurately identify the position on the musical score where the user is playing, using a performance tracking function described below. The following describes the data output device and a system including the data output device.

[0011] [System Configuration] Fig. 1 is a diagram illustrating the system configuration of the first embodiment. The system shown in Fig. 1 includes a data output device 10 and a data management server 90 connected via a network NW such as the Internet. In this example, an electronic musical instrument 80 is connected to the data output device 10. In this example, the data output device 10 is a computer such as a smartphone, tablet PC, laptop PC, or desktop PC. In this example, the electronic musical instrument 80 is an electronic keyboard device such as an electronic piano.

[0012] As described above, when a user plays a predetermined piece of music using the electronic musical instrument 80, the data output device 10 has a function (hereinafter referred to as a performance follow-up function) for executing an automatic performance that follows the performance and outputting data based on the automatic performance. A detailed description of the data output device 10 will be given later.

[0013] The data management server 90 includes a control unit 91, a storage unit 92, and a communication unit 98. The control unit 91 includes a processor such as a CPU and a storage unit such as RAM. The control unit 91 executes a program stored in the storage unit 92 using the CPU, thereby performing processing according to instructions written in the program. The storage unit 92 includes a storage unit such as a non-volatile memory and a hard disk drive. The communication unit 98 includes a communication module for connecting to a network NW and communicating with other devices. The data management server 90 provides music data to the data output device 10. The music data is data related to automatic performance, and details will be described later. If music data is provided to the data output device 10 by some other method, the data management server 90 need not exist.

[0014] [Electronic Instruments] 2 is a diagram illustrating the configuration of an electronic musical instrument according to the first embodiment. In this example, the electronic musical instrument 80 is an electronic keyboard device such as an electronic piano, and includes performance controls 84, a sound source unit 85, a speaker 87, and an interface 89. The performance controls 84 include multiple keys, and output signals to the sound source unit 85 in response to the operation of each key.

[0015] The sound source unit 85 includes a DSP (Digital Signal Processor) and generates sound data (performance sound data) including sound waveform signals in response to operation signals. The operation signals correspond to signals output from the performance operators 84. The sound source unit 85 converts the operation signals into sequence data (hereinafter referred to as operation data) in a predetermined format for controlling the generation of sound (hereinafter referred to as sound generation), and outputs the converted data to the interface 89. In this example, the predetermined format is MIDI. This enables the electronic musical instrument 80 to transmit operation data corresponding to performance operations on the performance operators 84 to the data output device 10. The operation data is information that specifies the content of sound generation, and is output sequentially as sound generation control information such as note on, note off, and note number. The sound source unit 85 provides the sound data to the interface 89, or can provide the sound data to the speaker 87 instead of providing it to the interface 89.

[0016] The speaker 87 can convert a sound waveform signal corresponding to sound data provided from the sound source unit 85 into air vibrations and provide the sound to the user. The sound data may be provided to the speaker 87 from the data output device 10 via an interface 89. The interface 89 includes a module for transmitting and receiving data to and from an external device wirelessly or via a wired connection. In this example, the interface 89 is connected to the data output device 10 via a wired connection and transmits operation data and sound data generated in the sound source unit 85 to the data output device 10. These data may be received from the data output device 10.

[0017] [Data output device] 3 is a diagram illustrating the configuration of a data output device according to the first embodiment. The data output device 10 includes a control unit 11, a storage unit 12, a display unit 13, an operation unit 14, a speaker 17, a communication unit 18, and an interface 19. The control unit 11 is an example of a computer equipped with a processor such as a CPU and a storage unit such as a RAM. The control unit 11 executes a program 12a stored in the storage unit 12 using the CPU (processor), causing the data output device 10 to realize functions for executing various processes. The functions realized in the data output device 10 include a performance follow-up function, which will be described later.

[0018] The storage unit 12 is a storage device such as a non-volatile memory or a hard disk drive. The storage unit 12 stores a program 12a executed by the control unit 11 and various data such as music data 12b required when executing this program 12a. The storage unit 12 stores three trained models obtained by machine learning. The trained models stored in the storage unit 12 include a musical score position model 210, a bar position model 230, and a beat position model 250.

[0019] The program 12a is downloaded from the data management server 90 or another server via the network NW and stored in the storage unit 12, whereby it is installed in the data output device 10. The program 12a may be provided in a state recorded on a non-transitory computer-readable recording medium (e.g., a magnetic recording medium, an optical recording medium, a magneto-optical recording medium, a semiconductor memory, etc.). In this case, the data output device 10 only needs to be equipped with a device for reading this recording medium. The storage unit 12 can also be considered an example of a recording medium.

[0020] Similarly, the song data 12b may be downloaded from the data management server 90 or another server via the network NW and stored in the storage unit 12, or may be provided in a state recorded on a non-transitory computer-readable recording medium. The song data 12b is data stored in the storage unit 12 for each song, and includes musical score parameter information 121, BPM information 125, singing sound data, and video data 129. Details of the song data 12b, musical score position model 210, intra-bar position model 230, and beat position model 250 will be described later.

[0021] The display unit 13 is a display having a display area that displays various screens under the control of the control unit 11. The operation unit 14 is an operation device that outputs signals to the control unit 11 in response to user operations. The speaker 17 generates sound by amplifying and outputting sound data supplied from the control unit 11. The communication unit 18 is a communication module that connects to the network NW under the control of the control unit 11 and communicates with other devices connected to the network NW, such as a data management server 90. The interface 19 includes a module for communicating with an external device via wireless communication, such as infrared communication or short-range wireless communication, or via wired communication. In this example, the external device includes an electronic musical instrument 80. The interface 19 is used for communication without going through the network NW.

[0022] [Trained model] Next, three trained models will be described. As described above, the trained models include the score position model 210, the bar position model 230, and the beat position model 250. Each trained model is an example of an estimation model that outputs an output value and a likelihood as estimation information for an input value. A known statistical estimation model is applied to each trained model, but different models may also be applied. The estimation model is a machine learning model that uses a neural network that uses, for example, a CNN (Convolutional Neural Network), an RNN (Recurrent Neural Network), or the like. The estimation model is a model that uses a LSTM (Long Short Term The estimation model may be a model using a neural network such as a neural network based on a neural network (NNN) or a gated recurrent unit (GRU), or may be a model that does not use a neural network such as a hidden Markov model (HMM). It is preferable that any estimation model is advantageous for handling time-series data.

[0023] The score position model 210 (first estimation model) is a trained model obtained by machine learning the correlation between performance data and a position on a predetermined score (hereinafter referred to as a score position). In this example, the predetermined score is score data indicating the score of the piano part of a target musical piece, and is described as time-series data in which time information and sound generation control information are associated with each other. The performance data is data obtained by various performers playing while looking at the target musical score, and is described as time-series data in which sound generation control information is associated with time information. The sound generation control information is information that specifies sound generation content such as note-on, note-off, and note number. The time information is, for example, information indicating playback timing relative to the start of the song, and is indicated by information such as delta time and tempo. The time information can also be considered information for identifying a position in the data, and corresponds to a score position.

[0024] The correlation between performance data and score positions indicates the correspondence between the sound generation control information arranged in time series in the performance data and the score data. In other words, this correlation can be said to indicate the data position of the score data corresponding to each data position of the performance data by the score position. The score position model 210 can also be said to be a trained model obtained by learning the performance content (e.g., piano playing style) of various performers when they play while looking at the score.

[0025] When input data corresponding to performance data is sequentially provided, the score position model 210 outputs estimation information (hereinafter referred to as score estimation information) including score positions and likelihoods corresponding to the input data. The input data corresponds, for example, to operation data sequentially output from the electronic musical instrument 80 in response to performance operations on the electronic musical instrument 80. Since the operation data is information sequentially output from the electronic musical instrument 80, it may contain information equivalent to sound generation control information but may not contain time information. In this case, time information corresponding to the time the input data was provided may be added to the input data.

[0026] The score position model 210 is a model obtained by machine learning for each target piece of music. Therefore, the score position model 210 can change the target piece of music by changing a parameter set (hereinafter referred to as score parameters) such as weight coefficients in the intermediate layer. If the score position model 210 is a model that does not use a neural network, the score parameters may be data corresponding to that model. If the score position model 210 uses, for example, dynamic programming (DP) matching to output score estimation information, the score parameters may be the score data itself. The score position model 210 does not have to be a trained model obtained by machine learning, but may be a model that indicates the relationship between performance data and score positions and outputs information corresponding to score positions and likelihoods when input data is provided sequentially.

[0027] The intra-bar position model 230 (second estimation model) is a trained model obtained by machine learning the correlation between performance data and a position in a bar (hereinafter referred to as an intra-bar position). The intra-bar position indicates, for example, any position from the start position to the end position in a bar, and is indicated, for example, by the number of beats and the inter-beat position. The inter-beat position indicates, for example, the position in an adjacent beat as a percentage. For example, if the performance data at a specific data position corresponds to the center between the second and third beats, the number of beats may be "2" and the inter-beat position may be "0.5," and the intra-bar position may be described as "2.5." The intra-bar position does not have to include the inter-beat position, in which case it represents information indicating which beat the bar is included in. The intra-bar position may also be described as a percentage, with the start position of a bar being "0" and the end position being "1."

[0028] The correlation between performance data and bar positions indicates the correspondence between the sound generation control information arranged in chronological order in the performance data and the bar positions. In other words, this correlation can be said to indicate the bar positions corresponding to each data position in the performance data. The bar position model 230 can also be said to be a trained model obtained by learning the bar positions when various performers play various pieces of music.

[0029] When input data corresponding to performance data is sequentially provided, the intra-bar position model 230 outputs estimation information including an intra-bar position and a likelihood (hereinafter referred to as "bar estimation information") corresponding to the input data. The input data corresponds, for example, to operation data sequentially output from the electronic musical instrument 80 in response to performance operations on the electronic musical instrument 80. The input data provided to the intra-bar position model 230 may be data from which information indicating sounding timing has been extracted, excluding information related to pitch such as note number from the operation data.

[0030] The bar position model 230 is a model obtained by machine learning, regardless of the musical piece. Therefore, the bar position model 230 can be used in common for any musical piece. The bar position model 230 may be a model obtained by machine learning for each time signature of the musical piece (duple time, triple time, etc.). In this case, the bar position model 230 can change the target time signature by changing a parameter set such as a weighting coefficient in the intermediate layer. The target time signature may be included in the musical piece data 12b. The bar position model 230 does not have to be a trained model obtained by machine learning; it may be any model that indicates the relationship between performance data and bar positions and outputs information corresponding to bar positions and likelihoods when input data is sequentially provided.

[0031] The beat position model 250 (third estimated information) is a trained model obtained by machine learning the correlation between performance data and positions within one beat (hereinafter referred to as beat positions). A beat position indicates any position within one beat, from the start position to the end position. For example, the beat position may be described as a ratio, with the start position of the beat being "0" and the end position being "1." The beat position may also be described, like a phase, with the start position of the beat being "0" and the end position being "2π."

[0032] The correlation between performance data and beat positions indicates the correspondence between the sound generation control information arranged in chronological order in the performance data and the beat positions. In other words, this correlation can be said to indicate the beat positions corresponding to each data position in the performance data. The beat position model 250 can also be said to be a trained model obtained by learning beat positions when various performers play various pieces of music.

[0033] When input data corresponding to performance data is sequentially provided, the beat position model 250 outputs estimation information including beat positions and likelihoods (hereinafter referred to as beat estimation information) corresponding to the input data. The input data corresponds, for example, to operation data sequentially output from the electronic musical instrument 80 in response to performance operations on the electronic musical instrument 80. The input data provided to the beat position model 250 may be data from which information indicating sound generation timing has been extracted, excluding information related to pitch such as note numbers from the operation data.

[0034] The beat position model 250 is a model obtained by machine learning, regardless of the music piece. Therefore, the beat position model 250 can be used for any music piece. In this example, the beat position model 250 corrects the beat estimation information based on BPM information 125. The BPM information 125 is information indicating the BPM (Beats Per Minute) of the music piece data 12b. The beat position model 250 may recognize the BPM determined from the performance data as an integer fraction or an integer multiple of the actual BPM. By using the BPM information 125, the beat position model 250 can exclude estimated values ​​derived from values ​​that are significantly different from the actual BPM (for example, by reducing the likelihood), thereby improving the accuracy of the beat estimation information. The BPM information 125 may also be used in the intra-bar position model 230. The beat position model 250 does not have to be a trained model obtained by machine learning; it may be a model that indicates the relationship between performance data and beat positions and outputs information corresponding to beat positions and likelihoods when input data is sequentially provided.

[0035] [Song data] Next, the song data 12b will be described. As described above, the song data 12b is data stored in the storage unit 12 for each song, and includes musical score parameter information 121, BPM information 125, singing sound data 127, and video data 129. In this example, the song data 12b includes data for reproducing the singing sound data in accordance with the user's performance.

[0036] As described above, the musical score parameter information 121 includes a parameter set corresponding to a musical piece and used in the musical score position model 210. As described above, the BPM information 125 is information provided to the beat position model 250, and indicates the BPM of the musical piece.

[0037] The singing sound data 127 is sound data including waveform signals of singing sounds corresponding to the vocal parts of a song, and time information is associated with each part of the data. The singing sound data 127 can also be said to be data defining the waveform signals of singing sounds in a time series. The moving image data 129 is moving image data including images of the singer of the vocal part, and time information is associated with each part of the data. The moving image data 129 can also be said to be data defining image data in a time series. The time information in the singing sound data 127 and the moving image data 129 is determined in correspondence with the above-mentioned musical score position. Therefore, a performance using the musical score data, the playback of the singing sound data 127, and the playback of the moving image data 129 can be synchronized via the time information.

[0038] The singing sounds included in the singing sound data may be generated using at least character information and pitch information. For example, the singing sound data may include time information and sound generation control information associated with the time information, similar to musical score data. The sound generation control information includes pitch information such as note numbers as described above, and further includes character information corresponding to lyrics. In other words, the singing sound data may not be data including waveform signals of the singing sounds, but may be control data for generating the singing sounds. The video data may also be control data including image control information for generating an image that resembles a singer.

[0039] [Performance tracking function] Next, the performance follow-up function realized by the control unit 11 executing the program 12a will be described.

[0040] 4 is a diagram illustrating the follow-performance function in the first embodiment. The follow-performance function 100 includes an input data acquisition unit 111, a calculation unit 113, a performance position identification unit 115, and a playback unit 117. The configuration for realizing the follow-performance function 100 is not limited to being realized by executing a program, and at least a part of the configuration may be realized by hardware.

[0041] The input data acquisition unit 111 acquires input data. In this example, the input data corresponds to operation data sequentially output from the electronic musical instrument 80. The input data acquired by the input data acquisition unit 111 is provided to the calculation unit 113.

[0042] The calculation unit 113 includes a musical score position model 210, a bar position model 230, and a beat position model 250, and provides input data to each model and estimates information (musical score estimation information, bar estimation information, and beat estimation information) output from each model to the performance position identification unit 115.

[0043] The score position model 210 functions as a trained model corresponding to a predetermined piece of music by setting weighting coefficients according to the score parameter information 121. As described above, the score position model 210 outputs score estimation information when input data is sequentially provided. This makes it possible to identify the likelihood of a score position for the provided input data. That is, the score estimation information makes it possible to indicate, based on the likelihood for each position, to which position on the score of the music piece the performance content of the user corresponding to the input data corresponds.

[0044] The bar position model 230 is a trained model that is independent of the music piece. When input data is sequentially provided, the bar position model 230 outputs bar estimation information. This makes it possible to identify the likelihood of a bar position for the provided input data. That is, the bar estimation information makes it possible to indicate, based on the likelihood for each position, to which position within one bar the user's performance content corresponding to the input data corresponds.

[0045] The beat position model 250 is a trained model that is independent of the music piece. When input data is sequentially provided, the beat position model 250 outputs beat estimation information. This makes it possible to identify the likelihood of a beat position for the provided input data. That is, the beat estimation information makes it possible to indicate, based on the likelihood for each position, to which position within one beat the content of the user's performance corresponding to the input data corresponds. As described above, the beat position model 250 may use the BPM information 125 as a parameter that is given in advance.

[0046] The playing position identification unit 115 identifies a score playing position based on the score estimation information, bar estimation information, and beat estimation information, and provides the identified position to the playback unit 117. The score playing position is a position on the score identified in response to a performance on the electronic musical instrument 80. The playing position identification unit 115 can also identify the score position with the highest likelihood in the score estimation information as the score playing position, but in this example, bar estimation information and beat estimation information are also used to further improve accuracy. The playing position identification unit 115 corrects the score position in the score estimation information with the bar position in the bar estimation information and the beat position in the beat estimation information.

[0047] The playing position identification unit 115 performs the correction in the following manner as a specific example. First, a first example will be described. The playing position identification unit 115 performs a predetermined calculation (multiplication, addition, etc.) using the likelihood determined for the musical score position, the likelihood determined for the position within the measure, and the likelihood determined for the beat position. The likelihood determined for the position within the measure is applied to each measure repeated in the musical score of the music piece. The likelihood determined for the beat position is applied to each beat repeated in each measure. As a result, the likelihood at each musical score position is corrected by applying the likelihood determined for the position within the measure and the likelihood determined for the beat position. The playing position identification unit 115 identifies the musical score position with the highest likelihood after correction as the musical score playing position.

[0048] Next, a second example will be described. The playing position identification unit 115 performs a predetermined calculation (multiplication, addition, etc.) using the likelihood determined for the position in the measure and the likelihood determined for the beat position of each beat repeated in the measure. The likelihood determined for the beat position is applied to each beat repeated in each measure. As a result, the likelihood determined for the position in the measure is corrected by applying the likelihood determined for the beat position. The playing position identification unit 115 identifies the position in the measure that has the highest likelihood after correction. The playing position identification unit 115 identifies the position in the measure thus identified from among the measures that include the musical score position with the highest likelihood as the musical score playing position.

[0049] When identifying a score performance position based solely on score estimation information, the accuracy of identifying the score performance position may be poor depending on the content of the music piece. For example, when a section with a clear melody is played, it is easy to identify an accurate score position. Therefore, it is possible to improve the accuracy of identifying the score performance position. On the other hand, when a section with little melody change is played, it is heavily influenced by the accompaniment. Since the accompaniment is often independent of the music piece, it is difficult to identify an accurate score position. Therefore, in this example, even if there is a section where an accurate score position cannot be identified, the score estimation information can be corrected to improve the accuracy of the ambiguous score position by identifying a detailed position using bar estimation information and beat estimation information that are independent of the music piece, thereby improving the accuracy of identifying the score performance position.

[0050] The playback unit 117 plays back the singing sound data 127 and the video data 129 based on the score performance position provided by the performance position identification unit 115, and outputs the data as playback data. The score performance position is a position on the score identified in response to a performance on the electronic musical instrument 80. Therefore, the score performance position is also related to the time information described above. The playback unit 117 plays back the singing sound data 127 and the video data 129 by referring to the singing sound data 127 and the video data 129 and reading out each portion of the data corresponding to the time information identified by the score performance position.

[0051] By performing playback in this manner, the playback unit 117 can synchronize the performance of the electronic musical instrument 80 by the user, the playback of the singing sound data 127, and the playback of the video data 129 via the musical score performance position and time information.

[0052] When the playback unit 117 reads out this sound data based on the score performance position, it may read out the sound data based on the relationship between the score performance position and time information, and adjust the pitch according to the readout speed. For example, the pitch may be adjusted so that it becomes the pitch when the sound data is read out at a predetermined readout speed.

[0053] The video data of the playback data is provided to the display unit 13, and an image of the singer is displayed on the display unit 13. The singing sound data of the playback data is provided to the speaker 17, and is output as singing sound from the speaker 17. The video data and singing sound data may be provided to an external device. For example, the singing sound data may be provided to the electronic musical instrument 80, and the singing sound may be output from the speaker 87 of the electronic musical instrument 80. In this way, the performance follow-up function 100 can accurately follow the singing, etc., of the user's performance. As a result, even if the user is playing alone, the user can get the feeling that multiple people are actually playing together. Therefore, a customer experience that gives the user a high sense of realism is provided. This concludes the description of the performance follow-up function.

[0054] [Data output method] Next, we will explain the data output method executed in the performance follow-up function 100. The data output method explained here starts when the program 12a is executed.

[0055] FIG. 5 is a diagram illustrating a data output method in the first embodiment. The control unit 11 acquires input data provided sequentially (step S101) and acquires estimation information from each estimation model (step S103). In this example, the estimation models include the score position model 210, the intra-bar position model 230, and the beat position model 250. The estimation information includes the score estimation information, bar estimation information, and beat estimation information. The control unit 11 identifies the score performance position based on this estimation information (step S105). The control unit 11 plays back the video data and sound data based on the score performance position (step S107) and outputs them as playback data (step S109). The control unit 11 repeats the processes from step S101 to step S109 until an instruction to end the process is input (step S111; No). When an instruction to end the process is input (step S111; Yes), the control unit 11 ends the process.

[0056] Second Embodiment In the second embodiment, a configuration will be described in which at least one of the estimation models separates input data into multiple ranges and has estimation models corresponding to the input data for each range. In this example, a configuration in which the range division configuration is applied to the musical score position model 210 will be described. Although not explained further, the range division configuration may also be applied to at least one of the intra-bar position model 230 and the beat position model 250.

[0057] FIG. 6 is a diagram illustrating a score position model in the second embodiment. The score position model 210A in the second embodiment includes a separation unit 211, a bass model 213, a treble model 215, and an estimation calculation unit 217. The separation unit 211 separates input data into two ranges. For example, the separation unit 211 separates the input data into treble input data, which is obtained by extracting sound generation control information related to note numbers on the treble side based on a predetermined pitch (e.g., C4), and bass input data, which is obtained by extracting sound generation control information related to note numbers on the bass side. The treble input data is data that mainly corresponds to the melody of the music piece, since it is an extracted performance in the treble pitch range. The bass input data is data that mainly corresponds to the accompaniment of the music piece, since it is an extracted performance in the bass pitch range. The input data provided to the score position model 210A can be said to include treble input data and bass input data.

[0058] The bass side model 213 has the same function as the score position model 210 in the first embodiment, but differs in that the performance data used for machine learning is in the same range as the bass side input data. When the bass side input data is provided, the bass side model 213 outputs bass side estimation information. The bass side estimation information is information similar to the score estimation information, but is information obtained using bass range data.

[0059] The treble side model 215 has the same function as the score position model 210 in the first embodiment, but differs in that the performance data used for machine learning is in the same range as the treble side input data. When the treble side input data is provided, the treble side model 215 outputs treble side estimation information. The treble side estimation information is information similar to the score estimation information, but is information obtained using treble range data.

[0060] The estimation calculation unit 217 generates score estimation information based on the bass side estimation information and the treble side estimation information. The likelihood for a score position in the score estimation information may be the larger of the likelihood of the bass side estimation information or the likelihood of the treble side estimation information at each score position, or may be calculated by a predetermined calculation (for example, addition) using each likelihood as a parameter.

[0061] By dividing the bass and treble sides in this way, the accuracy of the treble side estimation information can be increased in sections where the melody of the song is present, while in sections where the melody is not present, the accuracy of the treble side estimation information decreases, but it is possible to use bass side estimation information that is less affected by the melody.

[0062] <Third embodiment> In the third embodiment, a data generation function will be described for generating singing sound data and musical score data from sound data representing a song (hereinafter referred to as song sound data) and registering them in the data management server 90. The generated singing sound data is used as the singing sound data 127 included in the song data 12b in the first embodiment. The generated musical score data is used for machine learning in the musical score position model 210. In this example, a control unit 91 in the data management server 90 realizes the data generation function by executing a predetermined program.

[0063] 7 is a diagram illustrating the data generation function in the third embodiment. Data generation function 300 includes a sound data acquisition unit 310, a vocal part extraction unit 320, a singing sound data generation unit 330, a vocal score data generation unit 340, an accompaniment pattern estimation unit 350, a chord / beat estimation unit 360, an accompaniment score data generation unit 370, a score data generation unit 380, and a data registration unit 390. Sound data acquisition unit 310 acquires music sound data. The music sound data is stored in storage unit 92 of data management server 90.

[0064] The vocal part extraction unit 320 analyzes the music sound data using a known sound source separation technique and extracts data of a portion corresponding to the singing sound corresponding to the vocal part from the music sound data. An example of a known sound source separation technique is the technique disclosed in Japanese Patent Application Laid-Open No. 2021-135446. The singing sound data generation unit 330 generates singing sound data indicating the singing sound extracted by the vocal part extraction unit 320.

[0065] The vocal score data generator 340 identifies information about each note included in the vocal sound, such as pitch and duration, and converts it into sound generation control information and time information that represent the vocal sound. The vocal score data generator 340 generates time-series data that associates the time information and sound generation control information obtained through the conversion, i.e., score data that represents the score of the vocal part of the target song. The vocal part corresponds, for example, to the part played by the right hand in a piano part, and includes the melody of the vocal sound, i.e., the melody notes. The melody notes are determined within a predetermined range.

[0066] The accompaniment pattern estimation unit 350 analyzes the music sound data using known estimation technology to estimate the accompaniment pattern for each section of the music. Examples of known estimation technology include the technology disclosed in Japanese Patent Application Laid-Open No. 2014-29425. The chord / beat estimation unit 360 estimates the beat positions and chord progressions (chords for each section) of the music using known estimation technology. Examples of known estimation technology include the technology disclosed in Japanese Patent Application Laid-Open No. 2015-114361 and Japanese Patent Application Laid-Open No. 2019-14485.

[0067] The accompaniment score data generator 370 generates the content of the accompaniment part based on the estimated accompaniment pattern, beat positions, and chord progression, and generates score data representing the score of the accompaniment part. This score data is time-series data that associates time information representing the accompaniment notes of the accompaniment part with sound generation control information, i.e., the generated score data represents the score of the accompaniment part of the target musical piece. The accompaniment part corresponds, for example, to the part played by the left hand in the piano part, and includes at least one of chords and bass notes corresponding to the chords. The chords and bass notes are each determined within a predetermined range.

[0068] The accompaniment score data generator 370 may not use the estimated accompaniment pattern. In this case, the accompaniment notes may be determined, for example, so that chords and bass notes corresponding to the chord progression are sounded only when the chord changes, in at least a portion of the music piece. In particular, determining the accompaniment notes in this way in a section containing melody notes increases redundancy in the user's performance, thereby improving the accuracy of the score estimation information in the score position model 210.

[0069] The musical score data generator 380 generates musical score data by combining the musical score data for the vocal part and the musical score data for the accompaniment part. As described above, the vocal part corresponds to the part of the piano part played with the right hand, and the accompaniment part corresponds to the part of the piano part played with the left hand. Therefore, this musical score data can be said to represent the musical score when the piano part is played with both hands.

[0070] The score data generation unit 380 may modify some of the score data when generating the score data. For example, the score data generation unit 380 may modify the score data for a vocal part by adding a note one octave apart from each note in at least some sections. Whether the added note is one octave higher or lower may be determined based on the range of the singing voice. That is, if the pitch of the singing voice is lower than a predetermined pitch, a note one octave higher may be added, and if the pitch is equal to or higher than the predetermined pitch, a note one octave lower may be added. In this case, the score represented by the score data can be said to have a pitch one octave lower than the highest pitch. This increases redundancy in the user's performance and improves the accuracy of the score estimation information in the score position model 210.

[0071] The data registration unit 390 registers the singing sound data generated by the singing sound data generation unit 330 and the musical score data generated by the musical score data generation unit 380 in a database stored in the memory unit 92 or the like, in association with information identifying the song.

[0072] In this way, the data generating function 300 can analyze the music sound data to extract singing sound data and generate musical score data corresponding to the music.

[0073] <Fourth embodiment> In the fourth embodiment, a model generation function for generating an estimation model obtained by machine learning will be described. In this example, a control unit 91 in a data management server 90 executes a predetermined program to realize the model generation function. In the above example, the estimation model includes a musical score position model 210, a measure position model 230, and a beat position model 250. Therefore, the model generation function is also realized for each estimation model. In the following description, "teacher data" may be replaced with "training data." The expression "learning a model" may be replaced with "training a model." For example, the expression "a computer learns a learning model using the teacher data" may be replaced with the expression "a computer trains a learning model using the training data."

[0074] FIG. 8 is a diagram illustrating a model generation function for generating a score position model in the fourth embodiment. The model generation function 910 includes a machine learning unit 911. Performance data 913, score position information 915, and score data 919 are provided to the machine learning unit 911. The score data 919 is score data obtained by the data generation function 300 described above. The performance data 913 is data obtained by a performer performing a piece of music while looking at a score corresponding to the score data 919, and is described as time-series data in which sound generation control information and time information are associated with each other. The score position information 915 is information indicating the correspondence between a position in the performance indicated by the performance data 913 (performance position) and a position in the score indicated by the score data 919 (score position). The score position information 915 can also be said to be information indicating the correspondence between the time series of the performance data 913 and the time series of the score data 919.

[0075] A pair of performance data 913 and score position information 915 corresponds to training data in machine learning. Multiple pairs are prepared in advance for each piece of music and provided to the machine learning unit 911. The machine learning unit 911 uses this training data to perform machine learning for each piece of score data 919, i.e., for each piece of music, and generates the score position model 210 by determining weighting coefficients for the intermediate layer. In other words, the computer can generate the score position model 210 by training a learning model using the training data. The weighting coefficients correspond to the above-mentioned score parameter information 121 and are determined for each piece of music data 12b.

[0076] FIG. 9 is a diagram illustrating a model generation function for generating a bar position model in the fourth embodiment. The model generation function 930 includes a machine learning unit 931. The machine learning unit 931 is provided with performance data 933 and bar position information 935. The performance data 933 is data obtained by a performer performing a piece of music while looking at a predetermined musical score, and is written as time-series data in which sound generation control information and time information are associated with each other. The predetermined musical score includes not only scores for specific pieces of music, but also scores for a variety of pieces of music. The bar position information 935 is information indicating the correspondence between a position in the performance indicated by the performance data 933 (performance position) and a bar position. The bar position information 935 can also be said to be information indicating the correspondence between the time series of the performance data 933 and a bar position.

[0077] A pair of performance data 933 and intra-bar position information 935 corresponds to training data in machine learning. A plurality of pairs are prepared in advance and provided to the machine learning unit 931. The training data used in the model generation function 930 does not depend on the musical piece. The machine learning unit 931 performs machine learning using this training data and generates the intra-bar position model 230 by determining weighting coefficients for the intermediate layer. In other words, the intra-bar position model 230 can be generated by having a computer train a learning model using the training data. Because the weighting coefficients do not depend on the musical piece, they can be used in a general-purpose manner.

[0078] FIG. 10 is a diagram illustrating a model generation function for generating a beat position model in the fourth embodiment. The model generation function 950 includes a machine learning unit 951. Performance data 953 and beat position information 955 are provided to the machine learning unit 951. The performance data 953 is data obtained by a performer performing while looking at a predetermined musical score, and is written as time-series data in which sound generation control information and time information are associated with each other. The predetermined musical score includes not only musical scores for specific pieces of music, but also musical scores for various pieces of music. The beat position information 955 is information indicating the correspondence between the positions in the performance indicated by the performance data 953 (performance positions) and the beat positions. The beat position information 955 can also be said to be information indicating the correspondence between the time series of the performance data 953 and the beat positions.

[0079] A pair of performance data 953 and beat position information 955 corresponds to training data in machine learning. A plurality of pairs are prepared in advance and provided to the machine learning unit 951. The training data used in the model generation function 950 does not depend on the music piece. The machine learning unit 951 executes machine learning using the training data and generates the beat position model 250 by determining weighting coefficients for the intermediate layer. In other words, the beat position model 250 can be generated by having a computer train a learning model using the training data. The weighting coefficients do not depend on the music piece and can therefore be used in a general-purpose manner.

[0080] <Modification> The present invention is not limited to the above-described embodiments, and includes various other modified examples. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and are not necessarily limited to those having all of the described configurations. Some modified examples will be described below. The first embodiment will be described as a modified example, but the modified examples can also be applied to other embodiments. Multiple modified examples can also be combined and applied to each embodiment.

[0081] (1) The multiple estimation models included in the calculation unit 113 are not limited to the three estimation models of the score position model 210, the intra-bar position model 230, and the beat position model 250, but may also use two estimation models. For example, the calculation unit 113 does not need to use either the intra-bar position model 230 or the beat position model 250. That is, in the performance following function 100, the playing position identification unit 115 may identify the score playing position using the score estimation information and the bar estimation information, or may identify the score playing position using the score estimation information and the beat position estimation information. The playing position identification unit 115 may identify the score playing position using only the score estimation information.

[0082] (2) The input data acquired by the input data acquisition unit 111 is not limited to time-series data including sound generation control information, but may also be sound data including waveform signals of performance sounds. In this case, the performance data used in the machine learning of the estimation model may also be sound data including waveform signals of performance sounds. In such a case, the score position model 210 may be realized by a known estimation technique. Examples of known estimation techniques include the techniques disclosed in Japanese Patent Application Laid-Open Nos. 2016-99512 and 2017-207615. The input data acquisition unit 111 may convert the operation data in the first embodiment into sound data and acquire it as input data.

[0083] (3) The sound generation control information included in the input data and performance data may be incomplete information that does not include all of the information, as long as it is information that can specify the sound generation content. For example, as sound generation content that does not include a mute instruction, the sound generation control information in the input data and performance data may include note-on and note number, but may not include note-off. The sound generation control information in the performance data may extract sounds from a portion of the musical range of the composition. The sound generation control information in the input data may extract performance operations from a portion of the musical range.

[0084] (4) At least one of the video data and the sound data included in the playback data may not be present. That is, at least one of the video data and the sound data may follow the user's performance as an automatic process.

[0085] (5) The video data included in the playback data may be still image data.

[0086] (6) The functions of the data output device 10 and the electronic musical instrument 80 may be included in a single device. For example, the data output device 10 may be incorporated as a function of the electronic musical instrument 80. A portion of the components of the electronic musical instrument 80 may be included in the data output device 10, or a portion of the components of the data output device 10 may be included in the electronic musical instrument 80. For example, components other than the performance controls 84 of the electronic musical instrument 80 may be included in the data output device 10. In this case, the data output device 10 may generate sound data using a sound source unit from the acquired operation data. A portion of the components of the data output device 10 may be included in a component other than the electronic musical instrument 80, such as a server connected via a network NW or a terminal with which direct communication is possible. For example, the components of the calculation unit 113 of the performance tracking function 100 in the data output device 10 may be included in a server. The performance position on the score may be corrected in accordance with the delay time by measuring the delay time due to communication via the network NW. The correction may include, for example, shifting the performance position on the score to a future position on the score by an amount corresponding to the delay time.

[0087] (7) The control unit 11 may record the playback data output from the playback unit 117 on a recording medium or the like. The control unit 11 may generate recording data for outputting the playback data and record the data on the recording medium. The recording medium may be the storage unit 12 or a computer-readable storage medium connected as an external device. The recording data may be transmitted to a server device connected via the network NW. For example, the recording data may be transmitted to the data management server 90 and stored in the storage unit 92. The recording data may include video data and sound data, or may include singing sound data 127, video data 129, and time-series information on the playing position on the musical score. In the latter case, the playback data may be generated from the recording data by a function corresponding to the playback unit 117.

[0088] (8) The playing position identifying unit 115 may identify a score playing position during a portion of the musical piece, regardless of the estimated information output from the calculation unit 113. In this case, the musical piece data 12b may specify a progression speed of the score playing position to be identified during the portion of the musical piece. The playing position identifying unit 130 may identify the score playing position so that the score playing position changes at the specified progression speed during this period.

[0089] The above is the explanation regarding the modified example.

[0090] As described above, one embodiment of the present invention provides a data output method including: sequentially acquiring input data related to performance operations; providing the input data to a plurality of estimation models, including a first estimation model and a second estimation model, to acquire a plurality of pieces of estimated information, including first estimated information and second estimated information; identifying a musical score performance position corresponding to the input data based on the plurality of pieces of estimated information; and reproducing and outputting predetermined data based on the musical score performance position. The first estimation model is a model that indicates the relationship between performance data related to performance operations and a musical score position in a predetermined musical score, and when the input data is provided, outputs the first estimated information related to the musical score position corresponding to the input data. The second estimation model is a model that indicates the relationship between the performance data and a position in a measure, and when the input data is provided, outputs the second estimated information related to the position in a measure corresponding to the input data.

[0091] The plurality of estimation models may include a third estimation model. The plurality of estimation information may include third estimation information. The third estimation model may be a model that has learned a relationship between the performance data and beat positions, and may output the third estimation information regarding beat positions corresponding to the input data when the input data is provided.

[0092] According to one embodiment of the present invention, there is provided a data output method including: sequentially acquiring input data related to performance operations; providing the input data to a plurality of estimation models, including a first estimation model and a third estimation model, thereby acquiring a plurality of pieces of estimated information, including first estimated information and third estimated information; identifying a musical score performance position corresponding to the input data based on the plurality of pieces of estimated information; and reproducing and outputting predetermined data based on the musical score performance position. The first estimation model is a model that indicates the relationship between performance data related to performance operations and a musical score position in a predetermined musical score, and when the input data is provided, outputs the first estimated information related to the musical score position corresponding to the input data. The third estimation model is a model that indicates the relationship between the performance data and beat positions, and when the input data is provided, outputs the third estimated information related to the beat positions corresponding to the input data.

[0093] At least one of the plurality of estimation models may include a trained model that has been trained on the relationship by machine learning.

[0094] Reproducing the predetermined data may include reproducing sound data.

[0095] The sound data may include singing sounds.

[0096] The reproducing of the sound data may include reading out a waveform signal in accordance with the performance position on the musical score to generate the singing sound.

[0097] Reproducing the sound data may include reading out sound generation control information including character information and pitch information in accordance with the performance position on the musical score, and generating the singing sound.

[0098] The predetermined musical score may have, in at least a portion of the section, a pitch one octave lower than the highest pitch.

[0099] The input data provided to the first estimation model may include first input data extracted from a performance in a first pitch range and second input data extracted from a performance in a second pitch range.

[0100] The first estimation model may generate the first estimated information based on estimation information according to a musical score position corresponding to the first input data and estimation information according to a musical score position corresponding to the second input data.

[0101] A program for causing a processor to execute the above-described data output method may be provided.

[0102] A data output device may be provided which includes a processor for executing the program described above.

[0103] An electronic musical instrument may be provided that includes the data output device described above, a performance operator for inputting the performance operation, and a sound source section that generates performance sound data in accordance with the performance operation. [Explanation of symbols]

[0104] 10: Data output device, 11: Control unit, 12: Memory unit, 12a: Program, 12b: Music data, 13: Display unit, 14: Operation unit, 17: Speaker, 18: Communication unit, 19: Interface, 80: Electronic musical instrument, 84: Performance operator, 85: Sound source unit, 87: Speaker, 89: Interface, 90: Data management server, 91: Control unit, 92: Memory unit, 98: Communication unit, 100: Performance tracking function, 111: Input data acquisition unit, 113: Calculation unit, 115: Performance position identification unit, 117: Playback unit, 121: Musical score parameter information, 125: BPM information, 127: Singing sound data, 129: Video data, 130: Performance position identification unit, 210, 210A: Musical score position model, 211: Separation unit, 213: Bass side model, 215: Treble side model, 217: estimation calculation unit, 230: bar position model, 250: beat position model, 300: data generation function, 310: sound data acquisition unit, 320: vocal part extraction unit, 330: singing sound data generation unit, 340: vocal score data generation unit, 350: accompaniment pattern estimation unit, 360: beat estimation unit, 370: accompaniment score data generation unit, 380: score data generation unit, 390: data registration unit, 910: model generation function, 911: machine learning unit, 913: performance data, 915: score position information, 919: score data, 930: model generation function, 931: machine learning unit, 933: performance data, 935: bar position information, 950: model generation function, 951: machine learning unit, 953: performance data, 955: beat position information

Claims

1. A data processing method executed by a processor, comprising: Get the input data, identifying pitch and timing information contained in said input data; generating output data including musical score data based on the identified information; data processing methods, including

2. The input data is music sound data, extracting sound data corresponding to a vocal part from the music sound data; generating musical score data for a vocal part and musical score data for an accompaniment part based on the musical piece sound data; The data processing method of claim 1 , further comprising:

3. A data processing method as described in claim 1, wherein the musical score data is time series data in which pronunciation control information and time information are associated based on the pitch and the note length corresponding to the pitch.

4. A data processing method as described in claim 1, further comprising a step of playing back singing sound data and video data corresponding to the sheet music data in synchronization based on time information contained in the sheet music data.

5. A data processing method as described in claim 4, wherein the step of playing the singing sound data and the video data further includes the steps of providing the video data to a display device for display, and providing the voice singing sound data to an audio output device for playback.

6. Acquire performance data relating to performance operations, 5. The data processing method according to claim 4, further comprising the step of estimating a performance position from said performance data using an estimation model including a neural network.

7. A data processing method as described in claim 6, further comprising a step of correcting estimated information regarding the playing position obtained by the estimation model based on auxiliary information regarding the beat position or the position within the bar.

8. A data processing method as described in claim 6, wherein the step of playing the singing sound data and the video data further includes a step of reading out the singing sound data based on the relationship between the musical score performance position identified based on the estimated performance position and time information, and adjusting the pitch according to the reading speed.

9. A data processing method as described in Claim 8, wherein the process of playing the singing sound data and the video data includes a process of synchronously playing the singing sound data in accordance with the user's performance based on the musical score performance position.

10. A data acquisition unit that acquires input data; an identification unit that identifies pitch and timing information included in the input data; a data generating unit that generates output data including musical score data based on the identified information; a data processing system including:

11. To a computer Obtaining input data; identifying pitch and timing information contained in said input data; generating output data including musical score data based on the identified information; A program that executes.