Performance information prediction device, effective string vibration determination model training device, performance information generation system, performance information prediction method, and effective string vibration determination model training method

The preprocessing unit and trained model in guitar controllers differentiate between intentional and unintentional string vibrations, enhancing the accuracy of performance information generation and musical output.

JP2025100899APending Publication Date: 2025-07-03CASIO COMPUTER CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2025072316
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing guitar controllers fail to accurately distinguish between intentional and unintentional string vibrations, leading to incorrect performance information generation, particularly in dynamic range, display, and monophonic mode, which affects the quality of musical performances and analysis.

Method used

A preprocessing unit generates volume and pitch characterization data from string vibration waveforms, and a trained effective string vibration determination model predicts performance information, such as MIDI messages, to differentiate between valid and invalid string vibrations.

Benefits of technology

Accurately converts electronic stringed instrument performances into performance information, eliminating unnecessary vibrations and improving the accuracy of musical sound production and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025100899000001_ABST
    Figure 2025100899000001_ABST
Patent Text Reader

Abstract

To provide a technique capable of precisely converting the performance of an electronic stringed instrument into a piece of performance information.SOLUTION: A performance information prediction device in an embodiment includes: a pre-processing unit that generates a piece of volume-pitch featured data including volume information and pitch information of each string from string vibration waveform data representing the performance of a stringed instrument; and a performance information prediction unit that predicts the performance information of the performance of the stringed instrument from the volume-pitch featured data using a trained effective string vibration determination model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a performance information prediction device, a valid string vibration determination model training device, a performance information generation system, a performance information prediction method, and a valid string vibration determination model training method.

Background Art

[0002] There is an electronic musical instrument called a guitar controller (or guitar synthesizer) that converts the string vibration waveform of an instrument such as a guitar into an electrical signal by a magnetic pickup or a piezo pickup and analyzes its pitch and volume to convert it into digital performance data such as MIDI (Musical Instrumental Digital Interface) messages. Different from a dedicated guitar controller that only plays a sound source, this type of controller has a great merit in that it can perform MIDI performances while retaining the shape and functions of a normal guitar and mounting independent pickups on each string for acquiring performance information, and can be said to be the most common form.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] In such non-keyboard musical instruments, one of the major problems that has not been solved for years is that string vibrations unintentionally generated by the player are also converted into performance information without discrimination. For example, when playing while pressing the third string, the palm may touch the open fourth string and cause it to vibrate slightly, or when plucking the first string with a pick, the momentum may cause it to also touch the second string and generate vibrations for a short while. This kind of situation often occurs even in the performances of highly skilled players. It would be ideal if all unnecessary string vibrations could be muted while leaving only the necessary ones, but in many cases, it is almost impossible depending on the performance.

[0005] There are also problems specific to guitar controllers. These phenomena, unlike the natural or electric guitar sounds in normal playing, become more serious when used as a guitar controller for the following reasons.

[0006] Case 1: Performance with a narrow dynamic range When playing an electric guitar, sometimes the volume is intentionally varied, but often it is preferred to produce the sound at the same volume regardless of the loudness or softness of the performance sound, such as in a distortion sound. However, in a normal electric guitar, since the signals of all strings are mixed and distorted, the sound of the small open strings is masked by the sound of the strings with large amplitudes, so it is not much of a problem. However, when playing a synthesizer or the like using a guitar controller, the problem becomes much more serious. It is not uncommon for a synthesizer to also produce sound without varying the strength. When playing a guitar synthesizer with independent sound sources for each string, the sound produced for the slightly vibrating open strings and the vibrations of the strings played normally will have the same volume, so the musical sounds of unnecessary string vibrations will be reproduced at almost the same volume as the intentionally played musical sounds, greatly hindering the performance.

[0007] Case 2: Display and performance analysis The guitar controller can not only sound the sound source, but also use performance information for various purposes, such as displaying the score of the performance, discriminating chords, and inputting phrases into the computer to create music data. At this time, if completely unnecessary note information is mixed in, it may result in a very difficult-to-read score or be misinterpreted as completely wrong chords, which can lead to fatal problems. To solve this, it is necessary to individually remove unnecessary information from the recorded performance information, which can be very time-consuming.

[0008] Case 3: Monophonic mode A guitar has multiple strings and can play chords, but there are cases where one might deliberately play monophonically (single notes) for solo or melody playing. In particular, it is also possible to play legato phrases that smoothly move to another pitch with a portamento effect. For the same reasons in keyboard instruments, many polyphonic keyboard instruments that can play chords also have a monophonic mode that produces only single notes. Some guitar synthesizers also have a built-in sound source that supports such a monophonic mode, or there are cases where an external monophonic sound source is driven by transmitting a MIDI signal externally. Similar to the situation where only one key becomes effective even if multiple keys on the keyboard are pressed, in the case of the monophonic mode of a guitar, only one of the strings vibrates effectively even if multiple strings vibrate simultaneously. In the case of a keyboard, it is common for one note to be determined by rules such as high note priority, low note priority, or last-come first-served priority. However, in the case of a guitar controller, the last-come first-served priority case is more common. However, in the case of a guitar, compared to a keyboard instrument, there are cases where notes with widely separated pitches are inadvertently generated, so if unnecessary notes are mixed into a single-note phrase, the performance phrase will be greatly disrupted. That is, another pronunciation is muted and replaced by the pronunciation of that inadvertent note. What is even more influential is the performance when applying the aforementioned portamento effect. For example, when playing using the high positions of the 1st and 2nd strings, it is quite common for the sound of the open 6th string, which is far away, to vibrate for an instant. In this case, the pitch during pronunciation suddenly moves in an unintended low direction, and a meaningless and large pitch change occurs until the correct note is generated and the pitch returns.

[0009] Conventionally, in a conventional guitar controller, as a countermeasure against these unnecessary string vibrations, the signal level of the string is monitored, and simply, a sound below a predetermined volume is invalidated.

[0010] However, as a side effect of this measure, musical sounds of playing techniques such as weak picking and soft tapping are also ignored together. The fact that what should be pronounced is not pronounced is a serious problem, and currently it is a major obstacle to high-speed phrase playing and the like.

[0011] In view of the above problems, an object of the present disclosure is to provide a technique for accurately converting the performance of an electronic stringed instrument into performance information.

Means for Solving the Problems

[0012] To solve the above problems, one aspect of the present disclosure includes a preprocessing unit that generates volume pitch characterization data including volume information and pitch information of each string from string vibration waveform data representing stringed instrument performance, and a trained effective string vibration determination model And a performance information prediction unit that predicts performance information of the stringed instrument performance from the volume pitch characterization data.

Effects of the Invention

[0013] According to the present disclosure, the performance of an electronic stringed instrument can be accurately converted into performance information.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Embodiments for Carrying Out the Invention

[0015] In the following embodiments, a guitar controller that generates performance information (for example, MIDI messages, etc.) from a string vibration waveform generated by playing a guitar is disclosed. Note that the present disclosure is not limited to a guitar controller, and may be applied to any other performance information generation device that generates performance information from the performance of a stringed instrument having a string vibration waveform extraction function. [Overview of the Present Disclosure] Outlining the embodiments described later, as shown in FIG. 1, a guitar controller 10 according to an embodiment of the present disclosure includes a guitar 50 and a control device 100. The guitar controller 10 generates performance information from a string vibration waveform generated by playing the guitar 50 using a valid string vibration determination model realized as a machine learning model such as a neural network.

[0016] According to an embodiment of the present disclosure, the performance information shows performance types of sounding information, muting information, and pitch change information, as shown in FIG. 2.

[0017] The sounding information indicates intentional plucking discriminated by an effective string vibration determination model as a classification model and the strength of plucking by envelope detection. When expressing the sounding by MIDI messages, the note number of the intentional plucking and the strength of plucking may be represented by Note On: 0x9n,kk,vv.

[0018] Also, the muting information is detected by envelope detection and represents 0) sound stop and 1) replacement. When expressing the muting by MIDI messages, it may be represented by Note Off: 0x8n,kk,vv.

[0019] Also, the pitch change information is detected by zero crossing count. When expressing the pitch change by MIDI messages, it may be represented by Pitch Bend: 0xEn,ll,mm.

[0020] In the embodiment shown in FIG. 1, the guitar controller 10 has two operation modes: a training mode for training the effective string vibration determination model and a performance mode for predicting performance information using the trained effective string vibration determination model. The control device 100 has an effective string vibration determination model training device 200 used in the training mode and a performance information prediction device 300 used in the performance mode.

[0021] First, in the training mode, the guitar controller 10 acquires training data from the training performance information database 80. The training data is composed of, for example, a pair of sheet music data (e.g., TAB score, etc.) and a MIDI file corresponding to the sheet music data. The TAB score may be described according to a well-known notation as shown in, for example, FIG. 3. When the user plays the guitar 50 according to the sheet music of the acquired training sheet music data, the effective string vibration determination model training device 200 inputs the string vibration information generated by the guitar 50 based on the user's performance into the effective string vibration determination model to be trained, compares the MIDI message as the performance information output from the effective string vibration determination model with the training MIDI file, and trains the effective string vibration determination model so that these errors are reduced. In the present disclosure, performance information is acquired from the string vibration determined to be effective by the effective string vibration determination model using the volume pitch characterization data characterized based on the volume and pitch extracted from the string vibration waveform data. When the training is completed, the effective string vibration determination model training device 200 provides the trained effective string vibration determination model to the performance information prediction device 300.

[0022] Next, in the performance mode, when the user plays the guitar 50, the performance information prediction device 300 inputs the string vibration information generated by the guitar 50 based on the user's performance into the trained effective string vibration determination model, and acquires performance information such as MIDI messages. The acquired performance information is transmitted to, for example, an external playback device or a computer, and the user can play back the performance by the user via the playback device or use the performance information on the computer.

[0023] Thereby, it becomes possible to eliminate unnecessary string vibrations when converting the performance of the electronic stringed instrument into performance information and generate performance information based on effective string vibrations.

[0024] Note that the guitar controller 10 according to the embodiments described below has the valid string vibration determination model training device 200, but the present disclosure is not limited thereto. For example, the valid string vibration determination model may be trained by an external computer or server, and the trained valid string vibration determination model and / or update information of the valid string vibration determination model may be provided from the external computer or server to the performance information prediction device 300. [Hardware Configuration] Next, with reference to FIG. 4, the physical configuration of the guitar controller 10 will be described. FIG. 4 is a diagram showing the appearance of the guitar controller 10 according to an embodiment of the present disclosure.

[0025] As shown in FIG. 4, the guitar controller 10 is a separate type performance information generation system composed of an interconnected guitar 50 and a control device 100.

[0026] The guitar 50 is an ordinary electric guitar equipped with a hexadecide pick-up for picking up the independent vibrations of each of the six strings, a MIDI volume for controlling the volume of performance information, and an up / down switch for switching the patch memory number up and down with respect to the control device 100. These pieces of information and the output of a normal pick-up are transmitted to the control device 100 by a multi-cable. Also, the power supply is supplied from the control device 100 via the multi-cable. The hexadecide pick-up of the present embodiment is the same magnetic pick-up as a normal pick-up.

[0027] On the one hand, the control device 100 receives the input of the vibration of the strings of the guitar and generates performance information in the MIDI format. The destination for transmitting the performance information may be, without limitation, a sound source unit, a computer, or the like. As shown in FIG. 1, the control device 100 has a footswitch for switching the bank number and number of the patch memory storing various settings, a CONTROL switch for assigning and transmitting an arbitrary performance message, and a foot pedal. The number of the currently selected patch memory is displayed on the BANK / NUM screen. There is an LCD as the main display device, and a touch panel is mounted on the screen. Also, a rotary encoder for inputting data is provided on the panel. As terminals, there are provided a GUITAR INPUT which is an input terminal for a multi-cable from the guitar 50, a GUITAR OUT which is an audio output terminal for a normal pickup, a MIDI OUT which is an output terminal for a MIDI performance signal, a USB to HOST which is a connection terminal to a host computer, and an AC POWER which is an AC power input terminal.

[0028] Next, referring to FIG. 5, the hardware configuration of the guitar 50 according to an embodiment of the present disclosure will be described. FIG. 5 is a block diagram showing the hardware configuration of the guitar 50 according to an embodiment of the present disclosure.

[0029] As shown in FIG. 5, for the guitar 50, signals through the buffer amplifier of the hexadecided pickup, the MIDI volume control, the up / down switch of the patch memory, and the signals of the normal pickup are transmitted to the control device 100 via a multi-cable. The three normal pickups are selected by a pickup selector, and those passing through the buffer amplifier via a tone control circuit and a volume control circuit are transmitted to the control device 100.

[0030] Next, referring to FIG. 6, the hardware configuration of the control device 100 according to an embodiment of the present disclosure will be described. FIG. 6 is a block diagram showing the hardware configuration of the control device according to an embodiment of the present disclosure.

[0031] As shown in FIG. 6, the control device 100 is composed of a CPU (Central Processing Unit) and a DSP (Digital Signal Processor). The CPU manages the functions and processes of the entire control device 100, and the DSP executes waveform analysis processing that requires high-speed processing. Connected to the CPU bus are a RAM used by the CPU, a Flash ROM, an LCD controller that controls the LCD, an I / O interface connected to various I / O devices, the DSP, a USB interface, and a MIDI interface. Further, connected to the I / O interface are a footswitch, a rotary encoder, an LCD touch panel, the MIDI volume of the guitar 50, an A / D converter for detecting the position of the pedal of the control device 100, and an LED for displaying the patch memory number. Although only one A / D converter is shown in the figure, the input source is switched in time division by a multiplexer and the value is read. A dedicated RAM and Flash ROM are connected to the DSP, and an independent A / D converter for quickly digitizing the outputs of the six strings of the hexadecimal divided pickup is connected, enabling high-speed analysis processing.

[0032] However, the guitar 50 and the control device 100 are not limited to the above-described hardware configuration, and may be realized by any other appropriate hardware configuration. [Performance Information Prediction Device] Next, with reference to FIGS. 7 to 12, a performance information prediction device 300 according to an embodiment of the present disclosure will be described. The performance information prediction device 300 uses an effective string vibration determination model trained by the effective string vibration determination model training device 200 to predict performance information (for example, a MIDI file, etc.) from the performance of the guitar 50 by a performer. Specifically, when the performance information prediction device 300 acquires string vibration waveform data representing a guitar performance from the guitar 50, it performs volume detection and pitch detection on each acquired string vibration waveform data frame of each string, and acquires volume pitch characterization data including the detected volume information and pitch information of each string. Then, the performance information prediction device 300 inputs the acquired volume pitch characterization data into the trained effective string vibration determination model to determine the effectiveness of the plucking of each string. The performance information prediction device 300 generates muting information and pitch change information respectively based on the volume and pitch of each string determined to be effective, and generates velocity information indicating the strength (speed) of the plucking based on the volume, and adds it to the pronunciation information. In this way, the performance information prediction device 300 generates performance information (for example, MIDI messages, etc.) at each time including pronunciation information, muting information, and / or pitch change information from the volume pitch characterization data at each time, and transmits it to an external device (for example, a playback device, a computer, etc.). For example, these performance information generation processes may be realized by a CPU.

[0033] FIG. 7A is a block diagram showing the functional configuration of the performance information prediction device 300 according to an embodiment of the present disclosure.

[0034] As shown in FIG. 7A, the performance information prediction device 300 includes a preprocessing unit 310 and a performance information prediction unit 320.

[0035] The preprocessing unit 210 generates volume-pitch characterization data including volume information and pitch information for each string from the string vibration waveform data representing the performance of a stringed instrument. Specifically, when the guitar 50 is played by a player, the guitar 50 acquires string vibration waveform data indicating time and the amplitude of each string, and transmits it to the performance information prediction device 300. That is, since the guitar 50 has six strings, six types of string vibration waveform data are generated. For example, the string vibration waveform data can be acquired at a sampling rate of 20 KHz (interval of 50 μsec) and a data word length of 16 bits, which is sufficient for the analysis of string vibration.

[0036] The preprocessing unit 310 performs volume detection and pitch detection on the string vibration waveform data of each string, and acquires a volume envelope and a pitch envelope for each string as shown in FIG. 8. For example, regarding volume detection, as shown in FIG. 9(a), the preprocessing unit 310 may absolute-value the string vibration waveform data, perform low-pass filtering, and determine the amplitude level of the filtered waveform as the volume at the reference time. When the trained effective string vibration determination model does not detect an effective plucking for the reference time, or when the detected volume is below a predetermined threshold recognized as a muted state, the preprocessing unit 310 determines that there is no pronunciation and outputs mute information as performance information. Otherwise, the preprocessing unit 310 determines that there is pronunciation, sets the detected volume as the velocity value of the pronunciation, and includes the velocity value in the pronunciation information together with the note number output from the effective string vibration determination model.

[0037] Also, regarding pitch detection, as shown in FIG. 9(b), the preprocessing unit 320 performs high-pass filtering and low-pass filtering on the string vibration waveform data, counts the zero-crossing points, and generates pitch information. Then, when there is a difference from the most recent pitch information or pronunciation information, the preprocessing unit 320 determines that there is a pitch change and outputs pitch change information.

[0038] Note that the preprocessing unit 310 executes the above-described processing in parallel for all strings, detects string vibrations at a predetermined processing period s, and generates string vibration waveform data frames that overlap with respect to the time axis as shown in FIG. 8. As shown in FIG. 8, the preprocessing unit 310 generates a volume pitch characterization data frame including the volume information and pitch information of each string.

[0039] The performance information prediction unit 320 predicts performance information of a string instrument performance from the volume pitch characterization data using a trained valid string vibration determination model. Specifically, the performance information prediction unit 320 inputs the volume pitch characterization data frame generated by the preprocessing unit 310 into the trained valid string vibration determination model, and determines whether each string is valid or invalid. For example, the trained valid string vibration determination model receives, as inputs, a volume pitch characterization data frame at a reference time, p volume pitch characterization data frames immediately before the reference time, and n volume pitch characterization data frames immediately after the reference time, and determines valid information indicating whether each string at the reference time is valid or invalid. Here, the predetermined numbers p and n may be the same or different predetermined values. For example, it is preferable that the predetermined numbers p and n are set to values such that the time lag between the plucking of the guitar 50 by the performer and the output of the performance information in the performance information prediction device 300 cannot be recognized by the performer. By using the volume pitch characterization data frames in a certain time range in this way, it is possible to determine whether each string is an intended plucking considering the sequential relationship between the frames, and it is also possible to determine the time change of the plucking.

[0040] In this way, when the validity of each string is determined, as shown in FIG. 10, the performance information prediction unit 320 determines the volume detected for the string for which the string vibration at the reference time is determined to be valid as the volume of the string, and sets the volume detected for the string for which the string vibration at the reference time is determined to be invalid to zero. Then, the performance information prediction unit 320 generates performance information by constructing pronunciation information only from valid string vibrations and constructing pitch change information based on the detected pitch.

[0041] In addition, when there is no sound at the reference time, that is, when the reference time is in a muted state, the effective string vibration determination model may be trained to output invalid for all strings.

[0042] In one embodiment, the effective string vibration determination model may be implemented by a neural network. For example, the effective string vibration determination model may be a neural network having a network architecture as shown in FIG. 11. In this case, the performance information prediction unit 320 inputs the volume pitch characterization data frame at the reference time and the (p + n) volume pitch characterization data frames before and after the reference time to the input layer of the neural network, and obtains, from the output layer via the operations in the intermediate layer, values indicating the effectiveness or ineffectiveness of each string.

[0043] Also, in other embodiments, the effective string vibration determination model may be implemented by a recurrent neural network as shown in FIG. 12. In this case, the performance information prediction unit 320 may utilize volume pitch characterization data frames in a time range wider than the time range based on p and n described above. For example, the volume pitch characterization data frame at the reference time t, b (b > p) volume pitch characterization data frames immediately before the reference time, and f (f > n) volume pitch characterization data frames immediately after the reference time are input to the input layers Xt-b, ···, Xt-1, Xt, Xt+1, ···, Xt+f of the recurrent neural network, and values indicating the effectiveness or ineffectiveness of each string are obtained from the output layer via the operations in the intermediate layer. The recurrent neural network is suitable for processing time series data, and it is considered that the effectiveness of each string can be predicted with high accuracy from volume pitch characterization data frames in a certain time range. [Performance Information Prediction Process] Next, with reference to FIG. 13, a performance information prediction process according to an embodiment of the present disclosure will be described. The performance information prediction process is implemented by the performance information prediction apparatus 300 described above, and may be implemented, for example, when a processor executes a program or instructions. FIG. 13 is a flowchart showing a performance information prediction process according to an embodiment of the present disclosure.

[0044] As shown in FIG. 13, in step S101, the performance information prediction device 300 initializes the string number s to 0. Since the guitar 50 is composed of six strings, the string number s takes values from 0 to 5.

[0045] In step S102, the performance information prediction device 300 detects the current volume l[s] of string s from the string vibration waveform data. For example, the performance information prediction device 300 may apply absolute value conversion and low-pass filtering to the string vibration waveform data to detect the volume l[s].

[0046] In step S103, the performance information prediction device 300 detects the current pitch p[s] of string s from the string vibration waveform. For example, the performance information prediction device 300 may apply high-pass filtering and low-pass filtering to the string vibration waveform data to detect the pitch p[s].

[0047] In step S104, the performance information prediction device 300 increments the string number s by 1.

[0048] In step S105, the performance information prediction device 300 determines whether the volume and pitch have been detected for all strings.

[0049] In step S106, the performance information prediction device 300 stores the current volumes l[0] to l[5] and pitches p[0] to p[5] of each of the six strings.

[0050] In step S107, the performance information prediction device 300 inputs the volume-pitch characterization data frame at the reference time, the p volume-pitch characterization data frames immediately before the reference time, and the n volume-pitch characterization data frames immediately after the reference time into the trained effective string vibration determination model.

[0051] In step S108, the performance information prediction device 300 stores the determination results for each string in a[0] to a[5], respectively.

[0052] In step S109, the performance information prediction device 300 supplies the digital signal processor (DSP) with the string vibration valid information a[0] to a[5] indicating the determination result of each string, and controls the level signal. That is, the DSP may maintain the level of the string vibration for the string whose determination result is valid string vibration, and set the level of the string vibration to silent for the string whose determination result is not valid string vibration.

[0053] In step S110, the performance information prediction device 300 resets the string number s to 0.

[0054] In step S111, the performance information prediction device 300 stores the volume information output by the DSP in l[s].

[0055] In step S112, the performance information prediction device 300 stores the pitch information output by the DSP in p[s].

[0056] In step S113, the performance information prediction device 300 determines whether string s is sounding. If string s is not sounding (S113: No), the performance information prediction device 300 proceeds to step S117. On the other hand, if string s is sounding (S113: Yes), the performance information prediction device 300 determines in step S114 whether the string vibration of string s is invalid, that is, whether the string vibration valid information of string s is a[s]=0.

[0057] If the string vibration of string s is invalid (S114: Yes), the performance information prediction device 300 proceeds to step S116. On the other hand, if the string vibration of string s is valid (S114: No), the performance information prediction device 300 determines in step S115 whether the volume l[s] is less than a predetermined minimum sounding volume.

[0058] In step S115, the performance information prediction device 300 determines whether the volume l[s] is greater than a predetermined pronunciation level. If the volume l[s] is greater than the predetermined pronunciation level (S115: Yes), the performance information prediction device 300 generates muting information for the current pronunciation of the string s in step S116. On the other hand, if the volume l[s] is less than or equal to the predetermined pronunciation level (S115: No), the performance information prediction device 300 proceeds to step S117.

[0059] In step S117, the performance information prediction device 300 determines whether l[s] is greater than the pronunciation level. If l[s] is greater than the pronunciation level (S117: Yes), the performance information prediction device 300 proceeds to step S118. On the other hand, if l[s] is less than or equal to the pronunciation level (S117: No), the performance information prediction device 300 proceeds to step S122.

[0060] In step S118, the performance information prediction device 300 determines whether the string vibration validity information of the string s is a[s]=0. If the string vibration validity information of the string s is a[s]=0 (S118: Yes), the performance information prediction device 300 proceeds to step S122. On the other hand, if the string vibration validity information of the string s is not a[s]=0 (S118: No), the performance information prediction device 300 calculates the note number k based on p[s] in step S119.

[0061] In step S120, the performance information prediction device 300 calculates the velocity information v based on l[s].

[0062] In step S121, the performance information prediction device 300 generates pronunciation information of the note number k and the velocity information v.

[0063] In step S122, the performance information prediction device 300 determines whether the pitch p [s] matches the pitch P0 [s] of the previous string s. If they match (S122: Yes), the performance information prediction device 300 proceeds to step S125. On the other hand, if they do not match (S122: No), the performance information prediction device 300 calculates the difference and generates pitch bend information.

[0064] In step S124, the performance information prediction device 300 updates P0 [s] with p.

[0065] In step S125, the performance information prediction device 300 increments the string number s by 1.

[0066] In step S126, the performance information prediction device 300 determines whether all strings have been processed. If not all strings have been processed (S126: Yes), it returns to step S111. If all strings have been processed (S126: No), the process ends. [Effective String Vibration Determination Model Training Device] Next, with reference to FIGS. 14 and 7B, an effective string vibration determination model training device 200 according to an embodiment of the present disclosure will be described. FIG. 14 is a schematic diagram showing the operation of the effective string vibration determination model training device 200 according to an embodiment of the present disclosure.

[0067] As shown in FIG. 14, the effective string vibration determination model training device 200 trains an effective string vibration determination model using the training data stored in the performance information database 80 for training that stores training data. Specifically, the training data is composed of a pair of training score data and training performance information corresponding to the score data. A score (for example, a TAB score, etc.) displayed based on the training score data is displayed to the performer, and the performer plays the guitar 50 under the tempo control by a metronome. The string vibration waveform data representing the performance is provided to the effective string vibration determination model training device 200. The effective string vibration determination model training device 200 performs volume detection and pitch detection on the acquired string vibration waveform data of each string, and generates a volume pitch characterization data frame composed of the volume and pitch of each string. Then, the effective string vibration determination model training device 200 inputs the volume pitch characterization data frames at and around the reference time into the effective string vibration determination model to be trained. Then, the effective string vibration determination model training device 200 compares the output from the effective string vibration determination model with the pronunciation information (for example, pronunciation information and muting information, etc.) of the training performance information, and updates the parameters of the effective string vibration determination model according to the error. The effective string vibration determination model training device 200 repeats the above-described processing until a predetermined end condition is satisfied, and optimizes the effective string vibration determination model so that the output from the effective string vibration determination model approaches the pronunciation information of the training performance information.

[0068] FIG. 7B is a block diagram showing a functional configuration of an effective string vibration determination model training device 200 according to an embodiment of the present disclosure.

[0069] As shown in FIG. 7B, the effective string vibration determination model training device 200 includes a preprocessing unit 210 and an effective string vibration determination model training unit 220.

[0070] The preprocessing unit 210 generates volume pitch characterization data composed of the volume and pitch of each string from the string vibration waveform data representing the string instrument performance played according to the performance information for training. Specifically, when the guitar 50 is played by a player, the guitar 50 acquires string vibration waveform data indicating time and the amplitude of each string, and transmits it to the effective string vibration determination model training device 200. That is, since the guitar 50 has six strings, six types of string vibration waveform data are generated.

[0071] The preprocessing unit 210 performs volume detection and pitch detection on the string vibration waveform data of each string, and acquires volume pitch characterization data composed of the detected volume and pitch of each string. For example, the preprocessing unit 210 may extract a string vibration waveform frame with a window width w overlapping with respect to the time axis from the string vibration waveform data, perform volume detection and pitch detection for each sampling, and convert each string vibration waveform frame into a volume pitch characterization data frame. The preprocessing unit 210 generates volume pitch characterization data and provides it to the effective string vibration determination model training unit 220.

[0072] The effective string vibration determination model training unit 220 uses the performance information for training to train an effective string vibration determination model that determines the effectiveness of each string in string instrument performance from the volume pitch characterization data frame at the reference time and the volume pitch characterization data frames before and after the volume pitch characterization data frame at the reference time. Here, when predicting the performance information at the reference time to be predicted, the effective string vibration determination model to be trained obtains not only the volume pitch characterization data frame at the reference time but also the volume pitch characterization data frames at the times before and after the reference time as inputs, and outputs the effectiveness information of each string at the reference time. For example, the effective string vibration determination model training unit 220 may input the volume pitch characterization data frame at the reference time, p volume pitch characterization data frames immediately before the reference time, and n volume pitch characterization data frames immediately after the reference time into the effective string vibration determination model. Here, the predetermined numbers p and n may be the same or different predetermined values. For example, it is preferable that the predetermined numbers p and n are set to values such that the time lag between the plucking of the guitar 50 by the performer and the output of the performance information in the performance information prediction device 300 cannot be recognized by the performer.

[0073] By using the volume pitch characterization data frames within a certain time range in this way, it is possible to determine whether a new plucking has occurred in consideration of the sequential relationship between the frames, and it is also possible to determine the time change of the plucking.

[0074] Also, the effective string vibration determination model to be trained may be a pre-trained machine learning model, and the effective string vibration determination model training unit 220 may fine-tune the pre-trained effective string vibration determination model by the above-described training process. Thereby, it becomes possible to construct a highly accurate effective string vibration determination model with less training data as compared with training the effective string vibration determination model from the machine learning model in the initial state.

[0075] When the effective string vibration determination model acquires the effective information of each string, the effective string vibration determination model training unit 220 compares the pronunciation information derived based on the acquired effective information with the pronunciation information of the training performance information, and updates the parameters of the effective string vibration determination model so that they match. For example, when the effective string vibration determination model is realized by a neural network, the effective string vibration determination model training unit 220 may update the parameters of the neural network according to the comparison result according to the well-known error backpropagation method. In addition, when pronunciation information is output even though there is no pronunciation information in the training performance information, the pronunciation information is regarded as unnecessary string vibration.

[0076] The effective string vibration determination model training unit 220 repeats the above-described process until a predetermined end condition is satisfied, trains the effective string vibration determination model, and when the predetermined end condition is satisfied, transfers the effective string vibration determination model at that time to the performance information prediction device 300 as a trained effective string vibration determination model. Here, the predetermined end condition may be that all the prepared training data has been processed. [Effective String Vibration Determination Model Training Process] Next, with reference to FIG. 15, the effective string vibration determination model training process according to an embodiment of the present disclosure will be described. The effective string vibration determination model training process is realized by the above-described effective string vibration determination model training device 200, and may be realized, for example, when a processor executes a program or an instruction. FIG. 15 is a flowchart showing the effective string vibration determination model training process according to an embodiment of the present disclosure.

[0077] As shown in FIG. 15, in step S201, the effective string vibration determination model training device 200 selects training performance information from the training performance information database 80. Specifically, the effective string vibration determination model training device 200 may automatically select the training performance information randomly, sequentially, or by user selection.

[0078] In step S202, the effective string vibration determination model training device 200 converts the performance information into display information of a TAB score.

[0079] In step S203, the effective string vibration determination model training device 200 displays the TAB score on the LCD of the control device 100 or the like.

[0080] In step S204, the effective string vibration determination model training device 200 starts the MIDI player according to the tempo of the performance information.

[0081] In step S205, the effective string vibration determination model training device 200 starts the metronome according to the tempo.

[0082] In step S206, the MIDI player reproduces the performance information.

[0083] In step S207, the metronome reproduces the performance information. Thereby, the preparation for acquiring the performance of the performer is completed, and the performer starts the performance.

[0084] In step S208, the effective string vibration determination model training device 200 stores the sound generation information or sound cancellation information of the s channel generated from the MIDI player in the memory p.

[0085] In step S209, the effective string vibration determination model training device 200 generates a volume pitch characterization data frame from the string vibration waveform representing the performance by the performer, inputs the effective information indicating the validity or invalidity of each string at the reference time into the effective string vibration determination model to be trained, generates sound generation information or sound cancellation information based on the determined effective information, and stores it in the memory o.

[0086] In step S210, the effective string vibration determination model training device 200 compares the sound generation information or sound cancellation information in the memory p with the sound generation information or sound cancellation information of the effective string vibration determination model in the memory o.

[0087] In step S211, the effective string vibration determination model training device 200 determines whether there is a difference between the sound generation information or sound cancellation information in the memory p and the sound generation information or sound cancellation information in the memory o.

[0088] If there is a significant difference (S211: Yes), the effective string vibration determination model training device 200 applies, in step S212, optimization information for updating the effective string vibration determination model from the difference to the effective string vibration determination model and proceeds to step S213. On the other hand, if there is no significant difference (S211: No), the effective string vibration determination model training device 200 proceeds to step S213 without updating the effective string vibration determination model.

[0089] In step S213, the effective string vibration determination model training device 200 determines whether the performance has ended. If not (S213: No), it returns to step S206.

[0090] In step S214, the effective string vibration determination model training device 200 stops the metronome and the MIDI player.

[0091] In step S215, the effective string vibration determination model training device 200 determines whether there has been an end operation by the user or the like. If there is no end operation (S215: No), it returns to step S201 to select the next performance information. If there is an end operation (S215: Yes), the process ends.

[0092] In the above-described embodiment, a performance information prediction system that trains an effective string vibration determination model for predicting performance information from string vibration waveform data of a stringed instrument such as a guitar 50 and predicts performance information using the trained effective string vibration determination model has been described. However, the present disclosure is not limited to this and may be applied to wind instruments. That is, the present disclosure may be applied to a performance information prediction system that trains an effective string vibration determination model for predicting performance information from air vibration waveform data of a wind instrument and predicts performance information using the trained effective string vibration determination model.

[0093] Although the embodiments of the present invention have been described in detail above, the present invention is not limited to the specific embodiments described above, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims.

[0094] The invention described in the original claims of the present application is appended below. [Appended Note] In one aspect of the present disclosure, a preprocessing unit that generates volume-pitch characterization data including volume information and pitch information for each string from string vibration waveform data representing string instrument performance; a performance information prediction unit that predicts performance information of the string instrument performance from the volume-pitch characterization data using a trained valid string vibration determination model; a performance information prediction device having the above is provided.

[0095] In one embodiment, the performance information may be composed of pronunciation information, muting information, and pitch change information.

[0096] In one embodiment, the trained valid string vibration determination model may output whether the vibration of each string is valid or invalid.

[0097] In one embodiment, the trained valid string vibration determination model may be realized by a neural network.

[0098] In one embodiment, the performance information may be described according to the MIDI protocol.

[0099] In another aspect of the present disclosure, a preprocessing unit that generates volume-pitch characterization data including volume information and pitch information for each string from string vibration waveform data representing string instrument performance performed according to training performance information; a valid string vibration determination model training unit that trains a valid string vibration determination model that determines the validity of the vibration of each string from the volume-pitch characterization data using the training performance information; a valid string vibration determination model training device having the above is provided.

[0100] In another aspect of the present disclosure, an electronic string instrument; the performance information prediction device described above; A performance information generation system including the above-described effective string vibration determination model training device is provided.

[0101] In another aspect of the present disclosure, one or more processors generate volume pitch characterization data including volume information and pitch information of each string from string vibration waveform data representing a string instrument performance, and the one or more processors predict performance information of the string instrument performance from the volume pitch characterization data using a trained effective string vibration determination model. A performance information prediction method including these steps is provided.

[0102] In another aspect of the present disclosure, one or more processors generate volume pitch characterization data including volume information and pitch information of each string from string vibration waveform data representing a string instrument performance performed according to training performance information, and the one or more processors train an effective string vibration determination model that determines the effectiveness of vibration of each string from the volume pitch characterization data using the training performance information. An effective string vibration determination model training method including these steps is provided.

Description of Reference Numerals

[0103] 10 Guitar controller 50 Guitar 100 Control device 200 Effective string vibration determination model training device 210 Preprocessing unit 220 Effective string vibration determination model training unit 300 Performance information prediction device 310 Preprocessing unit 320 Performance information prediction unit

Claims

1. A preprocessing unit that generates volume pitch characterization data including volume information and pitch information for each string from string vibration waveform data representing string instrument performance, A performance information prediction device having a performance information prediction unit that predicts performance information of the string instrument performance from the volume pitch characterization data using a trained effective string vibration determination model.

2. The performance information prediction device according to claim 1, wherein the performance information is composed of pronunciation information, muting information, and pitch change information.

3. The performance information prediction device according to claim 1 or 2, wherein the trained effective string vibration determination model outputs whether the vibration of each string is effective or ineffective.

4. The performance information prediction device according to any one of claims 1 to 3, wherein the trained effective string vibration determination model is realized by a neural network.

5. The performance information prediction device according to any one of claims 1 to 4, wherein the performance information is described according to the MIDI protocol.

6. A preprocessing unit that generates volume pitch characterization data including volume information and pitch information for each string from string vibration waveform data representing string instrument performance played according to training performance information, An effective string vibration determination model training device having an effective string vibration determination model training unit that trains an effective string vibration determination model that determines the effectiveness of the vibration of each string from the volume pitch characterization data using the training performance information.

7. An electronic string instrument, The performance information prediction device according to any one of claims 1 to 5, And an effective string vibration determination model training device according to claim 6, a performance information generation system.

8. One or more processors generate volume pitch characterization data including volume information and pitch information for each string from string vibration waveform data representing string instrument performance, A performance information prediction method, wherein the one or more processors predict performance information of the string instrument performance from the volume pitch characterization data using a trained effective string vibration determination model.

9. One or more processors generate volume pitch characterization data including volume information and pitch information for each string from string vibration waveform data representing string instrument performance played according to training performance information, An effective string vibration determination model training method, wherein the one or more processors train an effective string vibration determination model that determines the effectiveness of the vibration of each string from the volume pitch characterization data using the training performance information.

Citation Information

Patent Citations

  • Playing information detector

    JP1997244637A

  • Playing information generator

    JP1997288482A

  • Trigger detecting device and method thereof

    JP2000105589A

  • Electronic stringed instrument

    JP2005189369A

  • Note sensing in M.I.D.I. guitars and the like

    US5033353A