Pronunciation control system, pronunciation control method, and program
The sound control system synchronizes sound production across devices by predicting timing based on network delay, ensuring consistent ensemble performance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- YAMAHA CORP
- Filing Date
- 2025-10-15
- Publication Date
- 2026-04-28
AI Technical Summary
The variation in network communication delay affects the ensemble performance feeling among musicians, leading to inconsistent musical experiences.
A sound control system that includes a first acquisition unit to measure network delay and a generation unit to predict sound production timing based on operation information, generating a control signal to synchronize sound production across connected devices.
Minimizes changes in the musical sensation experienced by performers despite fluctuations in communication environment.
Smart Images

Figure 2026071184000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for controlling pronunciation.
Background Art
[0002] Keyboard instruments such as electronic pianos generate a control signal including pronunciation intensity, pronunciation timing, etc., using a measurement signal obtained by measuring the movement of keys or the like. The sound source device generates a sound signal by receiving the control signal. The measurement signal is also being considered for other uses. For example, a technique has been proposed in which, instead of generating a control signal in a keyboard instrument that generated the measurement signal, the measurement signal is transmitted to another communication base and the control signal is generated on the receiving side. This technique is disclosed in, for example, Patent Document 1. By this technique, it is possible to reduce the problem of delay that occurs in ensemble performances with performers via a network.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The delay time in network communication varies greatly depending on the communication environment. Therefore, the feeling when performers perform in an ensemble changes depending on the difference in the communication environment.
[0005] One object of the present invention is to reduce the change in the feeling when performers perform in an ensemble even when the communication environment changes.
Means for Solving the Problems
[0006] A sound control system in one embodiment includes a first acquisition unit, a second acquisition unit, and a generation unit. The first acquisition unit acquires a delay time, including network delay, between a first keyboard device that includes a first key and causes a first sound-producing device to produce sound in accordance with the movement of the first key, and a second sound-producing device connected via a network. The second acquisition unit acquires operation information corresponding to the pressing of the first key in the first keyboard device. The generation unit generates a first control signal for causing the second sound-producing device to produce sound based on the delay time and the predicted timing of sound production in the first sound-producing device calculated using the operation information. [Effects of the Invention]
[0007] According to the present invention, even if the communication environment fluctuates, the change in the sensation experienced by musicians when playing together can be minimized. [Brief explanation of the drawing]
[0008] [Figure 1] This is a diagram illustrating the communication system configuration in the first embodiment. [Figure 2] This is a diagram illustrating the internal configuration of an automatic playing piano in the first embodiment. [Figure 3] This is a diagram illustrating the configuration of the control unit in the first embodiment. [Figure 4] This is a diagram illustrating the configuration of the sound generation control system in the first embodiment. [Figure 5] This is a diagram illustrating the method of predictive calculation in the first embodiment. [Figure 6] This is a diagram illustrating the configuration of the transmission function in the first embodiment. [Figure 7] This is a diagram illustrating the configuration of the receiving function in the first embodiment. [Figure 8] This diagram illustrates the operation of the automatic playing piano 1 at communication base C1 and communication base C2 in the first embodiment. [Figure 9] This is a flowchart illustrating the sound control method in the first embodiment. [Figure 10]This is a diagram for explaining the operation of the automatic-playing piano 1 at communication base C1 and communication base C2 in the second embodiment. [Figure 11] This is a diagram for explaining the configuration of the transmission function in the third embodiment. [Figure 12] This is a diagram for explaining the configuration of the reception function in the third embodiment. [Figure 13] This is a diagram for explaining the operation of the automatic-playing piano 1 at communication base C1 and communication base C2 in the third embodiment. [Figure 14] This is a diagram for explaining the configuration of the reception function in the fourth embodiment. [Figure 15] This is a diagram for explaining the configuration of the reception function in the fifth embodiment. [Figure 16] This is a diagram for explaining the operation of the automatic-playing piano 1 at communication base C1 and communication base C2 in the sixth embodiment. [Figure 17] The key trajectory and hammer trajectory in the proportionality of the seventh embodiment. [Figure 18] The key trajectory and hammer trajectory in the seventh embodiment. [Figure 19] This is an explanatory diagram of the estimation model in the seventh embodiment. [Figure 20] This is an explanatory diagram of the operation of the control unit in the eighth embodiment. [Figure 21] This is an explanatory diagram of the operation of the control unit in a modification of the eighth embodiment. [Figure 22] This is a block diagram exemplifying the functional configuration of the control unit in the ninth embodiment. [Figure 23] This is a flowchart of the conversion table generation process.
Modes for Carrying Out the Invention
[0009] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. The following embodiments are examples, and the present invention is not construed as being limited to these embodiments. The configurations described in each embodiment can also be applied to other embodiments. In the drawings referred to in the following plurality of embodiments, the same parts or parts having the same function are denoted by the same reference numerals or similar reference numerals (reference numerals with A, B, etc. attached after the numbers), and the repeated description may be omitted. The drawings may be schematically described with some parts of the configuration omitted for clarity of explanation.
[0010] <First Embodiment> [Communication System] FIG. 1 is a diagram for explaining the configuration of a communication system in the first embodiment. The communication system includes a server 1000 connected to a network NW such as the Internet. The server 1000 includes a control unit such as a CPU, a storage unit, and a communication unit. The control unit provides a service for realizing an ensemble between communication bases by executing a predetermined program.
[0011] The server 1000 controls the communication between a plurality of communication bases connected to the network NW, and executes the processing necessary for the automatic-playing pianos 1 at each communication base to realize P2P-type communication with each other. This processing may be realized by a known method. In the following description, when the communication bases C1 and C2 are not distinguished, they may simply be referred to as communication bases.
[0012] In this example, between the communication base C1 and the communication base C2, information related to the performance at each communication base is exchanged with each other by P2P communication. By this communication, an ensemble between a plurality of communication bases is realized. At each communication base, an automatic-playing piano 1 is arranged. The automatic-playing piano 1 includes a keyboard instrument 10, a control unit 20, a sensor 30, and a driving device 40.
[0013] This document provides a detailed explanation of a system that enables ensemble performances with minimal impact from network latency between communication points. In this example, ensemble performance is not limited to coordinating bidirectional performances; it is sufficient if any communication point can synchronize its performance with that of another communication point. For example, if communication point C2 reproduces the performance at communication point C1 using the automatic piano 1, the performer at communication point C2 can play along with the reproduced performance, thus constituting an ensemble performance. The performer at communication point C2 is not limited to playing the automatic piano 1; they may also play another instrument or sing.
[0014] [Automatic playing piano] Next, I will explain the configuration of the automatic playing piano 1.
[0015] Figure 2 is a diagram illustrating the internal configuration of the automatic playing piano 1 in the first embodiment. The keyboard instrument 10 is an example of a keyboard device, corresponding to, for example, a grand piano. The keyboard instrument 10 includes a plurality of keys 12. The keyboard instrument 10 includes hammers 14, strings 15, and dampers 18 provided corresponding to each key 12. The keyboard instrument 10 includes a plurality of pedals 13. The plurality of pedals 13 are, for example, a damper pedal, a shift pedal, and a sostenuto pedal. The keyboard instrument 10 further includes a soundboard 17, etc., through which the vibrations of the strings 15 are transmitted via a bridge 16.
[0016] In Figure 2, the configurations provided for each key 12 and pedal 13 are shown focusing on the configuration provided for one key 12 and pedal 13. Therefore, the configurations provided for other keys 12 and other pedals 13 are omitted from the description.
[0017] Sensor 30 includes a key sensor 32, a pedal sensor 33, and a hammer sensor 34. The key sensor 32 is provided for each key 12 and outputs a measurement signal to the control unit 20 according to the behavior of the key 12. In this example, the key sensor 32 outputs a measurement signal to the control unit 20 according to the amount the key 12 is pressed (sometimes called the pressed position of the key 12). The pressed position of the key 12 may be measured as a continuous amount (with fine resolution), or it may be measured by detecting when the key 12 has passed a predetermined pressed position. The pressed positions in which the key 12 is detected may be any of several positions within the range from the rest position to the end position.
[0018] The hammer sensor 34 is provided for each hammer 14 and outputs a measurement signal to the control unit 20 according to the behavior of the hammer 14. In this example, the hammer sensor 34 measures the position (rotation amount) of the hammer shank immediately before the hammer 14 strikes the string 15 and outputs a measurement signal to the control unit 20 according to the measurement result. The position of the hammer shank may be measured as a continuous quantity (fine resolution), or it may be measured by detecting that the hammer shank has passed a predetermined position. The positions in which the hammer shank is detected may be any of several positions within the range immediately before the hammer 14 strikes the string 15.
[0019] Each pedal sensor 33 is provided for each pedal 13 and outputs a measurement signal to the control unit 20 according to the behavior of the pedal 13. In this example, the pedal sensor 33 outputs a measurement signal to the control unit 20 according to the position (amount of depression) of the pedal 13. The position of the pedal 13 may be detected as a continuous amount (fine resolution), or it may be detected when the pedal 13 passes a predetermined position. The positions in which the pedal 13 is detected may be any of several positions within the depression range of the pedal 13 (the range from the rest position to the end position).
[0020] The drive unit 40 includes a key drive unit 42, a pedal drive unit 43, a stopper 44, an exciter 47, and a damper drive unit 48. The key drive unit 42 is provided in correspondence with each key 12 and drives the key 12 to be pressed based on control from the control unit 20. This mechanically reproduces the same situation as when a performer presses the key 12. The pedal drive unit 43 is provided in correspondence with each pedal 13 and drives the pedal 13 to be pressed based on control from the control unit 20. This mechanically reproduces the same situation as when a performer presses the pedal 13. The damper drive unit 48 is provided in correspondence with each damper 18 and drives the damper 18 to be released from the string 15 based on control from the control unit 20. The damper drive unit 48 may have a configuration that drives all dampers 18 simultaneously.
[0021] The stopper 44 is driven based on control from the control unit 20 to either a position where it collides with the hammer shank (blocking position) or a position where it does not collide with the hammer shank (retracted position). When the stopper 44 is in the blocking position, even if the key 12 is pressed, the movement of the hammer shank is restricted and the hammer 14 does not strike the string 15. When the stopper 44 is in the retracted position, when the key 12 is pressed, the hammer 14, which is linked to the key 12, strikes the string 15. The keyboard instrument 10 produces sound when the string 15 is struck. When the sound production control function described below is implemented, the stopper 44 is controlled to be in the retracted position. The sound production control function may be implemented when the stopper 44 is controlled to be in the blocking position and sound production using the sound source device 25 is realized.
[0022] In this example, the vibrator 47 is supported by a support connected to a straight column 19 so as to contact the opposite side of the soundboard 17 from the part where the bridge 16 is placed. The vibrator 47 vibrates the soundboard 17 based on control from the control unit 20. For example, when the control unit 20 supplies a drive signal including a piano sound, the vibrator 47 applies vibrations to the soundboard 17 in accordance with the drive signal. This causes the piano sound to be emitted from the soundboard 17. Multiple vibrators 47 may be arranged to contact the soundboard 17. Instead of the vibrators 47 that vibrate the soundboard 17, speakers that emit sound may be used.
[0023] Sound production by the keyboard instrument 10 includes cases where a hammer 14 strikes a string 15, which is a sound-producing element, and cases where an exciter 47 vibrates a sound-producing element, which is a sound-producing element, 17. Therefore, it can also be said that the keyboard instrument 10 includes a sound-producing device that generates a string-striking sound by driving a key 12, and a sound-producing device that generates sound from the sound-producing element 17 by driving an exciter 47. Since the keyboard instrument 10 has multiple keys 12, it can also be said that it includes a keyboard device. In other words, in this example, the automatic playing piano 1 includes a keyboard device and a sound-producing device.
[0024] The configuration of the control unit 20 will now be described. In this example, the control unit 20 is attached to the keyboard instrument 10.
[0025] Figure 3 illustrates the configuration of the control unit 20 in the first embodiment. The control unit 20 includes a control device 21, a storage device 22, an operating device 23, a communication device 24, a sound source device 25, and an interface 26. These components are connected via a bus 27.
[0026] The control device 21 is an example of a computer equipped with a processor such as a CPU and a memory device such as RAM. The control device 21 executes programs stored in the memory device 22 using the CPU (processor) and implements functions for performing various processes in the control unit 20. The functions implemented in the control unit 20 include the sound control function, which will be described later. This sound control function controls each part of the control unit 20 and each component connected to the interface 26 in various ways. As will be described later, the sound control function also supports transmission and reception functions.
[0027] The storage device 22 is a device such as a non-volatile memory or a hard disk drive. The storage device 22 includes a program storage area 22a. The program storage area 22a stores the program executed by the control device 21 and various data necessary when executing this program.
[0028] The operating device 23 has operation buttons, etc., that receive user input. When user input is received via these operation buttons, an operation signal corresponding to the input is output to the control device 21. The operating device 23 may also have a display screen. In this case, the operating device 23 may be a touch panel in which a touch sensor is combined with the display screen.
[0029] The communication device 24 is a communication module that communicates with other devices wirelessly or via wired connection. It is a communication module that communicates with other devices such as servers and smartphones. In this example, the other devices with which the communication device 24 communicates are the server 1000 and the automatic playing piano 1 at another communication point. In this example, the data communicated between communication points includes data corresponding to the performance of the keyboard instrument 10.
[0030] The sound source device 25 generates sound signals under control from the control device 21. These sound signals are used as drive signals for driving the vibrator 47, etc. In this example, the sound signals include signals indicating piano sounds. The sound source device 25 is controlled by sound generation control signals described in MIDI format, such as note on, note off, note number, and velocity. The signals for generating sound signals in the sound source device 25 are generated in the sound generation control function described later.
[0031] The sound source device 25 may be controlled by the control device 21 to generate sound signals indicating piano sounds corresponding to the performance content. The performance content is indicated by operation information. The operation information may be information generated based on measurement signals generated by the sensor 30, or it may be information directly indicated by the measurement signals. In this example, the operation information includes information indicating the pressed key 12 (e.g., note number, pitch information, etc.) and the position where the key 12 is pressed.
[0032] Interface 26 is an interface that connects the control unit 20 to each external component. As described above, each component connected to interface 26 in this example includes the sensor 30 and the drive unit 40. Interface 26 outputs the signal output from sensor 30 to the control device 21. Interface 26 outputs a signal (hereinafter sometimes referred to as a drive signal) to the drive unit 40 for driving each component. The drive signal is generated in the sound generation control function described later. The sound source device 25 described above may be configured to be connected to the control unit 20 via interface 26. Interface 26 may include a headphone jack or the like to which the sound signal generated by the sound source device 25 is supplied.
[0033] [Pronunciation control system] Figure 4 is a diagram illustrating the configuration of the sound production control system in the first embodiment. The sound production control system has a sound production control function 500. The sound production control function 500 has the function of realizing the performance content of the automatic piano 1 at the transmitting communication site on the automatic piano 1 at the receiving communication site. In this example, communication site C1 is the transmitting site and communication site C2 is the receiving site. The sound production control function 500 may be realized by a control device 21 at communication site C1 or a control device 21 at communication site C2, or it may be realized by the cooperation of the control device 21 at communication site C1 and the control device 21 at communication site C2. Therefore, the sound production control system including the sound production control function 500 includes the case where it is realized by a control device 21 at either communication site.
[0034] The keyboard device 581 includes the keys 12 and key sensors 32 of the automatic playing piano 1 at the communication base C1. The sound-producing device 583 includes a configuration for producing sound in response to the movement of the keys 12 of the automatic playing piano 1 at the communication base C1. The sound-producing device 583 may include, for example, strings 15 or sound source device 25.
[0035] The keyboard device 591 includes the keys 12 and the key drive device 42 of the automatic playing piano 1 at the communication base C2. Therefore, the keyboard device 581 and the keyboard device 591 are connected via a network NW. The sound-producing device 593 includes a configuration for producing sound in response to the movement of the keys 12 of the automatic playing piano 1 at the communication base C2. The sound-producing device 593 may include, for example, strings 15 or sound source device 25.
[0036] The sound generation control function 500 includes a delay acquisition unit 501, an operation information acquisition unit 511, and a control signal generation unit 513.
[0037] The delay acquisition unit 501 acquires the delay time due to network delay between the automatic piano 1 at communication base C1 and the automatic piano 1 at communication base C2. The delay time can be acquired by any known method. For example, the delay acquisition unit 501 performs time synchronization using NTP (Network Time Protocol) at both the control units 20 at communication bases C1 and C2. Then, the delay acquisition unit 501 measures the time from the transmission time at communication base C1 to the reception time at communication base C2 as the time corresponding to the network delay (hereinafter referred to as delay time TD2).
[0038] In this example, the delay time TD2 is measured when communication is established between communication point C1 and communication point C2. Subsequent measurements may be performed periodically, when a predetermined period of time has passed without key 12 being pressed, or when instructed to be performed by the user.
[0039] The delay acquisition unit 501 acquires the total delay time (hereinafter referred to as delay time TD) by adding a predetermined time to the delay time TD2. The time added corresponds to the time that occurs due to factors other than network delay, and in this example, delay times TD1, TD3, and TD4 are assumed.
[0040] The delay time TD1 is the processing time required by the control unit 20 at communication base C1, and corresponds to, for example, the time from when the control unit 20 receives the measurement signal until it transmits the information obtained based on the measurement signal (operation information or control signal, described later) to communication base C2. The delay time TD3 is the processing time required by the control unit 20 at communication base C2, and corresponds to, for example, the time from when the control unit 20 receives information from communication base C1 until it generates a drive signal. The delay time TD3 also includes the receive buffer. The delay time TD4 corresponds to the time from when the key 12 in the automatic playing piano 1 starts to be driven in response to the drive signal until it produces sound.
[0041] The operation information acquisition unit 511 acquires operation information corresponding to the pressing of the keys 12 in the keyboard device 581. As described above, the operation information is information indicating the pressed key 12 and the pressing position of the key 12, as indicated by measurement signals at each point in time. Therefore, the behavior of the key 12 is determined by sequentially acquiring the operation information at a predetermined sampling period.
[0042] The control signal generation unit 513 generates a control signal based on the delay time acquired by the delay acquisition unit 501 and the predicted timing of sound production in the sound production device 583. The control signal is a signal to cause the sound production device 593 to produce sound. In this example, the sound production device 593 produces sound when the key 12 in the keyboard device 591 is driven. That is, the control signal is a signal to drive the key 12 in the keyboard device 591.
[0043] The predicted timing is calculated by a predetermined prediction calculation. In this prediction calculation, delay time and operation information are used as calculation parameters. The method of the prediction calculation will now be explained.
[0044] Figure 5 is a diagram illustrating the prediction calculation method in the first embodiment. Figure 5 shows the relationship between the elapsed time from the start of pressing the key 12 and the pressed position. At the pressed position, "Rp" corresponds to the rest position, "Ep" corresponds to the end position, and "Hp" corresponds to the position when the string 15 is sounded.
[0045] In the prediction calculation, the pressed position at the predicted time after a delay time TD from the current time is calculated based on the behavior of key 12 during the measurement period TP, i.e., the change in the pressed position. The measurement period TP is the period from the start time T0 to the current time. For example, if the current time is time T1, the measurement period TP is the period from time T0 to time T1, and the predicted time corresponds to time T2, which is time T1 plus the delay time TD. Therefore, in the prediction calculation, the pressed position at time T2 is predicted based on the change in the pressed position from time T0 to time T1.
[0046] As time progresses, the predicted button press position is calculated sequentially. As time passes, the measurement period TP lengthens, improving the accuracy of the predicted button press position at the predicted time. In the example shown in Figure 5, the current time is Tr, and the predicted button press position at predicted time Th is Hp. The predicted timing described above corresponds to predicted time Th. In other words, the prediction calculation uses the operation information during the measurement period TP to calculate the predicted timing (predicted time Th). If the press time from the rest position to the end position is the same, the longer the delay time TD, the shorter the measurement period TP becomes. If the delay time TD is short, the measurement period TP becomes longer, and since only the near future prediction is needed, the prediction accuracy improves.
[0047] In this example, the control signal generation unit 513 generates a control signal at the predicted timing (predicted time Th) to produce sound at the predicted time Th, which is Tr. This control signal includes, for example, information indicating a note-on of the pitch corresponding to the target key 12. The control signal generation unit 513 may also include velocity information in the control signal by performing a calculation to predict the speed of the key 12 at the predicted time Th. The control signal generated by the control signal generation unit 513 is used to drive the key 12 of the keyboard device 591. When the key 12 is pressed based on the control signal, the sound-producing device 593 produces sound at the predicted time Th.
[0048] Thus, when a key 12 of the keyboard device 581 is pressed, the sound generation control function 500 calculates the predicted timing of sound generation in the sound generation device 583 while the key 12 is pressed, and drives the key 12 of the keyboard device 591 so that the sound generation device 593 produces sound at this predicted timing. Therefore, according to the sound generation control function 500, even if there are various delay factors including network delay, it is possible to control the sound generation in the sound generation device 593 at communication base C2 to occur at the same timing as the sound generation in the sound generation device 583 caused by the performance operation at communication base C1. At this time, by measuring the network delay in advance, the sound generation control function 500 can control the sound generation in accordance with the fluctuations even if the delay time TD2 fluctuates.
[0049] Next, we will explain the functional configurations implemented in the automatic piano 1 at communication base C1 and the automatic piano 1 at communication base C2. The automatic piano 1 at communication base C1 implements a transmission function. The automatic piano 1 at communication base C2 implements a reception function.
[0050] Each component included in the above-described pronunciation control function 500 may be included together in either the transmission function or the reception function, or it may be distributed between both. In the first embodiment, an example in which the transmission function corresponds to the pronunciation control function 500 will be described.
[0051] This section describes the transmission function implemented by the control unit 20 when the control device 21 executes a program at communication base C1. The configuration for implementing the transmission function is not limited to implementation by program execution; at least some of the configuration may be implemented by hardware. Alternatively, the transmission function may be implemented not by the control unit 20, but by a device connected to interface 26 (for example, a computer on which this program is installed).
[0052] Figure 6 is a diagram illustrating the configuration of the transmission function in the first embodiment. The transmission function 100 includes a delay acquisition unit 101, an operation information generation unit 111, a control signal generation unit 113, and a control signal transmission unit 115.
[0053] The delay acquisition unit 101 corresponds to the delay acquisition unit 501 described above and acquires the delay time TD. The delay acquisition unit 101 may receive at least one of the delay times TD2, TD3, and TD4 from the automatic playing piano 1 at the communication base C2. The operation information generation unit 111 corresponds to the operation information acquisition unit 511 described above and acquires operation information corresponding to the pressing of the key 12 by generating operation information based on the measurement signal from the key sensor 32.
[0054] The control signal generation unit 113 corresponds to the control signal generation unit 513 described above and generates a control signal based on the delay time TD acquired by the delay acquisition unit 101 and the operation information generated by the operation information generation unit 111. Based on the operation information, the control signal generation unit 113 predicts the sound after the delay time TD and generates a control signal including a note-on from the predicted timing of the sound (time Th in the example shown in Figure 5) to the timing before the delay time TD (time Tr in the example shown in Figure 5). The control signal transmission unit 115 transmits the control signal generated by the control signal generation unit 113 to the control unit 20 at the communication base C2.
[0055] This section describes the receiving function implemented by the control unit 20 when the control device 21 executes a program at communication base C2. The configuration for implementing the receiving function is not limited to implementation by program execution; at least some of the configuration may be implemented by hardware. Alternatively, the receiving function may be implemented not by the control unit 20, but by a device connected to interface 26 (for example, a computer on which this program is installed).
[0056] Figure 7 is a diagram illustrating the configuration of the receiving function in the first embodiment. The receiving function 200 includes a control signal receiving unit 211 and a drive signal generation unit 215.
[0057] The control signal receiving unit 211 receives the control signal transmitted by the control signal transmitting unit 115. The drive signal generating unit 215 generates a drive signal based on the control signal received by the control signal receiving unit 211. The generated drive signal is supplied to the key driving device 42. This drive signal causes the key driving device 42 to drive the key 12, and when the key 12 is pressed, the string 15 sounds.
[0058] Figure 8 is a diagram illustrating the operation of the automatic playing piano 1 at communication bases C1 and C2 in the first embodiment. In Figure 8, the state at communication base C1 is shown when key 12 is pressed, for example, as shown in Figure 5.
[0059] At communication station C1, when key 12 is pressed and time Tr arrives, it is determined by prediction calculation that it will be sounded at time Th. After a delay time TD1 has elapsed since the control signal was generated at time Tr, the control signal is transmitted from communication station C1 to communication station C2. After another delay time TD2 has elapsed, the control signal is received at communication station C2. At communication station C2, at time Ts, after a delay time TD3 has elapsed since the control signal was received, the key 12 is started to be driven based on the control signal. Subsequently, at time Te, after a delay time TD4 has elapsed, the string 15 is sounded in conjunction with the pressing of key 12.
[0060] Time Tr is the timing before delay time TD compared to time Th. Delay time TD is the sum of delay times TD1 to TD4. Therefore, although there will be errors in the prediction calculation of time Th and discrepancies due to fluctuations after network delay measurement, the sound at communication point C1 and the sound at communication point C2 will be synchronized to occur almost simultaneously.
[0061] As described above, the sound control function 500 enables sound generation at communication base C2 to be synchronized with the timing of sound generation caused by pressing key 12 at communication base C1. At this time, by acquiring the delay time TD and using it in prediction calculations, the impact of fluctuations in the amount of delay due to the network environment can be suppressed.
[0062] Figure 9 is a flowchart illustrating the pronunciation control method in the first embodiment. The pronunciation control method is a method executed by the pronunciation control function 500.
[0063] The delay acquisition unit 501 acquires the delay time (step S501). When key 12 is pressed, the operation information acquisition unit 511 acquires operation information (step S511). The control signal generation unit 513 predicts the sound from the operation information and generates a control signal based on the prediction (step S513). Each time key 12 is pressed, the acquisition of operation information (step S511) and the generation of a control signal (step S513) are performed.
[0064] In the first embodiment, the transmission function 100 includes a function corresponding to the sound generation control function 500. Therefore, the sound generation control method is performed by the control device 21 that implements the transmission function 100 at the communication base C1.
[0065] <Second Embodiment> In the second embodiment, an example is described in which the control signal is transmitted as a signal including the press position predicted by predictive calculation. The sound control function 500 (transmission function 100 and reception function 200) is generally the same as in the first embodiment, but the control signal generated in the control signal generation unit 513 (control signal generation unit 113) is different from that in the first embodiment.
[0066] In the second embodiment, the control signal generation unit 513 predicts the pressing position after a delay time TD based on sequentially acquired operation information. The control signal generation unit 513 generates a control signal indicating the predicted pressing position. The keys 12 of the keyboard device 591 are driven according to this control signal. That is, the keys 12 are driven to the pressing position indicated by the control signal.
[0067] Figure 10 is a diagram illustrating the operation of the automatic playing piano 1 at communication bases C1 and C2 in the second embodiment. In Figure 10, the state at communication base C1 is shown when key 12 is pressed, for example, as shown in Figure 5.
[0068] At communication point C1, when key 12 is pressed (time T0), a control signal indicating that the pressing has begun is sent to communication point C2. After delay times TD1 and TD2 have elapsed from time T0, the control signal is received at communication point C2. The control signal sent at this point does not include the predicted pressing position, but it is sent to initiate the operation of key 12. At time Ts1, after a further delay time TD3 has elapsed, the operation of key 12 is initiated at communication point C2 based on the control signal.
[0069] At communication base C1, as shown in Figure 5, at time T1, the pressing position at time T2, after a delay time TD from time T1, is predicted by a predictive calculation. At communication base C2, a control signal indicating the predicted pressing position is received at time Ts2, and key 12 is controlled to that pressed position. Control signals are transmitted sequentially from communication base C1 to communication base C2. At communication base C2, the pressing position of key 12 is controlled according to the received control signal. The time at communication base C2 when key 12 is actually controlled to this pressed position is when a delay time TD3 has elapsed from time Ts2.
[0070] At time Tr, a prediction calculation determines that the sound will be produced at time Th at communication point C1. The control signal generated at time Tr is received at communication point C2. This control signal indicates the key press position (corresponding to Hp in Figure 5) at the time of sound production. Therefore, at communication point C2, the key press position of key 12 is controlled so that the sound is produced at time Te, after a delay time TD3 has elapsed since the control signal was received. Time Te and time Th are approximately the same. In other words, key 12 at communication point C2 is controlled to the press position Hp at time Th. Even in this way, the sound production at communication point C1 and communication point C2 are synchronized so that they occur almost simultaneously.
[0071] In the first embodiment, the delay time TD4 is taken into consideration. On the other hand, in the second embodiment, the pressed position of key 12 at communication site C2 is synchronized with the pressed position of key 12 at communication site C1. Therefore, since key 12 is already activated, the above-mentioned delay time TD4 does not need to be taken into consideration. Accordingly, in the second embodiment, the delay time TD is the time obtained by adding the delay times TD1 and TD3 to the delay time TD2.
[0072] The control signal transmitted from communication base C1 to communication base C2 may have a timestamp attached. In this case, the key press position at communication base C2 is controlled according to the time indicated by the timestamp. In this way, even if the timing of receiving the control signal is shifted due to jitter or the like, the key press position of key 12 can be controlled according to the timestamp, as long as it is within the range of the receive buffer (delay time TD3).
[0073] <Third Embodiment> In the third embodiment, an example is described in which the receiving function implemented in the control unit 20 at communication base C2 corresponds to the sound production control function 500. The automatic playing piano 1 at communication base C1 implements a transmission function. The automatic playing piano 1 at communication base C2 implements a receiving function. This is the same as in the first embodiment.
[0074] Figure 11 is a diagram illustrating the configuration of the transmission function in the third embodiment. The transmission function 300 includes an operation information generation unit 311 and an operation information transmission unit 315. The operation information generation unit 311 generates operation information based on the measurement signal from the key sensor 32. This operation information is time-stamped. The operation information transmission unit 315 transmits the operation information generated by the operation information generation unit 311 to the control unit 20 of the communication base C2.
[0075] Figure 12 is a diagram illustrating the configuration of the receiving function in the third embodiment. The receiving function 400 includes a delay acquisition unit 401, an operation information receiving unit 411, a control signal generation unit 413, and a drive signal generation unit 415.
[0076] The delay acquisition unit 401 corresponds to the delay acquisition unit 501 described above and acquires the delay time TD. The operation information receiving unit 411 corresponds to the operation information acquisition unit 511 described above and acquires operation information by receiving it from the communication base C1.
[0077] The control signal generation unit 413 corresponds to the control signal generation unit 513 described above and generates a control signal based on the delay time TD acquired by the delay acquisition unit 401 and the operation information received by the operation information receiving unit 411. Based on the operation information, the control signal generation unit 413 predicts the sound after the delay time TD, using the time in the timestamp as a reference. The control signal generation unit 413 identifies the predicted timing by receiving operation information from the predicted timing of sound (Th in the example shown in Figure 5) to the timing before the delay time TD (Tr in the example shown in Figure 5), and generates a control signal including note-on.
[0078] The drive signal generation unit 415 generates a drive signal based on the control signal generated by the control signal generation unit 413. The generated drive signal is supplied to the key drive device 42. This drive signal causes the key drive device 42 to drive the key 12, and when the key 12 is pressed, the string 15 sounds.
[0079] Figure 13 is a diagram illustrating the operation of the automatic playing piano 1 at communication bases C1 and C2 in the third embodiment. In Figure 13, the state when key 12 is pressed at communication base C1, for example, as shown in Figure 5.
[0080] When key 12 is pressed at communication point C1 (time T0), operation information indicating that the key has been pressed is sent to communication point C2. This operation information is timestamped to indicate time T0. After delay times TD1 and TD2 have elapsed from time T0, the operation information is received at communication point C2. After further delay time TD3 has elapsed, at time Tb, prediction calculations based on the operation information are started at communication point C2.
[0081] Operation information is transmitted sequentially from communication base C1 to communication base C2. Based on the sequentially received operation information, a predictive calculation is performed to predict the button press position after a delay time TD has elapsed, using the time of the timestamp attached to the operation information as a reference.
[0082] Operation information transmitted from communication base C1 at time Tr is received at communication base C2 after delay times TD1 and TD2 have elapsed. At time Ts, after a further delay time TD3 has elapsed, the control signal generation unit 413 performs a prediction calculation using the operation information at time Tr, indicated by the time time Tr. As a result, the control signal generation unit 413 identifies time Th after delay time TD as the predicted timing for sound generation from time Tr indicated by the time time, and generates a control signal including note-on.
[0083] In this example, the measurement period TP when the predicted timing is identified corresponds to the period from time Tb to time Ts, but if the timestamp is used as the basis, it corresponds to the period from time T0 to time Tr, as in the first embodiment.
[0084] At time Ts, when the control signal generation unit 413 generates a control signal including a note-on signal, the key 12 is started to be driven at communication station C2 based on the control signal. Then, at time Te, after the delay time TD4 has elapsed, the string 15 is sounded in conjunction with the pressing of the key 12. Even in this manner, the sounding at communication station C1 and the sounding at communication station C2 are synchronized to occur almost simultaneously.
[0085] <Fourth and fifth embodiments> Sound generation at communication base C2 is not limited to being achieved by the string 15; it may also be achieved by the sound source device 25 generating an audio signal. Similar to the first embodiment, in the sound generation control function 500 shown in Figure 4, the control signal generated by the control signal generation unit 513 is a signal to cause the sound generation device 593 to produce sound. In this example, the sound generation device 593 produces sound when the sound source device 25 generates an audio signal and that audio signal is output. That is, the control signal is a signal to cause the sound source device 25 in the sound generation device 593 to generate an audio signal.
[0086] In this example, since the key 12 does not need to be driven, the delay time TD does not need to include the portion corresponding to the delay time TD4. The delay time TD4 may be set as the time required for the sound source device 25 to generate the sound signal, but it will be an extremely short time.
[0087] In the fourth embodiment, an example is described in which the configuration for sound generation in the receiving function 200 of the first embodiment is applied when the sound source device 25 generates an audio signal.
[0088] Figure 14 is a diagram illustrating the configuration of the receiving function in the fourth embodiment. The receiving function 200A in the fourth embodiment has a configuration in which the drive signal generation unit 215 of the receiving function 200 in the first embodiment is replaced with a sound signal generation unit 217 and a sound signal output unit 219. The sound signal generation unit 217 and the sound signal output unit 219 are included in the sound generation device 593 described above.
[0089] The sound signal generation unit 217 generates a sound signal based on a control signal. The sound source device 25 is used to generate the sound signal. Therefore, the control signal can also be said to be a signal that causes the sound source device 25 to generate a sound signal. The sound signal output unit 219 outputs the generated sound signal. In this example, the exciter 47 is used to output the sound signal. The sound signal output unit 219 may be a terminal that outputs the sound signal as an electrical signal, or it may be a speaker that outputs the sound signal as air vibrations.
[0090] In the fifth embodiment, an example is described in which the configuration for sound generation in the receiving function 400 of the third embodiment is applied when the sound source device 25 generates an audio signal.
[0091] Figure 15 is a diagram illustrating the configuration of the receiving function in the fifth embodiment. The receiving function 400A in the fifth embodiment has a configuration in which the drive signal generation unit 415 of the receiving function 400 in the third embodiment is replaced with a sound signal generation unit 417 and a sound signal output unit 419. The sound signal generation unit 417 and the sound signal output unit 419 are included in the sound generation device 593 described above. The sound signal generation unit 417 has the same configuration as the sound signal generation unit 217 described above. The sound signal output unit 419 has the same configuration as the sound signal output unit 419 described above. Therefore, a detailed explanation is omitted.
[0092] <Sixth Embodiment> In the sixth embodiment, the prediction calculation calculates the button press position at the predicted time not after a delay time TD from the current time, but after a reserve time TL. That is, compared to the first embodiment, the button press position at the predicted time is calculated when the delay time TD + reserve time TL has elapsed from the current time.
[0093] Figure 16 is a diagram illustrating the operation of the automatic playing piano 1 at communication stations C1 and C2 in the sixth embodiment. Compared to the example described in Figure 8 in the first embodiment, time Th is later than time Te. That is, the sound is produced at communication station C1 later than at communication station C2. In other words, the control signal transmitted from communication station C1 to communication station C2 includes a signal to cause the piano to be produced at communication station C2 at a timing (time Te) earlier than the predicted timing (time Th).
[0094] In this way, when an ensemble plays along with the sound produced by the automatic piano 1 at communication base C2, the amount of delay felt by the performer at communication base C1 when data representing the sound of the ensemble is transmitted from communication base C2 and received at communication base C1 can be reduced. The reduction in delay time corresponds to the time from time Te to time Th, i.e., the buffer time TL.
[0095] For example, data obtained by recording a musical performance sound at communication point C2 in sync with the sound produced at time Te is received at communication point C1, and that musical performance sound is output at communication point C1. In this case, the discrepancy between the sound produced by pressing key 12 at time Th at communication point C1 and the musical performance sound reproduced from the received data can be reduced according to the reserve time TL. Increasing the reserve time TL will further reduce the amount of discrepancy.
[0096] <Seventh Embodiment> The seventh and eighth embodiments are designed to shorten the delay time (particularly the delay time TD1 at the transmitting communication base C1) from the performance of the keyboard instrument 10 by the performer at communication base C1 to the sound produced by the keyboard instrument 10 at communication base C2.
[0097] Figure 17 shows the key trajectory Qk and hammer trajectory Qh in proportion to the seventh embodiment. The key trajectory Qk and hammer trajectory Qh at the transmitting communication site C1 and the receiving communication site C2 are illustrated in Figure 17.
[0098] The key trajectory Qk is the time evolution of the position of key 12 in the keyboard instrument 10. The key trajectory Qk is expressed by a time series of the position of key 12 (e.g., the displacement of key 12 relative to the rest position). The hammer trajectory Qh is the time evolution of the position of hammer 14. The hammer trajectory Qh is expressed by a time series of the position of hammer 14 (e.g., the rotation of hammer 14 relative to its initial position).
[0099] Time point t1 in Figure 17 is the moment when the performer begins pressing the key. The hammer 14 rotates in conjunction with the displacement of the key 12 caused by the pressing, and at time point t2, the hammer 14 strikes the string 15. That is, the keyboard instrument 10 at communication station C1 produces a musical note at time point t2 corresponding to the pitch of the key 12 pressed by the performer. The time from time point t1 to time point t2 corresponds to the delay time TD1.
[0100] In the proportional control system, the control unit 20 of communication station C1 transmits a control signal to communication station C2 at time t2, which includes a sound-producing event corresponding to a key press at communication station C1. The sound-producing event specifies the pitch corresponding to the key 12 pressed by the performer as the note number, and a numerical value corresponding to the rotation speed of the hammer 14 as the velocity. The rotation speed of the hammer 14 is calculated, for example, from the hammer trajectory Qh immediately before time t2. In the above explanation, we have focused on pressing one key 12, but the above processing is performed for each key 12 pressed by the performer.
[0101] The control unit 20 at communication base C2 receives a control signal transmitted from communication base C1 at time t3 and drives the keyboard instrument 10 in accordance with the sound-producing event represented by the control signal. Specifically, the control unit 20 displaces one of the multiple keys 12 of the keyboard instrument 10, the key 12 corresponding to the note number specified by the sound-producing event, at a speed corresponding to the velocity specified by the sound-producing event. The hammer 14 rotates in conjunction with the displacement of the key 12 in accordance with the sound-producing event, and the hammer 14 strikes the string 15 at time t4. As described above, the automatic playing piano 1 at communication base C2 produces a musical note at time t4 corresponding to the pitch of the key 12 pressed by the performer at communication base C1 at time t1.
[0102] Figure 18 is an explanatory diagram of the key trajectory Qk and hammer trajectory Qh of the automatic playing piano 1 at communication stations C1 and C2 in the seventh embodiment. As illustrated in Figure 18, the control unit 20 (control signal generation unit 113) of the seventh embodiment estimates the adjustment hammer trajectory Qp corresponding to the key trajectory Qk. The adjustment hammer trajectory Qp is the trajectory in which the hammer 14 reaches the position of the string 15 at a time t2a that is earlier on the time axis than the time t2 of sound production in the actual trajectory of the hammer 14 (hammer trajectory Qh) at communication station C1. In other words, the adjustment hammer trajectory Qp is a trajectory in which the time of striking the string is earlier than the normal hammer trajectory Qh corresponding to the key trajectory Qk.
[0103] The control unit 20 (control signal generation unit 113 and control signal transmission unit 115) of communication base C1 transmits a control signal including the sound production event to communication base C2 at time t2a, which is earlier than the actual sound production time t2. The sound production event specifies the pitch corresponding to the key 12 pressed by the performer as the note number, and a numerical value corresponding to the rotation speed of the hammer 14 as the velocity. The rotation speed of the hammer 14 is calculated, for example, from the adjustment hammer trajectory Qp immediately before time t2a.
[0104] As can be understood from the above explanation, according to the seventh embodiment, the delay time TD1 from the time t1 when the key is pressed to the time t2a when the sound-producing event is transmitted is shortened compared to the proportional method. Therefore, according to the seventh embodiment, the delay time from the time t1 when the performer at communication base C1 presses a key to the time t4 when the musical tone corresponding to that key press is produced at communication base C2 can be shortened.
[0105] A machine learning-based estimation model 61 is used to estimate the adjustment hammer trajectory Qp. Figure 19 is an explanatory diagram of the estimation model 61. The estimation model 61 generates output data Y(t) by processing input data X(t). The input data X(t) includes a portion of the key trajectory Qk. Specifically, the input data X(t) includes one key position from the key trajectory Qk at time t. The key position is a sample value of the measurement signal generated by the key sensor 32. On the other hand, the output data Y(t) includes a portion of the adjustment hammer trajectory Qp. Specifically, the output data Y(t) includes one hammer position from the adjustment hammer trajectory Qp at time t. The estimation model 61 is a statistical model that has learned the relationship between the input data X(t) including the key position and the output data Y(t) including the hammer position through machine learning. The input data X(t) may include time series of multiple key positions. The output data Y(t) may also include time series of multiple hammer positions.
[0106] The control unit 20 (control signal generation unit 113 and control signal transmission unit 115) processes input data X(t) including the key position represented by the measurement signal using the estimation model 61 to generate output data Y(t) including the hammer position in the adjustment hammer trajectory Qp. As a result of repeating the above processing, the adjustment hammer trajectory Qp is generated from the time series of output data Y(t) generated by the estimation model 61.
[0107] The control unit 21 executes a program stored in the memory device 22 to realize the estimation model 61. The estimation model 61 is composed of, for example, a deep neural network (DNN). Any form of deep neural network, such as a convolutional neural network (CNN) or a recurrent neural network (RNN), can be used as the estimation model 61. The estimation model 61 may also be composed of a combination of multiple types of deep neural networks. In addition, additional elements such as long short-term memory (LSTM) may be incorporated into the estimation model 61.
[0108] In Figure 19, the elements enclosed by the dashed line are the elements for machine learning of the estimation model 61. The learning processing unit 62 is implemented, for example, by a machine learning system, and establishes the estimation model 61 through machine learning using multiple training data T. Each of the multiple training data T includes training input data Xa and training output data Ya.
[0109] Each of the multiple training data sets T is generated using key trajectories Qk and hammer trajectories Qh measured in parallel by actual performance of the keyboard instrument 10. Specifically, the training processing unit 62 generates an adjusted hammer trajectory Qp by contracting the actually measured hammer trajectory Qh on the time axis. The time of the string strike in the adjusted hammer trajectory Qp is earlier than the time of the string strike in the hammer trajectory Qh before contraction. The training processing unit 62 includes one key position of key trajectory Qk in the training input data Xa, and includes the hammer position of the adjusted hammer trajectory Qp at the same time as that key position in the training output data Ya.
[0110] As can be understood from the above explanation, the training output data Ya of the training data T is the correct value that the estimation model 61 should output for the training input data Xa of the training data T. Although Figure 19 shows one pair of key trajectories Qk and hammer trajectories Qh (adjusted hammer trajectories Qp) for convenience, in reality, multiple pairs of key trajectories Qk and hammer trajectories Qh are collected and used to generate the training data T.
[0111] The learning processing unit 62 repeatedly performs unit processing to update the initial or provisional estimation model 61 using each of the multiple training data T. In unit processing, the learning processing unit 62 inputs the training input data Xa of one training data T into the estimation model 61 and updates the variables of the estimation model 61 (e.g., weights and biases) so that the error (loss function) between the output data Y(t) generated by the estimation model 61 and the training output data Ya of the training data T is reduced. As the above unit processing is repeated, the estimation model 61 learns the relationship between the training input data Xa and the training output data Ya of the training data T. That is, the estimation model 61 outputs statistically valid output data Y(t) for unknown input data X(t) under the latent relationship between the training input data Xa and the training output data Ya in multiple training data T.
[0112] <Eighth Embodiment> In the eighth embodiment, the control unit 20 of the communication base C1 estimates the hammer trajectory Qh from the key trajectory Qk represented by the measurement signal generated by the key sensor 32. In the seventh embodiment, the adjustment hammer trajectory Qp, which is contracted on the time axis, was estimated, but in the eighth embodiment, the hammer trajectory Qh that approximates the actual trajectory of the hammer 14 in the keyboard instrument 10 is estimated. The hammer trajectory Qh is transmitted from the communication base C1 to the communication base C2 as a control signal and is applied to the control of the keyboard instrument 10 at the communication base C2.
[0113] Figure 20 is an explanatory diagram of the operation of the control unit 20 at the transmitting communication base C1. Figure 20 shows the hammer position y(t) estimated for time t among the hammer trajectory Qh estimated by the control unit 20 (hereinafter referred to as "estimated hammer position"). As a result of the control unit 20 estimating the estimated hammer position y(t) for each different time point t on the time axis, the hammer trajectory Qh, which is composed of a time series of estimated hammer positions y(t), is estimated.
[0114] In the eighth embodiment, the control unit 20 (control signal generation unit 113 and control signal transmission unit 115) uses a trained estimation model 61 to estimate the hammer trajectory Qh. Similar to the seventh embodiment, the estimation model 61 is composed of, for example, a deep neural network.
[0115] As illustrated in Figure 20, the input data X(t) input to the estimation model 61 for estimating the estimated hammer position y(t) includes multiple key positions (multiple sample values of the measurement signal) located within a predetermined period of time (hereinafter referred to as the "input period") of the key trajectory Qk. The input period is a predetermined period of time located in the past with respect to the time t corresponding to the estimated hammer position y(t) to be estimated. The input period of input data X(t+1) is a period shifted backward by a time equivalent to a predetermined number of samples (e.g., one sample) compared to the input period of the immediately preceding input data X(t).
[0116] Furthermore, the output data Y(t) generated by the estimation model 61 includes multiple hammer positions located within a predetermined period of time (hereinafter referred to as the "output period") of the estimated hammer trajectory Qh. The output period is a period after the input period. For example, the output period is a period of predetermined length starting from the end point of the input period. The output period is shorter than the input period. However, forms in which the output period is longer than the input period, or where the output period and the input period are of the same length, are also conceivable. The input period of the output data Y(t+1) is a period shifted backward by a predetermined amount of time (e.g., one sample) from the output period of the immediately preceding output data Y(t).
[0117] One of the multiple hammer positions within the output period represented by the output data Y(t) is selected as the estimated hammer position y(t). Specifically, as illustrated in Figure 20, the hammer position located at the end of the output period is selected as the estimated hammer position y(t).
[0118] The estimation model 61 of the eighth embodiment is a statistical model that has learned, through prior machine learning, the relationship between input data X(t) representing the time series of multiple key positions within the input period and output data Y(t) representing the time series of multiple hammer positions within the output period. In other words, the estimation model 61 outputs statistically valid output data Y(t) for unknown input data X(t) under the latent relationships in the multiple training data used in machine learning.
[0119] As illustrated above, in the eighth embodiment, the hammer trajectory Qh in the output period, which is located after the input period, is estimated from the input data X(t) which includes multiple key positions within the input period in the key trajectory Qk. In other words, it is possible to advance the time at which the control unit 20 of communication station C2 can acquire the hammer trajectory Qh relative to the key trajectory Qk of the keyboard instrument 10 at communication station C1. Therefore, according to the eighth embodiment, similar to the seventh embodiment, the delay time from the time t1 when the performer at communication station C1 presses a key to the time t4 when the musical tone corresponding to that key press is produced at communication station C2 can be shortened. Also, in the seventh embodiment, the hammer trajectory Qh is estimated from the key trajectory Qk. Therefore, it is possible to eliminate the need for the hammer sensor 34 that detects the displacement of each hammer 14.
[0120] In the above explanation, a configuration in which the hammer trajectory Qh is estimated from the key trajectory Qk was illustrated. However, the key trajectory Qk in the output period located after the input period may also be estimated from input data X(t) which includes multiple key positions within the input period of the key trajectory Qk. That is, the output data Y(t) may include multiple key positions within the output period. The key trajectory Qk, which is composed of the time series of output data Y(t), is transmitted as a control signal from communication station C1 to communication station C2 and applied to the control of the keyboard instrument 10 at communication station C2. Even with the above configuration, it is possible to speed up the sound production at communication station C2.
[0121] In the above explanation, the hammer position located at the end of the output period among the multiple hammer positions represented by the output data Y(t) was selected as the estimated hammer position y(t). However, the target of selection as the estimated hammer position y(t) from among the multiple hammer positions within the output period is not limited to the examples above. For example, as illustrated in Figure 21, a hammer position located in the middle of the output period among the multiple hammer positions within the output period may be selected as the estimated hammer position y(t).
[0122] As illustrated in Figure 20, the further back the estimated hammer position y(t) is in the output period (i.e., closer to the end of the output period), the shorter the delay time from the moment the performer at communication station C1 presses the key to the moment the corresponding musical note is produced at communication station C2. On the other hand, as illustrated in Figure 21, the further forward the estimated hammer position y(t) is in the output period (i.e., closer to the start of the output period), the higher the estimation accuracy of the estimated hammer position y(t). In other words, the reduction of delay time and the estimation accuracy of the estimated hammer position y(t) are mutually exclusive.
[0123] Therefore, the control unit 20 may dynamically control the position of the estimated hammer position y(t) within the output period according to the operating environment, such as the communication delay between communication base C1 and communication base C2. For example, in an operating environment with a large communication delay, reducing the communication delay is prioritized, so the control unit 20 selects a hammer position close to the end point within the output period as the estimated hammer position y(t). On the other hand, in an operating environment with a sufficiently small communication delay, improving the estimation accuracy is prioritized, so the control unit 20 selects a hammer position close to the start point within the output period as the estimated hammer position y(t).
[0124] In the above description, an example was given in which the control unit 20 of communication base C1 estimates the hammer trajectory Qh from the key trajectory Qk. However, the main component of the process for estimating the hammer trajectory Qh may be changed as needed. For example, the control unit 20 of communication base C2 may estimate the hammer trajectory Qh from the key trajectory Qk using the method exemplified in the eighth embodiment.
[0125] <Summary of Embodiments 7 and 8> The sound production control system in the seventh and eighth embodiments employs a configuration (hereinafter referred to as "Configuration 1") comprising: an acquisition unit that acquires input data X(t) representing a part of the key trajectory Qk in the keyboard device 581; and a generation unit that processes the input data X(t) using an estimation model 61 to generate output data Y(t) for controlling sound production by the keyboard device 591. With this configuration, it is possible to link the sound production by the keyboard device 591 to the keyboard device 581.
[0126] The sound production control system of the seventh embodiment is characterized in that, in configuration 1, the output data Y(t) represents a portion of the adjustment hammer trajectory Qp at an earlier time of striking the string than the hammer trajectory Qh corresponding to the key trajectory Qk (hereinafter referred to as "configuration 2"). According to configuration 2, since output data Y(t) is generated that represents a portion of the adjustment hammer trajectory Qp at an earlier time of striking the string than the hammer trajectory Qh corresponding to the key trajectory Qk, it is possible to shorten the delay time from pressing a key in the keyboard device 581 to sound production by the keyboard device 591.
[0127] In the eighth embodiment of the sound production control system, in configuration 1, the output data Y(t) represents the portion of the hammer trajectory Qh corresponding to the key trajectory Qk that is behind a part of the key trajectory Qk. According to the above embodiment, since output data Y(t) is generated that represents the portion of the hammer trajectory Qh corresponding to the key trajectory Qk that is behind a part of the key trajectory Qk, it is possible to shorten the delay time from pressing a key in the keyboard device 581 to sound production by the keyboard device 591.
[0128] <Ninth Embodiment> The ninth embodiment is a configuration in which a control signal transmitted from communication base C1 is corrected at the receiving communication base C2. In the ninth embodiment, the measurement signal generated by the key sensor 32 at communication base C1 is transmitted from communication base C1 to communication base C2 as a control signal. That is, the control signal transmitted from communication base C1 to communication base C2 is a measurement signal representing the key trajectory in the keyboard device 581.
[0129] The keyboard device 581 at communication base C1 and the keyboard device 591 at communication base C2 may have different operating characteristics. Therefore, even if the keyboard device 591 at communication base C2 is driven according to the control signal, its behavior may differ from that of the keyboard device 591 at communication base C1. In the ninth embodiment, the control signal is corrected so that the keyboard device 591 at communication base C2 operates in the same way as the keyboard device 581 at communication base C1.
[0130] Figure 22 is a block diagram illustrating the functional configuration of the control unit 20 at the communication base C2. The control device 21 of the control unit 20 functions as a conversion processing unit 71 by executing a program stored in the storage device 22. The conversion processing unit 71 converts the control signal Z1 received from the communication base C1 into a control signal Z2. The control signal Z2 is a signal that instructs the key speed V2 in the keyboard device 591 in a time series.
[0131] Specifically, the conversion processing unit 71 calculates the time series of key velocity V1 from the key trajectory represented by the control signal Z1, and generates the control signal Z2 by converting each key velocity V1 to key velocity V2. The keyboard device 591 operates in accordance with the control signal Z2, causing the sound-producing device 593 to produce sound. Specifically, the key drive device 42 drives the key 12 according to the key velocity V2 represented by the control signal Z2, causing the sound-producing device 593 to produce sound. The driving of the key 12 by the key drive device 42 is a servo control that brings the movement speed of the key 12 detected by the key sensor 32 closer to the key velocity V2 represented by the control signal Z2.
[0132] The conversion processing unit 71 uses a conversion table 72 to generate the control signal Z2. The conversion table 72 is a data table that defines the relationship between the key speed V1 before conversion and the key speed V2 after conversion for each of the multiple keys 12, and is stored in the storage device 22 for each of the multiple keys 12. The conversion processing unit 71 sequentially calculates the key speed V1 from the control signal Z1, searches the conversion table 72 for the key speed V2 corresponding to the key speed V1, and outputs the time series of the key speed V2 as the control signal Z2 to the key drive device 42. As described above, the conversion table 72 defines the relationship between the key speed V1 represented by the control signal Z1 and the key speed V2 represented by the control signal Z2.
[0133] Figure 23 is a flowchart of the operation for generating the conversion table 72 (hereinafter referred to as the "conversion table generation process"). The conversion table generation process is executed, for example, by the control unit 20 (control device 21) of the communication base C2.
[0134] When the conversion table generation process begins, the control device 21 selects one of the multiple keys 12 of the keyboard instrument 10 (hereinafter referred to as the "selected key 12") (S1). The control device 21 also selects one of several different reference speeds (hereinafter referred to as the "target speed") (S2). Each of the multiple reference speeds corresponds to, for example, a different keystroke intensity. For example, eight levels of reference speeds corresponding to keystroke intensity from "ppp (pianississimo)" to "fff (fortississimo)" are assumed.
[0135] The control device 21 instructs the key drive device 42 to drive the selection key 12 at the target speed (S3). The key drive device 42 drives the selection key 12 to strike it repeatedly at a constant speed at the target speed in response to the instruction from the control device 21. However, due to differences in mechanical characteristics between the keyboard instrument 10 at communication station C1 and the keyboard instrument 10 at communication station C2, the actual speed of the selection key 12 (hereinafter referred to as "actual key speed") does not necessarily match the target speed. The control device 21 calculates the actual key speed of the selection key 12 from the measurement signal generated for the selection key 12 by the key sensor 32 (S4). That is, a pair of target speed and actual key speed is generated for the selection key 12.
[0136] The control device 21 determines whether the above processing has been performed for all of the multiple reference speeds (S5). If there are any unprocessed reference speeds (S5: NO), the control device 21 selects the unprocessed reference speed as a new target speed (S2), and then performs the driving of the selected key 12 (S3) and the calculation of the movement speed (S4). Thus, for each of the multiple reference speeds, the relationship between the reference speed (target speed) and the actual key speed is measured.
[0137] If processing is performed for all reference speeds (S5:YES), the control device 21 generates a conversion table 72 for the selected key 12 using the above measurement results (S6). Specifically, the control device 21 generates a conversion table 72 that associates the reference speed with the key speed before conversion V1 and the actual key speed of the selected key 12 with the key speed after conversion V2.
[0138] The control device 21 determines whether or not it has generated a conversion table 72 for all the keys 12 of the keyboard instrument 10 (S7). If there are any unprocessed keys 12 (S7: NO), the control device 21 selects the unprocessed keys 12 as new selected keys 12 (S1), and then generates the conversion table 72 (S2-S6). If the conversion table 72 has been generated for all the keys 12 (S7: YES), the control device 21 terminates the conversion table generation process.
[0139] As can be understood from the above explanation, the conversion table 72 is a data table that registers the relationship between the key speed V1 of the keyboard instrument 10 at communication base C1 and the key speed V2 required to actually drive each key 12 of the keyboard instrument 10 at communication base C2 at a speed equivalent to the key speed V1. Therefore, by driving the keyboard instrument 10 at communication base C2 with the converted control signal Z2, it is possible to displace each key 12 of the keyboard instrument 10 at a speed equivalent to that of each key 12 of the keyboard instrument 10 at communication base C1.
[0140] As described above, according to the ninth embodiment, a control signal Z2 is generated to bring the key trajectory in the receiving keyboard device 591 closer to the key trajectory in the transmitting keyboard device 581. Therefore, it is possible to operate the keyboard device 591 at communication base C2 in the same way as the keyboard device 581 at communication base C1. In other words, it is possible to compensate for the difference between the time change of key position in the transmitting keyboard device 581 and the time change of key position in the receiving keyboard device 591. The ninth embodiment can also be described as a form in which the time change of key position in the receiving keyboard device 591 is calibrated to approach the time change of key position in the transmitting keyboard device 581.
[0141] In the above explanation, the control signal Z1 is shown as representing the key trajectory, but the control signal Z1 may also be a signal that represents the time change of the key velocity V1. As can be understood from the above example, the control signal Z1 is comprehensively expressed as a signal that represents the time change of the key position in the keyboard device 581, and may be a signal that directly represents the time change of the key position (i.e., the key trajectory), or a signal that represents the time change of the key velocity V1.
[0142] Furthermore, although the above explanation illustrates a form in which the control signal Z2 represents the time series of the key velocity V2, the control signal Z2 may also be a signal representing the key trajectory. As can be understood from the above explanation, the control signal Z2 is comprehensively expressed as a signal representing the time change of the key position in the keyboard device 591, and may be a signal representing the time change of the key velocity V2, or a signal representing the time change of the key position (i.e., the key trajectory).
[0143] The sound production control system of the ninth embodiment is described as a system comprising: an acquisition unit that acquires a control signal Z1 representing the time change of key position in the keyboard device 581; and a generation unit that generates a control signal Z2 to bring the time change of key position in the keyboard device 591 closer to the time change of key position represented by the control signal Z1, using the relationship between the time change of key position represented by the control signal Z1 and the time change of key position in the keyboard device 591. With the above configuration, since a control signal Z2 is generated to bring the time change of key position in the keyboard device 591 closer to the time change of key position represented by the control signal Z1, it is possible to operate the keyboard device 591 of the communication base C2 in the same way as the keyboard device 581 of the communication base C1.
[0144] <Variation> The present invention is not limited to the embodiments described above, and includes various other modifications. For example, the embodiments described above are described in detail for the purpose of clearly illustrating the present invention, and are not necessarily limited to those having all the configurations described. Some modifications are described below. Unless otherwise specified, the examples described are modifications of the first embodiment, but they can also be applied as modifications of other embodiments. Multiple modifications can also be combined and applied to each embodiment.
[0145] (1) In the example shown in Figure 1, two communication points C1 and C2 are illustrated, but the number is not limited to this, and there may be many more communication points. If there is a communication point C3, the network delay between communication point C1 and communication point C3 should be measured by the delay acquisition unit 501. The control signal transmitted to communication point C2 and the control signal transmitted to communication point C3 may be generated by different prediction calculations according to the delay time TD acquired by the delay acquisition unit 501.
[0146] (2) The operating device 23 may accept, for example, operations to realize ensemble playing between multiple communication points, operations to determine the operating mode, etc. The operating modes include, for example, ensemble mode and remote performance mode. The ensemble mode is a mode in which sound is controlled by the sound control function 500 described above. The remote performance mode is a mode in which, after it is detected at communication point C1 that sound is actually being produced by the sound production device 583, a control signal is generated to cause sound to be produced by the sound production device 593.
[0147] In remote performance mode, there is no need to consider the effects of network latency. This situation is, for example, when the automatic piano 1 at communication point C2 is used solely to listen to the performance at communication point C1. On the other hand, in remote performance mode, since control signals are transmitted without using the predictive calculations described above, the accuracy of reproducing the performance content at communication point C1 at communication point C2 is improved.
[0148] (3) The prediction calculation is not limited to cases where the time at which the key is pressed to Hp is calculated as the predicted timing. For example, the time at which the hammer 14 strikes the string 15 may be calculated as the predicted timing based on the behavior of the key 12 during the measurement period TP.
[0149] (4) The method for obtaining the delay time TD is not limited to adding the delay times TD1, TD3, and TD4 to the delay time TD2 obtained by measuring the network delay. For example, the delay acquisition unit 501 may acquire the delay time TD as follows: The delay acquisition unit 501 obtains the time Tc1 when the string 15 is sounded by pressing the key 12 at the communication base C1. The delay acquisition unit 501 sends a control signal including note-on from the communication base C1 to the communication base C2 at time Tc1, and obtains the time Tc2 when the string 15 is sounded at the communication base C2. The delay acquisition unit 501 acquires the time from time Tc1 to time Tc2 as the delay time TD.
[0150] (5) The automatic playing piano 1 at communication base C1 does not need to include some of the functions necessary to achieve automatic playing (for example, the drive device 40).
[0151] (6) In the fourth embodiment, the communication base C2 does not need to have a configuration equivalent to the keyboard device 591, but it does need to have a configuration equivalent to the sound-producing device 593. In this case, a control signal to cause the sound-producing device 593 to produce sound should be transmitted to the sound-producing device 593 via the network.
[0152] (7) The keyboard device 581 and the sound-producing device 583 do not have to be configured as the same device, such as the keyboard instrument 10. For example, the keyboard device 581 may be a device such as a MIDI keyboard that does not have a sound-producing function, and the sound-producing device 583 may be connected to the keyboard device 581 by wire or wireless connection.
[0153] (8) The operation information indicated the pressed position of the key 12, but it may also indicate the position of a component that is linked to the key 12. The component linked to the key 12 may be, for example, a hammer 14.
[0154] (9) In the second embodiment, the control signal is not limited to indicating the pressed position, but may also be information indicating a change in the pressed position over time. For example, the control signal may be the pressing speed or pressing acceleration when driving the key 12. In controlling the key 12 at the communication base C2, for example, if data loss occurs in network communication in the case of the second embodiment, if the control signal is an pressed position, the next pressed position cannot be determined. However, if the control signal is information indicating a change in time, the driving of the key 12 can be continued based on the last received control signal.
[0155] (10) The receiving function 200 is not limited to causing the string 15 to sound by driving the key 12, but may also cause the string 15 to sound by using a different configuration, such as using a device to drive the hammer 14.
[0156] (11) The conditions immediately before the key 12 is pressed may be included as a parameter used in the prediction calculation. For example, by providing a pressure sensor on the surface of the key 12, the pressure on the key 12 immediately before it is pressed may be obtained and incorporated into the prediction calculation.
[0157] (12) The automatic piano 1 at communication base C1 and the automatic piano 1 at communication base C2 both implement the transmission function 100 and the reception function 200, so that control by the sound production control function 500 can be achieved bidirectionally. Sound production at communication base C2 occurs at approximately the same timing as sound production by performance operation at communication base C1. Furthermore, sound production at communication base C1 occurs at approximately the same timing as sound production by performance operation at communication base C2. In this case, operation information is generated only for the keys 12 played by the performer, and keys 12 driven by the key drive device 42 can be excluded from the generation of operation information. If sound production is performed using the sound source device 25 as shown in the fourth and fifth embodiments, it is not necessary to drive the keys 12, so such processing related to the generation of operation information is unnecessary.
[0158] (13) The method in the sound control function 500 may also be applied to a component other than the key 12, for example, the pedal 13. In this way, the pressing position of the pedal 13 at the communication base C1 can be predicted and used for driving control of the pedal 13 at the communication base C2.
[0159] (14) The keyboard instrument 10 in the automatic playing piano 1 is not limited to an acoustic piano such as a grand piano, but may also be an electronic keyboard instrument.
[0160] (15) The control unit 20 does not have to be a device attached to the keyboard instrument 10, but may be, for example, a personal computer, a tablet computer, a smartphone, etc.
[0161] (16) The network NW connecting the communication bases may be a dedicated line implemented by optical cable or the like. [Explanation of Symbols]
[0162] 1: Automatic playing piano, 10: Keyboard instrument, 12: Key, 13: Pedal, 14: Hammer, 15: String, 16: Bridge, 17: Soundboard, 18: Damper, 19: Straight support, 20: Control unit, 21: Control device, 22: Memory device, 23: Operating device, 24: Communication device, 25: Sound source device, 26: Interface, 27: Bus, 30: Sensor, 32: Key sensor, 33: Pedal sensor, 34: Hammer sensor, 40: Drive device, 42: Key drive device, 43: Pedal drive device, 44: Stopper, 47: Vibrator, 48: Damper drive device, 100: Transmission function, 101: Delay acquisition unit, 111: Operation information generation unit, 113: Control signal generation unit, 115: Control signal transmission unit, 200, 200A: Receiving function, 211: Control signal reception unit, 215: Drive signal generation unit, 217: Sound signal generation unit, 219: Sound signal output unit, 300: Transmission function, 311: Operation information generation unit, 315: Operation information transmission unit, 400, 400A: Receiving function, 401: Delay acquisition unit, 411: Operation information reception unit, 413: Control signal generation unit, 415: Drive signal generation unit, 417: Sound signal generation unit, 419: Sound signal output unit, 500: Sound generation control function, 501: Delay acquisition unit, 511: Operation information acquisition unit, 513: Control signal generation unit, 581: Keyboard device, 583: Sound generation device, 591: Keyboard device, 593: Sound generation device, 1000: Server
Claims
1. A first acquisition unit acquires a delay time, including network delay, between a first keyboard device which includes a first key and causes a first sound-producing device to produce sound in accordance with the movement of the first key, and a second sound-producing device which is connected via a network. A second acquisition unit acquires operation information corresponding to the pressing of the first key in the first keyboard device, A generation unit that generates a first control signal for causing the second sound-producing device to produce sound based on the aforementioned delay time and the predicted timing of sound production in the first sound-producing device calculated using the aforementioned operation information, A pronunciation control system, including a sound control system.
2. The first keyboard device is connected to the second keyboard device via the network, The second keyboard device includes a second key, a second sound-producing device, and a drive device for driving the second key. The second sound-producing device produces sound in accordance with the movement of the second key. The first control signal includes a signal to cause the drive device to drive the second key. The sound control system according to claim 1.
3. The second sound-producing device includes a sound signal generation unit that generates a sound signal using a sound source device and a sound signal output unit that outputs the sound signal, The first control signal includes a signal for causing the sound signal generation unit to generate a sound signal. The sound control system according to claim 1.
4. The generating unit is If the first operating mode is selected from a plurality of operating modes, including the first operating mode and the second operating mode, the first control signal is generated. When the second operating mode is selected, after it is detected that sound is to be produced in the first sound-producing device in response to the pressing of the first key, the second sound-producing device generates a second control signal to realize the sound. The sound control system according to claim 1.
5. The first control signal is transmitted to the second sound-producing device via the network. The sound control system according to claim 1.
6. The first control signal is transmitted to the second keyboard device via the network. The signal for driving the second key includes information indicating the speed or acceleration when driving the second key. The sound control system according to claim 2.
7. The generation unit calculates the predicted timing using the operation information acquired from the first acquisition unit during the measurement period after the first key is pressed. The longer the aforementioned delay time, the shorter the aforementioned measurement period becomes. The sound control system according to claim 1.
8. The first control signal includes a signal to cause the second sound-producing device to produce sound at a timing earlier than the predicted timing. The sound control system according to claim 1.
9. To obtain a delay time, including network delay, between a first keyboard device that includes a first key and causes a first sound-producing device to produce sound in accordance with the movement of the first key, and a second sound-producing device connected via a network, To acquire operation information corresponding to the pressing of the first key in the first keyboard device, A first control signal is generated to cause the second sound-producing device to produce sound based on the aforementioned delay time and the predicted timing of sound production in the first sound-producing device calculated using the aforementioned operation information. A method of controlling pronunciation, including...
10. The first keyboard device is connected to the second keyboard device via the network, The second keyboard device includes a second key, a second sound-producing device, and a drive device for driving the second key. The second sound-producing device produces sound in accordance with the movement of the second key. The first control signal includes a signal to cause the drive device to drive the second key. The method for controlling sound production according to claim 9.
11. The second sound-producing device includes a sound signal generation unit that generates a sound signal using a sound source device and a sound signal output unit that outputs the sound signal, The first control signal includes a signal for causing the sound signal generation unit to generate a sound signal. The method for controlling sound production according to claim 9.
12. Generating the first control signal means When the first operating mode is selected from a plurality of operating modes, including the first operating mode and the second operating mode, the first control signal is generated. When the second operating mode is selected, after it is detected that sound is to be produced in the first sound-producing device in response to the pressing of the first key, a second control signal is generated in the second sound-producing device to realize the sound. including, The method for controlling sound production according to claim 9.
13. The first control signal is transmitted to the second sound-producing device via the network. The method for controlling sound production according to claim 9.
14. The first control signal is transmitted to the second sound-producing device via the network. The signal for driving the second key includes information indicating the speed or acceleration when driving the second key. The method for controlling sound production according to claim 10.
15. Generating the first control signal includes calculating the predicted timing using the operation information acquired during the measurement period after the first key is pressed. The longer the aforementioned delay time, the shorter the aforementioned measurement period becomes. The method for controlling sound production according to claim 10.
16. The first control signal includes a signal to cause the second sound-producing device to produce sound at a timing earlier than the predicted timing. The method for controlling sound production according to claim 9.
17. A program for causing a computer to execute the pronunciation control method described in any one of claims 9 to 16.
Citation Information
Patent Citations
Control device
JP2023154288A