Sound control device and control method thereof, program, and electronic musical instrument
The sound control device addresses the issue of inconsistent syllable production by using performance information to control syllable pronunciation, ensuring alignment with the performer's intentions through precise note-on and note-off timing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-03-10
AI Technical Summary
Existing sound control devices for musical instruments do not adequately account for the performer's intention in controlling syllable production, particularly during note-offs, leading to inconsistencies in syllable pronunciation.
A sound control device that includes an acquisition unit for performance information, a judgment unit for note-ons and note-offs, an identification unit for syllables, and an instruction unit to control syllable pronunciation based on lyric data, allowing for precise syllable production aligned with the performer's intentions.
Enables syllable production that accurately follows the performer's intentions by starting and ending syllable pronunciation at appropriate times, enhancing musical expression and control.
Smart Images

Figure 2026042060000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a sound control device, a control method thereof, a program, and an electronic musical instrument. [Background technology]
[0002] In sound control devices for musical instruments, etc., in addition to generating electronic sounds that simulate musical instrument sounds, etc., synthetic singing sounds are also generated by synthesizing singing sounds. Patent Documents 1, 2, and 3 disclose technologies for generating synthetic singing sounds in real time in response to performance operations. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-206496 [Patent Document 2] Japanese Patent Application Laid-Open No. 2014-98801 [Patent Document 3] Patent No. 7036141 Summary of the Invention [Problem to be solved by the invention]
[0004] However, simply starting sound production in response to a note-on in a performance operation and ending sound production in response to a note-off in a performance operation may not always conform to the performer's intention, depending on the syllable to be produced. For example, little consideration has been given to controlling the production of syllables in response to a note-off. Therefore, there is room for improvement in producing syllables in accordance with the performer's intention.
[0005] One object of the present invention is to provide a sound control device that enables the player to produce syllables in accordance with his or her intention. [Means for solving the problem]
[0006] According to one aspect of the present invention, there is provided a sound control device having an acquisition unit that acquires performance information, a judgment unit that judges note-on and note-off based on the performance information, an identification unit that identifies a syllable corresponding to the timing at which the judgment unit judges the note-on from lyric data in which multiple syllables to be pronounced are arranged in chronological order, and an instruction unit that instructs the identification unit to start pronouncing the syllable identified at the timing corresponding to the note-on, and instructs the identification unit to pronounce some of the phonemes that make up the identified syllable at the timing corresponding to the note-off. [Effects of the Invention]
[0007] According to one aspect of the present invention, it is possible to produce syllables in accordance with the performer's intention. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a block diagram of a sound control system including a sound control device. [Figure 2] FIG. 10 is a diagram showing lyric data. [Figure 3] FIG. 2 is a functional block diagram of the sound control device. [Figure 4] 10 is a timing chart showing an example of sound control in response to a performance signal. [Figure 5] 10 is a flowchart showing a sound control process. [Figure 6] 10 is a flowchart showing an instruction process. [Figure 7] 10 is a flowchart showing an English language support process. [Figure 8] 10 is a flowchart showing Japanese language support processing. [Figure 9] 10 is a timing chart showing an example of sound control in the second embodiment. [Figure 10] 10 is a flowchart showing an English language support process. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0010] (First embodiment) 1 is a block diagram of a sound control system including a sound control device according to a first embodiment of the present invention. This sound control system includes a sound control device 100 and an external device 20. As an example, the sound control device 100 is an electronic musical instrument, and may be an electronic wind instrument in the form of, for example, a saxophone.
[0011] The sound control device 100 includes a control unit 11, an operation unit 12, a display unit 13, a memory unit 14, a performance operation unit 15, a sound generation unit 18, and a communication I / F (interface) 19. These elements are connected to each other via a communication bus 10.
[0012] The control unit 11 includes a CPU 11a, a ROM 11b, a RAM 11c, and a timer (not shown). The ROM 11b stores a control program executed by the CPU 11a. The CPU 11a loads the control program stored in the ROM 11b into the RAM 11c and executes it to realize various functions of the sound control device 100.
[0013] The control unit 11 includes a DSP (Digital Signal Processor) for generating an audio signal. The storage unit 14 is a non-volatile memory. The storage unit 14 stores setting information used when generating an audio signal representing a synthetic singing voice, as well as voice segments and the like for generating the synthetic singing voice. The setting information includes, for example, tone color and acquired lyric data.
[0014] The operation unit 12 includes a plurality of operators for inputting various information and receives instructions from the user. The display unit 13 displays various information. The sound generation unit 18 includes a sound source circuit, an effect circuit, and a sound system.
[0015] The performance operation unit 15 includes a plurality of operation keys 16 and a breath sensor 17 as elements for inputting a performance signal (performance information). The input performance signal includes pitch information indicating the pitch and volume information indicating the volume detected as a continuous quantity, and is supplied to the control unit 11. The main body of the sound control device 100 is provided with a plurality of tone holes (not shown). When a user (performer) plays the plurality of operation keys 16, the open / closed states of the tone holes change, thereby specifying the desired pitch.
[0016] A mouthpiece (not shown) is attached to the main body of the sound control device 100, and the breath sensor 17 is provided near the mouthpiece. The breath sensor 17 is a blowing pressure sensor that detects the pressure of the breath blown by the user through the mouthpiece. The breath sensor 17 detects whether or not breath is being blown, and during performance, it detects the strength and speed (force) of the blowing pressure. The volume is specified according to the change in pressure detected by the breath sensor 17. The magnitude of the pressure that changes over time, detected by the breath sensor 17, is treated as volume information detected as a continuous quantity.
[0017] The communication I / F 19 is connected to a communication network wirelessly or by wire. The sound control device 100 is connected, for example, by the communication I / F 19 to be able to communicate with an external device 20 via the communication network. The communication network may be, for example, the Internet, and the external device 20 may be a server device. The communication network may be a short-range wireless communication network using Bluetooth (registered trademark), infrared communication, LAN, etc. The number and types of connected external devices are not important. The communication I / F 19 may include a MIDI I / F that transmits and receives MIDI (Musical Instrument Digital Interface) signals.
[0018] The external device 20 stores music data necessary for providing karaoke, associated with a music ID. This music data includes data related to the song to be sung in karaoke, such as lead vocal data, chorus data, accompaniment data, and karaoke subtitle data. The accompaniment data is data indicating the accompaniment to the song. The lead vocal data, chorus data, and accompaniment data may be expressed in MIDI format. The karaoke subtitle data is data for displaying lyrics on the display unit 13.
[0019] The external device 20 also stores setting data in association with the song ID. This setting data is data that is set in the sound control device 100 according to the song to realize the synthesis of singing sounds. The setting data includes lyric data corresponding to each part of the song corresponding to the song ID. This lyric data is, for example, lyric data corresponding to the lead vocal part. The song data and setting data are associated with each other in terms of time.
[0020] This lyrics data may be the same as or different from the karaoke subtitle data. That is, the lyrics data is the same in that it is data that specifies the lyrics (characters) to be spoken, but it is adjusted to a format that is easy to use in the sound control device 100.
[0021] For example, the karaoke subtitle data is a character string of "ko", "n", "ni", "chi", and "ha". In contrast, the lyrics data may be a character string that matches the actual pronunciation of "ko", "n", "ni", "chi", and "wa" so that it can be easily used by the sound control device 100. In addition, this format may include, for example, information that identifies when two characters are sung per note, information that identifies the division of phrases, etc.
[0022] In performing the sound control process, the control unit 11 acquires music data and setting data designated by the user from the external device 20 via the communication I / F 19 and stores them in the storage unit 14. As described above, the music data includes accompaniment data, and the setting data includes lyric data. Moreover, the accompaniment data and the lyric data are associated with each other in terms of time.
[0023] FIG. 2 is a diagram showing lyric data. Hereinafter, each lyric (character) to be uttered, that is, each phonetic unit (a group of sounds) may be referred to as a "syllable." Lyric data is data that specifies the syllables to be uttered. Lyric data has text data in which multiple syllables to be uttered are arranged in chronological order. The syllables to be pronounced are identified in order according to the progress of the performance. Therefore, in the lyric data shown in FIG. 2, characters M(i)=M(1) to M(n) are uttered in order.
[0024] As shown in Figure 2, the lyrics data includes text data representing "ko," "n," "ni," "chi," "wa," "christ," "mas," "make," "fast," "desks," "ma," "su," and so on. The syllables representing "ko," "n," "ni," "chi," "wa," "christ," "mas," "make," "fast," "desks," "ma," "su," and so on, are associated with M(i), and the order of the syllables in the lyrics is determined by "i" (i = 1 to n). For example, M(5) corresponds to the fifth syllable in the lyrics. As explained below, the vocalization duration of each syllable included in the synthesized singing sound is controlled based on the performance information.
[0025] 3 is a functional block diagram of the sound control device 100 for realizing sound generation processing. The sound control device 100 includes, as functional units, an acquisition unit 31, a determination unit 32, a generation unit 33, an identification unit 34, a singing sound synthesis unit 35, and an instruction unit 36. The functions of these functional units are realized by cooperation between a CPU 11a, a ROM 11b, a RAM 11c, a timer, a communication I / F 19, etc. Note that the generation unit 33 and the singing sound synthesis unit 35 are not essential.
[0026] The acquisition unit 31 acquires a performance signal. The determination unit 32 determines the occurrence of a note-on (note start) and a note-off (note end) based on the comparison result between the performance signal and a threshold value. The generation unit 33 generates notes based on the determination of note-on and note-off. The identification unit 34 identifies, from the lyrics data, a syllable corresponding to the timing at which the determination unit 32 determines that a note is on.
[0027] The singing sound synthesis unit 35 synthesizes the identified syllables based on the setting data to generate a singing sound. The instruction unit 36 instructs the singing sound of the identified syllable to start being produced at a pitch and timing corresponding to a note-on, and to end the production at a timing corresponding to a note-off. Based on the instruction from the instruction unit 36, the singing sound obtained by synthesizing the syllables is produced by the pronunciation unit 18 (FIG. 1).
[0028] The instruction unit 36 instructs that some of the phonemes constituting the identified syllable be pronounced at a timing corresponding to a note-off, not a note-on. An example of the pronunciation control of some of the phonemes constituting the identified syllable will be described with reference to FIG. 4.
[0029] Next, an overview of the sound control process will be given. Lyric data and accompaniment data corresponding to a song designated by the user are stored in the memory unit 14. When the user issues a command to start playing using the operation unit 12, playback of the accompaniment data begins. That is, the sound generation unit 18 generates sounds corresponding to the accompaniment data. At this time, the lyrics of the lyric data (or karaoke subtitle data) are displayed on the display unit 13 in accordance with the progression of the accompaniment data. Note that the setting data may include musical score data, and in that case, the musical score of the main melody corresponding to the lead vocal data may also be displayed on the display unit 13 in accordance with the progression of the accompaniment data. The user plays using the performance operation unit 15 while listening to the accompaniment data. As the performance progresses, the acquisition unit 31 acquires a performance signal. Note that it is not essential that the accompaniment data be played back.
[0030] FIG. 4 is a timing chart showing an example of sound control in response to a performance signal.
[0031] The horizontal axis of Figure 4 represents elapsed time t, and the vertical axis represents the "performance depth" indicated by the performance signal. Here, the larger the value detected by the breath sensor 17, the stronger the blowing pressure, i.e., the deeper the performance. When not playing, the blowing pressure is "0." The volume information is determined by the performance depth.
[0032] In addition to the sound generation threshold TH0, a first threshold THA and a second threshold THB are provided as thresholds for mute control to be compared with the performance depth. The performance depth of the second threshold THB is shallower than the performance depth of the first threshold THA. The magnitude relationship between the sound generation threshold TH0 and the thresholds THA and THB does not matter, but in the example shown in FIG. 4, the performance depth of the sound generation threshold TH0 is deeper than the performance depth of the first threshold THA. The second threshold THB may be the same as "0".
[0033] In the example of Figure 4, from the non-playing state, the performance depth temporarily becomes deeper than the sounding threshold TH0, then crosses the thresholds THA and THB to shallower sides in that order, returning to the non-playing state. The time when the performance depth crosses the sounding threshold TH0 to deeper side is T1. The time when the performance depth crosses the first threshold THA to shallower side is T2. The time when the performance depth crosses the second threshold THB to shallower side is T3.
[0034] At time T1, the control unit 11 identifies a syllable to be pronounced and starts pronouncing the syllable. At this time, the control unit 11 controls the sound differently depending on whether the identified syllable has a consonant at the end or not. Hereinafter, a syllable with a consonant at the end is referred to as a "special syllable," and a syllable without a consonant at the end is referred to as a "non-special syllable."
[0035] For example, in the case of the syllable "see [si]" which is not a special syllable, the control is performed as follows: [si] is a phonetic notation. The control unit 11 starts pronouncing [si] at time T1 and ends pronouncing [si] at time T3.
[0036] On the other hand, the special syllable "mas[ma][s]" includes the consonant [s] at the end. Therefore, the control unit 11 starts pronouncing the first phoneme [ma] of the syllable "mas" at time T1. Then, at time T2, the control unit 11 finishes pronouncing [ma] and starts pronouncing the remaining phonemes [s] including the final consonant (start of pronouncing consonants, etc.), and further finishes pronouncing [s] at time T3.
[0037] Therefore, the duration of the pronunciation of [ma] is from time T1 to T2, and the duration of the pronunciation of [s] (the duration of the pronunciation of consonants, etc.) is from time T2 to T3. Here, the change in performance depth from time T2 to T3 indicates the degree of temporal change in performance depth, and therefore essentially corresponds to the velocity of note-off in performance (note-off velocity). Therefore, by speeding up / slowing down the operation by the user to shallow the performance depth, it is possible to make the pronunciation of [s] shorter / longer. In conventional control, when pronouncing the syllable "mas," there are cases where the pronunciation of [ma] starts when a note-on is detected, and ends when a note-off is detected. In this control, the pronunciation of [s] is omitted, and therefore it cannot be said that it is fully in line with the intention of the performer. In contrast, in this embodiment, it is possible to control the pronunciation of syllables according to note-off. This has made it possible for the performer to control the pronunciation of final consonants in particular.
[0038] Next, the sound control process will be described using a flowchart. In the sound control process, an instruction to generate or stop an audio signal corresponding to each syllable is output based on a performance operation on the performance operation unit 15.
[0039] 5 is a flowchart showing the sound control process. This process is realized by the CPU 11a loading a control program stored in the ROM 11b into the RAM 11c and executing it. This process starts when the user instructs playback of a song.
[0040] In step S101, the control unit 11 acquires lyric data from the storage unit 14. Next, in step S102, the control unit 11 executes initialization processing. In this initialization, a count value tc=0 is set, and various register values and flags are set to their initial values. Furthermore, the control unit 11 sets a character count value i=1 in M(i) (character M(i)=M(1)). As described above, "i" indicates the pronunciation order of syllables in the lyrics.
[0041] Next, in step S103, the control unit 11 increments the count value tc by setting the count value tc=tc+1. Furthermore, on the condition that the pronunciation instruction for the last identified syllable has been completed in step S108 (described later), the control unit 11 increments "i" to advance the syllable indicated by M(i) one by one among the syllables constituting the lyrics. In step S104, the control unit 11 reads out the data of the portion of the accompaniment data that corresponds to the count value tc.
[0042] In step S105, the control unit 11 determines whether or not the reading of the accompaniment data has finished. If the reading of the accompaniment data has not finished, in step S106, the control unit 11 determines whether or not the user has input an instruction to stop playing the music. If the user has not input an instruction to stop playing the music, in step S107, the control unit 11 determines whether or not a performance signal has been received. The performance signal here also includes the performance depth passing a threshold value. If a performance signal has not been received, the control unit 11 returns to step S105.
[0043] If the reading of the accompaniment data is completed in step S105, or if the user inputs an instruction to stop the performance of the music piece in step S106, the control unit 11 ends the processing shown in Fig. 5. If a performance signal is received from the performance operation unit 15 in step S107, the control unit 11 executes instruction processing for generating an audio signal by the DSP (step S108). Details of the instruction processing for generating an audio signal will be described later with reference to Fig. 6. When the instruction processing for generating an audio signal is completed, the control unit 11 returns to step S103.
[0044] FIG. 6 is a flowchart showing the instruction process executed in step S108 of FIG.
[0045] First, in step S201, the control unit 11 determines whether the syllable to be pronounced this time has already been identified. This syllable is the syllable corresponding to the timing determined as a note-on, and is identified in step S305 (FIG. 7) or step S405 (FIG. 8), which will be described later.
[0046] If the syllable to be pronounced this time has already been identified, the control unit 11 proceeds to step S203, and if the syllable to be pronounced this time has not yet been identified, the control unit 11 proceeds to step S202. In step S202, the control unit 11 provisionally identifies the syllable to be pronounced this time. As described above, the order in which the syllables to be pronounced are identified is determined by the character count value i. Therefore, except for the beginning of a song, the syllable next to the syllable pronounced immediately before is provisionally identified as the syllable to be pronounced this time. After step S202, the control unit 11 proceeds to step S203.
[0047] In step S203, the control unit 11 determines the language of the identified syllables, and further determines whether the determined language is English. Note that any language determination method may be used, and a known method such as that disclosed in Japanese Patent No. 6553180 may be adopted. Note that the user may specify a language in advance for each song, each section of a song, or each syllable that constitutes a song, and the control unit 11 may determine the language for each syllable based on the specification.
[0048] If the language of the identified syllable is English, the control unit 11 proceeds to step S205, otherwise the control unit 11 proceeds to step S204. In step S205, the control unit 11 executes an English support process (FIG. 7) to be described later, and ends the process shown in FIG.
[0049] In step S204, the control unit 11 determines whether the language of the identified syllable is Japanese. Here, the above-described language determination method is used. If the language of the identified syllable is Japanese, the control unit 11 proceeds to step S206, and if the language of the identified syllable is not Japanese, the control unit 11 proceeds to step S207.
[0050] In step S206, the control unit 11 executes a Japanese language handling process (FIG. 8) to be described later, and then ends the process shown in Fig. 6. In step S207, the control unit 11 executes an "other language handling process" (not shown) according to the language of the identified syllable, and then ends the process shown in Fig. 6.
[0051] Fig. 7 is a flowchart showing the English adaptation process executed in step S205 of Fig. 6. In this process, the identification unit 34 identifies one syllable for one note-on.
[0052] First, in step S301, the control unit 11 determines whether or not the flag F is set to 1 (flag F=1). Here, the flag F is a flag that indicates, when it is "1," that the pronunciation of a special syllable has started. The flag F is set to "1" in step S308. Then, if the flag F is not "1," the control unit 11 proceeds to step S302.
[0053] In step S302, the control unit 11 determines whether a new note-off has occurred based on the performance depth indicated by the performance signal. That is, the control unit 11 determines whether the performance depth determined by the detection result of the breath sensor 17 has newly crossed the second threshold value THB to the shallower side (whether time T3 in FIG. 4 has arrived).
[0054] If the control unit 11 determines that the performance depth has not newly crossed the second threshold value THB to the shallower side, the control unit 11 proceeds to step S303, where it determines whether a new note-on has occurred based on the performance depth indicated by the performance signal. That is, the control unit 11 determines whether the performance depth determined by the detection result of the breath sensor 17 has newly crossed the sound generation threshold value TH0 to the deeper side (whether time T1 in FIG. 4 has arrived).
[0055] If the control unit 11 determines that a new note-on has not occurred, the control unit 11 proceeds to step S317, executes other processing, and ends the processing shown in Fig. 7. In this "other processing," the control unit 11 outputs, for example, if vocalization is in progress, an instruction to change the pronunciation volume or pitch in response to the acquired change in performance depth. On the other hand, if the control unit 11 determines that a new note-on has occurred, the control unit 11 proceeds to step S304.
[0056] In step S304, the control unit 11 sets the pitch indicated by the acquired performance signal. In step S305, the control unit 11 identifies the syllable to be pronounced this time according to the identification order of the syllables to be pronounced. This syllable is the syllable corresponding to the timing determined as note-on in step S303.
[0057] In step S306, the control unit 11 determines whether the syllable identified in step S305 is a syllable having a consonant at the end (i.e., a special syllable). If the identified syllable is not a special syllable, the control unit 11 proceeds to step S309.
[0058] In step S309, the control unit 11 instructs the DSP to start producing the identified syllable at the pitch and timing corresponding to the current note-on. That is, the control unit 11 outputs an instruction to start generating an audio signal based on the set pitch and the pronunciation of the identified syllable to the DSP. This instruction to start producing is a normal instruction to continue producing until note-off. For example, if the identified syllable is "see," which is not a special syllable, the control unit 11 starts producing [si]. Thereafter, the control unit 11 ends the processing shown in FIG. 7.
[0059] If the control unit 11 determines in step S302 that the performance depth has newly crossed the second threshold value THB to the shallower side, the control unit 11 proceeds to step S316. In step S316, the control unit 11 instructs the currently identified syllable to end pronunciation at the timing corresponding to the current note-off. For example, if the identified syllable is the syllable "see," pronunciation of [si] ends. Thereafter, the control unit 11 ends the processing shown in FIG. 7.
[0060] If the identified syllable is a special syllable as a result of the determination in step S306, the control unit 11 proceeds to step S307. In step S307, the control unit 11 instructs the control unit 11 to start pronouncing the identified syllable except for "some phonemes including a final consonant." Therefore, the control unit 11 instructs the control unit 11 to start pronouncing the first phoneme of the identified syllable, but does not instruct the control unit 11 to start pronouncing the remaining phonemes including the final consonant. For example, if the identified syllable is the special syllable "mas," the control unit 11 starts pronouncing the first phoneme, [ma], of the syllable "mas" at time T1 (FIG. 4). However, the control unit 11 does not start pronouncing the remaining phoneme, [s], including the final consonant.
[0061] In step S308, the control unit 11 sets the flag F to "1" (flag F=1), and ends the processing shown in FIG.
[0062] If the result of the determination in step S301 is that flag F=1, the control unit 11 proceeds to step S310. In step S310, the control unit 11 determines whether a new note-off has occurred based on the performance depth indicated by the performance signal. That is, the control unit 11 determines whether the performance depth determined by the detection result of the breath sensor 17 has newly crossed the first threshold value THA to the shallower side (whether time T2 in FIG. 4 has arrived). Note that in this embodiment, for convenience, the case where the performance depth newly crosses the second threshold value THB to the shallower side (S302) and the case where the performance depth newly crosses the first threshold value THA to the shallower side (S310) are both referred to as note-offs.
[0063] Then, if the control unit 11 determines that the performance depth has newly crossed the first threshold value THA to the shallower side, the control unit 11 proceeds to step S311. In step S311, the control unit 11 instructs the start of pronunciation of "some phonemes including a final consonant" of the identified syllable, that is, the remaining phonemes including the final consonant. At this time, the control unit 11 ends the pronunciation started in step S307. For example, if the identified syllable is the special syllable "mas," the control unit 11 ends the pronunciation of [ma] and starts the pronunciation of the remaining phoneme [s] including the final consonant at time T2 (FIG. 4). Thereafter, the control unit 11 ends the processing shown in FIG. 7. .
[0064] On the other hand, if the control unit 11 determines in step S310 that the performance depth has not newly crossed the first threshold value THA to the shallower side, it determines in step S312 whether a new note-off has occurred. That is, the control unit 11 determines whether the performance depth determined by the detection result of the breath sensor 17 has newly crossed the second threshold value THB to the shallower side (whether time T3 in FIG. 4 has arrived).
[0065] If the control unit 11 determines that the performance depth has not newly crossed the second threshold value THB to the shallower side, the control unit 11 proceeds to step S314, executes other processing, and ends the processing shown in Fig. 7. In the "other processing" here, the control unit 11 outputs, for example, an instruction to change the sound volume or pitch in response to the acquired change in performance depth.
[0066] If the control unit 11 determines in step S312 that the performance depth has newly crossed the second threshold value THB to the shallower side, the control unit 11 proceeds to step S313. In step S313, the control unit 11 instructs the control unit 11 to end pronunciation of "some phonemes including a final consonant" of the identified syllable, that is, the remaining phonemes including the final consonant.
[0067] For example, if the identified syllable is the special syllable "mas," the control unit 11 ends the pronunciation of the remaining phoneme [s] including the final consonant at time T3 (FIG. 4). As a result, the pronunciation of [s] continues only during the period from time T2 to T3. The user can adjust the period from time T2 to T3 through performance, so that the disappearance of the remaining phonemes including the final consonant can be controlled, thereby expanding the range of musical expression.
[0068] Note that the control unit 11 essentially instructs that the pronunciation of the vowels, which was started from the first phoneme in step S307, should be continued until an instruction to pronounce the remaining phonemes is given in step S313.
[0069] In step S315, the control unit 11 sets the flag F to "0" (flag F=0), and ends the processing shown in FIG.
[0070] FIG. 8 is a flowchart showing the Japanese language support process executed in step S206 of FIG.
[0071] In this process, the identification unit 34 may identify two or more syllables for one note-on. A setting specific to this process is a "collective pronunciation setting." For example, a user can set the collective pronunciation setting when instructing playback of a song. The collective pronunciation setting is a setting in which multiple syllables are identified as a set for one note-on, and only the consonant is pronounced for the last syllable of these multiple syllables.
[0072] For example, the "ma" in M(11) and the "su" in M(12) shown in FIG. 3 are each one syllable. Consider a case where "ma" and "su" are a pair of syllables specified for one note-on due to the batch pronunciation setting. In this case, in response to one note-on, the first syllable "ma" is pronounced as usual, but the last syllable "su" does not have a vowel pronounced, and only the consonant [s] is pronounced. The instruction unit 36 instructs the pronunciation to start from the first phoneme [ma] of "ma" at the timing corresponding to the note-on, and also instructs the pronunciation of the consonant [s] of "su" at the timing corresponding to the note-off. The following is a description based on the flowchart.
[0073] In steps S401 to S404, the control unit 11 executes the same processes as steps S301 to S304 in Fig. 7. In step S405, the control unit 11 specifies a syllable to be pronounced this time according to the specified order of the syllables to be pronounced. At that time, if a syllable in the specified order corresponds to the first syllable in a group set by the collective pronunciation setting, the control unit 11 specifies multiple syllables in the group including the first syllable as the syllables to be pronounced this time.
[0074] In step S406, the control unit 11 determines whether the identified syllables are a pair according to the collective pronunciation setting. If the identified syllables are not a pair according to the collective pronunciation setting, the control unit 11 executes the same process as in step S309 in step S410. On the other hand, if the identified syllables are a pair according to the collective pronunciation setting, the control unit 11 proceeds to step S407.
[0075] In step S407, the control unit 11 instructs to start pronunciation from the first phoneme of the first syllable of the identified pair of syllables. That is, the identified syllables are started to be pronounced except for the consonant phoneme of the last syllable. For example, if "ma" and "su" are a pair in the collective pronunciation setting, the control unit 11 instructs to start pronunciation of the first phoneme [ma] of "ma" (time point T1).
[0076] In step S408, the control unit 11 executes the same process as in step S308. In steps S417 and S409, the control unit 11 executes the same processes as in steps S316 and S317, respectively. In steps S411, S413, S415, and S416, the control unit 11 executes the same processes as in steps S310, S312, S314, and S315.
[0077] In step S412, the control unit 11 instructs the start of pronunciation of the consonant in the last syllable of the identified syllables. At this time, the control unit 11 ends the pronunciation started in step S407. For example, if "ma" and "su" are a pair in the collective pronunciation setting, the control unit 11 instructs the end of pronunciation of [ma] and the start of pronunciation of the consonant [s] in "su" (time point T2). Thereafter, the control unit 11 ends the processing shown in FIG. 7.
[0078] In step S414, an instruction is given to end the pronunciation of the consonant in the last syllable among the identified syllables. For example, if "ma" and "su" are a pair in the collective pronunciation setting, the control unit 11 gives an instruction to end the pronunciation of the consonant [s] in "su" (time T3).
[0079] According to this embodiment, note-ons and note-offs are determined based on the acquired performance signal (performance information), and the syllable corresponding to the timing of the determined note-on is identified from the lyric data. The control unit 11 (instruction unit 36) instructs the system to start producing the identified syllable at the timing corresponding to the note-on, and also instructs the system to produce some of the phonemes that make up the identified syllable at the timing corresponding to the note-off. This makes it possible to produce syllables in accordance with the performer's intention.
[0080] In particular, when the language is English, if the identified syllable has a consonant at the end, the control unit 11 instructs the start of pronunciation from the first phoneme at a timing corresponding to note-on. Furthermore, the control unit 11 instructs the pronunciation of the remaining phonemes, including the final consonant, at a timing corresponding to note-off. Therefore, the final consonant can also be pronounced by operation 1.
[0081] Furthermore, the control unit 11 instructs the control unit 11 to start producing the remaining phonemes in response to the performance depth newly passing (crossing) the first threshold value THA to the shallower side. Furthermore, the control unit 11 instructs the control unit 11 to end producing the final consonants in the remaining phonemes in response to the performance depth newly passing the second threshold value THB to the shallower side. Therefore, the production duration of the consonants can be adjusted by the performance operation.
[0082] Furthermore, when the language is Japanese and multiple syllables (such as "masu" and "su") specified for one note-on are set as targets for collective pronunciation setting, the control is performed as follows: The control unit 11 instructs the start of pronunciation from the first phoneme of the first syllable among the specified syllables at the timing corresponding to the note-on, and instructs the start of pronunciation of the consonant of the last syllable at the timing corresponding to the note-off. Therefore, even for Japanese lyrics, the final consonant can be pronounced with operation 1, and the pronunciation length of the consonant can be adjusted by performance operations, making it possible to pronounce syllables according to the performer's intentions.
[0083] In addition to "mas," the "special syllables" that are the target of the processing in FIG. 7 include "teeth," "make," "rice," "fast," and "desks."
[0084] One syllable may contain two vowels. For a "special syllable" having two vowels, the control unit 11 may instruct in step S307 to start pronunciation from the first phoneme of the identified syllable so that the first of the two vowels is included. In this case, in step S311, the control unit 11 may instruct to pronounce the second vowel and the final consonant as the remaining phonemes.
[0085] For example, in the case of "make," the phonemes excluding "some phonemes including a final consonant" in step S307 correspond to [me], and the "some phonemes including a final consonant" in step S311 correspond to [i] and [k]. Therefore, the pronunciation of [me] starts at time T1, and ends at time T2, and the pronunciation of [i] starts. The pronunciation of [i] ends at time T3, and [k] is pronounced for a certain period of time. Alternatively, after [i] is pronounced for a certain period of time at time T2, the pronunciation of [k] may start, and the pronunciation of [k] may end at time T3.
[0086] In addition, a third threshold may be set in addition to the thresholds THA and THB as a threshold for silencing control, so that the pronunciation of [i] starts at the first threshold THA, the pronunciation of [i] ends and the pronunciation of [k] starts at the second threshold THB, and the pronunciation of [k] ends at the third threshold.
[0087] In the case of "rice," which has two vowels, the phonemes that correspond to it are [ra] except for "some phonemes that include a final consonant," and the phonemes that correspond to it are [i] and [s].
[0088] Note that some syllables have two or more consonant phonemes. For example, in the case of "fast," [fa] corresponds to the phonemes excluding "some phonemes including final consonants," and [s] and [t] correspond to "some phonemes including final consonants." Regarding [s] and [t], the pronunciation of [s] begins at time T2. The pronunciation of [s] ends at time T3, and [t] is pronounced for a certain period of time. Note that after [s] is pronounced for a certain period of time at time T2, the pronunciation of [t] may begin, and the pronunciation of [t] may end at time T3.
[0089] In addition, a third threshold may be set in addition to the thresholds THA and THB as a threshold for silencing control, so that the pronunciation of [s] starts at the first threshold THA, the pronunciation of [s] ends and the pronunciation of [t] starts at the second threshold THB, and the pronunciation of [t] ends at the third threshold.
[0090] For a syllable with three or more consonant phonemes (for example, "desks"), four thresholds may be set to determine the timing for starting and ending the pronunciation of each consonant phoneme.
[0091] In this embodiment, the number of threshold values for mute control may be one. In this case, for example, the pronunciation length of a consonant phoneme may be a fixed value.
[0092] (Second embodiment) The second embodiment of the present invention differs from the first embodiment in the sound control processing. The English language support processing in this embodiment will be mainly described with reference to Figures 9 and 10 instead of Figures 4 and 7.
[0093] 9 is a timing chart showing an example of sound control in response to a performance signal in the second embodiment of the present invention, and FIG. 10 is a flowchart showing the English language support process executed in step S205 of FIG.
[0094] In the first embodiment, the time period from T2 to T3 substantially corresponds to the note-off velocity. In contrast, in the present embodiment, the pronunciation duration of "some phonemes including a final consonant" is determined based on the actually acquired note-off velocity.
[0095] The significance of times T11, T12, and T13 shown in FIG. 9 is the same as that of times T1, T2, and T3 shown in FIG. 4. The definitions of "special syllable" and "non-special syllable" are also the same as in the first embodiment. The thresholds TH0, THA, and THB may be the same as in the first embodiment, but the individual values may be set differently. As in the first embodiment, the control unit 11 identifies a syllable to be pronounced at time T11 and starts pronouncing the syllable.
[0096] The instruction unit 36 acquires the note-off velocity from the time period from time T12 to time T13. The instruction unit 36 determines the pronunciation length of the final consonant in the remaining phonemes ("some phonemes including a final consonant") according to the acquired note-off velocity. The determined pronunciation length is the length from time T13 to T14. For example, the pronunciation length is shorter the faster the note-off velocity. In other words, the shorter the length from time T12 to T13, the shorter the pronunciation length. At time T13, the instruction unit 36 starts pronouncing some phonemes including the final consonant for the determined pronunciation length (start of pronouncing consonants, etc.).
[0097] For example, for the special syllable "mas," the control unit 11 starts pronunciation of the first phoneme [ma] of the syllable "mas" at time T11. Then, the control unit 11 finishes pronouncing [ma] at time T13 and starts pronouncing the remaining phonemes [s] including the final consonant, and further finishes pronouncing [s] at time T14. Therefore, the pronunciation duration of [ma] is from time T11 to T13, and the pronunciation duration of [s] (pronunciation period of consonants, etc.) is from time T13 to T14.
[0098] The process of FIG. 10 will be described. First, in steps S501 to S509, S514, and S515, the control unit 11 executes the same processes as steps S301 to S309, S316, and S317 of FIG. 7. In step S510, the control unit 11 starts acquiring the note-off velocity. Specifically, the control unit 11 continues to monitor the performance depth. The control unit 11 then acquires time T12 in response to determining that the performance depth has newly crossed the first threshold value THA to the shallower side, and further acquires time T13 in response to determining that the performance depth has newly crossed the second threshold value THB to the shallower side. After acquiring time T13, the control unit 11 acquires the note-off velocity from the time difference between time T13 and time T12. After step S510, the control unit 11 ends the process shown in FIG. 10.
[0099] If the result of the determination in step S501 is that flag F = 1, the control unit 11 proceeds to step S511. In step S511, the control unit 11 determines whether the note-off velocity has already been acquired and a new note-off has occurred (i.e., whether the playing depth has newly crossed the second threshold value THB to the shallower side).
[0100] In this embodiment, only two thresholds for mute control are provided: a first threshold THA and a second threshold THB. Therefore, in step S511, if the playing depth crosses the second threshold THB to the shallower side, a corresponding note-off velocity is acquired, and the answer is Yes.
[0101] If it is determined in step S511 that the note-off velocity has not been acquired or that a new note-off has not occurred, the control unit 11 ends the process shown in Fig. 10. On the other hand, if it is determined that the note-off velocity has been acquired and a new note-off has occurred, the control unit 11 proceeds to step S512.
[0102] In step S512, the control unit 11 determines the pronunciation period (pronunciation length) of the final consonant in the remaining phoneme according to the acquired note-off velocity. Furthermore, the control unit 11 specifies the determined pronunciation period and instructs to start pronouncing "some phonemes including the final consonant." At this time, the control unit 11 ends the pronunciation started in step S507.
[0103] For example, if the identified syllable is "mas," the control unit 11 ends the pronunciation of [ma] at time T13, and then designates the period from time T13 to time T14 as a pronunciation period, and starts the pronunciation of the remaining phoneme, [s], including the final consonant. Therefore, the pronunciation of [s] ends at time T14.
[0104] In step S513, the control unit 11 executes the same process as in step S315.
[0105] Note that three or more thresholds for mute control may be provided. In this case, two of the thresholds may be used to obtain the note-off velocity, and one of the thresholds (a predetermined threshold) may be used to determine the occurrence of a new note-off. For example, the control unit 11 may obtain the note-off velocity from the time difference at which the performance depth crosses two deeper thresholds, and may instruct the controller 11 to start producing the remaining phonemes when the performance depth has newly crossed a predetermined threshold (for example, the shallowest threshold) to the shallower side.
[0106] This embodiment can achieve the same effect as the first embodiment in that it allows the player to pronounce syllables in accordance with their intentions. Furthermore, the note-off velocity is acquired based on the performance signal, and the duration of the final consonants in the remaining phonemes is determined based on the acquired note-off velocity. Therefore, the duration of the final consonants can be determined before detecting the timing to start pronouncing the final consonants, thereby reducing the processing load when starting to pronounce the consonants.
[0107] This embodiment can also be applied to Japanese language support processing.
[0108] In each of the above embodiments, the volume may be determined by the note-on velocity. In this case, two or more sound generation thresholds may be provided to determine the note-on velocity.
[0109] When the instruction unit 36 pronounces some of the phonemes that make up the identified syllable, it is not essential that the phonemes pronounced at the timing corresponding to the note-off include a consonant. Conventionally, little consideration has been given to controlling the pronunciation of syllables in response to the note-off. Therefore, even if the phonemes pronounced at the timing corresponding to the note-off do not include a consonant, some of the phonemes that make up the identified syllable may be pronounced at the timing corresponding to the note-off. By doing so, the effect of pronouncing syllables according to the performer's intention can be achieved.
[0110] The "performance depth" indicated by the performance signal differs depending on the instrument. The sound control device 100 is not limited to a wind instrument type, and may be of other forms, such as a keyboard instrument. For example, when the present invention is applied to a keyboard instrument, a key sensor may be provided to detect the stroke position of each key, and the passage of positions corresponding to thresholds TH0, THA, and THB may be detected. The configuration of the key sensor is not important, and pressure-sensitive sensors, optical sensors, etc., may be used. In the case of a keyboard instrument, the key position in the non-operated state is "0," and the deeper the key is depressed, the deeper the "performance depth" becomes.
[0111] The sound control device 100 does not necessarily have to have the functions and form of a musical instrument, and may be a device that can detect pressing operations, such as a touchpad. Furthermore, the present invention can also be applied to devices that can obtain "performance depth" by detecting the strength of operations on an on-screen control, such as a smartphone.
[0112] The performance signal (performance information) may be acquired from an external device via communication, so providing the performance operation unit 15 is not essential.
[0113] In each of the above embodiments, at least a part of the functional units shown in FIG. 3 may be realized by AI (Artificial Intelligence).
[0114] Although the present invention has been described in detail above based on preferred embodiments thereof, the present invention is not limited to these specific embodiments, and various forms within the scope of the gist of the present invention are also included in the present invention. Parts of the above-described embodiments may be combined as appropriate.
[0115] A storage medium storing a control program represented by software for achieving the present invention may be read into the device to achieve the same effect as the present invention. In this case, the program code read from the storage medium itself realizes the novel functions of the present invention, and the non-transitory computer-readable storage medium storing the program code constitutes the present invention. The program code may also be supplied via a transmission medium, in which case the program code itself constitutes the present invention. In these cases, storage media such as ROMs, floppy disks, hard disks, optical disks, magneto-optical disks, CD-ROMs, CD-Rs, magnetic tape, and non-volatile memory cards can be used. Non-transitory computer-readable storage media also include volatile memory (e.g., DRAM (Dynamic Random Access Memory)) within computer systems that serve as servers or clients when a program is transmitted via a network such as the Internet or a communication line such as a telephone line, which retains the program for a certain period of time. [Explanation of symbols]
[0116] 11 control unit, 31 acquisition unit, 32 determination unit, 34 identification unit, 36 instruction unit
Claims
1. an acquisition unit for acquiring performance information; a determination unit that determines a note-off based on the performance information; an identification unit that identifies a part of a phoneme that is to be pronounced corresponding to the timing at which the determination unit determines that a note is off, from lyric data in which a plurality of syllables to be pronounced are arranged in time series; and an instruction unit that instructs the device to start producing sound of the part of the phoneme identified by the identification unit at a timing corresponding to the note-off.
2. The sound control device according to claim 1 , wherein the part of the phonemes is a phoneme including a consonant located at the end of the syllable.
3. The sound control device according to claim 2 , wherein some of the phonemes are phonemes that do not include vowels and are made up of only consonants.
4. 3. The sound control device according to claim 2, wherein the instruction unit instructs the start of pronunciation of a part of the phoneme in response to the performance depth indicated by the performance information passing a first threshold to the shallower side, and instructs the end of pronunciation of the final consonant in the part of the phoneme in response to the performance depth indicated by the performance information passing a second threshold to the shallower side, the second threshold corresponding to a performance depth shallower than the first threshold.
5. The sound control device according to claim 2 , wherein the instruction unit acquires a velocity of the note-off based on the performance information, and determines a pronunciation length of the final consonant in the part of the phoneme according to the acquired velocity.
6. The sound control device described in claim 5, wherein the instruction unit acquires the velocity using a plurality of thresholds that are compared with the performance depth indicated by the performance information, and instructs the start of sound production of a part of the phoneme when the performance depth indicated by the performance information passes a predetermined threshold among the plurality of thresholds to the shallower side.
7. The sound control device of claim 1 , wherein the lyrics data includes lyrics in English.
8. the lyrics data includes Japanese lyrics, The sound control device according to claim 1 , wherein the part of the phonemes is a consonant corresponding to a final syllable of a plurality of syllables.
9. A sound control device according to any one of claims 1 to 8; a performance operation unit for a user to input the performance information.
10. the performance operation unit includes a breath sensor that detects pressure changes; 10. The electronic musical instrument according to claim 9, wherein the performance information is acquired based on a pressure change detected by the breath sensor.
11. A program for causing a computer to execute a control method for a sound control device, The control method of the sound control device includes: Get performance information, determining a note-off based on the performance information; Identifying a part of the phoneme that is pronounced corresponding to the timing determined as the note-off from lyric data in which a plurality of syllables to be pronounced are arranged in time series; a program that instructs the part of the identified phonemes to start sounding at a timing corresponding to the note-off;
12. A control method for a sound control device implemented by a computer, comprising: Get performance information, determining a note-off based on the performance information; Identifying a part of the phoneme that is pronounced corresponding to the timing determined as the note-off from lyric data in which a plurality of syllables to be pronounced are arranged in time series; A control method for a sound control device that instructs a sound control device to start producing a part of the identified phonemes at a timing corresponding to the note-off.
Citation Information
Patent Citations
Voice synthesizing apparatus
JP2014098801A
Controller, synthetic singing sound creation device and program
JP2016206496A
Electronic musical instrument, method and program
JP7036141B2