Information processing device, electronic musical instrument system, electronic musical instrument, method and program

The information processing device in electronic musical instruments controls syllable progression based on the number of active operators, addressing the issue of inappropriate syllable advancement and ensuring harmonious voice output.

JP7838595B2Active Publication Date: 2026-04-01CASIO COMPUTER CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-06-19
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

Existing electronic musical instruments may advance lyrics too far ahead of the user's desired progression when syllables are triggered by key presses, leading to inappropriate syllable progression in synthesized voices.

Method used

An information processing device with a control unit that controls syllable progression based on the number of operators currently being operated on, preventing advancement to the next syllable if a set number is reached or allowing progression if fewer operators are active, and ensuring syllables align with desired harmonic changes.

Benefits of technology

This solution allows for appropriate control of syllable progression, maintaining vowel sounds in melody parts while allowing other parts to change pitch, thus enhancing the flexibility and accuracy of synthesized voice output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007838595000001
    Figure 0007838595000001
  • Figure 0007838595000002
    Figure 0007838595000002
  • Figure 0007838595000003
    Figure 0007838595000003
Patent Text Reader

Abstract

To appropriately control syllable progression when reproducing harmony in, for example, a chorus group, based on operation of electronic instruments.SOLUTION: When operation to a second operator is detected after a set time elapses from detection of operation to a first operator, a CPU of an electronic instrument controls whether to advance the syllable to be pronounced from a first syllable to the next second syllable or not depending on the number of operators that are still being operated at timing when operation to the second operator is detected.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an electronic musical instrument system, an electronic musical instrument, a method, and a program.

Background Art

[0002] In recent years, the usage scenes of synthesized voices have been expanding. Among these, it is preferable to have an electronic musical instrument that can advance lyrics according to a key-pressing operation by a user (performer) and output a synthesized voice corresponding to the lyrics, in addition to automatic performance, so that more flexible expression of synthesized voices becomes possible.

[0003] For example, Patent Document 1 discloses a technique for advancing lyrics in synchronization with performance based on user operations using a keyboard or the like.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] By the way , press If the syllables of the lyrics are advanced for each key, There is a possibility that the lyrics may progress too far ahead of the user's desired progression.

[0006] The present invention has been made in view of the above problems, and aims to appropriately control the measure progression based on the operation of an electronic musical instrument. te sound

Means for Solving the Problems

[0007] ​To solve the above problems, the information processing device of the present invention includes a control unit that, when an operation on a second operator specifying a second pitch is detected after a set time has elapsed since an operation on a first operator specifying a first pitch was detected, obtains the number of pitch-specifying operators that are currently being operated on based on the detection of the operation on the second operator, controls the system so that the syllable to be pronounced in response to the detection of the operation on the second operator does not advance from the first syllable to the next second syllable if the number of obtained operators reaches a set number, and controls the system so that the syllable to be pronounced in response to the detection of the operation on the second operator advances from the first syllable to the next second syllable if the number of obtained operators is less than the set number. Each syllable to be pronounced includes multiple frames to be pronounced sequentially, and the control unit determines, when the number of operators reaches a set number, whether the position of the next frame to be pronounced exceeds the vowel ending position of the first syllable which is the current target of pronunciation, and if it does not, it starts pronouncing from the next frame position within the first syllable, and if it does exceed it, it controls the system so that it cannot proceed from the first syllable to the second syllable. . [Effects of the Invention]

[0008] According to the present invention, based on the operation of an electronic musical instrument te sound This makes it possible to appropriately control the progression of the clauses. [Brief explanation of the drawing]

[0009] [Figure 1] This figure shows an example of the overall configuration of the electronic musical instrument system of the present invention. [Figure 2] This figure shows the external appearance of the electronic musical instrument shown in Figure 1. [Figure 3] Figure 1 is a block diagram showing the functional configuration of an electronic musical instrument. [Figure 4] Figure 1 is a block diagram showing the functional configuration of the terminal device. [Figure 5] This figure shows the configuration related to the production of singing sounds in response to key presses in the singing sound production mode of the electronic musical instrument shown in Figure 1. [Figure 6] This is an image illustrating the relationship between frames and syllables. [Figure 7] Figure 3 is a flowchart showing the flow of the pronunciation control process executed by the CPU. [Figure 8] Figure 3 is a flowchart showing the flow of syllable progression control processing executed by the CPU. [Figure 9] This figure shows an example of syllable progression using the syllable progression control process shown in Figure 8. [Modes for carrying out the invention]

[0010] The embodiments for carrying out the present invention will be described below with reference to the drawings. However, the embodiments described below are subject to various technically preferred limitations for carrying out the present invention. Therefore, the technical scope of the present invention is not limited to the embodiments and illustrated examples below.

[0011] [Configuration of Electronic Musical Instrument System 1] Figure 1 shows an example of the overall configuration of the electronic musical instrument system 1 according to the present invention. As shown in Figure 1, the electronic musical instrument system 1 is configured by connecting the electronic musical instrument 2 and the terminal device 3 via a communication interface I (or communication network N).

[0012] [Configuration of Electronic Instrument 2] In addition to a normal mode in which the instrument sounds are output in response to the user's key presses on the keyboard 101, the electronic instrument 2 also has a singing voice mode in which it produces singing voices in response to the key presses on the keyboard 101, making it possible to produce polyphonic sounds of harmonies consisting of multiple parts, such as those of a chorus.

[0013] Figure 2 shows an example of the appearance of the electronic instrument 2. The electronic instrument 2 includes a keyboard 101 consisting of multiple keys as controls, a first switch panel 102 and a second switch panel 103 for instructing various settings, and an LCD 104 (Liquid Crystal Display) for displaying various information. The electronic instrument 2 also includes a speaker 214 on the back, side, or rear to emit musical sounds and voices (singing voices) generated by playing.

[0014] FIG. 3 is a block diagram showing the functional configuration of the control system of the electronic musical instrument 2 in FIG. 1. As shown in FIG. 3, the electronic musical instrument 2 includes a CPU (Central Processing Unit) 201 connected to a timer 210, a ROM (Read Only Memory) 202, a RAM (Random Access Memory) 203, a sound source unit 204, a voice synthesis unit 205, a keyboard 101 in FIG. 2, a first switch panel 102, and a key scanner 206 to which the second switch panel 103 is connected, an LCD controller 207 to which the LCD 104 in FIG. 2 is connected, and a communication unit 208, which are respectively connected to a bus 209. In the present embodiment, the first switch panel 102 includes a singing sound generation mode switch described later. Also, the second switch panel 103 includes a timbre setting switch described later. In addition, D / A converters 211 and 212 are respectively connected to the sound source unit 204 and the voice synthesis unit 205. The waveform data of the musical instrument sound output from the sound source unit 204 and the voice waveform data of the singing voice (singing voice waveform data) output from the voice synthesis unit 205 are respectively converted into analog signals by the D / A converters 211 and 212, amplified by an amplifier 213, and then output from a speaker 214.

[0015] The CPU 201 executes the control operation of the electronic musical instrument 2 in FIG. 1 by executing the program stored in the ROM 202 while using the RAM 203 as a work memory. The CPU 201 realizes the function of the control unit of the information processing apparatus of the present invention by executing the tone generation control processing and the syllable progression control processing described later in cooperation with the program stored in the ROM 202. The ROM 202 stores programs and various fixed data and the like.

[0016] The sound source unit 204 has a waveform ROM that stores waveform data of various timbres, such as human voices, dog barks, and cat barks, as well as waveform data for vocalization sources in singing mode (instrument sound waveform data) for instrument sounds such as piano, organ, synthesizer, string instruments, and wind instruments. Instrument sound waveform data can also be used as vocalization waveform data.

[0017] In normal mode, the sound source unit 204 reads instrument sound waveform data from a waveform ROM (not shown) based on the pitch information of the keys pressed on the keyboard 101, according to control instructions from the CPU 201, and outputs it to the D / A converter 211. In singing voice generation mode, the sound source unit 204 reads waveform data from a waveform ROM (not shown) based on the pitch information of the keys pressed on the keyboard 101, according to control instructions from the CPU 201, and outputs it to the speech synthesis unit 205 as waveform data for vocalization. The sound source unit 204 can output waveform data for multiple channels simultaneously. Alternatively, waveform data corresponding to the pitch of the keys pressed on the keyboard 101 may be generated based on the pitch information and the waveform data stored in the waveform ROM. The sound source unit 204 is not limited to the PCM (Pulse Code Modulation) sound source method, but may use other sound source methods, such as the FM (Frequency Modulation) sound source method.

[0018] The speech synthesis unit 205 has a synthesis filter 205a and generates singing waveform data based on singing parameters provided by the CPU 201 and waveform data for the vocalization source input from the sound source unit 204, and outputs it to the D / A converter 212.

[0019] The sound source unit 204 and the speech synthesis unit 205 may be configured using dedicated hardware such as an LSI (Large-Scale Integration), or they may be realized by software through the cooperation of the CPU 201 and a program stored in the ROM 202.

[0020] The key scanner 206 continuously scans the pressed / released state of each key on the keyboard 101 in Figure 2, the switch operation state of the first switch panel 102 and the second switch panel 103, and outputs the pitch of the operated key, pressed / released information (performance operation information), and switch operation information to the CPU 201.

[0021] The LCD controller 207 is an integrated circuit (IC) that controls the display state of the LCD 104.

[0022] The communication unit 208 connects to communication networks such as the Internet and USB (Universal Serial). Data is transmitted and received with external devices such as terminal devices 3 connected via a communication interface I such as a bus cable.

[0023] [Configuration of Terminal Device 3] Figure 4 is a block diagram showing the functional configuration of terminal device 3 in Figure 1. As shown in Figure 4, the terminal device 3 is a computer comprising a CPU 301, ROM 302, RAM 303, storage unit 304, operation unit 305, display unit 306, communication unit 307, etc., and each unit is connected by a bus 308. Examples of terminal devices 3 include tablet PCs (Personal Computers), notebook PCs, and smartphones.

[0024] The ROM 302 of terminal device 3 is equipped with a pre-trained model 302a. The pre-trained model 302a is generated by machine learning on multiple datasets consisting of musical score data (lyrics data (text information of lyrics) and pitch data (including information on note length)) of multiple songs, and vocal waveform data when a certain singer sings each song. When the pre-trained model 302a receives lyric data and pitch data of any song (even a phrase), it infers a set of vocal parameters (called vocal information) to produce a vocal sound equivalent to that of the singer who generated the pre-trained model 302a when singing the input song.

[0025] [Singing voice pronunciation mode operation] Figure 5 shows the configuration related to the production of singing voices in response to key presses on the keyboard 101 in singing voice production mode. The operation of the electronic instrument 2 in producing singing voices in response to key presses on the keyboard 101 in singing voice production mode will be explained below with reference to Figure 5.

[0026] If the user wishes to perform in vocal mode, they press the vocal mode switch on the first switch panel 102 of the electronic instrument 2 to initiate the transition to vocal mode. When the singing voice mode switch is pressed, the CPU 201 switches to singing voice mode. Also, when the user selects the desired voice tone using the tone selection switch on the second switch panel 103, the CPU 201 sets the information of the selected tone in the sound source unit 204.

[0027] Next, the user inputs the lyrics data and pitch data of any song they want the electronic instrument 2 to pronounce in singing mode into the terminal device 3 using a dedicated application or the like. Alternatively, the lyrics data and pitch data of songs can be stored in the storage unit 304, and the user can select the lyrics data and pitch data of any song from those stored in the storage unit 304. In terminal device 3, when lyrics data and pitch data of any song to be played in singing mode are input, CPU 301 inputs the lyrics data and pitch data of the input song into the trained model 302a, causes the trained model 302a to infer a group of singing parameters, and transmits the singing information, which is the inferred group of singing parameters, to the electronic instrument 2 via the communication unit 307.

[0028] Now, let's discuss the vocal information. Each section of a song, divided into predetermined time units in the time direction, is called a frame, and the trained model 302a generates vocal parameters on a frame-by-frame basis. That is, the vocal information of one song is composed of multiple vocal parameters (vocal parameter group) on a frame-by-frame basis. In this embodiment, one frame is defined as the length of one sample multiplied by 225 when the song is sampled at a predetermined sampling frequency (e.g., 44.1 kHz).

[0029] Frame-by-frame vocal parameters include spectral parameters (frequency spectrum of the voice being produced) and fundamental frequency F0 parameters (pitch frequency of the voice being produced).

[0030] Additionally, the frame-by-frame singing parameters include syllable information. Figure 6 is an illustrative diagram showing the relationship between frames and syllables (Note that Figure 6 does not use any registered trademarks). As shown in Figure 6, the audio of a song is composed of multiple syllables (the first to third syllables in Figure 6). Each syllable is generally composed of one vowel, or a combination of one vowel and one or more consonants. Each syllable is pronounced over multiple consecutive frame intervals in the time direction, and the syllable start position, syllable end position, vowel start position, and vowel end position (all positions in the time direction) of each syllable included in a song can be identified by the frame position (which frame it is from the beginning). In the singing information, the singing parameters of the frames corresponding to the syllable start position, syllable end position, vowel start position, and vowel end position of each syllable include information such as the nth syllable start frame, the nth syllable end frame, the nth vowel start frame, and the nth vowel end frame (where n is a natural number).

[0031] Returning to Figure 5, when the electronic instrument 2 receives singing information from the terminal device 3 via the communication unit 208, the CPU 201 stores the received singing information in the RAM 203. When the user operates the keyboard 101 and performance operation information is input from the key scanner 206, the CPU 201 inputs the pitch information of the pressed key to the sound source unit 204. The sound source unit 204 reads waveform data corresponding to the input pitch information of a preset timbre from the waveform ROM as waveform data for the voice source and inputs it to the synthesis filter 205a of the speech synthesis unit 205. Furthermore, when performance operation information is input from the key scanner 206, the CPU 201 executes the syllable progression control process (see Figure 8), which will be described later, to identify the frame to be sounded according to the performance operation, reads the spectral parameters of the identified frame from the RAM 203, and inputs them to the synthesis filter 205a.

[0032] The synthesis filter 205a generates singing waveform data based on the input spectral parameters and waveform data for the vocal sound source, and outputs it to the D / A converter 212. The singing waveform data output to the D / A converter 212 is converted into an analog audio signal, amplified by the amplifier 213, and output from the speaker 214.

[0033] In choral harmony, the soprano and other melody parts often maintain their vowels without changing pitch, while the alto and bass parts change pitch using melisma. However, if the syllables of the lyrics advance with each key press, it becomes impossible to reproduce such harmonic changes.

[0034] Therefore, in the singing voice pronunciation mode, the CPU 201 controls the syllable progression to be appropriate when reproducing harmonies such as those of a chorus, by executing a pronunciation control process, including the syllable progression control process shown in Figure 8, in response to the input of performance operation information from the key scanner 206.

[0035] Figure 7 is a flowchart showing the flow of the pronunciation control process. The pronunciation control process is executed, for example, when the communication unit 208 stores the singing voice information received from the terminal device 3 in the RAM 203, through the cooperation of the CPU 201 and the program stored in the ROM 202.

[0036] First, CPU201 initializes the variables used in the syllable progression control process (step S1). Next, the CPU 201 determines whether or not performance operation information has been input by the key scanner 206 (step S2). If it is determined that performance operation information has been input (step S2; YES), the CPU 201 executes syllable progression control processing (step S3).

[0037] Figure 8 is a flowchart showing the flow of syllable progression control processing. Syllable progression control processing is performed through the cooperation of the CPU 201 and the program stored in ROM 202.

[0038] In the syllable progression control process, the CPU 201 detects a key press or release operation based on the performance operation information input from the key scanner 206 (step S31). If a key press operation is detected (step S31; YES), the CPU 201 sets KeyOnCounter to KeyOnCounter+1 (step S32). Here, KeyOnCounter is a variable that stores the number of keys currently being pressed (the number of operators whose operation is ongoing).

[0039] Next, CPU201 determines whether KeyOnCounter is 1 or not (step S33). In other words, it determines whether the detected key press was performed while no other keys were pressed.

[0040] If it is determined that KeyOnCounter is 1 (step S33; YES), CPU201 obtains SystemTime, sets the obtained SystemTime to FirstKeyOnTime (step S34), and proceeds to step S37. Here, FirstKeyOnTime is a variable that stores the time when the first key pressed (the first operator) among the currently pressed keys was pressed. In other words, when CPU201 determines that KeyOnCounter is 1, it determines that it has detected an operation on the first operator (referred to as the first key press) and sets FirstKeyOnTime.

[0041] If it is determined that KeyOnCounter is not 1 (step S33; NO), CPU201 obtains SystemTime and determines whether SystemTime - FirstKeyOnTime > M (step S35). Here, M is a pre-set simultaneous determination period (approximately a few milliseconds; corresponding to the set time in this invention) used to determine whether the detected key press (operation on the second operator) was performed approximately simultaneously with the first key press. If SystemTime - FirstKeyOnTime > M is not true (i.e., the elapsed time since the first key press is within the simultaneous determination period), the detected key press is considered to be a simultaneous key press with the first key press. If SystemTime - FirstKeyOnTime > M (i.e., the elapsed time since the first key press is outside the simultaneous determination period), the detected key press is not considered to be a simultaneous key press with the first key press.

[0042] If it is determined that SystemTime - FirstKeyOnTime is not greater than M (i.e., it is within the simultaneous determination period) (step S35; NO), CPU201 proceeds to step S41. In this case, the keys pressed that result in a NO judgment in step S35 are the first key press and simultaneous key presses. In the case of multiple simultaneous key presses, the control is such that the entire sequence, including the first key press, advances to one syllable. In this embodiment, since the first key press advances the syllable, for other simultaneous key presses, the control proceeds to step S41 to prevent the syllable from advancing.

[0043] If SystemTime - FirstKeyOnTime > M (i.e., outside the simultaneous determination period) (step S35; YES), the CPU 201 determines whether KeyOnCounter < 4, i.e., whether the number of keys currently being pressed is less than 4 (step S36). Here, the number of settings compared with KeyOnCounter in step S36 (4 in this case) is the number of parts to be played in singing mode. In this embodiment, four parts—soprano, alto, tenor, and bass—are played in singing mode, so the number of settings compared with KeyOnCounter in step S36 is set to 4. Note that this number of settings can be changed according to user operation.

[0044] If it is determined that KeyOnCounter < 4 (step S36; YES), that is, if it is determined that the number of keys currently pressed is less than the number of parts, the CPU 201 proceeds to step S37.

[0045] If it is determined that KeyOnCounter < 4 (step S36; NO), that is, if it is determined that the number of keys currently being pressed has reached the number of parts, the CPU 201 proceeds to step S41.

[0046] In step S37, the CPU 201 determines whether CurrentFramePos is the frame position of the last syllable (step S37). CurrentFramePos is a variable that stores the frame position of the currently sounded frame, and until it is replaced with the frame position of the next sounded frame in step S43 or S44, it stores the frame position of the previously sounded frame.

[0047] If the CPU determines that CurrentFramePos is the frame position of the last syllable (step S37; YES), the CPU 201 sets NextFramePos, a variable that stores the frame position of the next frame to be pronounced, to the syllable start position of the first syllable (step S38), and proceeds to step S43.

[0048] If it is determined that CurrentFramePos is not the frame position of the last syllable (step S37; NO), the CPU 201 sets NextFramePos to the syllable start position of the next syllable (step S39) and proceeds to step S43.

[0049] In step S43, the CPU 201 sets CurrentFramePos to NextFramePos (step S43) and proceeds to step S4 in Figure 7. In other words, if the previously pronounced frame is not the last syllable, the position of the frame being pronounced advances to the syllable start position of the next syllable. If the previously pronounced frame is the last syllable, there is no syllable following the previously pronounced syllable, so the position of the frame being pronounced advances to the frame at the start position of the first syllable.

[0050] On the other hand, if it is determined in step S31 that a key release has been detected (step S31; NO), the CPU 201 sets KeyOnCounter to KeyOnCounter - 1 (step S40) and proceeds to step S41.

[0051] In step S41, CPU201 sets NextFramePos to CurrentFramePos + playback rate / 120 (step S41). Here, 120 is the default tempo value, but it is not limited to this. The playback rate is a value set by the user. For example, if the playback rate is set to 240, the position of the next sound will be set to two frames ahead of the current frame position. If the playback rate is set to 60, the position of the next sound will be set to 0.5 frames ahead of the current frame position.

[0052] Next, CPU201 determines whether NextFramePos > vowel end position (step S42). That is, it determines whether the position of the next frame to be pronounced exceeds the vowel end position of the currently pronounced syllable (i.e., the vowel end position of the previously pronounced syllable). If it is determined that NextFramePos is not the vowel end position (step S42; NO), the CPU 201 proceeds to step S43, sets CurrentFramePos to NextFramePos (step S43), and proceeds to step S4 in Figure 7. That is, the frame position of the frame to be pronounced is advanced to NextFramePos, but since NextFramePos is before the vowel end position of the previously pronounced syllable, it does not proceed to the next syllable.

[0053] If it is determined that NextFramePos is equal to the vowel end position (step S42; YES), the CPU 201 sets CurrentFramePos to the vowel end position of the currently pronounced syllable (step S44) and proceeds to step S4 in Figure 7. In other words, the frame position of the frame to be pronounced is set to the vowel end position of the previously pronounced syllable, so it does not proceed to the next syllable.

[0054] Figure 9 schematically illustrates the syllable control process performed by the syllable progression control process described above. In Figure 9, the black inverted triangles indicate the timing when all keys are released. The KeyOnCounter values ​​represent the KeyOnCounter values ​​at each of the timings T1 to T6. In the performance shown in Figure 9, pressing the keys at timing T1 is a simultaneous press of four parts, so the syllable advances by one. Pressing the keys at timing T2 is outside the simultaneous judgment period, and the number of keys pressed at this timing has reached the number of parts (4), so the syllable does not advance. Pressing the keys at timing T3 is a simultaneous press of four parts, so the syllable advances by one. Pressing the keys at timing T4 is a simultaneous press of four parts, so the syllable advances by one. Pressing the keys at timing T5 is outside the simultaneous judgment period, and the number of keys pressed at this timing has reached the number of parts (4), so the syllable does not advance. Pressing the keys at timing T6 is outside the simultaneous judgment period, and the number of keys pressed simultaneously at this timing is less than the number of parts (4), so the syllable advances by one.

[0055] Thus, according to the syllable progression control process described above, even if a key press is detected, if it is a key press outside the simultaneous judgment period (i.e., not the first key press or a key press simultaneously with the first key press), and the number of keys pressed at the time of this key press has reached the number of parts, the syllable to be pronounced will not progress to the next syllable. Therefore, in cases where the melody part (soprano) maintains its vowel without changing its pitch, while only the alto and bass parts change their pitch with melismas, it is possible to prevent the syllables of the lyrics from progressing, and the syllable progression when reproducing harmony can be appropriately controlled.

[0056] In step S4 of Figure 7, the CPU 201 determines whether the operation detected based on the performance operation information input in step S1 is a key press operation (step S4).

[0057] If the detected operation is determined to be a key press operation (step S4; YES), the CPU 201 performs a sound generation process to generate sound for the frame at the frame position stored in CurrentFramePos (step S5), and then proceeds to step S7.

[0058] In step S5, the CPU 201 causes the speech synthesis unit 205 to synthesize and output singing voice based on the pitch information of the key that was detected to have been pressed and the spectral parameters of the frame at the frame position stored in CurrentFramePos. Specifically, the CPU 201 inputs pitch information of keys pressed on the keyboard 101 and keys currently being pressed to the sound source unit 204. The sound source unit 204 reads waveform data corresponding to the input pitch information for a preset timbre from the waveform ROM and inputs it as waveform data for the voice synthesis unit 205 to the synthesis filter 205a. The CPU 201 also obtains the spectral parameters of the frame at the frame position stored in CurrentFramePos from the singing voice information stored in RAM 203 and inputs them to the synthesis filter 205a. The synthesis filter 205a then generates singing voice waveform data based on the input spectral parameters and the waveform data for the voice synthesis unit. The generated singing voice waveform data is converted into an analog audio signal by the D / A converter 212 and output (sounds) via the amplifier 213 and speaker 214.

[0059] If the detected operation is determined to be a key release operation (step S4; NO), the CPU 201 performs a sound muting process for the key that was released (step S6) and proceeds to step S7. In step S7, the CPU 201 synthesizes and outputs a singing voice based on the pitch information of the currently pressed key (excluding the key that has been released) and the spectral parameters of the frame at the frame position stored in CurrentFramePos. Specifically, the CPU 201 inputs pitch information of the currently pressed key (excluding the released key) to the sound source unit 204. The sound source unit 204 then inputs waveform data corresponding to the input pitch information and a preset timbre as waveform data for the voice synthesis unit 205 to the synthesis filter 205a. The CPU 201 also obtains the spectral parameters of the frame at the frame position stored in CurrentFramePos from the singing voice information stored in RAM 203 and inputs them to the synthesis filter 205a. The synthesis filter 205a then generates singing voice waveform data based on the input spectral parameters and the waveform data for the voice synthesis unit. The generated singing voice waveform data is converted into an analog audio signal by the D / A converter 212 and output (sounds) via the amplifier 213 and speaker 214.

[0060] In step S7, the CPU 201 determines whether or not the singing voice pronunciation mode has been instructed to end (step S7). For example, if the singing voice mode switch is pressed while singing voice mode is active, the CPU 201 determines that the singing voice mode should be terminated.

[0061] If it is determined that the singing voice pronunciation mode has not been instructed to end (step S7; NO), CPU201 returns to step S2. If the CPU 201 determines that it has been instructed to end the singing voice pronunciation mode (step S7; YES), the CPU 201 terminates the singing voice pronunciation mode.

[0062] As explained above, according to the CPU 201 of the electronic instrument 2, if a key press operation is detected after the simultaneous judgment period has elapsed, the CPU controls whether or not to advance the syllable to be produced from the first syllable (not limited to the first syllable) to the next second syllable, depending on the number of operators that are still operating at the time the key press operation is detected. For example, CPU201 controls the system so that it cannot proceed from the first syllable to the second syllable if the number of operators currently in operation reaches a set number, and controls the system so that it can proceed from the first syllable to the second syllable if the number of operators currently in operation is less than the set number. Therefore, for example, when the melody part maintains its vowels without changing its pitch, while only the alto and bass parts change their pitch with melismas, it is possible to prevent the syllables of the lyrics from progressing, and to appropriately control the syllable progression when reproducing the harmony.

[0063] Furthermore, if no operator is currently operating at the detected timing, the CPU 201 controls the syllable corresponding to the speech to be pronounced to advance from the first syllable to the second syllable. Thus, syllable progression can be appropriately controlled.

[0064] Furthermore, the CPU 201 starts counting the simultaneous judgment period when it detects an operation on any of the operators while no operation has been performed on any of them. Therefore, syllable progression can be appropriately controlled.

[0065] The descriptions in the above embodiments are merely preferred examples of the information processing device, electronic musical instrument, syllable progression control method, and program according to the present invention, and are not limited thereto. For example, in the above embodiment, the information processing device of the present invention was described as being included in the electronic musical instrument 2, but the invention is not limited to this. For example, the functions of the information processing device of the present invention may be provided in an external device (for example, the terminal device 3 described above (PC (Personal Computer), tablet terminal, smartphone, etc.)) connected to the electronic musical instrument 2 via a wired or wireless communication interface. In this case, the information processing device transmits parameters (in this case, spectral parameters) corresponding to the position control of syllables to the electronic musical instrument 2, and the electronic musical instrument 2 produces synthesized speech based on the received parameters.

[0066] Furthermore, although the above embodiment describes the trained model 302a as being provided in the terminal device 3, it may also be provided in the electronic instrument 2. In that case, the trained model 302a may infer singing information based on the lyric data and pitch data input to the electronic instrument 2.

[0067] Furthermore, although the above embodiment was described using the example of an electronic keyboard instrument 2, it is not limited to this, and may be other electronic instruments such as electronic string instruments or electronic wind instruments.

[0068] Furthermore, while the above embodiments disclose examples in which semiconductor memory such as ROM or hard disks are used as computer-readable media for the program according to the present invention, the invention is not limited to these examples. Other computer-readable media that can be used include SSDs and portable recording media such as CD-ROMs. In addition, carrier waves can also be used as a medium for providing the data of the program according to the present invention via a communication line.

[0069] Furthermore, the detailed configuration and operation of the electronic musical instrument, information processing device, and electronic musical instrument system can also be modified as appropriate without departing from the spirit of the invention.

[0070] Although embodiments of the present invention have been described above, the technical scope of the present invention is not limited to the embodiments described above, but is determined based on the claims. Furthermore, equivalent scopes of the present invention that have been modified from the claims but are not related to the essence of the present invention are also included in the technical scope of the present invention. The invention described in the claims initially attached to the application for this patent is listed below. The claim numbers listed below are the same as those in the claims initially attached to the application for this patent. [Note] <Claim 1> An information processing device comprising a control unit that, when an operation on a second operator is detected after a set time has elapsed since an operation on a first operator was detected, controls whether or not to advance the syllable to be pronounced from the first syllable to the next second syllable, depending on the number of operators that are still being operated, at the time the operation on the second operator is detected. <Claim 2> The control unit, If the number of operators currently performing the operation reaches a set number, control is performed to prevent the operation from progressing from the first syllable to the second syllable. If the number of operators currently performing the operation is less than the set number, control is performed to proceed from the first syllable to the second syllable. The information processing apparatus according to claim 1. <Claim 3> The control unit, if there is no operator currently in operation at the detected timing, controls the syllables corresponding to the sound to be pronounced to advance from the first syllable to the second syllable. The information processing apparatus according to claim 1 or 2. <Claim 4> The control unit, when it detects an operation on any of the controls while no operation has been performed on any of the controls, determines that an operation on the first control has been detected and starts counting the set time. The information processing apparatus according to any one of claims 1 to 3. <Claim 5> An information processing device according to any one of claims 1 to 4, Electronic musical instruments, Equipped with, The information processing device transmits parameters corresponding to the syllable position control to the electronic musical instrument. The electronic musical instrument emits a sound synthesized based on the received parameters. Electronic musical instrument system. <Claim 6> An information processing device according to any one of claims 1 to 4, Multiple operators, An electronic musical instrument equipped with [specific features / features]. <Claim 7> The control unit of the information processing device, A method for controlling whether or not to advance the syllable to be pronounced from the first syllable to the next second syllable, depending on the number of operators that are still being operated at the time the operation on the second operator is detected, after a set time has elapsed since the operation on the first operator was detected. <Claim 8> The control unit of the information processing device, If an operation on the second operator is detected after a set time has elapsed since an operation on the first operator was detected, the system controls whether or not to advance the syllable to be pronounced from the first syllable to the next second syllable, depending on the number of operators that are still being operated, at the time the operation on the second operator is detected. A program for executing a process. [Explanation of Symbols]

[0071] 1. Electronic musical instrument system 2 Electronic musical instruments 101 Keyboard 102 First switch panel 103 Second switch panel 104 LCD 201 CPU 202 ROM 203 RAM 204 Sound Source Section 205 Speech Synthesis Unit 205a Composite filter 206 Key Scanners 208 Communications Department Bus 209 210 timers 211 D / A Converter 212 D / A Converters 213 Amplifier 214 speakers 3 Terminal devices 301 CPU 302 ROM 302a Pre-trained model 303 RAM 304 Storage section 305 Operation section 306 Display section 307 Communications Department 308 Bus

Claims

1. If an operation on the first operator specifying the first pitch is detected, and a set time has elapsed since the first operator specifying the first pitch was detected, then an operation on the second operator specifying the second pitch is detected, and based on the detection of the operation on the second operator, the number of pitch-specifying operators that are currently being operated on is obtained. If the number of acquired operators reaches a set number, the system controls the syllable to be pronounced in response to the detection of an operation on the second operator, preventing it from progressing from the first syllable to the next second syllable. If the number of acquired operators is less than the set number, the system controls the syllable to be pronounced in response to the detection of an operation on the second operator to advance from the first syllable to the next second syllable. Equipped with a control unit, Each of the aforementioned syllables to be pronounced includes multiple frames to be pronounced sequentially. The control unit, If the number of operators reaches the set number, it is determined whether the position of the next frame to be pronounced exceeds the vowel end position of the first syllable which is the current target of pronunciation. If it does not exceed this, the next sound will be pronounced from the frame position within the first syllable. If it exceeds this limit, control is implemented to prevent progression from the first syllable to the second syllable. Information processing device.

2. The control unit, When an operation is detected on the first operator while no operation is being performed on any of the other operators, the countdown of the set time begins. The information processing apparatus according to claim 1.

3. The information processing apparatus according to claim 1 or 2, Electronic musical instruments, Equipped with, Electronic musical instrument system.

4. The information processing apparatus according to claim 1 or 2, Multiple operators, An electronic musical instrument equipped with [specific features / features].

5. The control unit of the information processing device, If an operation on the first operator specifying the first pitch is detected, and a set time has elapsed since the first operator specifying the first pitch was detected, then an operation on the second operator specifying the second pitch is detected, and based on the detection of the operation on the second operator, the number of pitch-specifying operators that are currently being operated on is obtained. If the number of acquired operators reaches a set number, the system controls the syllable to be pronounced in response to the detection of an operation on the second operator, preventing it from progressing from the first syllable to the next second syllable. If the number of acquired operators is less than the set number, the system controls the syllable to be pronounced in response to the detection of an operation on the second operator to advance from the first syllable to the next second syllable. Each of the aforementioned syllables to be pronounced includes multiple frames to be pronounced sequentially. The control unit, If the number of operators reaches the set number, it is determined whether the position of the next frame to be pronounced exceeds the vowel end position of the first syllable which is the current target of pronunciation. If it does not exceed this, the next sound will be pronounced from the frame position within the first syllable. If it exceeds this limit, control is implemented to prevent progression from the first syllable to the second syllable. method.

6. The control unit of the information processing device, If an operation on the first operator specifying the first pitch is detected, and a set time has elapsed since the first operator specifying the first pitch was detected, then an operation on the second operator specifying the second pitch is detected, and based on the detection of the operation on the second operator, the number of pitch-specifying operators that are currently being operated on is obtained. If the number of acquired operators reaches a set number, the system controls the syllable to be pronounced in response to the detection of an operation on the second operator, preventing it from progressing from the first syllable to the next second syllable. If the number of acquired operators is less than the set number, the system controls the syllable to be pronounced in response to the detection of an operation on the second operator to advance from the first syllable to the next second syllable. Each of the aforementioned syllables to be pronounced includes multiple frames to be pronounced sequentially. The control unit, If the number of operators reaches the set number, it is determined whether the position of the next frame to be pronounced exceeds the vowel end position of the first syllable which is the current target of pronunciation. If it does not exceed this, the next sound will be pronounced from the frame position within the first syllable. If it exceeds this limit, control is implemented to prevent progression from the first syllable to the second syllable. A program for executing a process.

Citation Information

Patent Citations

  • Waveform regeneration device

    JP2000066680A

  • Electronic musical instrument, method, and program

    JP2021099461A

  • Electronic musical instrument, method, and program

    JP2021099462A

  • Apparatus and program for vocal synthesis

    JP4735544B2

  • JPP4735544B