Singing sound output system and method, musical instrument

The singing voice output system generates synchronized singing sounds from devices like drums by analyzing sound velocity and accent, addressing the challenge of pitchless input devices.

JP7835313B2Active Publication Date: 2026-03-25YAMAHA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Existing systems struggle to generate singing voices in response to performance inputs from devices that cannot provide pitch information, such as drums.

Method used

A singing voice output system that analyzes sound information for velocity and accent, generates phrases from syllables, and synthesizes singing sounds based on performance input, allowing synchronization with accompaniment.

Benefits of technology

The system outputs singing sounds that correspond to the level and timing of performance input, synchronizing with accompaniment, even from devices lacking pitch information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007835313000001
    Figure 0007835313000001
  • Figure 0007835313000002
    Figure 0007835313000002
  • Figure 0007835313000003
    Figure 0007835313000003
Patent Text Reader

Abstract

To output singing sound corresponding to timing and strength of performance input.SOLUTION: A singing sound output system capable of outputting singing sound corresponding to timing and strength of performance input is provided. The singing sound output system includes: an acquisition part 42 for acquiring a series of sound information N including at least information indicating timing and information indicating velocity; a phrase generation part 47 for analyzing accents of the series of sound information N from velocity of individual sound information in the series of acquired sound information N and generating a phrase formed of a plurality of syllables corresponding to the series of sound information N based on the accents; a synthesis part 45 for synthesizing singing sound based on the syllables of the generated phrase; and an output part 45 for outputting the synthesized singing sound.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a singing voice output system that outputs singing voices. and methods, instruments

Background Art

[0002] Techniques for generating singing voices in response to performance operations are known. For example, the singing voice synthesizer disclosed in Patent Document 1 generates singing voices by automatically advancing lyrics one character or one syllable at a time in response to real-time performance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The possibilities are endless if we can generate vocal sounds using devices that cannot input pitch information, such as drums.

[0005] One object of the present invention is to provide a singing voice output system that can output singing voices according to performance input. The strength and methods, instruments

Means for Solving the Problems

[0006] According to one aspect of the present invention, Be an acquisition unit that acquires a series of sound information including at least information indicating velocity, and an accent of the series of sound information is analyzed from the velocity of each sound information in the series of sound information acquired by the acquisition unit, and a phrase composed of a plurality of syllables corresponding to the series of sound information is generated based on the accent, and a phrase generated by the phrase generation unit Report Zu ​​​​A singing sound output system is provided, comprising a synthesis unit that synthesizes singing sounds based on [a certain condition], and an output unit that outputs the singing sounds synthesized by the synthesis unit. [Effects of the Invention]

[0007] According to one embodiment of the present invention, performance input The strength It can output singing sounds that correspond to the level of singing. [Brief explanation of the drawing]

[0008] [Figure 1] This diagram shows the overall configuration of the singing sound output system according to the first embodiment. [Figure 2] This is a block diagram of the singing sound output system. [Figure 3] This is a functional block diagram of the singing sound output system. [Figure 4] This is a timing chart for the process of outputting vocal sounds through instrumental performance. [Figure 5] This is a flowchart of the system processing. [Figure 6] This is a timing chart for the process of outputting vocal sounds through instrumental performance. [Figure 7] This is a flowchart of the system processing. [Modes for carrying out the invention]

[0009] Embodiments of the present invention will be described below with reference to the drawings.

[0010] (First Embodiment) Figure 1 shows the overall configuration of a singing sound output system according to the first embodiment of the present invention. This singing sound output system 1000 includes a PC (personal computer) 101, a cloud server 102, and a sound output device 103. The PC 101 and the sound output device 103 are connected to the cloud server 102 via a communication network 104 such as the Internet. Within the environment in which the PC 101 is used, there are items and devices for inputting sound, such as a keyboard 105, a wind instrument 106, and a drum 107.

[0011] The keyboard 105 and drum 107 are electronic instruments used to input MIDI (Musical Instrument Digital Interface) signals. The wind instrument 106 is an acoustic instrument used to input monophonic analog sound. The keyboard 105 and wind instrument 106 can also input pitch information. The wind instrument 106 may be an electronic instrument, and the keyboard 105 and drum 107 may be acoustic instruments. These instruments are examples of devices for inputting sound information and are played by the user on the PC 101 side. The voice of the user on the PC 101 side may also be used as a means of inputting analog sound, in which case the human voice is input as analog sound. Therefore, the concept of "performance" for inputting sound information in this embodiment also includes input of the human voice. Furthermore, the device for inputting sound information does not have to be in the form of a musical instrument.

[0012] Although details will be described later, a typical process performed by the singing sound output system 1000 is outlined below. The user on PC 101 plays an instrument while listening to the accompaniment. PC 101 sends singing data 51, timing information 52, and accompaniment data 53 (all described later in Figure 3) to the cloud server 102. The cloud server 102 synthesizes the singing sound based on the sound produced by the user's performance on PC 101. The cloud server 102 sends the singing sound, timing information 52, and accompaniment data 53 to the sound output device 103. The sound output device 103 is a device equipped with speaker functionality. The sound output device 103 outputs the received singing sound and accompaniment data 53. At that time, the sound output device 103 synchronizes the singing sound and accompaniment data 53 based on the timing information 52 and outputs them accordingly. The form of "output" here is not limited to playback, but also includes transmission to external devices and recording to recording media.

[0013] Figure 2 is a block diagram of the singing sound output system 1000. The PC 101 includes a CPU 11, ROM 12, RAM 13, memory unit 14, timer 15, operation unit 16, display unit 17, sound generation unit 18, input unit 8, and various I / F (interfaces) 19. These components are connected to each other by a bus 10.

[0014] The CPU 11 controls the entire PC 101. The ROM 12 stores various data in addition to the program executed by the CPU 11. The RAM 13 provides a work area when the CPU 11 executes the program. The RAM 13 temporarily stores various information. The storage unit 14 includes non-volatile memory. The timer 15 measures time. The timer 15 may be a counter type. The operation unit 16 includes multiple controls for inputting various information and accepts instructions from the user. The display unit 17 displays various information. The sound generation unit 18 includes a sound source circuit, an effects circuit, and a sound system.

[0015] The input unit 8 includes an interface for acquiring sound information from devices that input electronic sound information such as the keyboard 105 and the drum 107. The input unit 8 also includes a device such as a microphone for acquiring sound information from devices that input acoustic sound information such as the wind instrument 106. The various I / Fs 19 are connected to the communication network 104 (FIG. 1) wirelessly or by wire.

[0016] The cloud server 102 has a CPU 21, a ROM 22, a RAM 23, a storage unit 24, a timer 25, an operation unit 26, a display unit 27, a sound generation unit 28, and various I / Fs 29. These components are connected to each other by a bus 20. The configuration of these components is the same as that shown by reference numerals 11 to 17 and 19 in the PC 101.

[0017] The sound output device 103 has a CPU 31, a ROM 32, a RAM 33, a storage unit 34, a timer 35, an operation unit 36, a display unit 37, a sound generation unit 38, and various I / Fs 39. These components are connected to each other by a bus 30. The configuration of these components is the same as that shown by reference numerals 11 to 19 in the PC 101.

[0018] FIG. 3 is a functional block diagram of the singing voice output system 1000. The singing voice output system 1000 has a functional block 110. The functional block 110 includes, as individual functional units, an instruction unit 41, an acquisition unit 42, a syllable identification unit 43, a timing identification unit 44, a synthesis unit 45, an output unit 46, and a phrase generation unit 47.

[0019] In this embodiment, as an example, the functions of the teaching unit 41 and the acquisition unit 42 are implemented by the PC 101. These functions are implemented in software by programs stored in the ROM 12. In other words, the CPU 11 loads the necessary programs into the RAM 13 and executes them, providing each function by controlling various calculations and hardware resources. To put it another way, these functions are mainly realized through the cooperation of the CPU 11, ROM 12, RAM 13, timer 15, display unit 17, sound generation unit 18, input unit 8, and various I / F 19. The programs executed here include sequence software.

[0020] Furthermore, the functions of the syllable identification unit 43, timing identification unit 44, synthesis unit 45, and phrase generation unit 47 are implemented by the cloud server 102. Each of these functions is implemented in software by a program stored in the ROM 22. These functions are mainly realized through the cooperation of the CPU 21, ROM 22, RAM 23, timer 25, and various I / F 29.

[0021] Furthermore, the functions of the output unit 46 are realized by the sound output device 103. The functions of the output unit 46 are realized in software by a program stored in the ROM 32. These functions are mainly realized through the cooperation of the CPU 31, ROM 32, RAM 33, timer 35, sound generation unit 38, and various I / F 39.

[0022] The singing sound output system 1000 refers to singing data 51, timing information 52, accompaniment data 53, and phrase database 54. The phrase database 54 is pre-stored in ROM 12, for example. Note that the phrase generation unit 47 and the phrase database 54 are not essential in this embodiment. These will be described in the third embodiment described later.

[0023] The singing data 51, timing information 52, and accompaniment data 53 are associated with each other for each song and are pre-stored in the ROM 12. The accompaniment data 53 is information for playing the accompaniment of each song recorded as sequence data. The singing data 51 contains multiple syllables. The singing data 51 includes lyric text data and a phonological information database. The lyric text data is data that describes the lyrics, and the lyrics for each song are described separated into syllable units. In each song, the accompaniment position in the accompaniment data 53 and the syllables in the singing data 51 are temporally associated by the timing information 52.

[0024] The processing performed by each functional unit in the functional block 110 will be explained in detail in Figures 4 and 5. Here, we will provide an overview. The teaching unit 41 indicates (instructs) the user on the position of progress in the singing data 51. The acquisition unit 42 acquires at least one sound information N (see Figure 4) input by the performance. The syllable identification unit 43 identifies the syllable corresponding to the acquired sound information N from among multiple syllables in the singing data 51. The timing identification unit 44 associates the difference ΔT (see Figure 4) with the sound information N as relative information indicating the timing relative to the identified syllable. The synthesis unit 45 synthesizes the singing sound based on the identified syllable. The output unit 46 outputs the synthesized singing sound and the accompaniment sound based on the accompaniment data 53 in synchronization based on the above relative information.

[0025] Figure 4 is a timing chart of the process that outputs singing sounds through performance. When a song is selected and processing begins, as shown in Figure 4, the PC 101 displays to the user the syllables corresponding to the progress position in the singing data 51. For example, syllables such as "sa", "ku", and "ra" are displayed in order. The pronunciation start timing t(t1~t3) is determined by the temporal correspondence with the accompaniment data 53 and is the pronunciation start timing of the original syllables defined in the singing data 51. For example, time t1 indicates the pronunciation start position of the syllable "sa" in the singing data 51. In parallel with the syllable progression guidance, the accompaniment based on the accompaniment data 53 also progresses.

[0026] The user plays along with the indicated syllable progression. Here, we give an example where MIDI signals are input by playing a keyboard 105 capable of inputting pitch information. The user, as the performer, presses the keys corresponding to each syllable in time with the start timing of each syllable "sa," "ku," and "ra." In this way, sound information N (N1~N3) is acquired sequentially. The duration of each sound information N is the time from the input start timing s (s1~s3) to the input end timing e (e1~e3). The input start timing s corresponds to note on, and the input end timing e corresponds to note off. Sound information N includes pitch information and velocity.

[0027] The user may intentionally shift the actual input start time s relative to the pronunciation start time t. In the cloud server 102, the time difference between the input start time s and the pronunciation start time t is calculated as a temporal difference ΔT (ΔT1~T3) (relative information). The difference ΔT is calculated for each syllable and associated with each syllable. The cloud server 102 synthesizes the singing sound based on the sound information N and sends it to the sound output device 103 along with the accompaniment data 53.

[0028] The sound output device 103 outputs the singing sound and the accompanying sound based on the accompanying data 53 in a synchronized manner. In doing so, the sound output device 103 outputs the accompanying sound at a set constant tempo. For the singing sound, the sound output device 103 outputs it while matching each syllable with the position of the accompanying sound based on the timing information 52. Note that processing time is required from the input of sound information N to the output of the singing sound. Therefore, the sound output device 103 delays the output of the accompanying sound using delay processing in order to match each syllable with the position of the accompanying sound.

[0029] For example, the sound output device 103 adjusts the output timing by referring to the difference ΔT corresponding to each syllable. As a result, the singing sound is output at the input timing (input start timing s). For example, the output (pronunciation) of the syllable "ku" starts at a timing that is difference ΔT2 earlier than the pronunciation start timing t2. Also, the output (pronunciation) of the syllable "ra" starts at a timing that is difference ΔT3 later than the pronunciation start timing t3. The pronunciation of each syllable ends (is muted) at a time that corresponds to the input end timing e. Therefore, the accompaniment sound is output at a fixed tempo, and the singing sound is output at a timing that corresponds to the performance timing. Thus, the singing sound can be output at the timing when the sound information N is input, in synchronization with the accompaniment.

[0030] Figure 5 is a flowchart showing the system processing that outputs singing sounds through performance executed by the singing sound output system 1000. In this system processing, PC processing executed on PC 101, cloud server processing executed on cloud server 102, and sound output device processing executed on sound output device 103 are executed in parallel. PC processing is realized by the CPU 11 loading a program stored in ROM 12 into RAM 13 and executing it. Cloud server processing is realized by the CPU 21 loading a program stored in ROM 22 into RAM 23 and executing it. Sound output device processing is realized by the CPU 31 loading a program stored in ROM 32 into RAM 33 and executing it. Each of these processes starts when the PC 101 is instructed to start system processing.

[0031] First, let's explain the PC processing. In step S101, the CPU 11 of PC 101 selects a song to be played (hereinafter referred to as the selected song) from among several prepared songs, based on the user's instructions. The tempo of each song is predetermined by default. However, when a song is selected, the CPU 11 may change the tempo based on the user's instructions.

[0032] In step S102, the CPU 11 sends relevant data (singing data 51, timing information 52, accompaniment data 53) corresponding to the selected song to the cloud server 102 via various I / F 19.

[0033] In step S103, the CPU 11 begins teaching the current position. Accordingly, the CPU 11 sends a notification to the cloud server 102 indicating that teaching the current position has begun. The teaching process here is implemented, for example, by executing sequence software. The CPU 11 (teaching unit 41) teaches the current position using timing information 52.

[0034] For example, the display unit 17 displays lyrics corresponding to the syllables in the singing data 51. The CPU 11 indicates the current position on the displayed lyrics. For example, the teaching unit 41 indicates the current position by changing the display method, such as the color, of the lyrics at the current position, or by moving the cursor position or the position of the lyrics themselves. Furthermore, the CPU 11 indicates the current position by playing the accompaniment data 53 at the set tempo. Note that the method of indicating the current position is not limited to these examples, and various methods that allow recognition by sight or hearing can be employed. For example, it may be a method of showing the note at the current position on the displayed musical score. Alternatively, the start timing may be indicated and then a metronome sound may be generated. Also, only at least one method needs to be employed, and multiple methods may be combined.

[0035] In step S104, the CPU 11 (acquisition unit 42) executes sound information acquisition processing. The user plays in time with the lyrics, for example, while confirming the taught position (for example, while listening to the accompaniment). The CPU 11 acquires the MIDI data or analog sound from the performance as sound information N. Sound information N usually includes input start timing s, input end timing e, pitch information, and velocity information. However, pitch information is not always included, as in the case of drum 107 being played. Velocity information may be canceled. The input start timing s and input end timing e are defined by relative time to the accompaniment progression. If analog sound such as a human voice is acquired by a microphone, audio data is acquired as sound information N.

[0036] In step S105, the CPU 11 sends the sound information N acquired in step S104 to the cloud server 102. In step S106, it is determined whether the selected song has finished, that is, whether the instruction of the progress position up to the last position in the selected song has been completed. If the selected song has not finished, the CPU 11 returns to step S104. Therefore, until the selected song finishes, the sound information N acquired in accordance with the performance as the song progresses is sent to the cloud server 102 as needed. When the selected song finishes, the CPU 11 sends a notification to the cloud server 102 to that effect and terminates the PC processing.

[0037] Next, the cloud server processing will be explained. In step S201, the CPU 21 of the cloud server 102 receives the relevant data corresponding to the selected song through the various I / F 29, and then proceeds to step S202. In step S202, the CPU 21 transmits the received relevant data to the sound output device 103 through the various I / F 29. Note that the singing data 51 does not need to be transmitted to the sound output device 103.

[0038] In step S203, the CPU 21 starts a series of processes (S204-S209). At the start of this series of processes, the CPU 21 executes sequencing software and uses the received related data to advance time while waiting for the next sound information N to be received. In step S204, the CPU 21 receives the sound information N.

[0039] In step S205, the CPU 21 (syllable identification unit 43) identifies the syllable corresponding to the received sound information N. First, the CPU 21 calculates the difference ΔT for each syllable between the input start timing s in the sound information N and the pronunciation start timing t for each of the multiple syllables in the singing data 51 corresponding to the selected song. Then, the CPU 21 identifies the syllable with the smallest difference ΔT among the multiple syllables in the singing data 51 as the syllable corresponding to the sound information N that was just received.

[0040] For example, in the example shown in Figure 4, for sound information N2, the difference ΔT2 between the input start timing s2 and the pronunciation start timing t2 of the syllable "ku" is the smallest compared to the differences with other syllables. Therefore, the CPU 21 identifies the syllable "ku" as the syllable corresponding to sound information N2. In this way, for each piece of sound information N, the syllable corresponding to the pronunciation start timing t that is closest to the input start timing s is identified as the corresponding syllable.

[0041] If the sound information N is audio data, the CPU 21 (syllable identification unit 43) determines the timing of sound production and cessation, pitch, and velocity of the sound information N through analysis.

[0042] In step S206, the CPU 21 (timing identification unit 44) performs timing identification processing. That is, the CPU 21 associates the difference ΔT with the sound information N that was just received and the syllable identified as the syllable corresponding to the sound information N.

[0043] In step S207, the CPU 21 (synthesis unit 45) synthesizes a singing sound based on the identified syllables. The pitch of the singing sound is determined by the pitch information of the corresponding sound information N. If the sound information N is a drum sound, the pitch of the singing sound may be, for example, a constant pitch. Regarding the output timing of the singing sound, the sounding timing and mute timing are determined by the sounding start timing t and input end timing e (or sounding length) of the corresponding sound information N. Therefore, the singing sound is synthesized from the syllables corresponding to the sound information N at the pitch determined by the performance. In some cases, if the mute during performance is too late, the sounding period of the current syllable may overlap with the original sounding timing of the next syllable in the singing data. In this case, the input end timing e may be modified so that the sound is forcibly mute before the original sounding timing of the next syllable.

[0044] In step S208, the CPU 21 performs data transmission. Specifically, the CPU 21 transmits the synthesized singing sound, the difference ΔT corresponding to the syllables, and velocity information during performance to the sound output device 103 via various I / F 29.

[0045] In step S209, the CPU 21 determines whether the selected song has finished, that is, whether it has received notification from the PC 101 that the selected song has finished. If the selected song has not finished, the CPU 21 returns to step S204. Therefore, singing sounds based on the syllables corresponding to the sound information N are synthesized and transmitted as needed until the selected song finishes. The CPU 21 may also determine that the selected song has finished when a predetermined time has elapsed since the processing of the last received sound information N data has finished. Once the selected song finishes, the CPU 21 terminates the cloud server processing.

[0046] Next, the sound output device processing will be explained. In step S301, the CPU 31 of the sound output device 103 receives relevant data corresponding to the selected song through various I / F 39, and then proceeds to step S302. In step S302, the CPU 31 receives the data (singing sound, difference ΔT, velocity) sent from the cloud server 102 in step S208.

[0047] In step S303, the CPU 31 (output unit 46) performs synchronized output of the singing sound and accompaniment based on the received singing sound and difference ΔT, the accompaniment data 53 already received, and the timing information 52.

[0048] As explained in Figure 4, the CPU 31 outputs accompaniment sounds based on the accompaniment data 53, and in parallel outputs vocal sounds while adjusting the output timing based on timing information and the difference ΔT. Here, playback is adopted as a typical synchronized output mode for the accompaniment sounds and vocal sounds. Therefore, the sound output device 103 can listen to the performance of the PC 101 user in synchronization with the accompaniment.

[0049] Furthermore, the mode of synchronized output is not limited to playback; it may also be recorded as an audio file in the storage unit 34, or transmitted to an external device via various I / F 39.

[0050] In step S304, the CPU 31 determines whether the selected song has finished, that is, whether it has received notification from the cloud server 102 that the selected song has finished. If the selected song has not finished, the CPU 31 returns to step S302. Therefore, the synchronized output of the received singing sound continues until the selected song finishes. The CPU 31 may also determine that the selected song has finished when a predetermined time has elapsed since the processing of the last received data was completed. Once the selected song has finished, the CPU 31 terminates the sound output device processing.

[0051] According to this embodiment, a syllable corresponding to the sound information N, which is acquired while showing the user the position in the singing data 51, is identified from multiple syllables in the singing data 51. Relative information (difference ΔT) is associated with the sound information N, and a singing sound is synthesized based on the identified syllable. Based on the relative information, the singing sound and the accompaniment sound based on the accompaniment data 53 are output in sync. Therefore, the singing sound can be output at the timing when the sound information N is input, in sync with the accompaniment.

[0052] Furthermore, if the sound information N includes pitch information, the singing sound can be output at the pitch input through performance. Also, if the sound information N includes velocity information, the singing sound can be output at a volume corresponding to the intensity of the performance.

[0053] The related data (singing data 51, timing information 52, accompaniment data 53) was sent to the cloud server 102 and the sound output device 103 after the selected song was determined, but this is not limited to that. For example, related data for multiple songs may be stored in advance on the cloud server 102 and the sound output device 103. Then, when the selected song is determined, information identifying the selected song may be sent to the cloud server 102 and further to the sound output device 103.

[0054] (Second Embodiment) In the second embodiment of the present invention, some of the system processing differs from that of the first embodiment. Therefore, the differences from the first embodiment will be mainly explained with reference to Figures 5 and 6. In the first embodiment, the performance tempo was fixed, but in this embodiment, the performance tempo is variable and changes depending on the performance by the musician.

[0055] Figure 6 is a timing chart of the process of outputting singing sounds through performance. The order of multiple syllables in the singing data 51 is predetermined. In Figure 6, when displaying syllable progression, the singing sound output system 1000 shows the user the next syllable in the singing data while waiting for the input of sound information N, and each time sound information N is input, it advances the syllable indicating the progression position by one syllable to the next syllable. Therefore, the syllable progression display waits until there is performance input corresponding to the next syllable. The progress indication of the accompaniment data also waits until there is performance input, in line with the syllable progression.

[0056] The cloud server 102 identifies the next syllable in the sequence of events at the time the sound information N is input as the syllable corresponding to the input sound information N. Therefore, the corresponding syllable is identified sequentially each time a key is pressed.

[0057] The actual input start timing s may differ from the pronunciation start timing t. Similar to the first embodiment, the cloud server 102 calculates the time difference between the input start timing s and the pronunciation start timing t as a temporal difference ΔT (ΔT1~T3) (relative information). The difference ΔT is calculated for each syllable and associated with each syllable. The cloud server 102 synthesizes the singing sound based on the sound information N and sends it to the sound output device 103 along with the accompaniment data 53.

[0058] In Figure 6, the syllable pronunciation start timing t'(t1'~t3') is the timing at which the syllables begin to be pronounced during output. The syllable pronunciation start timing t' is determined by the input start timing s. The progression of the accompanying sound during output also changes as it progresses, depending on the syllable pronunciation start timing t'.

[0059] The sound output device 103 synchronizes the output of the singing sound and the accompanying sound based on the accompanying data 53 by adjusting the output timing based on the timing information and the difference ΔT. At that time, the sound output device 103 outputs the singing sound at the syllable pronunciation start timing t'. For the accompanying sound, the sound output device 103 outputs while matching each syllable with the accompaniment position based on the difference ΔT. In order to match each syllable with the accompaniment position, the sound output device 103 delays the output of the accompanying sound using delay processing. Therefore, the singing sound is output at a timing corresponding to the performance timing, and the tempo of the accompanying sound changes in accordance with the performance timing.

[0060] The system processing in this embodiment will be explained following the flowchart in Figure 5. Unless otherwise specified, the process is the same as in the first embodiment.

[0061] In PC101, during the teaching process initiated in step S103, the CPU11 (teaching unit 41) uses timing information 52 to indicate the current position. In step S104, the CPU11 (acquisition unit 42) executes sound information acquisition processing. The user inputs the sound corresponding to the next syllable while confirming the position. The CPU11 waits for the input of the next sound information N before proceeding with the teaching and accompaniment progression of the syllable. Therefore, the CPU11 teaches the next syllable while waiting for the input of sound information N, and each time sound information N is input, it advances the syllable indicating the position by one syllable to the next syllable. The CPU11 also synchronizes the accompaniment progression with the teaching progression of the syllable.

[0062] In the cloud server 102, during the series of processes initiated in step S203, the CPU 21 advances time while waiting for the reception of sound information N. In step S204, the CPU 21 receives sound information N as it progresses, and advances time when sound information N is received. Therefore, it waits for time to advance until the next sound information N is received.

[0063] Upon receiving sound information N, in step S205, the CPU 21 (syllable identification unit 43) identifies the syllable corresponding to the received sound information N. Here, the CPU 21 identifies the syllable that was the next syllable in the progression sequence at the time sound information N was input as the syllable corresponding to the sound information N that was just received. Therefore, each time a key is pressed during performance, the corresponding syllable is identified in order.

[0064] After identifying the syllables, in step S206, the CPU 21 calculates the difference ΔT and associates it with the identified syllables. That is, as shown in Figure 6, the CPU 21 determines the difference ΔT as the time difference between the input start timing s and the pronunciation start timing t corresponding to the identified syllable. Then, the CPU 21 associates the calculated difference ΔT with the identified syllables.

[0065] In step S208, the CPU 21 transmits the synthesized singing sound, the difference ΔT corresponding to the syllables, and the velocity during performance to the sound output device 103 via various I / F 29.

[0066] In the sound output device 103, during the synchronous output processing performed in step S303, the CPU 31 (output unit 46) performs synchronous output of the singing sound and accompaniment based on the received singing sound and difference ΔT, the already received accompaniment data 53, and the timing information 52. At that time, the CPU 31 adjusts the output timing of the accompaniment sound and singing sound by referring to the difference ΔT, thereby matching each syllable with the accompaniment position during the output processing.

[0067] As a result, as shown in Figure 6, the singing sound is output at the time of input start (input start time s) according to the input timing. For example, the output (pronunciation) of the syllable "ku" starts at a time ΔT2 earlier than the pronunciation start time t2. Also, the output (pronunciation) of the syllable "ra" starts at a time ΔT3 later than the pronunciation start time t3. The pronunciation of each syllable ends at the time corresponding to the input end time e.

[0068] On the other hand, the tempo of the accompanying sound changes in accordance with the performance timing. For example, CPU31 corrects the position of the sound start timing t2 for the accompanying sound and outputs it at the position of the sound start timing t2'.

[0069] Therefore, the accompaniment sound is output at a variable tempo, and the vocal sound is output at a timing corresponding to the performance timing. Consequently, the vocal sound can be output in synchronization with the accompaniment, at the timing when sound information N is input.

[0070] According to this embodiment, the teaching unit 41 indicates the next syllable while waiting for the input of sound information N, and each time sound information N is input, it advances the syllable indicating the current position by one to the next syllable. The syllable identification unit 43 then identifies the syllable that was the next syllable in the progression order at the time sound information N was input as the syllable corresponding to the input sound information N. Therefore, the same effect as in the first embodiment can be achieved in terms of outputting the singing sound in synchronization with the accompaniment at the timing when sound information N is input. Furthermore, even if the user plays at a free tempo, the singing sound can be output in synchronization with the accompaniment according to the user's playing tempo.

[0071] In the first and second embodiments, the relative information associated with the sound information N is not limited to the difference ΔT. For example, the relative information indicating the timing relative to a specified syllable may be the relative time of the sound information N and the relative time of each syllable, based on a certain time defined by the timing information 52.

[0072] (Third embodiment) A third embodiment of the present invention will be described with reference to Figures 1 to 3 and Figure 7. The enjoyment of singing can be expanded if singing sounds can be produced using a device that cannot input pitch information, such as a drum. Therefore, in this embodiment, a drum 107 is used as the performance input. In this embodiment, without being taught accompaniment or syllable progression, when the user freely strikes and plays the drum 107, a singing phrase is generated for each unit of sound information N acquired thereby. The basic configuration of the singing sound output system 1000 is the same as in the first embodiment. In this embodiment, performance input using a drum 107 is assumed, and since it is assumed that there is no pitch information, a different control is applied than in the first embodiment.

[0073] In this embodiment, the teaching unit 41, timing identification unit 44, singing data 51, timing information 52, and accompaniment data 53 shown in Figure 3 are not essential. The phrase generation unit 47 analyzes the accent of a series of sound information N from the velocity of each individual sound information N in the series of sound information N, and generates a phrase consisting of multiple syllables corresponding to the series of sound information N based on the accent. The phrase generation unit 47 generates a phrase corresponding to the series of sound information N by extracting a phrase that matches the accent from a phrase database 54 containing multiple pre-prepared phrases. A phrase having the number of syllables that constitute the series of sound information N is extracted.

[0074] Here, the accent of a series of sound information N refers to a dynamic accent based on the relative intensity of the sounds. The accent of a phrase refers to a pitch accent based on the relative pitch of each syllable. Therefore, the intensity of the sounds in the sound information N corresponds to the pitch of the phrase.

[0075] Figure 7 is a flowchart showing the system processing that outputs singing sounds through performance performed by the singing sound output system 1000. The execution entities, execution conditions, and start conditions for the PC processing, cloud server processing, and sound output device processing in this system processing are the same as those shown in Figure 5.

[0076] First, let's explain the PC processing. In step S401, the CPU 11 of PC 101 transitions to the playback start state based on the user's instructions. At that time, the CPU 11 sends a notification to the cloud server 102 via various I / F 19 indicating that it has transitioned to the playback start state.

[0077] In step S402, the CPU 11 (acquisition unit 42) acquires sound information N corresponding to the user striking the drum 107. The sound information N is either MIDI data or analog sound. The sound information N includes at least information indicating the input start timing (strike on) and information indicating the velocity.

[0078] In step S403, the CPU 11 (acquisition unit 42) determines whether the current series of sound information N has been finalized. For example, if the first sound information N is input within a first predetermined time after transitioning to the playback start state, the CPU 11 determines that the series of sound information N has been finalized if a second predetermined time has elapsed since the last sound information N was input. The series of sound information N is assumed to be a set of multiple sound information N, but it may also be a single sound information N.

[0079] In step S404, the CPU 11 sends the acquired series of sound information N to the cloud server 102. In step S405, the CPU 11 determines whether or not the user has instructed the end of the playback state. If the user has not instructed the end of playback, the CPU 11 returns to step S402. If the user has instructed the end of playback, the CPU 11 sends a notification to that effect to the cloud server 102 and terminates the PC processing. Therefore, each time a set of series of sound information N is confirmed, that set of sound information N is sent.

[0080] Next, the cloud server processing will be explained. When CPU21 receives notification that it has transitioned to the playback start state, it starts a series of processes (S502~S506) in step S501. In step S502, CPU21 receives a series of sound information N sent from PC101 in step S404.

[0081] In step S503, the CPU 21 (phrase generation unit 47) generates a single phrase for the current series of sound information N. The method is illustrated below. For example, the CPU 21 analyzes the accent of the series of sound information N from the velocity of each individual sound information N, and extracts a phrase from the phrase database 54 that matches the accent and the number of syllables that make up the series of sound information N. At that time, the extraction range may be narrowed by conditions. For example, the phrase database 54 may be classified by conditions, and the user may be able to set at least one of conditions such as "noun," "fruit," "stationery," "color," and "size."

[0082] For example, consider the case where there are 4 sound information elements N and the condition is "fruit". If the analyzed accent is "strong-weak-weak-weak", then "durian" is extracted. If the accent is "weak-strong-weak-weak", then "orange" is extracted. Consider the case where there are 4 sound information elements N and the condition is "stationery". If the analyzed accent is "strong-weak-weak-weak", then "compass" is extracted. If the accent is "weak-strong-weak-weak", then "crayon" is extracted. Note that setting a condition is not mandatory.

[0083] In step S504, the CPU 21 (synthesis unit 45) synthesizes a singing sound from the generated phrase. The pitch of the singing sound may be in accordance with the pitch of each syllable set in the phrase. In step S505, the CPU 21 transmits the singing sound to the sound output device 103 via various I / F 29.

[0084] In step S506, the CPU 21 determines whether or not it has received a notification from the PC 101 that playback has been instructed to end. If the CPU 21 has not received a notification that playback has been instructed to end, it returns to step S502. If the CPU 21 has received a notification that playback has been instructed to end, it sends a notification that playback has been instructed to end to the sound output device 103 and terminates the cloud server processing.

[0085] Next, the sound output device processing will be described. In step S601, when the CPU 31 of the sound output device 103 receives the singing sound through the various I / F 39, it proceeds to step S602. In step S602, the CPU 31 (output unit 46) outputs the received singing sound. The output timing of each syllable depends on the input timing of the corresponding sound information N. The mode of output here is not limited to playback, as in the first embodiment.

[0086] In step S603, the CPU 31 determines whether or not it has received a notification from the cloud server 102 that it has been instructed to end the performance. If the CPU 31 has not received a notification that it has been instructed to end the performance, it returns to step S601. If it has received a notification that it has been instructed to end the performance, it terminates the sound output device processing. Therefore, the CPU 31 outputs each time it receives a singing sound of a phrase.

[0087] According to this embodiment, it is possible to output singing sounds that correspond to the timing and intensity of the performance input.

[0088] In this embodiment, since the timbre differs between striking the drumhead and striking the rim (rimshot), this difference in timbre may also be used as a parameter for phrase generation. For example, the above conditions for phrase extraction may be different for striking the drumhead and striking the rim.

[0089] Furthermore, the sound produced by striking is not limited to drums; it could also be handclaps. If using electronic drums, the striking position on the drumhead can be detected, and the difference in striking position may also be used as a parameter for phrase generation.

[0090] In this embodiment, if the acquired sound information N includes pitch information, the pitch can be replaced with accents, and processing can be performed in a similar manner to that performed when drums are struck. For example, when a piano plays "C-E-C", a phrase corresponding to when drums play "weak-strong-weak" can be extracted.

[0091] In each of the above embodiments, if the sound output device 103 is equipped with multiple singing voices (multiple genders, etc.), the singing voice to be used may be switched according to the sound information N. For example, if the sound information N is audio data, the singing voice may be switched according to its timbre. If the sound information N is MIDI data, the singing voice may be switched according to the timbre or other parameters set on the PC 101.

[0092] In each of the above embodiments, it is not essential that the singing sound output system 1000 includes the PC 101, the cloud server 102, and the sound output device 103. Nor is it limited to a system that goes via a cloud server. In other words, each of the functional units shown in Figure 3 may be implemented in any of the devices, or in a single device. If the above functional units are implemented in a single integrated device, that device does not have to be called a singing sound output system, but may be called a singing sound output device.

[0093] In addition, in each of the above embodiments, at least a part of each functional unit shown in Figure 3 may be realized by AI (Artificial Intelligence).

[0094] Although the present invention has been described in detail above based on its preferred embodiments, the present invention is not limited to these specific embodiments, and various forms that do not depart from the spirit of the invention are also included in the present invention. Some of the above embodiments may be combined as appropriate.

[0095] Furthermore, the same effects as the present invention may be achieved by reading a storage medium containing a control program represented by software for achieving the present invention into this system. In this case, the program code read from the storage medium itself will realize the novel function of the present invention, and the non-transient computer-readable recording medium storing that program code will constitute the present invention. Alternatively, the program code may be supplied via a transmission medium or the like, in which case the program code itself will constitute the present invention. In these cases, the storage medium can be ROM, floppy disk, hard disk, optical disk, magneto-optical disk, CD-ROM, CD-R, magnetic tape, non-volatile memory card, etc. The non-transient computer-readable recording medium also includes volatile memory (e.g., DRAM (Dynamic Random Access Memory)) inside a computer system that acts as a server or client when a program is transmitted via a network such as the Internet or a communication line such as a telephone line, which retains the program for a certain period of time. [Explanation of symbols]

[0096] 42 Acquisition unit, 45 Synthesis unit, 46 Output unit, 47 Phrase generation unit, 1000 Singing sound output system

Claims

1. An acquisition unit that acquires a series of sound information including at least velocity information, A phrase generation unit analyzes the accent of the series of sound information from the velocity of each sound information in the series of sound information acquired by the acquisition unit, and generates a phrase consisting of multiple syllables corresponding to the series of sound information based on the accent, A synthesis unit synthesizes singing sounds based on the phrases generated by the phrase generation unit, A singing sound output system comprising: an output unit that outputs the singing sound synthesized by the synthesis unit.

2. The singing sound output system according to claim 1, wherein the phrase generation unit generates a phrase corresponding to the series of sound information by extracting a phrase that matches the accent from a pre-prepared database of phrases.

3. The singing sound output system according to claim 2, wherein the series of sound information further includes information indicating timing.

4. The singing sound output system according to claim 2, wherein the phrase generation unit narrows the extraction range based on conditions when extracting the phrase.

5. The aforementioned series of sound information further includes timbre information, The singing sound output system according to claim 4, wherein the phrase generation unit varies the conditions when extracting the phrase depending on the difference in timbre.

6. The aforementioned series of sound information is generated by the performer's striking operation. The unit includes a detection unit for detecting the striking position during the striking operation, The singing sound output system according to claim 4, wherein the phrase generation unit varies the conditions depending on the difference in the striking position when extracting the phrase.

7. A series of sound information is obtained that includes at least velocity information, The accent of the series of sound information is analyzed from the velocity of each sound information in the acquired series of sound information, and a phrase consisting of multiple syllables corresponding to the series of sound information is generated based on the accent. Based on the generated phrase, the singing sound is synthesized, Output the synthesized singing sound. A method for outputting singing sounds, performed by a singing sound output system.

8. An acquisition unit that acquires a series of sound information including at least velocity information, A phrase generation unit analyzes the accent of the series of sound information from the velocity of each sound information in the series of sound information acquired by the acquisition unit, and generates a phrase consisting of multiple syllables corresponding to the series of sound information based on the accent, A synthesis unit synthesizes singing sounds based on the phrases generated by the phrase generation unit, A musical instrument having an output unit that outputs a singing sound synthesized by the aforementioned synthesis unit.

Citation Information

Patent Citations

  • Electronic instrument

    JP1997281970A

  • Phoneme information synthesis device and voice synthesis device

    JP2016080827A

  • Singing sound synthesis device

    JP2016206323A