Karaoke device

The karaoke device addresses discomfort by aligning guide vocals with the singer's voice quality and preventing silent gaps using vocal range and phoneme data to adjust guide vocal playback.

JP2025151029APending Publication Date: 2025-10-09DAIICHI KOSHO COMPANY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024052245
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-27
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing karaoke machines emit guide vocals that have different vocal qualities from the singer's own voice, causing discomfort during performances, and can create silence when the singer stops singing, also leading to discomfort.

Method used

A karaoke device that acquires a singer's vocal range and phoneme data to determine suitable guide vocal emission times and pitches, adjusting guide vocal playback to match the singer's voice quality and avoid silent gaps.

Benefits of technology

The device reduces discomfort for both the singer and audience by ensuring guide vocals align with the singer's voice quality and preventing silent gaps during performances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025151029000001_ABST
    Figure 2025151029000001_ABST
Patent Text Reader

Abstract

To provide a karaoke device capable of emitting guide vocal sound of reducing discomfort of singers and audiences.SOLUTION: A karaoke device includes an acquisition unit for acquiring voice range data and phoneme data, a determination unit for determining whether a reference pitch of each note in reference data of a music piece selected by a singer is included in a range from a highest pitch to a lowest pitch represented by the voice range data of the singer, a generation unit for generating guide vocal timing data, and a performance control unit for controlling performance means to emit voice of a guide vocal based on the phoneme data and the reference pitch associated with a sounding start time of one note recorded in the guide vocal timing data at the sounding start time of the one note, and to stop emitting the voice of the guide vocal at a sounding end time of the one note.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a karaoke machine. [Background technology]

[0002] Karaoke machines are equipped with a guide vocal function to assist karaoke singing. The guide vocal function plays pre-registered guide vocals (for example, vocals by a professional singer) along with the accompaniment of the song. By using the guide vocals as a model, singers using the karaoke machine can easily sing even unfamiliar songs.

[0003] Patent Document 1 discloses a music playback device with an automatic performance switching function that automatically switches from karaoke to performance with vocals when the singer's voice stops, without the singer having to perform an instantaneous switch operation, and automatically switches from performance with vocals to karaoke when the singer starts singing again. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Publication No. 4-298793 Summary of the Invention [Problem to be solved by the invention]

[0005] The technology described in Patent Document 1 starts vocal performance when there is a few seconds of silence from the singer. In other words, even though the singing section of the karaoke performance is in progress, neither the singing voice nor the guide vocals are emitted, which creates a sense of discomfort for the singer singing the karaoke and the audience listening to the karaoke performance. Furthermore, the guide vocals emitted by the technology described in Patent Document 1 have different vocal qualities from the singer's own singing voice. Therefore, if the guide vocals are emitted in the middle of a karaoke performance, the singer and the audience will feel uncomfortable.

[0006] An object of the present invention is to provide a karaoke apparatus capable of emitting guide vocals that reduce the sense of discomfort felt by singers and audiences. [Means for solving the problem]

[0007] One invention for achieving the above object includes an acquisition unit that acquires vocal range data indicating the highest and lowest pitches that a singer can sing and phoneme data indicating the singer's voice segments for each character; a determination unit that determines whether the reference pitch of each note in reference data for a song selected by the singer is within the range from the highest pitch to the lowest pitch indicated by the acquired vocal range data of the singer; and a guide vocal recording unit that records the pronunciation start time and pronunciation end time of a note whose reference pitch is determined not to be within the range from the highest pitch to the lowest pitch, the reference pitch of the note, and the lyric character corresponding to the note, in association with each other. and a performance control unit that controls the performance means to, after karaoke performance of the song has started, identify phoneme data of the same character as the lyric character associated with the pronunciation start time of one note recorded in the guide vocal timing data from the phoneme data acquired by the acquisition unit at the pronunciation start time of the one note, emit a guide vocal voice based on the identified phoneme data and the reference pitch associated with the pronunciation start time of the one note, and stop emitting the guide vocal voice at the pronunciation end time of the one note. Other features of the present invention will become apparent from the following description and drawings. [Effects of the Invention]

[0008] According to the present invention, it is possible to emit a guide vocal sound that reduces the sense of discomfort felt by the singer and the audience. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a diagram showing a karaoke device according to an embodiment; [Figure 2] 1 is a diagram showing a karaoke machine main body according to an embodiment; [Figure 3] FIG. 10 is a diagram showing reference data according to the embodiment. [Figure 4] FIG. 2 is a diagram showing lyrics data according to the embodiment; [Figure 5] 4 is a flowchart showing the process of the karaoke device according to the embodiment. [Figure 6] FIG. 10 is a diagram showing guide vocal timing data according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] A karaoke device according to an embodiment will be described with reference to FIGS.

[0011] ==Karaoke Equipment== The karaoke device K is a device for performing karaoke music and for singers to sing karaoke. As shown in Fig. 1, the karaoke device K includes a karaoke main unit 10, a speaker 20, a display device 30, a microphone 40, and a remote control device 50.

[0012] The karaoke machine 10 controls various aspects of karaoke performance and singing, such as controlling the karaoke performance of the selected song, controlling the display of lyrics and background images, and processing audio signals input through the microphone 40. The speaker 20 is configured to emit sound based on the sound emission signal from the karaoke machine 10. The display device 30 is configured to display videos and images on a screen based on the signal from the karaoke machine 10. The microphone 40 is configured to convert the singer's singing voice into an analog audio signal and input it to the karaoke machine 10. The remote control device 50 is a device for performing various operations on the karaoke machine 10.

[0013] 2, the karaoke machine 10 according to this embodiment includes a storage unit 10a, a communication unit 10b, an input unit 10c, a performance unit 10d, and a control unit 10e. Each component is connected to a bus B via an interface (not shown).

[0014] [Storage means] The storage means 10a is a large-capacity storage device that stores various types of data, including music data.

[0015] The song data is provided with song identification information for identifying each song. The song identification information is information unique to each song, such as a song ID for identifying the song. The song data includes accompaniment data, reference data, etc. The accompaniment data is data that forms the basis of the karaoke performance sound. The reference data is data that indicates the singing melody of the song performed karaoke, and is data used to evaluate the singer's karaoke singing.

[0016] The reference data consists of multiple notes (musical notes), and for each note, a note number, a predetermined pitch (i.e., reference pitch), a sounding start time (so-called note-on) and a sounding end time (so-called note-off) are set. The note number indicates the order of the notes. For example, the note number corresponding to the first note of a song is "1." The sounding start time and sounding end time are indicated by a reference time (0) that is the start time of the karaoke performance of the song.

[0017] 3 is an example of reference data for music piece X. For example, note No. 1 of music piece X has a reference pitch of "E3," an onset time of "16000 msec," and an onset time of "16750 mec."

[0018] The storage means 10a stores lyric data for displaying lyrics corresponding to each song on the display device 30 or the like in sync with the karaoke performance, background video data such as background video to be displayed on the display device 30 or the like during the karaoke performance, attribute information of the song (song title, singer name, genre, performance time, etc.), and guide vocal data for playing guide vocals for the song.

[0019] Lyric data is composed of multiple lyric characters that make up the lyrics, and includes the note number corresponding to the note at which a certain lyric character is pronounced, the type of lyric character (lyrics or ruby ​​characters), the color-change start time when a certain lyric character begins to change color as the karaoke performance is performed, and the color-change end time when the color-change ends for that certain lyric character. If the lyric characters are kanji or English characters, the note number is set to the ruby ​​character corresponding to that kanji or English character. The color-change start time and color-change end time are indicated using the time (0) as the base time from the start of the karaoke performance of the song.

[0020] 4 is an example of lyrics data for song X. For example, the lyrics character "ya" on note No. 1 of song X is typed as "ruby characters," the color change start time is "16000 msec," and the color change end time is "16750 mec."

[0021] [Communication means / input means] The communication means 10b provides an interface for communicating with the remote control device 50. The input means 10c is configured to allow the singer to input various operations. The input means 10c is a button or the like provided on the karaoke main unit 10. Alternatively, the remote control device 50 may function as the input means 10c.

[0022] [Means of performance] Based on the control of the control means 10e, the performance means 10d performs karaoke performance of music pieces and processes audio signals input through the microphone 40. The performance means 10d includes a sound source, a mixer, an amplifier, etc. (none of which are shown).

[0023] [Control means] The control means 10e performs various controls in the karaoke device K. The control means 10e includes a CPU and a memory (neither of which is shown). The CPU executes programs stored in the memory to realize various functions.

[0024] In this embodiment, the CPU executes a program stored in the memory, causing the control means 10e to function as an acquisition unit 100, a determination unit 200, a generation unit 300, and a performance control unit 400 (see FIG. 2).

[0025] (Acquisition Department) The acquisition unit 100 acquires vocal range data and phoneme data.

[0026] Vocal range data indicates the highest and lowest pitches that a singer can sing. The highest pitch is the highest pitch that a singer can sing. The lowest pitch is the lowest pitch that a singer can sing. In other words, a singer can produce any pitch that falls within the range from the highest pitch to the lowest pitch (conversely, a singer cannot produce any pitch that falls outside that range).

[0027] Phoneme data indicates speech segments for each character of a singer. Speech segments are fragments extracted from the singing voice of a particular singer. Phoneme data differs for each singer. In the phoneme data, speech segments are associated with each character (50 sounds, voiced consonants, palatalized consonants, etc.). For example, the character "a" is associated with a speech segment extracted from the singing voice of a particular singer when he / she pronounces "a," and the character "gyo" is associated with a speech segment extracted from the singing voice of a particular singer when he / she pronounces "gyo."

[0028] The vocal range data and phoneme data can be acquired in various ways. For example, the singer operates the remote control device 50, enters his / her singer identification information and password, and then selects the login icon. The singer identification information is information unique to each singer, such as a singer ID for identifying the singer.

[0029] The karaoke device K transmits a login request including the input singer identification information and password to a server device (not shown). The server device has a singer database. The singer database stores singer information for each singer. The singer information includes singer identification information, a preset login password, singing history, vocal range data, phoneme data, etc.

[0030] The server device completes the login by storing the singer identification information and password included in the received login request in a storage means. The server device also references a singer database and extracts singer information including the combination of singer identification information and password included in the received login request. The server device transmits the extracted singer information to the karaoke device K that sent the login request (at this time, the server device also transmits a signal indicating that login has been completed). The acquisition unit 100 acquires vocal range data and phoneme data included in the singer information received from the server device.

[0031] Alternatively, the acquisition unit 100 can acquire vocal range data directly from the singer's singing voice using known technology (for example, processing similar to the processing by an evaluation means for identifying the singer's vocal range as described in Japanese Patent Application Laid-Open No. 2003-15672). Also, the acquisition unit 100 can acquire phoneme data directly from the singer's singing voice using known technology (for example, processing similar to the processing by a user-specific phoneme data extraction means as described in Japanese Patent Application Laid-Open No. 2009-244789).

[0032] (Judgment Department) The determination unit 200 determines whether the reference pitch of each note in the reference data of the song selected by the singer is included in the range from the highest pitch to the lowest pitch indicated by the acquired vocal range data of the singer.

[0033] If the reference pitch of a note is not within the range from the highest pitch to the lowest pitch indicated by the vocal range data, the singer corresponding to that vocal range data cannot produce the reference pitch of that note. On the other hand, if the reference pitch of a note is within the range from the highest pitch to the lowest pitch indicated by the vocal range data, the singer corresponding to that vocal range data can produce the reference pitch of that note.

[0034] The singer operates the remote control device 50 to select a song that the singer wishes to sing karaoke. In this case, the determination unit 200 refers to the storage means 10a and extracts reference data included in the song data corresponding to the song identification information of the song selected by the singer. The determination unit 200 compares the reference pitch of each note in the extracted reference data with the singer's vocal range data acquired by the acquisition unit 100, and determines for each note whether the reference pitch is within the range from the highest pitch to the lowest pitch indicated by the vocal range data. The determination unit 200 outputs the determination result to the generation unit 300.

[0035] (Generation part) The generating unit 300 generates guide vocal timing data that records the pronunciation start time and pronunciation end time of a note whose reference pitch is determined not to be included in the range from the highest pitch to the lowest pitch, the reference pitch of the note, and the lyric characters corresponding to the note, in association with each other.

[0036] When the determination unit 200 outputs a determination result that the reference pitch of a certain note is not included in the range from the highest pitch to the lowest pitch indicated by the vocal range data, the generation unit 300 reads out the pronunciation start time and pronunciation end time of the certain note, the reference pitch of the certain note, and the lyric characters corresponding to the certain note from the storage means 10a, and records them in association with each other.

[0037] On the other hand, if the determination unit 200 outputs a determination result that the reference pitch of a certain note is within the range from the highest pitch to the lowest pitch indicated by the vocal range data, the generation unit 300 will not read out the pronunciation start time and pronunciation end time of that certain note, the reference pitch of that certain note, or the lyric characters corresponding to that certain note. In other words, the pronunciation start time and pronunciation end time of that certain note, etc. will not be recorded.

[0038] The generation unit 300 repeats the above process each time the determination unit 200 outputs a determination result for a certain note. When determination results have been obtained for all notes included in the reference data for a certain song, the generation unit 300 generates guide vocal timing data that records the onset time and onset end time of a note determined to have a reference pitch that does not fall within the range from the highest pitch to the lowest pitch, the reference pitch of the note, and the lyric characters corresponding to the note, in association with each other.

[0039] (Performance control unit) The performance control unit 400 controls the performance means 10d to perform a karaoke performance of the song selected by the singer.

[0040] The performance control unit 400 reads out the accompaniment data of the song selected by the singer from the storage unit 10a based on the song identification information of the song. The performance control unit 400 controls the performance unit 10d to perform karaoke based on the accompaniment data, and causes the speaker 20 to emit karaoke performance sounds.

[0041] Here, after the karaoke performance of the song has started, the performance control unit 400 of this embodiment identifies, from the phoneme data acquired by the acquisition unit 100, the phoneme data of the same character as the lyric character associated with the pronunciation start time of one note recorded in the guide vocal timing data, at the pronunciation start time of the one note, controls the performance means 10d to emit a guide vocal voice based on the identified phoneme data and the reference pitch associated with the pronunciation start time of the one note, and stops emitting the guide vocal voice at the pronunciation end time of the one note.

[0042] Suppose that when singing karaoke of a certain song, the singer operates the remote control device 50 and selects to use the guide vocal function. In this case, the performance control unit 400 references the guide vocal timing data generated by the generation unit 300 in time with the start of karaoke performance of the certain song, and checks the pronunciation start time and pronunciation end time in the order of the recorded notes.

[0043] When the pronunciation start time of a note included in the guide vocal timing data arrives, the performance control unit 400 identifies, from the phoneme data acquired by the acquisition unit 100, phoneme data of the same character as the lyric character associated with the pronunciation start time of the note. The performance control unit 400 generates a guide vocal audio signal based on the identified phoneme data and the reference pitch associated with the pronunciation start time of the note. A known technique can be used to generate the guide vocal audio signal based on the phoneme data and the reference pitch. For example, when generating an audio signal from the phoneme data, the performance control unit 400 generates the guide vocal audio signal by converting the phoneme data into a fundamental frequency corresponding to the reference pitch.

[0044] The performance control unit 400 outputs the generated audio signal to the performance means 10d and controls the performance means 10d to emit the guide vocal voice based on the audio signal. The performance means 10d generates the guide vocal voice based on the input audio signal and emits it from the speaker 20. On the other hand, when the pronunciation end time of one note included in the guide vocal timing data arrives, the performance control unit 400 controls the performance means 10d to stop emitting the guide vocal voice. The performance control unit 400 repeats the above process until the karaoke performance of the song is completed.

[0045] ==About the operation of the karaoke device K== Next, a specific example of the operation of the karaoke device K in this embodiment will be described with reference to Figures 5 and 6. Figure 5 is a flowchart showing an example of the operation of the karaoke device K. Figure 6 shows an example of guide vocal timing data generated by the generation unit 300. In this example, it is assumed that a singer V sings karaoke. It is also assumed in this example that the storage means 10a stores the reference data for song X shown in Figure 3 and the lyric data for song X shown in Figure 4.

[0046] The acquiring unit 100 acquires vocal range data indicating the highest and lowest pitches that the singer V can sing, and phoneme data indicating the voice segments for each character of the singer V (acquire vocal range data and phoneme data; step 10).

[0047] The singer V operates the remote control device 50 to select the song X that he or she wishes to sing (selecting song X; step 11).

[0048] The determination unit 200 determines whether the reference pitch of each note in the reference data for song X selected by singer V is within the range from the highest pitch to the lowest pitch indicated by the vocal range data of singer V acquired in step 10 (determines whether the reference pitch of each note in the reference data is within the range from the highest pitch to the lowest pitch indicated by the vocal range data; step 12).

[0049] The generating unit 300 generates guide vocal timing data that records the sounding start time and sounding end time of a note whose reference pitch is determined not to be included in the range from the highest pitch to the lowest pitch, the reference pitch of the note, and the lyric characters corresponding to the note in association with each other (generating guide vocal timing data; step 13).

[0050] The performance control unit 400 starts the karaoke performance of the piece of music X (start karaoke performance; step 14).

[0051] At the pronunciation start time of one note recorded in the guide vocal timing data generated in step 13 (if Y in step 15), the performance control unit 400 identifies phoneme data of the same character as the lyric character associated with the pronunciation start time of that one note from the phoneme data acquired by the acquisition unit 100, and controls the performance means 10d to emit a guide vocal voice based on the identified phoneme data and the reference pitch associated with the pronunciation start time of that one note (emit a guide vocal voice based on the phoneme data and the reference pitch; step 16).

[0052] On the other hand, at the time when the pronunciation of one note recorded in the guide vocal timing data generated in step 13 ends (if Y in step 17), the performance control unit 400 controls the performance means 10d to stop emitting the guide vocal voice that started to be emitted in step 16 (stop emitting the guide vocal voice; step 18).

[0053] The karaoke device K repeats the processes of steps 15 to 18 based on the guide vocal timing data generated in step 13 until the karaoke performance of the song X is completed (if Y in step 19).

[0054] Specifically, singer V operates remote control device 50, inputs his / her singer ID and password, and then selects a login icon. Karaoke device K transmits a login request including the input singer ID and password to server device. The server device completes the singer V's login by storing the singer ID and password included in the received login request in storage means. The server device also references a singer database and extracts singer information of singer V, including the combination of singer ID and password included in the received login request. The server device transmits the singer information of singer V to karaoke device K that transmitted the login request. The acquisition unit 100 acquires vocal range data VR and phoneme data VP of singer V included in the singer information received from the server device. It is assumed that the vocal range data VR has a highest pitch of "B4" and a lowest pitch of "D3."

[0055] The singer V operates the remote control device 50 to select the song X that he / she wishes to sing as karaoke and to select the use of the guide vocal function.

[0056] In this case, the determination unit 200 refers to the storage means 10a and extracts reference data included in the song data corresponding to the song ID of song X (see FIG. 3).

[0057] Thereafter, the determination unit 200 determines, in order of note number, whether the reference pitch of each note is included in the range from the highest pitch "B4" to the lowest pitch "D3" indicated by the vocal range data VR.

[0058] 3, the reference pitch of note No. 1 is “E3.” Therefore, the determination unit 200 determines that the reference pitch of note No. 1 is included in the range from the highest pitch to the lowest pitch.

[0059] On the other hand, the reference pitch of note No. 2 is "C3." Therefore, the determination unit 200 determines that the reference pitch of note No. 2 is not included in the range from the highest pitch to the lowest pitch. In this case, the generation unit 300 records the note pronunciation start time "16750 msec" and pronunciation end time "17000 msec" of note No. 2, the reference pitch "C3" of note No. 2, and the lyric character "mi" corresponding to note No. 2 in association with each other.

[0060] Furthermore, the reference pitch of note No. 3 is “F3.” Therefore, the determination unit 200 determines that the reference pitch of note No. 3 is included in the range from the highest pitch to the lowest pitch.

[0061] Similarly, the reference pitch of note No. n-1 is “A4.” Therefore, the determination unit 200 determines that the reference pitch of note No. n-1 is included in the range from the highest pitch to the lowest pitch.

[0062] On the other hand, the reference pitch of note No. n is "C5." Therefore, the determination unit 200 determines that the reference pitch of note No. n is not included in the range from the highest pitch to the lowest pitch. In this case, the generation unit 300 records the note pronunciation start time "37500 msec" and pronunciation end time "38000 msec" of note No. n, the reference pitch "C5" of note No. n, and the lyric character "ki" corresponding to note No. n in association with each other.

[0063] Furthermore, the reference pitch of note No. n+1 is “G4.” Therefore, the determination unit 200 determines that the reference pitch of note No. n+1 is included in the range from the highest pitch to the lowest pitch.

[0064] The above process is repeated until determination results are obtained for all notes included in the reference data for song X. The generation unit 300 then generates guide vocal timing data that records the onset and end times of notes determined not to fall within the range of the reference pitch from the highest to the lowest pitch, the reference pitch of the notes, and the lyrics corresponding to the notes, in association with each other. In this example, it is assumed that the guide vocal timing data GD shown in FIG. 6 has been generated.

[0065] After the guide vocal timing data GD is generated, the performance control unit 400 reads the accompaniment data of the song X from the storage means 10a based on the song ID of the song X selected by the singer V. The performance control unit 400 controls the performance means 10d to perform karaoke based on the accompaniment data, and causes the speaker 20 to emit karaoke performance sounds. The singer V then sings karaoke along with the karaoke performance sounds.

[0066] In this example, the performance control unit 400 references the guide vocal timing data GD generated by the generation unit 300 in time with the start of karaoke performance of song X, and checks the pronunciation start time and pronunciation end time in the order of the recorded notes.

[0067] When the pronunciation start time "16750 msec" of note No. 2 included in the guide vocal timing data GD arrives, the performance control unit 400 identifies, from the phoneme data VP acquired by the acquisition unit 100, the phoneme data of the same character as the lyric character "mi" associated with the pronunciation start time "16750 msec" of note No. 2. The performance control unit 400 generates an audio signal for guide vocal GV2 based on the phoneme data of the identified character "mi" and the reference pitch "C3" associated with the pronunciation start time "16750 msec" of note No. 2. The performance control unit 400 outputs the generated audio signal for guide vocal GV2 to the performance means 10d and controls the performance means 10d to emit the audio of guide vocal GV2 based on the audio signal. The performance means 10d generates the audio of guide vocal GV2 based on the input audio signal and emits it from the speaker 20.

[0068] The reference pitch "C3" of note No. 2 is a pitch that is not included in the range from the highest pitch "B4" to the lowest pitch "D3" indicated by singer V's vocal range data, i.e., a pitch that singer V cannot produce. Therefore, by emitting the voice of guide vocal GV2 at the timing when the pronunciation of note No. 2 begins, singer V's karaoke singing can be supported. Furthermore, because the voice of guide vocal GV2 is emitted based on the time when the pronunciation of the note begins, a situation will not occur in which neither the singing voice nor the voice of guide vocal GV2 is emitted when switching from singer V's singing voice to the voice of guide vocal GV2. Furthermore, the voice of guide vocal GV2 that is emitted is based on singer V's phoneme data VP. Therefore, even when singer V's singing voice is switched to the voice of guide vocal, it is difficult to distinguish audibly.

[0069] On the other hand, when the sounding end time "17000 msec" of note No. 2 included in the guide vocal timing data GD arrives, the performance control section 400 controls the performance means 10d to stop emitting the sound of the guide vocal GV2.

[0070] The reference pitch "F3" of note No. 3, the note following note No. 2, is a pitch included in the range from the highest pitch "B4" to the lowest pitch "D3" indicated by singer V's vocal range data VR, i.e., a pitch that singer V can produce. Therefore, by stopping the emission of the guide vocalist GV2's voice at the timing when the pronunciation of note No. 2 ends, singer V can easily sing karaoke in his or her own voice. Also, because the emission of the guide vocalist GV2's voice is stopped based on the time when the pronunciation of the note ends, the voice of guide vocalist GV2 does not affect subsequent karaoke singing.

[0071] Similarly, by referring to the guide vocal timing data GD, when the pronunciation start time of note No. n, "37500 msec," arrives, the performance control unit 400 identifies, from the phoneme data VP acquired by the acquisition unit 100, the phoneme data of the same character as the lyric character "ki" associated with the pronunciation start time of note No. n, "37500 msec." The performance control unit 400 generates an audio signal for guide vocal GVn based on the phoneme data of the identified character "ki" and the reference pitch "C5" associated with the pronunciation start time of note No. n, "37500 msec." The performance control unit 400 outputs the generated audio signal for guide vocal GVn to the performance means 10d and controls the performance means 10d to emit the audio of guide vocal GVn based on the audio signal. The performance means 10d generates the audio of guide vocal GVn based on the input audio signal and emits it from the speaker 20.

[0072] On the other hand, when the sounding end time "38000 msec" of note No. n included in the guide vocal timing data GD arrives, the performance control section 400 controls the performance means 10d to stop emitting the sound of the guide vocal GVn.

[0073] As is clear from the above, the karaoke device K according to this embodiment comprises an acquisition unit 100 that acquires vocal range data indicating the highest and lowest pitches that a singer can sing and phoneme data indicating the singer's voice segments for each character; a determination unit 200 that determines whether the reference pitch of each note in the reference data of a song selected by the singer is within the range from the highest pitch to the lowest pitch indicated by the acquired vocal range data of the singer; and a determination unit 200 that records, in association with each other, the pronunciation start time and pronunciation end time of a note whose reference pitch is determined not to be within the range from the highest pitch to the lowest pitch, the reference pitch of the note, and the lyric character corresponding to the note. and a performance control unit 400 that, after karaoke performance of the song has started, identifies phoneme data of the same character as the lyric character associated with the pronunciation start time of one note from the phoneme data acquired by the acquisition unit 100 at the pronunciation start time of the one note recorded in the guide vocal timing data, emits a guide vocal voice based on the identified phoneme data and the reference pitch associated with the pronunciation start time of the one note, and controls the performance means 10d to stop emitting the guide vocal voice at the pronunciation end time of the one note.

[0074] With this karaoke device K, whether or not to emit a guide vocal can be adjusted using guide vocal timing data that records the onset and end times of notes with reference pitches determined not to fall within the range from the highest pitch to the lowest pitch among the reference pitches of notes in the reference data for the song selected by the singer, the reference pitch of the note, and the corresponding lyric text. By emitting the guide vocal based on the onset and end times of the note, a situation will never occur during karaoke singing where neither the singer's voice nor the guide vocal is emitted (unless the singer intentionally stops singing). Furthermore, the guide vocal generated based on the reference pitch of the note and the corresponding lyric text has the same vocal quality as the singer's own singing voice. In other words, the karaoke device K according to this embodiment can emit a guide vocal that reduces the sense of discomfort felt by the singer and the audience.

[0075] <Other> The program can also be supplied to a computer using a non-transitory computer-readable medium with an executable program stored thereon. Examples of non-transitory computer-readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), CD-ROMs (Read Only Memory), etc.

[0076] The above-described embodiments are presented as examples and do not limit the scope of the invention. The above configurations can be implemented in appropriate combinations, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. The above-described embodiments and their modifications are included in the scope and spirit of the invention, as well as in the inventions described in the claims and their equivalents. [Explanation of symbols]

[0077] 100 Acquisition Department 200 Judgment section 300 Generation part 400 Performance control unit K Karaoke equipment

Claims

[Claim 1] an acquisition unit that acquires vocal range data indicating the highest and lowest pitches that a singer can sing, and phoneme data indicating the singer's speech segments for each character; a determination unit that determines whether the reference pitch of each note in the reference data of the song selected by the singer is within the range from the highest pitch to the lowest pitch indicated by the acquired vocal range data of the singer; a generation unit that generates guide vocal timing data that records the onset and end times of notes whose reference pitch is determined not to be within the range from the highest pitch to the lowest pitch, the reference pitch of the notes, and the lyrics corresponding to the notes, in association with each other; a performance control unit that controls the performance means to, after the karaoke performance of the song has started, identify phoneme data of the same character as the lyric character associated with the pronunciation start time of one note recorded in the guide vocal timing data from the phoneme data acquired by the acquisition unit, at the pronunciation start time of the one note, emit a guide vocal voice based on the identified phoneme data and the reference pitch associated with the pronunciation start time of the one note, and stop emitting the guide vocal voice at the pronunciation end time of the one note; A karaoke device having:

Citation Information

Patent Citations

  • Music reproduction device with automatic performance switching function

    JP1992298793A