Karaoke device

The karaoke device adjusts tempo based on mouth shape analysis, addressing the challenge of microphone usage difficulties by automatically synchronizing performance tempo with user singing speed.

JP2025108300APending Publication Date: 2025-07-23DAIICHI KOSHO COMPANY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024002146
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-23

AI Technical Summary

Technical Problem

Existing karaoke devices require users to sing into a microphone to adjust the performance tempo, which can be challenging for individuals with difficulty using microphones.

Method used

A karaoke device that adjusts performance tempo based on mouth shape analysis from video footage, without requiring the user's singing voice, using mouth shape images and associated identification information to match with stored data for tempo adjustment.

Benefits of technology

Enables automatic tempo adjustment of music pieces without relying on the user's singing voice, accommodating users who have difficulty using microphones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025108300000001_ABST
    Figure 2025108300000001_ABST
Patent Text Reader

Abstract

To provide a Karaoke device which automatically adjusts a performance tempo of music without using singing voice of a user.SOLUTION: In a Karaoke device, a Karaoke body 10 acquires a video of a user, acquires mouth shape identification information and extraction time, determines whether first mouth shape identification information matches first mouth shape identification information, sets previous extraction time, previous sound length and a previous matching flag (matching) if they match, sets a previous matching flag (non-matching) if they do not match, executes a first total value processing unit for performing first to fourth processing and a second total value processing unit for performing fifth to eighth processing according to a result of determining whether (the first+n)th mouth shape identification information match with (1+n)th mouth type information, and when finish time for color change of a last lyric character included in the prescribed section has come, calculates a singing tempo in a prescribed section based on the first and second total values acquired up to that point, to adjust the performance tempo according to the calculated singing tempo.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a karaoke device.

Background Art

[0002] When performing karaoke singing using a karaoke device, there may be a deviation between the performance tempo of a music piece and the singing speed of a singer.

[0003] Therefore, there is a technology in a karaoke device to make the performance tempo of a music piece follow the singing speed of a singer. For example, Patent Document 1 discloses a technology that tracks the pitch of singing voice captured from a microphone for a predetermined period to recognize the pattern of pitch change of the singing voice, determines the slowness or rapidity of singing based on the difference between the recognized pitch change pattern and the pitch change pattern indicated by a guide melody, and controls the tempo of karaoke according to the determination result.

[0004] In addition, a karaoke device installed in a nursing facility, a welfare facility, etc. (hereinafter referred to as "facility") and capable of providing various contents using music is known. When performing karaoke singing using the karaoke device described in Non-Patent Document 1, the staff of the facility checks whether the user can follow the karaoke performance and manually adjusts the performance tempo of the music piece.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Non-Patent Documents

[0006]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0007] When adjusting the performance tempo of a music piece using the technology described in Patent Document 1, it is necessary to capture the singing voice through a microphone held by the user. On the other hand, among the users in the facility, there are some who have difficulty performing karaoke singing using a microphone.

[0008] An object of the present invention is to provide a karaoke apparatus capable of automatically adjusting the performance tempo of a music piece without using the singing voice of the user.

Means for Solving the Problems

[0009] A first storage unit that stores, in association with each other, a plurality of mouth shape images showing the mouth shape when pronouncing vowels and the mouth shape in a closed-mouth state, and mouth shape identification information indicating the mouth shape corresponding to the mouth shape image; a pronunciation time indicating the timing at which each lyric character included in the lyric data should be pronounced, a predetermined section including the pronunciation time, mouth shape identification information indicating the mouth shape to be pronounced at the pronunciation time, and a sound length indicating the ratio of the length of pronunciation of each lyric character to the length of one beat based on the performance tempo of the music; a second storage unit that stores, for each music, lyric mouth shape data in which the above are associated with each other; an acquisition unit that acquires, in association with each other, mouth shape identification information specified based on a mouth shape image extracted from a video obtained by photographing a user performing karaoke singing of a certain music, and the extraction time which is the timing at which the mouth shape image is extracted; a first determination unit that, when the first mouth shape image is extracted in a certain predetermined section of the certain music, determines whether or not the first mouth shape identification information specified based on the first mouth shape image matches the first mouth shape identification information in the certain predetermined section specified by referring to the second storage unit; a setting unit that, when it is determined by the first determination unit that they match, sets the first extraction time associated with the first mouth shape identification information as the previous extraction time, refers to the second storage unit, sets the sound length associated with the mouth shape identification information corresponding to the first mouth shape identification information as the previous sound length, and sets a value indicating match as the previous match flag, and when it is determined by the first determination unit that they do not match, sets a value indicating non-match as the previous match flag; a second determination unit that, when the (1 + n)th mouth shape image is extracted in a certain predetermined section of the certain music, determines whether or not the (1 + n)th mouth shape identification information specified based on the (1 + n)th mouth shape image matches the (1 + n)th mouth shape identification information in the certain predetermined section specified by referring to the second storage unit; a first process that, when it is determined by the second determination unit that they match and a value indicating match is set in the previous match flag, adds the pronunciation time obtained by subtracting the previous extraction time from the (1 + n)th extraction time associated with the (1 + n)th mouth shape identification information specified based on the (1 + n)th mouth shape image to a first total value, and re-sets the (1 + n)th extraction time as the previous extraction time, and when it is determined by the second determination unit that they match,When a value indicating a mismatch is set in the previous match flag, a second process is executed to reset the previous extraction time to the (1 + n)-th extraction time associated with the (1 + n)-th mouth shape identification information identified based on the (1 + n)-th mouth shape image. When it is determined by the second determination unit that there is no match and a value indicating a match is set in the previous match flag, the pronunciation time obtained by subtracting the previous extraction time from the (1 + n)-th extraction time associated with the (1 + n)-th mouth shape identification information identified based on the (1 + n)-th mouth shape image is added to the first total value, and a third process is executed to reset the previous extraction time. When it is determined by the second determination unit that there is no match and a value indicating a mismatch is set in the previous match flag, a first total value processing unit that executes a fourth process to reset the previous extraction time; when it is determined by the second determination unit that there is a match and a value indicating a match is set in the previous match flag, the previous sound length is added to the second total value, and the second storage unit is referred to. The sound length associated with the mouth shape identification information corresponding to the (1 + n)-th mouth shape identification information identified based on the (1 + n)-th mouth shape image is set as the (1 + n)-th sound length, and a fifth process is executed to reset the (1 + n)-th sound length as the previous sound length. When it is determined by the second determination unit that there is a match and a value indicating a mismatch is set in the previous match flag, the second storage unit is referred to. The sound length associated with the mouth shape identification information corresponding to the (1 + n)-th mouth shape identification information identified based on the (1 + n)-th mouth shape image is set as the (1 + n)-th sound length, the (1 + n)-th sound length is reset as the previous sound length, and a sixth process is executed to reset the value indicating a match in the previous match flag. When it is determined by the second determination unit that there is no match and a value indicating a match is set in the previous match flag, the previous sound length is added to the second total value, the previous sound length is reset, and a seventh process is executed to reset the value indicating a mismatch in the previous match flag. When it is determined by the second determination unit that there is no match and a value indicating a mismatch is set in the previous match flag, a second total value processing unit that executes an eighth process to reset the previous sound length; when the end time of the color change of the last lyric character included in a certain predetermined section arrives,Based on the first total value and the second total value obtained so far, a calculation unit that calculates the singing tempo of the user in a certain predetermined section, and a performance processing unit that adjusts the performance tempo of the next predetermined section of the certain predetermined section according to the calculated singing tempo. It is a karaoke device having. Other features of the present invention will be clarified by the description in the specification and drawings described later.

Effect of the Invention

[0010] According to the present invention, the performance tempo of a music piece can be automatically adjusted without using the singing voice of the user.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Mode for Carrying Out the Invention

[0012] <Embodiment> With reference to FIGS. 1 to 7, the karaoke device K according to the embodiment will be described.

[0013] ==Karaoke Device== The karaoke device K is a device for karaoke performance of music and for users to perform karaoke singing. As shown in FIG. 1, the karaoke device K includes a karaoke main body 10, a speaker 20, a display device 30, a remote control device 40, and a camera 50.

[0014] The karaoke main body 10 performs various controls related to karaoke performance and karaoke singing, such as performance control of the selected music and display control of lyrics, background images, etc. The speaker 20 is configured to emit sound based on the sound emission signal from the karaoke main body 10. The display device 30 is configured to display images and pictures on the screen based on the signal from the karaoke main body 10. The display device 30 may be a large display separate from the karaoke device K. The remote control device 40 is a device for performing various operations on the karaoke main body 10. The camera 50 is configured to photograph the user. The camera 50 starts photographing, for example, at the timing when the use of the karaoke device K is started or at the timing when the karaoke performance is started. The camera 50 is an example of the "photographing means". The photographing means does not necessarily have to be provided in the karaoke device K itself, such as a camera for indoor photography provided in the karaoke room where the karaoke device K is installed. Although not used in the invention according to this embodiment, the karaoke device K may be equipped with a microphone.

[0015] As shown in FIG. 2, the karaoke main body 10 according to this embodiment includes a storage means 10a, a communication means 10b, an input means 10c, a performance means 10d, and a control means 10e. Each component is connected to the bus B via an interface (not shown).

[0016] [Storage means] The storage means 10a is a large-capacity storage device that stores various data. The storage means 10a stores music data. The music data is provided with music identification information. The music identification information is information unique to each music, such as a music ID for identifying the music. The music data includes accompaniment data, reference data, section information, etc.

[0017] Accompaniment data is the data that serves as the basis for karaoke performance sounds. The accompaniment data includes tempo information indicating the tempo for karaoke performance. The tempo is set to a predetermined value (for example, 100 BPM, 120 BPM) for each piece of music. Reference data is the data used as a reference when scoring a user's karaoke singing. Interval information indicates the performance interval. The performance interval is the interval during which karaoke performance is conducted. The performance interval includes a singing interval and a non-singing interval. The singing interval is the interval (for example, the A melody, B melody, and refrain of the first verse) in which the lyrics to be sung are set in the music. The non-singing interval is the interval in which no lyrics to be sung are set in the music, such as the introduction, interlude, and coda.

[0018] Further, the storage means 10a stores lyric data, background image data such as background images displayed on the display device 30 etc. during karaoke performance, performance time data indicating the karaoke performance time for each piece of music, and attribute information of the music (information regarding the music such as the singer's name, lyricist's name, composer's name, genre, etc.).

[0019] Lyric data is the data for causing the lyric characters of each piece of music to be displayed on the display device 30 etc. The lyric data includes each displayed lyric character, the character type indicating whether the lyric character is a main text lyric or a ruby character, the start time and end time of the color change of each lyric character. Also, a predetermined interval is associated with each lyric character. The predetermined interval is a unit such as a phrase separated by a character string where consecutive syllables to be pronounced are interrupted, or a line separated by a character string for one horizontal row of lyrics to be displayed.

[0020] The start time and end time of the color change are the elapsed time when the start time of the karaoke performance is set to 0. Also, generally, there is a time lag from when the user sees the displayed lyrics until they start singing. Therefore, the start time and end time of the color change may be set to a time slightly (for example, 200 ms) ahead of the karaoke performance.

[0021] Figure 3 shows a part of the lyric data of song X. For example, for the first lyric character "a" of song X, the character type "lyric text", the start time of color change "15800ms", the end time of color change "16350ms", and the phrase F1 which is a predetermined section are associated with it.

[0022] Lyric characters are displayed based on a predetermined section. The lyric characters included in one predetermined section are displayed on the same screen. Here, generally, one phrase includes a plurality of lines. In this case, each line is displayed at a different position on the same screen. For example, in the case of song X, phrase F1 includes two lines, "atama o kumo no" and "ue ni dasi". Therefore, on the same screen, "atama o kumo no" is displayed in the upper row and "ue ni dasi" is displayed in the lower row. When the user sees the lyric characters included in phrase F1 of song X being displayed, they try to sing karaoke in accordance with the performance tempo of song X from the first lyric character "a" to the last lyric character "si".

[0023] In the present embodiment, a part of the storage area of the storage means 10a functions as a first storage unit 11a and a second storage unit 12a.

[0024] (First storage unit) The first storage unit 11a stores a plurality of mouth shape images in association with mouth shape identification information.

[0025] The mouth shape image is an image showing the mouth shape when pronouncing vowels and the mouth shape in a closed state. Specifically, the mouth shape image corresponds to the mouth shapes when pronouncing the five vowels "a", "i", "u", "e", "o" and the mouth shape of the closed state "n". That is, there are six types of mouth shape images (see, for example, FIG. 5 of Japanese Patent Laid-Open No. 2008-310382).

[0026] The mouth shape identification information is information indicating the mouth shape corresponding to the mouth shape image. When the mouth shape image is "a", "i", "u", "e", "o", "n", the mouth shape identification information is "A", "I", "U", "E", "O", "N", respectively.

[0027] (Second storage unit) The second storage unit 12a stores the lyric mouth shape data for each piece of music. The lyric mouth shape data is data that associates a pronunciation time, a predetermined section including the pronunciation time, mouth shape identification information indicating the mouth shape when pronouncing at that timing, and a sound length.

[0028] The pronunciation time is the time indicating the timing at which each lyric character included in the lyric data should be pronounced. The pronunciation time can be indicated by the elapsed time when the start time of the karaoke performance is set to 0. The predetermined section is, similar to the predetermined section included in the lyric data, a phrase or the like. The mouth shape identification information is, similar to the mouth shape identification information stored in the first storage unit 11a, "A", "I", "U", "E", "O", "N", etc. The sound length indicates the ratio of the length of pronouncing each lyric character to the length of one beat based on the performance tempo of the music.

[0029] FIG. 4 shows a part of the lyric mouth shape data of music X. Assume that the performance tempo of music X is 120 bpm (that is, the length of one beat is 500 ms). In this example, the predetermined section corresponds to a phrase. The lyric mouth shape data included in one phrase is numbered in the order of pronunciation along with karaoke singing.

[0030] For example, the sound length associated with the first mouth shape identification information "A" (pronunciation time 16000 ms) of phrase F1 is "1.5". Therefore, the pronunciation time of the mouth shape identification information "A" (that is, the time to be pronounced in the mouth shape corresponding to the mouth shape identification information "A") is 750 ms. Also, the sound length associated with the second mouth shape identification information "IA" (pronunciation time 16550 ms) of phrase F1 is "0.5". Therefore, the pronunciation time of the mouth shape identification information "IA" (that is, the time to be pronounced in the mouth shape corresponding to the mouth shape identification information "IA") is 250 ms.

[0031] [Communication means · Input means] The communication means 10b provides an interface for communicating with the remote control device 40. The input means 10c is configured for a facility staff or the like to perform an instruction input. The input means 10c is a button or the like provided on the karaoke main body 10. Alternatively, the remote control device 40 may function as the input means 10c.

[0032] [Performance means] The performance means 10d performs karaoke performance of a music piece based on the control of the control means 10e. The performance means 10d includes a sound source, a mixer, an amplifier, etc. (all not shown in the figure).

[0033] [Control means] The control means 10e performs various controls in the karaoke device K. The control means 10e includes a CPU and a memory (both not shown in the figure). The CPU realizes various functions by executing a program stored in the memory.

[0034] In the present embodiment, when the CPU executes a program stored in the memory, the control means 10e functions as an acquisition unit 100, a first determination unit 200, a setting unit 300, a second determination unit 400, a first total value processing unit 500, a second total value processing unit 600, a calculation unit 700, and a performance processing unit 800 (see Figure 2).

[0035] (Acquisition unit) The acquisition unit 100 acquires, in association with each other, the lip shape identification information specified based on a lip shape image extracted from a video obtained by photographing a user who performs karaoke singing of a certain music piece, and the extraction time which is the timing when the lip shape image is extracted.

[0036] The acquisition unit 100 extracts a mouth shape image from a video obtained by filming a user singing karaoke of a song. The mouth shape image is an image that makes it possible to distinguish the shape of the user's mouth. The process of extracting the mouth shape image from the user's video can use known technology. The extraction time is the time when the mouth shape image is extracted. The extraction time is obtained by referring to the elapsed time, with the start time of the karaoke performance measured by the performance means 10d set as 0, when the acquisition unit 100 extracts the mouth shape image from the filmed video.

[0037] For example, assume that a user sings a karaoke piece selected by a facility staff member. The karaoke device K starts playing the selected piece. The camera 50 starts shooting the user as the karaoke performance of the piece starts. The camera 50 sequentially outputs the captured images to the karaoke unit 10.

[0038] The karaoke device K causes the display device 30 to display lyrics for each predetermined section based on the lyrics data of the selected song while the karaoke is being performed. The user vocalizes each character in accordance with the displayed lyrics (in this embodiment, the user vocalizes all the displayed lyrics).

[0039] The acquisition unit 100 extracts a mouth shape image from a video obtained by shooting a user singing karaoke of a song. The acquisition unit 100 refers to the first storage unit 11a and identifies one piece of mouth shape identification information associated with the same mouth shape image as the extracted mouth shape image. The acquisition unit 100 also identifies one time when the mouth shape image is extracted as an extraction time.

[0040] The acquisition unit 100 associates the identified piece of mouth shape identification information with the extraction time and acquires it. The acquisition unit 100 repeats the same process every time a mouth shape image is extracted.

[0041] (First judgment section) When the first lip image is extracted in a certain predetermined section of a certain piece of music, the first determination unit 200 determines whether the first lip identification information specified based on the first lip image matches the first lip identification information in the certain predetermined section specified by referring to the second storage unit 12a.

[0042] The first lip identification information is lip identification information specified by the acquisition unit 100 based on the first lip image.

[0043] When the first (i.e., the first) lip image is extracted in a certain predetermined section (for example, the first phrase) of a certain piece of music, the first determination unit 200 compares the one lip identification information (i.e., the first lip identification information) specified based on the first lip image with the first lip identification information in the certain predetermined section specified by referring to the second storage unit 12a, and determines whether they match or not. The first determination unit 200 outputs the determination result to the setting unit 300.

[0044] For example, when the first lip identification information is "A" and the first lip identification information in the certain predetermined section specified by referring to the second storage unit 12a is "A", the first determination unit 200 outputs a determination result of "match" to the setting unit 300. On the other hand, when the first lip identification information is "A" and the first lip identification information in the certain predetermined section specified by referring to the second storage unit 12a is "U", the first determination unit 200 outputs a determination result of "mismatch" to the setting unit 300.

[0045] (Setting unit) When it is determined by the first determination unit 200 that they match, the setting unit 300 sets the first extraction time associated with the first lip identification information as the previous extraction time, refers to the second storage unit 12a, sets the sound length associated with the lip identification information corresponding to the first lip identification information as the previous sound length, and sets a value indicating a match as the previous match flag. When it is determined by the first determination unit 200 that they do not match, the setting unit 300 sets a value indicating a mismatch as the previous match flag.

[0046] The first extraction time is the time associated with the first lip shape identification information. The value indicating a match and the value indicating a mismatch are values for identifying the determination result by the first determination unit 200, such as the value "1" indicating a match and the value "0" indicating a mismatch.

[0047] When the determination result of "match" is output from the first determination unit 200, the setting unit 300 specifies the extraction time associated with the first lip shape identification information as the first extraction time. The setting unit 300 sets the first extraction time as the previous extraction time.

[0048] Also, when the determination result of "match" is output from the first determination unit 200, the setting unit 300 refers to the second storage unit 12a and specifies the sound length associated with the lip shape identification information corresponding to the first lip shape identification information. The setting unit 300 sets the specified sound length as the previous sound length.

[0049] Furthermore, when the determination result of "match" is output from the first determination unit 200, the setting unit 300 sets the value indicating a match as the previous match flag.

[0050] On the other hand, when the determination result of "mismatch" is output from the first determination unit 200, the setting unit 300 sets the value indicating a mismatch as the previous match flag. In this case, the setting unit 300 does not set the previous extraction time and the previous sound length.

[0051] Note that at the start point of karaoke performance of a certain piece of music, the previous extraction time and the previous sound length are "0", and the previous match flag is not set. Also, the previous extraction time, the previous sound length, and the previous match flag are reset every predetermined interval (that is, the previous extraction time and the previous sound length are reset to "0" every predetermined interval, and the setting of the previous match flag is cancelled every predetermined interval).

[0052] (Second determination unit) When the (1 + n)-th mouth shape image is extracted in a certain predetermined section of a certain piece of music, the second determination unit 400 determines whether or not the (1 + n)-th mouth shape identification information specified based on the (1 + n)-th mouth shape image matches the (1 + n)-th mouth shape identification information in the certain predetermined section specified with reference to the second storage unit 12a.

[0053] "n" is a natural number. The (1 + n)-th mouth shape identification information is the mouth shape identification information specified by the acquisition unit 100 based on the (1 + n)-th mouth shape image.

[0054] When the (1 + n)-th mouth shape image is extracted in a certain predetermined section (for example, the first phrase) of a certain piece of music, the second determination unit 400 compares the one mouth shape identification information (that is, the (1 + n)-th mouth shape identification information) specified based on the (1 + n)-th mouth shape image with the (1 + n)-th mouth shape identification information in the certain predetermined section specified with reference to the second storage unit 12a, and determines whether they match or not. The second determination unit 400 outputs the determination result to the first total value processing unit 500 and the second total value processing unit 600.

[0055] Each time a mouth shape image is extracted in a certain predetermined section of a certain piece of music, the second determination unit 400 repeats the same process.

[0056] (First total value processing unit) When it is determined by the second determination unit 400 that they match and the value indicating a match is set in the previous match flag, the first total value processing unit 500 executes the first process. When it is determined by the second determination unit 400 that they match and the value indicating a non-match is set in the previous match flag, the first total value processing unit 500 executes the second process. When it is determined by the second determination unit 400 that they do not match and the value indicating a match is set in the previous match flag, the first total value processing unit 500 executes the third process. When it is determined by the second determination unit 400 that they do not match and the value indicating a non-match is set in the previous match flag, the first total value processing unit 500 executes the fourth process.

[0057] The first process is to add the pronunciation time obtained by subtracting the previous extraction time from the (1 + n)-th extraction time associated with the (1 + n)-th mouth shape identification information specified based on the (1 + n)-th mouth shape image to the first total value, and to reset the (1 + n)-th extraction time as the previous extraction time. The second process is to reset the (1 + n)-th extraction time associated with the (1 + n)-th mouth shape identification information specified based on the (1 + n)-th mouth shape image as the previous extraction time. The third process is to add the pronunciation time obtained by subtracting the previous extraction time from the (1 + n)-th extraction time associated with the (1 + n)-th mouth shape identification information specified based on the (1 + n)-th mouth shape image to the first total value, and to reset the previous extraction time. The fourth process is to reset the previous extraction time. The first total value processing unit 500 performs any of the above processes each time the (1 + n)-th mouth shape image is extracted in a predetermined section.

[0058] "n" is a natural number. The (1 + n)-th mouth shape identification information is the mouth shape identification information specified by the acquisition unit 100 based on the (1 + n)-th mouth shape image as described above. The (1 + n)-th extraction time is the time associated with the (1 + n)-th mouth shape identification information. The first total value is the value obtained by summing the pronunciation time (the time when the user pronounced) in a predetermined section. At the start point of the karaoke performance of a certain piece of music, the first total value is "0". Also, the first total value is reset for each predetermined section (that is, reset to "0" for each predetermined section).

[0059] -Details of the first process- When the first lip image is extracted in a certain predetermined section of a certain piece of music, assume that the first determination unit 200 determines that the first lip identification information specified based on the first lip image matches the first lip identification information in the certain predetermined section specified by referring to the second storage unit 12a. Then, assume that the setting unit 300 sets the previous extraction time ET1 and the previous sound length TL1, and sets the value "1" indicating the match to the previous match flag. After that, when the second (n = 1) lip image is extracted in a certain predetermined section, assume that the second determination unit 400 determines that the second lip identification information specified based on the second lip image matches the second lip identification information in the certain predetermined section specified by referring to the second storage unit 12a.

[0060] In this case, the first total value processing unit 500 specifies the extraction time associated with the second lip identification information (the lip identification information specified by the acquisition unit 100 based on the second lip image) as the second extraction time ET2. The first total value processing unit 500 subtracts the previous extraction time ET1 from the second extraction time ET2 to obtain the pronunciation time PT1, and adds it (0 + PT1) to the first total value S1. Also, the first total value processing unit 500 re-sets the second extraction time ET2 as the previous extraction time (that is, the previous extraction time becomes "ET2").

[0061] -Details of the second process- When the first lip image is extracted in a certain predetermined section of a certain piece of music, assume that the first determination unit 200 determines that the first lip identification information specified based on the first lip image does not match the first lip identification information in the certain predetermined section specified by referring to the second storage unit 12a. Then, assume that the setting unit 300 sets the value "0" indicating the non-match to the previous match flag. After that, when the second (n = 1) lip image is extracted in a certain predetermined section, assume that the second determination unit 400 determines that the second lip identification information specified based on the second lip image matches the second lip identification information in the certain predetermined section specified by referring to the second storage unit 12a.

[0062] In this case, the first total value processing unit 500 specifies the extraction time associated with the second mouth shape identification information (the mouth shape identification information specified by the acquisition unit 100 based on the second mouth shape image) as the second extraction time ET2. The first total value processing unit 500 resets the second extraction time ET2 as the previous extraction time (that is, the previous extraction time becomes "ET2"). In this case, the process of adding the previous extraction time ET2 to the first total value S1 is not performed.

[0063] - Details of the third process - When the first mouth shape image is extracted in a certain predetermined section of a certain piece of music, and it is assumed that the first determination unit 200 determines that the first mouth shape identification information specified based on the first mouth shape image matches the first mouth shape identification information in the certain predetermined section specified by referring to the second storage unit 12a. Then, it is assumed that the setting unit 300 sets the previous extraction time ET1 and the previous sound length TL1, and sets the value "1" indicating the match to the previous match flag. After that, when the second (n = 1) mouth shape image is extracted in a certain predetermined section, it is assumed that the second determination unit 400 determines that the second mouth shape identification information specified based on the second mouth shape image does not match the second mouth shape identification information in the certain predetermined section specified by referring to the second storage unit 12a.

[0064] In this case, the first total value processing unit 500 specifies the extraction time associated with the second mouth shape identification information (the mouth shape identification information specified by the acquisition unit 100 based on the second mouth shape image) as the second extraction time ET2. The first total value processing unit 500 subtracts the previous extraction time ET1 from the second extraction time ET2 to obtain the pronunciation time PT1, and adds it (0 + PT1) to the first total value S1. Also, the first total value processing unit 500 resets the previous extraction time (that is, the previous extraction time becomes "0").

[0065] - Details of the fourth process - When the first mouth shape image is extracted in a certain predetermined section of a certain piece of music, assume that the first determination unit 200 determines that the first mouth shape identification information specified based on the first mouth shape image does not match the first mouth shape identification information in the certain predetermined section specified with reference to the second storage unit 12a. Then, assume that the setting unit 300 sets the value "0" indicating a mismatch to the previous match flag. After that, when the second (n = 1) mouth shape image is extracted in a certain predetermined section, assume that the second determination unit 400 determines that the second mouth shape identification information specified based on the second mouth shape image does not match the second mouth shape identification information in the certain predetermined section specified with reference to the second storage unit 12a.

[0066] In this case, the first total value processing unit 500 resets the previous extraction time (that is, the previous extraction time becomes "0"). In this case, the process of adding the previous extraction time ET2 to the first total value S1 is not performed.

[0067] (Second total value processing unit) When it is determined by the second determination unit 400 that they match and a value indicating a match is set in the previous match flag, the second total value processing unit 600 executes the fifth process. When it is determined by the second determination unit 400 that they match and a value indicating a mismatch is set in the previous match flag, the second total value processing unit 600 executes the sixth process. When it is determined by the second determination unit 400 that they do not match and a value indicating a match is set in the previous match flag, the second total value processing unit 600 executes the seventh process. When it is determined by the second determination unit 400 that they do not match and a value indicating a mismatch is set in the previous match flag, the second total value processing unit 600 executes the eighth process.

[0068] The fifth process is to add the previous sound length to the second total value, refer to the second storage unit 12a, set the sound length associated with the mouth shape identification information corresponding to the (1 + n)-th mouth shape identification information specified based on the (1 + n)-th mouth shape image as the (1 + n)-th sound length, and reset the previous sound length with the (1 + n)-th sound length. The sixth process is to refer to the second storage unit 12a, set the sound length associated with the mouth shape identification information corresponding to the (1 + n)-th mouth shape identification information specified based on the (1 + n)-th mouth shape image as the (1 + n)-th sound length, reset the previous sound length with the (1 + n)-th sound length, and reset the value indicating a match to the previous match flag. The seventh process is to add the previous sound length to the second total value, reset the previous sound length, and reset the value indicating a non-match to the previous match flag. The eighth process is to reset the previous sound length. The second total value processing unit 600 performs any one of the above processes each time the (1 + n)-th mouth shape image is extracted in a predetermined section.

[0069] "n" is a natural number. The (1 + n)-th mouth shape identification information is the mouth shape identification information specified by the acquisition unit 100 based on the (1 + n)-th mouth shape image as described above. The (1 + n)-th sound length is the sound length associated with the mouth shape identification information corresponding to the (1 + n)-th mouth shape identification information. The second total value is the value obtained by summing the sound lengths in a predetermined section. At the start point of the karaoke performance of a certain piece of music, the second total value is "0". Also, the second total value is reset for each predetermined section (i.e., reset to "0" for each predetermined section).

[0070] -Details of the Fifth Process- When the first lip image is extracted in a certain predetermined section of a certain piece of music, assume that the first determination unit 200 determines that the first lip identification information specified based on the first lip image matches the first lip identification information in the certain predetermined section specified with reference to the second storage unit 12a. Then, assume that the setting unit 300 sets the previous extraction time ET1 and the previous sound length TL1, and sets the value "1" indicating a match to the previous match flag. After that, when the second (n = 1) lip image is extracted in a certain predetermined section, assume that the second determination unit 400 determines that the second lip identification information specified based on the second lip image matches the second lip identification information in the certain predetermined section specified with reference to the second storage unit 12a.

[0071] In this case, the second total value processing unit 600 adds the previous sound length TL1 to the second total value S2 (0 + TL1). Also, the second total value processing unit 600 refers to the second storage unit 12a and specifies the sound length associated with the lip identification information corresponding to the second lip identification information (the lip identification information specified by the acquisition unit 100 based on the second lip image) as the second sound length TL2. The second total value processing unit 600 re-sets the second sound length TL2 as the previous sound length (that is, the previous sound length becomes "TL2"). Note that in this case, the previous match flag is set with the value "1" indicating a match. Therefore, the second total value processing unit 600 does not re-set the previous match flag.

[0072] -Details of the Sixth Process- When the first mouth shape image is extracted in a certain predetermined section of a certain piece of music, assume that the first determination unit 200 determines that the first mouth shape identification information specified based on the first mouth shape image does not match the first mouth shape identification information in the certain predetermined section specified by referring to the second storage unit 12a. Then, assume that the setting unit 300 sets the value "0" indicating a mismatch to the previous match flag. After that, when the second (n = 1) mouth shape image is extracted in a certain predetermined section, assume that the second determination unit 400 determines that the second mouth shape identification information specified based on the second mouth shape image matches the second mouth shape identification information in the certain predetermined section specified by referring to the second storage unit 12a.

[0073] In this case, the second total value processing unit 600 refers to the second storage unit 12a and specifies the sound length associated with the mouth shape identification information corresponding to the second mouth shape identification information (the mouth shape identification information specified by the acquisition unit 100 based on the second mouth shape image) as the second sound length TL2. The second total value processing unit 600 re - sets the second sound length TL2 as the previous sound length (that is, the previous sound length becomes "TL2"). Also, the second total value processing unit 600 re - sets the value "1" indicating a match to the previous match flag (that is, the previous match flag becomes "1"). In this case, the process of adding the previous sound length TL2 to the second total value S2 is not performed.

[0074] -Details of the Seventh Process- When the first lip image is extracted in a certain predetermined section of a certain piece of music, assume that the first determination unit 200 determines that the first lip identification information specified based on the first lip image matches the first lip identification information in the certain predetermined section specified by referring to the second storage unit 12a. Then, assume that the setting unit 300 sets the previous extraction time ET1 and the previous sound length TL1, and sets the value "1" indicating a match to the previous match flag. After that, when the second (n = 1) lip image is extracted in a certain predetermined section, assume that the second determination unit 400 determines that the second lip identification information specified based on the second lip image does not match the second lip identification information in the certain predetermined section specified by referring to the second storage unit 12a.

[0075] In this case, the second total value processing unit 600 adds the previous sound length TL1 to the second total value S2 (0 + TL1). Also, the second total value processing unit 600 resets the previous sound length (that is, the previous sound length becomes "0"). Further, the second total value processing unit 600 resets the previous match flag to the value "0" indicating a non-match (that is, the previous match flag becomes "0").

[0076] -Details of the Eighth Process- When the first lip image is extracted in a certain predetermined section of a certain piece of music, assume that the first determination unit 200 determines that the first lip identification information specified based on the first lip image does not match the first lip identification information in the certain predetermined section specified by referring to the second storage unit 12a. Then, assume that the setting unit 300 sets the value "0" indicating a non-match to the previous match flag. After that, when the second (n = 1) lip image is extracted in a certain predetermined section, assume that the second determination unit 400 determines that the second lip identification information specified based on the second lip image does not match the second lip identification information in the certain predetermined section specified by referring to the second storage unit 12a.

[0077] In this case, the second total value processing unit 600 resets the previous note length (that is, the previous note length becomes "0"). In this case, the process of adding the previous note length TL2 to the second total value S2 is not performed.

[0078] (Calculation unit) When the end time of the color change of the last lyric character included in a predetermined section arrives, the calculation unit 700 calculates the singing tempo of the user in a predetermined section based on the first total value and the second total value obtained so far.

[0079] Whether or not the end time of the color change of the last lyric character included in a predetermined section has arrived can be determined by referring to the lyric data of a certain piece of music (the end time associated with the lyric characters included in a predetermined section).

[0080] Known techniques can be used to calculate the singing tempo. Specifically, the calculation unit 700 calculates the length of one beat in the user's karaoke singing in a predetermined section by dividing the first total value by the second total value. The calculation unit 700 calculates the singing tempo of the user in a predetermined section from the calculated length of one beat and the performance tempo of a certain piece of music.

[0081] (Performance processing unit) The performance processing unit 800 controls the performance means 10d to perform the karaoke performance of the music.

[0082] The staff of the facility where the karaoke device K is installed operates the remote control device 40 to select a music piece for which the karaoke performance is to be performed. The remote control device 40 registers the music identification information of the selected music piece in the reservation queue.

[0083] Based on the music identification information registered in the reservation queue, the performance processing unit 800 acquires the accompaniment data of the music from the storage means 10a. The performance processing unit 800 reproduces the acquired accompaniment data and controls the performance means 10d to emit the karaoke performance sound from the speaker 20. The user can perform karaoke singing in accordance with the karaoke performance sound emitted from the speaker 20.

[0084] The performance processing unit 800 according to this embodiment adjusts the performance tempo of the next predetermined section in a predetermined section according to the calculated singing tempo. The adjustment of the performance tempo is to increase the performance tempo or decrease the performance tempo with respect to the performance tempo preset in the music piece.

[0085] The adjustment of the performance tempo can be performed by various methods. For example, the performance processing unit 800 determines whether the singing tempo of the user in a predetermined section calculated by the calculation unit 700 is within a preset threshold. When it is not within the threshold (that is, when the deviation between the singing tempo and the performance tempo is large), the performance processing unit 800 makes an adjustment from the next predetermined section of a predetermined section so that the performance tempo of the music piece matches the calculated singing tempo of the user.

[0086] Alternatively, the performance processing unit 800 obtains the difference between the singing tempo and the performance tempo of a music piece. The performance processing unit 800 adjusts the performance tempo of the music piece in the next predetermined section according to the obtained difference.

[0087] As described above, the lyric characters included in one predetermined section are displayed on the same screen. Therefore, when the karaoke device K performs karaoke performance of the next phrase, the lyric characters included in the previous phrase are erased and the lyric characters included in the next phrase are displayed. That is, according to the karaoke device K according to this embodiment, the performance tempo of the music piece can be adjusted at the timing when the display of the lyric characters is switched.

[0088] ==Regarding the processing of the karaoke device K== Next, a specific example of the processing of the karaoke device K in this embodiment will be described with reference to Fig. 5 to Fig. 7. Fig. 5 to Fig. 7 are flow charts showing the processing of the karaoke device K. In this example, it is assumed that a user U sings karaoke alone without using a microphone in a facility where the karaoke device K is installed. It is also assumed that the lyric data of the song X shown in Fig. 3 is stored in the storage means 10a, and the lyric mouth shape data of the song X shown in Fig. 4 is stored in the second storage unit 12a. It is also assumed that the performance tempo of the song X is 120 bpm.

[0089] A staff member of the facility operates the remote control device 40 to select a piece of music X suitable for karaoke singing by the user U. The karaoke device K (performance processing unit 800) starts the karaoke performance of the selected piece of music X (karaoke performance start; step 10).

[0090] The karaoke device K displays lyrics for each predetermined section based on the lyrics data of the song X on the display device 30 as the song X is being played (displaying lyrics for each predetermined section; step 11). The user U sings each character of the lyrics displayed. The camera 50 starts filming the user U as the karaoke performance of the song X begins. The camera 50 outputs the captured images to the karaoke main unit 10 in sequence. The karaoke main unit 10 acquires the image of the user U singing karaoke (acquires the user's image; step 12).

[0091] The acquisition unit 100 acquires mouth shape identification information determined based on a mouth shape image extracted from a video obtained by filming a user U singing karaoke of song X, in association with the extraction time, which is the timing at which the mouth shape image was extracted (acquires mouth shape identification information and extraction time; step 13).

[0092] When the first lip image is extracted in a certain predetermined section of music X, the first determination unit 200 determines whether the first lip identification information specified based on the first lip image matches the first lip identification information in the certain predetermined section specified by referring to the second storage unit 12a (determines whether the first lip identification information matches the first lip identification information. Step 14).

[0093] When it is determined in step 14 that they match, the setting unit 300 sets the first extraction time associated with the first lip identification information as the previous extraction time, refers to the second storage unit 12a, and sets the sound length associated with the lip identification information corresponding to the first lip identification information as the previous sound length, and sets a value indicating a match as the previous match flag. When it is determined in step 14 that they do not match, the setting unit 300 sets a value indicating a mismatch as the previous match flag (when it is determined that they match, sets the previous extraction time, the previous sound length, and the previous match flag (match). When it is determined that they do not match, sets the previous match flag (mismatch). Step 15).

[0094] When the (1 + n)th lip image is extracted in a certain predetermined section of music X, the second determination unit 400 determines whether the (1 + n)th lip identification information specified based on the (1 + n)th lip image matches the (1 + n)th lip identification information in the certain predetermined section specified by referring to the second storage unit 12a (determines whether the (1 + n)th lip identification information matches the (1 + n)th lip identification information. Step 16).

[0095] The first total value processing unit 500 executes any one of the first to fourth processes according to the result of step 16 (executes any one of the first to fourth processes. Step 17). Also, the second total value processing unit 600 executes any one of the fifth to eighth processes according to the result of step 16 (executes any one of the fifth to eighth processes. Step 18).

[0096] More specifically, as shown in FIG. 7, when it is determined that there is a match in step 16 and a value indicating a match is set in the previous match flag, the first total value processing unit 500 adds the pronunciation time obtained by subtracting the previous extraction time from the (1 + n)-th extraction time associated with the (1 + n)-th mouth shape identification information specified based on the (1 + n)-th mouth shape image to the first total value, and executes a first process of resetting the (1 + n)-th extraction time as the previous extraction time (add the pronunciation time to the first total value and reset the previous extraction time. Step 17a). Further, the second total value processing unit 600 adds the previous sound length to the second total value, refers to the second storage unit 12a, sets the sound length associated with the mouth shape identification information corresponding to the (1 + n)-th mouth shape identification information specified based on the (1 + n)-th mouth shape image as the (1 + n)-th sound length, and executes a fifth process of resetting the (1 + n)-th sound length as the previous sound length (add the sound length to the second total value and reset the previous sound length. Step 18a).

[0097] Also, when it is determined that there is a match in step 16 and a value indicating a non-match is set in the previous match flag, the first total value processing unit 500 executes a second process of resetting the (1 + n)-th extraction time associated with the (1 + n)-th mouth shape identification information specified based on the (1 + n)-th mouth shape image as the previous extraction time (reset the previous extraction time. Step 17b). Further, the second total value processing unit 600 refers to the second storage unit 12a, sets the sound length associated with the mouth shape identification information corresponding to the (1 + n)-th mouth shape identification information specified based on the (1 + n)-th mouth shape image as the (1 + n)-th sound length, resets the (1 + n)-th sound length as the previous sound length, and executes a sixth process of resetting a value indicating a match in the previous match flag (reset the previous sound length and the previous match flag (match). Step 18b).

[0098] Also, when it is determined in step 16 that there is no match and a value indicating a match is set in the previous match flag, the first total value processing unit 500 adds the pronunciation time obtained by subtracting the previous extraction time from the (1 + n)th extraction time associated with the (1 + n)th mouth shape identification information specified based on the (1 + n)th mouth shape image to the first total value, and executes a third process of resetting the previous extraction time (add the pronunciation time to the first total value and reset the previous extraction time. Step 17c). Also, the second total value processing unit 600 executes a seventh process of adding the previous sound length to the second total value, resetting the previous sound length, and resetting a value indicating a mismatch to the previous match flag (add the sound length to the second total value, reset the previous sound length, and reset the previous match flag (mismatch). Step 18c).

[0099] Also, when it is determined in step 16 that there is no match and a value indicating a mismatch is set in the previous match flag, the first total value processing unit 500 executes a fourth process of resetting the previous extraction time (reset the previous extraction time. Step 17d). Also, the second total value processing unit 600 executes an eighth process of resetting the previous sound length (reset the previous sound length. Step 18d).

[0100] The karaoke device K repeats the processes of steps 12 to 18 until the end time of the color change of the last lyric character included in a certain predetermined section arrives (when N in step 19) (however, the processes of steps 14 and 15 are skipped after the second time).

[0101] When the end time of the color change of the last lyric character included in a certain predetermined section arrives (when Y in step 19), the calculation unit 700 calculates the singing tempo of the user U in a certain predetermined section based on the first total value and the second total value obtained so far (calculate the singing tempo in a certain predetermined section. Step 20).

[0102] The performance processing unit 800 adjusts the performance tempo of the next predetermined section in a predetermined section according to the singing tempo calculated in step 20 (adjust the performance tempo of the next predetermined section. Step 21).

[0103] The karaoke device K repeats the processing from step 11 to step 21 for each predetermined section until the karaoke performance ends (when it is Y in step 22). Note that the karaoke device K may repeat the processing from step 11 to step 21 until a certain point (for example, until the performance section of the first song ends) instead of until the karaoke performance ends. Also, the processing by the first total value processing unit 500 and the second total value processing unit 600 may be performed in the reverse order or simultaneously. In this example, an example having two total value processing units (the first total value processing unit 500 and the second total value processing unit 600) is described, but one total value processing unit may be provided, and the one total value processing unit may execute the processing of both the first total value processing unit 500 and the second total value processing unit 600.

[0104] Here, the adjustment of the performance tempo will be specifically described. At the start point of the karaoke performance of song X, the previous extraction time, the previous sound length, the first total value, and the second total value are all "0". Also, the previous match flag is not set. In this example, it is assumed that the predetermined section is a phrase.

[0105] - Specific examples of step 17a and step 18a - The karaoke device K causes the lyric character "lift your head above the clouds" of the phrase F1 based on the lyric data of song X to be displayed on the display device 30 along with the karaoke performance of song X. The user U makes a sound of "a" in accordance with the displayed lyric character. The camera 50 outputs the video I1 of the user U making the sound of "a" to the karaoke main body 10.

[0106] The acquisition unit 100 extracts the lip image MI1 from the output video I1. Then, the acquisition unit 100 refers to the first storage unit 11a and identifies one lip identification information "A" associated with the same lip image as the extracted lip image MI1. Also, the acquisition unit 100 identifies a certain time "16014 ms", which is the timing when the lip image MI1 is extracted, as the extraction time.

[0107] The acquisition unit 100 acquires the identified one lip identification information "A" and the extraction time "16014 ms" in association with each other.

[0108] The first determination unit 200 determines whether or not the identified first lip identification information "A" matches the first lip identification information in the phrase F1 of the music X identified by referring to the second storage unit 12a. As shown in FIG. 4, the first lip identification information in the phrase F1 of the music X is "A". Therefore, the first determination unit 200 outputs a determination result of "match" to the setting unit 300.

[0109] The setting unit 300 sets the extraction time "16014 ms" associated with the identified first lip identification information "A" as the previous extraction time ET1, refers to the second storage unit 12a, sets the sound length "1.5" associated with the lip identification information corresponding to the first lip identification information "A" as the previous sound length TL1, and sets a value "1" indicating a match as the previous match flag MF.

[0110] Next, the camera 50 outputs a video I2 of the user U who utters "ta" in accordance with the displayed lyric characters to the karaoke machine 10.

[0111] The acquisition unit 100 extracts the lip image MI2 from the output video I2. Then, the acquisition unit 100 refers to the first storage unit 11a and identifies one lip identification information "IA" associated with the same lip image as the extracted lip image MI2. Also, the acquisition unit 100 identifies a certain time "16910 ms", which is the timing when the lip image MI2 is extracted, as the extraction time.

[0112] The acquisition unit 100 acquires by associating the identified single lip shape identification information "IA" with the extraction time "16910 ms".

[0113] When the second lip shape image is extracted in the phrase F1 of the music X, the second determination unit 400 compares the single lip shape identification information "IA" (i.e., the second lip shape identification information) specified based on the second lip shape image MI2 with the second lip shape identification information in the phrase F1 of the music X specified by referring to the second storage unit 12a, and determines whether they match or not. As shown in FIG. 4, the second lip shape identification information in the phrase F1 of the music X is "IA". Therefore, the second determination unit 400 outputs the determination result of "match" to the first total value processing unit 500 and the second total value processing unit 600.

[0114] The first total value processing unit 500 specifies the extraction time "16910 ms" associated with the lip shape identification information "IA" as the extraction time ET2. The first total value processing unit 500 subtracts the previous extraction time ET1 (16014 ms) from the extraction time ET2 (16910 ms) to obtain the pronunciation time PT1 (896 ms), and adds it to the first total value S1 (0 ms + 896 ms = 896 ms). Further, the first total value processing unit 500 resets the extraction time ET2 as the previous extraction time (i.e., the previous extraction time becomes "ET2 (16910 ms)").

[0115] Also, the second total value processing unit 600 adds the previous sound length TL1 (1.5) to the second total value S2 (0 + 1.5 = 1.5). Further, the second total value processing unit 600 refers to the second storage unit 12a and specifies the sound length "0.5" associated with the lip shape identification information corresponding to the lip shape identification information "IA" as the sound length TL2. The second total value processing unit 600 resets the sound length TL2 as the previous sound length (i.e., the previous sound length becomes "TL2 (0.5)").

[0116] - Specific examples of step 17b and step 18b - The karaoke device K causes the display device 30 to display the lyrics text "lift your head above the clouds" of the phrase F1 based on the lyrics data of the music X along with the karaoke performance of the music X. The user U makes a pronunciation of "ta" which is different from the displayed lyrics text. The camera 50 outputs a video I1 of the user U making the pronunciation of "ta" to the karaoke main body 10.

[0117] The acquisition unit 100 extracts a mouth shape image MI1 from the output video I1. Then, the acquisition unit 100 refers to the first storage unit 11a and identifies one mouth shape identification information "IA" associated with the same mouth shape image as the extracted mouth shape image MI1. Also, the acquisition unit 100 identifies a certain time "16014 ms", which is the timing when the mouth shape image MI1 is extracted, as the extraction time.

[0118] The acquisition unit 100 acquires the identified one mouth shape identification information "IA" and the extraction time "16014 ms" by associating them.

[0119] The first determination unit 200 determines whether or not the identified first mouth shape identification information "IA" matches the first mouth shape identification information in the phrase F1 of the music X identified by referring to the second storage unit 12a. As shown in FIG. 4, the first mouth shape identification information in the phrase F1 of the music X is "A". Therefore, the first determination unit 200 outputs a determination result of "not matching" to the setting unit 300.

[0120] The setting unit 300 sets a value "0" indicating non - match to the previous match flag MF.

[0121] Next, the camera 50 outputs a video I2 of the user U making a pronunciation of "ta" in accordance with the displayed lyrics text to the karaoke main body 10.

[0122] The acquisition unit 100 extracts the lip image MI2 from the output video I2. Then, the acquisition unit 100 refers to the first storage unit 11a and identifies one lip identification information "IA" associated with the same lip image as the extracted lip image MI2. Also, the acquisition unit 100 identifies the extraction time "16910 ms", which is the timing when the lip image MI2 is extracted, as the extraction time.

[0123] The acquisition unit 100 acquires the identified one lip identification information "IA" and the extraction time "16910 ms" in association with each other.

[0124] When the second lip image is extracted in the phrase F1 of the music X, the second determination unit 400 compares one lip identification information "IA" (i.e., the second lip identification information) specified based on the second lip image MI2 with the second lip identification information in the phrase F1 of the music X specified by referring to the second storage unit 12a, and determines whether they match or not. As shown in FIG. 4, the second lip identification information in the phrase F1 of the music X is "IA". Therefore, the second determination unit 400 outputs the determination result of "match" to the first total value processing unit 500 and the second total value processing unit 600.

[0125] The first total value processing unit 500 identifies the extraction time "16910 ms" associated with the lip identification information "IA" as the extraction time ET2. The first total value processing unit 500 resets the extraction time ET2 (16910 ms) as the previous extraction time (i.e., the previous extraction time becomes "ET2 (16910 ms)").

[0126] Also, the second total value processing unit 600 refers to the second storage unit 12a and identifies the sound length "0.5" associated with the lip identification information corresponding to the lip identification information "IA" specified based on the second lip image MI2 as the sound length TL2. The second total value processing unit 600 resets the sound length TL2 as the previous sound length (i.e., the previous sound length becomes "TL2 (0.5)"). Also, the second total value processing unit 600 resets the value "1" indicating a match to the previous match flag MF.

[0127] - Specific Examples of Step 17c and Step 18c - With the karaoke performance of music X, the karaoke device K causes the display device 30 to display the lyric characters "Put your head above the clouds" of phrase F1 based on the lyric data of music X. User U makes a sound of "a" according to the displayed lyric characters. The camera 50 outputs the video I1 of user U making the sound of "a" to the karaoke main body 10.

[0128] The acquisition unit 100 extracts the mouth shape image MI1 from the output video I1. Then, the acquisition unit 100 refers to the first storage unit 11a and identifies one mouth shape identification information "A" associated with the same mouth shape image as the extracted mouth shape image MI1. Also, the acquisition unit 100 identifies the extraction time "16014 ms", which is the timing when the mouth shape image MI1 is extracted, as the extraction time.

[0129] The acquisition unit 100 acquires the identified mouth shape identification information "A" and the extraction time "16014 ms" in association with each other.

[0130] The first determination unit 200 determines whether the identified first mouth shape identification information "A" matches the first mouth shape identification information in phrase F1 of music X identified by referring to the second storage unit 12a. As shown in FIG. 4, the first mouth shape identification information in phrase F1 of music X is "A". Therefore, the first determination unit 200 outputs the determination result of "match" to the setting unit 300.

[0131] The setting unit 300 sets the extraction time "16014 ms" associated with the identified first mouth shape identification information "A" as the previous extraction time ET1, refers to the second storage unit 12a, sets the sound length "1.5" associated with the mouth shape identification information corresponding to the first mouth shape identification information "A" as the previous sound length TL1, and sets the value "1" indicating a match as the previous match flag MF.

[0132] Next, the camera 50 outputs the video I2 of user U making a sound of "ka", which is different from the displayed lyric characters, to the karaoke main body 10.

[0133] The acquisition unit 100 extracts the lip image MI2 from the output video I2. Then, the acquisition unit 100 refers to the first storage unit 11a and identifies one lip identification information "A" associated with the same lip image as the extracted lip image MI2. Also, the acquisition unit 100 identifies the extraction time "16910 ms" as the extraction time at the timing when the lip image MI2 is extracted.

[0134] The acquisition unit 100 acquires the identified one lip identification information "A" and the extraction time "16910 ms" in association with each other.

[0135] When the second lip image is extracted in the phrase F1 of the music X, the second determination unit 400 compares one lip identification information "A" (i.e., the second lip identification information) specified based on the second lip image MI2 with the second lip identification information in the phrase F1 of the music X specified by referring to the second storage unit 12a, and determines whether they match or not. As shown in FIG. 4, the second lip identification information in the phrase F1 of the music X is "IA". Therefore, the second determination unit 400 outputs the determination result of "mismatch" to the first total value processing unit 500 and the second total value processing unit 600.

[0136] The first total value processing unit 500 identifies the extraction time "16910 ms" associated with the lip identification information "IA" as the extraction time ET2. The first total value processing unit 500 subtracts the previous extraction time ET1 (16014 ms) from the extraction time ET2 (16910 ms) to obtain the pronunciation time PT1 (896 ms), and adds it to the first total value S1 (0 ms + 896 ms = 896 ms). Also, the first total value processing unit 500 resets the previous extraction time (i.e., the previous extraction time becomes "0").

[0137] Also, the second total value processing unit 600 adds the previous note length TL1 (1.5) to the second total value S2 (0 + 1.5 = 1.5). Also, the second total value processing unit 600 resets the previous note length (that is, the previous note length becomes "0"). Also, the second total value processing unit 600 resets the value "0" indicating a mismatch to the previous match flag MF (that is, the previous match flag becomes "0").

[0138] - Specific examples of step 17d and step 18d - The karaoke device K causes the display device 30 to display the lyric character "lift your head above the clouds" of the phrase F1 based on the lyric data of the music X during the karaoke performance of the music X. The user U pronounces "ta", which is different from the displayed lyric character. The camera 50 outputs the video I1 of the user U pronouncing "ta" to the karaoke main body 10.

[0139] The acquisition unit 100 extracts the mouth shape image MI1 from the output video I1. Then, the acquisition unit 100 refers to the first storage unit 11a and identifies one mouth shape identification information "IA" associated with the same mouth shape image as the extracted mouth shape image MI1. Also, the acquisition unit 100 identifies the extraction time "16014 ms", which is the timing when the mouth shape image MI1 is extracted, as the extraction time.

[0140] The acquisition unit 100 acquires the identified one mouth shape identification information "IA" and the extraction time "16014 ms" in association with each other.

[0141] The first determination unit 200 determines whether or not the identified first mouth shape identification information "IA" matches the first mouth shape identification information in the phrase F1 of the music X identified by referring to the second storage unit 12a. As shown in FIG. 4, the first mouth shape identification information in the phrase F1 of the music X is "A". Therefore, the first determination unit 200 outputs a determination result of "mismatch" to the setting unit 300.

[0142] The setting unit 300 sets the value "0" indicating a mismatch to the previous match flag MF.

[0143] Next, the camera 50 outputs a video I2 that captures the user U making a pronunciation of "ka" different from the displayed lyrics to the karaoke main unit 10.

[0144] The acquisition unit 100 extracts a lip image MI2 from the output video I2. Then, the acquisition unit 100 refers to the first storage unit 11a and identifies one lip identification information "A" associated with the same lip image as the extracted lip image MI2. Also, the acquisition unit 100 identifies a certain time "16910 ms", which is the timing when the lip image MI2 was extracted, as the extraction time.

[0145] The acquisition unit 100 acquires the identified one lip identification information "A" and the extraction time "16910 ms" in association with each other.

[0146] When the second lip image is extracted in the phrase F1 of the music X, the second determination unit 400 compares the one lip identification information "A" (i.e., the second lip identification information) specified based on the second lip image MI2 with the second lip identification information in the phrase F1 of the music X specified by referring to the second storage unit 12a, and makes a determination of match or mismatch. As shown in FIG. 4, the second lip identification information in the phrase F1 of the music X is "IA". Therefore, the second determination unit 400 outputs the determination result of "mismatch" to the first total value processing unit 500 and the second total value processing unit 600.

[0147] The first total value processing unit 500 resets the previous extraction time (i.e., the previous extraction time becomes "0"). Also, the second total value processing unit 600 resets the previous tone length (i.e., the previous tone length becomes "0").

[0148] Each time the user vocalizes, the karaoke device K performs any one of the processes of step 17a and step 18a, step 17b and step 18b, step 17c and step 18c, step 17d and step 18d. In this example, when the color change end time "23300 ms" of the last lyric character "shi" included in the phrase F1 arrives, it is assumed that the first total value S1 is "7188 ms" and the second total value S2 is "12.0".

[0149] The calculation unit 700 divides the first total value S1 (7188 ms) by the second total value S2 (12.0) to calculate the length " " of one beat in the user's karaoke singing in the phrase F1. The calculation unit 700 calculates the singing tempo "100 bpm" of the user U in the phrase F1 from the calculated length of one beat and the performance tempo "120 bpm" of the music X.

[0150] The performance processing unit 800 obtains the difference between the calculated singing tempo "100 bpm" and the performance tempo "120 bpm" of the music X. The performance processing unit 800 adjusts the performance tempo of the music X in the phrase F2 following the phrase F1 (for example, -20 bpm) according to the obtained difference "20".

[0151] As is clear from the above, the karaoke device K according to the present embodiment stores in association a plurality of mouth shape images showing the mouth shape when pronouncing a vowel and the mouth shape in a state where the mouth is closed, and mouth shape identification information indicating the mouth shape corresponding to the mouth shape image in a first storage unit 11a. A second storage unit 12a stores, for each piece of music, lyric mouth shape data in which a pronunciation time indicating the timing at which each lyric character included in the lyric data should be pronounced, a predetermined section including the pronunciation time, mouth shape identification information indicating the mouth shape to be pronounced at the pronunciation time, and a sound length indicating the ratio of the length of each lyric character pronounced to the length of one beat based on the performance tempo of the music are associated with each other. An acquisition unit 100 acquires, in association with each other, mouth shape identification information specified based on a mouth shape image extracted from a video obtained by photographing a user performing karaoke singing of a certain piece of music, and an extraction time which is the timing at which the mouth shape image is extracted. When the first mouth shape image is extracted in a certain predetermined section of a certain piece of music, a first determination unit 200 determines whether or not the first mouth shape identification information specified based on the first mouth shape image matches the first mouth shape identification information in the certain predetermined section specified by referring to the second storage unit 12a. When it is determined by the first determination unit 200 that they match, the first extraction time associated with the first mouth shape identification information is set as the previous extraction time, the second storage unit 12a is referred to, the sound length associated with the mouth shape identification information corresponding to the first mouth shape identification information is set as the previous sound length, and a value indicating a match is set as the previous match flag. When it is determined by the first determination unit 200 that they do not match, a value indicating a mismatch is set as the previous match flag. When the (1 + n)th mouth shape image is extracted in a certain predetermined section of a certain piece of music, a second determination unit 400 determines whether or not the (1 + n)th mouth shape identification information specified based on the (1 + n)th mouth shape image matches the (1 + n)th mouth shape identification information in the certain predetermined section specified by referring to the second storage unit 12a. When it is determined by the second determination unit 400 that they match and a value indicating a match is set in the previous match flag, a first process is executed in which the pronunciation time obtained by subtracting the previous extraction time from the (1 + n)th extraction time associated with the (1 + n)th mouth shape identification information specified based on the (1 + n)th mouth shape image is added to a first total value, and the (1 + n)th extraction time is reset as the previous extraction time.When it is determined by the second determination unit 400 that there is a match and a value indicating no match is set in the previous match flag, a second process is executed to reset the (1 + n)-th extraction time associated with the (1 + n)-th lip shape identification information identified based on the (1 + n)-th lip shape image as the previous extraction time. When it is determined by the second determination unit 400 that there is no match and a value indicating a match is set in the previous match flag, the pronunciation time obtained by subtracting the previous extraction time from the (1 + n)-th extraction time associated with the (1 + n)-th lip shape identification information identified based on the (1 + n)-th lip shape image is added to the first total value, and a third process is executed to reset the previous extraction time. When it is determined by the second determination unit 400 that there is no match and a value indicating no match is set in the previous match flag, a fourth process is executed to reset the previous extraction time. A first total value processing unit 500 that executes the fourth process, when it is determined by the second determination unit 400 that there is a match and a value indicating a match is set in the previous match flag, adds the previous sound length to the second total value, refers to the second storage unit 12a, sets the sound length associated with the lip shape identification information corresponding to the (1 + n)-th lip shape identification information identified based on the (1 + n)-th lip shape image as the (1 + n)-th sound length, and executes a fifth process to reset the (1 + n)-th sound length as the previous sound length. When it is determined by the second determination unit 400 that there is a match and a value indicating no match is set in the previous match flag, refers to the second storage unit 12a, sets the sound length associated with the lip shape identification information corresponding to the (1 + n)-th lip shape identification information identified based on the (1 + n)-th lip shape image as the (1 + n)-th sound length, resets the (1 + n)-th sound length as the previous sound length, and executes a sixth process to reset the value indicating a match to the previous match flag. When it is determined by the second determination unit 400 that there is no match and a value indicating a match is set in the previous match flag, adds the previous sound length to the second total value, resets the previous sound length, and executes a seventh process to reset the value indicating no match to the previous match flag. When it is determined by the second determination unit 400 that there is no match and a value indicating no match is set in the previous match flag, a second total value processing unit 600 that executes an eighth process to reset the previous sound length, when the end time of the color change of the last lyric character included in a certain predetermined section arrives,Based on the first total value and the second total value obtained so far, a calculation unit 700 that calculates the singing tempo of a user in a predetermined section, and a performance processing unit 800 that adjusts the performance tempo of the next predetermined section according to the calculated singing tempo.

[0152] According to such a karaoke device K, by using the lip image extracted from the video obtained by photographing the user who performs karaoke singing of a certain song, the singing tempo of a certain predetermined section can be calculated, and the performance tempo of the next predetermined section can be adjusted. Therefore, even when the user performs karaoke singing without using a microphone, the performance tempo of the song can be adjusted according to the singing tempo of the user. Furthermore, according to the karaoke device K according to the present embodiment, even when the user makes a voice different from the displayed lyrics, the performance tempo of the song can be adjusted. That is, according to the karaoke device K according to the present embodiment, the performance tempo of the song can be automatically adjusted without using the singing voice of the user.

[0153] <Modification Example 1> In the above embodiment, an example of a single user has been described. On the other hand, in a nursing facility or the like, a single song may be karaoke-sung by a plurality of people. In such a case, the karaoke device K can adjust the performance tempo of the song in consideration of the singing tempo of each user.

[0154] (Calculation Unit) The calculation unit 700 according to this modification example calculates the singing tempo for each user when the karaoke device K is used by a plurality of users. The calculation of the singing tempo can be performed in the same manner as in the embodiment. The identification of the user can be performed based on, for example, a face image identified using a known face recognition technique from the photographed video.

[0155] (Performance Processing Unit) The performance processing unit 800 according to this modification example adjusts the performance tempo of the next predetermined section according to the calculated singing tempo of a plurality of users.

[0156] For example, as a result of karaoke singing of phrase F1 of music X in the embodiment by three users U1 to U3, it is assumed that the singing tempo of user U1 is calculated as "120 bpm", the singing tempo of user U2 is calculated as "110 bpm", and the singing tempo of user U3 is calculated as "100 bpm".

[0157] In this case, the performance processing unit 800 can adjust the performance tempo of phrase F2 according to "110 bpm", which is the average value of the singing tempos of the three users. Also, the performance processing unit 800 can adjust the performance tempo of phrase F2 according to "100 bpm", which is the slowest singing tempo among the singing tempos of the three users. Alternatively, the performance processing unit 800 can adjust the performance tempo of phrase F2 according to "105 bpm", which is the average value of the singing tempos of users U2 and U3 whose singing tempos are slower than the performance tempo "120 bpm" of music X among the singing tempos of the three users.

[0158] As is clear from the above, when the karaoke device K according to this modification example is used by a plurality of users, the calculation unit 700 calculates the singing tempo for each user, and the performance processing unit 800 can adjust the performance tempo of the next predetermined section in a predetermined section according to the calculated singing tempos of the plurality of users. According to such a karaoke device K, even when a single piece of music is karaoke-sung by a plurality of people, the performance tempo of the music can be automatically adjusted without using the singing voices of each user.

[0159] <Modification Example 2> A user whose singing tempo is faster than the performance tempo of the music may start karaoke singing of the next predetermined section during a certain predetermined section. In this case, even if the lip identification information is specified in a certain predetermined section (that is, even if the lip image is extracted), the corresponding lip identification information (lip identification information in a certain predetermined section) is not included in the second storage unit 12a. In this modification example, even in such a case, the performance tempo of the music can be adjusted.

[0160] [Control means] In this modification example, when the CPU executes a program stored in the memory, the control means 10e functions as an acquisition unit 100, a first determination unit 200, a setting unit 300, a second determination unit 400, a first total value processing unit 500, a second total value processing unit 600, a calculation unit 700, a performance processing unit 800, and a specifying unit 900 (see FIG. 8).

[0161] (Specifying unit) The specifying unit 900 refers to the second storage unit 12a and specifies the number of lip shape identification information corresponding to the lyric characters included in a predetermined section, which is displayed on the display means along with the karaoke performance of a certain song.

[0162] The karaoke device K causes the display device 30 to display lyric characters for each predetermined section based on the lyric data of the selected song along with the karaoke performance of the song.

[0163] The specifying unit 900 refers to the lyric lip shape data of a certain song stored in the second storage unit 12a and specifies the number of lip shape identification information corresponding to the lyric characters included in a predetermined section, which is displayed on the display device 30.

[0164] For example, in the example of the embodiment, the karaoke device K causes the display device 30 to display the lyric characters "lift your head above the clouds" of the phrase F1 based on the lyric data of the song X along with the karaoke performance of the song X.

[0165] The specifying unit 900 refers to the lyric lip shape data of the song X stored in the second storage unit 12a and specifies the number "12" of the lip shape identification information corresponding to the lyric characters included in the phrase F1, which is displayed on the display device 30.

[0166] (First total value processing unit) The first total value processing unit 500 according to this modification example executes any one of the first to fourth processes when the (1 + n)th lip shape image is extracted in a predetermined section and the value of 1 + n is less than or equal to the number of specified lip shape identification information.

[0167] (Second total value processing unit) When the 1 + n-th mouth shape image is extracted in a certain predetermined section and the value of 1 + n is less than or equal to the number of identified mouth shape identification information, the second total value processing unit 600 according to this modification example executes any one of the fifth to eighth processes.

[0168] For example, it is assumed that the number "12" of the mouth shape identification information corresponding to the lyric characters included in the phrase F1 is identified as described above.

[0169] In this case, according to the example of the embodiment, the second mouth shape image MI2 is extracted in the phrase F1 of the music X, and the value of "2" is less than or equal to the number "12" of the identified mouth shape identification information. Therefore, the first total value processing unit 500 and the second total value processing unit 600 execute any of the processes described in the embodiment (the first process and the fifth process, the second process and the sixth process, the third process and the seventh process, or the fourth process and the eighth process).

[0170] On the other hand, in the example of the embodiment, it is assumed that the singing tempo of the user U is fast and the user U has pronounced the first lyric character "shi" of phrase 2 in phrase F1. Even in this case, the camera 50 outputs the video I13 that has captured the user U who pronounced "shi" to the karaoke main body 10.

[0171] The acquisition unit 100 extracts the mouth shape image MI13 from the output video I13. Then, the acquisition unit 100 refers to the first storage unit 11a and identifies one mouth shape identification information "IA" associated with the same mouth shape image as the extracted mouth shape image MI13. In addition, the acquisition unit 100 identifies the extraction time as a certain time "22000 ms" which is the timing when the mouth shape image MI13 is extracted.

[0172] The acquisition unit 100 acquires the identified one mouth shape identification information "IA" and the extraction time "22000 ms" in association with each other.

[0173] Here, the 13th mouth shape image MI13 is extracted from the phrase F1 of the music X. However, the value of "13" is equal to or greater than the number "12" of the specified mouth shape identification information. In this case, the first total value processing unit 500 and the second total value processing unit 600 do not execute the processes described in the embodiment.

[0174] As is clear from the above, the karaoke device K according to this modification example refers to the second storage unit 12a and has an identification unit 900 that identifies the number of mouth shape identification information corresponding to the lyric characters included in a predetermined section, which is displayed on the display means along with the karaoke performance of a certain music. The first total value processing unit 500 according to this modification executes any one of the first to fourth processes when the (1 + n)th mouth shape image is extracted in a predetermined section and the value of (1 + n) is equal to or less than the number of the specified mouth shape identification information. Further, the second total value processing unit 600 according to this modification example executes any one of the fifth to eighth processes when the (1 + n)th mouth shape image is extracted in a predetermined section and the value of (1 + n) is equal to or less than the number of the specified mouth shape identification information. According to such a karaoke device K, even when the singing tempo of the user is too fast, the performance tempo of the music can be automatically adjusted according to the singing tempo of the user.

[0175] <Modification Example 3> In the example of the embodiment, an example in which the lyric mouth shape data is set in advance for each music has been described. On the other hand, the karaoke device K may generate lyric mouth shape data each time a certain music is selected and store it in the second storage unit 12a.

[0176] Specifically, the karaoke device K can generate lyric mouth shape data for each music by associating the pronunciation time indicating the timing at which each character included in the lyric data should be pronounced, a predetermined section including the pronunciation time, the mouth shape identification information indicating the mouth shape when pronouncing at the timing, and the sound length indicating the ratio of the length of each character pronounced to the length of one beat based on the performance tempo of the music, and store it in the second storage unit 12a.

[0177] The predetermined section is associated with each lyric character. The pronunciation time can be the start time of the color change (or the time obtained by adding a predetermined time (for example, 200 ms) to the start time of the color change) associated with each lyric character. The mouth shape identification information can be specified for each lyric character by referring to a table in which preset mouth shape identification information is associated with characters (see, for example, FIG. 3 of Japanese Patent Application Laid-Open No. 2008-310382). The note length can be specified based on the performance tempo of the music, the start time of the color change associated with each lyric character, and the end time of the color change. For example, when the performance tempo is "120 bpm", the length of one beat is "500 ms". Here, if the time from the start of the color change to the end of the color change of a certain lyric character is "750 ms", the note length is "1.5".

[0178] <Others> It is also possible to supply a program to a computer using a non-transitory computer-readable medium storing the above program (non-transitory computer readable medium with an executable program thereon). Examples of non-transitory computer-readable media include magnetic recording media (such as flexible disks, magnetic tapes, hard disk drives), CD-ROM (Read Only Memory), and the like.

[0179] The above embodiments are presented as examples and do not limit the scope of the invention. The above configurations can be implemented in appropriate combinations, and various omissions, replacements, and changes can be made without departing from the gist of the invention. The above embodiments and their modifications are included in the scope and gist of the invention, and are also included in the invention described in the claims and its equivalent scope.

Explanation of Reference Numerals

[0180] K Karaoke device 10 Karaoke main body 11a First storage unit 12a Second storage unit 100 Acquisition Unit 200 First Determination Unit 300 Setting Unit 400 Second Determination Unit 500 First Total Value Processing Unit 600 Second Total Value Processing Unit 700 Calculation Unit 800 Performance Processing Unit 900 Identification Unit

Claims

1. A first storage unit that stores, in association with each other, a plurality of mouth shape images showing the mouth shape when pronouncing a vowel and the mouth shape in a closed-mouth state, and mouth shape identification information indicating the mouth shape corresponding to the mouth shape image; A second storage unit that stores, for each piece of music, lyric mouth shape data in which a pronunciation time indicating the timing at which each lyric character included in the lyric data should be pronounced, a predetermined section including the pronunciation time, mouth shape identification information indicating the mouth shape to be pronounced at the pronunciation time, and a sound length indicating the ratio of the length of pronunciation of each lyric character to the length of one beat based on the performance tempo of the music are associated with each other; An acquisition unit that acquires, in association with each other, mouth shape identification information specified based on a mouth shape image extracted from a video obtained by photographing a user performing karaoke singing of a certain piece of music, and an extraction time that is the timing at which the mouth shape image was extracted; When the first mouth shape image is extracted in a certain predetermined section of the certain piece of music, a first determination unit that determines whether or not the first mouth shape identification information specified based on the first mouth shape image matches the first mouth shape identification information in the certain predetermined section specified by referring to the second storage unit; When it is determined by the first determination unit that they match, a setting unit that sets the first extraction time associated with the first mouth shape identification information as the previous extraction time, refers to the second storage unit, sets the sound length associated with the mouth shape identification information corresponding to the first mouth shape identification information as the previous sound length, and sets a value indicating a match as the previous match flag, and when it is determined by the first determination unit that they do not match, sets a value indicating a mismatch as the previous match flag; When the (1 + n)th mouth shape image is extracted in a certain predetermined section of the certain piece of music, a second determination unit that determines whether or not the (1 + n)th mouth shape identification information specified based on the (1 + n)th mouth shape image matches the (1 + n)th mouth shape identification information in the certain predetermined section specified by referring to the second storage unit; When it is determined by the second determination unit that there is a match and a value indicating a match is set in the previous match flag, the pronunciation time obtained by subtracting the previous extraction time from the (1 + n)-th extraction time associated with the (1 + n)-th mouth shape identification information specified based on the (1 + n)-th mouth shape image is added to the first total value, and a first process of resetting the (1 + n)-th extraction time as the previous extraction time is executed. When it is determined by the second determination unit that there is a match and a value indicating a non-match is set in the previous match flag, a second process of resetting the (1 + n)-th extraction time as the previous extraction time is executed. When it is determined by the second determination unit that there is no match and a value indicating a match is set in the previous match flag, the pronunciation time obtained by subtracting the previous extraction time from the (1 + n)-th extraction time associated with the (1 + n)-th mouth shape identification information specified based on the (1 + n)-th mouth shape image is added to the first total value, and a third process of resetting the previous extraction time is executed. When it is determined by the second determination unit that there is no match and a value indicating a non-match is set in the previous match flag, a first total value processing unit that executes a fourth process of resetting the previous extraction time When it is determined by the second determination unit that there is a match and a value indicating a match is set in the previous match flag, add the previous sound length to the second total value, and refer to the second storage unit. The sound length associated with the mouth shape identification information corresponding to the (1 + n)-th mouth shape identification information specified based on the (1 + n)-th mouth shape image is set as the (1 + n)-th sound length, and execute a fifth process of resetting the (1 + n)-th sound length as the previous sound length. When it is determined by the second determination unit that there is a match and a value indicating a mismatch is set in the previous match flag, refer to the second storage unit. The sound length associated with the mouth shape identification information corresponding to the (1 + n)-th mouth shape identification information specified based on the (1 + n)-th mouth shape image is set as the (1 + n)-th sound length, and execute a sixth process of resetting the (1 + n)-th sound length as the previous sound length and resetting a value indicating a match to the previous match flag. When it is determined by the second determination unit that there is no match and a value indicating a match is set in the previous match flag, add the previous sound length to the second total value, reset the previous sound length, and execute a seventh process of resetting a value indicating a mismatch to the previous match flag. When it is determined by the second determination unit that there is no match and a value indicating a mismatch is set in the previous match flag, execute an eighth process of resetting the previous sound length, a second total value processing unit; When the end time of the color change of the last lyric character included in a certain predetermined section arrives, based on the first total value and the second total value obtained so far, calculate the singing tempo of the user in the certain predetermined section, a calculation unit; A performance processing unit that adjusts the performance tempo of the next predetermined section of the certain predetermined section according to the calculated singing tempo; A karaoke device having the above.

2. When the karaoke device is used by a plurality of users, the calculation unit calculates the singing tempo for each user. The performance processing unit adjusts the performance tempo of the next predetermined section of the certain predetermined section according to the calculated singing tempos of the plurality of users. The karaoke device according to claim 1, characterized in that.

3. It has a specifying unit that refers to the second storage unit and specifies the number of mouth shape identification information corresponding to the lyric characters included in the certain predetermined section, which is displayed on the display means along with the karaoke performance of the certain music. When the 1 + n-th mouth shape image is extracted in the certain predetermined section and the value of 1 + n is less than or equal to the number of mouth shape identification information for which the value has been specified, the first total value processing unit executes any one of the first to fourth processes. The karaoke device according to claim 1 or 2, wherein when the 1 + n-th mouth shape image is extracted in the certain predetermined section and the value of 1 + n is less than or equal to the number of mouth shape identification information for which the value has been specified, the second total value processing unit executes any one of the fifth to eighth processes.

Citation Information

Patent Citations

  • Tempo controller for karaoke

    JP1998149180A