Karaoke device
The karaoke device uses mouth shape recognition to automatically adjust music tempo, addressing the challenge of microphone usage difficulties by synchronizing the music tempo with the user's singing pace.
Patent Information
- Application Number
- JP2023219754
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-07-08
AI Technical Summary
Existing karaoke devices require users to sing into a microphone to adjust the music tempo, which can be challenging for individuals with difficulty using microphones.
A karaoke device that uses mouth shape recognition to automatically adjust the music tempo based on mouth shape images and associated identification information, without requiring the user's singing voice.
Enables automatic tempo adjustment of music pieces based on the user's singing tempo, allowing individuals to perform karaoke without using a microphone.
Smart Images

Figure 2025102358000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a karaoke device.
Background Art
[0002] When performing karaoke singing using a karaoke device, there may be a deviation between the performance tempo of the music and the singing speed of the singer.
[0003] Therefore, there is a technique for causing the performance tempo of the music in the karaoke device to follow the singing speed of the singer. For example, Patent Document 1 discloses a technique of tracking the pitch of the singing voice captured from a microphone for a predetermined period to recognize the pattern of pitch changes in the singing voice, determining the slowness or rapidity of singing based on the difference between the recognized pattern of pitch changes and the pattern of pitch changes indicated by the guide melody, and controlling the tempo of the karaoke according to the determination result.
[0004] In addition, a karaoke device installed in a nursing facility, a welfare facility, etc. (hereinafter referred to as "facility") and capable of providing various contents using music is known. When performing karaoke singing using the karaoke device described in Non-Patent Document 1, the staff of the facility checks whether the user can follow the karaoke performance and manually adjusts the performance tempo of the music.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Non-Patent Documents
[0006]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0007] When adjusting the performance tempo of a music piece using the technology described in Patent Document 1, it is necessary to capture the singing voice through a microphone held by the user. On the other hand, among the users in the facility, there are some who have difficulty performing karaoke singing using a microphone.
[0008] An object of the present invention is to provide a karaoke device capable of automatically adjusting the performance tempo of a music piece without using the singing voice of the user.
Means for Solving the Problems
[0009] The invention for achieving the above object includes a first storage unit that stores in association a plurality of mouth shape images indicating the mouth shape when pronouncing vowels and the mouth shape in a closed-mouth state, and mouth shape identification information indicating the mouth shape corresponding to the mouth shape image; a pronunciation time indicating the timing at which each lyric character included in the lyric data should be pronounced, a predetermined section including the pronunciation time, mouth shape identification information indicating the mouth shape to be pronounced at the pronunciation time, and a syllable length indicating the ratio of the length of pronouncing each lyric character to the length of one beat based on the performance tempo of the music. A second storage unit stores the lyric mouth shape data associated with each other for each music; an acquisition unit that acquires, in association, the mouth shape identification information specified based on the mouth shape image extracted from the video obtained by photographing a user performing karaoke singing of a certain music, and the extraction time which is the timing when the mouth shape image is extracted; when the first mouth shape image is extracted in a certain predetermined section of the certain music, the first extraction time associated with the first mouth shape identification information specified based on the first mouth shape image is set as the previous extraction time, and the second storage unit is referred to, and the syllable length associated with the mouth shape identification information corresponding to the first mouth shape identification information is set as the previous syllable length. A setting unit; when the (1 + n)th mouth shape image is extracted in the certain predetermined section, the pronunciation time obtained by subtracting the previous extraction time from the (1 + n)th extraction time associated with the (1 + n)th mouth shape identification information specified based on the (1 + n)th mouth shape image is added to the first total value, and the (1 + n)th extraction time is reset as the previous extraction time. A first total value processing unit; when the (1 + n)th mouth shape image is extracted in the certain predetermined section, the previous syllable length is added to the second total value, and the second storage unit is referred to, and the syllable length associated with the mouth shape identification information corresponding to the (1 + n)th mouth shape identification information specified based on the (1 + n)th mouth shape image is set as the (1 + n)th syllable length, and the (1 + n)th syllable length is reset as the previous syllable length. A second total value processing unit; when the end time of the color change of the last lyric character included in the certain predetermined section arrives, a calculation unit that calculates the singing tempo of the user in the certain predetermined section based on the first total value and the second total value obtained so far. A karaoke device having a performance processing unit that adjusts the performance tempo of the next predetermined section of the predetermined section according to the calculated singing tempo. Other features of the present invention will be clarified by the description of the specification and drawings described later.
Effects of the Invention
[0010] According to the present invention, the performance tempo of a music piece can be automatically adjusted without using the singing voice of the user.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Modes for Carrying Out the Invention
[0012] <Embodiment> With reference to FIGS. 1 to 6, a karaoke device K according to an embodiment will be described.
[0013] ==Karaoke Device== The karaoke device K is a device for karaoke performance of music pieces and for the user to perform karaoke singing. As shown in FIG. 1, the karaoke device K includes a karaoke main body 10, a speaker 20, a display device 30, a remote control device 40, and a camera 50.
[0014] The karaoke main body 10 performs various controls related to karaoke performance and karaoke singing, such as performance control of the selected music and display control of lyrics, background images, etc. The speaker 20 is configured to emit sound based on the sound emission signal from the karaoke main body 10. The display device 30 is configured to display videos and images on the screen based on the signal from the karaoke main body 10. The display device 30 may be a large display separate from the karaoke device K. The remote control device 40 is a device for performing various operations on the karaoke main body 10. The camera 50 is configured to photograph the user. The camera 50 starts photographing, for example, at the timing when the use of the karaoke device K is started or at the timing when the karaoke performance is started. The camera 50 is an example of the "photographing means". The photographing means does not necessarily have to be provided in the karaoke device K itself, such as a camera for indoor photography provided in the karaoke room where the karaoke device K is installed. Although not used in the invention according to this embodiment, the karaoke device K may be equipped with a microphone.
[0015] As shown in FIG. 2, the karaoke main body 10 according to this embodiment includes a storage means 10a, a communication means 10b, an input means 10c, a performance means 10d, and a control means 10e. Each component is connected to the bus B via an interface (not shown).
[0016] [Storage means] The storage means 10a is a large-capacity storage device that stores various data. The storage means 10a stores music data. The music data is provided with music identification information. The music identification information is information unique to each music, such as a music ID for identifying the music. The music data includes accompaniment data, reference data, section information, etc.
[0017] Accompaniment data is the data that serves as the basis for karaoke performance sounds. The accompaniment data includes tempo information indicating the tempo for karaoke performance. The tempo is set to a predetermined value (for example, 100 BPM, 120 BPM) for each song. Reference data is the data used as a reference when scoring the user's karaoke singing. Interval information indicates the performance interval. The performance interval is the interval during which karaoke performance is conducted. The performance interval includes a singing interval and a non-singing interval. The singing interval is the interval (for example, the A melody, B melody, and refrain of the first verse) in which the lyrics to be sung in the song are set. The non-singing interval is the interval in which no lyrics to be sung in the song are set, such as the introduction, interlude, and coda.
[0018] Also, the storage means 10a stores lyric data, background image data such as background images to be displayed on the display device 30 etc. during karaoke performance, performance time data indicating the karaoke performance time for each song, and attribute information of the song (information related to the song such as the singer's name, lyricist's name, composer's name, genre, etc.).
[0019] Lyric data is the data for causing the display device 30 etc. to display the lyric characters of each song. The lyric data includes each displayed lyric character, the character type indicating whether the lyric character is the main lyric or a ruby character, the start time and end time of the color change of each lyric character. Also, each lyric character is associated with a predetermined interval. The predetermined interval is a unit such as a phrase separated by a character string where consecutive syllables to be pronounced are interrupted, or a line separated by a character string for one horizontal row of displaying lyrics.
[0020] The start time and end time of the color change are the elapsed time when the start time of the karaoke performance is set to 0. Also, generally, there is a time lag from when the user sees the displayed lyrics until starting to sing. Therefore, the start time and end time of the color change may be set to a time slightly (for example, 200 ms) ahead of the karaoke performance.
[0021] 3 shows a portion of the lyrics data for song X. For example, the first lyric character "a" of song X is associated with the character type "lyrics", the color change start time "15800 ms", the color change end time "16350 ms", and phrase F1, which is a predetermined section.
[0022] The lyrics are displayed based on a predetermined section. The lyrics included in one predetermined section are displayed on the same screen. Generally, one phrase includes multiple lines. In this case, each line is displayed at a different position on the same screen. For example, in the case of music X, phrase F1 includes two lines, "Atama wo Kumo no" (head on the clouds) and "Ue dashi" (put on above). Therefore, on the same screen, "Atama wo Kumo no" is displayed on the upper line, and "Ue dashi" is displayed on the lower line. When the lyrics included in phrase F1 of music X are displayed, the user attempts to sing karaoke in time with the tempo of music X, from the first lyric character "a" to the last lyric character "shi."
[0023] In this embodiment, a part of the storage area of the storage means 10a functions as a first storage unit 11a and a second storage unit 12a.
[0024] (First memory unit) The first storage unit 11a stores a plurality of mouth shape images and mouth shape identification information in association with each other.
[0025] The mouth shape images are images showing the mouth shapes when pronouncing a vowel and when the mouth is closed. Specifically, the mouth shape images are images corresponding to the mouth shapes when pronouncing the five vowels "a", "i", "u", "e", and "o" and the mouth shape when the mouth is closed. In other words, there are six types of mouth shape images (for example, see FIG. 5 in JP2008-310382A).
[0026] The mouth shape identification information is information that indicates the mouth shape corresponding to the mouth shape image. When the mouth shape images are "a", "i", "u", "e", "o", and "n", the mouth shape identification information is "A", "I", "U", "E", "O", and "N", respectively.
[0027] (Second storage unit) The second storage unit 12a stores the lyric mouth shape data for each piece of music. The lyric mouth shape data is data that associates a pronunciation time, a predetermined section including the pronunciation time, mouth shape identification information indicating the mouth shape when pronouncing at that timing, and a sound length.
[0028] The pronunciation time is the time indicating the timing at which each lyric character included in the lyric data should be pronounced. The pronunciation time can be indicated by the elapsed time when the start time of the karaoke performance is set to 0. The predetermined section is, like the predetermined section included in the lyric data, a phrase or the like. The mouth shape identification information is the same as the mouth shape identification information stored in the first storage unit 11a, such as "A", "I", "U", "E", "O", "N", etc. The sound length indicates the ratio of the length of pronouncing each lyric character to the length of one beat based on the performance tempo of the music.
[0029] FIG. 4 shows a part of the lyric mouth shape data of music X. Assume that the performance tempo of music X is 120 bpm (that is, the length of one beat is 500 ms). In this example, the predetermined section corresponds to a phrase. The lyric mouth shape data included in one phrase is numbered in the order of pronunciation along with karaoke singing.
[0030] For example, the sound length associated with the first mouth shape identification information "A" (pronunciation time 16000 ms) of phrase F1 is "1.5". Therefore, the pronunciation time of the mouth shape identification information "A" (that is, the time to be pronounced in the mouth shape corresponding to the mouth shape identification information "A") is 750 ms. Also, the sound length associated with the second mouth shape identification information "IA" (pronunciation time 16550 ms) of phrase F1 is "0.5". Therefore, the pronunciation time of the mouth shape identification information "IA" (that is, the time to be pronounced in the mouth shape corresponding to the mouth shape identification information "IA") is 250 ms.
[0031] [Communication means · Input means] The communication means 10b provides an interface for communicating with the remote control device 40. The input means 10c is configured for a facility staff or the like to perform an instruction input. The input means 10c is a button or the like provided on the karaoke main body 10. Alternatively, the remote control device 40 may function as the input means 10c.
[0032] [Performance means] The performance means 10d performs karaoke performance of a music piece based on the control of the control means 10e. The performance means 10d includes a sound source, a mixer, an amplifier, etc. (all not shown in the figure).
[0033] [Control means] The control means 10e performs various controls in the karaoke device K. The control means 10e includes a CPU and a memory (both not shown in the figure). The CPU realizes various functions by executing a program stored in the memory.
[0034] In the present embodiment, by the CPU executing the program stored in the memory, the control means 10e functions as an acquisition unit 100, a setting unit 200, a first total value processing unit 300, a second total value processing unit 400, a calculation unit 500, and a performance processing unit 600 (see FIG. 2).
[0035] (Acquisition unit) The acquisition unit 100 acquires, in association with each other, the lip shape identification information specified based on a lip shape image extracted from a video obtained by photographing a user performing karaoke singing of a certain music piece, and the extraction time which is the timing when the lip shape image is extracted.
[0036] The acquisition unit 100 extracts a lip shape image from a video obtained by photographing a user performing karaoke singing of a music piece. The lip shape image is an image capable of discriminating the user's lip shape. A technique known in the art can be used for the process of extracting the lip shape image from the user's video. The extraction time is the time when the lip shape image is extracted. Note that the extraction time can be obtained by the acquisition unit 100 referring to the elapsed time with the start time of the karaoke performance timed by the performance means 10d being 0 when extracting the lip shape image from the photographed video.
[0037] For example, assume that a user performs karaoke singing of a piece of music selected by the staff of the facility. The karaoke device K starts karaoke performance of the selected piece of music. When the karaoke performance of the music starts, the camera 50 starts photographing the user. The camera 50 sequentially outputs the photographed video to the karaoke main body 10.
[0038] The karaoke device K causes the display device 30 to display the lyric characters for each predetermined section based on the lyric data of the music along with the karaoke performance of the selected music. The user vocalizes one character at a time in accordance with the displayed lyric characters (in this embodiment, it is assumed that the user vocalizes for all the displayed lyric characters).
[0039] The acquisition unit 100 extracts a lip image from the video obtained by photographing the user performing karaoke singing of the music. The acquisition unit 100 refers to the first storage unit 11a and specifies one piece of lip identification information associated with the same lip image as the extracted lip image. Further, the acquisition unit 100 specifies the time at which the lip image is extracted as the extraction time.
[0040] The acquisition unit 100 acquires the specified one piece of lip identification information and the extraction time in association with each other. The acquisition unit 100 repeatedly performs the same process each time a lip image is extracted.
[0041] (Setting unit) When the first lip image is extracted in a certain predetermined section of a certain piece of music, the setting unit 200 sets the first extraction time associated with the first lip identification information specified based on the first lip image as the previous extraction time, and refers to the second storage unit 12a to set the sound length associated with the lip identification information corresponding to the first lip identification information as the previous sound length.
[0042] The first lip identification information is the lip identification information specified by the acquisition unit 100 based on the first lip image. The first extraction time is the time associated with the first lip identification information.
[0043] When the first (i.e., the 1st) mouth shape image is extracted in a certain predetermined section (for example, the first phrase) of a certain piece of music, the setting unit 200 specifies the extraction time associated with a certain mouth shape identification information (i.e., the first mouth shape identification information) specified based on the first mouth shape image as the first extraction time. The setting unit 200 sets the first extraction time as the previous extraction time.
[0044] Also, when the first (i.e., the 1st) mouth shape image is extracted in a certain predetermined section (for example, the first phrase) of a certain piece of music, the setting unit 200 refers to the second storage unit 12a and specifies the sound length associated with the mouth shape identification information corresponding to the certain mouth shape identification information (i.e., the first mouth shape identification information) that has been specified. The setting unit 200 sets the specified sound length as the previous sound length.
[0045] Note that at the start point of the karaoke performance of a certain piece of music, the previous extraction time and the previous sound length are "0". Also, the previous extraction time and the previous sound length are reset for each predetermined section (i.e., reset to "0" for each predetermined section).
[0046] (First total value processing unit) When the (1 + n)th mouth shape image is extracted in a certain predetermined section, the first total value processing unit 300 adds the pronunciation time obtained by subtracting the previous extraction time from the (1 + n)th extraction time associated with the (1 + n)th mouth shape identification information specified based on the (1 + n)th mouth shape image to the first total value, and resets the (1 + n)th extraction time as the previous extraction time.
[0047] "n" is a natural number. The (1 + n)-th mouth shape identification information is the mouth shape identification information specified by the acquisition unit 100 based on the (1 + n)-th mouth shape image. The (1 + n)-th extraction time is the time associated with the (1 + n)-th mouth shape identification information. The first total value is the value obtained by summing the pronunciation times (the times when the user uttered sounds) in a certain predetermined interval. At the start point of the karaoke performance of a certain piece of music, the first total value is "0". Also, the first total value is reset for each predetermined interval (that is, reset to "0" for each predetermined interval).
[0048] For example, as described above, assume that in a certain predetermined interval of a certain piece of music, the first mouth shape image is extracted, and the previous extraction time ET1 and the previous sound length TL1 are set.
[0049] After that, when the second (n = 1) mouth shape image is extracted in a certain predetermined interval, the first total value processing unit 300 specifies the extraction time associated with the second mouth shape identification information (the mouth shape identification information specified by the acquisition unit 100 based on the second mouth shape image) as the second extraction time ET2. The first total value processing unit 300 subtracts the previous extraction time ET1 from the second extraction time ET2 to obtain the pronunciation time PT1, and adds it (0 + PT1) to the first total value S1. Also, the first total value processing unit 300 resets the second extraction time ET2 as the previous extraction time (that is, the previous extraction time becomes "ET2"). The first total value processing unit 300 repeats the above process each time the (1 + n)-th mouth shape image is extracted in a certain predetermined interval.
[0050] (Second total value processing unit) When the (1 + n)-th mouth shape image is extracted in a certain predetermined interval, the second total value processing unit 400 adds the previous sound length to the second total value, and refers to the second storage unit 12a, and sets the sound length associated with the mouth shape identification information corresponding to the (1 + n)-th mouth shape identification information specified based on the (1 + n)-th mouth shape image as the (1 + n)-th sound length, and resets the (1 + n)-th sound length as the previous sound length.
[0051] "n" is a natural number. The (1 + n)-th mouth shape identification information is the mouth shape identification information specified by the acquisition unit 100 based on the (1 + n)-th mouth shape image as described above. The (1 + n)-th note length is the note length associated with the mouth shape identification information corresponding to the (1 + n)-th mouth shape identification information. The second total value is the value obtained by summing the note lengths in a certain predetermined interval. At the start time of the karaoke performance of a certain piece of music, the second total value is "0". Also, the second total value is reset for each predetermined interval (that is, reset to "0" for each predetermined interval).
[0052] For example, as described above, assume that in a certain predetermined interval of a certain piece of music, the first mouth shape image is extracted and the previous extraction time ET1 and the previous note length TL1 are set.
[0053] After that, when the second (n = 1) mouth shape image is extracted in a certain predetermined interval, the second total value processing unit 400 adds the previous note length TL1 to the second total value S2 (0 + TL1). Also, the second total value processing unit 400 refers to the second storage unit 12a and specifies the note length associated with the mouth shape identification information corresponding to the second mouth shape identification information (the mouth shape identification information specified by the acquisition unit 100 based on the second mouth shape image) as the second note length TL2. The second total value processing unit 400 resets the second note length TL2 as the previous note length (that is, the previous note length becomes "TL2"). The second total value processing unit 400 repeats the above process each time the (1 + n)-th mouth shape image is extracted in a certain predetermined interval.
[0054] (Calculation unit) When the end time of the color change of the last lyric character included in a certain predetermined interval arrives, the calculation unit 500 calculates the singing tempo of the user in a certain predetermined interval based on the first total value and the second total value obtained so far.
[0055] Whether the end time of the color change of the last lyric character included in a certain predetermined interval has arrived can be determined by referring to the lyric data of a certain piece of music (the end time associated with the lyric characters included in a certain predetermined interval).
[0056] The calculation of the singing tempo can use known techniques. Specifically, the calculation unit 500 calculates the length of one beat in the user's karaoke singing within a predetermined section by dividing the first total value by the second total value. The calculation unit 500 calculates the user's singing tempo within a predetermined section from the calculated length of one beat and the performance tempo of a certain piece of music.
[0057] (Performance processing unit) The performance processing unit 600 controls the performance means 10d to perform the karaoke performance of the music.
[0058] The staff of the facility where the karaoke device K is installed operates the remote control device 40 to select a music piece for the karaoke performance. The remote control device 40 registers the music identification information of the selected music piece in the reservation queue.
[0059] The performance processing unit 600 acquires the accompaniment data of the music from the storage means 10a based on the music identification information registered in the reservation queue. The performance processing unit 600 reproduces the acquired accompaniment data and controls the performance means 10d to emit the karaoke performance sound from the speaker 20. The user can perform karaoke singing in accordance with the karaoke performance sound emitted from the speaker 20.
[0060] The performance processing unit 600 according to this embodiment adjusts the performance tempo of the next predetermined section of a certain predetermined section according to the calculated singing tempo. The adjustment of the performance tempo is to increase the performance tempo or decrease the performance tempo with respect to the performance tempo preset for the music.
[0061] The adjustment of the performance tempo can be performed by various methods. For example, the performance processing unit 600 determines whether the user's singing tempo within a predetermined section calculated by the calculation unit 500 is within a preset threshold. If it is not within the threshold (that is, when the deviation between the singing tempo and the performance tempo is large), the performance processing unit 600 makes an adjustment from the next predetermined section of a certain predetermined section so that the performance tempo of the music matches the calculated user's singing tempo.
[0062] Alternatively, the performance processing unit 600 calculates the difference between the singing tempo and the performance tempo of a certain piece of music. The performance processing unit 600 adjusts the performance tempo of the music in the next predetermined section according to the calculated difference.
[0063] As described above, the lyric characters included in one predetermined section are displayed on the same screen. Therefore, when the karaoke device K performs karaoke performance of the next phrase, the lyric characters included in the previous phrase are erased and the lyric characters included in the next phrase are displayed. That is, according to the karaoke device K according to the present embodiment, the performance tempo of the music can be adjusted at the timing when the display of the lyric characters is switched.
[0064] ==Regarding the processing of the karaoke device K== Next, with reference to FIGS. 5 and 6, a specific example of the processing of the karaoke device K in the present embodiment will be described. FIG. 5 is a flowchart showing the processing of the karaoke device K. FIG. 6 is a diagram showing the lip recognition information and extraction times obtained in the specific example in time series. In this example, it is assumed that in the facility where the karaoke device K is installed, the user U performs karaoke singing alone without using a microphone. Also, it is assumed that the lyric data of the music X shown in FIG. 3 is stored in the storage means 10a, and the lyric lip data of the music X shown in FIG. 4 is stored in the second storage unit 12a. Also, it is assumed that the performance tempo of the music X is 120 bpm.
[0065] The staff of the facility operates the remote control device 40 to select the music X suitable for the karaoke singing of the user U. The karaoke device K (performance processing unit 600) starts the karaoke performance of the selected music X (karaoke performance starts. Step 10).
[0066] The karaoke device K causes the display device 30 to display the lyric characters for each predetermined section based on the lyric data of the music piece X along with the karaoke performance of the music piece X (display the lyric characters for each predetermined section. Step 11). The user U vocalizes one character at a time in accordance with the displayed lyric characters. The camera 50 starts photographing the user U upon the start of the karaoke performance of the music piece X. The camera 50 sequentially outputs the photographed video to the karaoke main body 10. The karaoke main body 10 acquires the video of the user U performing karaoke singing (acquire the video of the user. Step 12).
[0067] The acquisition unit 100 acquires, in association with each other, the mouth shape identification information specified based on the mouth shape image extracted from the video obtained by photographing the user U performing karaoke singing of the music piece X and the extraction time which is the timing at which the mouth shape image was extracted (acquire the mouth shape identification information and the extraction time. Step 13).
[0068] When the first mouth shape image is extracted in a certain predetermined section of the music piece X, the setting unit 200 sets the first extraction time associated with the first mouth shape identification information specified based on the first mouth shape image as the previous extraction time, and refers to the second storage unit 12a to set the sound length associated with the mouth shape identification information corresponding to the first mouth shape identification information as the previous sound length (when the first mouth shape image is extracted, set the previous extraction time and the previous sound length. Step 14).
[0069] When the (1 + n)-th mouth shape image is extracted in a certain predetermined section, the first total value processing unit 300 adds the pronunciation time obtained by subtracting the previous extraction time from the (1 + n)-th extraction time associated with the (1 + n)-th mouth shape identification information specified based on the (1 + n)-th mouth shape image to the first total value, and re-sets the (1 + n)-th extraction time as the previous extraction time (when the (1 + n)-th mouth shape image is extracted, add the pronunciation time to the first total value and re-set the previous extraction time. Step 15).
[0070] When the (1 + n)-th mouth shape image is extracted in a certain predetermined section, the second total value processing unit 400 adds the previous note length to the second total value, refers to the second storage unit 12a, and sets the note length associated with the mouth shape identification information corresponding to the (1 + n)-th mouth shape identification information specified based on the (1 + n)-th mouth shape image as the (1 + n)-th note length, and resets the previous note length as the (1 + n)-th note length (when the (1 + n)-th mouth shape image is extracted, add the note length to the second total value and reset the previous note length. Step 16). Until the end time of the color change of the last lyric character included in a certain predetermined section arrives (when it is N in step 17), the karaoke device K repeatedly performs the processes from step 12 to step 17 (however, the process of step 14 is skipped).
[0071] When the end time of the color change of the last lyric character included in a certain predetermined section arrives (when it is Y in step 17), the calculation unit 500 calculates the singing tempo of the user U in a certain predetermined section based on the first total value and the second total value obtained so far (calculate the singing tempo in a certain predetermined section. Step 18).
[0072] The performance processing unit 600 adjusts the performance tempo of the next predetermined section according to the singing tempo calculated in step 18 (adjust the performance tempo of the next predetermined section. Step 19).
[0073] Until the karaoke performance ends (when it is Y in step 20), the karaoke device K repeatedly performs the processes from step 11 to step 19 for each predetermined section. Note that the karaoke device K may repeatedly perform the processes from step 11 to step 19 until a certain point (for example, until the end of the first performance section of the music) instead of until the end of the karaoke performance.
[0074] Here, the adjustment of the performance tempo will be specifically described. At the start point of the karaoke performance of music X, the previous extraction time, the previous note length, the first total value, and the second total value are all "0". Also, in this example, it is assumed that the predetermined section is a phrase.
[0075] The karaoke device K causes the display device 30 to display the lyric text "Put your head above the clouds" of the phrase F1 based on the lyric data of the music X along with the karaoke performance of the music X. The user U makes a sound of "a" in accordance with the displayed lyric text. The camera 50 outputs the video I1 that has captured the user U making the sound of "a" to the karaoke main body 10.
[0076] The acquisition unit 100 extracts the mouth shape image MI1 from the output video I1. Then, the acquisition unit 100 refers to the first storage unit 11a and specifies one mouth shape identification information "A" associated with the same mouth shape image as the extracted mouth shape image MI1. Also, the acquisition unit 100 specifies the extraction time "16014 ms", which is the timing when the mouth shape image MI1 is extracted, as the extraction time.
[0077] The acquisition unit 100 acquires the specified one mouth shape identification information "A" and the extraction time "16014 ms" in association with each other (see Fig. 6).
[0078] Since the first mouth shape image MI1 in the phrase F1 of the music X is extracted, the setting unit 200 sets the extraction time "16014 ms" associated with the mouth shape identification information "A" specified based on the mouth shape image MI1 as the previous extraction time ET1, and refers to the second storage unit 12a to set the sound length "1.5" associated with the mouth shape identification information corresponding to the mouth shape identification information "A" as the previous sound length TL1.
[0079] Next, the camera 50 outputs the video I2 that has captured the user U making the sound of "ta" to the karaoke main body 10.
[0080] The acquisition unit 100 extracts the mouth shape image MI2 from the output video I2. Then, the acquisition unit 100 refers to the first storage unit 11a and specifies one mouth shape identification information "IA" associated with the same mouth shape image as the extracted mouth shape image MI2. Also, the acquisition unit 100 specifies the extraction time "16910 ms", which is the timing when the mouth shape image MI2 is extracted, as the extraction time.
[0081] The acquisition unit 100 acquires by associating the identified single mouth shape identification information "IA" with the extraction time "16910 ms" (see FIG. 6).
[0082] Since the second mouth shape image MI2 was extracted in the phrase F1 of the music piece X, the first total value processing unit 300 specifies the extraction time "16910 ms" associated with the mouth shape identification information "IA" as the extraction time ET2. The first total value processing unit 300 subtracts the previous extraction time ET1 (16014 ms) from the extraction time ET2 (16910 ms) to obtain the pronunciation time PT1 (896 ms), and adds it to the first total value S1 (0 ms + 896 ms = 896 ms). Also, the first total value processing unit 300 re - sets the extraction time ET2 as the previous extraction time (that is, the previous extraction time becomes "ET2 (16910 ms)").
[0083] Also, since the second mouth shape image MI2 was extracted in the phrase F1 of the music piece X, the second total value processing unit 400 adds the previous sound length TL1 (1.5) to the second total value S2 (0 + 1.5 = 1.5). Also, the second total value processing unit 400 refers to the second storage unit 12a and specifies the sound length "0.5" associated with the mouth shape identification information corresponding to the mouth shape identification information "IA" as the sound length TL2. The second total value processing unit 400 re - sets the sound length TL2 as the previous sound length (that is, the previous sound length becomes "TL2 (0.5)").
[0084] Next, the camera 50 outputs the video I3 that captures the user U who pronounces "ma" to the karaoke body 10.
[0085] The acquisition unit 100 extracts the mouth shape image MI3 from the output video I3. Then, the acquisition unit 100 refers to the first storage unit 11a and specifies the single mouth shape identification information "NA" associated with the same mouth shape image as the extracted mouth shape image MI3. Also, the acquisition unit 100 specifies the single time "17212 ms", which is the timing when the mouth shape image MI3 was extracted, as the extraction time.
[0086] The acquisition unit 100 acquires by associating the identified single lip shape identification information "NA" with the extraction time "17212 ms" (see FIG. 6).
[0087] Since the third lip shape image MI3 was extracted in the phrase F1 of the music X, the first total value processing unit 300 identifies the extraction time "17212 ms" associated with the lip shape identification information "NA" as the extraction time ET3. The first total value processing unit 300 subtracts the previous extraction time ET2 (16910 ms) from the extraction time ET3 (17212 ms) to obtain the pronunciation time PT2 (302 ms), and adds it to the first total value S1 (0 ms + 896 ms + 302 ms = 1198 ms). Also, the first total value processing unit 300 resets the extraction time ET3 as the previous extraction time (that is, the previous extraction time becomes "ET3 (17212 ms)").
[0088] Also, since the third lip shape image MI3 was extracted in the phrase F1 of the music X, the second total value processing unit 400 adds the previous sound length TL2 (0.5) to the second total value S2 (0 + 1.5 + 0.5 = 2.0). Also, the second total value processing unit 400 refers to the second storage unit 12a and identifies the sound length "1.0" associated with the lip shape identification information corresponding to the lip shape identification information "NA" as the sound length TL3. The second total value processing unit 400 resets the sound length TL3 as the previous sound length (that is, the previous sound length becomes "TL3 (1.0)").
[0089] In this example, by repeating the above processing, when the color change end time "23300 ms" of the last lyric character "shi" included in the phrase F1 arrives, the first total value S1 becomes "7188 ms" and the second total value S2 becomes "12.0".
[0090] The calculation unit 500 calculates the length of one beat " " in the user's karaoke singing in the phrase F1 by dividing the first total value S1 (7188 ms) by the second total value S2 (12.0). The calculation unit 500 calculates the singing tempo "100 bpm" of the user U in the phrase F1 from the calculated length of one beat and the performance tempo "120 bpm" of the music X.
[0091] The performance processing unit 600 obtains the difference between the calculated singing tempo "100 bpm" and the performance tempo "120 bpm" of music X. The performance processing unit 600 adjusts (for example, -20 bpm) the performance tempo of music X in phrase F2 following phrase F1 according to the obtained difference "20".
[0092] As is clear from the above, the karaoke device K according to the present embodiment stores in association a plurality of mouth shape images showing the mouth shape when pronouncing a vowel and the mouth shape in a closed mouth state, and mouth shape identification information indicating the mouth shape corresponding to the mouth shape image in a first storage unit 11a. The karaoke device K also stores, for each music, lyric mouth shape data in which a pronunciation time indicating the timing at which each lyric character included in the lyric data should be pronounced, a predetermined section including the pronunciation time, mouth shape identification information indicating the mouth shape to be pronounced at the pronunciation time, and a sound length indicating the ratio of the length of pronouncing each lyric character to the length of one beat based on the performance tempo of the music are associated with each other in a second storage unit 12a. An acquisition unit 100 acquires, in association, mouth shape identification information specified based on a mouth shape image extracted from a video obtained by photographing a user performing karaoke singing of a certain music, and an extraction time which is the timing at which the mouth shape image is extracted. When the first mouth shape image is extracted in a certain predetermined section of a certain music, a setting unit 200 sets the first extraction time associated with the first mouth shape identification information specified based on the first mouth shape image as the previous extraction time, and refers to the second storage unit 12a to set the sound length associated with the mouth shape identification information corresponding to the first mouth shape identification information as the previous sound length. When the (1 + n)th mouth shape image is extracted in a certain predetermined section, a first total value processing unit 300 adds the pronunciation time obtained by subtracting the previous extraction time from the (1 + n)th extraction time associated with the (1 + n)th mouth shape identification information specified based on the (1 + n)th mouth shape image to the first total value, and re-sets the (1 + n)th extraction time as the previous extraction time. When the (1 + n)th mouth shape image is extracted in a certain predetermined section, a second total value processing unit 400 adds the previous sound length to the second total value, refers to the second storage unit 12a, sets the sound length associated with the mouth shape identification information corresponding to the (1 + n)th mouth shape identification information specified based on the (1 + n)th mouth shape image as the (1 + n)th sound length, and re-sets the (1 + n)th sound length as the previous sound length. When the end time of the color change of the last lyric character included in a certain predetermined section arrives, a calculation unit 500 calculates the singing tempo of the user in a certain predetermined section based on the first total value and the second total value obtained so far. A performance processing unit 600 adjusts the performance tempo of the next predetermined section of a certain predetermined section according to the calculated singing tempo.
[0093] According to such a karaoke device K, by using a lip image extracted from an image obtained by photographing a user performing karaoke singing of a certain song, the singing tempo of a certain predetermined section can be calculated, and the performance tempo of the next predetermined section can be adjusted. Therefore, even when a user performs karaoke singing without using a microphone, the performance tempo of the song can be adjusted according to the singing tempo of the user. That is, according to the karaoke device K according to the present embodiment, the performance tempo of the song can be automatically adjusted without using the singing voice of the user.
[0094] <Modification Example 1> In the above embodiment, an example of a single user has been described. On the other hand, in a nursing facility or the like, a single song may be karaoke-sung by a plurality of people. In such a case, the karaoke device K can adjust the performance tempo of the song in consideration of the singing tempo of each user.
[0095] (Calculation unit) The calculation unit 500 according to this modification example calculates the singing tempo for each user when the karaoke device K is used by a plurality of users. The calculation of the singing tempo can be performed in the same manner as in the embodiment. The identification of the user can be performed based on a face image specified using a known face recognition technique, for example, from the photographed image.
[0096] (Performance processing unit) The performance processing unit 600 according to this modification example adjusts the performance tempo of the next predetermined section of a certain predetermined section according to the calculated singing tempos of a plurality of users.
[0097] For example, as a result of three users U1 to U3 performing karaoke singing of phrase F1 of song X in the embodiment, it is assumed that the singing tempo of user U1 is calculated as "120 bpm", the singing tempo of user U2 is calculated as "110 bpm", and the singing tempo of user U3 is calculated as "100 bpm".
[0098] In this case, the performance processing unit 600 can adjust the performance tempo of the phrase F2 according to "110 bpm", which is the average value of the singing tempos of three people. Further, the performance processing unit 600 can adjust the performance tempo of the phrase F2 according to "100 bpm", which is the slowest singing tempo among the singing tempos of three people. Alternatively, the performance processing unit 600 can adjust the performance tempo of the phrase F2 according to "105 bpm", which is the average value of the singing tempos of the users U2 and U3 whose singing tempos are slower than the performance tempo "120 bpm" of the music X among the singing tempos of three people.
[0099] As is clear from the above, when the karaoke device K according to this modification example is used by a plurality of users, the calculation unit 500 calculates the singing tempo for each user, and the performance processing unit 600 can adjust the performance tempo of the next predetermined section in a certain predetermined section according to the calculated singing tempos of the plurality of users. According to such a karaoke device K, even when a single piece of music is karaoke-sung by a plurality of people, the performance tempo of the music can be automatically adjusted without using the singing voices of each user.
[0100] <Modification Example 2> A user whose singing tempo is faster than the performance tempo of the music may start karaoke-singing the next predetermined section during a certain predetermined section. In this case, even if the lip recognition information is specified in a certain predetermined section (that is, even if the lip image is extracted), the corresponding lip recognition information (the lip recognition information in a certain predetermined section) is not included in the second storage unit 12a. In this modification example, even in such a case, the performance tempo of the music can be adjusted.
[0101] [Control means] In this modification example, by the CPU executing the program stored in the memory, the control means 10e functions as the acquisition unit 100, the setting unit 200, the first total value processing unit 300, the second total value processing unit 400, the calculation unit 500, the performance processing unit 600, and the specifying unit 700 (see FIG. 7).
[0102] (Specifying unit) The specific part 700 refers to the second storage part 12a and specifies the number of lip shape identification information corresponding to the lyric characters included in a certain predetermined section, which is displayed on the display means along with the karaoke performance of a certain piece of music.
[0103] The karaoke device K causes the display device 30 to display lyric characters for each predetermined section based on the lyric data of the selected piece of music along with the karaoke performance of the selected piece of music.
[0104] The specific part 700 refers to the lyric lip shape data of a certain piece of music stored in the second storage part 12a and specifies the number of lip shape identification information corresponding to the lyric characters included in a certain predetermined section, which is displayed on the display device 30.
[0105] For example, in the example of the embodiment, the karaoke device K causes the display device 30 to display the lyric characters "Put your head above the clouds" of the phrase F1 based on the lyric data of the music X along with the karaoke performance of the music X.
[0106] The specific part 700 refers to the lyric lip shape data of the music X stored in the second storage part 12a and specifies the number "12" of lip shape identification information corresponding to the lyric characters included in the phrase F1, which is displayed on the display device 30.
[0107] (First total value processing part) When the (n + 1)-th lip shape image is extracted in a certain predetermined section and the value of (n + 1) is less than or equal to the number of lip shape identification information for which the value of (n + 1) is specified, the first total value processing part 300 according to this modification example adds the pronunciation time obtained by subtracting the previous extraction time from the (n + 1)-th extraction time associated with the (n + 1)-th lip shape identification information specified based on the (n + 1)-th lip shape image to the first total value, and resets the (n + 1)-th extraction time as the previous extraction time.
[0108] (Second total value processing part) When the 1 + n-th mouth shape image is extracted in a certain predetermined section and the value of 1 + n is less than or equal to the number of identified mouth shape identification information, the second total value processing unit 400 according to this modification example adds the previous sound length to the second total value, refers to the second storage unit 12a, and sets the sound length associated with the mouth shape identification information corresponding to the (1 + n)-th mouth shape identification information specified based on the (1 + n)-th mouth shape image as the (1 + n)-th sound length, and re-sets the (1 + n)-th sound length as the previous sound length.
[0109] For example, it is assumed that the number "12" of the mouth shape identification information corresponding to the lyric characters included in the phrase F1 as described above is identified.
[0110] In this case, according to the example of the embodiment, the second mouth shape image MI2 is extracted in the phrase F1 of the music X, and the value of "2" is less than or equal to the number "12" of the identified mouth shape identification information. Therefore, the first total value processing unit 300 and the second total value processing unit 400 execute the processing described in the embodiment.
[0111] On the other hand, in the example of the embodiment, it is assumed that the singing tempo of the user U is fast and the user U has pronounced the first lyric character "shi" of the phrase 2 in the phrase F1. Even in this case, the camera 50 outputs the video I13 that captures the user U who pronounces "shi" to the karaoke main body 10.
[0112] The acquisition unit 100 extracts the mouth shape image MI13 from the output video I13. Then, the acquisition unit 100 refers to the first storage unit 11a and identifies one mouth shape identification information "IA" associated with the same mouth shape image as the extracted mouth shape image MI13. In addition, the acquisition unit 100 specifies the extraction time as one time "22000 ms" which is the timing when the mouth shape image MI13 is extracted.
[0113] The acquisition unit 100 acquires the specified one mouth shape identification information "IA" and the extraction time "22000 ms" in association with each other.
[0114] Here, the 13th mouth shape image MI13 is extracted from the phrase F1 of the music piece X, but the value of "13" is equal to or greater than the number "12" of the specified mouth shape identification information. In this case, the first total value processing unit 300 and the second total value processing unit 400 do not execute the processing described in the embodiment.
[0115] As is clear from the above, the karaoke device K according to this modification example refers to the second storage unit 12a and has an identification unit 700 that identifies the number of mouth shape identification information corresponding to the lyric characters included in a predetermined section and displayed on the display means along with the karaoke performance of a certain music piece. The first total value processing unit 300 according to this modification example, when the (1 + n)th mouth shape image is extracted in a certain predetermined section and the value of (1 + n) is less than or equal to the number of the specified mouth shape identification information, subtracts the previous extraction time from the pronunciation time obtained by subtracting the previous extraction time from the extraction time corresponding to the (1 + n)th mouth shape identification information specified based on the (1 + n)th mouth shape image, and adds the result to the first total value, and resets the (1 + n)th extraction time as the previous extraction time. The second total value processing unit 400 according to this modification example, when the (1 + n)th mouth shape image is extracted in a certain predetermined section and the value of (1 + n) is less than or equal to the number of the specified mouth shape identification information, adds the previous sound length to the second total value, refers to the second storage unit 12a, sets the sound length corresponding to the mouth shape identification information corresponding to the (1 + n)th mouth shape identification information specified based on the (1 + n)th mouth shape image as the (1 + n)th sound length, and resets the (1 + n)th sound length as the previous sound length. According to such a karaoke device K, even when the singing tempo of the user is too fast, the performance tempo of the music piece can be automatically adjusted according to the singing tempo of the user.
[0116] <Modification Example 3> In the example of the embodiment, an example in which the lyric mouth shape data is set in advance for each music piece has been described. On the other hand, the karaoke device K may generate the lyric mouth shape data each time a certain music piece is selected and store it in the second storage unit 12a.
[0117] Specifically, the karaoke device K refers to the lyric data of a music piece, and for each music piece, generates lyric mouth shape data in which the pronunciation time indicating the timing at which each character included in the lyric data should be pronounced, a predetermined section including the pronunciation time, mouth shape identification information indicating the mouth shape when pronouncing at the timing, and the sound length indicating the ratio of the length of each character pronounced to the length of one beat based on the performance tempo of the music piece are associated with each other, and can store the data in the second storage unit 12a.
[0118] The predetermined section is associated with each lyric character. The pronunciation time can be set as the start time of color change (or the time obtained by adding a predetermined time (for example, 200 ms) to the start time of color change) associated with each lyric character. The mouth shape identification information can be specified for each lyric character by referring to a table (for example, refer to FIG. 3 of Japanese Patent Application Laid-Open No. 2008-310382) associating preset mouth shape identification information with characters. The sound length can be specified based on the performance tempo of the music piece, the start time of color change and the end time of color change associated with each lyric character. For example, when the performance tempo is "120 bpm", the length of one beat is "500 ms". Here, if the time from the start of color change to the end of color change of a certain lyric character is "750 ms", the sound length is "1.5".
[0119] <Others> It is also possible to supply a program to a computer using a non-transitory computer-readable medium (non-transitory computer readable medium with an executable program thereon) storing the above program. Examples of the non-transitory computer-readable medium include magnetic recording media (for example, flexible disks, magnetic tapes, hard disk drives), CD-ROM (Read Only Memory), and the like.
[0120] The above embodiments are presented as examples and do not limit the scope of the invention. The above configurations can be implemented in appropriate combinations, and various omissions, replacements, and changes can be made without departing from the gist of the invention. The above embodiments and their modifications are included in the scope and gist of the invention, and are also included in the invention described in the claims and its equivalent scope.
Explanation of Reference Numerals
[0121] K Karaoke device 10 Karaoke main body 11a First storage unit 12a Second storage unit 100 Acquisition unit 200 Setting unit 300 First total value processing unit 400 Second total value processing unit 500 Calculation unit 600 Performance processing unit 700 Identification unit
Claims
1. A first storage unit that stores, in association with each other, a plurality of mouth shape images showing the mouth shape when pronouncing a vowel and the mouth shape in a closed-mouth state, and mouth shape identification information indicating the mouth shape corresponding to the mouth shape image; A second storage unit that stores, for each piece of music, lyric mouth shape data in which a pronunciation time indicating the timing at which each lyric character included in the lyric data should be pronounced, a predetermined section including the pronunciation time, mouth shape identification information indicating the mouth shape to be pronounced at the pronunciation time, and a sound length indicating the ratio of the length of pronunciation of each lyric character to the length of one beat based on the performance tempo of the music are associated with each other; An acquisition unit that acquires, in association with each other, mouth shape identification information specified based on a mouth shape image extracted from a video obtained by photographing a user performing karaoke singing of a certain piece of music, and an extraction time that is the timing at which the mouth shape image was extracted; When the first mouth shape image is extracted in a certain predetermined section of the certain piece of music, a first extraction time associated with the first mouth shape identification information specified based on the first mouth shape image is set as the previous extraction time, and the second storage unit is referred to, and a sound length associated with the mouth shape identification information corresponding to the first mouth shape identification information is set as the previous sound length; When the (1 + n)th mouth shape image is extracted in the certain predetermined section, a pronunciation time obtained by subtracting the previous extraction time from the (1 + n)th extraction time associated with the (1 + n)th mouth shape identification information specified based on the (1 + n)th mouth shape image is added to a first total value, and the (1 + n)th extraction time is reset as the previous extraction time; When the (1 + n)th mouth shape image is extracted in the certain predetermined section, the previous sound length is added to a second total value, and the second storage unit is referred to, and a sound length associated with the mouth shape identification information corresponding to the (1 + n)th mouth shape identification information specified based on the (1 + n)th mouth shape image is set as the (1 + n)th sound length, and the (1 + n)th sound length is reset as the previous sound length; When the end time of color change of the last lyric character included in the certain predetermined section arrives, a calculation unit that calculates the singing tempo of the user in the certain predetermined section based on the first total value and the second total value obtained so far; A performance processing unit that adjusts the performance tempo of the next predetermined section of the certain predetermined section according to the calculated singing tempo; A karaoke device having the above components.
2. When the karaoke device is used by a plurality of users, the calculation unit calculates the singing tempo for each user, The performance processing unit adjusts the performance tempo of the next predetermined section of the certain predetermined section according to the calculated singing tempos of the plurality of users. The karaoke device according to claim 1, characterized in that.
3. A specifying unit that refers to the second storage unit and specifies the number of lip shape identification information corresponding to the lyric characters included in the certain predetermined section, which is displayed on the display means along with the karaoke performance of the certain piece of music. When the (1 + n)-th lip shape image is extracted in the certain predetermined section and the value of 1 + n is less than or equal to the number of lip shape identification information whose value has been specified, the first total value processing unit subtracts the previous extraction time from the (1 + n)-th extraction time associated with the (1 + n)-th lip shape identification information specified based on the (1 + n)-th lip shape image, adds the resulting pronunciation time to the first total value, and resets the (1 + n)-th extraction time as the previous extraction time. When the (1 + n)-th lip shape image is extracted in the certain predetermined section and the value of 1 + n is less than or equal to the number of lip shape identification information whose value has been specified, the second total value processing unit adds the previous sound length to the second total value, refers to the second storage unit, sets the sound length associated with the lip shape identification information corresponding to the (1 + n)-th lip shape identification information specified based on the (1 + n)-th lip shape image as the (1 + n)-th sound length, and resets the (1 + n)-th sound length as the previous sound length. The karaoke device according to claim 1 or 2, characterized in that.
Citation Information
Patent Citations
Tempo controller for karaoke
JP1998149180A