Karaoke System
The karaoke system connects a server and karaoke device to identify singers through singing voice and video data, allowing users to engage in lip-sync guessing and singer identification, enhancing user experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-27
- Publication Date
- 2026-03-11
AI Technical Summary
Existing karaoke systems lack the ability to identify or guess the singer based on their singing voice and lip-syncing performance.
A karaoke system with a server device and karaoke device connected via a network, where the server stores content information including song identification, singing voice data, and video data, and the karaoke device generates composite video and sound to display and emit singing voices for lip-sync guessing or singer identification.
Enables users to guess the singer by comparing singing voices and video data, providing new content that enhances user engagement and enjoyment.
Smart Images

Figure 2026042560000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a karaoke system. [Background technology]
[0002] The karaoke system can provide various contents using the karaoke device.
[0003] For example, Patent Document 1 discloses a technology for a karaoke system that includes a central device that transmits karaoke song data and a karaoke terminal device that receives and plays the karaoke song data, in which content data is generated by combining video and audio data transmitted from the karaoke terminal device with a selected piece of audio data and video data paired with the audio data, and the content data is made available for viewing. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2011-59619 Summary of the Invention [Problem to be solved by the invention]
[0005] The object of the present invention is to provide a karaoke system capable of implementing new content that identifies a singer who is lip-syncing (moving their mouths without vocalizing along with the karaoke performance) or a singer who is singing karaoke, based on the singing voice emitted from the sound emitting means of the karaoke device and the image displayed on the display means. [Means for solving the problem]
[0006] One invention for achieving the above object is a karaoke system in which a server device and a karaoke device are communicably connected, the server device having an information storage unit that stores a plurality of pieces of content information, each piece of content information including song identification information of a song, singing voice data corresponding to a singer singing the song in karaoke, and video data corresponding to a video of the singer singing the song in karaoke, an identification unit that identifies one piece of content information from the content information including song identification information of a song selected by a user of the karaoke device, an extraction unit that extracts at least one other piece of content information based on a similarity between the singing voice data included in the identified piece of content information and singing voice data included in other content information other than the identified piece of content information, the content information including the song identification information of the song, and a video data storage unit that stores a plurality of pieces of content information, each piece of content information including song identification information of a song, the information storage unit storing a plurality of pieces of content information, each piece of content information including song identification information of a song selected by a user of the karaoke device, and a video data storage unit that stores a plurality of pieces of content information, each piece of content information including song identification information of a song selected by a user of the karaoke device, the information storage unit storing ... the information storage unit storing a plurality of pieces of content information, each piece of content information including song identification information of a song selected by a user of the karaoke device, and a transmission processing unit that transmits other singing voice data and other video data included in the user information to the karaoke device, wherein the karaoke device has: a generation unit that generates composite video data by combining each video data so that a video based on the transmitted one video data and a video based on the transmitted other video data are displayed on a display means in a predetermined layout; a sound emission processing unit that emits singing voice based on the transmitted other singing voice data from a sound emission means in time with a karaoke performance of the selected one song; a display processing unit that displays a composite video based on the generated composite video data on the display means in the predetermined layout in time with a karaoke performance of the selected one song; and a correct answer determination unit that, when a user selects a video from the displayed composite videos and no singing voice data is associated with the video data corresponding to the selected video, notifies the user that the selection is correct. Further, one invention for achieving the above object is a karaoke system in which a server device and a karaoke device are communicatively connected, the server device comprising: an information storage unit that stores a plurality of pieces of content information, each piece including song identification information of a song, singing voice data corresponding to a singer singing the song in karaoke, and video data corresponding to a video of the singer singing the song in karaoke; an identification unit that identifies one piece of content information from the content information including song identification information of a song selected by a user of the karaoke device; an extraction unit that extracts at least one other piece of content information based on a similarity between the singing voice data included in the identified piece of content information and singing voice data included in other content information other than the identified piece of content information, the singing voice data and video data included in the identified piece of content information; and a transmission processing unit that transmits the extracted other video data included in the other content information to the karaoke device, wherein the karaoke device has: a generation unit that generates composite video data by combining the transmitted one video data and the transmitted other video data so that the video based on the one video data and the video based on the transmitted other video data are displayed on a display means in a predetermined layout; a sound emission processing unit that emits singing voice based on the transmitted one singing voice data from a sound emission means in time with the karaoke performance of the selected one song; a display processing unit that displays a composite video based on the generated composite video data in the predetermined layout in time with the karaoke performance of the selected one song; and a correct answer determination unit that, when a user selects a video from the displayed composite videos and singing voice data is associated with the video data corresponding to the selected video, notifies the user that the selection is correct. Other features of the present invention will become apparent from the following description and drawings. [Effects of the Invention]
[0007] According to the present invention, it is possible to implement new content in which the singer who is lip-syncing or singing karaoke can be guessed from the singing voice emitted from the sound emitting means of the karaoke device and the image displayed on the display means. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a diagram showing a karaoke system according to a first embodiment. [Figure 2] FIG. 2 is a diagram illustrating a server device according to the first embodiment. [Figure 3] 1 is a diagram showing a karaoke device according to a first embodiment. [Figure 4] 1 is a diagram showing a karaoke main unit according to a first embodiment. FIG. [Figure 5] 3 is a flowchart showing the processing of the karaoke system according to the first embodiment. [Figure 6] FIG. 2 is a diagram showing a composite image displayed on the display device according to the first embodiment. [Figure 7A] FIG. 4 is a diagram showing a correct answer screen displayed on the display device according to the first embodiment. [Figure 7B] FIG. 10 is a diagram showing an incorrect answer screen displayed on the display device according to the first embodiment. [Figure 8] FIG. 10 is a diagram illustrating a server device according to a second embodiment. [Figure 9] FIG. 10 is a diagram showing a karaoke main unit according to a second embodiment. [Figure 10] 10 is a flowchart showing the processing of the karaoke system according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] First Embodiment The karaoke system according to this embodiment will be described with reference to FIGS. 1 to 7B.
[0010] ==Karaoke System== The karaoke system includes a server device and karaoke devices. The server device is a computer that manages various information. The karaoke device is a device for playing karaoke songs and for singers to sing karaoke. In this embodiment, the karaoke device can implement lip-sync content. The lip-sync content allows a user to select a video that they think is lip-synced from multiple videos displayed on the display means of the karaoke device and notifies the user whether their selection is correct or incorrect. The karaoke device is communicably connected to the server device via a network. The network is, for example, a transmission path such as a public telephone network or the Internet. The server device can also communicate with other karaoke devices (not shown) via the network.
[0011] 1, the karaoke system 1 according to this embodiment includes a karaoke device K and a server device S. The karaoke device K is connected to the server device S via a network N so as to be able to communicate with each other.
[0012] ==Server Device== 2, the server device S according to this embodiment includes a storage unit 10, a communication unit 20, and a control unit 30. Each component is connected to a bus B via an interface (not shown).
[0013] [Storage means] The storage means 10 is a large-capacity storage device that stores various types of data. In this embodiment, a part of the storage area of the storage means 10 functions as an information storage unit 11.
[0014] (Information storage unit) The information storage unit 11 stores a plurality of pieces of content information. The content information according to this embodiment is information that may be used to implement lip-sync content in the karaoke device K.
[0015] The content information includes song identification information of the song, singing voice data, and video data.
[0016] The song identification information is information specific to each song, such as a song ID for identifying each song. The singing voice data is data corresponding to the singing voice of the singer singing the song as karaoke (specifically, audio data of the singing voice). The video data is data corresponding to video of the singer singing the song as karaoke (specifically, video data of the singer). The singer can be filmed, for example, by a camera installed in the karaoke machine.
[0017] For example, when a singer U sings a song X at a karaoke device communicatively connected to the server device S, the karaoke device acquires song identification information for the song X, singing voice data corresponding to the singer U's singing voice when singing the song X at karaoke, and video data corresponding to a video of the singer U singing the song X at karaoke. The karaoke device generates content information that associates the acquired data and transmits it to the server device S. The server device S stores the received content information in the information storage unit 11.
[0018] [Means of communication] The communication means 20 provides an interface for communicating with the karaoke device K.
[0019] [Control means] The control means 30 performs various controls in the server device S. The control means 30 includes a CPU and a memory (neither of which are shown). The CPU executes programs stored in the memory to realize various functions.
[0020] In this embodiment, the CPU executes a program for lip-syncing content stored in the memory, causing the control means 30 to function as an identifying unit 31, an extracting unit 32, and a transmission processing unit 33 (see FIG. 2).
[0021] (Specific part) The identifying unit 31 identifies one piece of content information from among the content information including the music identification information of one piece of music selected by the user of the karaoke device K.
[0022] A user who wishes to perform lip-sync content operates a remote control device 80 (described later) of the karaoke device K to select the lip-sync content. The user then operates the remote control device 80 to select a song. The karaoke device K transmits song identification information of the selected song to the server device S.
[0023] Based on the received song identification information of a song, the identification unit 31 identifies one piece of content information that includes the song identification information from among the content information stored in the information storage unit 11. Identification can be performed by various methods. For example, the identification unit 31 can randomly identify one piece of content information from among the content information that includes the song identification information of a song. Alternatively, the identification unit 31 can identify, from among the content information that includes the song identification information of a song, the content information that was most recently stored in the information storage unit 11 as the piece of content information (in this case, the content information is stored in the information storage unit 11 in association with information on the date and time of storage).
[0024] (Extraction part) The extraction unit 32 extracts at least one other piece of content information based on the similarity between the singing voice data included in the identified piece of content information and the singing voice data included in the other piece of content information.
[0025] The similarity indicates the degree to which two different singing voice data are similar. The similarity can be expressed, for example, by a number between 0 and 100%. The other content information is content information including the song identification information of one song, but other than the one content information. In other words, the other content information is content information that has not been identified by the identification unit 31 among the content information including the song identification information of one song selected by the user of the karaoke device K.
[0026] The extraction unit 32 calculates a similarity between the singing voice data included in the identified piece of content information and the singing voice data included in the other piece of content information. The extraction unit 32 extracts the other piece of content information whose calculated similarity is equal to or greater than a predetermined value.
[0027] The predetermined value is a reference value for extracting other content information. When the similarity is expressed as a number between 0 and 100% as described above, the predetermined value is set as a single number, such as 80% or 60%. The number of other content information to be extracted only needs to be at least one. The number of other content information to be extracted may also be set in advance. Furthermore, if other content information having a similarity equal to or greater than the set predetermined value cannot be extracted, the extraction unit 32 may change the predetermined value and then extract other content information again.
[0028] More specifically, the extraction unit 32 calculates the similarity based on the timing of pronunciation of each syllable extracted from singing voice data included in one piece of content information and the timing of pronunciation of each syllable extracted from singing voice data included in another piece of content information, and extracts other piece of content information for which the similarity is equal to or greater than a predetermined value.
[0029] The extraction unit 32 extracts each syllable contained in the singing voice data using a known voice recognition technique, and determines the vocalization timing of each syllable by linking the vocalization time of each syllable. The vocalization time can be expressed as elapsed time, with the start of the karaoke performance being set as 0. The vocalization timing may be determined for the entire song, or may be determined only for a predetermined section (for example, the verse or chorus of the first verse).
[0030] The extraction unit 32 compares the vocalization timing of the first syllable extracted from the singing voice data included in one piece of content information with the vocalization timing of the first syllable extracted from the singing voice data included in the other piece of content information, and checks whether the difference in vocalization time is within a predetermined range (for example, within 100 msec). If the difference in vocalization time is within the predetermined range, the extraction unit 32 determines that the syllable is a "match." On the other hand, if the difference in vocalization time is not within the predetermined range, the extraction unit 32 determines that the syllable is a "mismatch." The extraction unit 32 determines whether each extracted syllable is a "match" or a "mismatch."
[0031] After determining whether all the extracted syllables are "matched" or "mismatched," the extraction unit 32 calculates the ratio of syllables determined to be "matched" to all the syllables as the degree of similarity.
[0032] The extraction unit 32 checks whether the obtained similarity is equal to or greater than a predetermined value, and extracts other content information whose obtained similarity is equal to or greater than the predetermined value.
[0033] Alternatively, the extraction unit 32 may determine the similarity by comparing the waveform of the singing voice corresponding to the singing voice data contained in the identified content information with the waveform of the singing voice corresponding to the singing voice data contained in the other content information.
[0034] Note that the user may select the level of difficulty when selecting lip-sync content. In this case, the extraction unit 32 can change the predetermined value depending on the selected level of difficulty. For example, suppose there are three levels of difficulty, "difficult," "normal," and "easy," and the user selects "difficult." In this case, the extraction unit 32 sets the predetermined value to, for example, 90%. On the other hand, suppose the user selects "easy." In this case, the extraction unit 32 sets the predetermined value to, for example, less than 70%.
[0035] (Transmission processing unit) The transmission processing unit 33 transmits to the karaoke device K one piece of video data included in the identified one piece of content information and other singing voice data and other video data included in the extracted other piece of content information.
[0036] The transmission processing unit 33 reads out the video data included in the one piece of content information identified by the identification unit 31 (i.e., the one piece of video data) from the information storage unit 11. The transmission processing unit 33 also reads out the singing voice data and video data included in the other piece of content information extracted by the extraction unit 32 (i.e., the other singing voice data and the other video data) from the information storage unit 11.
[0037] The transmission processing unit 33 transmits the read one video data to the karaoke device K in association with the other singing voice data and other video data.
[0038] ==Karaoke Equipment== As shown in FIG. 3, the karaoke device K includes a karaoke main unit 40, a speaker 50, a display device 60, a microphone 70, and a remote control device 80.
[0039] The karaoke machine main unit 40 performs various controls related to karaoke performance and singing, such as controlling the performance of selected songs, displaying lyrics and background images, and processing audio signals input through the microphone 70. The speaker 50 is configured to emit sound based on signals from the karaoke machine main unit 40. The speaker 50 is an example of "sound emitting means." A speaker (not shown) installed in the karaoke room or other room where the karaoke machine K is installed may also be used as the "sound emitting means." The display device 60 is configured to display videos and images on a screen based on signals from the karaoke machine main unit 40. The display device 60 is an example of "display means." A display (not shown) installed in the karaoke room or other room where the karaoke machine K is installed may also be used as the "display means." The microphone 70 is configured to convert the singer's singing voice based on the karaoke performance into an analog audio signal and input it to the karaoke machine main unit 40. The remote control device 80 is a device for performing various operations on the karaoke machine main unit 40. The remote control device 80 is an example of "display means."
[0040] 4, the karaoke unit 40 includes a storage unit 40a, a communication unit 40b, an input unit 40c, a performance unit 40d, and a control unit 40e. Each component is connected to a bus B via an interface (not shown).
[0041] [Storage means] The storage means 40a is a large-capacity storage device that stores various types of data, including music data.
[0042] The song data is provided with song identification information. The song data includes accompaniment data and reference data. The accompaniment data is data that is the source of the karaoke performance sound. The reference data is data that indicates the main melody of the song. In addition, the storage means 40a stores, for each song, lyric subtitle data for displaying song lyrics on the display device 60 or the like in sync with the karaoke performance, and background video data such as background video to be displayed on the display device 60 or the like during the karaoke performance.
[0043] Moreover, the storage means 40a according to the present embodiment stores one piece of video data, other singing voice data, and other video data transmitted from the server device S (transmission processing unit 33) in association with each other.
[0044] [Communication means / input means] The communication means 40b provides an interface for communicating with the remote control device 80 and the server device S. The input means 40c is configured to allow the user to input various instructions. The input means 40c is a button or the like provided on the karaoke main unit 40. Alternatively, the remote control device 80 may function as the input means 40c.
[0045] [Means of performance] Based on the control of the control means 40e, the performance means 40d performs karaoke performance of the music piece and processes signals based on the singing voice input through the microphone 70. The performance means 40d includes a sound source, a mixer, an amplifier, etc. (none of which are shown).
[0046] [Control means] The control means 40e performs various controls in the karaoke device K. The control means 40e includes a CPU and a memory (neither of which is shown). The CPU executes programs stored in the memory to realize various functions.
[0047] In this embodiment, the control means 40e functions as a generation unit 100, a sound emission processing unit 200, a display processing unit 300, and a correct answer determination unit 400 by the CPU executing a program of lip-sync content stored in memory.
[0048] (Generation part) The generator 100 generates the composite video data.
[0049] The composite video data is generated by combining video data transmitted from the server device S, in such a way that video based on one video data and video based on another video data are displayed on the display means in a predetermined layout. That is, a composite video based on the composite video data is a video in which multiple different videos are displayed simultaneously. The predetermined layout is made up of multiple areas for displaying each video. Each area is assigned an area number. For the predetermined layout, multiple templates are set in advance according to the number of videos included in the composite video.
[0050] Specifically, the generation unit 100 determines one layout depending on the number of other video data. The generation unit 100 sets an area for displaying video based on one video data and an area for displaying video based on other video data depending on the determined one layout. The areas can be set randomly. The generation unit 100 generates composite video data by combining each video data so that the video based on the one video data and the video based on the other video data are displayed in the set areas.
[0051] (Sound emission processing unit) The sound emitting processing unit 200 causes the sound emitting means to emit singing voice based on the other transmitted singing voice data in time with the karaoke performance of the selected piece of music.
[0052] The sound emitting processor 200 reads out the accompaniment data of a song selected by the user from the storage means 40a based on the song identification information of the song. The sound emitting processor 200 controls the performance means 40d to perform karaoke based on the accompaniment data, and causes the speaker 50 to emit karaoke performance sounds.
[0053] Furthermore, the sound emission processing unit 200 according to this embodiment reads out other singing voice data from the storage means 40a. The sound emission processing unit 200 causes the speaker 50 to emit singing voices based on the read singing voice data in time with the karaoke performance of the selected song. When the storage means 40a stores a plurality of other singing voice data and other video data, the sound emission processing unit 200 causes the speaker 50 to emit a plurality of singing voices based on each of the plurality of other singing voice data.
[0054] (Display processing unit) The display processing unit 300 displays a composite image based on the generated composite image data on the display means in a predetermined layout in time with the karaoke performance of the selected piece of music.
[0055] The display processing unit 300 displays a composite image based on the composite image data in a predetermined layout on the display device 60 or the like in time with the karaoke performance of one song by the performance unit 40d. When displaying the composite image, the display processing unit 300 may display the respective area numbers in each area.
[0056] (Correct answer determination section) When a user selects a certain image from the displayed composite images and singing voice data is not associated with the image data corresponding to the certain image, the correct answer determination unit 400 notifies the user that the selected image is correct.
[0057] The user selects the video that they consider to be lip-synced while viewing the composite video displayed on the display device 60. The video may be selected by operating the remote control device 80 and inputting an area number, or by directly selecting one video from the composite videos displayed on the touch panel display screen of the remote control device 80.
[0058] The correct answer determination unit 400 determines whether the video selected by the user is a lip-sync video. Specifically, the correct answer determination unit 400 refers to the storage means 40a and checks whether singing voice data is associated with the video data of the video selected by the user.
[0059] If singing voice data is not associated with the video data of the video selected by the user, the singing voice corresponding to the video is not emitted. In other words, the video is in a lip-sync state. Therefore, the correct answer determination unit 400 notifies the user that the answer is correct. On the other hand, if singing voice data is associated with the video data of the video selected by the user, the singing voice corresponding to the video is emitted. In other words, the video is not in a lip-sync state. Therefore, the correct answer determination unit 400 notifies the user that the answer is incorrect.
[0060] Notification of correct and incorrect answers can be performed in various ways. For example, the correct answer determination unit 400 displays the words "correct" or "incorrect" on the display device 60 or the remote control device 80. The correct answer determination unit 400 also causes the speaker 50 to emit a sound indicating "correct" or "incorrect." The correct answer determination unit 400 also controls the performance means 40d to stop the karaoke performance if the answer is "correct," and to continue the karaoke performance if the answer is "incorrect." If the answer is "incorrect," the user can select from the remaining videos that they believe to be lip-synced. Alternatively, the correct answer determination unit 400 can notify correct and incorrect answers by combining the above processes.
[0061] ==About the Karaoke System== Next, the processing in the karaoke system 1 according to this embodiment will be described with reference to Fig. 5 to Fig. 7B. Fig. 5 is a flowchart showing the processing in the karaoke system 1. Fig. 6 shows a composite image displayed on the display device 60. Fig. 7A shows a correct answer screen displayed on the display device 60. Fig. 7B shows an incorrect answer screen displayed on the display device 60. In this example, the information storage unit 11 stores content information C1 to Cn, each of which includes a song ID ***X of song X, singing voice data corresponding to the singing voice of a singer singing song X as karaoke, and video data corresponding to video of the singer singing song X as karaoke.
[0062] When the user operates the remote control device 80 of the karaoke device K and selects the lip-synching content, the karaoke device K starts the lip-synching content (start lip-synching content; step 10).
[0063] The user operates the remote control device 80 to select a song X to be used in the lip-sync content. The karaoke device K transmits the song ID ***X of the selected song X to the server device S (transmitting the song ID of the selected song; step 11).
[0064] The identifying unit 31 of the server device S identifies one piece of content information from the content information C1 to Cn that includes the song ID ***X of the song X (identifying one piece of content information; step 12).
[0065] The extraction unit 32 of the server device S extracts at least one other piece of content information based on the similarity between the singing voice data included in the one piece of content information identified in step 12 and the singing voice data included in other content information other than the one piece of content information, which includes the song ID ***X of song X (extracting other content information; step 13).
[0066] The transmission processing unit 33 of the server device S transmits the one video data included in the one content information identified in step 12 and the other singing voice data and other video data included in the other content information extracted in step 13 to the karaoke device K (transmitting the one video data, the other singing voice data and the other video data; step 14).
[0067] The generation unit 100 of the karaoke device K generates composite video data by combining the video data so that the video based on one video data transmitted in step 14 and the video based on the other video data transmitted in step 14 are displayed on the display device 60 in a predetermined layout (generating composite video data; step 15).
[0068] The sound emission processing unit 200 of the karaoke device K emits singing voice based on the other singing voice data transmitted in step 14 from the speaker 50 in time with the karaoke performance of the song X selected in step 11 (emits karaoke performance sounds and singing voice; step 16).
[0069] The display processing unit 300 of the karaoke device K displays the composite image based on the composite image data generated in step 15 in a predetermined layout on the display device 60 in time with the karaoke performance of the song X selected in step 11 (display the composite image; step 17).
[0070] While viewing the composite video displayed on the display device 60, the user selects the video that he or she considers to be lip-synced.
[0071] If singing voice data is not associated with the video data corresponding to the video selected by the user (N in step 18), the correct answer determination unit 400 of the karaoke device K notifies the user that the answer is correct (notification that the answer is correct; step 19). On the other hand, if singing voice data is associated with the video data corresponding to the video selected by the user (Y in step 18), the correct answer determination unit 400 of the karaoke device K notifies the user that the answer is incorrect (notification that the answer is incorrect; step 20).
[0072] Specifically, it is assumed that the server device S receives from the karaoke device K the song ID ***X of the song X selected by the user.
[0073] In this case, the identifying unit 31 identifies the content information C1 from the content information C1 to Cn that includes the song ID ***X and is stored in the information storage unit 11, based on the song ID ***X of the received song X.
[0074] The extraction unit 32 extracts other content information C2 to C4 based on the similarity between the singing voice data included in the identified content information C1 and the singing voice data included in the other content information C2 to Cn.
[0075] The transmission processing unit 33 reads out the video data MD1 included in the content information C1 identified by the identification unit 31 from the information storage unit 11. The transmission processing unit 33 also reads out the singing voice data VD2 to VD4 and the video data MD2 to MD4 included in the other content information C2 to C4 extracted by the extraction unit 32 from the information storage unit 11. The transmission processing unit 33 associates the read video data MD1 with the other singing voice data VD2 to VD4 and the other video data MD2 to MD4 and transmits them to the karaoke device K.
[0076] The karaoke device K stores the video data MD1 in association with the other singing voice data VD2 to VD4 and the other video data MD2 to MD4 in the storage means 40a.
[0077] In this example, the number of video data is four. Therefore, the generation unit 100 of the karaoke device K determines a layout that can display four videos. The generation unit 100 sets areas to display videos based on each video data according to the determined layout. In this example, the generation unit 100 displays video M1 based on video data MD1 in area No. 1, video M2 based on video data MD2 in area No. 2, video M3 based on video data MD3 in area No. 3, and video M4 based on video data MD4 in area No. 4. The generation unit 100 generates composite video data CID by combining the video data so that the videos M1 to M4 based on the video data MD1 to MD4 are displayed in the set areas.
[0078] Thereafter, the sound emission processing unit 200 reads out the accompaniment data of the song X from the storage means 40a based on the song ID ***X of the song X selected by the user. The sound emission processing unit 200 controls the performance means 40d to perform karaoke performance based on the accompaniment data, and causes the speaker 50 to emit karaoke performance sounds.
[0079] The sound emission processing unit 200 also reads out other singing voice data VD2 to VD4 from the storage means 40a. The sound emission processing unit 200 causes the speaker 50 to emit singing voices V2 to V4 based on the read singing voice data VD2 to VD4 in time with the karaoke performance of the song X.
[0080] Furthermore, the display processing unit 300 causes the display device 60 to display a composite image CI based on the composite image data CID in a predetermined layout in time with the karaoke performance of the song X by the performance means 40d. As described above, in this example, the display processing unit 300 causes the area No. 1 to display an image M1 based on the image data MD1, the area No. 2 to display an image M2 based on the image data MD2, the area No. 3 to display an image M3 based on the image data MD3, and the area No. 4 to display an image M4 based on the image data MD4 (see FIG. 6).
[0081] While viewing the composite video CI displayed on the display device 60, the user selects a video that he or she considers to be lip-synced.
[0082] For example, suppose the user selects video M1 (shown in gray in FIG. 7A). The correct answer determination unit 400 refers to the storage means 40a and checks whether singing voice data is associated with the video data MD1 of the video M1 selected by the user. In this example, singing voice data is not associated with the video data MD1. Therefore, the correct answer determination unit 400 notifies the user that the answer is correct by displaying the word "correct" on the display device 60 (see FIG. 7A).
[0083] On the other hand, suppose the user selects video M2 (shown in gray in FIG. 7B). The correct answer determination unit 400 refers to the storage means 40a and checks whether singing voice data is associated with the video data MD2 of video M2 selected by the user. In this example, singing voice data V2 is associated with the video data MD2. Therefore, the correct answer determination unit 400 notifies the user that the answer is incorrect by displaying the word "incorrect" on the display device 60 (see FIG. 7B).
[0084] As is clear from the above, the karaoke system 1 according to this embodiment includes a server device S and a karaoke device K connected to each other so as to be able to communicate with each other. The server device S includes an information storage unit 11 that stores a plurality of pieces of content information, each piece including song identification information of a song, singing voice data corresponding to a singer singing the song to karaoke, and video data corresponding to a video of the singer singing the song to karaoke, an identification unit 31 that identifies a piece of content information from the content information including song identification information of a song selected by a user of the karaoke device K, an extraction unit 32 that extracts at least one other piece of content information based on the similarity between the singing voice data included in the identified piece of content information and singing voice data included in other content information including the song identification information of the song other than the identified piece of content information, and a transmission processing unit 33 that transmits the video data included in the identified piece of content information and the other singing voice data and other video data included in the extracted other piece of content information to the karaoke device K. The karaoke device K has a generation unit 100 that generates composite video data by combining each of the video data so that a video based on one transmitted video data and a video based on the other transmitted video data are displayed on the display means in a predetermined layout; a sound emission processing unit 200 that emits a singing voice based on the other transmitted singing voice data from a sound emission means in time with the karaoke performance of the one selected song; a display processing unit 300 that displays a composite video based on the generated composite video data in a predetermined layout in time with the karaoke performance of the one selected song; and a correct answer determination unit 400 that notifies the user that the selection is correct when the user selects a video from the displayed composite videos and no singing voice data is associated with the video data corresponding to the selected video.
[0085] With this karaoke system 1, a user can listen to the karaoke performance sound of a selected song and singing voices based on other singing voice data, select a video that he or she thinks is lip-syncing from the synthesized video, and find out whether the answer is correct or incorrect, thereby enjoying lip-sync guessing content. That is, the karaoke system 1 according to this embodiment can provide new content in which the user can guess the singer who is lip-syncing from the singing voice emitted from the sound emitting means of the karaoke device and the video displayed on the display means.
[0086] Furthermore, the extraction unit 32 in the karaoke system 1 according to this embodiment can calculate a similarity between the timing of each syllable extracted from the singing voice data included in one piece of content information and the timing of each syllable extracted from the singing voice data included in another piece of content information, and extract another piece of content information for which the similarity is equal to or greater than a predetermined value. According to this karaoke system 1, the timing of each syllable can be used to extract another piece of content information.
[0087] Second Embodiment Next, the karaoke system according to this embodiment will be described with reference to Figs. 8 to 10. The karaoke system according to this embodiment can implement singer guessing content. The singer guessing content allows the user to select a video of the singer they think is actually singing karaoke from among multiple videos displayed on the display means of the karaoke device, and notifies the user whether the video is correct or incorrect. Descriptions of the same components as those in the first embodiment will be omitted.
[0088] ==Server Device== 8, the server device S according to this embodiment includes a storage unit 10, a communication unit 20, and a control unit 30. Each component is connected to a bus B via an interface (not shown).
[0089] [Storage means] The storage means 10 is a large-capacity storage device that stores various types of data. In this embodiment, a part of the storage area of the storage means 10 functions as an information storage unit 11.
[0090] (Information storage unit) The information storage unit 11 stores a plurality of pieces of content information. The content information according to this embodiment is information that may be used to implement singer guessing content in the karaoke device K. As in the first embodiment, the content information according to this embodiment includes song identification information of the song, singing voice data, and video data.
[0091] [Means of communication] The communication means 20 provides an interface for communicating with the karaoke device K.
[0092] [Control means] The control means 30 performs various controls in the server device S. The control means 30 includes a CPU and a memory (neither of which are shown). The CPU executes programs stored in the memory to realize various functions.
[0093] In this embodiment, the CPU executes a program for singer guessing content stored in the memory, causing the control means 30 to function as an identifying unit 34, an extracting unit 35, and a transmission processing unit 36 (see FIG. 8).
[0094] (Specific part) The identification unit 34 has the same configuration as the identification unit 31 of the first embodiment. That is, the identification unit 34 identifies one piece of content information from among content information including song identification information of one song selected by a user of the karaoke device K.
[0095] A user who wishes to play the singer guessing content operates the remote control device 80 of the karaoke device K and selects the singer guessing content. Then, the user operates the remote control device 80 to select one song. The karaoke device K transmits song identification information of the selected song to the server device S.
[0096] The identifying unit 34 identifies one piece of content information including the received song identification information from among the content information stored in the information storage unit 11 based on the song identification information of the received song.
[0097] (Extraction part) The extraction unit 35 has the same configuration as the extraction unit 32 of the first embodiment. That is, the extraction unit 35 extracts at least one other piece of content information based on the similarity between the singing voice data included in the identified piece of content information and the singing voice data included in the other piece of content information.
[0098] When selecting singer guessing content, the user may select the difficulty level. In this case, the extraction unit 35 can change the predetermined value depending on the selected difficulty level. For example, suppose there are three difficulty levels, "difficult," "normal," and "easy," and the user selects "difficult." In this case, the extraction unit 35 sets the predetermined value to, for example, 90%. On the other hand, suppose the user selects "easy." In this case, the extraction unit 35 sets the predetermined value to, for example, less than 70%.
[0099] (Transmission processing unit) The transmission processing unit 36 transmits to the karaoke device K one singing voice data and one video data included in the identified one piece of content information, and other video data included in the extracted other piece of content information.
[0100] The transmission processing unit 36 reads out the singing voice data and video data (i.e., one singing voice data and one video data) included in the one piece of content information identified by the identification unit 34 from the information storage unit 11. In addition, the transmission processing unit 33 reads out the video data (i.e., other video data) included in the other piece of content information extracted by the extraction unit 35 from the information storage unit 11.
[0101] The transmission processing unit 36 transmits the read one singing voice data and one video data to the karaoke device K in association with the other video data.
[0102] ==Karaoke Equipment== 9, the karaoke main unit 40 includes a storage means 40a, a communication means 40b, an input means 40c, a performance means 40d, and a control means 40e. Each component is connected to a bus B via an interface (not shown).
[0103] [Storage means] The storage means 40a according to the present embodiment stores one piece of singing voice data and one piece of video data transmitted from the server device S (transmission processing unit 36) in association with other video data.
[0104] [Control means] The control means 40e performs various controls in the karaoke device K. The control means 40e includes a CPU and a memory (neither of which is shown). The CPU executes programs stored in the memory to realize various functions.
[0105] In this embodiment, the CPU executes the program of the singer guessing content stored in the memory, and the control means 40e functions as a generation unit 500, a sound emission processing unit 600, a display processing unit 700, and a correct answer determination unit 800.
[0106] (Generation part) The generating unit 500 has the same configuration as the generating unit 100 of the first embodiment. That is, the generating unit 500 generates composite video data.
[0107] (Sound emission processing unit) The sound emitting processing unit 600 causes the sound emitting means to emit a singing voice based on the transmitted singing voice data in time with the karaoke performance of the selected piece of music.
[0108] The sound emitting processor 600 reads out the accompaniment data of a song selected by the user from the storage means 40a based on the song identification information of the song. The sound emitting processor 600 controls the performance means 40d to perform karaoke based on the accompaniment data, and causes the speaker 50 to emit karaoke performance sounds.
[0109] The sound emission processing unit 600 according to the present embodiment reads out one singing voice data from the storage means 40a. The sound emission processing unit 600 causes the speaker 50 to emit a singing voice based on the read out one singing voice data in time with the karaoke performance of the selected piece of music.
[0110] (Display processing unit) The display processing unit 700 has the same configuration as the display processing unit 300 of the first embodiment. That is, the display processing unit 700 displays a composite image based on the generated composite image data on the display means in a predetermined layout in accordance with the karaoke performance of the selected song.
[0111] (Correct answer determination section) When a user selects a certain image from the displayed composite images and singing voice data is associated with the image data corresponding to the certain image, the correct answer determination unit 800 notifies the user that the selected image is correct.
[0112] While viewing the composite images displayed on the display device 60, the user selects an image that he or she thinks represents a singer actually singing karaoke.
[0113] The correct answer determination unit 800 determines whether the video selected by the user is a video of a singer actually singing karaoke. Specifically, the correct answer determination unit 800 refers to the storage means 40a and checks whether singing voice data is associated with the video data of the video selected by the user.
[0114] If singing voice data is not associated with the video data of the video selected by the user, the singing voice corresponding to the video is not emitted. In other words, the video is in a lip-sync state. Therefore, the correct answer determination unit 800 notifies the user that the answer is incorrect. On the other hand, if singing voice data is associated with the video data of the video selected by the user, the singing voice corresponding to the video is emitted. In other words, the video shows a singer actually singing karaoke. Therefore, the correct answer determination unit 800 notifies the user that the answer is correct.
[0115] ==About the Karaoke System== Next, processing in the karaoke system 1 according to this embodiment will be described with reference to Fig. 10. Fig. 10 is a flowchart showing processing in the karaoke system 1. In this example, the information storage unit 11 stores content information C1 to Cn, each of which includes a song ID ***X of song X, singing voice data corresponding to the singing voice of a singer singing song X as karaoke, and video data corresponding to video of the singer singing song X as karaoke.
[0116] When the user operates the remote control device 80 of the karaoke device K and selects the singer guessing content, the karaoke device K starts the singer guessing content (start singer guessing content; step 30).
[0117] The user operates the remote control device 80 to select a song X to be used in the singer guessing content. The karaoke device K transmits the song ID ***X of the selected song X to the server device S (transmitting the song ID of the selected song; step 31).
[0118] The identifying unit 34 of the server device S identifies one piece of content information from the content information C1 to Cn that includes the song ID ***X of the song X (identifying one piece of content information; step 32).
[0119] The extraction unit 35 of the server device S extracts at least one other piece of content information based on the similarity between the singing voice data included in the one piece of content information identified in step 32 and the singing voice data included in other content information other than the one piece of content information, which includes the song ID ***X of song X (extracting other content information; step 33).
[0120] The transmission processing unit 36 of the server device S transmits the one singing voice data and one video data contained in the one content information identified in step 32, and the other video data contained in the other content information extracted in step 33, to the karaoke device K (transmitting the one singing voice data and one video data, and the other video data; step 34).
[0121] The generation unit 500 of the karaoke device K generates composite video data by combining the video data so that the video based on one video data transmitted in step 34 and the video based on the other video data transmitted in step 34 are displayed on the display device 60 in a predetermined layout (generating composite video data; step 35).
[0122] The sound emission processing unit 600 of the karaoke device K emits the singing voice based on the singing voice data transmitted in step 34 from the speaker 50 in time with the karaoke performance of the song X selected in step 31 (emits the karaoke performance sound and singing voice; step 36).
[0123] The display processing unit 700 of the karaoke device K displays the composite image based on the composite image data generated in step 35 in a predetermined layout on the display device 60 in time with the karaoke performance of the song X selected in step 31 (display composite image; step 37).
[0124] While viewing the composite video displayed on the display device 60, the user selects a video that he or she believes shows the singer actually singing karaoke.
[0125] If singing voice data is not associated with the video data corresponding to the video selected by the user (N in step 38), the correct answer determination unit 800 of the karaoke device K notifies the user that the answer is incorrect (notification of incorrect answer, step 39). On the other hand, if singing voice data is associated with the video data corresponding to the video selected by the user (Y in step 38), the correct answer determination unit 800 of the karaoke device K notifies the user that the answer is correct (notification of correct answer, step 40).
[0126] Specifically, it is assumed that the server device S receives from the karaoke device K the song ID ***X of the song X selected by the user.
[0127] In this case, the identifying unit 34 identifies the content information C1 from the content information C1 to Cn that includes the song ID ***X and is stored in the information storage unit 11, based on the song ID ***X of the received song X.
[0128] The extraction unit 35 extracts other content information C2 to C4 based on the similarity between the singing voice data included in the identified content information C1 and the singing voice data included in the other content information C2 to Cn.
[0129] The transmission processing unit 36 reads out the singing voice data VD1 and the video data MD1 included in the content information C1 identified by the identification unit 34 from the information storage unit 11. The transmission processing unit 36 also reads out the video data MD2 to MD4 included in the other content information C2 to C4 extracted by the extraction unit 35 from the information storage unit 11. The transmission processing unit 36 associates the read singing voice data VD1 and the video data MD1 with the other video data MD2 to MD4 and transmits them to the karaoke device K.
[0130] The karaoke device K stores the singing voice data VD1 and the video data MD1 in association with the other video data MD2 to MD4 in the storage means 40a.
[0131] The generation unit 500 sets areas to display images based on each video data according to the determined layout in which the four videos can be displayed. In this example, as in the first embodiment, it is assumed that an image M1 based on video data MD1 is displayed in area No. 1, an image M2 based on video data MD2 is displayed in area No. 2, an image M3 based on video data MD3 is displayed in area No. 3, and an image M4 based on video data MD4 is displayed in area No. 4. The generation unit 500 generates composite video data CID by combining the video data so that the images M1 to M4 based on the video data MD1 to MD4 are displayed in the set areas.
[0132] Thereafter, the sound emission processing unit 600 reads out the accompaniment data of the song X from the storage means 40a based on the song ID ***X of the song X selected by the user. The sound emission processing unit 600 controls the performance means 40d to perform karaoke performance based on the accompaniment data, and causes the speaker 50 to emit karaoke performance sounds.
[0133] The sound emission processing unit 600 also reads out one singing voice data VD1 from the storage means 40a. The sound emission processing unit 600 causes the speaker 50 to emit a singing voice V1 based on the read singing voice data VD1 in time with the karaoke performance of the song X.
[0134] Furthermore, the display processing unit 700 causes the display device 60 to display a composite image CI based on the composite image data CID in a predetermined layout in time with the karaoke performance of the song X by the performance means 40d. In this example, similar to the first embodiment, the display processing unit 700 causes an image M1 based on the image data MD1 to be displayed in area No. 1, an image M2 based on the image data MD2 to be displayed in area No. 2, an image M3 based on the image data MD3 to be displayed in area No. 3, and an image M4 based on the image data MD4 to be displayed in area No. 4.
[0135] While viewing the composite image CI displayed on the display device 60, the user selects an image that he or she thinks shows a singer actually singing karaoke.
[0136] For example, suppose the user selects video M1. The correct answer determination unit 800 refers to the storage means 40a and checks whether singing voice data is associated with the video data MD1 of the video M1 selected by the user. In this example, singing voice data V1 is associated with the video data MD1. Therefore, the correct answer determination unit 800 notifies the user that the answer is correct by displaying the word "correct" on the display device 60.
[0137] On the other hand, suppose the user selects video M2. The correct answer determination unit 800 refers to the storage means 40a and checks whether singing voice data is associated with the video data MD2 of video M2 selected by the user. In this example, singing voice data is not associated with the video data MD2. Therefore, the correct answer determination unit 800 notifies the user that the answer is incorrect by displaying the word "incorrect" on the display device 60.
[0138] As is clear from the above, the karaoke system 1 according to this embodiment includes a server device S and a karaoke device K connected to each other so as to be able to communicate with each other. The server device S includes an information storage unit 11 that stores a plurality of pieces of content information, each piece including song identification information of a song, singing voice data corresponding to a singer singing the song to karaoke, and video data corresponding to a video of the singer singing the song to karaoke, an identification unit 34 that identifies one piece of content information from the content information including song identification information of a song selected by a user of the karaoke device K, an extraction unit 35 that extracts at least one other piece of content information based on the similarity between the singing voice data included in the identified piece of content information and singing voice data included in other content information including the song identification information of the song other than the identified piece of content information, and a transmission processing unit 36 that transmits the one piece of singing voice data and one piece of video data included in the identified piece of content information and the other video data included in the extracted other piece of content information to the karaoke device K. The karaoke device K has a generation unit 500 that generates composite video data by combining each of the video data so that a video based on one transmitted video data and a video based on another transmitted video data are displayed on the display means in a predetermined layout; a sound emission processing unit 600 that emits a singing voice based on the transmitted singing voice data from a sound emission means in time with the karaoke performance of the selected song; a display processing unit 700 that displays a composite video based on the generated composite video data in a predetermined layout in time with the karaoke performance of the selected song on the display means; and a correct answer determination unit 800 that notifies the user that the selection is correct when the user selects a certain video from the displayed composite videos and when singing voice data is associated with the video data corresponding to the certain video.
[0139] With this karaoke system 1, a user can listen to the karaoke performance sound of a selected song and singing voices based on other singing voice data, select a video of the singer they think is actually singing karaoke from the composite video, and find out whether their answer is correct or incorrect, allowing them to enjoy singer guessing content. That is, the karaoke system 1 according to this embodiment can provide new content in which the user can guess the singer singing karaoke from the singing voice emitted from the sound emitting means of the karaoke device and the video displayed on the display means.
[0140] <Other> The program can also be supplied to a computer using a non-transitory computer-readable medium with an executable program stored thereon. Examples of non-transitory computer-readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), CD-ROMs (Read Only Memory), etc.
[0141] The above-described embodiments are presented as examples and do not limit the scope of the invention. The above configurations can be implemented in appropriate combinations, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. The above-described embodiments and their modifications are included in the scope and spirit of the invention, as well as in the inventions described in the claims and their equivalents. [Explanation of symbols]
[0142] 1 Karaoke system 31, 34 Specific part 32, 35 Extraction part 33, 36 Transmission processing unit 100, 500 generation section 200, 600 Sound emission processing section 300, 700 Display processing unit 400, 800 Correct answer determination section K Karaoke equipment S Server device
Claims
1. A karaoke system in which a server device and a karaoke device are communicably connected, The server device an information storage unit that stores a plurality of pieces of content information, including song identification information of songs, singing voice data corresponding to the singing voice of a singer singing the songs as karaoke, and video data corresponding to a video of the singer singing the songs as karaoke; an identifying unit that identifies one piece of content information from the content information including song identification information of one song selected by a user of the karaoke device; an extracting unit that extracts at least one other piece of content information based on a similarity between singing voice data included in the identified piece of content information and singing voice data included in content information other than the identified piece of content information, the singing voice data being content information including song identification information of the identified piece of song; a transmission processing unit that transmits the one video data included in the identified one content information and the other singing voice data and other video data included in the extracted other content information to the karaoke device; and The karaoke device a generating unit that generates composite video data by combining the video data so that a video based on the one video data and a video based on the other video data are displayed on a display device in a predetermined layout; a sound emitting processing unit for emitting a singing voice based on the transmitted other singing voice data from a sound emitting means in accordance with the karaoke performance of the selected one piece of music; a display processing unit that displays a composite image based on the generated composite image data on the display means in the predetermined layout in sync with the karaoke performance of the selected piece of music; a correct answer determination unit that notifies the user that the selected image is correct when the user selects a certain image from the displayed composite images and when singing voice data is not associated with video data corresponding to the certain image; A karaoke system having:
2. A karaoke system in which a server device and a karaoke device are communicably connected, The server device an information storage unit that stores a plurality of pieces of content information, including song identification information of songs, singing voice data corresponding to the singing voice of a singer singing the songs as karaoke, and video data corresponding to a video of the singer singing the songs as karaoke; an identifying unit that identifies one piece of content information from the content information including song identification information of one song selected by a user of the karaoke device; an extracting unit that extracts at least one other piece of content information based on a similarity between singing voice data included in the identified piece of content information and singing voice data included in content information other than the identified piece of content information, the singing voice data being content information including song identification information of the identified piece of song; a transmission processing unit that transmits the one singing voice data and one video data included in the identified one piece of content information and the other video data included in the extracted other piece of content information to the karaoke device; and The karaoke device a generating unit that generates composite video data by combining the video data so that a video based on the one video data and a video based on the other video data are displayed on a display device in a predetermined layout; a sound emitting processing unit for emitting a singing voice based on the transmitted singing voice data from a sound emitting means in accordance with the karaoke performance of the selected piece of music; a display processing unit that displays a composite image based on the generated composite image data on the display means in the predetermined layout in sync with the karaoke performance of the selected piece of music; a correct answer determination unit that notifies the user that the selected image is correct when the user selects a certain image from the displayed composite images and when singing voice data is associated with video data corresponding to the certain image; A karaoke system having:
3. The karaoke system of claim 1 or 2, characterized in that the extraction unit calculates a similarity based on the pronunciation timing of each syllable extracted from the singing voice data included in the one content information and the pronunciation timing of each syllable extracted from the singing voice data included in the other content information, and extracts other content information for which the similarity is equal to or greater than a predetermined value.
Citation Information
Patent Citations
Karaoke system, central device and content data creation method
JP2011059619A