Karaoke equipment

JP7866433B2Active Publication Date: 2026-05-27DAIICHI KOSHO COMPANY

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
DAIICHI KOSHO COMPANY
Filing Date
2022-05-27
Publication Date
2026-05-27

AI Technical Summary

Technical Problem

Existing karaoke devices lack the capability to generate singing video data that replicates the live performance of a professional singer, lacking the sense of presence and power found in general-purpose background videos.

Method used

A karaoke device equipped with a live video data processing unit that extracts angle and zoom information from live video data, a video data acquisition unit that acquires multiple video data from different shooting angles, and a singing video data generation unit that combines these data in chronological order to create a singing video.

Benefits of technology

Enables the generation of singing video data that resembles a live performance of the original singer, providing a more engaging and immersive karaoke experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007866433000001
    Figure 0007866433000001
  • Figure 0007866433000002
    Figure 0007866433000002
  • Figure 0007866433000003
    Figure 0007866433000003
Patent Text Reader

Abstract

To provide a Karaoke device that can generate singing moving image data for displaying a singing moving image like a live video of a singer of an original music piece.SOLUTION: The Karaoke device includes: a live video data processing unit that extracts angle information indicating a shooting angle from live video data on a singer of an original music piece associated with a music piece selected by a singing person and associates the angle information with a performance time of the music piece corresponding to a frame from which the angle information has been extracted, thereby generating shooting information; a video data acquisition unit that based on video signals transmitted from a plurality of shooting means set at different shooting angles, acquires a plurality of pieces of video data; and a singing moving image data generation unit that based on the generated shooting information, generates singing moving image data in which some of the plurality of pieces of acquired video data are combined in a time-series manner.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a karaoke device.

Background Art

[0002] There is a known technique for photographing a user performing karaoke singing and generating a singing video. For example, Patent Document 1 discloses a technique for recording the singing posture of a singer together with singing voices using a karaoke device connected to a video camera.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When performing a karaoke performance of a music piece, the karaoke device can display a live video by a professional singer (hereinafter referred to as "original song singer") who is actually singing the music piece as a background video for the karaoke performance. Since the live video is a combination of videos shot from different angles at the live venue, it has a sense of presence and power that cannot be found in general-purpose background videos, and can be enjoyed.

[0005] An object of the present invention is to provide a karaoke device capable of generating singing video data for displaying a singing video such as a live video of an original song singer.

Means for Solving the Problems

[0006] One invention to achieve the above objective is a karaoke device comprising: a live video data processing unit that extracts angle information indicating the shooting angle from live video data of the original singer associated with a song selected by the singer, and generates shooting information by associating the angle information with the performance time of the song corresponding to the frame from which the angle information was extracted; a video data acquisition unit that acquires multiple video data based on video signals transmitted from multiple shooting means set to different shooting angles; and a singing video data generation unit that generates singing video data by combining a portion of the acquired multiple video data in chronological order based on the generated shooting information. Other features of the present invention will be revealed in the specification and drawings described below. [Effects of the Invention]

[0007] According to the present invention, it is possible to generate singing video data for displaying singing videos, such as live performance videos of the original singer. [Brief explanation of the drawing]

[0008] [Figure 1] This figure shows a karaoke apparatus and a shooting means according to an embodiment. [Figure 2] This is a diagram showing a karaoke device according to an embodiment. [Figure 3] This is a diagram showing a karaoke machine according to an embodiment. [Figure 4] This is a flowchart showing the processing of the karaoke apparatus according to the embodiment. [Figure 5] This figure shows the shooting information according to the embodiment. [Figure 6] This diagram shows the structure of the singing video data according to the embodiment. [Figure 7] This figure shows the photographic information related to the modified example 1. [Figure 8] This diagram shows the structure of the singing video data related to Modification Example 1. [Figure 9] This figure shows the photographic information related to the modified example 2. [Figure 10]This figure shows the structure of the singing video data related to the modified example 2. [Figure 11] This is a diagram showing a karaoke machine according to modified example 3. [Modes for carrying out the invention]

[0009] <Embodiment> A karaoke apparatus according to an embodiment will be described with reference to Figures 1 to 6.

[0010] ==Karaoke equipment and filming means== A karaoke machine is a device that provides karaoke playback of songs selected by the singer, and allows the singer to perform karaoke. A karaoke machine is typically installed in a karaoke room in a karaoke establishment. A singer is a user of the karaoke machine who performs karaoke. In the example in Figure 1, karaoke machine K is installed in karaoke room KR. User U uses karaoke machine K and performs karaoke on the stage. In other words, user U is the singer.

[0011] In this embodiment, a karaoke room is equipped with multiple shooting devices. The shooting devices are connected to the karaoke machine in a communicative manner. The shooting devices are devices such as cameras that are used to photograph the stage where the singer is located or the inside of the karaoke room KR. Each shooting device is installed to photograph from a different shooting angle. In the example in Figure 1, the karaoke room KR is equipped with three shooting devices SE1 to SE3. Each of the shooting devices SE1 to SE3 is connected to the karaoke machine K in a communicative manner. Shooting device SE1 is installed to photograph from a shooting angle facing the front of the stage. Shooting device SE2 is installed to photograph from a shooting angle facing the left side of the stage. Shooting device SE3 is installed to photograph from a shooting angle facing the right side of the stage.

[0012] Each imaging means starts imaging in synchronization with the karaoke performance based on an instruction from the karaoke device, and sequentially transmits the obtained video signal to the karaoke device. In the example of FIG. 1, the imaging means SE1 to SE3 start imaging in synchronization with the karaoke performance based on an instruction from the karaoke device K, and sequentially transmit the obtained video signal to the karaoke device K.

[0013] ==Karaoke Device== As shown in FIG. 2, the karaoke device K includes a karaoke main body 10, a speaker 20, a display device 30, a microphone 40, and a remote control device 50.

[0014] The karaoke main body 10 performs various controls related to karaoke performance and karaoke singing, such as karaoke performance control of the selected music, display control of lyrics, background images, etc., and processing of the audio signal input through the microphone 40. The speaker 20 is a configuration for emitting sound based on the sound emission signal from the karaoke main body 10. The display device 30 is a configuration for displaying video and images on the screen based on the signal from the karaoke main body 10. The microphone 40 is a configuration for converting the singing voice of the singer's karaoke singing into an analog audio signal and inputting it to the karaoke main body 10. The remote control device 50 is a device for performing various operations on the karaoke main body 10.

[0015] As shown in FIG. 3, the karaoke main body 10 according to the present embodiment includes a storage means 10a, a communication means 10b, an input means 10c, a performance means 10d, and a control means 10e. Each component is connected to the bus B via an interface (not shown).

[0016] [Storage Means] The storage means 10a is a large-capacity storage device that stores various types of data. The storage means 10a stores music data. The music data is provided with music identification information. The music identification information is information unique to each music, such as a music ID for identifying the music. The music data includes karaoke performance data, reference data, section information, etc. The karaoke performance data is MIDI-formatted data that serves as the source of karaoke performance sounds. The reference data is data indicating the main melody of the music performed in karaoke. The section information indicates the performance section. The performance section is the section in which karaoke performance is carried out. The performance section includes a singing section and a non-singing section. The singing section is a section (for example, the A melody, B melody, and refrain of No. 1) in which the lyrics to be sung in a certain music are set. The non-singing section is a section in which no lyrics to be sung in a certain music are set, such as an intro, an interlude, or an outro.

[0017] In addition, the storage means 20 stores, for each music, background video data corresponding to the background video displayed during karaoke performance, lyric telop data for displaying the lyrics of the music, and attribute information of the music (music name, singer name, genre, etc.).

[0018] In the present embodiment, the background video data includes live video data. The live video data is data corresponding to a live video of the original song singer's live performance. The live video is a combination of videos shot from different angles in chronological order. The live video data is associated with the elapsed time from the start of the music performance (i.e., the music performance time) for each frame constituting the live video.

[0019] [Communication means · Input means] The communication means 10b provides an interface for communicating with the remote control device 50. The input means 10c is a configuration for the user to input various instructions. The input means 10c is buttons or the like provided on the karaoke machine 10. Alternatively, the remote control device 50 may function as the input means 10c.

[0020] [Performance means] The performance means 10d performs karaoke performance of a song and processes signals based on singing voice input through the microphone 40, based on the control of the control means 10e. The performance means 10d includes a sound source, mixer, amplifier, etc. (none of which are shown in the figure).

[0021] [Control means] The control means 10e performs various controls on the karaoke machine K. The control means 10e includes a CPU and memory (neither of which are shown in the figure). The CPU realizes various functions by executing programs stored in the memory.

[0022] In this embodiment, the CPU executes a program stored in memory, and the control means 10e functions as a live video data processing unit 100, a video data acquisition unit 200, and a singing video data generation unit 300.

[0023] (Live video data processing unit) The live video data processing unit 100 extracts angle information indicating the shooting angle from the live video data of the original singer associated with the song selected by the singer, and generates shooting information by associating the angle information with the performance time of the song corresponding to the frame from which it was extracted.

[0024] Angle information is information that indicates the shooting angle of the image for each frame that makes up the live video. Angle information can be extracted by various methods. For example, the live video data processing unit 100 can extract angle information using a pre-built learning model.

[0025] The trained model consists of training data that links the images used to analyze the orientation of people and objects with the shooting angle of those images. Known techniques can be used to analyze the orientation of people and objects in the images.

[0026] In the following explanation, the left and right of a person's face are defined as the left and right from the person's perspective, and the left and right of the shooting angle and stage are defined as the left and right from the perspective facing the stage, as described above. For example, suppose karaoke machine K analyzes an image and determines that "the front of a person's face is visible" or "the front of the stage is visible." In this case, karaoke machine K associates the shooting angle "front" with that image and trains it. Similarly, suppose karaoke machine K analyzes an image and determines that "mainly the left side of a person's face is visible" or "the right side of the stage is in the foreground." In this case, karaoke machine K associates the shooting angle "right" with that image and trains it. Also, suppose karaoke machine K analyzes an image and determines that "mainly the right side of a person's face is visible" or "the left side of the stage is in the foreground." In this case, karaoke machine K associates the shooting angle "left" with that image and trains it. Karaoke machine K builds a trained model by repeatedly performing these processes. If multiple people are in the image, the karaoke machine K will perform facial image analysis on one person at a time. For example, the karaoke machine K will compare the size of each person in the image and perform facial image analysis on the person who appears largest. The construction of the trained model may also be performed on an external server device (not shown) separate from the karaoke machine K. In this case, the external server device can provide the trained model to the karaoke machine K.

[0027] Now, let's assume the singer has selected a song by operating the remote control device 50. Also, let's assume the singer has selected "Singing Video Mode" to generate singing video data by operating the remote control device 50.

[0028] In this case, the live video data processing unit 100 reads live video data of the original singer associated with a song selected by the singer from the storage means 10a.

[0029] The live video data processing unit 100 inputs images of frames that make up the live video into a trained model. The trained model outputs the shooting angles associated with images that match or are similar to the input images. The live video data processing unit 100 extracts the shooting angles output from the trained model as angle information.

[0030] Furthermore, the live video data processing unit 100 identifies the performance time of the song associated with the frame corresponding to the image input to the trained model. The live video data processing unit 100 generates shooting information by associating the extracted angle information with the performance time of the identified song.

[0031] Furthermore, angle information may be pre-associated with the live video data as metadata. In this case, the live video data processing unit 100 can extract the angle information from the metadata.

[0032] Furthermore, there may be cases where live video data is not associated with the song selected by user U. In this case, the live video data processing unit 100 displays a message to that effect on the display screen of the remote control device 50 or the like (for example, "A singing video cannot be created for the selected song because there is no live video data. Please select a different song."). In this case, the karaoke device K does not generate singing video data.

[0033] (Video data acquisition unit) The video data acquisition unit 200 acquires multiple video data based on video signals transmitted from multiple shooting devices set to different shooting angles.

[0034] Each shooting device starts shooting in synchronization with the karaoke performance based on instructions from the karaoke machine K. Each shooting device transmits the video signal obtained from shooting the stage to the karaoke machine K. The video data acquisition unit 200 acquires multiple video data based on the received video signal.

[0035] (Singing video data generation unit) The singing video data generation unit 300 generates singing video data by combining parts of multiple acquired video data in chronological order based on the generated shooting information.

[0036] The singing video data generation unit 300 selects a video data corresponding to footage shot at a shooting angle that matches the angle information, based on the angle information and performance time included in the generated shooting information, as video data from one performance time to the next. The singing video data generation unit 300 associates the performance time as the playback time for the selected video data.

[0037] The singing video data generation unit 300 generates singing video data by combining multiple selected video data in chronological order, by performing the same processing for each performance time included in the shooting information.

[0038] The singing video data generation unit 300 can play back the generated singing video data and display the singing video on the display device 30.

[0039] ==Regarding the operation of karaoke machine K== Next, a specific example of the operation of the karaoke device K in this embodiment will be described with reference to Figures 4 to 6. Figure 4 is a flowchart showing an example of the operation of the karaoke device K. Figure 5 is a diagram showing an example of shooting information. Figure 6 is a diagram showing an example of singing video data. In this example, as shown in Figure 1, it is assumed that user U uses the karaoke device K to sing karaoke. It is also assumed that three shooting means SE1 to SE3 are installed in the karaoke room KR.

[0040] User U operates the remote control device 50 to select song X. User U also operates the remote control device 50 to select "Singing Video Mode" (Selection of Singing Video Mode and song selection. Step 10).

[0041] The live video data processing unit 100 extracts angle information indicating the shooting angle from the live video data of the original singer associated with the song X selected by the user U, and generates shooting information by associating the angle information with the performance time of the song corresponding to the frame from which it was extracted (generate shooting information; step 11).

[0042] Specifically, the live video data processing unit 100 reads live video data of the original singer associated with the song X selected by the user U from the storage means 10a.

[0043] The live video data processing unit 100 detects scene changes from the read live video data. For detecting scene changes, known techniques such as "Analysis of Long-Term Soccer Broadcast Video for Scene Search Systems" (Information Processing Society of Japan CVIM Workshop Report (2004-CVIM-144, pp.125-132) can be used.

[0044] The live video data processing unit 100 identifies the frame corresponding to the beginning of a scene from among the frames that make up the live video, based on the detected scene transition. The live video data processing unit 100 inputs the image of the identified frame into a trained model. The trained model outputs the shooting angle associated with images that match or are similar to the input image. The live video data processing unit 100 extracts the shooting angle output from the trained model as angle information. The reason for identifying the frame corresponding to the beginning of a scene and extracting angle information is that, in the case of live video, the shooting angle and zoom state usually do not change within a single scene.

[0045] Furthermore, the live video data processing unit 100 identifies the performance time of the song associated with the frame corresponding to the image input to the trained model. The live video data processing unit 100 associates the extracted angle information with the identified performance time of the song.

[0046] The live video data processing unit 100 generates shooting information by performing the same process on all images of the first frame of each detected scene. The live video data processing unit 100 stores the generated shooting information in the storage means 10a.

[0047] Here, according to the shooting information shown in Figure 5, for example, the angle information at the performance time "00:00:00" is "front." This indicates that the image in the frame associated with that performance time is a frontal image. On the other hand, the angle information at the performance time "00:00:12" is "left." This indicates that the image in the frame associated with that performance time is a left image.

[0048] Subsequently, the karaoke device K begins playing song X (starts karaoke performance of song. Step 12). User U sings song X using microphone 40.

[0049] The video data acquisition unit 200 acquires multiple video data based on video signals transmitted from multiple shooting means SE1 to SE3 that are set to different shooting angles (acquisition of multiple video data; step 13).

[0050] Specifically, the video data acquisition unit 200 instructs the shooting means SE1 to SE3 to start shooting in synchronization with the karaoke performance when the karaoke performance of song X begins. In other words, the elapsed time from the start of shooting at a given point in time coincides with the elapsed time from the start of the karaoke performance of song X.

[0051] Based on instructions from the karaoke machine K, the shooting means SE1 starts shooting at a shooting angle from the front of the stage. Shooting means SE1 transmits the obtained video signal (the video signal corresponding to the "front" shooting angle) to the karaoke machine K. Similarly, based on instructions from the karaoke machine K, shooting means SE2 starts shooting at a shooting angle from the left side of the stage. Shooting means SE2 transmits the obtained video signal (the video signal corresponding to the "left" shooting angle) to the karaoke machine K. Shooting means SE3 starts shooting at a shooting angle from the right side of the stage, based on instructions from the karaoke machine K. Shooting means SE3 transmits the obtained video signal (the video signal corresponding to the "right" shooting angle) to the karaoke machine K.

[0052] The video data acquisition unit 200 receives video signal VS1 from the shooting means SE1, video signal VS2 from the shooting means SE2, and video signal VS3 from the shooting means SE3, in synchronization with the karaoke performance of song X. The video data acquisition unit 200 acquires video data VD1 to VD3 based on video signals VS1 to VS3. The video data acquisition unit 200 stores the acquired video data VD1 to VD3 in the storage means 10a. The video data acquisition unit 200 performs the same process until the karaoke performance of song X ends (if the result is Y in step 14).

[0053] After the karaoke performance of song X is finished, the singing video data generation unit 300 generates singing video data by combining parts of the multiple video data acquired in step 13 in chronological order, based on the shooting information generated in step 11 (generate singing video data; step 15). The singing video data generation unit 300 plays the singing video data generated in step 15 and displays the singing video on the display device 30 (play singing video data; step 16).

[0054] Specifically, the storage means 10a stores the shooting information shown in Figure 5. Furthermore, the storage means 10a stores video data VD1 to VD3 based on the video signals received from the shooting means SE1 to SE3.

[0055] The singing video data generation unit 300 selects video data VD1 from video data VD1 to VD3 as video data for the period from the performance time "00:00:00" to the performance time "00:00:08" (excluding the performance time "00:00:08"; the same applies hereafter), corresponding to the video shot at a shooting angle that matches the angle information "front" shown in Figure 5. The singing video data generation unit 300 associates the playback time "00:00:00 to 00:00:08" with the selected video data VD1. Similarly, the singing video data generation unit 300 selects video data VD1 from video data VD1 to VD3 as video data for the period from the performance time "00:00:08" to the performance time "00:00:12" shown in Figure 5, corresponding to the video shot at a shooting angle that matches the angle information "front" shown in Figure 5. The singing video data generation unit 300 associates the playback time "00:00:08~00:00:12" with the selected video data VD1. The singing video data generation unit 300 also selects video data VD2 from video data VD1 to VD3 as video data for the performance time from "00:00:12" to "00:00:14", which corresponds to the video shot at a shooting angle that matches the angle information "left". The singing video data generation unit 300 associates the playback time "00:00:12~00:00:14" with the selected video data VD2.

[0056] By repeatedly performing this process, the singing video data generation unit 300 can generate the singing video data shown in Figure 6. For example, in the singing video data shown in Figure 6, video data VD1 is associated with the playback time "00:00:08~00:00:12", and video data VD2 is associated with the playback time "00:00:12~00:00:14". This indicates that when playing the singing video data, from the playback time "00:00:08", video data VD1 corresponding to the video shot at the front shooting angle is played, while from the playback time "00:00:12", video data VD2 corresponding to the video shot at the left shooting angle is played.

[0057] Furthermore, in cases where singing video data is played after the karaoke performance has finished, as in this example, it is preferable for the karaoke device K to emit the karaoke performance sound of the song and the singer's voice, recorded in sync with the karaoke performance, from the speaker 20 in conjunction with the playback of the singing video data. In addition, the karaoke device K may upload the generated singing video data to an external server device and make it viewable not only to user U but also to others.

[0058] Furthermore, the singing video data generation unit 300 can sequentially generate singing video data by performing the above processing in accordance with the karaoke performance. In this case, the singing video data generation unit 300 can display the singing video on another display device while displaying the live video on the display device 30.

[0059] Furthermore, in live video footage, during non-singing sections of a song, the audience or other objects may be shown instead of the original singer. Therefore, the singing video data generation unit 300 may select predetermined video data without using the shooting information during non-singing sections. The predetermined video data may be, for example, data based on a video signal transmitted from one of the shooting means, or data pre-stored in the storage means 10a.

[0060] Furthermore, the above example describes a case where the type of angle information extracted from live video data matches the shooting angles of multiple shooting devices, but it is not limited to this. For example, suppose in the above example, the types of angle information extracted from live video data are "front," "left," and "right," while there are two shooting devices installed in the karaoke room KR (shooting device SE2 and shooting device SE3). In this case, there is no video data corresponding to a video shot at a shooting angle that matches the extracted angle information "front." Therefore, the singing video data generation unit 300 selects video data corresponding to either the video signal transmitted from shooting device SE2 or shooting device SE3.

[0061] As is clear from the above, the karaoke device K according to this embodiment includes: a live video data processing unit 100 that generates shooting information by extracting angle information indicating the shooting angle from live video data of the original singer associated with the song selected by the singer, and associating the angle information with the performance time of the song corresponding to the frame from which the angle information was extracted; a video data acquisition unit 200 that acquires multiple video data based on video signals transmitted from multiple shooting means set to different shooting angles; and a singing video data generation unit 300 that generates singing video data by combining a portion of the acquired video data in chronological order based on the generated shooting information.

[0062] According to this karaoke device K, video signals transmitted from multiple shooting means with different shooting angles are acquired as video data, and from this video data, video data shot at a shooting angle that matches the angle information extracted from live video data is selected, and singing video data is generated by combining the selected video data in chronological order. The singing video based on the generated singing video data is displayed with the same shooting angle as the live video. In other words, according to the karaoke device K of this embodiment, singing video data can be generated to display a singing video that resembles a live video of the original singer.

[0063] <Example 1> The shooting information may include information other than angle information and the duration of the music performance. The live video data processing unit 100 in this modified example can generate shooting information that includes zoom information. The zoom information indicates the zoom state in the frame from which the angle information was extracted. Several zoom states are pre-set, such as "in," "slightly in," "slightly out," and "out."

[0064] Specifically, the live video data processing unit 100 extracts angle information and identifies the duration of the music performance by performing the same processing as in the embodiment.

[0065] Here, the live video data processing unit 100 in this modified version calculates the proportion of the original singer's face in the entire image by analyzing the image of the frame used to extract angle information using known techniques. The live video data processing unit 100 identifies the zoom state by applying the calculated proportion to a predetermined standard. The live video data processing unit 100 determines the identified zoom state as zoom information.

[0066] The predetermined criteria are set in advance, for example, "If the ratio is R1 (40%) or higher: zoom is in", "If the ratio is R2 (20%) or higher but less than R1: zoom is slightly in", "If the ratio is R3 (10%) or higher but less than R2: zoom is slightly out", and "If the ratio is less than R3: zoom is out".

[0067] The live video data processing unit 100 associates the extracted angle information, the performance time of the identified song, and the determined zoom information.

[0068] The live video data processing unit 100 generates shooting information by performing the same process on all images of the first frame of each detected scene. The live video data processing unit 100 stores the generated shooting information in the storage means 10a.

[0069] Figure 7 shows an example of the shooting information related to this modified example. According to the shooting information shown in Figure 7, for example, the angle information at the performance time "00:00:00" is "front" and the zoom information is "out". This indicates that the image of the frame associated with that performance time is a front-facing image with a zoom level of "out". On the other hand, the angle information at the performance time "00:00:12" is "left" and the zoom information is "in". This indicates that the image of the frame associated with that performance time is a left-facing image with a zoom level of "in".

[0070] Furthermore, the singing video data generation unit 300 in this modified version generates singing video data using video data that has undergone digital zoom processing based on zoom information.

[0071] Digital zoom processing is a process of trimming the original video data. In this modified example, digital zoom processing is performed based on the zoom state indicated by the zoom information included in the shooting information. It is desirable to identify the area containing the singer's face image using known techniques and perform the digital zoom processing on that area. In this modified example, in response to the determination of zoom information by the live video data processing unit 100, if the "zoom state is in," the image is trimmed so that the proportion of the singer's face in the entire image is R1 (40%) or more; if the "zoom state is slightly in," the image is trimmed so that it is R2 (20%) or more and less than R1; if the "zoom state is slightly out," the image is trimmed so that it is R3 (10%) or more and less than R2; and if the "zoom state is out," the image is trimmed so that it is less than R3. On the other hand, if the "zoom state is out," the video data may be used as is without performing digital zoom processing.

[0072] Specifically, the storage means 10a stores the shooting information shown in Figure 7. Also, similar to the embodiment, the storage means 10a stores video data VD1 to VD3 based on the video signals received from the shooting means SE1 to SE3. The shooting means SE1 to SE3 are installed to be able to photograph the entire stage in order to support all zoom states: "in," "slightly in," "slightly out," and "out."

[0073] The singing video data generation unit 300 selects video data VD1 from among video data VD1 to VD3 as video data for the performance time from "00:00:00" to "00:00:08" as shown in Figure 7, and selects video data VD1 that corresponds to the video shot at a shooting angle that matches the angle information "front" shown in Figure 7. The singing video data generation unit 300 performs digital zoom processing on the selected video data VD1 based on the zoom information "out" shown in Figure 7 to generate video data ZD1-1. The singing video data generation unit 300 associates the playback time "00:00:00 to 00:00:08" with the generated video data ZD1-1. Similarly, the singing video data generation unit 300 selects video data VD1 from among video data VD1 to VD3 as video data for the performance time from "00:00:08" to "00:00:12" shown in Figure 7, and selects video data VD1 that corresponds to the video shot at a shooting angle that matches the angle information "front" shown in Figure 7. The singing video data generation unit 300 performs digital zoom processing on the selected video data VD1 based on the zoom information "in" shown in Figure 7 to generate video data ZD1-2. The singing video data generation unit 300 associates the playback time "00:00:08 to 00:00:12" with the generated video data ZD1-2. Furthermore, the singing video data generation unit 300 selects video data VD2 from among video data VD1 to VD3 as video data for the performance time from "00:00:12" to "00:00:14", corresponding to the video shot at a shooting angle that matches the angle information "left". The singing video data generation unit 300 performs digital zoom processing on the selected video data VD2 based on the zoom information "in" shown in Figure 7, and generates video data ZD2-3. The singing video data generation unit 300 associates the playback time "00:00:12 to 00:00:14" with the generated video data ZD2-3.

[0074] By repeatedly performing this process, the singing video data generation unit 300 can generate the singing video data shown in Figure 8. For example, in the singing video data shown in Figure 8, video data ZD1-2 is associated with the playback time "00:00:08~00:00:12", and video data ZD2-3 is associated with the playback time "00:00:12~00:00:14". This indicates that when playing the singing video data, from the playback time "00:00:08", video data ZD1-2, which corresponds to the video shot from a front shooting angle and digitally zoomed in to the "in" state, is played, while from the playback time "00:00:12", video data ZD2-3, which corresponds to the video shot from a left shooting angle and digitally zoomed in to the "in" state, is played.

[0075] As is clear from the above, in the karaoke device K according to this modified example, the live video data processing unit 100 generates shooting information including zoom information indicating the zoom state in the frame from which angle information has been extracted. The singing video data generation unit 300 generates singing video data using video data that has undergone digital zoom processing based on the zoom information.

[0076] According to this karaoke device K, it is possible to select video data shot at a shooting angle that matches the angle information extracted from live video data, perform digital zoom processing on the selected video data, and then generate singing video data by combining them in chronological order. The singing video based on the generated singing video data is displayed with the same shooting angle and zoom state as the live video. In other words, according to the karaoke device K according to this modified example, it is possible to generate singing video data for displaying a singing video that reflects the shooting angle and zoom state of the live video of the original singer.

[0077] <Modification 2> The shooting information may also include person information. The live video data processing unit 100 in this modified example can generate shooting information that includes person information. The person information is the person (for example, the original singer, a professional guitarist, bassist, drummer, or keyboardist) who appears in the image of the frame from which the angle information has been extracted.

[0078] Specifically, the live video data processing unit 100 performs the same processing as in the embodiment and modified example 1 to extract angle information, identify the duration of the music performance, and determine zoom information.

[0079] In this modified version, the live video data processing unit 100 identifies a person located at the center of the image by analyzing the image of the frame used when extracting angle information and determining zoom information using known techniques. The live video data processing unit 100 then identifies the identified person as person information.

[0080] For example, suppose the live video data processing unit 100 analyzes the image and determines that the person in the center of the image is holding a microphone. In this case, the live video data processing unit 100 determines that the person is the "original singer." Similarly, if the live video data processing unit 100 analyzes the image and determines that the person in the center of the image is holding a guitar, it determines that the person is the "guitarist." If the live video data processing unit 100 analyzes the image and determines that the person in the center of the image is surrounded by a drum set, it determines that the person is the "drummer."

[0081] The live video data processing unit 100 associates the extracted angle information, the performance time of the identified song, zoom information, and person information.

[0082] The live video data processing unit 100 generates shooting information by performing the same process on all images of the first frame of each detected scene. The live video data processing unit 100 stores the generated shooting information in the storage means 10a.

[0083] Figure 9 shows an example of the shooting information related to this modified example. According to the shooting information shown in Figure 9, for example, at the performance time "00:00:00", the angle information is "front", the zoom information is "out", and the person information is "original singer". This indicates that the image of the frame associated with that performance time is a frontal shot, the zoom is "out", and the image is centered on the original singer. On the other hand, at the performance time "00:00:12", the angle information is "left", the zoom information is "in", and the person information is "guitarist". This indicates that the image of the frame associated with that performance time is a left shot, the zoom is "in", and the image is centered on the guitarist.

[0084] Furthermore, in this modified example, the video corresponding to the video signals transmitted from multiple shooting means is obtained by filming the singer and the instrumentalists. The instrumentalists are those who play their instruments live in time with the karaoke track when the singer performs karaoke. The instruments include guitars, basses, drums, keyboards, etc. For example, in the embodiment, when user U performs karaoke of song X, the guitarist, bassist, drummer, and keyboardist each play their instruments live in time with the karaoke track.

[0085] It is assumed that the singer and instrumentalists are within the shooting range of the shooting device (for example, on the stage in Figure 1). Therefore, the video corresponding to the video signal transmitted from the shooting device will show the singer and instrumentalists. The video data acquisition unit 200 can acquire multiple video data based on the video signals transmitted from multiple shooting devices, similar to the embodiment.

[0086] Furthermore, the singing video data generation unit 300 related to this modification generates singing video data using video data that has undergone digital zoom processing based on person information and zoom information.

[0087] Specifically, the storage means 10a stores the shooting information shown in Figure 9. Furthermore, the storage means 10a stores video data VD'1 to VD'3 based on the video signals received from the shooting means SE1 to SE3. The video data VD'1 to VD'3 includes not only the singer but also the musicians.

[0088] The singing video data generation unit 300 selects video data VD'1 to VD'3 from among the video data VD'1 to VD'3 as video data for the performance time from "00:00:00" to "00:00:08" shown in Figure 9, and selects video data VD'1 that corresponds to the video shot at a shooting angle that matches the angle information "front" shown in Figure 9. The singing video data generation unit 300 then performs a digital zoom process on the selected video data VD'1, centering on the singer (the person corresponding to the original singer) as the focus, based on the person information "original singer" and zoom information "out" shown in Figure 9, and generates video data ZD'1-1. The singing video data generation unit 300 associates the playback time "00:00:00 to 00:00:08" with the generated video data ZD'1-1. Similarly, the singing video data generation unit 300 selects video data VD'1 to VD'3 from among the video data VD'1 to VD'3 as video data for the performance time from "00:00:08" to "00:00:12" shown in Figure 9, and selects video data VD'1 that corresponds to the video shot at a shooting angle that matches the angle information "front" shown in Figure 9. The singing video data generation unit 300 then performs a digital zoom process centered on the singer on the selected video data VD'1, based on the person information "original singer" and zoom information "in" shown in Figure 9, and generates video data ZD'1-2. The singing video data generation unit 300 associates the playback time "00:00:08 to 00:00:12" with the generated video data ZD'1-2. Furthermore, the singing video data generation unit 300 selects video data VD'2 from among video data VD'1 to VD'3 as video data for the performance time from "00:00:12" to "00:00:14", corresponding to the video shot at a shooting angle that matches the angle information "left". The singing video data generation unit 300 then performs a digital zoom process on the selected video data VD'2, centering it on the guitarist (the person corresponding to the guitarist) based on the person information "guitarist" and zoom information "in" shown in Figure 9, and generates video data ZD'2-3. The singing video data generation unit 300 associates the playback time "00:00:12 to 00:00:14" with the generated video data ZD'2-3.

[0089] By repeatedly performing this process, the singing video data generation unit 300 can generate the singing video data shown in Figure 10. For example, in the singing video data shown in Figure 10, video data ZD'1-2 is associated with the playback time "00:00:08~00:00:12", and video data ZD'2-3 is associated with the playback time "00:00:12~00:00:14". This indicates that when playing the singing video data, from the playback time "00:00:08", video data ZD'1-2 is played, which corresponds to a video shot from a frontal angle and digitally zoomed in on the singer, while from the playback time "00:00:12", video data ZD'2-3 is played, which corresponds to a video shot from a left angle and digitally zoomed in on the guitarist.

[0090] As is clear from the above, in the karaoke device K according to this modified example, the live video data processing unit 100 generates shooting information including person information indicating a person appearing in the image of the frame from which angle information has been extracted, and the video corresponding to the video signals transmitted from multiple shooting means is obtained by shooting the singer and the musicians who play instruments live in time with the karaoke music when the singer performs karaoke, and the singing video data generation unit 300 generates singing video data using the video data that has undergone digital zoom processing based on the person information and zoom information.

[0091] According to this karaoke device K, it is possible to select video data shot at a shooting angle that matches the angle information extracted from live video data, perform digital zoom processing on the selected video data centered on a specific person, and then generate singing video data by combining them in chronological order. The singing video based on the generated singing video data displays the person with the same shooting angle and zoom state as the live video. In other words, according to the karaoke device K according to this modified example, it is possible to generate singing video data for displaying a singing video that reflects the shooting angle, zoom state, and the person shown in the video of the original singer's live performance. In addition, if the singing video data generation unit 300 of this modified example cannot identify a person corresponding to the person information in any of the selected video data VD'1 to VD'3, it may perform digital zoom processing centered on a predetermined person (for example, the singer).

[0092] <Variation 3> When multiple users share a karaoke machine, audience members may use their mobile devices as a means of taking photos or videos. The audience members are those other than the singers. The mobile devices are smartphones or other devices with camera capabilities. In this modified example, each mobile device has a dedicated application software (hereinafter referred to as the "photography app") pre-installed for taking photos of the stage.

[0093] [Control means] The control means 10e performs various controls on the karaoke machine K. The control means 10e includes a CPU and memory (neither of which are shown in the figure). The CPU realizes various functions by executing programs stored in the memory.

[0094] In this modified example, the CPU executes a program stored in memory, causing the control means 10e to function as a live video data processing unit 100, a video data acquisition unit 200, a singing video data generation unit 300, and a pairing unit 400 (see Figure 11).

[0095] (Pairing section) The pairing unit 400 pairs with the mobile device held by the audience.

[0096] When multiple users use the karaoke machine K, one user sings karaoke. In this case, the other users launch a camera application on their mobile devices. The mobile devices transmit their device identification information to the karaoke machine K. The device identification information is unique to each mobile device, such as a device ID used to identify the mobile device. The pairing unit 400 pairs with the mobile device by storing the device ID received from the mobile device in the storage means 10a. The paired mobile device functions as a camera.

[0097] For example, suppose four users U1 to U4 use the karaoke machine K in the karaoke room KR shown in Figure 1. In this example, assume that the camera devices SE1 to SE3 are not installed. Also, assume that each user has a mobile device M1 to M4 with a camera application installed.

[0098] When user U1 performs karaoke as the singer, users U2 to U4, who are the audience, launch a photo app on their mobile devices M2 to M4. Each mobile device transmits its own device ID to the karaoke device K. The pairing unit 400 pairs the mobile devices M2 to M4 by storing the received device IDs in the storage means 10a.

[0099] Specifically, the pairing unit 400 stores the terminal ID of mobile terminal M2 as a shooting method for shooting from a shooting angle facing the front of the stage, stores the terminal ID of mobile terminal M3 as a shooting method for shooting from a shooting angle facing the left side of the stage, and stores the terminal ID of mobile terminal M4 as a shooting method for shooting from a shooting angle facing the right side of the stage.

[0100] Furthermore, users can specify any shooting angle via the shooting app. For example, in the above example, suppose user U2 launches the shooting app on mobile device M2 and then specifies "left" as the desired shooting angle. Mobile device M2 transmits the device ID and the specified shooting angle "left" to karaoke device K.

[0101] The pairing unit 400 performs pairing with the mobile terminal M2 by storing the terminal ID of the mobile terminal M2 as a shooting means to shoot from the left side of the stage, according to the received shooting angle.

[0102] (Video data acquisition unit) The video data acquisition unit 200 in this modified example acquires at least a portion of multiple video data based on the video signal transmitted from the paired mobile terminal.

[0103] Each paired mobile device starts recording in sync with the karaoke performance based on instructions from the karaoke machine K. Each mobile device records the stage and transmits the resulting video signal to the karaoke machine K. The video data acquisition unit 200 acquires multiple video data based on the received video signal. Each mobile device can record in sync with the karaoke performance based on the synchronization signal transmitted from the karaoke machine K.

[0104] Specifically, as a result of the pairing described above, mobile device M2 functions as a shooting device for shooting from a front-on angle of the stage, mobile device M3 functions as a shooting device for shooting from a left-side angle of the stage, and mobile device M4 functions as a shooting device for shooting from a right-side angle of the stage.

[0105] The video data acquisition unit 200 instructs the mobile terminals M2 to M4 to start recording in sync with the karaoke performance when the karaoke performance of song X begins.

[0106] Mobile terminal M2 starts shooting from a front-facing angle of the stage based on instructions from karaoke machine K. Mobile terminal M2 transmits the obtained video signal (video signal corresponding to the "front" shooting angle) to karaoke machine K. Similarly, mobile terminal M3 starts shooting from a left-facing angle of the stage based on instructions from karaoke machine K. Mobile terminal M3 transmits the obtained video signal (video signal corresponding to the "left" shooting angle) to karaoke machine K. Mobile terminal M4 starts shooting from a right-facing angle of the stage based on instructions from karaoke machine K. Mobile terminal M4 transmits the obtained video signal (video signal corresponding to the "right" shooting angle) to karaoke machine K.

[0107] The video data acquisition unit 200 receives video signal VS1 from mobile terminal M2, video signal VS2 from mobile terminal M3, and video signal VS3 from mobile terminal M4 in synchronization with the karaoke performance of song X. The video data acquisition unit 200 acquires video data VD1 to VD3 based on video signals VS1 to VS3. The video data acquisition unit 200 stores the acquired video data VD1 to VD3 in the storage means 10a.

[0108] The video data acquisition unit 200 may also input the video signal transmitted from the paired mobile terminal into a pre-trained model similar to the image pre-trained model described in the embodiment (a pre-trained model that links video and shooting angle) to extract the shooting angle, and acquire video data according to the extracted shooting angle.

[0109] For example, the video data acquisition unit 200 inputs video corresponding to a video signal transmitted from a certain mobile terminal into a trained model. The trained model outputs "front" as the shooting angle associated with video that matches or is similar to the input video. Based on the shooting angle "front" output from the trained model, the video data acquisition unit 200 determines that the video signal transmitted from a certain mobile terminal is video signal VS1 transmitted from mobile terminal M2. Based on the received video signal VS1, the video data acquisition unit 200 acquires video data VD1.

[0110] Furthermore, a combination of the shooting means described in the embodiment and a mobile terminal may be used as the shooting means. For example, shooting means SE1 can be used as the shooting means to shoot from the front of the stage as shown in Figure 1, mobile terminal M3 can be used as the shooting means to shoot from the left side of the stage, and shooting means SE3 can be used as the shooting means to shoot from the right side of the stage.

[0111] As is clear from the above, the karaoke device K according to this modified example has a pairing unit 400 that pairs with a mobile terminal held by an audience member, who is a user other than the singer. The video data acquisition unit 200 acquires at least a portion of a plurality of video data based on the video signal transmitted from the paired mobile terminal.

[0112] With such a karaoke device K, the video data used to generate singing video data can be video data obtained by shooting with a mobile device owned by the user.

[0113] <Other> The above embodiments are presented as examples and do not limit the scope of the invention. The above configurations can be combined as appropriate, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. The above embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims of the invention and its equivalents. [Explanation of Symbols]

[0114] 100 Live video data processing unit 200 Video data acquisition unit 300 Singing Video Data Generation Unit 400 Pairing section K Karaoke machine

Claims

1. A live video data processing unit generates shooting information by extracting angle information indicating the shooting angle from live video data of the original singer associated with the song selected by the singer, and associating the extracted angle information with the performance time of the song corresponding to the frame from which the angle information was extracted. A shooting means for shooting the stage on which the singer is located, comprising a video data acquisition unit that acquires multiple video data based on video signals transmitted from multiple shooting means, each set to a different shooting angle, A singing video data generation unit selects one video data corresponding to a video shot at a shooting angle that matches the angle information included in the generated shooting information, and generates singing video data by combining the selected multiple video data in chronological order. A karaoke machine having the following features.

2. The live video data processing unit generates the shooting information, which includes zoom information indicating the zoom state in the frame from which the angle information was extracted. The karaoke device according to claim 1, characterized in that the singing video data generation unit generates the singing video data using a plurality of video data that have undergone digital zoom processing based on the zoom information.

3. The live video data processing unit generates the shooting information, which includes person information indicating a person appearing in the image of the frame from which the angle information has been extracted. The video corresponding to the video signals transmitted from the aforementioned multiple shooting means is obtained by filming the singer on the stage and the musicians who play instruments live in sync with the karaoke music when the singer performs karaoke. The karaoke device according to claim 2, characterized in that the singing video data generation unit generates singing video data using a plurality of video data that have undergone digital zoom processing based on the person information and the zoom information, respectively.

4. It has a pairing unit that pairs with mobile devices owned by audience members, who are users other than the singer. The karaoke device according to any one of claims 1 to 3, characterized in that the video data acquisition unit acquires at least a portion of the plurality of video data based on a video signal corresponding to a video of the stage where the singer is located, transmitted from the paired mobile terminal.