Server equipment

JP7897729B2Active Publication Date: 2026-07-30DAIICHI KOSHO COMPANY
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
DAIICHI KOSHO COMPANY
Filing Date
2022-06-28
Publication Date
2026-07-30

AI Technical Summary

Benefits of technology

【0008】 本発明によれば、歌唱動画の撮影開始が遅れた場合であっても、前奏全てのカラオケ演奏音を含む歌唱動画データを生成できる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007897729000001
    Figure 0007897729000001
  • Figure 0007897729000002
    Figure 0007897729000002
  • Figure 0007897729000003
    Figure 0007897729000003
Patent Text Reader

Abstract

To provide a server device which can generate singing video data including all karaoke performance sounds of introduction even if a photographing start of singing video is delayed.SOLUTION: A server device includes: an acquisition part for acquiring singing video data corresponding to singing video from a portable terminal; a specification part for specifying music identification information of music which is karaoke-performed in singing video corresponding to the acquired singing video data; an extraction part for extracting accompaniment data corresponding to introduction from accompaniment data of music, which corresponds to specified music identification information; and an editing processing part for generating edited video data by editing the acquired singing video data by using the extracted accompaniment data.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a server device.

Background Art

[0002] A user of a karaoke device may shoot a singing video using a mobile terminal such as a smartphone while performing karaoke singing.

[0003] For example, Patent Document 1 discloses a mobile terminal that generates a singing video based on karaoke performance sound based on music performance data, singing voice collected by a sound collection means, and a moving image obtained by shooting a singer performing karaoke singing by a shooting means.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] By the way, a certain user can shoot a singing video of another user performing karaoke singing using their own mobile terminal. However, when another user performs karaoke singing compared to when oneself performs karaoke singing, it is difficult to grasp the timing when the karaoke performance starts, so it may be possible to start shooting the singing video after the prelude of the music has started. The singing video obtained in this way lacks a part of the prelude.

[0006] An object of the present invention is to provide a server device that enables the generation of singing video data including all karaoke performance sounds of the prelude even when the start of shooting the singing video is delayed.

Means for Solving the Problems

[0007] One invention for achieving the above objective is a server device that is communicatively connected to a plurality of karaoke devices, and comprises: an acquisition unit that acquires singing video data corresponding to a singing video obtained by filming a user singing karaoke with the mobile terminal, an identification unit that identifies song identification information of a song being played in karaoke in the singing video corresponding to the acquired singing video data, an extraction unit that extracts accompaniment data corresponding to the intro from the accompaniment data of the song corresponding to the identified song identification information, and an editing processing unit that generates edited video data by editing the acquired singing video data using the extracted accompaniment data. Other features of the present invention will be revealed in the specification and drawings described below. [Effects of the Invention]

[0008] According to the present invention, even if the recording of the singing video is delayed, it is possible to generate singing video data that includes the entire karaoke performance sound of the intro. [Brief explanation of the drawing]

[0009] [Figure 1] This is a diagram showing a communication karaoke system according to an embodiment. [Figure 2] This is a diagram showing a server device according to an embodiment. [Figure 3] This is a flowchart showing the processing of the server device according to the embodiment. [Modes for carrying out the invention]

[0010] <Embodiment> The server device according to this embodiment will be described with reference to Figures 1 to 3.

[0011] ==Communication Karaoke System== The communication karaoke system includes multiple karaoke devices and a server device. Each karaoke device and the server device are connected to each other via a network so that they can communicate with each other.

[0012] A karaoke machine has the function of playing songs on karaoke and singing karaoke. A karaoke machine has various components such as a microphone, display, speaker, and remote control. A karaoke machine is assigned device identification information. Device identification information is information unique to each karaoke machine, such as a device ID used to identify the karaoke machine.

[0013] The server device is a computer that manages various information and performs various processes. The server device according to this embodiment can edit singing video data corresponding to singing videos captured by a mobile terminal (details will be described later). The singing video includes karaoke performance sound based on song accompaniment data (described later), the singing voice of the user performing karaoke, and video footage obtained by capturing the user performing karaoke.

[0014] Mobile devices are smartphones, tablets, and other devices owned by users of karaoke equipment. These devices are equipped with microphones to collect karaoke music and singing voices, and cameras to film users singing karaoke. Each mobile device is assigned device identification information. This device identification information includes a device ID and other unique information specific to each device.

[0015] The mobile terminal according to this embodiment has dedicated application software (hereinafter referred to as the "singing video editing app") installed for editing singing videos on the server device. The singing video editing app can be obtained, for example, by downloading it from a designated website provided by the server device. By launching the singing video editing app on the mobile terminal, communication between the mobile terminal and the server device is established.

[0016] As shown in Figure 1, the communication karaoke system 1 according to this embodiment includes karaoke devices K1 to Kn and a server device S. Each karaoke device and the server device S are connected to each other via a network N so as to be able to communicate.

[0017] The karaoke device K1 is installed in the karaoke store L1, and the karaoke device Kn is installed in the karaoke store Ln. There are two users, user U1 and user U2, staying in the karaoke store L1. User U1 owns a mobile terminal M1. A singing video editing app is installed on the mobile terminal M1.

[0018] Note that the number of karaoke devices and the number of users using the karaoke devices are not limited to the example in FIG. 1.

[0019] ==Server Device== As shown in FIG. 2, the server device S includes a storage unit 10, a communication unit 20, and a control unit 30. Each component is connected to the bus B via an interface (not shown).

[0020] [Storage Unit] The storage unit 10 is a large-capacity storage device that stores various types of data. The storage unit 10 according to the present embodiment stores music data. The music data is provided with music identification information. The music identification information is information unique to each music, such as a music ID for identifying the music. The music data includes accompaniment data, reference data, section information, etc. The accompaniment data is MIDI-formatted data that serves as the source of the karaoke performance sound. The reference data is data indicating the main melody of the music performed in karaoke. The section information indicates the performance section. The performance section is the section in which karaoke performance is performed. The performance section includes a singing section and a non-singing section. The singing section is a section (for example, the A melody, B melody, and refrain of the first song) in which the lyrics to be sung in a certain music are set. The non-singing section is a section in which the lyrics to be sung in a certain music are not set, such as a prelude, interlude, or postlude. The music data is distributed to each karaoke device connected via the network N.

[0021] In addition, the storage unit 10 stores, for each music, background video data corresponding to the background video displayed during karaoke performance, lyric telop data for displaying the lyrics of the music, and attribute information of the music (music name, singer name, genre, etc.).

[0022] [g means] The communication means 20 provides an interface for communicating with karaoke machines and mobile terminals.

[0023] [Control means] The control means 30 performs various controls on the server device S. The control means 30 includes a CPU and memory (neither of which are shown in the figure). The CPU implements various functions by executing programs stored in the memory.

[0024] In this embodiment, the CPU executes a program stored in memory, and the control means 30 functions as an acquisition unit 100, a identification unit 200, an extraction unit 300, and an editing processing unit 400.

[0025] (Acquisition Department) The acquisition unit 100 acquires singing video data corresponding to the singing video obtained by filming a user singing karaoke with the mobile device, from the said mobile device.

[0026] A user who has recorded a singing video using a mobile device launches a singing video editing application installed on the mobile device. Server device S instructs the mobile device that has launched the singing video editing application to send the singing video data corresponding to the singing video that the user wishes to edit.

[0027] Upon receiving instructions, the mobile terminal transmits the singing video data corresponding to the singing video to be edited, along with terminal identification information, to the server device S. The acquisition unit 100 acquires the transmitted singing video data and terminal identification information.

[0028] The above explanation describes an example in which a mobile terminal transmits singing video data to the server device S based on instructions from the server device S. However, the system may also be configured to allow the mobile terminal to transmit singing video data to the server device S even without instructions from the server device S. Specifically, the user launches a singing video editing application on their mobile terminal, operates the mobile terminal to select the singing video they wish to edit, and requests editing from the server device S. The mobile terminal transmits the singing video data corresponding to the selected singing video and terminal identification information to the server device S along with the editing request signal. In this case as well, the acquisition unit 100 can acquire the transmitted singing video data and terminal identification information.

[0029] Alternatively, instead of using a singing video editing app, the mobile device can access a dedicated website provided by the server device S based on user input and upload singing video data and device identification information. In this case, the acquisition unit 100 acquires the uploaded singing video data and device identification information.

[0030] (Specific part) The identification unit 200 identifies the song identification information of the song being played in the karaoke video corresponding to the acquired singing video data.

[0031] Identifying song titles can be done in various ways. For example, when a singing video is acquired, the identification unit 200 connects to a known external site where song titles can be searched from audio data corresponding to karaoke performance sounds. The identification unit 200 transmits the audio data corresponding to the karaoke performance sounds contained in the acquired singing video to the external site. The external site uses the received audio data to search for song titles and transmits the search results (i.e., the song titles of the songs being played in the singing video) to the identification unit 200. The identification unit 200 refers to the attribute information of each song stored in the storage means 10 and extracts songs that match the song titles received from the external site. The identification unit 200 identifies the song titles of the extracted songs as the song titles of the songs being played in the singing video corresponding to the acquired singing video data.

[0032] Alternatively, the server device S may perform the same processing as an external site. In this case, the identification unit 200 can identify the song identification information of a song without connecting to an external site. Furthermore, if an external site is not used, the identification unit 200 can also identify the song identification information without relying on the song title. Specifically, the identification unit 200 compares the audio data corresponding to the karaoke performance sound included in the singing video data with the song data for each song stored in the storage means 10 (specifically, audio data rendered by software from the accompaniment data). The identification unit 200 identifies the song identification information attached to the song data that is most similar to the audio data corresponding to the karaoke performance sound included in the singing video as the song identification information of the song being played in the karaoke performance in the singing video corresponding to the acquired singing video data.

[0033] (Extraction part) The extraction unit 300 extracts accompaniment data corresponding to the intro from the accompaniment data of the song corresponding to the identified song identification information.

[0034] The extraction unit 300 reads accompaniment data and section information from the song data to which the identified song identification information has been assigned. Based on the section information, the extraction unit 300 extracts the accompaniment data corresponding to the intro from the read accompaniment data. In this embodiment, the extraction unit 300 extracts the accompaniment data (MIDI format data) corresponding to the intro in the form of audio data rendered by software.

[0035] (Editing and processing) The editing processing unit 400 generates edited video data by editing the acquired singing video data using the extracted accompaniment data. The edited video data is singing video data that includes the karaoke performance sound for the entire intro.

[0036] Edited video data can be generated by various methods. For example, the editing processing unit 400 sequentially compares the waveform of the karaoke performance sound audio data included in the acquired singing video data for a predetermined period (for example, 0 msec to 100 msec) from the beginning with the waveform of the extracted accompaniment data from the beginning, and determines the timing when the waveforms match. Waveform comparison can be performed using known techniques. The editing processing unit 400 determines that the timing when the waveforms match is the timing when the shooting of the singing video corresponding to the singing video data started.

[0037] In this case, the acquired singing video data does not include the karaoke performance sound from the start of the intro (0 msec) up to the determined timing. Therefore, the editing processing unit 400 edits the singing video data by inserting the accompaniment data from the start of the intro (0 msec) up to the determined timing from the extracted accompaniment data to the beginning of the audio data corresponding to the karaoke performance sound included in the acquired singing video data. In this case, edited video data that includes the entire karaoke performance sound of the intro is generated.

[0038] Furthermore, the editing processing unit 400 can edit the acquired singing video data using background video data of the song corresponding to the identified song identification information.

[0039] The editing processing unit 400 reads the background video data of the song being played in the karaoke version of the song in the song video that corresponds to the acquired song video data. From the read background video data, the editing processing unit 400 extracts only the background video data corresponding to the accompaniment data inserted at the beginning of the audio data corresponding to the karaoke performance sound included in the acquired song video data. The editing processing unit 400 edits the song video data by inserting the extracted background video data at the beginning of the video data included in the acquired song video data. In this case, edited video data is generated that includes the background video corresponding to the entire karaoke performance sound of the intro and the inserted accompaniment data.

[0040] Furthermore, the editing processing unit 400 can transmit the generated edited video data to the mobile device that filmed the singing video. Transmission is performed based on the acquired device identification information. Alternatively, the editing processing unit 400 can upload the generated edited video data to a dedicated site and publish the edited singing video. It is preferable to obtain prior permission from the user who filmed the singing video and the user who is singing karaoke in the singing video before publishing the edited singing video. For example, the user can choose in advance to allow the publication of the edited singing video through the singing video editing application.

[0041] ==Regarding the operation of the karaoke system== Next, a specific example of the operation of the server device S according to this embodiment will be described with reference to Figure 3. Figure 3 is a flowchart showing the operation of the server device S. In this example, as shown in Figure 1, it is assumed that users U1 and U2 are using the karaoke machine K1 at the karaoke establishment L1. It is also assumed that a singing video editing application is installed on the mobile terminal M1 owned by user U1.

[0042] The acquisition unit 100 acquires singing video data VD corresponding to the singing video V obtained by filming user U2 singing karaoke with the mobile terminal M1 from the mobile terminal M1 (acquisition of singing video data from the mobile terminal. Step 10).

[0043] The identification unit 200 identifies the song identification information of the song being played in the karaoke video V corresponding to the acquired singing video data VD (identification of song identification information; step 11).

[0044] The extraction unit 300 extracts accompaniment data corresponding to the intro from the accompaniment data of the song corresponding to the song identification information identified in step 11 (extraction of accompaniment data corresponding to the intro; step 12).

[0045] The editing processing unit 400 edits the singing video data VD acquired in step 10 using the accompaniment data extracted in step 12 (editing the singing video data using the extracted accompaniment data; step 13).

[0046] Furthermore, the editing processing unit 400 edits the singing video data VD acquired in step 10 using the background video data of the song corresponding to the song identification information identified in step 11 (editing the singing video data using the background video data; step 14).

[0047] The editing processing unit 400 generates edited video data VD' by executing the processes in steps 13 and 14 (generates edited video data; step 15).

[0048] The editing processing unit 400 sends the edited video data VD' generated in step 15 to the mobile terminal M1 (sending edited video data to the mobile terminal; step 16).

[0049] Specifically, user U2 operates the remote control of karaoke machine K1 and selects song X, which they wish to sing. Karaoke machine K1 starts playing song X, which user U2 selected. User U2 sings along to the karaoke music. Meanwhile, user U1, after hearing the intro to song X and realizing that the karaoke performance of song X has started, begins recording a video of the singing with their mobile device M1. In this case, the resulting video V will have a portion of the intro to song X missing.

[0050] Therefore, after user U1 has finished recording the singing video, U1 launches the singing video editing application installed on the mobile device M1. The server device S instructs the mobile device M1, which has launched the singing video editing application, to send the singing video data VD corresponding to the singing video V that U1 wishes to edit.

[0051] Upon receiving instructions, the mobile terminal M1 transmits the singing video data VD corresponding to the singing video V that user U1 wishes to edit, and the terminal ID ***M1, to the server device S.

[0052] The acquisition unit 100 acquires the transmitted singing video data VD and terminal ID ***M1.

[0053] When a singing video data VD is acquired, the identification unit 200 connects to a known external site where song titles can be searched from audio data corresponding to karaoke performance sounds. The identification unit 200 transmits the audio data corresponding to the karaoke performance sounds contained in the singing video data VD to the external site. The external site uses the received audio data to search for song titles and transmits the song title of song X as a search result to the identification unit 200. The identification unit 200 refers to the attribute information of each song stored in the storage means 10 and extracts song X that matches the song title received from the external site. The identification unit 200 identifies the song ID***X of the extracted song X as song identification information for the song being played in karaoke in the singing video V corresponding to the singing video data VD.

[0054] The extraction unit 300 reads accompaniment data and section information from the song data to which the identified song ID***X has been assigned. Based on the section information, the extraction unit 300 extracts the accompaniment data corresponding to the intro from the read accompaniment data.

[0055] The editing processing unit 400 sequentially compares the waveform W of the karaoke performance sound audio data contained in the singing video data VD from the beginning (0 mse) to 100 msec with the waveform of the extracted accompaniment data at 100 msec intervals from the beginning, and determines the timing at which the waveforms match.

[0056] Here, assume that waveform W matches the waveform Wn of the extracted accompaniment data from 4,500 msec to 4,600 msec. In this case, the editing processing unit 400 determines that the timing from the start of the intro to 4,500 msec is the timing when the recording of the singing video V corresponding to the singing video data VD began.

[0057] The editing processing unit 400 edits the singing video data VD by inserting the accompaniment data from the start of the intro (0 msec) to 4,500 msec from the extracted accompaniment data into the beginning of the audio data corresponding to the karaoke performance sound included in the singing video data VD.

[0058] Furthermore, the editing processing unit 400 reads the background video data of the song X being played in the singing video V from the storage means 10. The editing processing unit 400 extracts only the background video data (data corresponding to the background video displayed from 0 msec to 4,500 msec) that corresponds to the accompaniment data (data from 0 msec to 4,500 msec) inserted at the beginning of the audio data corresponding to the karaoke performance sound included in the singing video data VD from the read background video data. The editing processing unit 400 edits the singing video data VD by inserting the extracted background video data at the beginning of the video data included in the singing video data VD.

[0059] The editing processing unit 400 generates edited video data VD' by executing the two editing processes described above.

[0060] The editing processing unit 400 sends the generated edited video data VD' to the mobile terminal M1 that filmed the singing video V.

[0061] Furthermore, the editing processing unit 400 can also perform a process such as crossfading when inserting a portion of the extracted accompaniment data into the beginning of the audio data corresponding to the karaoke performance sound included in the singing video data. For example, in the above example, the editing processing unit 400 may edit the video so that, between 4,500 msec and 4,800 msec (a 300 msec period) after the start of shooting the singing video V, the volume level of the extracted accompaniment data is gradually reduced from 100% to 0%, and the volume level of the audio data of the karaoke performance sound included in the singing video data is gradually increased from 0% to 100%.

[0062] As is clear from the above, the server device S according to this embodiment is connected to a plurality of karaoke devices K1 to Kn in a communicative manner. The server device S includes an acquisition unit 100 that acquires singing video data corresponding to a singing video obtained by filming a user singing karaoke with the mobile terminal, an identification unit 200 that identifies song identification information of the song being played in the singing video corresponding to the acquired singing video data, an extraction unit 300 that extracts accompaniment data corresponding to the intro from the accompaniment data of the song corresponding to the identified song identification information, and an editing processing unit 400 that generates edited video data by editing the acquired singing video data using the extracted accompaniment data.

[0063] With such a server device S, even if singing video data with a missing intro is acquired from a mobile terminal, edited video data can be generated that supplements the missing intro using the song's accompaniment data. In other words, with the server device S according to this embodiment, even if the start of shooting the singing video is delayed, singing video data including the entire karaoke performance sound of the intro can be generated.

[0064] Furthermore, in the server device S according to this embodiment, the editing processing unit 400 can transmit the generated edited video data to a mobile terminal. Therefore, a mobile terminal that has recorded a singing video in which part of the intro is missing can obtain singing video data that includes the entire karaoke performance sound of the intro.

[0065] Furthermore, in the server device S according to this embodiment, the editing processing unit 400 can edit the acquired singing video data using background video data of the song corresponding to the identified song identification information. With such a server device S, it is possible to generate edited video data that includes the entire karaoke performance sound of the intro and background video corresponding to the supplemented intro.

[0066] <Example 1> As described in the embodiments, song identification information can be identified in various ways. For example, the identification unit 200 can identify song identification information based on performance history information stored for each karaoke machine. Performance history information is various information related to karaoke performances performed in the karaoke machine. Specifically, the performance history information includes song identification information for the song that was performed in the karaoke machine, the start time and actual performance time of the karaoke performance (the difference between the start time and the time when the performance actually ended), key information indicating the key value set in the karaoke performance, tempo information indicating the tempo value set in the karaoke performance, and so on.

[0067] (Acquisition Department) The acquisition unit 100 in this modified version acquires location information indicating the location of the mobile terminal, and shooting time information including the start time of shooting the singing video, from the mobile terminal. The location information can be obtained using the GPS function that is generally available on mobile terminals.

[0068] For example, in the embodiment, user U1 launches a singing video editing application installed on mobile terminal M1. Server device S instructs mobile terminal M1, which has launched the singing video editing application, to send singing video data VD corresponding to the singing video V to be edited, location information, and shooting time information.

[0069] Upon receiving the instructions, the mobile terminal M1 transmits to the server device S the singing video data VD corresponding to the singing video V to be edited, along with terminal ID***M1, location information P obtained using the GPS function, and shooting time information T including the start time t of shooting the singing video V.

[0070] The acquisition unit 100 acquires the transmitted singing video data VD, terminal ID***M1, location information P, and shooting time information T. In this case, the location information P is approximately equal to the location where the singing video V was filmed, i.e., the location of the karaoke store where the karaoke machine K1 is installed.

[0071] (Specific part) In this modified example, the identification unit 200 identifies a karaoke machine based on acquired location information and identifies song identification information based on acquired shooting time information and performance history information stored in the karaoke machine. In this modified example, the server device S stores in advance device location information (for example, the address of the karaoke establishment) indicating the location where each karaoke machine capable of communication is installed.

[0072] As described above, when singing video data VD is acquired, the identification unit 200 refers to the device location information for each karaoke machine stored in the server device S and finds the device location information that matches or most closely approximates the received location information P. The identification unit 200 identifies the karaoke machine corresponding to the device location information as the karaoke machine.

[0073] The identification unit 200 refers to the performance history information (specifically, the performance start time and performance duration) obtained from a single karaoke machine and extracts the songs that were being played at the time of the start of filming included in the filming time information. The identification unit 200 identifies the song identification information of the extracted songs as the song identification information of the songs being played in the karaoke video corresponding to the acquired singing video data.

[0074] For example, suppose karaoke device K1 is identified as one of the karaoke devices. In this case, the identification unit 200 refers to the performance history information obtained from karaoke device K1 and extracts song X as the song that was being played at the start time t of the recording included in the recording time information T. The identification unit 200 identifies the song ID***X of the extracted song X as the song identification information for the song being played in the recording video V corresponding to the acquired recording video data VD.

[0075] As is clear from the above, in the server device S according to this modified example, the acquisition unit 100 acquires location information indicating the location of the mobile terminal and shooting time information including the start time of shooting the singing video from the mobile terminal. The identification unit 200 identifies a karaoke machine based on the acquired location information and identifies song identification information based on the acquired shooting time information and the performance history information stored in the karaoke machine. With such a server device S, song identification information can be identified based on the performance history information stored in a karaoke machine. In addition, the storage means of the server device S may store the performance history information of each connected karaoke machine. In this case, the server device S updates the information stored in the storage means as needed based on the performance history information transmitted from each karaoke machine.

[0076] <Modification 2> The determination of the timing at which recording of a singing video corresponding to singing video data began is not limited to the examples of this embodiment. For example, the editing processing unit 400 can determine the timing at which recording of a singing video corresponding to singing video data began based on the acquired recording time information and performance history information.

[0077] In this modified example, as in Modified Example 1, the acquisition unit 100 acquires location information and shooting time information from the mobile terminal. The identification unit 200 identifies a karaoke machine based on the acquired location information and identifies song identification information based on the acquired shooting time information and the performance history information stored in the karaoke machine.

[0078] The editing processing unit 400 determines the timing when the singing video was started based on the difference between the start time of the singing video recording included in the acquired shooting time information and the start time of the performance of the song corresponding to the identified song identification information.

[0079] For example, as described in Modification 1, suppose the acquisition unit 100 acquires singing video data VD, terminal ID***M1, location information P, and shooting time information T. Also, suppose the identification unit 200 identifies the song ID***X of the song X being played in karaoke in the singing video V corresponding to the singing video data VD.

[0080] In this case, the editing processing unit 400 calculates the difference D between the start time t of filming the singing video V included in the filming time information T and the start time t' of playing the song X. The editing processing unit 400 determines that the time indicated by the difference D is the timing when filming of the singing video V began.

[0081] <Variation 3> The extraction unit 300 may consider key and tempo values ​​when rendering accompaniment data (MIDI format data) corresponding to the intro using software. For example, the extraction unit 300 can render accompaniment data corresponding to the intro based on key information and / or tempo information contained in the performance history information stored in a specific karaoke device.

[0082] In this modified example, as in Modified Examples 1 and 2, the acquisition unit 100 acquires location information and shooting time information from the mobile terminal. The identification unit 200 identifies a karaoke machine based on the acquired location information and identifies song identification information based on the acquired shooting time information and the performance history information stored in the karaoke machine.

[0083] The extraction unit 300 obtains key information KS and / or tempo information TP set at the recording start time t indicated by the recording time information T from the performance history information of a identified karaoke device.

[0084] The extraction unit 300 uses software to render MIDI-formatted accompaniment data corresponding to the prelude, based on key information KS and / or tempo information TP.

[0085] By performing this process, the editing processing unit 400 can insert accompaniment data that matches the key and tempo values ​​set in the karaoke performance within the singing video.

[0086] <Modification 4> When editing the acquired singing video data, the editing processing unit 400 can use flash cut video data extracted from the singing video instead of using the background video data described in the embodiment.

[0087] (Editing and processing) The editing processing unit 400 in this modified version extracts multiple flash cut video data corresponding to multiple flash cut videos from the acquired singing video data, and edits the acquired singing video data using the extracted multiple flash cut video data.

[0088] Flash cut footage is footage used for flash cuts. The editing processing unit 400 extracts multiple flash cut footage data from the video data included in the singing video data. The number of flash cut footage data and the length of the corresponding flash cut footage are not particularly limited, but are preferably determined according to the length of the accompaniment data inserted into the singing video data.

[0089] The editing processing unit 400 edits the singing video data by sequentially inserting the extracted flash cut video data into the beginning of the video data contained in the singing video data.

[0090] For example, as described in the embodiment, suppose the editing processing unit 400 determines that the timing 4,500 msec from the start of the intro is the timing at which it started shooting the singing video V corresponding to the singing video data VD.

[0091] In this case, the editing processing unit 400 extracts the first 500 msec of the A section of the first verse of song X as flash cut video data FCD1 from the video data VD, extracts the first 500 msec of the B section of the first verse of song X as flash cut video data FCD2, extracts the first 1,000 msec of the chorus of the first verse of song X as flash cut video data FCD3, extracts the first 500 msec of the A section of the second verse of song X as flash cut video data FCD4, extracts the first 500 msec of the B section of the second verse of song X as flash cut video data FCD5, extracts the first 1,000 msec of the chorus of the second verse of song X as flash cut video data FCD6, and extracts the first 500 msec of the A section of the third verse of song X as flash cut video data FCD7.

[0092] The editing processing unit 400 then edits the singing video data VD by assigning flash cut video data FCD1 to the accompaniment data from the start of the intro (0 msec) to 4,500 msec at the start of the intro, flash cut video data FCD2 at 500 msec after the start of the performance, flash cut video data FCD3 at 1,000 msec after the start of the performance, flash cut video data FCD4 at 2,000 msec after the start of the performance, flash cut video data FCD5 at 2,500 msec after the start of the performance, flash cut video data FCD6 at 3,000 msec after the start of the performance, and flash cut video data FCD7 at 4,000 msec after the start of the performance.

[0093] The above example describes the extraction of a flash cut video from a singing section, but it is not limited to this. For example, the editing processing unit 400 may extract the first 500 msec of the first verse of song X as flash cut video data FCD1 from the video data included in the singing video data VD, and also extract the data from 2,000 msec to 2,500 msec of the first verse of song X as flash cut video data FCD2. The editing processing unit 400 may also extract flash cut videos from the interlude or outro, which are non-singing sections similar to the intro. Furthermore, the editing processing unit 400 may use a combination of the background video data described in the embodiment and the flash cut video data extracted from the singing video. Specifically, the editing processing unit 400 may use the background video data including the title image from the start of the intro until 1,500 msec, and use the extracted flash cut video data from 1,500 msec to 4,500 msec.

[0094] As is clear from the above, in the server device S according to this modified example, the editing processing unit 400 can extract multiple flash cut video data corresponding to multiple flash cut videos from the acquired singing video data, and edit the acquired singing video data using the extracted multiple flash cut video data. The flash cut videos are videos extracted from singing videos acquired from a mobile terminal. In a singing video edited using such videos, the video of the user who sang karaoke will be displayed even in the part where the intro has been added. Therefore, the discomfort felt by users watching singing videos based on edited video data can be reduced.

[0095] <Other> The above embodiments are presented as examples and do not limit the scope of the invention. The above configurations can be combined as appropriate, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. The above embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims of the invention and its equivalents. [Explanation of Symbols]

[0096] 1. Communication Karaoke System 100 Acquisition Department 200 Specific section 300 Extraction part 400 Editing Processing Unit M1 Mobile Device K1, Kn karaoke equipment S Server Device

Claims

1. A server device that is communicatively connected to multiple karaoke devices, An acquisition unit that acquires singing video data corresponding to a singing video obtained by filming a user singing karaoke with the mobile device, An identification unit identifies song identification information of the song being played in karaoke in the singing video corresponding to the acquired singing video data, An extraction unit extracts accompaniment data corresponding to the intro from the accompaniment data of the song corresponding to the identified song identification information, An editing processing unit generates edited video data by inserting the accompaniment data corresponding to the extracted intro, specifically the accompaniment data from the start of the intro up to the point when the singing video recording begins, into the beginning of the audio data corresponding to the karaoke performance sound included in the acquired singing video data. A server device having the following features.

2. The server device according to claim 1, characterized in that the editing processing unit transmits the generated edited video data to the mobile terminal.

3. The acquisition unit acquires location information indicating the location of the mobile terminal, and shooting time information including the start time of shooting the singing video from the mobile terminal. The server device according to claim 1, characterized in that the identifying unit identifies a karaoke device based on the acquired location information, and identifies the song identification information based on the acquired shooting time information and the performance history information stored in the karaoke device.

4. The server device according to any one of claims 1 to 3, characterized in that the editing processing unit edits the acquired singing video data using background video data of a song corresponding to the identified song identification information.

5. The server device according to any one of claims 1 to 3, characterized in that the editing processing unit extracts multiple flash cut video data corresponding to multiple flash cut videos from the acquired singing video data, and edits the acquired singing video data using the extracted multiple flash cut video data.