Information processing device, moving image synthesis method and moving image synthesis program
The information processing apparatus and method address the limitations of existing karaoke recording systems by enabling users to re-record specific parts of a performance and synthesize the videos at a specified start position, resulting in improved user satisfaction and video editing flexibility.
Patent Information
- Application Number
- JP2025060651
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-06-12
- Estimated Expiration
- 2038-07-25
AI Technical Summary
Existing karaoke recording systems struggle to allow users to re-record specific parts of a performance without restarting from the beginning and do not enable the creation of performance videos like singing videos.
An information processing apparatus and method that allows users to determine a synthesis start position for re-recording a video, synthesize the original video with the re-recorded video at the specified start position, and display guidance images to ensure consistent recording conditions.
Enables users to easily create a desirable video by allowing seamless re-recording and synthesis of video segments, improving user satisfaction and flexibility in video editing.
Smart Images

Figure 2025089595000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, a video synthesis method, and a video synthesis program.
Background Art
[0002] In recent years, online services such as websites (web pages, web services), online games, and application software (hereinafter referred to as "apps") distributed to information processing apparatuses such as mobile terminals via computer networks have become widely popular.
[0003] One of the online services is a video posting site where users can post videos they have taken themselves and allow other users to view the videos. On this video posting site, users also post performance videos such as singing videos by themselves. Therefore, in order to post a better singing video of themselves, users may edit the singing videos to be posted.
[0004] Here, Patent Document 1 discloses a karaoke recording system for dealing with performance interruption, which aims to obtain the recording data of the final singing voice without having to sing the karaoke song from the beginning again even when the performance of the karaoke song is interrupted.
[0005] This karaoke recording system for dealing with performance interruption includes a replay instruction means for instructing a replay of a karaoke song whose performance has been interrupted, and when the performance of any karaoke song for which singing voice recording is being performed by the function of the recording means is interrupted, at least the music of the karaoke song, the recorded performance range data, and the mid-recording data are associated and recorded by recording information recording means. When a replay of a karaoke song is instructed, based on the recorded performance range data, it controls so as not to record the singing voice for the recorded performance range.
Prior Art Documents
Patent Documents
[0006] [Patent Document 1] Japanese Patent Application Laid-Open No. 2010-237389 [Summary of the Invention] [Problems to be Solved by the Invention]
[0007] In the system disclosed in Patent Document 1, recording is additionally performed after the point in time when the recording is interrupted, and the user cannot determine the timing of the karaoke song for which additional recording is to be performed by himself / herself. For this reason, even if the user fails in singing, the user cannot re-record the failed part. Furthermore, in the karaoke recording system for dealing with performance interruption disclosed in Patent Document 1, although the singing voice is recorded, a performance video such as a singing video is not shot.
[0008] The present invention has been made in view of such circumstances, and an object thereof is to provide an information processing apparatus, a video synthesis method, and a video synthesis program that can easily create a video that the user feels desirable when the user records a video using a camera. [Means for Solving the Problems]
[0009] In order to solve the above problems, the information processing apparatus, video synthesis method, and video synthesis program of the present invention employ the following means.
[0010] In order to solve the above problems, an "information processing apparatus" which is an aspect of the present invention is an information processing apparatus that records an image captured by a camera as a video, and includes a determination means for determining a synthesis start position for starting synthesis of a second video to be recorded later with respect to a first video recorded previously, a recording control means for recording the second video with the camera, and a video synthesis means for synthesizing the first video and the second video at the synthesis start position of the first video.
[0011] To solve the above problems, a "video synthesis method" according to one aspect of the present invention includes: a first step of recording a first video with a camera; a second step of determining a synthesis start position for starting the synthesis of a second video to be recorded later with respect to the first video; a third step of recording the second video with the camera; and a fourth step of synthesizing the first video and the second video at the synthesis start position of the first video.
[0012] To solve the above problems, a "video synthesis program" according to one aspect of the present invention causes a computer included in an information processing apparatus that records an image captured by a camera as a video to function as: a determination means for determining a synthesis start position for starting the synthesis of a second video to be recorded later with respect to a first video recorded previously; a recording control means for recording the second video with the camera; and a video synthesis means for synthesizing the first video and the second video at the synthesis start position of the first video.
[0013] As exemplified below, various technical limitations may be imposed on the above "information processing apparatus". Also, technical limitations of the same gist may be added to the processing steps executed by the "video synthesis method" and the functions of the "video synthesis program".
[0014] An image display control means is provided that displays, on a screen, a synthesis start position image that is an image at the synthesis start position of the first video at the start of recording of the second video.
[0015] The image display control means superimposes the synthesis start position image on an image captured by the camera for recording the second video and displays it on the screen.
[0016] The video synthesis means performs image processing to make the synthesis part of the first video and the second video less conspicuous.
[0017] The video synthesis means superimposes the first video and the second video for a predetermined period from the synthesis start position and synthesizes the first video and the second video.
[0018] The first video and the second video are videos recorded while a music piece is being played.
[0019] The recording control means starts playing the music piece before the synthesis start position, and starts recording the second video when the played music piece reaches the timing corresponding to the synthesis start position.
[0020] The recording control means starts playing the music piece before starting the recording of the second video and counts down to the start of the recording of the second video.
[0021] The position that can be set as the synthesis start position is a predetermined position of the music piece.
Advantages of the Invention
[0022] According to the present invention, when a user records a video using a camera, there is an effect that a video that the user feels desirable can be easily created.
Brief Description of the Drawings
[0023]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Embodiments for Carrying Out the Invention
[0024] Hereinafter, an embodiment of an information processing apparatus, a video synthesis method, and a video synthesis program according to the present invention will be described with reference to the drawings.
[0025] In this embodiment, a performer records his / her performance as a video using a mobile terminal and uploads it to a video posting site by sending it to a server. The video (posted video) uploaded to the video posting site can be viewed via the mobile terminal. In this embodiment, the performance is singing, the performer is a singing user, and the video uploaded to the video posting site is a singing video. Also, a user who views the singing video uploaded to the video posting site is called a viewing user.
[0026] [1. Configuration of Karaoke System] FIG. 1 is a schematic configuration diagram of a karaoke system 1 according to this embodiment. The karaoke system 1 includes a communication line 2, a plurality of mobile terminals 3 (mobile terminals 3A, 3B), and a server 4.
[0027] The communication line 2 forms a computer network, and for example, is a wide area communication line provided by an electric utility company.
[0028] The mobile terminal 3 is an information processing terminal such as a smartphone, a tablet terminal, or a notebook computer, and is used for a user to utilize an online service. The mobile terminal 3 includes a touch panel display 3a for displaying an image, a speaker 3b for outputting sound, a microphone 3c for inputting sound, a camera 3d for photographing a subject, and an earphone terminal 3e (not shown) to which an earphone is connected. Here, the photographing means that the camera 3d functions and the subject is displayed on the touch panel display 3a regardless of whether recording is performed. The touch panel display 3a includes, for example, an LCD (Liquid Crystal Display) and a touch sensor. The LCD displays various images, and the touch sensor receives various input operations performed using an indicator such as a finger, a stylus, or a pen. In the following description, the touch panel display 3a is also referred to as the screen 3a.
[0029] The server 4 is an information processing device that provides an online service to the mobile terminal 3 via the communication line 2. In the example of FIG. 1, the singing user uploads the singing video (singing video data) to the video posting site by transmitting the singing video from the mobile terminal 3A to the server 4. Then, the viewing user accesses the video posting site using the mobile terminal 3B and views the singing video. The singing user can also view the singing video uploaded by himself / herself by accessing the video posting site using the mobile terminal 3A. In addition, the user of the mobile terminal 3B can become a singing user and upload the singing video to the video posting site.
[0030] [2. Configuration of Server] FIG. 2 is a block diagram showing the electrical configuration of the server 4 according to the present embodiment.
[0031] The server 4 according to this embodiment includes a CPU (Central Processing Unit) 20 which is a main control unit that controls the operation of the entire server 4, a ROM (Read Only Memory) 22 in which various programs and various data are stored in advance, a RAM (Random Access Memory) 24 used as a work area when the CPU 20 executes various programs, and an HDD (Hard Disk Drive) 26 as a storage means for storing various programs and various data.
[0032] The HDD 26 stores the singing video data transmitted from the mobile terminal 3A, that is, the singing video data uploaded to the video posting site, the music data indicating the music that the singing user can sing, and the like. Note that the storage means is not limited to the HDD 26, and may be other storage media such as a semiconductor memory such as a flash memory.
[0033] Furthermore, the server 4 includes an operation input unit 28 composed of a keyboard and a mouse and the like for receiving inputs of various operations, a monitor 30 such as a liquid crystal display device for displaying various images, and is connected to other information processing devices such as the mobile terminal 3 via the communication line 2, and has an external interface 32 for transmitting and receiving various data to and from other information processing devices.
[0034] These CPU 20, ROM 22, RAM 24, HDD 26, operation input unit 28, monitor 30, and external interface 32 are electrically connected to each other via a system bus 34. Therefore, the CPU 20 can access the ROM 22, RAM 24, and HDD 26, grasp the operation state of the operation input unit 28, display an image on the monitor 30, and transmit and receive various data to and from other information processing devices via the external interface 32.
[0035] [3. Electrical Configuration of Mobile Terminal] FIG. 3 is a functional block diagram showing the electrical configuration of the mobile terminal 3.
[0036] In addition to the configuration shown in FIG. 1, the mobile terminal 3 includes a main control unit 40, a main memory unit 42, an auxiliary memory unit 44, a communication unit 46, and operation buttons 48.
[0037] The main control unit 40 is, for example, a CPU, a microprocessor, a DSP (Digital Signal Processor), etc., and controls the overall operation of the mobile terminal 3.
[0038] The main memory unit 42 is composed of, for example, a RAM or a DRAM (Dynamic Random Access Memory), etc., and is used as a work area, etc., when executing processing based on various programs by the main control unit 40.
[0039] The auxiliary memory unit 44 is, for example, a non-volatile memory such as a flash memory, and stores various data such as images and programs used for the processing of the main control unit 40. The programs stored in the auxiliary memory unit 44 are, for example, an OS (Operating System) for realizing the basic functions of the mobile terminal 3, drivers for controlling various hardware, programs for realizing functions such as e-mail and web browsing, and other various functions. In addition, in the auxiliary memory unit 44, there is stored an application (hereinafter referred to as the "video posting and viewing application").
[0040] The communication unit 46 is, for example, a NIC (Network Interface Controller), and has a function of connecting to the communication line 2. Note that the communication unit 46 may have a function of connecting to a wireless LAN (Local Area Network), a function of connecting to a wireless WAN (Wide Area Network), a function of enabling short-range wireless communication such as Bluetooth (registered trademark), and infrared communication, instead of or together with the NIC.
[0041] The operation button 48 is provided on the side surface of the mobile terminal 3 and is a power button for starting or stopping the mobile terminal 3, a volume adjustment button for the sound output from the speaker 3b, etc.
[0042] These main control unit 40, main memory unit 42, auxiliary storage unit 44, communication unit 46, operation button 48, touch panel display 3a, speaker 3b, microphone 3c, camera 3d, and earphone terminal 3e are electrically connected to each other via the system bus 49. Therefore, the main control unit 40 can access the main memory unit 42 and the auxiliary storage unit 44, display an image on the touch panel display 3a, grasp the operation state of the touch panel display 3a and the operation button 48 by the user, input sound to the microphone 3c, output sound from the speaker 3b or the earphone connected to the earphone terminal 3e, control the camera 3d, and access various communication networks and other information processing devices via the communication unit 46, etc.
[0043] [4. Shooting of a singing video by a singing user] The case where a singing user shoots a singing video using the mobile terminal 3A will be described.
[0044] When shooting a singing video, the singing user starts a video posting and viewing application on the mobile terminal 3A. When the video posting and viewing application starts, the mobile terminal 3A accesses a server 4 that stores a plurality of music data. Then, the singing user arbitrarily selects a music for singing by himself / herself from the video posting and viewing application and downloads the music data from the server 4 to the mobile terminal 3A. Then, the singing user uses the video posting and viewing application to play the music at an arbitrary timing and sing. The video posting and viewing application starts playing the music and at the same time starts shooting a video by the camera 3d. That is, the singing video is a video shot by the mobile terminal 3A while the music is being played from the mobile terminal 3A.
[0045] Note that lyric data is also associated with the music data, and when the music data is downloaded from the server 4 to the mobile terminal 3A, the associated lyric data is also downloaded to the mobile terminal 3A. In the following description, it is assumed that the music data includes lyric data.
[0046] FIG. 4 is an example of a display state (hereinafter referred to as "screen display") on the screen 3a of the mobile terminal 3A when shooting a singing video.
[0047] As shown in FIG. 4, the screen 3a is divided into a lyric display area 50A and a captured image display area 50B. The lyric display area 50A includes a lyric image 52 showing the lyrics of the music that the user sings, a pitch image 54 showing the pitch of the music, and a progress bar 56 showing the progress of shooting.
[0048] The lyric image 52 and the pitch image 54 are updated according to the progress of the music. In this embodiment, as an example, the lyric image 52 and the pitch image 54 are updated by several phrases and displayed in the lyric display area 50A. Note that the update timings of the lyric image 52 and the pitch image 54 may be the same or different.
[0049] As an example, the lyric image 52 displays the lyrics in multiple lines (two lines in the example of FIG. 4), and the color of the upper-line lyrics changes from the left end to the right end according to the progress of the music so that the singing user can grasp the lyrics to be sung currently. When the color change of the upper-line lyrics reaches the right end, the lower-line lyrics rise and are displayed in the upper line, and new lyrics are displayed in the lower line. Then, the color of the upper-line lyrics changes from the left end to the right end again according to the progress of the music.
[0050] As an example, in the pitch image 54, a plurality of pitch bars 54A are displayed in a stepped manner in the left-right direction according to the strength of the pitch. Then, so that the singing user can grasp the pitch of the lyrics to be sung currently, the color of the pitch bar 54A changes from the left end to the right end according to the progress of the music, and the pointer 54B moves from the left end to the right end. When the color change of the pitch bar 54A and the pointer 54B reach the right end, a pitch image 54 showing the next pitch is updated and displayed.
[0051] The progress bar 56, as an example, has a length from the left end to the right end indicating the length of the entire piece of music. When the playback of the music starts, a pointer 56A indicating the playback position of the music moves from the left end to the right end, and when the pointer 56A reaches the right end, the music ends. In addition, the progress bar 56 passed by the pointer 56A is displayed thicker than the previous position.
[0052] The recording of the singing video starts after a predetermined time (for example, 10 seconds later) after the singing user clicks a recording start button (not shown) displayed on the screen 3a after selecting a piece of music. Also, the start and end of the video recording may coincide with the start and end of the music, but it is not limited to this. The video recording may start a predetermined time before the start of the music (for example, 5 seconds before), or the video recording may end a predetermined time after the end of the music (for example, 5 seconds after).
[0053] The singing user connects earphones to the earphone terminal 3e and listens to the music played using the earphones and sings along with the music. The mobile terminal 3A captures the singing user with the camera 3d and records the singing of the singing user with the microphone 3c. That is, the microphone 3c does not acquire the sound of the music being played. Then, the mobile terminal 3A records the singing voice of the singing user acquired by the microphone 3c as singing data.
[0054] Note that the singing data may be obtained by extracting the frequency band of the human voice through filtering processing. With this filtering processing, noise caused by the surrounding environment of the singing user is removed from the singing data, so that the singing voice of the recorded singing user becomes clearer.
[0055] Then, the video submission viewing application combines the music data and the singing data with the recording data to obtain singing video data that can be transmitted to the server 4. Note that the user can select one of the following two types as the timing of transmitting the singing video data to the server 4, that is, the timing of uploading to the video submission site.
[0056] One is a live stream where a singing user uploads singing video data to a video posting site in real time while singing. In a live stream, viewing users will be able to watch the singing by the singing user in real time. The other is a non-live stream where, after the singing of a song is completed, the singing user uploads the singing video data to the video posting site at an arbitrary timing.
[0057] When a singing user conducts a live stream, the user makes settings for conducting the live stream before recording the singing video, so that the singing video data is uploaded to the video posting site along with the start of video recording. In the case of a live stream, the singing video data may be uploaded to the video posting site without being stored in the mobile terminal 3A.
[0058] As settings for conducting a live stream, either a first live stream setting that enables viewing users to view the singing video during the live stream or a second live stream setting that enables viewing users to view the singing video even after the live stream can be set by the singing user. That is, in the first live stream setting, the singing video data is deleted from the server 4 when the live stream ends, and viewing users cannot view the singing video live-streamed after the end of the live stream. On the other hand, in the second live stream, since the server 4 continues to store the singing video data even after the live stream ends, viewing users can view the singing video as a non-live stream even after the end of the live stream.
[0059] In the case of a non-live stream, the singing video data is temporarily stored in the mobile terminal 3A, and the singing user uploads the singing video to the video posting site at an arbitrary timing by operating the video posting viewing application.
[0060] [5. Viewing of Singing Videos by Viewing Users] The case where a viewing user views a singing video using the mobile terminal 3B will be described.
[0061] When a viewing user watches a singing video, the user launches a video posting and viewing app on the mobile terminal 3B. When the video posting and viewing app is launched, the mobile terminal 3B accesses a server 4 that stores a plurality of singing video data, that is, a video posting site. Then, the viewing user selects a singing video to be viewed via the video posting and viewing app and displays it on the screen 3a. As an example of the method for delivering the singing video from the server 4 to the mobile terminal 3B, it is streaming delivery.
[0062] FIG. 5 is an example of the screen display of the mobile terminal 3B when watching a singing video, and shows the screen display when live delivery is being performed.
[0063] On the screen 3a, a singing video is displayed, and a singing user display area 50C, a lyrics display area 50D, and a comment input display area 50E are provided. The singing user display area 50C, the lyrics display area 50D, and the comment input display area 50E may be displayed superimposed on the singing video.
[0064] In the singing user display area 50C, the user name of the singing user who posted the singing video, a display indicating whether it is a live delivery, and the name of the song being sung are displayed.
[0065] In the lyrics display area 50D, the lyrics of the singing video are displayed. As an example, the displayed lyrics are in multiple phrases, and the color of the lyrics changes from the left end to the right end in accordance with the progress of the song. As an example, the lyrics display area 50D may display the lyrics in multiple lines. In this case, when the color change of the lyrics in the upper line reaches the right end, the lyrics in the lower line rise and are displayed in the upper line, and new lyrics are displayed in the lower line, and the color of the lyrics in the upper line changes from the left end to the right end again in accordance with the progress of the song.
[0066] In the comment input display area 50E, an input field for comments is displayed, and comments from the viewing users watching the singing video are displayed together with the user names. As an example, every time a comment is input from a viewing user, the comment is additionally displayed at the uppermost row of the comment input display area 50E, and the comments displayed until then are scrolled downwards. And when the comments cannot be fully displayed in the comment input display area 50E, a scroll bar (not shown) is displayed on the right side of the comment display area, and by operating the scroll bar by the viewing user, comments that have not been displayed on the screen 3a until then are displayed.
[0067] Furthermore, on the screen 3a, operation icons 58A to 58D for the viewing user to perform various operations are displayed.
[0068] The operation icon 58A is an icon that is clicked when the viewing user empathizes with the singing video being watched. The total number of clicks of the operation icon 58A for the singing video is displayed above the operation icon 58A.
[0069] The operation icon 58B is an icon that is clicked when the viewing user applies for a battle (hereinafter referred to as "battle singing") against the singing user who is live-streaming the singing video displayed on the screen 3a. Battle singing simultaneously displays a plurality of singing videos (the first singing video, the second singing video) by different singing users on the screen 3a of the viewing user's mobile terminal 3B, and the singing videos sing the same song alternately. That is, the viewing user who clicks the operation icon 58B becomes the singing user who performs battle singing.
[0070] The operation icon 58C is an icon that is clicked when the viewing user makes various settings for the video posting viewing application.
[0071] The operation icon 58D is an icon that is clicked by the viewing user when superimposing a decoration image on the singing video displayed on the screen 3a. Note that the decoration image according to this embodiment has a price determined by its type and can be purchased by the viewing user through payment. Then, by clicking the operation icon 58D, the viewing user superimposes a decoration image on the singing video that they are viewing. The singing user of the singing video with the superimposed decoration image receives money corresponding to the superimposed decoration image from the operator of the video posting site. That is, the superimposition (display instruction) of the decoration image on the singing video by the viewing user corresponds to what is called a tip for the singing user.
[0072] [6. Video synthesis function] The video posting and viewing app according to this embodiment has a video synthesis function for synthesizing a plurality of singing videos. The video synthesis function is a function of synthesizing a second recorded singing video with a first recorded singing video. Note that the songs sung in the first singing video and the second singing video are the same song.
[0073] [6-1. Synthesis start position] FIG. 6 is a schematic diagram showing the content of the video synthesis function according to this embodiment. The video synthesis function according to this embodiment determines a synthesis start position for starting the synthesis of a second recorded singing video with a first recorded singing video (FIG. 6(A)). After determining the synthesis start position, the second singing video is recorded with the camera 3d (FIG. 6(B)), and the first singing video and the second singing video are synthesized at the synthesis start position of the first singing video (FIG. 6(C)). Note that the positions used for the synthesis start position, playback position, etc. are represented by, for example, the elapsed time from the start of the singing video or the song.
[0074] With such a video synthesis function, when a singing user records a singing video (first singing video) but feels that the result is not satisfactory, the user sets the synthesis start position to a position before the video part that is considered unsatisfactory, and then reshoots the singing video (second singing video) from the synthesis start position. By synthesizing the first singing video and the second singing video at the synthesis start position of the first singing video, the singing user can easily create a singing video that they feel is desirable.
[0075] As described above, the video synthesis function according to this embodiment starts recording the second singing video after determining the synthesis start position for the first singing video. Therefore, the video synthesis function starts recording the second singing video from the playback position of the music corresponding to the synthesis start position in the first singing video. Accordingly, the singing user does not need to start singing the music from the beginning when recording the second singing video and can start from the synthesis start position, so that the synthesis of the singing video can be performed more simply.
[0076] FIG. 7 is an image displayed on the screen 3a of the mobile terminal 3A that a singing user operates when determining the synthesis start position. The screen 3a is divided into a playback control area 50F and a playback video display area 50G. The playback video display area 50G displays the first singing video to be played back. The playback control area 50F displays a lyric image 52 showing the lyrics of the music of the first singing video, a pitch image 54 showing the pitch of the music, and a slide bar 60 for selecting the playback position of the video.
[0077] The left end of the slide bar 60 indicates the start of the first singing video, and the right end indicates the end of the first singing video. When the singing user moves the pointer 60A left and right, the playback position of the first singing video displayed on the screen 3a changes accordingly, and the first singing video is displayed on the screen 3a in a paused state. Also, as the pointer 60A moves, the lyric image 52 and the pitch image 54 are updated.
[0078] When the singing user clicks the determination button 62, the playback position of the first singing video displayed in the playback video display area 50G is determined as the synthesis start position.
[0079] Note that the position that can be set as the synthesis start position may be a predetermined position of the music piece. This predetermined position is, for example, a position where it is easy to start singing from the middle of the lyrics, such as the boundary between phrases of the lyrics, and a plurality of positions are set. In other words, the singing user cannot select a position where it is difficult to start singing, such as in the middle of a phrase, as the synthesis start position. Note that a plurality of positions that can be set as the synthesis start position may be displayed on the slider 60 and the lyric image 52 so that the singing user can recognize them.
[0080] Also, not limited to the example of FIG. 7, buttons for playing (resuming) the first singing video, buttons for stopping the playback, buttons for fast-forwarding or rewinding, buttons for slow playback, etc. may be displayed on the screen 3a.
[0081] [6-2. Recording of the Second Singing Video] When the synthesis start position is determined, the video synthesis function records the second singing video.
[0082] As shown in FIG. 8, the video synthesis function according to the present embodiment has a guide function of displaying a synthesis start position image 64 (still image), which is an image at the synthesis start position of the first singing video, on the screen 3a at the start of recording of the second singing video. The synthesis start position image 64 is displayed thinly as an example so that the singing user can recognize that the image displayed on the screen 3a is the synthesis start position image 64. Then, as shown in FIG. 9, the video synthesis function superimposes the synthesis start position image 64 on the image (hereinafter referred to as "current captured image") captured by the camera 3d for recording the second singing video and displays it on the screen 3a.
[0083] In this way, the guidance function causes the singing user to inevitably check the image at the synthesis start position of the first singing video at the start of recording the second singing video by displaying the synthesis start position image 64 on the screen 3a. Therefore, the singing user can have the same image composition, pose, and expression as the synthesis start position in the first singing video at the start of recording the second singing video. As a result, a synthesis with less discomfort can be achieved in the synthesis of the first singing video and the second singing video. Further, as shown in FIG. 9, by superimposing the synthesis start position image 64 and the current captured image, the singing user can more easily make the image at the start of recording the second singing video the same as the synthesis start position image 64.
[0084] The guidance function may include a detection function for detecting, for example, the coincidence rate between the image of the singing user at the synthesis start position of the first singing video and the image of the singing user at the start of recording the second singing video. The guidance function may include a function of notifying the user when the coincidence rate between the image of the singing user at the synthesis start position of the first singing video and the image of the singing user at the start of recording the second singing video is lower than a predetermined threshold value.
[0085] In the examples of FIGS. 8 and 9, the synthesis start position image 64 is displayed over the entire playback video display area 50G on the screen 3a. However, the present invention is not limited to this, and the synthesis start position image 64 may be window-displayed and superimposed on a part of the playback video display area 50G. Then, the window-displayed synthesis start position image 64 and the current captured image may be superimposed and displayed.
[0086] Next, with reference to FIG. 10, the recording of the second singing video will be described. FIG. 10 is a schematic diagram showing the display timing of the synthesis start position image 64 on the screen 3a and the superimposed synthesis of the first singing video and the second singing video, the details of which will be described later. The bar indicated by hatching shows the image displayed on the screen 3a when the second singing video is recorded.
[0087] As shown in FIG. 10, the synthesis start position image 64 is displayed as a still image on screen 3a before the recording of the second singing video starts, and the display on screen 3a stops when the recording of the second singing video starts.
[0088] Then, the video synthesis function starts playing the music from before the synthesis start position, and starts recording the second singing video when the played music reaches the timing corresponding to the synthesis start position. The video synthesis function according to the present embodiment starts playing the music, for example, 30 seconds before the synthesis start position as an example. Thereby, the singing user can record the second singing video while matching the singing start of the second singing video with the timing of the recording start with a margin. Further, the video synthesis function starts playing the music before the recording of the second singing video starts, and also performs a countdown display for the start of the recording of the second singing video (see FIG. 9). Thereby, the singing user can clearly recognize the timing of the singing start of the second singing video.
[0089] Furthermore, as shown in FIG. 10, after the display of the synthesis start position image 64 stops and the recording of the second singing video starts, as a guide function, the first singing video may be superimposed and displayed on screen 3a for a predetermined period from the synthesis start position. Thereby, the singing user can record the second singing video while matching the image composition, pose, and expression with the first singing video, so that it is possible to synthesize the first singing video and the second singing video with less sense of incongruity. Note that this predetermined period is, for example, 5 seconds, which is an overlapping synthesis period to be described in detail later as an example, but the overlapping synthesis period may be set based on the number of frames instead of the time as a reference unit.
[0090] Here, with reference to FIG. 10, the overall flow of the recording of the second singing video will be described.
[0091] When the synthesis start position is determined by the operation of the singing user, the synthesis start position image 64 is displayed on the screen 3a and the shooting by the camera 3d is started, and the synthesis start position image 64 and the current shooting image are superimposed and displayed on the screen 3a. At this time, the same music data as the music of the first singing video is downloaded from the server 4 to the mobile terminal 3A. Then, when the singing user inputs an instruction to start recording the second singing video to the mobile terminal 3A, the recording countdown starts along with the playback of the music.
[0092] When the recording countdown ends, the recording of the second singing video starts and the display of the synthesis start position image 64 stops. Then, the first singing video during a predetermined period (superimposed synthesis period) from the synthesis start position is superimposed on the current shooting image and displayed on the screen 3a. Note that the recorded second singing video does not include the first singing video that is superimposed and displayed. When the predetermined period (superimposed synthesis period) has elapsed, the display of the superimposed first singing video stops, and the recording of the second singing video continues until the music ends.
[0093] [6-3. Synthesis of the First Singing Video and the Second Singing Video] When the recording of the second singing video ends, the video synthesis function synthesizes the first singing video and the recorded second singing video at the synthesis start position of the first singing video. In the following description, the video obtained by synthesizing the first singing video and the second singing video is referred to as a synthesized singing video.
[0094] The video synthesis function may perform image processing (hereinafter referred to as "effect processing") to make the synthesis part of the first singing video and the second singing video less noticeable. The effect processing is, for example, a predetermined period before and after the synthesis part (for example, 5 seconds before and after, or 5 frames before and after). In the example of FIG. 11, a plurality of predetermined images (star images) are randomly scattered as the effect processing, but it is not limited to this, and other image processing such as mosaic processing or blurring processing may be performed.
[0095] Due to this effect processing, the boundary between the first singing video and the second singing video becomes unclear, so that it is possible to synthesize the first singing video and the second singing video with less discomfort.
[0096] Also, as an effect process, the entire image of the composite part may be instantaneously replaced with a predetermined color. The predetermined color is, for example, white or black. For example, by setting the predetermined color to white (or black), a white (or black) jump is intentionally caused on the screen, making the boundary between the first singing video and the second singing video unclear.
[0097] Also, as shown in FIG. 10, the video composite function may superimpose the first singing video and the second singing video for a predetermined period (the above-described superimposed composite period) from the composite start position to composite (hereinafter referred to as "superimposed composite") the first singing video and the second singing video. This superimposed composite is an image process such as a so-called cross dissolve or overlap. For example, while fading out the first singing video to be superimposed with time, the second singing video may be faded in. Due to this superimposed composite, the boundary between the first singing video and the second singing video becomes unclear, enabling a more seamless composite of the first singing video and the second singing video.
[0098] Note that the video composite function according to the present embodiment does not perform superimposed composite on the recorded singing voice, but is not limited thereto, and superimposed composite may also be performed on the singing voice.
[0099] [7. Functional Blocks of Video Composite Function] FIG. 12 is a functional block diagram related to the video composite function according to the present embodiment. The main control unit 40 included in the mobile terminal 3 includes an image display control unit 70, a recording control unit 72, a composite start position determination unit 74, and a video composite unit 76. The processes executed by the respective functions included in the main control unit 40 are realized by a program stored in the auxiliary storage unit 44.
[0100] The image display control unit 70 controls the display of images on the screen 3a. For example, when the video posting and viewing application is launched, it causes the video distributed from the server 4 to be displayed on the screen 3a, or causes the image captured by the camera 3d to be displayed on the screen 3a. Note that, when shooting the second singing video, the image display control unit 70 according to the present embodiment superimposes the synthesis start position image 64 and the first singing video for a predetermined period from the synthesis start position, which has been captured by the camera, on the image being captured, and then displays the result on the screen 3a.
[0101] The recording control unit 72 records the image captured by the camera 3d. Note that, the recording control unit 72 according to the present embodiment plays a music piece from the mobile terminal 3A by launching the video posting and viewing application, and records the first singing video or the second singing video together with the sound of the played music piece. Further, when recording the second singing video, the recording control unit 72 starts playing the music piece before the synthesis start position and performs a countdown for starting the recording of the second singing video. When the timing at which the played music piece reaches the position corresponding to the synthesis start position is reached, the recording of the second singing video is started.
[0102] The synthesis start position determination unit 74 determines the synthesis start position for starting the synthesis of the second singing video to be recorded later with respect to the first singing video recorded earlier, based on the operation of the mobile terminal 3A by the singing user.
[0103] The video synthesis unit 76 synthesizes the first singing video and the second singing video at the synthesis start position of the first singing video to obtain a synthesized singing video. Further, the video synthesis unit 76 according to the present embodiment performs image processing for making the synthesis portion of the first singing video and the second singing video less prominent, or superimposes the first singing video and the second singing video for a predetermined period (superimposed synthesis period) from the synthesis start position to synthesize the first singing video and the second singing video.
[0104] [8. Flowchart of Video Synthesis Processing] FIG. 13 is a flowchart showing the flow of video synthesis processing executed by the main control unit 40 provided in the mobile terminal 3. A program (video posting and viewing application) for executing the video synthesis processing is stored in advance in a predetermined area of the auxiliary storage unit 44.
[0105] First, in step S100, the composition start position determination unit 74 receives the selection of the first singing video by the singing user. The first singing video is stored in the auxiliary storage unit 44 of the mobile terminal 3A. Note that the singing user may download the singing video uploaded by himself / herself from a video posting site and use it as the first singing video.
[0106] In the next step S102, the image display control unit 70 displays the first singing video on the screen 3a as shown in FIG. 7, and the composition start position determination unit 74 determines the composition start position based on the operation on the screen 3a by the singing user.
[0107] In the next step S104, the recording control unit 72 determines whether an input for instructing the start of shooting the second singing video is received. If the determination is affirmative, the process proceeds to step S106. If the determination is negative, the recording control unit 72 waits until an instruction to start shooting the second singing video is input.
[0108] In step S106, the image display control unit 70 displays the composition start position image 64 and the captured image (current captured image) by the camera 3d on the screen 3a.
[0109] In the next step S108, the recording control unit 72 determines whether an input for instructing the start of recording the second singing video is received. If the determination is affirmative, the process proceeds to step S110. If the determination is negative, the recording control unit 72 waits until an instruction to start recording the second singing video is input.
[0110] In step S110, the recording control unit 72 starts playing the music from before the composition start position. When the played music reaches the timing corresponding to the composition start position, the recording of the second singing video is started, and when the music ends, the recording of the second singing video is ended. The recorded second singing video is stored in the auxiliary storage unit 44.
[0111] In step S112, the video synthesizing unit 76 synthesizes the first singing video and the second singing video at the synthesis start position of the first singing video to generate a synthesized singing video. Note that superimposing synthesis may be performed as the synthesis of the first singing video and the second singing video. The synthesized singing video is stored in the auxiliary storage unit 44.
[0112] In the next step S114, the video synthesizing unit 76 performs an effect process on the synthesized singing video to end this video synthesis process. Note that the type of the effect process to be executed is preset by the singing user.
[0113] Then, the singing user uploads the synthesized singing video generated in this way to a video posting site via a video posting viewing application, making the synthesized singing video viewable by viewing users. Note that when the first singing video has been uploaded to the video posting site, the first singing video may be replaced with the synthesized singing video by the upload. Note that when the singing user is not satisfied with the degree of completion of the generated synthesized singing video, the video synthesis process shown in FIG. 13 is performed again from the beginning to generate the synthesized singing video again.
[0114] As described above, the mobile terminal 3A that records the image captured by the camera 3d as a video determines the synthesis start position for starting the synthesis of the second singing video to be recorded later with respect to the first singing video recorded earlier, and then records the second singing video with the camera 3d, and synthesizes the first singing video and the second singing video at the synthesis start position. Therefore, when the singing user records a video using the camera 3d, the mobile terminal 3A can easily create a video that the singing user feels desirable.
[0115] [9. Other Embodiments] As described above, the present invention has been described using the above embodiments, but the technical scope of the present invention is not limited to the scope described in the above embodiments. Various changes or improvements can be made to the above embodiments without departing from the gist of the invention, and the forms to which such changes or improvements are added are also included in the technical scope of the present invention. Also, the above embodiments may be combined as appropriate.
[0116] For example, in the above embodiment, the case where the performance is singing and the video uploaded to the video posting site is a singing video has been described. However, the present invention is not limited to this. For example, the performance may be other than singing, such as dance, and the video uploaded to the video posting site may be a dance video.
[0117] Also, in the above embodiment, the case where the video synthesis function is executed by the mobile terminal 3A has been described. However, the present invention is not limited to this. For example, part or all of the video synthesis function may be executed by the server 4. Even when part or all of the video synthesis function is executed by the server 4, the shooting and recording of the video are performed by the camera 3d of the mobile terminal 3A.
[0118] Also, the video synthesis function may generate a new first singing video by synthesizing a first singing video and a second singing video, and then generate a new synthesized singing video by synthesizing a new second singing video with the first singing video.
[0119] Also, the flow of the video synthesis process described in the above embodiment is also an example, and unnecessary steps may be deleted, new steps may be added, or the processing order may be changed within the scope not departing from the gist of the present invention.
Explanation of Reference Numerals
[0120] 3 Mobile terminal (information processing device) 3a Screen 3d Camera 64 Composite start position image 70 Image display control unit (image display control means) 72 Recording control unit (recording control means) 74 Composite start position determination unit (determination means) 76 Video synthesis unit (video synthesis means)
Claims
1. An information processing device that records images captured by a camera as video, a determining means for determining a combination start position at which a second moving image to be recorded later should be combined with a first moving image to be recorded earlier; a recording control means for recording the second moving image with the camera; and a moving image synthesizing means for synthesizing the first moving image and the second moving image at the synthesis start position of the first moving image.
2. 2 . The information processing apparatus according to claim 1 , further comprising an image display control means for displaying a combination start position image, which is an image of the first moving image at the combination start position, on a screen when recording of the second moving image starts.
3. 3 . The information processing apparatus according to claim 2 , wherein the image display control means displays, on the screen, the synthesis start position image superimposed on an image captured by the camera for recording the second moving image.
4. 4. The information processing apparatus according to claim 1, wherein the moving image synthesizing means performs image processing to make a synthesized portion of the first moving image and the second moving image less noticeable.
5. 5 . The information processing device according to claim 1 , wherein the moving image synthesizing means synthesizes the first moving image and the second moving image by superimposing the first moving image and the second moving image for a predetermined period from the synthesis start position.
6. The information processing device according to claim 1 , wherein the first moving image and the second moving image are moving images recorded while a piece of music is being played back.
7. 7. The information processing apparatus according to claim 6, wherein the recording control means starts playing the music before the synthesis start position, and starts recording the second moving image when the played music reaches a timing corresponding to the synthesis start position.
8. 8. The information processing apparatus according to claim 6, wherein the recording control means starts playing the music before the recording of the second moving image starts, and performs a countdown to the start of recording of the second moving image.
9. 9. The information processing apparatus according to claim 6, wherein the position that can be set as the synthesis start position is a predetermined position in the music piece.
10. A first step of recording a first video with a camera; a second step of determining a synthesis start position at which synthesis of a second moving image to be recorded later starts with respect to the first moving image; a third step of recording the second video with the camera; and a fourth step of combining the first moving image and the second moving image at the combination start position of the first moving image.
11. A computer included in an information processing device that records images captured by a camera as video, a determining means for determining a combination start position at which a second moving image to be recorded later should be combined with a first moving image to be recorded earlier; a recording control means for recording the second moving image with the camera; a moving image compositing program for causing the device to function as moving image compositing means for compositing the first moving image and the second moving image at the composition start position of the first moving image;
Citation Information
Patent Citations
Method and device for AV data recording
JP2005293749A
Reproducing device
JP2006079748A
Image processing unit
JP2006222640A
Video audio recording apparatus and method
JP2007088932A
Imaging device, method for photographing moving image, and moving image photography control program
JP2008017377A