Video Generation System

The video generation system addresses the limitation of still images by processing captured video data to generate composite videos that follow user movements, enhancing the service experience with dynamic and personalized content.

JP7782213B2Active Publication Date: 2025-12-09DAI NIPPON PRINTING CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021181265
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-05
Publication Date
2025-12-09
Estimated Expiration
2041-11-05

AI Technical Summary

Technical Problem

Existing systems provide only still images from user photographs, limiting the scope of services offered by tourist spots and leisure facilities.

Method used

A video generation system that processes captured video data to create composite videos based on user movements, using a shooting device, memory unit, and image processing unit to synthesize material frames with shot video frames, determining key frames for synthesis and material frame selection.

Benefits of technology

Enables the creation of videos that follow user movements, enhancing the service experience by providing dynamic and personalized video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007782213000001
    Figure 0007782213000001
  • Figure 0007782213000002
    Figure 0007782213000002
  • Figure 0007782213000003
    Figure 0007782213000003
Patent Text Reader

Abstract

To process, according to the movement of the user, a moving image obtained by shooting a user.SOLUTION: A moving image generation system comprises: an imaging device which images a user to generate an imaged moving image; a storage unit which stores a plurality of material frames according to a frame rate of the imaged moving image and the maximum imaging time taken by the imaging device; and an image processing unit which generates a combined moving image by combining the material frame to an imaged moving image frame constituting the imaged moving image. The image processing unit extracts a first key frame starting combining of the material frame and a second key frame ending combining of the material frame from the plurality of imaged moving image frames constituting the shot moving image, and decides a material frame to be combined to the imaged moving image frame on the basis of the number of imaged moving image frames from the first key frame to the second key frame and the number of material frames stored in the storage unit.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a video generation system. [Background technology]

[0002] Tourist spots and leisure facilities offer services in which customers (tourists, visitors) are photographed by photographers or automatically using sensors, and the photographed images are then printed out. Conventionally, all photographed data was provided to customers as still images (printed materials), which limited the scope of the service. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-46162 Summary of the Invention [Problem to be solved by the invention]

[0004] An object of the present invention is to provide a video generation system that processes a video captured of a user in accordance with the user's movements. [Means for solving the problem]

[0005] The video generation system of the present invention comprises a shooting device that shoots a user and generates a shot video, a memory unit that stores a plurality of material frames according to the maximum shooting time of the shooting device and the frame rate of the shot video, and an image processing unit that synthesizes the material frames with the shot video frames that make up the shot video to generate a composite video, wherein the image processing unit extracts a first key frame that starts synthesizing the material frames and a second key frame that ends synthesizing the material frames from the plurality of shot video frames that make up the shot video, and determines the material frames to be synthesized with the shot video frames based on the number of shot video frames from the first key frame to the second key frame and the number of material frames stored in the memory unit. [Effects of the Invention]

[0006] According to the present invention, a video of a user can be processed in accordance with the user's movements. [Brief explanation of the drawings]

[0007] [Figure 1] 1 is a schematic configuration diagram of a video generation system according to an embodiment of the present invention. [Figure 2] FIG. 10 is a diagram showing an example of combining a captured video with a material. [Figure 3] FIG. 2 is a functional block diagram of the server device. [Figure 4] FIG. 10 is a diagram illustrating an example of material data. [Figure 5] FIG. 10 is a diagram illustrating an example of a key frame. [Figure 6] FIG. 10 is a diagram illustrating a method for combining a captured video with a material. [Figure 7] 7a and 7b are diagrams showing examples of materials. [Figure 8] 8a and 8b are diagrams showing examples of composite moving images. DETAILED DESCRIPTION OF THE INVENTION

[0008] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0009] 1, the photography system according to this embodiment includes a server device 1, a photography device 2, a display device 3, a reader 4, a control device 5, and a sales terminal 6. The photography device 2 photographs a user U, and provides the user U with a composite video by combining the photographed video with animation materials such as effects and characters (material video, hereinafter sometimes simply referred to as "material"). The photography device 2 is, for example, a digital camera.

[0010] This photography system is installed in leisure facilities and the like, and for example, the photography device 2 captures scenes of a user U (a visitor) enjoying an attraction. The attraction is an attraction in which the user U moves in a predetermined direction at a substantially constant speed, such as a roller coaster or zip line. For example, when a sensor (not shown) detects the user U's approach, the photography device 2 begins filming. The captured video data is transferred to the control device 5. Filming by the photography device 2 may end after a certain period of time has elapsed, or may end when another sensor detects that the user U has left (left the photography range).

[0011] The control device 5 is a computer with communication capabilities. The control device 5 displays the captured video on the display device 3. After enjoying the attraction, the user U checks the video of himself / herself displayed on the display device 3 and holds the ticket T, on which the two-dimensional code is printed, over the reader 4.

[0012] This two-dimensional code indicates identification information for identifying user U, and also indicates the URL of the website from which the composite video can be downloaded, as will be described later. For example, the last ten digits of the URL are alphanumeric characters for user identification. A ticket T is distributed to each user U upon entry to the leisure facility.

[0013] The reader 4 reads the two-dimensional code printed on the ticket T and transmits the identification information to the control device 5.

[0014] The control device 5 associates the captured video with the identification information and transmits them to the server device 1. The server device 1 stores the captured video received from the control device 5 in a storage folder corresponding to the identification information.

[0015] The two-dimensional code printed on the ticket T may be read before photographing the user U. After reading the two-dimensional code, the control device 5 associates the video data received from the photographing device 2 with the identification information and transmits them to the server device 1.

[0016] The server device 1 generates a composite video by combining elements with the shot video. For example, as shown in FIG. 2, each frame of the shot video is combined with an element frame corresponding to the position of the subject (user) in the frame to generate a composite video. This results in a composite video in which animation matching the movement of the subject is combined. The method for generating the composite video will be described in detail later.

[0017] The server device 1 stores the generated composite moving image in a storage folder corresponding to the identification information.

[0018] The user U holds the ticket T over the sales terminal 6. The sales terminal 6 reads the two-dimensional code printed on the ticket T and transmits the identification information to the server device 1. The server device 1 notifies the sales terminal 6 whether or not there is a composite video corresponding to the identification information received from the sales terminal 6.

[0019] When the sales terminal 6 receives a notification from the server device 1 that a composite video is available, it displays a video purchase screen and accepts the purchase of the composite video from the user U. The sales terminal 6 notifies the server device 1 that the composite video has been purchased. The server device 1 sets a downloadable flag for the composite video data notified of the purchase.

[0020] The user uses a user terminal 7 such as a smartphone to obtain the URL of the composite video download site from the two-dimensional code printed on the ticket T, and accesses the download site provided by the server device 1.

[0021] The server device 1 detects the identification information of the user U from the URL accessed by the user terminal 7, and transmits the composite moving image to the user terminal 7 if the composite moving image corresponding to the detected identification information is downloadable.

[0022] The user plays the downloaded composite video on the user terminal 7. The composite video is created by combining materials in accordance with the user's movements, and the user can enjoy the composite video.

[0023] Next, we will explain the configuration of the server device 1. The server device 1 is a computer having a CPU, a storage unit, a communication unit, etc., and the CPU executes a video generation program stored in the storage unit to realize the functions of a shooting data receiving unit 11, an image processing unit 12, a management unit 13, and a composite video transmitting unit 14, as shown in Fig. 3.

[0024] Materials to be combined with the captured video are stored in advance in the storage unit 10 of the server device 1. The materials are prepared in an amount corresponding to the frame rate and maximum capture time of the image capture device 2.

[0025] For example, in the attraction shown in Figure 1, if the frame rate of the camera 2 is 30 fps, and the time it takes for a user of any build to pass through the camera's shooting range is less than 20 seconds, and the maximum shooting time is 20 seconds, then 600 (= 30 fps x 20 seconds) raw frames (frames that make up the raw video) are prepared, as shown in Figure 4.

[0026] The 600 material frames are the same size as the captured video frames, and each frame features a different character posture and position. The character's position gradually moves in accordance with the user's movement direction in the attraction. For example, in an attraction where the user moves from right to left within the capture range of the camera 2, as shown in FIG. 4, from the first material frame F1 to the 600th material frame F600, a bird character gradually appears from the right edge and shifts position to the left by a predetermined amount.

[0027] The photographed data receiving unit 11 receives the photographed video and the identification information from the control device 5 and stores them in the storage unit 10.

[0028] The image processing unit 12 synthesizes the material with the captured moving image to generate a composite moving image.

[0029] First, the image processing unit 12 extracts from the captured video a first key frame (start frame) that is the first frame to be composited with the material, and a second key frame (end frame) that is the last frame to be composited with the material. The second key frame is a frame that comes after the first key frame.

[0030] For example, the first key frame is the first frame that includes the entire body of the subject (user), as shown in FIG.

[0031] Also, for example, the second key frame is the first frame after the first key frame in which part of the subject's body disappears from the frame, as shown in Fig. 5. The first frame in which the entire subject disappears from the frame (out of the frame) may be set as the second key frame.

[0032] The image processing unit 12 calculates the number of frames (video time) from the first key frame to the second key frame, and determines the material frames to be composited with each video frame according to the ratio between the calculated number of frames and the number of frames of the prepared material.

[0033] Specifically, if the number of frames from the first key frame to the second key frame is g and the number of material frames stored in the memory unit 10 is h, the material frames are synthesized with the captured video frames at h / g frame intervals.

[0034] For example, if the number of frames from the first key frame to the second key frame is 300 and the material stored in the memory unit 10 is 600 frames, the material frames are synthesized with the captured video frames at intervals of two frames (skipping one frame), as shown in Figure 6.

[0035] In this case, the first shot video frame (first key frame) is composited with the first material frame. The second shot video frame is composited with the third material frame. The third shot video frame is composited with the fifth material frame. The 300th shot video frame (second key frame) is composited with the 599th material frame.

[0036] This allows the character to be composited at a position corresponding to the subject (user) moving within the shooting range, as shown in Fig. 2. For example, a composite video is generated in which the character is composited to follow the user moving in a predetermined direction.

[0037] If the number of captured video frames from the first key frame to the second key frame is not a divisor of the number of material frames stored in the memory unit 10, the first key frame or the second key frame may be reset to be a divisor, and the number of captured video frames to be synthesized may be increased.

[0038] For example, if the number of shot video frames from the first key frame to the second key frame is 190 and the number of material frames stored in the memory unit 10 is 600, the first key frame or the second key frame is reset so that the number of shot video frames to be synthesized is 200, which is greater than 190 and is the closest divisor of the number of material frames to 190. This allows the material frames to be synthesized at equal intervals with respect to the shot video frames, making the movements of characters in the synthesized video more natural.

[0039] When the image processing unit 12 generates a composite moving image by combining the captured moving image frames with the material frames, the image processing unit 12 stores the composite moving image in the storage unit 10 in association with the identification information.

[0040] Management unit 13 manages whether or not a composite video has been purchased. When management unit 13 receives identification information from sales terminal 6, it notifies sales terminal 6 whether or not there is a composite video corresponding to this identification information. When management unit 13 receives notification from sales terminal 6 that a composite video has been purchased, it sets a downloadable flag for the composite video data.

[0041] When the user terminal 7 accesses the download site, the composite video sending unit 14 acquires identification information from the URL at the time of access, and if a downloadable flag is set for the composite video data corresponding to this identification information, it sends the composite video data to the user terminal 7.

[0042] In this way, according to this embodiment, a video of a user can be processed in accordance with the user's movements.

[0043] You can prepare multiple patterns of material to be combined with the video depending on the size (height, etc.) of the subject. For example, prepare material for adults as shown in Figure 7a and material for children as shown in Figure 7b.

[0044] The image processing unit 12 detects the height and head position of the subject in the captured video, selects material that matches the size of the subject, and composites it into the captured video. An example of the composite of the material shown in Figure 7a is shown in Figure 8a. An example of the composite of the material shown in Figure 7b is shown in Figure 8b. A character with a pose that matches the size of the subject is composited, and a high-quality composite video is generated.

[0045] In the above embodiment, an example has been described in which a material frame of the same size and in which the position of the character is fixed is synthesized with a shot video frame, but a material frame that does not have position information and is smaller in size than the shot video frame may also be synthesized. The material frames are prepared in an amount corresponding to the frame rate of the image capture device 2 and the shot video duration.

[0046] If the frame rate of the image capture device 2 is 30 fps and the time it takes for a user of any build to pass through the capture range of the image capture device 2 is less than 30 seconds, then 900 material frames are prepared. The nth material frame (n is an integer between 1 and 899) and the n+1th material frame have slightly different character posture (posture), facial expression, etc.

[0047] The image processing unit 12 detects key points within the captured video frames. Key points are, for example, predetermined parts of the subject, such as the face or hands. The image processing unit 12 determines the composition position of the material frames using relative coordinates from the key points, performs composition, and generates a composite video. This makes it possible to generate a composite video in which a character is composited at a certain distance from the subject.

[0048] When the server device 1 generates multiple composite videos for the same identification information, the server device 1 may also notify the sales terminal 6 of information on the shooting locations and shooting times of the original shot videos that were used for the synthesis.

[0049] The server device 1 may transmit sample data of the composite moving image to the sales terminal 6. By displaying the sample moving image, the sales terminal 6 can increase the user's desire to purchase.

[0050] The server device 1 may extract one or more frames from the composite video and transmit them to the sales terminal 6. When a user purchases the composite video, the sales terminal 6 can print out the frames (still images) received from the server device 1 from its built-in printer.

[0051] The sales terminal 6 may allow the user to select whether to purchase an unprocessed shot video or a composite video created by combining materials. In this case, the server device 1 may generate a composite video by combining the shot video and materials after the purchase process for the composite video has been completed at the sales terminal 6. This eliminates the need to generate a composite video that will not be purchased, thereby reducing the processing load on the server device 1.

[0052] Although the present invention has been described in detail with reference to specific embodiments, it will be apparent to those skilled in the art that various modifications can be made without departing from the spirit and scope of the invention. [Explanation of symbols]

[0053] 1. Server device 2. Imaging equipment 3 Display device 4 Reader 5. Control device 6 Sales terminals 7 User terminal

Claims

1. a photographing device that photographs a user and generates a photographed video; a storage unit for storing a number of material frames corresponding to the product of the maximum shooting time of the shooting device and the frame rate of the shot video; an image processing unit that generates a composite video by combining the material frames with captured video frames that constitute the captured video; Equipped with The image processing unit extracts a first key frame that starts compositing the material frames and a second key frame that ends compositing the material frames from a plurality of captured video frames that constitute the captured video, and determines the material frames to be composited with the captured video frames based on the number of captured video frames from the first key frame to the second key frame and the number of material frames stored in the memory unit, wherein the first key frame is the first frame that includes the entire user's body, and the second key frame is the first frame that does not include part of the user's body or the first frame that does not include the entire user's body, and if the number of captured video frames from the first key frame to the second key frame is not a divisor of the number of material frames, the first key frame or the second key frame is reset so that it does include a divisor, thereby increasing the number of captured video frames to be composited with the material frames.

2. The video generation system according to claim 1 , wherein the captured video frames and the material frames have the same size.

3. The material frame is smaller in size than the captured video frame, The video generation system according to claim 1 , wherein the image processing unit detects a predetermined part of the user in a captured video frame, and synthesizes the material frame at a position of predetermined relative coordinates with the predetermined part as a reference.

4. The video generation system according to claim 1 , wherein the image processing unit detects a height or a head position of the user in the captured video frames, and synthesizes material frames according to the detection results.

5. a control device that acquires the user's identification information via a reader, acquires the captured video from the photographing device, and transmits the identification information and the captured video to a server device having the storage unit and the image processing unit; The moving image generating system according to claim 1 , wherein the storage unit stores the captured moving image and the composite moving image in association with the identification information.

Citation Information

Patent Citations

  • Personal video capture system

    JP1997504928A

  • Time varying image data edit method and system

    JP1999177922A

  • Image distribution system, image distribution method, image distribution program, and recording medium storing the program and capable of being read by computer

    JP2003032647A

  • Moving image generation system, moving image display method, web server and program

    JP2013128328A

  • Synthetic video data generation system and program

    JP2016151975A