Still image generation system and still image generation method

The still image generation system addresses the challenges of remote shooting by generating high-quality, exclusive still images from decomposed video frames, ensuring viewer satisfaction and preventing unwanted image sharing.

JP2026135717APending Publication Date: 2026-08-25CAP CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025021391
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing remote shooting systems struggle with generating high-quality still images that meet diverse viewer demands, as they often include mis-shots, lack exclusive ownership, and fail to accommodate simultaneous shutter operations, especially when a large number of viewers participate.

Method used

A still image generation system that decomposes video into multiple still images, removes mis-shots, and generates high-quality still images based on viewer shutter operations, allowing exclusive ownership and preventing simultaneous image sharing.

Benefits of technology

Ensures high-quality still images are produced without mis-shots, enabling exclusive ownership and preventing unwanted image circulation, while providing a realistic shooting experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026135717000001_ABST
    Figure 2026135717000001_ABST
Patent Text Reader

Abstract

This is a still image generation service designed to provide a simulated photography experience. It ensures that no misfires are included regardless of shutter timing, and allows for the creation of products using these images through a series of online procedures. [Solution] When outputting still images based on captured video data, only still images provided to viewers are exported from shutter-ready still images, which are created by removing or correcting those treated as mis-shots from a set of multiple still images based on a predetermined frame rate. Therefore, no matter when the shutter is operated, the exported still images will not include mis-shots or unusable cuts. Even amateur viewers can take photos at the same level as professional photographers, and only still images permitted by the copyright management of the subject images can be provided to viewers. If each viewer wishes to keep the shots they have taken for themselves, they are guaranteed exclusive ownership that will not be shared with other viewers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a still image generation system and a still image generation method.

Background Art

[0002] In recent years, an application has emerged that allows users to receive and view moving images captured by a camera via a network, and send an instruction corresponding to pressing the shutter at a desired timing, thereby obtaining an experience as if photographing a subject right in front of them. Even if the viewer does not actually go to the shooting location, they can simulate the shooting action and obtain a sense of presence as if they were taking the picture themselves. Moreover, since it is a remote operation, it is possible to shoot from any physically distant location, and there is no risk of being unable to shoot due to a limit on the number of people at the shooting site. Therefore, the need for remote shooting is expected to increase in the future.

[0003] As prior art based on a similar concept, for example, there are the following Patent Documents 1-4. The shooting meeting system described in Patent Document 1 is configured such that when shooting conditions such as zooming in are sent to a shooting agent (human or device) that actually shoots a subject at the shooting site, the proxy shooter adjusts the shooting conditions based on the shooting instruction. The remote shooting system of Patent Document 2 does not use pre-recorded moving image data, but instead proposes a remote shooting means that distributes moving images shot at a live venue or the like to viewer terminals in real time via a server, and executes a finger action corresponding to a shutter operation on the viewer terminal for a favorite scene. The image data providing device described in Patent Document 3 stores moving image data in association with one or more frame data (or frame data), and while reproducing the moving image data by a display means, extracts and displays still image data corresponding to an arbitrary operation timing performed by the user. The remote shooting system described in Patent Document 4 is an invention filed by the applicant of the present application.

[0004] Furthermore, at real-world photo shoots, in addition to photobooks and portraits featuring the subject, various related products such as towels, T-shirts, and personal accessories are often sold. Merchandise is especially popular when the subject is an idol or celebrity. Products that are not available in regular stores and have limited distribution channels become branded items, often selling out completely. Even in the case of remote photo shoots, there is a high demand for items featuring images taken by the photographer themselves. [Prior art documents] [Patent Documents]

[0005] [Patent Document 1] Japanese Patent Publication No. 2003-209741 [Patent Document 2] Patent No. 7288641 [Patent Document 3] International Publication No. 2004 / 077830 [Patent Document 4] Patent Application No. 2024-187199 [Overview of the project] [Problems that the invention aims to solve]

[0006] The invention described in Patent Document 1 is basically intended for photo shoots with a very small audience. Therefore, the substitute photographer can accept requests from the audience and change the shooting conditions each time. However, since the still image generation system covered by the present invention is based on the premise of a remote photo shoot in which a large number of viewers participate, it is practically impossible to accept shooting conditions that match the diverse requests of each viewer. Unlike Patent Document 1, as the number of viewers participating in the remote shoot increases, it becomes impossible to proceed with the photo shoot if all viewers are asked to specify the subject to be photographed, the shooting angle, the magnification, etc.

[0007] The invention described in Patent Document 2 is based on the premise that a large number of viewers will participate, and each viewer can freely decide on their desired shooting scene and change the shooting angle, magnification, etc. However, there is a problem that viewers who are not experts in photography do not know when it is appropriate to take a picture, so the resulting still images are often not what was initially expected and are unsatisfactory.

[0008] The invention described in Patent Document 3 associates one or more still image data with video data, so a still image corresponding to the user's shutter operation is extracted. However, the video data is uniformly divided at predetermined time intervals (e.g., 1 / 100th of a second), and one frame corresponding to the shutter timing is generated as a still image. Therefore, there is a high possibility that the trajectory of the moving subject will be captured in the extracted still image. Similar to the system users in Patent Document 2, for viewers who are not experts in photography, there is a problem that when streaming videos of people, etc., the shutter may be pressed at the moment the subject's eyes are closed, or the trajectory of the moving subject, as described above, may be captured, resulting in blurry footage. These are essentially mis-shots that should be removed, and are not still images that viewers would want to obtain. Furthermore, copyright and related rights apply to the artist's photographs, images, and performances, and it is known that there is a strong desire to avoid, as much as possible, the exposure of unusable shots to the market, especially for major artists.

[0009] Furthermore, if multiple viewers' shutter timings are almost simultaneous, the same image will be provided to each viewer. With conventional still image generation systems used for remote photo shoots, it was not possible to exclusively own an image. Moreover, unlike mass-produced goods sold at actual photo shoot venues, the products purchased through this invention are made using the viewer's favorite image selected from the shots they themselves have taken. Although each viewer wants to obtain a one-of-a-kind product using the image they have taken and selected, since the image is shared with other viewers, it is not much different from conventional products that are mass-produced and circulate in the market, and thus fails to satisfy the desires of fans.

[0010] Therefore, the present invention aims to provide a remote shooting service that allows viewers to virtually experience the act of shooting without actually going to the shooting location, and that enables the exclusive acquisition of high-quality still images (photographs) that do not include any mis-shots or unusable cuts regardless of the shutter timing, and enables the production of products using those images through a series of online procedures. [Means for solving the problem]

[0011] To achieve the above objective, the present invention provides a still image generation system for providing a simulated shooting experience by distributing a video for viewing, created based on a video image captured by a shooting means, to multiple users' image display terminals via a communication network, characterized in that (a) the video image is decomposed into a plurality of still images based on a predetermined frame rate, and (b) a reconstructed still image is generated based on a still image obtained by removing some of the multiple still images or a still image obtained by modifying some of the multiple still images, thereby generating a still image for shutter output, and when it is determined that the plurality of users have performed a finger action equivalent to a shutter operation on the image display terminal while viewing the video for viewing, a still image corresponding to the timing of the shutter operation is identified from the still images for shutter output, and during or after viewing the video for viewing, it is displayed on the image display terminal so that it can be identified whether or not one or more of the identified still images can be exclusively purchased, and if any of the plurality of users select a still image that can be exclusively purchased, the other users will not be able to purchase the selected still image.

[0012] Furthermore, the present invention relates to a still image generation system for providing a simulated shooting experience by distributing a video for viewing, created based on captured video images acquired by a shooting means, to image display terminals of multiple users via a communication network, characterized in that, at predetermined time intervals from the start of acquisition of the captured video images, (a) the captured video images within the predetermined time interval are decomposed into a plurality of still images based on a predetermined frame rate, (b) a still image for shutter output is generated by generating a reconstructed still image based on a still image obtained by removing some of the still images from the plurality of still images or a still image obtained by modifying some of the still images, and (c) in parallel with the generation of the shutter output still image, the video for viewing is distributed to the image display terminals of the multiple users after a predetermined time has elapsed from the start of acquisition, and when a finger action equivalent to a shutter operation is detected, a still image corresponding to the timing of the shutter operation is identified from the shutter output still image, and during or after the acquisition of the captured video images, it is displayed on the image display terminal so that it can be identified whether or not one or more identified still images can be exclusively purchased, and if any of the multiple users select a still image that can be exclusively purchased, the other users will not be able to purchase the selected still image. [Effects of the Invention]

[0013] The still image generation system and method according to the present invention output still images based on high-quality video data captured and edited, so that low-resolution images like so-called screenshots are not output as shutter images. Furthermore, from a plurality of still images composed based on a predetermined frame rate, images that are treated as mis-shots or unacceptable cuts are removed, and only still images that have been corrected to be high-quality are written out for the viewer. Therefore, no matter when the shutter is operated, the written still images will not include mis-shot images or unacceptable cuts. Even amateur viewers without advanced shooting skills can take photos at the same level as professional photographers, and only still images that are permitted by the copyright management of the subject images can be provided to the viewer. In addition, if each viewer wishes to make the captured shot image their own, they can select the still image and send a purchase instruction to the still image generation system, ensuring exclusive ownership without being shared with other viewers.

Brief Description of the Drawings

[0014] [Figure 1] It is a diagram showing the overall configuration in one embodiment of the still image generation system. [Figure 2] It is a flowchart showing the generation process of the still image for shutter output. [Figure 3] FIG. 3(A) shows the relationship between the video for viewing and the frame image with the cut to be removed specified, and FIG. 3(B) is a diagram showing the relationship between the shutter timing and the still image to be acquired. [Figure 4] It is a diagram showing an example of the multi-angle screen displayed on the viewing terminal. [Figure 5] It is a diagram for explaining switching different display areas based on the same moving image. [Figure 6] It is a diagram showing an example of the shutter indicator. [Figure 7] It is a diagram showing an example of the screen when proceduring the exclusive purchase of an image. [Figure 8] It is a diagram showing the temporal deviation between the distribution of the viewing image and the generation of the still image for shutter output.

Embodiments for Carrying Out the Invention

[0015] Hereinafter, one embodiment of the still image generation system according to the present invention will be described in detail with reference to the drawings. In all the drawings for explaining the following embodiments, the same parts are generally denoted by the same reference numerals, and the repeated explanations thereof are omitted. Also, a part of the configuration that is not important for explanation in each drawing is omitted and shown: Needless to say, the present invention can be implemented in many different forms and is not limited to the description content disclosed below, and various changes, substitutions, permutations of the processing order, and omissions are possible without departing from the gist thereof.

[0016] FIG. 1 shows the overall configuration of the still image generation system 100. The still image generation system 100 is composed of the relationship among a photography event operation company 1, a platform operation company 2, and a plurality of users 3, and is configured to transmit and receive various data via a communication network 4 such as the Internet. In this embodiment, the moving image to be processed by the still image generation system 100 is taken at a shooting venue as an example for explanation, but it does not have to be an indoor space such as a shooting venue (including studios and live venues), and it may be taken at any outdoor location (for example, a filming location).

[0017] The photography event operation company 1 plans the date and time of the shooting event, the shooting location, models and artists to be subjects at the shooting venue, etc., and announces that a remote shooting event will be held through a homepage or SNS. The platform operation company 2 that has contracted with the photography event operation company 1 in advance is an entity that provides the platform of the still image generation system 100 and is responsible for all technical support necessary to realize the planned shooting event. Since the platform operation company 2 provides the platform to the photography event operation company 1 on an OEM basis, the shooting event is recognized by the user 3 as being carried out by the photography event operation company 1. Note that the still image generation system 100 shown in FIG. 1 treats the photography event operation company 1 and the platform operation company 2 as separate entities, but there is no change in the effects of the present invention even if the photography event operation company 1 and the platform operation company 2 constitute the still image generation system 100 as the same entity. Also, the event notice to the user may be made by the platform operation company 2 instead of the photography event operation company 1.

[0018] The filming event management company 1 accepts applications from users to participate in remote filming sessions via the communication network 4. The filming event management company 1 edits the filmed video and accompanying audio (hereinafter referred to as "filmed video") and saves the resulting video for viewing on an information processing device (for example, a server or a television editing machine; hereinafter referred to as "server"). The video for viewing uploaded from the server is provided to user 3 via the communication network 4. In this embodiment, platform management company 2 owns and manages the server, but it may be a server within the filming event management company 1, or it may be a virtual server on the cloud. Hereinafter, the server will also be referred to as a "cloud server".

[0019] At the shooting location, preparations are made for setting up one or more shooting devices 6 (hereinafter referred to as "camera 6") for shooting the subject, and for shooting video and recording audio. At the shooting location, a person may hold the camera 6 and shoot, or the camera 6 may be automatically controlled without human intervention by command signals from a server. In this embodiment, an example is shown in which the subject is shot with multiple cameras 6 (for example, including setting up multiple cameras around the subject, such as 180 degrees or 360 degrees), but shooting may also be done with a single camera 6.

[0020] While promotional videos and music videos are edited, the filmed video footage stored on the server in the present invention is also edited using a video / still image editor to process both video and still images (hereinafter, "still images" are also referred to as "still images"). The filmed video footage, edited as appropriate, is provided to user 3 as a video for viewing. There are various editing methods, but typical ones include multicam editing, which involves switching scenes while simultaneously playing footage shot from various angles with multiple cameras 6, noise reduction, overlaying arbitrary footage, and text insertion. Other editing processes include image stabilization and the addition of special visual and sound effects. In this embodiment, the captured video footage is edited and then generated to produce the viewing image. However, there are also cases where the captured video footage is distributed directly to viewers as the viewing image, which will be described later as a second embodiment.

[0021] Participant 3, a user participating in the photo shoot, possesses a viewing device 3-1 to 3-N (hereinafter referred to as "viewing device 3N") and will view the images via the communication network 4. When receiving the images for viewing, the viewing device 3N must be connected to the communication network 4. When the distribution of the images for viewing begins, the registered participant 3 will view the images on their respective viewing device 3N, and an icon 7, which corresponds to the shutter operation, will be displayed on the app screen of the viewing device 3N. Participant 3 will press icon 7 at any time to perform the so-called shutter operation.

[0022] By operating the shutter, participant 3 will obtain a still image (sometimes called a "frame image") of their favorite scene. The present invention is characterized by its method of exporting these still images, which will be explained below. First, it is necessary to prepare still images for shutter output, which are different from the video to be viewed, in advance before viewing. Figure 2 is a flowchart showing the process of generating still images for shutter output.

[0023] As shown in Figure 2, the video for viewing stored on the server is read (step S20). As mentioned above, the video for viewing has undergone various video editing, so a video is generated with the effects applied to the image and sound removed (step S21). Effects include video compositing, color expression, and echo processing, but specifically, they include adding environmental effects such as thunder and rain, applying blurring or bumpy effects to the finished image, or using special sounds such as scratches or horns. In step S21, the process of removing effects related to visual effects is performed (however, the process of removing effects related to sound is not excluded). These processes may be performed fully automatically or manually. In contrast, conventional still image generation systems do not bother to extract and provide high-resolution still images (frame images) if the LCD screen of the viewing terminal 3N is FHD (1920 x 1080 pixels). The still image generation system 100 of this embodiment shoots based on the image quality of ultra-high resolution / high-definition video such as 8K (7680 x 4320 pixels) or 4K (3840 x 2160 pixels). Even if the LCD screen of the viewing terminal 3N is FHD, in order to provide high-definition still images such as 8K or 4K to participant 3, a high-resolution video with effects removed is generated from the high-resolution viewing video.

[0024] Next, multiple still images based on a predetermined frame rate are output from the video generated in step S21 (step S22). In other words, the video is decomposed into multiple still images. If the streamed video is 4K and 30fps (a frame rate consisting of 30 images per second), each still image can be identified by the elapsed time obtained by adding 1 / 30 = 0.0333 seconds from the start time of the video streaming for viewing. In other embodiments, the effects may not be removed (i.e., step 21 may not be performed), and the process may proceed directly from step 20 to step 22.

[0025] Next, some still images are removed from the multiple frame images created in step S22 (step S23). The still images to be removed include those with closed eyes or inappropriate posing. Also, NG cut images that are not acceptable to the copyright holders of artists, etc., are removed. Hereafter, these will be collectively referred to as NG cuts. Figure 3(A) shows the relationship between an edited video for viewing and multiple frame images designated as NG cuts that should be removed. Figure 3(A) uses still images exported from a video for viewing with 30 frames per second and no effects removed as an example, with frames 4, 5, 11, 12, 13, 21, 22, and 27 designated as NG cuts due to reasons such as eyes being closed. The server stores a reconstructed still image composed of multiple frame images from which such NG cuts have been removed (step S24). Alternatively, instead of actually removing some still images that correspond to NG cuts, the reconstructed still image may be created by using software processing with flags or tables to identify which frames are NG cuts and simulate the removal of NG cuts.

[0026] In this embodiment, the selection of which frames to designate as NG cuts is done manually, but the selection of NG cuts may be automated using image recognition applications such as face recognition software or eye-tracking software, or an AI system (AI application).

[0027] Next, the reconstructed still image created in step S24 is retouched (step S25) to construct a still image for shutter output and save it to the server (step S26). Retouching includes, for example, image processing such as adjusting the skin tone of the subject, refining the skin texture, correcting facial or body parts such as enlarging the eyes, improving resolution, and even changing the background color or clothing color. It may also include processing to add and delete objects not included in the reconstructed still image, such as brightening only specific parts, removing unwanted objects, compositing absentees into a group photo, or compositing the background, and there are no particular restrictions on the type or content of retouching. Furthermore, to increase resolution, a low-resolution image may be input and processed using super-resolution techniques to improve the resolution limit. Examples of super-resolution methods include case-based (learning-type) super-resolution, fractal super-resolution, and reconstruction-type super-resolution.

[0028] The retouching and super-resolution processing described above can be applied fully automatically to all reconstructed still images, manually to each still image, or to some still images and then processed similarly using AI software for the remaining images. It should be noted that retouching and super-resolution processing are not mandatory; they should be performed only as needed.

[0029] When steps S21, S23, and S24 are executed using AI software via ASP, their operation methods are optimized based on the track record of the still image generation system 100.

[0030] After preparing the still images for shutter output in this manner, the viewing video is distributed to participant 3. The start time of the viewing video distribution is centrally managed by a server connected to the communication network 4, and the shutter operation time for each participant 3 is managed in relation to the start time of the distributed video. In other words, the elapsed time from the start time of the distributed video represents the shutter operation time. When the viewing video is distributed to multiple participants simultaneously, the server manages one common start time for all participants, but when the start times are staggered for each participant, the server manages the start time corresponding to each participant.

[0031] The time at which each participant 3 takes a shutter operation on the video received by their viewing device 3N during the photo shoot is transmitted to the server in real time and recorded on the server side. This time is treated as a UNIX® timestamp in seconds or milliseconds format, and based on the elapsed time starting from the start time of the distributed video (which is also recorded as a UNIX® timestamp), it is possible to identify the frame image corresponding to the time when participant 3 took a shutter operation.

[0032] Furthermore, even if latency delays occur due to the communication environment of the viewing terminal 3N, the still image generation system 100 of this embodiment maintains low latency so that communication for recording the shutter operation time is kept within milliseconds. Even if a participant operates the shutter while the video is being streamed, and a problem occurs such as being on the move or entering a building with a poor communication environment, causing the connection to the internet to be lost, the platform installed on the viewing terminal 3N for video streaming temporarily caches the shutter operation time on the viewing terminal side, and automatically resynchronizes once the connection is restored. Therefore, measures are taken to ensure that the server does not fail to record the operation time even if the participant operates the shutter.

[0033] Participant 3 can view a video on the viewing terminal 3N, and by pressing icon 7 on the viewing terminal 3N at their desired timing while watching the video, they can acquire a still image of their favorite scene. However, a characteristic of the present invention is that the acquired still image does not always precisely coincide with the shutter timing. Figure 3(B) shows the relationship between the shutter timing by participant 3 and the acquired still image.

[0034] The 30 frames in Figure 3(B) are related to Figure 3(A), and frames 4, 5, 11, 12, 13, 21, 22, and 27 are NG shots. If participant 3's shutter timing T1 matches frame 4 (or can be considered to match within a predetermined time range, and so on), since frame 4 is an NG shot, the still image of frame 3, which is the closest non-NG shot to T1, is extracted. Similarly, if participant 3's shutter timing T2 matches frame 11, the still image of frame 10, which is the closest non-NG shot to T2, is extracted. The same applies to shutter timing T4. Of course, if participant 3's shutter timing corresponds to a still image of a non-NG shot, the still image of that frame is extracted (for example, frame 17 of T3 or frame 29 of T6). Furthermore, if participant 3's shutter timing T5 is at frame 27, then frame 27 is an unacceptable shot. Therefore, the system determines which of the non-unacceptable shots, frame 26 or frame 28, is closer to that shutter timing by a fraction of a second and makes a decision.

[0035] Taking 4K30fps as an example, the time lag between frames is only 0.0333 seconds, as mentioned above. Therefore, participant 3 perceives that a still image at the exact shutter timing has been extracted, and does not recognize that frames around the shutter timing, which do not strictly match, have been extracted. Consequently, when participant 3 presses icon 7 to take a picture while watching the video, they can experience taking a high-resolution still image that perfectly matches the shutter timing. Moreover, since the acquired still image is extracted from the shutter output still image, which does not include any NG shots such as closed eyes, it is guaranteed that no mis-shots are included. Furthermore, this also prevents images of NG shots from circulating in the market for artist management.

[0036] The still images (frame images) extracted by the shutter operation are displayed in real time on the viewing terminal 3N, showing the most recent still image. The reason for displaying the image corresponding to the time the shutter was operated immediately during shooting, rather than after the end of the video distribution for viewing, is to allow participant 3 to immediately check the frame images extracted on the viewing terminal 3N and to retake the picture immediately if the expected image was not captured. The same feeling as being able to check the image immediately after shooting on the camera's LCD monitor when shooting with a digital camera can be experienced in the shooting operation of the still image generation system 100 of this embodiment, making it possible to obtain a sense of realism in shooting.

[0037] However, instead of just displaying the most recent still image, the system may also display the captured still images sequentially. Furthermore, if the screen size of the viewing terminal 3N is small and the display area for the video is small, the still images captured after viewing are displayed together as a scrolling image. It goes without saying that although it is explained that one still image is extracted when the shutter is pressed, if icon 7 is held down, the corresponding still images will be extracted in succession, effectively creating a video.

[0038] Furthermore, when viewer 3 presses shutter 7 at a desired time, a charge may be applied based on the number of shutter presses, or a limit on the number of shutter presses may be set for the distribution of the viewing video 1, for example, 100 or 200 times. Also, the shutter fee may be free, or a bulk fee may be set for every 100 shutter presses, for example.

[0039] Next, I will explain the exclusive purchase and commercialization of still images obtained through photography. After watching the video, the still images captured and extracted are displayed as a list on the viewing terminal 3N, or displayed sequentially by scrolling, etc. Participant 3 can then purchase the images they wish to exclusively own. Figure 7 shows an example of the captured still images being displayed as a list on the viewing terminal 3N. Still images that can be exclusively purchased will have text or symbols such as "Exclusive Sale: Available for Purchase" added inside or near each still image, as shown in the figure (any text or symbols are acceptable as long as it is visually clear that the image can be exclusively purchased).

[0040] On the other hand, for still images that have been exclusively purchased by other participants, or for still images whose copyrights or other rights are not permitted for exclusive purchase by the party managing the subject image, text or symbols such as "Exclusive Sale: Sold Out" will be placed near each still image as shown in the diagram (any text or symbols are acceptable as long as it is visually clear that exclusive purchase is not possible). In this embodiment, the basis is that exclusive purchases are granted on a first-come, first-served basis from the perspective of fairness among participants, but it is also possible to configure the system so that some participants or groups of participants are given priority in exclusive purchases.

[0041] In other embodiments, in order to enable exclusive purchases during viewing of the video rather than after viewing, the system may add text or symbols such as "Exclusive Sale: Available for Purchase" in response to the purchase record immediately after the still images extracted by the shutter operation are sequentially displayed in real time on the viewing terminal 3N. The amount required for exclusive purchase may be added to the shutter fee, or it may be included in the shutter fee.

[0042] The still image generation system 100 not only allows users to purchase and download still images, but also performs the process of creating original products using their favorite still images. Specifically, you select the still image you wish to purchase (either exclusive or non-exclusive), and then select the product. Examples of products include, but are not limited to, acrylic stands, clear files, mugs, pass cases, fans, calendars, towels, keychains, straps, badges, stickers, clothing such as T-shirts, tote bags, and photobooks. The still image generation system 100 passes the still image data to a means that completes a product by transferring the selected still image onto the selected product, and the completed product is sent to the participant. The size and transfer position of the still image may be fixed, or the participant may be able to set them arbitrarily. In particular, if the participant specifies that the product be made using an image that has been exclusively purchased, the participant will be able to obtain an original product.

[0043] Furthermore, the still image generation system 100 of this embodiment can also incorporate information identifying the participant 3 who operated the shutter into each still image before providing that still image to each participant 3. For example, the text "Captured by ABC" (where ABC is each participant's unique ID or name, etc.) can be superimposed in small font at the bottom of each still image. Alternatively, the system can be configured to allow the participant to add and display any desired information, such as characters or symbols. This gives each participant 3 a sense of ownership over their own still image, and this identification information may also be displayed on the aforementioned products. Moreover, if the above identification information or desired information is processed in a way that makes it impossible for participant 3 to erase it, even if the still image or the product created is distributed without the permission of the copyright holder, the participant 3 who leaked it without permission can be identified from the identification information. Note that the superimposed display may be in the form of a watermark or hidden so that it is not immediately apparent that the identification information is incorporated into the still image.

[0044] Incidentally, once a still image is released to the market, it circulates freely, as seen in sales by unspecified individuals on communication networks, making the relationship between the transferor and transferee unclear. Furthermore, there is a risk that the image can be easily tampered with (modified). Even if the still image generation system 100 prepares still images for shutter output to prevent the circulation of unusable shots, it becomes meaningless if the images are modified to the same extent as unusable shots after being provided. Even if information identifying the participant who operated the shutter is incorporated into each still image as described above, many people may still believe that a modified fake image is genuine.

[0045] Therefore, the still image generation system 100 of this embodiment affixes an electronic signature to the purchased still image, thereby proving that the still image was taken using the still image generation system 100 and has not been tampered with (modified). This electronic signature indicates that each still image was created by the still image generation system 100 (in practice, by the platform operating company 2 of the still image generation system 100 or the entity that manages the copyright of the photographic image). If there is a still image without an electronic signature, it can be proven that it was not taken using the still image generation system 100, or that it has been tampered with by unauthorized or unknown means. As a result, it becomes possible to take legal action against the seller of still images without an electronic signature.

[0046] Furthermore, to prove the legitimacy of the digital signature attached to the still image itself, the system may be configured to attach a digital certificate as needed before sending it to participants. By using a digital certificate, any certification authority that issues digital certificates can prove that the digital signature attached to the still image was indeed made by the still image generation system 100. In other words, by using digital signatures and / or digital certificates, even after still images or products made using still images have entered the market, it becomes easy and reliable to verify whether they are genuine products created by the execution of the still image generation system 100 and whether they have not been tampered with by unauthorized or unknown means. The methods for applying digital signatures and using digital certificates are well-known mechanisms and can be implemented using various applications, so we will omit further explanation here.

[0047] Furthermore, by managing digital certificates attached to captured video and related still images on the blockchain, the ownership and uniqueness of the images can be guaranteed. For example, by converting an image into an NFT (Non-Fungible Token) as digital art and recording it on the blockchain, even if the image is copied, its uniqueness can be ensured by the digital certificate. The owner of the image can be verified. As a result, NFT trading becomes possible in the market, enabling secure and reliable rights management for artworks and commercial use.

[0048] Furthermore, the still image generation system 100 of this embodiment is designed so that the still images output by the shutter do not include any rejected shots, but the series of images includes what are considered to be the highest quality, so-called best shots, even to a professional photographer. Professional photographers, based on their extensive experience and skills, select the best shots from a vast number of shots taken from various angles, and these are used in photo books and the like. The fact that the same still image as this best shot is displayed on participant 3's viewing terminal 3N means that participant 3 took the picture at the same time as a professional photographer, and that it is a superb still image.

[0049] Therefore, in this embodiment, the still image generation system 100 has one or more best shot images predetermined and set by a professional photographer from among the still images for shutter output. If the still image corresponding to the shutter timing by participant 3 is the best shot image, the viewing terminal 3N will display a decorative indication that it is the best shot image (for example, the outer frame of the still image may be displayed in a specific color, or words such as "Grand Prize" or "Best Shot" or a celebratory ball may be displayed within the still image, or a sound may be output to indicate that it is the best shot image). Furthermore, if the still image captured by the shutter operation is the best shot image, various preferential treatments may be provided, such as being given a specific gift like a concert ticket for the featured artist, or being allowed unlimited shooting on the next shoot.

[0050] Furthermore, it is possible to add a competitive element by having multiple participants compete to see how many best shots they can input from all the still images used for shutter output. Alternatively, different points could be assigned to some or all of the still images in advance, and the points for the still images for which the shutter was operated could be accumulated to create a game-like element where participants compete against each other.

[0051] Figure 4 shows an example of the screen displayed on a viewing terminal 3N for a participant who is taking part in a photo shoot and watching a video for viewing. The still image generation system 100 of this embodiment photographs the subject with multiple cameras 6. Therefore, multiple videos for viewing, each capturing the subject from an angle corresponding to the number of cameras, are displayed on the viewing terminal 3N. Specifically, as shown in Figure 4, different videos for viewing, divided into multiple sections 41-45, are displayed on the screen. In other words, this is an example of a multi-angle screen where multiple viewing videos from different cameras are simultaneously displayed on the viewing terminal 3N in a split-screen format. This corresponds to the number of cameras used in the photo shoot, and since the position, orientation, and angle of each camera relative to the subject are different, each video will be displayed for viewing.

[0052] Participant 3 selects an image from a desired camera by tapping the display area of ​​the video feed from multiple cameras 6 shown on the viewing terminal 3N. Only the image from that camera is then enlarged and displayed on the entire screen or a portion of it 45, and Participant 3 can press the shutter button 7 at any time they like. If they wish to select an image from another camera 6, they can repeat the process by clicking on the multiple video feeds shown in Figure 4.

[0053] The multi-angle screen display shown in Figure 4 is extremely convenient for photo shoots involving a large number of participants. This is because each viewer can select video from their preferred camera angle (shooting angle and position), thus accommodating the preferences of almost all participants. Furthermore, the present invention adjusts the angle by switching the displayable area, from the perspective of accurately meeting the needs of a large number of participants. In other words, although the screen resolution on the participants' viewing terminals is generally set to FHD (1920 x 1080 pixels), the actual video shooting is based on ultra-high resolution and high-definition video quality, such as 8K (7680 x 4320 pixels) or 4K (3840 x 2160 pixels), allowing for changes to the displayable area.

[0054] Figure 5 shows examples of image ranges for FHD, 4K, and 8K. Each participant 3 can move between displayable areas, including the second area (4K) 52 and the third area (8K) 53, which are outside the first area (FHD) 51, by using finger operations such as tapping and pinching or mouse operations on their viewing terminal to change the area. This allows for flexible adaptation to the desired field of view (wide-angle, standard, telephoto) for each viewer. As a result, it is possible to provide viewers with an experience as if they were adjusting the field of view themselves using the camera 6 at the shooting location.

[0055] Furthermore, the resolution of the recorded video is not fixed or limited to 8K or 4K, and it goes without saying that it will increase in line with advancements in communication technology and LCD display technology. In addition, the settings for the shooting size and framing ratio when viewers change the displayable area on their viewing device can be changed to any value desired by each viewer, in accordance with the framing of the LCD monitor on the viewing device.

[0056] The ability to change the displayable area by switching image quality, as described above, offers significant convenience, especially when multiple viewers are taking pictures based on images captured by a single camera. This is like a photo shoot with one camera and many viewers. A fixed angle from a single camera makes it difficult to accommodate the wishes of many viewers, but according to the present invention, each viewer can freely select the display area size on their viewing terminal based on the moving image from a single camera. However, this is also effective in cases where multiple cameras are used, as in this embodiment. Since it is equivalent to multiple cameras shooting at an N-fold magnification, it will be possible to respond precisely to each viewer's instructions for shooting operations such as zooming in and out, no matter how large the number of viewers.

[0057] The present invention allows viewers to take a picture at their preferred timing from a streamed video, but it is based on the premise that a large number of viewers are watching "simultaneously." In such a viewing situation, the images taken at the moment when multiple participants 3 are taking pictures are often considered to be the optimal scenes to capture, i.e., to be extracted images.

[0058] However, each participant in a remotely filmed video for viewing does not know when other participants are taking pictures. If you are actually at a filming location, you can see that many other people are taking pictures, so you can recognize when it's a good opportunity to take a picture and take your own. However, with a still image generation system like the present invention, you cannot know when others are taking pictures, so you may regret later that you should have taken a picture.

[0059] Therefore, the system may include a configuration that utilizes the fact that the time of each participant's shutter operation is transmitted to the server in real time, allowing the shutter status of other participants to be viewed in real time. For example, as shown in Figure 6, each of the multi-angle screens 41-45 is provided with shutter indicators 61-65, and an indicator value representing the current total number of shutters is displayed. This indicates, for example, the total number of shutters from the present time to 5 seconds ago, and is updated every 5 seconds. Here, 5 seconds is just an example and can be set to any number of seconds. Suppose participant A planned to take a picture from the viewing video with the camera angle corresponding to section 41, but upon seeing the shutter indicator 62 on the multi-angle screen, realized that many participants were taking pictures from the viewing video with the camera angle corresponding to section 42. Viewer A can then switch from the camera angle of section 41 to the camera angle of section 42 and take a picture of the scene that other viewers are focusing on in the same way.

[0060] Other methods for displaying the indicator besides those shown in Figure 6 include, for example, highlighting the corresponding screen frames 41-45 on the multi-angle screen if a certain percentage of participants (e.g., more than half) are taking pictures with a camera, or making the size of screen frames 41-45 relatively larger to make them more noticeable. Alternatively, the indicator could be a curve graph showing the number of shutter clicks.

[0061] The still image generation system uses volumetric technology to capture video data from multiple cameras 6 that photograph the subject's surroundings in 360 degrees. The server then generates 3D data from this data and distributes it to the viewing terminal 3N. Volumetric technology involves setting up a 360-degree green screen as the shooting background, photographing the subject, and then generating 3D data afterward. By shooting against a green screen, video data of the subject can be composited into a virtual background such as VR (Virtual Reality), AR (Augmented Reality), or MR (Mixed Reality), allowing for flexible image processing and the addition of effects. Alternatively, even without shooting against a green screen, the system can recognize a person from any background, crop it with optimal framing, and composite it with other images.

[0062] This allows participants to experience the freedom to move freely in all directions—forward, backward, left, right, up, and down—from any angle and viewpoint while operating the shutter. Using volumetric shooting data, it's possible to switch between free-viewpoint footage and multiple free-viewpoint footage, such as having the camera pass through subjects as if in a video game, or viewing them from above.

[0063] The still image generation system, based on the multiple embodiments described above, enables a single video stream distributed to a large number of viewers to respond to all viewer requests, such as instructions for shooting operations including shooting angle, position, and zoom in / out. It also assists viewers in freely selecting the optimal shot. Furthermore, it enables virtual photo shoots in virtual reality worlds (such as the metaverse) and the operation of such photo studios. It can provide a realistic shooting experience even in metaverse space, enabling shooting and collaboration in the digital world.

[0064] (Second embodiment) The still image generation system 100 described above has been explained as a system that first saves the video footage shot at the shooting location to a server, edits it as needed to create a video for viewing, and then provides it to user 3 in the form of recorded data. Furthermore, it was assumed that still images for shutter output would be prepared in advance before viewing, that is, that reconstructed still images consisting of frame images with rejected shots removed would be prepared in advance before the video for viewing was distributed.

[0065] On the other hand, there is a need to watch live video footage shot at concert venues in real time, rather than recording it, and to obtain desired shots by remotely controlling the shutter. In this case, the captured video footage is not saved to a server but is delivered directly to the viewer as images for viewing. Even with live streaming, it is possible to divide the viewing image into multiple still images at a predetermined frame rate, automatically removing any unsuitable shots, and composing still images for shutter output in real time, extracting still images according to the shutter timing. Furthermore, if high-speed image processing is possible, editing such as noise reduction, image stabilization, overlays, and removal of effects can be performed automatically. An example of this procedure is shown below.

[0066] Figure 8 shows the time lag between the distribution of images for viewing and the generation of still images for shutter output. Let's assume that images for viewing, taken at a live venue, etc., are sent to a server and received by the server at time T0. Between time T0 and a predetermined time α (for example, 15 or 30 seconds), the server automatically performs the NG cut removal process described in Figure 3 in the first embodiment on the received images for viewing. Specifically, for example, for each still image divided at a predetermined frame rate, if the face recognition software determines that there is a person's face, it identifies whether the person's eyes are closed or not. If the eyes are closed, the still image is considered an NG cut and is not included in the still images for shutter output (corresponding to the frame images at the dotted lines). Furthermore, instead of simply removing NG cuts, AI software may be used to correct NG cuts. For example, if an image with closed eyes is detected, the software can identify the subject's eyes from the preceding and succeeding images and automatically correct them to an open-eyed state, resulting in a better expression. In other words, AI image processing is performed in real time to correct (modify) still images that may become NG cuts.

[0067] Furthermore, if the contrast value of a part of the image, such as a person or face, deviates by a certain amount compared to the average contrast value of the entire still image, the image will be deemed too bright or too dark and will be rejected, or the overly bright or dark parts will be corrected to harmonize with the rest of the image. The same applies to sharpness; if a part of the image, such as a person or face, is judged to be blurred below a predetermined value, the image will be rejected, or the blurred areas will be corrected to make them clearer. In addition, an AI application is used to identify still images with inappropriate posing and reject them or correct them. This automatic removal of unacceptable shots, as well as arbitrary correction processes such as blur correction, skin correction, background compositing, effect addition, and resolution improvement, are performed on a group of still images divided at a predetermined frame rate over a time interval α. For example, at 30fps, there are 30 images per second, so if the time interval α is 15 seconds, 450 images will be processed. The automatic removal of unacceptable shots and arbitrary correction processes described above can also be applied to the first embodiment. It is efficient to automate the automatic removal of unacceptable shots and arbitrary correction processes during the process of constructing still images for shutter output, starting from viewing videos stored on the server.

[0068] Next, at time T1, after a time interval α has elapsed, the server begins distributing the viewing image via the communication network so that it can be viewed on user 3's viewing terminal 3N. In other words, the automatic removal of NG cuts is performed α time in advance, and the viewing image is distributed with a time delay of α. While distributing the viewing image at time T1, the server simultaneously performs the automatic removal of NG cuts for the next time interval α from time T1 in parallel. With this configuration, the effects of the present invention can be achieved without having to save the reconstructed image with NG cuts removed on the server and prepare still images for shutter output in advance before viewing. The distribution of the viewing image may be adjusted to time T0+2α or time T0+3α, etc., depending on the time required for the automatic removal and correction of NG cuts. In addition, if the effect removal process shown in step S21 and the retouching process (including super-resolution processing) shown in step S25 of Figure 2 can be completed together with the automatic removal of NG cuts before the distribution of the viewing image, they may be included as additional processes.

[0069] In all of the embodiments described above, the subject is captured through the user's own shutter operation, and multiple still images are obtained that do not include rejected shots, etc. However, if obtaining the still images is subject to a fee, the participant will have to choose a limited number of images from among the multiple images. The still image generation system 100 has the following mechanism to facilitate this selection, according to the participant's request.

[0070] (1) Batch selection after shooting After the photo shoot, AI software that analyzes the subject's expressions and poses is used to extract the best shots from the multiple still images taken by the participants. If a participant specifies a desired number of images, the best-scoring images will be extracted in a batch. The extraction criteria are registered in advance. (2) Real-time selection Immediately after a shutter operation, AI software is used to select the still image extracted from that shutter operation as a best shot candidate if it meets or exceeds a predetermined standard. The best shot standard is registered in advance. If the still image extracted from the next shutter operation meets or exceeds the predetermined standard and is also a best shot candidate, it is compared with the previous best shot candidate to determine which is the best shot image. This process is repeated sequentially. In the above repetition, it is also possible to compare multiple best shot candidate images, rank them, and select the desired number of images for the participant from the top. This is suitable for real-time viewing image distribution, as in the second embodiment. (3) Suggestions for your favorite shots The system acquires participants' preferences as data in advance and suggests still images that match each participant's preferences after shooting or during real-time viewing. (4) Random shutter In addition to the shutter operation by the participants (or with the participants' shutter operation disabled without their knowledge), the AI ​​software automatically operates the shutter according to predetermined rules. From the still images extracted by the shutter operation, the following are selected based on (1) or (2) above.

[0071] The above (1) to (4) will meet the needs of participants who find photography enjoyable but find selecting still images difficult and would prefer to leave it to the professionals.

[0072] The still image generation system of the present invention is realized by processes, means, and functions executed by a computer in accordance with instructions from a program (software) launched on the system. The program can send commands to various components of the computer, causing them to execute predetermined processes, functions, etc., according to the present invention as described above. In other words, each process, means, and function in the present invention is realized by specific means in which the program and the computer work together. Furthermore, all or part of the program is provided on, for example, a magnetic disk, optical disk, semiconductor memory, or any other computer-readable recording medium, and the program read from the recording medium is installed and executed on the computer. Alternatively, the program can be loaded directly onto the computer via a communication line without using a recording medium and executed thereafter. Therefore, the program installed or loaded onto the computer, and these storage media, are included within the scope of the invention.

[0073] Furthermore, the processing performed by a server (e.g., a single personal computer) in the still image generation system according to the present invention may be performed by a mobile phone, personal computer, or data relay device. It can also be configured with a single server or with multiple servers (e.g., a group of multiple server computers). The above embodiments are merely examples to illustrate the present invention clearly. Moreover, the viewing terminal that communicates with the still image generation system via a communication network is a computer connected to a network such as the Internet or a dedicated line. Specifically, examples include PCs (Personal Computers), mobile phones and smartphones, PDAs (Personal Digital Assistants), tablets, and wearable devices. A business scheme including the still image generation system is formed by configuring PC terminals and mobile terminals connected to the communication network via wired or wireless connections to enable communication with each other. In the embodiments described above, the still image generation system may be configured to cooperate with an ASP (Application Service Provider). [Explanation of Symbols]

[0074] 1. Filming event management company 2. Platform operating company 3 participants 3N viewing terminal 4. Communication Network 6 cameras 7 Shutter icon 100 Still Image Generation Systems

Claims

1. A still image generation system for providing a simulated shooting experience by distributing a video for viewing, created based on captured video images obtained by a shooting device, to multiple users' image display terminals via a communication network, (a) The video is decomposed into a plurality of still images based on a predetermined frame rate, (b) A reconstructed still image is generated based on still images obtained by removing some of the multiple still images or still images obtained by modifying some of the still images. This generates a still image for shutter output, If it is determined that multiple users perform a finger action equivalent to a shutter operation on the image display terminal while viewing the aforementioned video, the still image corresponding to the timing of the shutter operation is identified from the shutter output still images. During or after viewing the aforementioned video for viewing, the image display terminal will display whether or not it is possible to exclusively purchase one or more specified still images. A still image generation system in which, if any of the aforementioned users select a still image that is exclusively available for purchase, the other users are prevented from purchasing that selected still image.

2. A still image generation system for providing a simulated shooting experience by distributing a video for viewing, created based on captured video images obtained by a shooting device, to multiple users' image display terminals via a communication network, At predetermined time intervals from the start of acquiring the aforementioned video footage, (a) The captured video image within the predetermined time interval is decomposed into a plurality of still images based on a predetermined frame rate, (b) A still image for shutter output is generated by generating a reconstructed still image based on a still image obtained by removing some of the multiple still images or a still image obtained by modifying some of the still images, (c) In parallel with the generation of the still images for shutter output, the video for viewing is distributed to the image display terminals of the multiple users after a predetermined time has elapsed from the start of acquisition, and when it is detected that a finger action equivalent to a shutter operation has been performed, a still image corresponding to the timing of the shutter operation is identified from the still images for shutter output. Repeatedly, During or after the acquisition of the aforementioned video footage, the image display terminal will display whether or not it is possible to exclusively purchase one or more of the identified still images. A still image generation system in which, if any of the aforementioned users select a still image that is exclusively available for purchase, the other users are prevented from purchasing that selected still image.

3. The still image generation system according to claim 1 or 2, wherein the reconstructed still image has had removed still images that were determined to show the user with their eyes closed or still images that are not permitted to be provided to the user.

4. The still image generation system according to claim 1 or 2, wherein the still image displayed on the user's image display terminal includes information that identifies the user or desired information specified by the user.

5. The still image generation system according to claim 1 or 2, wherein the still images purchased by the user are affixed with a digital signature or digital certificate.

6. The still image generation system according to claim 1 or 2, wherein the still image purchased by the user is transferred to a selected item.

7. The still image generation system according to claim 1 or 2, wherein if the still image corresponding to the timing of the shutter operation matches a pre-set good shot image, a predetermined display or sound is output on the image display terminal of the user who performed the shutter operation to allow recognition.

Citation Information

Patent Citations

  • Virtual photo session system and photographing system

    JP2003209741A

  • Remote photography system and remote photography method

    JP7288641B1

  • Remote photography system and remote photography method

    JP7640159B1

  • Image data providing device and image data providing method

    WO2004077830A1