Still image creating system

The still image generation system addresses network and mis-shot issues in remote shooting by processing video footage to generate high-quality, copyright-compliant still images based on viewer inputs, ensuring professional-level photography for multiple participants.

WO2026088918A1PCT designated stage Publication Date: 2026-04-30CAP CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
CAP CO LTD
Filing Date
2025-10-20
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Existing remote shooting systems struggle with network load issues due to real-time video streaming, poor network environments affecting video quality, and the inclusion of mis-shots or unusable images, especially when a large number of viewers participate, which complicates individual shooting preferences and copyright management.

Method used

A still image generation system that processes captured video footage to remove effects, decomposes it into still images, filters out mis-shots and NG cuts, and generates high-quality still images based on viewer shutter operations, ensuring high-resolution output and compliance with copyright restrictions.

Benefits of technology

The system provides high-quality, mis-shot-free still images to viewers, regardless of network conditions, allowing amateur users to achieve professional-level photography and managing copyright-compliant image distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025036893_30042026_PF_FP_ABST
    Figure JP2025036893_30042026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a video distribution imaging service that is not affected by a viewing communication environment of a viewer, and that provides high-quality still images (photographs) excluding bad shots, regardless of shutter timing. Since the still images are output on the basis of high-quality moving image data that have been imaged and edited, the shutter output does not produce coarse images like so-called screen shots. Shutter still images constructed by removing still images considered to be bad shots from among a plurality of still images formed on the basis of a predetermined frame rate are prepared in advance, and images with the eyes closed, for example, are excluded from still images to be exported in accordance with the time at which the shutter was operated. Therefore, even an amateur viewer who does not have professional-level photographic skills can obtain still images (photographs) without bad shots, and only still images permitted by a party managing a copyright and the like of subject images can be provided to the viewers.
Need to check novelty before this filing date? Find Prior Art

Description

Still image generation system

[0001] The present invention relates to a still image generation system.

[0002] In recent years, an application has emerged that allows users to receive and view moving images captured by a camera via a network, and send an instruction equivalent to pressing the shutter at a desired timing, providing an experience as if photographing a subject right in front of them. Even without actually going to the shooting location, viewers can have a pseudo-experience of the shooting act and feel a sense of presence as if they were taking the photo themselves. Moreover, since it is a remote operation, it is possible to shoot from any location physically far away, and there is no situation where shooting becomes impossible due to the number limit at the shooting site. Therefore, the need for remote shooting is expected to increase in the future.

[0003] As prior art based on a similar concept, for example, there is the following patent document (see Patent Document 1). The shooting meeting system described in Patent Document 1 is configured such that when shooting conditions such as zooming in are sent to a shooting agent (human or device) that actually shoots a subject at the shooting site, the proxy shooter adjusts the shooting conditions based on the shooting instruction. In addition, a remote shooting system has been proposed that distributes moving images captured at a live venue or the like to viewer terminals via a server in a batch, and executes a finger action equivalent to a shutter operation on the viewer terminal for a favorite scene (see Patent Document 2).

[0004] Japanese Patent Application Laid-Open No. 2003-209741, Patent No. 7288641

[0005] The invention described in Patent Document 1 is basically intended for photo shoots with a very small number of viewers. Therefore, the proxy photographer can accept requests from viewers and change the shooting conditions each time. However, the remote photo shoot system that the present invention is intended for is based on the premise that a large number of viewers will participate, so it is practically impossible to accept shooting conditions that match the diverse requests of each viewer. Unlike Patent Document 1, as the number of viewers participating in the photo shoot system increases, it becomes impossible to proceed with the photo shoot if all viewers are asked to specify the subject to be photographed, the shooting angle, the magnification, etc.

[0006] The invention described in Patent Document 2 is based on the premise that a large number of viewers will participate, and each viewer can freely decide on their desired shooting scene and change the shooting angle, magnification, etc. However, the configuration involves uploading the recorded video data to a communication network in real time, receiving it on a server, and then instantly downloading the video data to the viewer's terminal. This upload and download is performed sequentially at predetermined time intervals (for example, 5 seconds, 30 seconds, etc.), and in addition, information on the shutter timing time by each viewer is uploaded to the server in real time each time the shutter is operated, which results in a large network load. Therefore, if the network communication environment is poor, problems arise in that it is difficult to deliver high-quality and stable video to viewers, such as the video being paused or the video being streamed discontinuously.

[0007] Furthermore, in live video streams of people, viewers often take photos when the subject's eyes are closed, or the movement of the subject is captured in the frame, resulting in blurry footage. These are generally considered mis-shots that should be removed, not still images that viewers want. In addition, copyrights and related rights apply to photographs, images, and performances of artists, and it is known that major artists, in particular, have a strong desire to avoid having unusable images exposed on the market.

[0008] Therefore, the present invention aims to provide a video streaming shooting service that is not affected by the viewer's viewing and communication environment, and that provides high-quality still images (photographs) that do not include mis-shots or unusable cuts regardless of the shutter timing.

[0009] To achieve the above objective, the still image generation system and method according to the present invention distributes a video for viewing, created based on captured video footage acquired by a shooting means, to multiple users' image display terminals via a communication network, characterized in that: (a) if arbitrary effect processing is applied to the captured video footage, a video footage with the effect processing removed is generated; (b) the video footage is decomposed into multiple still images based on a predetermined frame rate and output; (c) a reconstructed still image is generated based on still images from which some still images have been removed from the multiple still images, thereby generating a still image for shutter output; and when it is determined that the multiple users have performed a finger action equivalent to a shutter operation on their image display terminals while viewing the video for viewing, a still image corresponding to the timing of the shutter operation is identified from the still images for shutter output; and the identified still image is displayed on the users' image display terminals.

[0010] Furthermore, the reconstructed still image is characterized in that still images in which the user's eyes are closed or still images that are not permitted to be provided to the user have been removed.

[0011] The still image generation system and method according to the present invention output still images based on high-quality video data captured and edited, so that the shutter output does not consist of low-resolution images like so-called screenshots. Furthermore, a set of still images for the shutter is prepared in advance by removing images that are treated as mis-shots and NG cuts from a plurality of still images composed based on a predetermined frame rate, and only still images from this set are written out to be provided to the viewer. Therefore, no matter when the shutter is operated, the still images written out will not include mis-shot images or NG cuts. Even amateur viewers without advanced shooting skills can take photos at the same level as professional photographers, and only still images permitted by those managing the copyrights of the subject images can be provided to the viewer.

[0012] This is a diagram showing the overall configuration of one embodiment of a still image generation system. This is a flowchart showing the process of generating still images for shutter output. Figure 3(A) shows the relationship between the video for viewing and the frame images in which cuts to be removed are specified, and Figure 3(B) shows the relationship between the shutter timing and the still image to be acquired. This is a diagram showing an example of a multi-angle screen displayed on a viewing terminal. This is a diagram to explain switching between different display areas based on the same video image. This is a diagram showing an example of a shutter indicator. This is a diagram showing an example of a screen when processing the exclusive purchase of an image. This is a diagram showing the time lag between the distribution of the video for viewing and the generation of the still image for shutter output.

[0013] Hereinafter, one embodiment of the still image generation system according to the present invention will be described in detail with reference to the drawings. In all the drawings used to describe the following embodiment, the same parts will be denoted by the same reference numerals in principle, and repeated descriptions will be omitted. In addition, some components that are not important for explanation will be omitted from each drawing. It goes without saying that the present invention can be implemented in many different forms and is not limited to the contents disclosed below, and various modifications, substitutions, changes in processing order, and omissions are possible without departing from the gist of the invention.

[0014] An embodiment of the still image generation system according to the present invention will be described below with reference to the drawings. Figure 1 shows the overall configuration of the still image generation system 100. The still image generation system 100 consists of the interaction between a filming event management company 1, a platform management company 2, and a number of users 3, and is configured to send and receive various types of data via a communication network 4 such as the Internet. In this embodiment, filming at a filming venue will be used as an example, but it is not necessary for the filming venue to be an indoor space such as a studio or live venue, and filming may take place at any outdoor location (for example, a filming location).

[0015] The photography event management company 1 plans the date and time of the photo shoot, the location, and the models or artists who will be photographed at the venue, and announces that the remote photo shoot will be held through its website and social media. The platform management company 2, which has a prior contract with the photography event management company 1, is a business entity that provides the platform for the still image generation system 100 and is responsible for all the technical support necessary to realize the planned photo shoot. Because the platform management company 2 provides the platform to the photography event management company 1 on an OEM basis, the user 3 perceives the photo shoot as being conducted by the photography event management company 1. Although the still image generation system 100 shown in Figure 1 treats the photography event management company 1 and the platform management company 2 as separate entities, there is no difference in the effects and advantages of the present invention even if the photography event management company 1 and the platform management company 2 constitute the same entity in the still image generation system 100. Furthermore, the announcement of the event to users may be made by the platform management company 2 on behalf of the photography event management company 1.

[0016] The filming event management company 1 accepts applications from users to participate in remote filming sessions via the communication network 4. The filming event management company 1 edits the filmed video and accompanying audio (hereinafter referred to as "filmed video") and stores the resulting video for viewing in an information processing device (for example, including a server or a television editing machine; hereinafter referred to as "server"). The video for viewing uploaded from the server is provided to user 3 via the communication network 4. In this embodiment, platform management company 2 owns and manages the server, but it may be a server within the filming event management company 1, or it may be a virtual server on the cloud. Hereinafter, the server will also be referred to as a "cloud server".

[0017] At the shooting location, preparations are made for setting up one or more shooting devices 6 (hereinafter referred to as "camera 6") for shooting the subject, and for shooting moving images and recording sound. At the shooting location, a person may hold the camera 6 and shoot, or the camera 6 may be automatically controlled without human intervention by command signals from a server. In this embodiment, an example is shown in which the subject is shot with multiple cameras 6 (for example, including setting up multiple cameras around the subject, such as 180 degrees or 360 degrees), but shooting may also be done with a single camera 6.

[0018] While promotional videos and music videos are edited, the recorded video footage stored on the server in the present invention is also edited using a video / still image editor. The appropriately edited recorded video footage is provided to user 3 as a video for viewing. There are various editing methods, but typical ones include multicam editing, which involves switching scenes while simultaneously playing footage taken from various angles with multiple cameras 6, noise reduction, overlaying arbitrary footage, and text insertion. Other editing processes include image stabilization and the addition of special visual and sound effects. In this embodiment, the video for viewing is generated by editing the recorded video footage, but it is also possible to distribute the recorded video footage directly to viewers as a video for viewing.

[0019] Participant 3, a user participating in the photo shoot, possesses a viewing terminal 3-1 to 3-N (hereinafter referred to as "viewing terminal 3N") and will view the images via the communication network 4. When receiving the distribution of viewing images, the viewing terminal 3N must be connected to the communication network 4. When the distribution of viewing images begins, registered participants 3 will view the images on their respective viewing terminals 3N, and an icon 7, which corresponds to the shutter operation, will be displayed on the application screen of the viewing terminal 3N. Participants will press the icon 7 at any time to perform the so-called shutter operation.

[0020] By operating the shutter, participant 3 will obtain still images (sometimes called "frame images") of their favorite scenes. The present invention is characterized by its method of exporting these still images, which will be explained below. First, it is necessary to prepare still images for shutter output in advance, which are different from the video for viewing. Figure 2 is a flowchart showing the process of generating still images for shutter output.

[0021] As shown in Figure 2, the video for viewing stored on the server is read (step S20). As mentioned above, the video for viewing has undergone various video editing, so a video is generated with the effects applied to the image and sound removed (step S21). Effects include video synthesis, color expression, and echo processing, but specifically, they include adding environmental effects such as thunder and rain, applying blurring or bumpy effects to the finished image, or using special sounds such as scratches or horns. In step S21, the process of removing effects related to visual effects is performed (however, the process of removing effects related to sound is not excluded). Note that in the case of conventional still image generation systems, if the LCD screen of the viewing terminal 3N is FHD (1920 x 1080 pixels) compatible, there is no need to extract and provide high-resolution still images (frame images). In contrast, the still image generation system 100 of this embodiment shoots based on the image quality of ultra-high resolution / high-definition video such as 8K (7680 x 4320 pixels) or 4K (3840 x 2160 pixels). Even if the LCD screen of the viewing terminal 3N is FHD quality, in order to provide the participant 3 with high-definition still images such as 8K or 4K, a high-quality video is generated by removing effects from a high-quality video for viewing.

[0022] Next, multiple still images based on a predetermined frame rate are output from the video generated in step S21 (step S22). In other words, the video is broken down into multiple still images. If the streamed video is 4K resolution and 30fps (a frame rate consisting of 30 images per second), each still image can be identified by the elapsed time obtained by adding 1 / 30 = 0.0333 seconds from the start time of the video streaming for viewing. In other embodiments, the effects may not be removed (i.e., step 21 may not be performed), and the process may proceed directly from step 20 to step 22.

[0023] Next, some still images are removed from the multiple frame images created in step S22 (step S23). The still images to be removed include those with closed eyes or inappropriate posing. In addition, NG cut images that are not acceptable to the copyright management party, such as the artist, are also removed. Hereafter, these will be collectively referred to as NG cuts. Figure 3(A) shows the relationship between the edited video for viewing and the multiple frame images that have been designated as NG cuts that should be removed. Figure 3(A) uses still images exported from a video for viewing with 30 frames per second and no effects removed as an example, and frames 4, 5, 11, 12, 13, 21, 22, and 27 have been designated as NG cuts due to reasons such as closed eyes. The server stores a reconstructed still image composed of multiple frame images from which such NG cuts have been removed (step S24). Alternatively, instead of actually removing some still images that would otherwise be considered NG cuts, it may be possible to simulate the removal of NG cuts by using software processing such as flags or tables to identify which frames are NG cuts, thereby creating a reconstructed still image equivalent to the original.

[0024] In this embodiment, the selection of which frames to designate as NG cuts is done manually, but the selection of NG cuts may be automated using image recognition applications such as face recognition software or eye-tracking software, or an AI system (AI application).

[0025] Next, the reconstructed still image created in step S24 is retouched (step S25) to construct a still image for shutter output and store it on the server (step S26). Retouching includes, for example, image processing such as adjusting the skin tone of the subject, refining the skin texture, correcting parts of the face or body such as making the eyes larger, improving the resolution, and changing the background color or the color of the clothing. It may also include processing to add and delete objects not included in the reconstructed still image, such as brightening only specific parts, removing unwanted objects, compositing absentees into a group photo, or compositing the background, and there are no particular restrictions on the type or content of retouching. Furthermore, in order to increase the resolution, an unresolved image may be input and processing may be performed to improve the resolution limit using super-resolution technology. Examples of super-resolution methods include case-based (learning-type) super-resolution, fractal super-resolution, and reconstruction-type super-resolution.

[0026] The retouching and super-resolution processing described above can be applied fully automatically to all reconstructed still images, or manually to each still image, or to some images and then to the remaining images using AI software. It should be noted that retouching and super-resolution processing are not mandatory; they should be performed only as needed.

[0027] When steps S21, S23, and S24 are executed using AI software via ASP, their operation methods are optimized based on the performance of the still image generation system 100.

[0028] After preparing the still images for shutter output in this manner, the viewing video is distributed to participant 3. The start time of the viewing video distribution is centrally managed by a server connected to the communication network 4, and the shutter operation time for each participant 3 is managed in relation to the start time of the distributed video. In other words, the elapsed time from the start time of the distributed video represents the shutter operation time. When the viewing video is distributed to multiple participants simultaneously, the server manages one common start time for all participants, but when the start times are staggered for each participant, the server manages the start time corresponding to each participant.

[0029] The time at which each participant 3 takes a shutter operation on the video received by their respective viewing terminal 3N during the photo shoot is transmitted to the server in real time and recorded on the server side. This time is treated as a UNIX® timestamp in seconds or milliseconds format, and based on the elapsed time starting from the start time of the distributed video (which is also recorded as a UNIX® timestamp), it is possible to identify the frame image corresponding to the time when participant 3 took a shutter operation.

[0030] Furthermore, even if latency delays occur due to the communication environment of the viewing terminal 3N, the still image generation system 100 of this embodiment maintains low latency so that communication for recording the shutter operation time is kept within milliseconds. Even if a participant operates the shutter while the video is being streamed, and a problem occurs such as being on the move or entering a building with a poor communication environment, causing the connection to the internet to be lost, the platform installed on the viewing terminal 3N for video streaming temporarily caches the shutter operation time on the viewing terminal side, and automatically resynchronizes once the connection is restored. Therefore, measures are taken to ensure that the server does not fail to record the operation time even if the participant operates the shutter.

[0031] Participant 3 can view a video on the viewing terminal 3N, and by pressing icon 7 on the viewing terminal 3N at their desired timing while watching the video, they can acquire a still image of their favorite scene. However, a characteristic of the present invention is that the acquired still image does not always precisely coincide with the shutter timing. Figure 3(B) shows the relationship between the shutter timing by participant 3 and the acquired still image.

[0032] The 30 frames in Figure 3(B) are related to Figure 3(A), and frames 4, 5, 11, 12, 13, 21, 22, and 27 are all unacceptable shots. If participant 3's shutter timing T1 matches frame 4 (or can be considered to match within a predetermined time range, and so on), since frame 4 is an unacceptable shot, the still image of frame 3, which is the closest non-unacceptable shot to T1, is extracted. Similarly, if participant 3's shutter timing T2 matches frame 11, the still image of frame 10, which is the closest non-unacceptable shot to T2, is extracted. The same applies to shutter timing T4. Of course, if participant 3's shutter timing corresponds to a still image of a non-unacceptable shot, the still image of that frame is extracted (for example, frame 17 at T3 or frame 29 at T6). Furthermore, if participant 3's shutter timing T5 is at frame 27, then frame 27 is an unacceptable shot. Therefore, it is determined which of the non-unacceptable shots, frame 26 or frame 28, is closer to that shutter timing by a fraction of a second, and the decision is made accordingly.

[0033] Taking 4K 30fps as an example, the time lag between frames is only 0.0333 seconds, as mentioned above. Therefore, participant 3 perceives that a still image at the exact shutter timing has been extracted, and does not recognize that frames around the shutter timing, which do not strictly match, have been extracted. Consequently, when participant 3 presses icon 7 to take a picture while watching the video, they can experience taking a high-resolution still image that perfectly matches the shutter timing. Moreover, since the acquired still image is extracted from the shutter output still image, which does not include any NG shots such as closed eyes, it is guaranteed that no mis-shots are included. Furthermore, this also prevents images of NG shots from circulating in the market for artist management.

[0034] The still images (frame images) extracted by the shutter operation are displayed in real time on the viewing terminal 3N, showing the most recent still image. The reason for displaying the image corresponding to the time the shutter was operated immediately during shooting, rather than after the end of the video distribution for viewing, is to allow participant 3 to immediately check the frame images extracted on the viewing terminal 3N and to immediately retake the picture if the expected image was not captured. The same feeling as being able to check the image immediately after shooting on the camera's LCD monitor when shooting with a digital camera can be experienced in the shooting operation of the still image generation system 100 of this embodiment, making it possible to obtain a sense of immediacy during shooting.

[0035] However, instead of just displaying the most recent still image, the system may also display the captured still images sequentially. Furthermore, if the screen size of the viewing terminal 3N is small and the display area for the video is small, the still images captured after viewing may be displayed as a scrolling image. It goes without saying that although it is explained that one still image is extracted when the shutter is pressed, if icon 7 is pressed and held down, the corresponding still images will be extracted in succession, effectively creating a video.

[0036] Furthermore, when viewer 3 presses the shutter 7 at a desired timing, a charge may be applied based on the number of shutter presses, or a limit on the number of shutter presses, such as 100 or 200, may be set for the distribution of the viewing video 1. Also, the shutter fee may be free, or a bulk fee may be set for every 100 shutter presses, for example.

[0037] The number of times participant 3 takes a photo while watching a video becomes statistical data based on participant 3's actual actions, allowing for the quantification of emotions expressed through the use of the shutter. For example, if the photo shoot is a fashion show, styling data can be compiled, making it possible to understand the clothing preferences of many participants. If the photo shoot is a public model audition, popularity data can be compiled, and if an influencer participates, the expressions and poses that receive the most shutter clicks can be analyzed to identify popular points, which can then be applied to other models. Therefore, the photo shoot event management company 1 can sell the number of shutter clicks transmitted to the server. Furthermore, it is possible to manufacture and sell various products or provide services based on the compiled data.

[0038] Next, we will explain the exclusive purchase and product creation of still images obtained through photography. After viewing the video, the still images extracted from the photos are displayed as a list on the viewing terminal 3N, or displayed sequentially by scrolling, etc. From these, participant 3 can purchase the images they wish to exclusively own. Figure 7 shows an example of the still images taken being displayed as a list on the viewing terminal 3N. Still images that can be exclusively purchased will have text or symbols such as "Exclusive Sale: Available for Purchase" added inside or near each still image, as shown in the figure (any text or symbols are acceptable as long as it is visually clear that exclusive purchase is possible).

[0039] On the other hand, for still images that have been exclusively purchased by other participants, or for still images whose copyrights or other rights are not permitted for exclusive purchase by the copyright holder, a message or symbol such as "Exclusive Sale: Sold Out" will be placed near each still image, as shown in the diagram (any message or symbol is acceptable as long as it is visually clear that exclusive purchase is not possible). In this embodiment, the basic principle is that exclusive purchases are granted on a first-come, first-served basis from the perspective of fairness among participants, but it is also possible to provide a configuration that allows some participants or groups of participants to have priority in exclusive purchases.

[0040] In other embodiments, to enable exclusive purchases during viewing of a video rather than after viewing, the text or symbols such as "Exclusive Sale: Available for Purchase" may be added in response to the purchase immediately after the still images extracted by the shutter operation are sequentially displayed in real time on the viewing terminal 3N. The amount required for exclusive purchase may be added to the shutter fee or included in the shutter fee.

[0041] The still image generation system 100 not only allows users to purchase and download still images, but also performs the process of creating original products using their favorite still images. Specifically, after selecting the still image they wish to purchase (either exclusive or non-exclusive purchase), they select a product. Examples of products include, but are not limited to, acrylic stands, clear files, mugs, pass cases, fans, calendars, towels, keychains, straps, badges, stickers, clothing such as T-shirts, tote bags, and photo books. The still image generation system 100 passes the still image data to a means that transfers the selected still image onto the selected product to complete the product, and the completed product is sent to the participant. The size and transfer position of the still image may be fixed, or participants may be able to set them arbitrarily. In particular, if the participant specifies the creation of a product using an exclusive purchased image, they will be able to obtain an original product.

[0042] In addition, the still image generation system 100 of the present embodiment can also provide each participant 3 with the still image after incorporating information for identifying the participant 3 who performed the shutter operation on each still image. For example, at the bottom of each still image, "Captured by ABC" (ABC is the ID or name unique to each participant, etc.) can be superimposed and displayed in small font. Alternatively, desired information such as characters or symbols specified by the participant can be added and displayed. From the perspective of each participant 3, a feeling of having their own unique still image can arise, and this identification information may also be displayed on the above-mentioned products. Furthermore, if the above-mentioned identification information and desired information are processed on the participant 3 side so that they cannot be deleted, even if the still image or the produced product is distributed without permission by the copyright owner or the like, there is a merit that the participant 3 who caused the unauthorized outflow can be identified from the identification information. Note that the superimposed display may be in the form of a watermark display or non-display so that it cannot be immediately determined that the identification information is incorporated into the still image.

[0043] By the way, once a still image enters the market, it circulates among various parties and is transferred, and not only does the relationship between the transferor and the transferee become unclear, but there is also a risk that the image will be easily altered (modified). Even if the still image generation system 100 prepares a still image for shutter output to prevent NG cuts from spreading, if a modification similar to an NG cut is made after the provision, it will become meaningless. Even if information for identifying the participant who performed the shutter operation as described above is incorporated into each still image, many people may believe that the altered fake image is genuine.

[0044] Therefore, the still image generation system 100 of this embodiment affixes an electronic signature to the purchased still image, thereby proving that the still image was taken using the still image generation system 100 and has not been tampered with (modified). This electronic signature indicates that each still image was created by the still image generation system 100 (in practice, by the platform operating company 2 of the still image generation system 100 or the entity that manages the copyright of the photographic image). If there is a still image without an electronic signature, it can be proven that it was not taken using the still image generation system 100, or that it has been tampered with by unauthorized or unknown means. As a result, it becomes possible to take legal action against the seller of still images without an electronic signature.

[0045] Furthermore, to prove the legitimacy of the digital signature attached to the still image itself, the system may be configured to attach a digital certificate as needed before sending it to the participant. By using a digital certificate, any certification authority that issues digital certificates can prove that the digital signature attached to the still image was indeed made by the still image generation system 100. In other words, by using digital signatures and / or digital certificates, even after still images and products made using still images have entered the market, it becomes easy and reliable to verify whether they are genuine products created by the execution of the still image generation system 100 and whether they have not been tampered with by unauthorized or unknown means. Note that the methods for attaching digital signatures and using digital certificates are known mechanisms and can be implemented using various applications, so an explanation is omitted here.

[0046] Furthermore, by managing the captured moving images and the electronic certificates attached to the related still images on the blockchain, the ownership and uniqueness of the images can be guaranteed. For example, by tokenizing the image as an NFT (Non-Fungible Token) for digital art and recording it on the blockchain, even if the image is copied, its uniqueness can be ensured by the electronic certificate. The ownership of the image can be proven. As a result, NFT transactions become possible in the market, enabling secure and reliable rights management in art works and commercial use.

[0047] In addition, the still image generation system 100 of the present embodiment is designed such that the shutter output still images do not include NG cuts. Among a series of these images, there are so-called best shot images of the highest quality that even a professional cameraman would consider excellent. Based on their rich experience and techniques, professional cameramen select the best shot images from a vast number of shot images taken from various angles, and these are used in photo albums and the like. The fact that the same still image as this best shot image is displayed on the viewing terminal 3N of Participant 3 means that Participant 3 took the shot at the same timing as a professional cameraman, resulting in wonderful still images.

[0048] Therefore, the still image generation system 100 of the present embodiment pre-sets one or more best shot images among the shutter output still images, determined in advance by a professional cameraman. When the still image corresponding to the shutter timing by Participant 3 is a best shot image, a decorative display that can be recognized as a best shot image is shown on the viewing terminal 3N (for example, the outer frame of the still image is displayed in a specific color, or characters such as "big win" or "best shot" or a display of a cracked ball are shown inside the still image, or a sound is output to indicate that it is a best shot image). Also, when the still image taken by the shutter operation is a best shot image, various preferential treatments may be provided, such as awarding a specific present like a live ticket for the artist of the subject, or making the next shooting unlimited.

[0049] Furthermore, it is possible to add a competitive element by having multiple participants compete to see how many best shots they can input from all the still images used for shutter output. Alternatively, different points could be assigned to some or all of the still images in advance, and the points for the still images for which the shutter was operated could be accumulated to create a game-like element where participants compete against each other.

[0050] Figure 4 shows an example of a screen displayed on a viewing terminal 3N for a participant who is taking part in a photo shoot and watching a video for viewing. The still image generation system 100 of this embodiment photographs the subject with multiple cameras 6. Therefore, multiple videos for viewing, each showing the subject from an angle corresponding to the number of cameras, are displayed on the viewing terminal 3N. Specifically, as shown in Figure 4, different videos for viewing are displayed on the screen, divided into multiple sections 41-45. In other words, this is an example of a multi-angle screen where videos for viewing from multiple cameras are simultaneously displayed on the viewing terminal 3N in a split-screen format. This corresponds to the number of cameras used in the photo shoot, and because the position, orientation, and angle of each camera relative to the subject are different, they result in different videos for viewing.

[0051] Participant 3 selects an image from a desired camera by tapping the display area of ​​the video feed from multiple cameras 6 shown on the viewing terminal 3N. Only the image from that camera is then enlarged and displayed on the entire screen or a portion of it 45, and participant 3 can press the shutter button 7 at any time they like. If they wish to select an image from another camera 6, they can repeat the process by clicking again from the multiple video feeds shown in Figure 4.

[0052] The multi-angle screen display shown in Figure 4 is extremely convenient for photo shoots involving a large number of participants. This is because each viewer can select video from their preferred camera angle (shooting angle and position), thus accommodating the preferences of almost all participants. Furthermore, the present invention adjusts the angle by switching the displayable area, from the perspective of accurately meeting the needs of a large number of participants. In other words, although the screen resolution on the participants' viewing terminals is generally set to FHD (1920 x 1080 pixels), the actual video shooting is based on ultra-high resolution and high-definition video quality, such as 8K (7680 x 4320 pixels) or 4K (3840 x 2160 pixels), allowing for changes to the displayable area.

[0053] Figure 5 shows examples of image ranges for FHD, 4K, and 8K. Each participant 3 can move between displayable areas, including the second area (4K) 52 and the third area (8K) 53, which are outside the first area (FHD range) 51, by performing finger operations such as tapping and pinching or mouse operations on the viewing terminal to change the area. This allows for flexible adaptation to the desired field of view (wide-angle, standard, telephoto) for each viewer. As a result, it is possible to provide viewers with an experience as if they were adjusting the field of view using the camera 6 at the shooting location.

[0054] Furthermore, the resolution of the recorded video is not fixed or limited to 8K or 4K, and it goes without saying that it will increase in line with advancements in communication technology and LCD display technology. In addition, the settings for the shooting size and framing ratio when viewers change the displayable area on their viewing device can be changed to any value desired by each viewer, in accordance with the framing of the LCD monitor on the viewing device.

[0055] The ability to change the displayable area by switching image quality, as described above, offers significant convenience, especially when multiple viewers are taking pictures based on images captured by a single camera. This is a photo shoot with one camera and many viewers. A fixed angle from a single camera makes it difficult to accommodate the wishes of many viewers, but according to the present invention, each viewer can freely select the area size on their viewing terminal based on the moving image from a single camera. However, this is also effective in cases where multiple cameras are used, as in this embodiment. Since it is equivalent to multiple cameras capturing images at a further multiplier of N, it becomes possible to precisely respond to each viewer's instructions for shooting operations such as zooming in / out, no matter how large the number of viewers.

[0056] The present invention allows viewers to take a picture at their preferred timing from a streamed video, but it is based on the premise that a large number of viewers are watching "simultaneously." In such a viewing situation, the images taken at the moment when multiple participants 3 are taking pictures are often considered to be the optimal scenes to capture, i.e., to be extracted images.

[0057] However, each participant in a remotely filmed video for viewing does not know when other participants are taking pictures. If you are actually at a filming location, you can see that many other people are taking pictures, so you can recognize when it's a good opportunity to take a picture and take your own. However, with a still image generation system like the present invention, you cannot know when others are taking pictures, so you may regret later that you should have taken a picture.

[0058] Therefore, the still image generation system may include a configuration that allows users to see the shutter status of other participants in real time by utilizing the fact that the time of each participant's shutter operation is transmitted to the server in real time. For example, as shown in Figure 6, shutter indicators 61-65 are provided on each of the multi-angle screens 41-45, and an indicator value representing the current total number of shutters is displayed. This indicates, for example, the total number of shutters from the present to 5 seconds ago, and is updated every 5 seconds. Here, 5 seconds is just an example and can be set to any number of seconds. Suppose participant A planned to take a shutter from the viewing video with the camera angle corresponding to section 41, but after looking at the shutter indicator 62 on the multi-angle screen, A realizes that many participants are taking shutters on the viewing video with the camera angle corresponding to section 42. Viewer A can then switch from the camera angle of section 41 to the camera angle of section 42 and take a shutter in the same way as other viewers who are focusing on that scene.

[0059] Other methods for displaying the indicator besides those shown in Figure 6 include, for example, highlighting the corresponding screen frames 41-45 on the multi-angle screen when a predetermined percentage of participants (e.g., more than half) are taking pictures with a camera, or making the size of screen frames 41-45 relatively larger to make them more noticeable. Alternatively, the indicator could be a curve graph showing the number of shutter clicks.

[0060] The still image generation system uses volumetric technology to capture 360-degree images of a subject using multiple cameras 6. The server generates 3D data from these images and distributes it to the viewing terminal 3N. The volumetric technology involves setting up a 360-degree green screen as the shooting background, photographing the subject, and then generating 3D data afterward. By shooting against a green screen, the video data of the subject can be composited into a virtual background such as VR (Virtual Reality), AR (Augmented Reality), or MR (Mixed Reality), allowing for flexible image processing and the addition of effects. Alternatively, even without shooting against a green screen, the system can recognize a person from an arbitrary background, crop it with optimal framing, and composite it with other images.

[0061] This allows participants to experience the freedom to move freely in all directions—forward, backward, left, right, up, and down—from any angle and viewpoint while operating the shutter. Using volumetric shooting data, it's possible to switch between free-viewpoint footage and multiple free-viewpoint footage, such as having the camera pass through subjects as if in a video game, or viewing them from above.

[0062] The still image generation system, based on the multiple embodiments described above, enables a single video stream distributed to a large number of viewers to respond to all viewer requests, such as instructions for shooting operations including shooting angle, position, and zoom in / out. It also assists viewers in freely selecting the optimal shot. Furthermore, it enables virtual photo shoots in virtual reality worlds (such as the metaverse) and the operation of such photo studios. It can provide a realistic shooting experience even in metaverse space, enabling shooting and collaboration in the digital world.

[0063] (Second Embodiment) The still image generation system 100 described above has been explained as saving the video footage taken at the shooting venue to a server, editing it as appropriate to create a video for viewing, and then providing it to the user 3 in the form of recorded data. Furthermore, it was assumed that still images for shutter output would be prepared in advance before viewing, that is, that reconstructed still images consisting of frame images with NG cuts removed would be prepared in advance before distributing the video for viewing.

[0064] On the other hand, there is a need to watch live video footage shot at a concert venue in real time, rather than recording it, and to obtain desired shots by remotely controlling the shutter. In this case, the captured video footage is not saved to a server but is delivered directly to the viewer as images for viewing. Even with live streaming, it is possible to divide the viewing image into multiple still images at a predetermined frame rate, automatically removing any unsuitable shots, and composing still images for shutter output in real time, extracting still images according to the shutter timing. Furthermore, if high-speed image processing is possible, editing such as noise reduction, image stabilization, overlays, and removal of effects can be performed automatically. An example of this procedure is shown below.

[0065] Figure 8 shows the time lag between the distribution of images for viewing and the generation of still images for shutter output. Let's assume that images for viewing, taken at a live venue, etc., are sent to a server and received by the server at time T0. Between time T0 and a predetermined time α (for example, 15 or 30 seconds), the server automatically performs the NG cut removal process described in Figure 3 in the first embodiment on the received images for viewing. Specifically, for example, for each still image divided at a predetermined frame rate, if the face recognition software determines that there is a person's face, it identifies whether the person's eyes are closed or not. If the eyes are closed, the still image is considered an NG cut and is not included in the still images for shutter output (corresponding to the frame images at the dotted lines). Furthermore, instead of simply removing NG cuts, AI software may be used to correct NG cuts. For example, if an image with closed eyes is detected, the software can identify the subject's eyes from the preceding and succeeding images and automatically correct them to an open-eyed state, resulting in a better expression. In other words, AI image processing is performed in real time to correct (modify) still images that may become NG cuts.

[0066] Furthermore, if the contrast value of a part of the still image, such as a person or face, deviates by a certain amount compared to the average contrast value of all still images, the still image is either too bright or too dark and is marked as an NG cut, or the part that is too bright or too dark is corrected to harmonize with the other parts. The same applies to sharpness; if it is determined that a part of the image, such as a person or face, is blurred below a predetermined value, it is marked as an NG cut, or the blurred area is corrected to make it clearer. In addition, an AI application may be used to mark any still images with inappropriate posing as NG cuts or to correct them. This automatic removal of NG cuts, as well as arbitrary correction processes such as blur correction, skin correction, background compositing, effect addition, and resolution improvement, are performed on a group of still images divided at a predetermined frame rate over a time interval α. For example, at 30fps, there are 30 images per second, so if the time interval α is 15 seconds, 450 images are processed. The automatic removal of NG cuts and arbitrary correction processes described above are also applicable to the first embodiment. It would be more efficient if the process of creating still images for shutter output, starting from the viewing video stored on the server, could include automatic removal of NG cuts and arbitrary correction processing.

[0067] Next, the server will determine the time T after the time interval α has elapsed. 1 At this time, the system begins distributing the viewing image via the communication network to be viewed on user 3's viewing terminal 3N. In other words, the system is configured to perform automatic removal of NG cuts α time in advance, and then distribute the viewing image with a time delay of α. The server is at time T 1 While streaming the viewing image, simultaneously at time T 1From there, the automatic removal of NG cuts at the next time interval α is executed in parallel. With this configuration, the effects of the present invention can be achieved without having to save the reconstructed image with NG cuts removed to the server and prepare the still image for shutter output in advance before viewing. Depending on the time required for the automatic removal and correction of NG cuts, the distribution of the viewing image may be adjusted to time T0+2α or time T0+3α, etc. Furthermore, if the effect removal process shown in step S21 and the retouching process (including super-resolution processing) shown in step S25 of Figure 2 can be completed together with the automatic removal of NG cuts before the distribution of the viewing image, they may be included as additional processes.

[0068] In all of the embodiments described above, the subject is captured through the user's own shutter operation, and multiple still images are obtained that do not include rejected shots, etc. However, if obtaining the still images is subject to a fee, the participant will have to choose a limited number of images from among the multiple still images. The still image generation system 100 has the following mechanism to facilitate this selection according to the participant's request.

[0069] (1) Batch Selection After the shooting session ends, AI software that analyzes the subject's facial expressions and poses is used to extract the best shot from the multiple still images taken by the participant. When a participant specifies a desired number of images, the best images are extracted in a batch, starting with those with the highest evaluation scores. The extraction criteria are registered in advance. (2) Real-time Selection Immediately after a shutter operation, AI software is used to select the still image extracted from that shutter operation as a best shot candidate if it meets or exceeds a predetermined standard. The best shot criteria are registered in advance. If the still image extracted from the next shutter operation meets or exceeds the predetermined standard and is also a best shot candidate, it is compared with the previous best shot candidate to determine which is the best shot. This is repeated sequentially. In the above repetition, it is also possible to compare multiple best shot candidate images, rank them, and select the desired number of images from the top. This is suitable for real-time viewing and image distribution as in the second embodiment. (3) Recommended Shot Suggestion Participant preferences are acquired as data in advance, and still images that match each participant's preferences are suggested after shooting or during real-time viewing. (4) Random Shutter In addition to the shutter operation by the participant (or with the participant's shutter operation disabled so that the participant is unaware of it), the AI ​​software automatically operates the shutter according to predetermined rules. From the still images extracted by the shutter operation, select based on (1) or (2) above.

[0070] The above (1) to (4) will meet the needs of participants who find photography enjoyable but find selecting still images difficult and would prefer to leave it to us.

[0071] The still image generation system of the present invention is realized by processes, means, and functions executed by a computer in accordance with instructions from a program (software) launched on the system. The program can send commands to each component of the computer, causing it to execute predetermined processes, functions, etc., according to the present invention as described above. In other words, each process, means, and function in the present invention is realized by specific means in which the program and the computer work together. Furthermore, all or part of the program is provided on a recording medium that can be read by any computer, such as a magnetic disk like a CD-ROM, an optical disk, a semiconductor memory, and so on, and the program read from the recording medium is installed and executed on the computer. Alternatively, the program can be loaded directly onto the computer via a communication line without using a recording medium and executed. Therefore, the program installed or loaded onto the computer, and these storage media, are included within the scope of the invention.

[0072] Furthermore, the processing performed by a server (e.g., a single personal computer) in the still image generation system according to the present invention may be performed by a mobile phone, personal computer, or data relay device. It can also be configured with a single server or with multiple servers (e.g., a group of multiple server computers). The above embodiments are merely examples to illustrate the present invention clearly. Moreover, the viewing terminal that communicates with the still image generation system via a communication network is a computer connected to a network such as the Internet or a dedicated line. Specifically, examples include PCs (Personal Computers), mobile phones and smartphones, PDAs (Personal Digital Assistants), tablets, and wearable devices. A business scheme including the still image generation system is formed by configuring PC terminals and mobile terminals connected to the communication network via wired or wireless connections to enable communication with each other. In the embodiments described above, the still image generation system may be configured to cooperate with an ASP (Application Service Provider).

[0073] 1. Event management company for filming 2. Platform management company 3. Participants 3N Viewing terminals 4. Communication network 6. Camera 7. Shutter icon 100 Still image generation system

Claims

1. A still image generation system that provides a simulated shooting experience by distributing a video for viewing, created based on a video image acquired by a shooting means, to multiple users' image display terminals via a communication network, wherein the system generates a shutter output still image by decomposing the video image into multiple still images based on a predetermined frame rate and outputting them, generating a reconstructed still image based on a still image obtained by removing some of the multiple still images or by modifying some of the still images, and when it is determined that one of the multiple users has performed a finger action equivalent to a shutter operation on the image display terminal while viewing the video for viewing, the system identifies a still image corresponding to the timing of the shutter operation from the shutter output still image, and displays the identified still image on the user's image display terminal.

2. The still image generation system according to claim 1, wherein the reconstructed still image has had removed any still images that were determined to show the user with their eyes closed or any still images that are not permitted to be provided to the user.

3. The still image generation system according to claim 1, comprising a plurality of shooting means, wherein viewing videos corresponding to each shooting means are simultaneously displayed on each of the plurality of users' image display terminals, and a viewing video from a desired shooting means can be selected, and the displayable area of ​​the viewing video is dynamically switched for each user according to the specifications of each of the plurality of users during the distribution of the viewing video.

4. The still image generation system according to claim 3, wherein an indicator that can recognize in real time when another user is operating the shutter on the video for viewing is displayed on the image display terminal.

5. The still image generation system according to claim 1, further comprising analyzing the preferences of the multiple users based on the number of finger actions corresponding to the shutter operation, or providing a product or service based on aggregated data of the number of actions.

6. The still image generation system according to claim 1, which synthesizes the captured video images acquired by the shooting means into a digital virtual space displaying an arbitrary background and distributes them to the image display terminal.

7. The still image generation system according to claim 1, wherein, during or after the acquisition of the captured video image, it is displayed on the image display terminal to identify whether or not one or more specified still images can be exclusively purchased, and if any of the multiple users select a still image that can be exclusively purchased, the other users are prevented from purchasing the selected still image.

8. The still image generation system according to claim 1, wherein the still image displayed on the user's image display terminal includes information that identifies the user or desired information specified by the user.

9. The still image generation system according to claim 1, wherein the still images purchased by the user are affixed with a digital signature or digital certificate.

10. The still image generation system according to claim 1, wherein the still image purchased by the user is transferred to a selected item.

11. The still image generation system according to claim 1, wherein if the still image corresponding to the timing of the shutter operation matches a pre-set good shot image, a predetermined display or sound is output on the image display terminal of the user who performed the shutter operation to allow recognition.

Citation Information

Patent Citations

  • Information processing apparatus, management device, information processing method, and program

    JP2016025633A

  • Remote photography system and remote photography method

    JP7288641B1

  • Image data providing device and image data providing method

    WO2004077830A1

  • Live broadcasting system

    WO2021246498A1

  • Information processing device, information processing method, and information processing system

    WO2022014170A1