Image processing system, image processing method, program, and recording medium
Patent Information
- Application Number
- PCT/JP2025/041110
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-19
- Filing Date
- 2025-11-26
- Publication Date
- 2026-08-27
Smart Images

Figure JP2025041110_27082026_PF_FP_ABST
Abstract
Description
Image processing system, image processing method, program, and recording medium
[0001] The present invention relates to an image processing system, an image processing method, a program, and a recording medium capable of processing a composite video composed of multiple videos.
[0002] Users may shoot multiple videos, combine them temporally to create a single composite video, and then use that composite video. Users may also select and print still images representing representative scenes from the videos. A technology to assist users in such cases is, for example, the technology described in Patent Document 1.
[0003] In the technology described in Patent Document 1, a video image printing device extracts characteristic sections from moving image data to create a digest video. The video image printing device also stores multiple still image data of the subject captured in the moving image data. Furthermore, based on information about the section used to create the digest video from the moving image data, the video image printing device selects still image data captured at a temporally or spatially close position to be used as representative still image data, and prints the representative still image data.
[0004] According to the above technology, when printing a representative image representing a digest of a video, it is possible to print a still image that is easy to understand the content of the digest and has high resolution, making it excellent for viewing.
[0005] Japanese Patent Publication No. 2010-258487
[0006] However, in order to utilize the technology described in Patent Document 1, it is necessary to capture moving images and separately capture still images that are temporally or spatially close to a digest of the moving images, which imposes a burden on the user who is the photographer. Therefore, when selecting and acquiring representative images, it is necessary to reduce the burden on the user.
[0007] One embodiment of the present invention has been made in view of the above circumstances, and aims to provide an image processing system, an image processing method, a program, and a recording medium that can appropriately and easily select a still image to be printed from a composite video composed of multiple videos.
[0008] The above objective is achieved by the image processing system described in [1] below. [1] An image processing system having a processor and a printer, wherein the processor acquires a composite video composed of multiple videos, analyzes the composite video or the multiple videos to calculate an evaluation value for a first image which is a still image included in the composite video, extracts a second image from the first image whose evaluation value is calculated and whose evaluation value satisfies the extraction conditions, and prints the second image to the printer.
[0009] Furthermore, in one embodiment of the present invention, an image processing system described in any of the following [2] to
[17] can be provided. [2] The image processing system described in [1], wherein the processor transmits a synthesized video to the destination. [3] The image processing system described in [2], wherein the processor transmits a synthesized video to the destination when the second image is printed by a printer. [4] The image processing system described in [2] or [3], wherein the processor transmits a derived image generated based on the synthesized video to the destination. [5] The image processing system described in any of the following [2] to [4], wherein the processor causes the printer to print a second image on which information regarding the storage area of the synthesized video at the destination is superimposed. [6] The image processing system described in any of the following [1] to [5], wherein the processor calculates evaluation values for two or more first images, extracts one or more second images from the two or more first images on which evaluation values have been calculated, and causes the printer to print a third image selected by the user from among the one or more second images. [7] The image processing system described in [6], wherein the processor has a storage device for storing images, and causes the printer to print one of the images stored in the storage device. [8] The image processing system according to [7], wherein a third image is stored in the storage device. [9] The image processing system according to [7] or [8], wherein a second image different from the third image is stored in the storage device from among two or more second images.
[10] The image processing system according to any one of [1] to [9], comprising a video generation device that shoots multiple videos and generates a composite video based on effect information selected from among multiple effect information, the video generation device classifies the composite video based on the effect information selected when generating the composite video.
[11] The image processing system according to
[10] , wherein the printer and the video generation device are integrated.
[12] The image processing system according to
[10] or
[11] , wherein the video generation device generates a generated image based on two or more of the composite videos classified into the same classification category.
[13] The image processing system according to any one of
[10] to
[12] , wherein the video generation device stores the effect information selected when generating a composite video in association with the time when the composite video was generated based on the effect information, and determines the effect information to be used when generating the next composite video from among multiple pieces of effect information based on the relationship between the effect information and the time.
[14] The image processing system according to any one of [1] to
[13] , wherein the processor displays a question on the screen regarding the image quality of the image output by the printer after the second image has been printed by the printer, and determines the print conditions for when the printer prints the next image according to the user's answer to the question.
[15] The image processing system according to
[14] , wherein the processor determines the print conditions based on the video generation information and derived information derived from the answer regarding the generation of a composite video including the printed second image.
[16] The image processing system according to any one of [1] to
[15] , wherein, by playing a composite video, multiple videos are displayed in order, the processor extracts two or more second images from the composite video, and the processor does not extract as a second image a first image that is displayed within a predetermined time from the time a first image extracted as a second image is displayed during the playback time of the composite video.
[17] The image processing system according to any one of [1] to
[15] , wherein the processor analyzes two or more videos from among multiple videos, calculates an evaluation value for the first image for each of the two or more videos, and extracts a second image from the first image for which an evaluation value has been calculated for each of the two or more videos.
[0010] Furthermore, the above-mentioned objectives are achieved by an image processing method described in any of the following
[18] to
[20] .
[18] An image processing method comprising: a step of a processor acquiring a composite video composed of multiple videos; a step of a processor analyzing the composite video or multiple videos to calculate an evaluation value for a first image which is a still image included in the composite video; a step of a processor extracting a second image from the first image for which an evaluation value has been calculated, the second image whose evaluation value satisfies the extraction conditions; and a step of a processor causing a printer to print the second image.
[19] The image processing method according to
[18] , wherein the processor transmits the composite video to a destination.
[20] The image processing method according to
[18] or
[19] , wherein in the step of calculating an evaluation value, an evaluation value is calculated for two or more first images; in the step of extracting a second image, one or more second images are extracted from the two or more first images for which an evaluation value has been calculated; and in the step of causing a printer to print a second image, a third image selected by the user from among the one or more second images is printed by the printer.
[0011] Furthermore, the program according to one embodiment of the present invention is a program that causes a computer to perform each step included in the image processing method described in any of
[18] to
[20] . Furthermore, the recording medium according to one embodiment of the present invention is a computer-readable recording medium on which a program that causes a computer to perform each step included in the image processing method described in any of
[18] to
[20] is recorded.
[0012] According to one embodiment of the present invention, it is possible to appropriately and easily select a still image to be printed from a composite video composed of multiple videos.
[0013] This is a diagram showing the configuration of an image processing system according to one embodiment of the present invention. This is an explanatory diagram about a composite video. This is an explanatory diagram about the processes performed in the image processing system. This is a diagram showing an example of a printed object with a second image printed on it. This is an explanatory diagram about the procedure for classifying composite videos and generating an integrated video. This is an explanatory diagram about the procedure for determining recommended effect information from among multiple effect information. This is a diagram showing an example of a screen displaying questions about the image quality of a printed image. This is an explanatory diagram about the evaluation of the first image. This is a diagram showing an example of a screen displayed to ask the user about the evaluation criteria they consider important. This is a diagram showing the processing flow for extracting the second image. This is a diagram showing an example of a selection screen for the extracted second image. This is an explanatory diagram about derived information. This is a diagram showing a modified example of the method for extracting the second image.
[0014] One embodiment of the present invention relates to an image processing system and an image processing method. Furthermore, one embodiment of the present invention can also be applied to programs and program products, as well as recording media on which programs are stored. Specific embodiments of the present invention will be described below. For convenience of explanation, the content of the present invention may be described from the perspective of a GUI (Graphical User Interface) in the following description.
[0015] Furthermore, the fundamental data processing technologies for realizing the present invention (communication / transmission technologies, data acquisition technologies, data recording technologies, data processing / analysis technologies, machine learning technologies, image processing technologies, image display technologies, and visualization technologies, etc.) are known technologies, and therefore, explanations thereof will be omitted. Regarding the equipment and devices used to apply the above-mentioned known technologies, it is advisable to appropriately select those available at the time of implementing the present invention.
[0016] Furthermore, in this specification, the concept of "device" includes both a single device that performs a specific function and a combination of multiple devices that exist independently and in a distributed manner while cooperating (linking) to perform a specific function. Furthermore, in this specification, the concept of "system" includes both a system composed of multiple devices connected in a manner that allows them to communicate with each other and a system composed of a single device.
[0017] Furthermore, in this invention, "user" refers to a person who uses the image processing system of the present invention for the purpose of enjoying the benefits of implementing the present invention. Specifically, a person who captures images (especially videos) through the image processing system of the present invention, makes images transmittable, displays images on a screen, views or appreciates images, or prints images (especially still images) is considered a user of the present invention.
[0018] Furthermore, in this specification, "person" means an entity that performs a specific action, and includes individuals, groups, corporations and other legal entities, and organizations, and may also include computers and devices that constitute artificial intelligence (AI). Artificial intelligence (AI) is a system that realizes intelligent functions such as reasoning, prediction, and judgment using hardware and software resources.
[0019] Furthermore, within this specification, machine learning algorithms may include neural networks, convolutional neural networks, recurrent neural networks, attention, transformers, variational autoencoders, generative adversarial networks, deep learning neural networks, Boltzmann machines, matrix factoryization, factoryization machines, MWAY factoryization machines, field-aware factoryization machines, field-aware neural factoryization machines, support vector machines, Bayesian networks, decision trees, and random forests, as well as other machine learning algorithms.
[0020] Furthermore, in this specification, unless otherwise specified, "image" refers to digital image data that can be processed (hereinafter simply referred to as "image data"). Images as image data include both videos and still images. Videos include, for example, image data of videos stored in file formats such as MPEG (Moving Picture Experts Group)-1, MPEG-2, MPEG-3, MPEG-4, MPEG-4 Part 14 (MP4), FLV (Flash Video), MOV (Movie), and AVI (Audio Video Interleaved). Videos also include videos with sound and videos without sound. Still images include, for example, image data of still images compressed using lossy compression such as JPEG (Joint Photographic Experts Group) format, and image data of still images compressed using lossless compression such as GIF (Graphics Interchange Format) or PNG (Portable Network Graphics) format.
[0021] <<About One Embodiment of the Invention>> Hereinafter, a specific embodiment of the present invention (hereinafter, this embodiment) will be described. (Basic configuration of the image processing system according to this embodiment) The basic configuration of the image processing system according to this embodiment (hereinafter, the image processing system 100) will be described with reference to Figures 1 and 2. Figure 1 is a diagram showing the configuration of the image processing system according to this embodiment. Figure 2 is an explanatory diagram of the composite video described later.
[0022] As shown in Figure 1, the image processing system 100 includes a video recorder 10, a printer 20, a user terminal 30, and a server 50. These devices are connected to each other in a manner that allows them to communicate with one another via external networks such as the Internet and mobile communication networks, and communication standards such as Wi-Fi® and Bluetooth®.
[0023] Furthermore, in this embodiment, the video recorder 10 and the printer 20 are integrated as a single device (i.e., a video recorder with a printer), and specifically, for example, the two devices are housed in the same enclosure. However, this is not limited to this, and the video recorder 10 and the printer 20 may be separated as separate devices. For the sake of explanation, the configurations of the video recorder 10 and the printer 20 will be described separately below.
[0024] [Video Recorder] The video recorder 10 is a video generation device that is operated by the user to shoot video, and is composed of, for example, a digital video camera. The video recorder 10 is equipped with a touch panel display. The user can set various contents related to the video recorder 10 (for example, shooting conditions and effect information described later) through the operation screen displayed on this display.
[0025] The video recorder 10 has the functions of video recording, adding effects, and video classification. These functions are realized through the cooperation of the hardware equipment (including the processor) provided in the video recorder 10 and the control program as software implemented in the video recorder 10.
[0026] The video recording function is a function that generates video image data by recording a video consisting of multiple frame images at a predetermined frame rate. The generated video image data file incorporates information such as the date and time of recording, the location of recording, the recording conditions, the equipment used for recording (i.e., the video recorder 10), and other incidental information related to the recording. The incidental information data corresponds to the metadata of the recorded video.
[0027] Furthermore, as shown in Figure 2, the video recorder 10 can generate image data for a single video (hereinafter referred to as a composite video) by combining multiple videos obtained by intermittently shooting multiple scenes in chronological order. Specifically, when the user presses the record button (not shown) of the video recorder 10 once, video recording begins. After recording begins, when the user presses the pause button (not shown) of the video recorder 10, video recording is paused. If the user then presses the pause button again, video recording resumes. The pausing and resumption of video recording is repeated according to the user's button operation, and finally, when the user presses the stop button (not shown) of the video recorder 10, video recording ends. Based on this series of operations, the video recorder 10 generates a composite video by concatenating the videos of the scenes shot up to the point of pause and the videos of the scenes shot after recording resumes.
[0028] As shown in Figure 2, the composite video is composed of multiple videos (hereinafter also referred to as unit videos) shot according to the procedure described above, and when the composite video is played, the multiple unit videos are displayed in order. In other words, when the composite video is played, each unit video is displayed in the order of their respective shooting dates and times, and the audio in each unit video is output according to the time axis of each unit video. Each of the multiple unit videos may be a relatively short video of a few seconds or a few minutes, or a relatively long video of several tens of minutes. Furthermore, there is no particular limit to the number of unit videos included in the composite video; two or more are sufficient. In addition, the multiple unit videos may be videos shot at the same shooting location but depicting different scenes (for example, landscapes or subjects at different times), or they may be videos shot at different shooting locations.
[0029] The effect addition function is a function that adds effects to each of the multiple unit videos that make up a composite video during the generation of a composite video. Effects refer to the video quality (specifically, the gradation, saturation, hue, contrast, resolution, presence or absence of noise, presence or absence of blur, and presence or absence of camera shake, etc.) of the frame image, the video playback speed and frame rate, and other visual effects that appear during video playback. The effects added to a video are changed by adjusting the shooting conditions during video recording (specifically, exposure conditions, zoom, type of lens used, angle of view, focus position, and focal length, etc.) and the image processing conditions (filtering and color correction, etc.) applied to the frame images that make up the video.
[0030] Regarding effects, multiple candidates (hereinafter referred to as effect information) may be prepared in advance, in which case the user can select one of the multiple effect information. The video recorder 10 then shoots multiple unit videos under shooting conditions corresponding to the effect information selected by the user (hereinafter referred to as selected effect information), and performs image processing corresponding to the selected effect information on the frame images in each unit video. This generates a composite video with the effect according to the selected effect information added. Note that the shooting conditions and image processing content may be set and changed directly by the user. In that case, the video recorder 10 will generate a composite video with the effect according to the shooting conditions and image processing set by the user.
[0031] The video classification function is a function for classifying composite videos. With this function, for example, already generated composite videos can be classified by category. A category is a classification division when classifying composite videos according to a predetermined procedure. Two or more composite videos classified into the same category can be linked together to generate a single video (hereinafter referred to as a "combined video") (see Figure 5). A combined video is an example of a generated image produced based on two or more composite videos classified into the same category. When a combined video is played, the multiple composite videos that make up the combined video are displayed in order of their generation date and time.
[0032] [Printer] The printer 20 is a device that prints images (more precisely, still images), and is, for example, an instant photograph type printer that houses a photosensitive film inside and forms (prints) an image on the image-forming surface of the photosensitive film using silver halide technology. However, the image printing method is not particularly limited and may be an inkjet method, a dye-sublimation thermal transfer method, or an electrophotographic method using toner, etc.
[0033] The photosensitive film on which the image is printed (hereinafter referred to as "printed material P") is ejected from the printer 20, and the user receives the ejected printed material P (see Figure 4). This allows the user to view the image printed on the printed material P (hereinafter referred to as "printed image"). Furthermore, if code information such as a two-dimensional barcode is printed as part of the printed image, the user can read that code information with the user terminal 30 and, for example, access the URL (Uniform Resource Locator) indicated by that code information to obtain the information present at that URL.
[0034] Furthermore, when the printer 20 starts printing an image (i.e., generating a printed object P), it communicates with the user terminal 30 and notifies the user terminal 30 that printing has started. However, it is not limited to this, and the printer 20 does not have to notify the user terminal 30 that printing has started.
[0035] [User Terminal] The user terminal 30 is a terminal operated by the user, and specifically consists of a PC (Personal Computer), smartphone, mobile phone, tablet terminal, wearable device, television receiver with communication function, or store-installed terminal, etc. The number of user terminals 30 used by one user may be one or two or more.
[0036] As shown in Figure 1, the user terminal 30 includes a processor 30A, memory 30B, communication interface 30C, storage 30D, input device 30E, and output device 30F, etc.
[0037] The processor 30A is constituted by, for example, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), a TPU (Tensor Processing Unit), or the like. The memory 30B is constituted by, for example, semiconductor memories such as a ROM (Read Only Memory) and a RAM (Random Access Memory). The communication interface 30C is constituted by, for example, a network interface card or a communication interface board or the like.
[0038] The storage 30D may be constituted by a non-volatile storage medium such as an HDD (Hard Disc Drive), an SSD (Solid State Drive), a memory card, or a USB memory (Universal Serial Bus memory). The storage 30D may be built in the user terminal 30 or may be externally attached to the terminal body. The input device 30E may be constituted by, for example, a keyboard, a mouse, a touch panel, a touch pad, a camera, and a microphone or the like. The output device 30F may be constituted by, for example, a display and a speaker or the like. Also, the type of the display is not particularly limited, but in the present embodiment, it is assumed to be an LCD (Liquid Crystal Display).
[0039] Further, an application program (hereinafter, an image processing application) for causing the user terminal 30 to function as an image processing device is installed. That is, the image processing device according to the present embodiment is realized by a processor and a program executable by the processor, and is constituted by, for example, a general-purpose computer.
[0040] The image processing application is a program for causing a computer, specifically the processor 30A of the user terminal 30, to execute each step included in the image processing method of the present invention. That is, the processor 30A reads out the image processing application and executes a series of data processing in the image processing flow described later. The image processing application may be obtained by being read from a computer-readable recording medium, or may be obtained by being received (downloaded) through a communication line such as the Internet or an intranet.
[0041] The user terminal 30 as an image processing apparatus has functions of image reception, image editing, image analysis, image extraction, printer control, image reading, image transmission, and video playback. These functions can be realized by the cooperation of the hardware devices installed in the user terminal 30 and the image processing application as software installed in the user terminal 30.
[0042] The function of image reception is a function of receiving image data such as a composite video generated by the video recorder 10, the integrated video described above, and a derivative image described later from the video recorder 10. In other words, it is a function of acquiring a composite video or the like.
[0043] The function of image editing is a function of editing the composite video received from the video recorder 10. For example, it is a function of extracting a lively scene in the composite video and generating a shortened video or a digest video. The shortened video or digest video generated by this function may be combined with the composite video so as to be displayed as an additional video (specifically, an opening video or an end roll video) at the beginning or the end of the composite video. Also, subtitles may be added to the composite video (strictly speaking, the frame images included in the composite video) by the function of image editing. In this case, the subtitles may be text information input by the user, or may be text information representing the content of the composite video (for example, the subject and the shooting scene, etc.) specified by analyzing the composite video.
[0044] The image analysis function analyzes the composite video received from the video recorder 10. More specifically, it is an analysis function for evaluating frame images (hereinafter also referred to as the first image), which are still images contained in each of the multiple unit videos that make up the composite video. The image extraction function extracts first images that satisfy predetermined extraction conditions from the first images to be evaluated as second images. The extracted second images are representative images that represent characteristic scenes in the composite video and can be edited using the image editing function described above. For example, by combining them with a frame-shaped template image, they can be edited into a second image with a frame. The evaluation of the first images and the extraction of the second images will be explained in detail in a later section.
[0045] The printer control function controls the aforementioned printer 20 to print images. More specifically, it generates print control data and sends it to the printer 20, causing the printer 20 to perform print operations according to the control data. The image to be printed by the printer 20 using this function is, for example, an image selected by the user from among the second images extracted by the image extraction function described above. In this case, the second image may be edited with a frame-shaped template image as described above and printed with a frame (for example, as an image for the jacket of a composite video).
[0046] The image reading function reads code information, such as a two-dimensional barcode, from the printed image in the printed material P generated by the printer 20. More specifically, the user terminal 30 reads the code information using a camera mounted on the user terminal 30 and analyzes the read image (still image). This allows the user terminal 30 to identify the URL indicated by the code information, access the storage area indicated by that URL, and request the distribution of information stored in that storage area. The information requested by the image reading function may include video stored in a predetermined storage area of the server 50 (more specifically, the storage area identified by the URL).
[0047] The image transmission function is a function that transmits (uploads) an image to the server 50 and stores the image in a predetermined memory area of the server 50. Images transmitted by this function include composite videos and integrated videos received from the video recorder 10, as well as second images extracted by the image extraction function of the user terminal 30. The video playback function is a function that receives a distributed video when a video stored in a predetermined memory area of the server 50 (for example, the memory area indicated by the URL identified by the image reading function) is distributed, and plays the distributed video through the output device 30F of the user terminal 30. With this function, the user can view or enjoy videos distributed from the server 50 using the user terminal 30.
[0048] [Server] Server 50 corresponds to a storage device that stores images. It receives the composite video and integrated video generated by the video recorder 10 via the user terminal 30, and stores the received video in a predetermined storage area within Server 50 so that it can be accessed. In other words, Server 50 corresponds to the destination of the composite video and integrated video.
[0049] Furthermore, if the user terminal 30 accesses the storage area of the server 50 where the video is stored and requests the server 50 to distribute the video, the server 50 distributes the video to the requesting user terminal 30. In addition, the server 50 stores still images as images, specifically the third image described later. The still images stored in the server 50 are images transmitted from the user terminal 30, or more specifically, images that the user terminal 30 has the printer 20 print. In other words, the user terminal 30 has the printer 20 print one of the images stored in the server 50.
[0050] (Regarding image processing performed in the image processing system) Next, the image processing performed in the image processing system 100 will be explained with reference to Figures 3 to 7.
[0051] In the image processing system 100, the user can shoot multiple unit videos using the video recorder 10. The video recorder 10 combines the multiple unit videos that have been shot to generate a composite video and transmits the composite video to the user terminal 30. The composite video, and each unit video that constitutes it, includes two or more frame images (i.e., a first image), as shown in Figure 3. The user may also select one of several effect information as the selected effect information when shooting a video. In this case, the video recorder 10 shoots a video under shooting conditions corresponding to the user's selected effect information and performs image processing corresponding to the selected effect information to generate a composite video with the effect according to the selected effect information added.
[0052] Subsequently, when the user operates the user terminal 30 to launch the image processing application and perform a predetermined operation, the processor 30A of the user terminal 30 (hereinafter simply referred to as processor 30A) executes an evaluation process on the composite video received from the video recorder 10. The evaluation process evaluates two or more first images contained in each of the multiple unit videos that make up the composite video according to multiple evaluation criteria and calculates an evaluation value.
[0053] The processor 30A then extracts from the evaluated first images (in other words, the first images from which evaluation values have been calculated) an image whose evaluation value satisfies predetermined extraction conditions, and uses that image as the second image representing the synthesized video. In this embodiment, as shown in Figure 3, two or more second images (hereinafter also referred to as the second image group) are extracted for each synthesized video. The extracted second image group is presented to the user by being displayed on the user terminal 30's screen, and the user performs an operation to select one of the presented second images from the second image group through the input device 30E of the user terminal 30.
[0054] When the processor 30A of the user terminal 30 receives a selection operation from the user, it generates control data to print the second image (hereinafter referred to as the third image) selected by the user to the printer 20. At this time, as shown in Figure 4, the processor 30A generates control data so that the third image, on which a coded image Pb such as a two-dimensional barcode is superimposed, is printed. The coded image Pb is readable by the camera of the user terminal 30.
[0055] In this embodiment, by reading the coded image Pb, it is possible to access a predetermined storage area (hereinafter referred to as the first storage area) in the server 50 and request the transmission of a video, more specifically a composite video, stored in the first storage area. In other words, the coded image Pb corresponds to information about the storage area of the composite video in the server 50, and the processor 30A causes the printer 20 to print a second image (more precisely, a third image) on which the coded image Pb is superimposed. In addition, decorative images other than the coded image Pb, such as a frame-shaped template image placed on the edge of the third image to surround it, may be further superimposed on the printed third image.
[0056] Subsequently, the processor 30A controls the printer 20 based on the control data, and the printer 20 prints the third image under the control of the user terminal 30. As a result, as shown in Figures 3 and 4, a printed object P with the third image, which is the second image selected by the user, is ejected from the printer 20, and the user can obtain the printed object P. In addition, as shown in Figure 4, a coded image Pb for accessing the first storage area is superimposed on the printed image Pa in the printed object P.
[0057] Then, as shown in Figure 3, when the processor 30A prints the third image to the printer 20, it sends (uploads) the composite video to the server 50. More specifically, when printing of the third image begins, the printer 20 notifies the user terminal 30 of this. When the processor 30A receives the print start notification, it uses this as a trigger to send the composite video to the server 50. By linking the printing of the third image with the sending of the composite video to the server 50 in this way, the effort required from the user is minimized, and the sharing of the composite video (i.e., making it available for transmission to other users) can be easily facilitated.
[0058] Server 50 stores the composite video received from the user terminal 30 in its first memory area. The first memory area, which is the storage location for the composite video, is pre-configured before the composite video is sent (uploaded) to server 50. In other words, the composite video is stored in a predetermined memory area on server 50.
[0059] The processor 30A may also transmit the printed third image to the server 50 and have the server 50 store the third image in a predetermined area. In other words, the processor 30A can have the printer 20 print the third image, which is one of the images stored in the server 50, and in that case, the server 50 may store the printed third image. In this case, it is preferable that the third image be stored in the server 50 in association with the composite video that includes the third image.
[0060] After the synthesized video is stored in the server 50, the user who owns the printed material P, or another user who has received the printed material P from the user, reads the coded image Pb in the printed image Pa of the printed material P using the camera of the user terminal 30. In this case, the processor 30A accesses the first storage area of the server 50, which is identified by the URL indicated by the coded image Pb, and requests the synthesized video stored in the first storage area from the server 50. In response to this request, the server 50 transmits the image data of the synthesized video to the requesting user terminal 30, as shown in Figure 3. The user terminal 30 plays the synthesized video received from the server 50, displays the synthesized video on the display, and outputs the sound incorporated into the synthesized video from the speaker.
[0061] The series of processes described above constitute the basic processes executed by the image processing system 100 according to this embodiment. In addition to the basic processes described above, the image processing system 100 can also execute the additional processes described below.
[0062] The first additional process is the process by which the processor 30A generates a derived image based on the composite video received from the video recorder 10. The derived image may be an opening video added to the beginning of the composite video, an end-roll video added to the end of the composite video, or both of these videos. The derived image may also be a jacket image that represents the composite video among multiple first images (frame images) included in the composite video and can be used as an image for printing.
[0063] One example of a procedure for generating derived images is as follows: For each of the multiple first images included in the composite video, an evaluation value is calculated by performing an evaluation according to predetermined evaluation criteria, and a derived image is generated using one or more first images selected based on the evaluation value. If an opening video or an end-credits video is generated as a derived image, it is desirable that the video (derived image) be composited into the composite video so that it is displayed at least at the start or end of the composite video.
[0064] Furthermore, derived images and composite videos formed by combining derived images (hereinafter referred to as "derived images, etc.") are preferably transmitted (uploaded) from the user terminal 30 to the server 50 and stored in a predetermined memory area on the server 50 in an accessible state. The memory area for derived images, etc. on the server 50 (hereinafter also referred to as the second memory area) is a different area from the first memory area where the composite video is stored. In other words, derived images, etc. are stored separately from the composite video on the server 50. It is preferable that derived images, etc. be stored in the second memory area in association with the composite video stored in the first memory area.
[0065] Furthermore, when the user terminal 30 prints a second image (third image) selected by the user to the printer 20, the coded image Pb superimposed on the second image may indicate the URL of the second memory area. In this case, the user terminal 30 can access the second memory area by reading the printed image Pa in the printed material P ejected from the printer 20 with its camera, and can request the derived images etc. stored in the second memory area from the server 50. As a result, the user terminal 30 receives the derived images etc. sent from the server 50 and displays them on the display, and the user can view or appreciate the displayed derived images etc.
[0066] The second additional process involves the video recorder 10 (more specifically, the processor installed in the video recorder 10) classifying the composite video according to the effects applied to it. More specifically, as mentioned above, the user selects one of several effect information as the selected effect information, and the video recorder 10 generates a composite video by concatenating multiple unit videos to which the effect corresponding to the selected effect information has been applied. Subsequently, as shown in Figure 5, the video recorder 10 classifies the composite video into categories based on the effects applied to the composite video, more specifically, the selected effect information selected when shooting the multiple unit videos.
[0067] The video recorder 10 then performs processing according to the classification result. The processing method according to the classification result is not particularly limited. For example, as shown in Figure 5, the video recorder 10 may save the composite video in a folder corresponding to the classification result from among the folders prepared for each category. Alternatively, the video recorder 10 may incorporate data indicating the classification result into the data file of the composite video. Specifically, it may change the flag data corresponding to the classification result from "0" to "1" among the flag data defined for each category.
[0068] Furthermore, as shown in Figure 5, the video recorder 10 can generate a single integrated video by combining two or more composite videos classified into the same category. Specifically, it can combine two or more composite images to generate a single integrated video. Note that the generated image, which is generated based on two or more composite videos classified into the same category, is not limited to a video (integrated video); for example, it may be a still image composed of frame images (first images) extracted from each of the two or more composite videos.
[0069] The integrated video is transmitted from the video recorder 10 to the user terminal 30, and then transmitted (uploaded) from the user terminal 30 to the server 50, where it is stored in a predetermined storage area on the server 50 in an accessible state. The storage area for the integrated video on the server 50 (hereinafter also referred to as the third storage area) is a different area from both the first storage area where the composite video is stored and the second storage area where derived images, etc., are stored. In other words, on the server 50, the integrated video is stored separately from the composite video and derived images, etc.
[0070] Furthermore, when the user terminal 30 prints the second image (third image) selected by the user to the printer 20, the coded image Pb superimposed on the second image may indicate the URL of the third storage area. In this case, the user terminal 30 can access the third storage area by reading the printed image Pa in the printed material P ejected from the printer 20 with its camera, and request the integrated video stored in the third storage area from the server 50. As a result, the user terminal 30 receives and plays the integrated video sent from the server 50, and the user can view the played integrated video.
[0071] The third additional process involves the video recorder 10 (more specifically, the processor installed in the video recorder 10) determining the effect information to be used when generating the next composite video from among multiple effect information. More specifically, as shown in Figure 6, each time a composite video is generated, the video recorder 10 stores the selected effect information used when generating the composite video, associating it with the time when the composite video was generated based on that selected effect information. As a result, the storage device in the video recorder 10 stores a history of effect information selections showing the relationship between the selected effect information and the time when the composite video was generated, and a database is built. The video recorder 10 also analyzes the history of effect information selections to identify the user's preferences regarding effect information, or more specifically, the tendency of which effect information the user will select.
[0072] Then, when the user shoots multiple videos to generate a composite video, the video recorder 10 determines recommended effect information from among multiple effect information based on the user's preferences regarding selected effect information, that is, the user's tendency to select effect information, as shown in Figure 6. The recommended effect information is the effect information to be used when generating the next composite video, and is, for example, the effect information that the user frequently selects. As shown in Figure 6, the video recorder 10 may display the determined recommended effect information on the screen and present it to the user (recommend), or it may automatically set the effect information to be used when shooting the video as the recommended effect information.
[0073] The fourth additional process is the process by which the processor 30A determines the print conditions for the next time the printer 20 prints an image (hereinafter referred to as the "next print conditions") after the second image (more precisely, the third image) selected by the user has been printed by the printer 20. Specifically, after the printer 20 prints the second image (third image), the processor 30A displays one or more questions on the screen regarding the image quality of the printed image Pa, which is the image output by the printer 20, as shown in Figure 7. The user answers the displayed questions while looking at the printed image Pa (i.e., the printed second image) in the printed material P ejected from the printer 20, and inputs the answers through the input device 30E of the user terminal 30.
[0074] When the processor 30A receives the user's answer to a question, it determines the print conditions for subsequent prints based on that answer. For example, if the user answers "dark" to the question "How is the color of the printed image?", the saturation may be lower than it actually is due to the performance of the user terminal 30's display (LCD). Taking this into account, the print conditions for subsequent prints are modified from the most recently used print conditions, specifically by correcting the gradation values (RGB values) for printer control to increase the saturation.
[0075] The determined print conditions for subsequent prints may be displayed on the user terminal 30's display and presented (recommended) to the user. Alternatively, when the printer 20 prints an image next, the processor 30A may automatically set or change the print conditions used for that print to the print conditions for subsequent prints and control the printer 20. The procedure for determining the print conditions for subsequent prints based on the user's response will be explained in detail in a later section.
[0076] As described above, the printing conditions for subsequent prints are determined based on the user's answers to questions regarding the print image quality. This allows the printer 20 to adjust the print quality of the next image to be printed, reflecting the user's answers. To illustrate this point with a specific example, there may be a difference between the color of the image to be printed when viewed on the display (LCD) of the user terminal 30 and the color of the same image when viewed on the printed product P. By determining the printing conditions for subsequent prints using the procedure described above and having the printer 20 print the image under those determined conditions, the difference can be reduced. In other words, the printing conditions can be set to take into account the characteristics of the LCD and make the color of the printed image closer to the color of the same image when displayed on the LCD.
[0077] Furthermore, the images to be printed under the determined print conditions for subsequent printings are not particularly limited; they may be the same image as the second image (strictly speaking, the third image) that has already been printed, or they may be different images from the second image that has already been printed.
[0078] (Details regarding image evaluation and extraction, and determination of print conditions) Next, we will explain in more detail the functions of the user terminal 30, specifically the function for evaluating the first image, the function for extracting the second image, and the function for determining print conditions.
[0079] [Regarding the evaluation of the first image] The evaluation of the first image is performed after the user terminal 30 receives (acquires) the synthesized video from the video recorder 10. For example, if the user launches an image processing application on the user terminal 30 and performs an operation for image evaluation, this will be used as a trigger for the evaluation.
[0080] In the evaluation of the first image, the processor 30A of the user terminal 30 analyzes the acquired composite video, uses the first image included in the composite video as the evaluation target, and calculates the evaluation value of the first image. In the present embodiment, as shown in FIG. 8, the processor 30A calculates the evaluation value for each of two or more first images included in the composite video. Here, the first image for which the evaluation value is calculated, that is, the first image to be evaluated, may be all the frame images included in the composite video. Alternatively, frame images in the composite video may be selected at regular time intervals from the start point of the composite video, and each selected frame image may be used as the first image to be evaluated.
[0081] An example of the method for calculating the evaluation value will be described. According to each of a plurality of evaluation criteria set from a plurality of viewpoints, the score of the first image is determined for each evaluation criterion, the scores for each evaluation criterion are totaled, and the total value is used as the comprehensive evaluation value of the first image. When totaling the scores for each evaluation criterion, weighting may be performed for each evaluation criterion.
[0082] Specifically, as shown in Equation 1 below, the score for each evaluation criterion is multiplied by a weighting factor corresponding to the evaluation criterion, the product is calculated for each evaluation criterion, the products for each evaluation criterion are totaled, and the total value may be used as the evaluation value of the first image. S=α 1 ×S 1 +α 2 ×S 2 +・・・・α n ×S n (n is a natural number) (1) In Equation 1 above, S is the evaluation value of the first image, α k (k is a natural number from 1 to n) is the weighting factor set for each evaluation criterion, and S k is the score obtained according to each evaluation criterion.
[0083] When calculating the evaluation value S of the first image according to Equation 1 above, the total value of the weighting factors α k set for each evaluation criterion, that is, Σα k is 1, and each weighting factor α k is set to a default value at the initial time point, but may be changed as appropriate. Each weighting factor α kThe initial value of α may be a uniform value among the weight coefficients. Alternatively, each weight coefficient α k The initial value of may be determined empirically and set to different values among the weight coefficients.
[0084] Furthermore, when performing image evaluation, the processor 30A may display the inquiry screen shown in Figure 9 on the display and ask the user through that screen which of the multiple evaluation criteria to prioritize. In this case, the user should input the answers to one or more questions displayed on the inquiry screen, i.e., the evaluation criteria that the user prioritizes, through the input device 30E of the user terminal 30. The processor 30A may then set or change the weight coefficient αk of each evaluation criterion according to the user's answers to the inquiries. For example, the weight coefficient for the evaluation criterion that is given more importance may be set to a higher value than the weight coefficients for the other evaluation criteria. This makes it possible to evaluate the first image while reflecting the user's preferences.
[0085] Furthermore, the evaluation value calculated for the first image may be stored in the user terminal 30, video recorder 10, or server 50, linked to that first image. This allows for the creation of a database of evaluation values, which can then be used for subsequent evaluations.
[0086] Next, specific examples of evaluation criteria will be explained. The first evaluation criterion concerns items related to the subject in the first image, which is the subject of evaluation. Specifically, known subject detection techniques may be applied to the first image to detect subjects within the first image, and scoring (pointing) may be performed according to the detected subjects. For example, if a person is detected in the first image, the image of that person may be analyzed and points may be assigned according to the number of people detected, their faces (specifically, whether they match the faces of specific people registered in advance), the facial expressions and skin color of the people, the size and orientation of the faces in the image, their gaze, and the content of the actions the people are performing.
[0087] Furthermore, in the first evaluation criterion, points may be added if the person in the first image is performing a predetermined action (a pre-registered action, such as a popular action at the time of shooting). This makes it possible to evaluate images in a way that is relevant to the times, for example, by reflecting current trends in the evaluation of the first image.
[0088] The second evaluation criterion concerns the image quality of the first image being evaluated. Specifically, known methods for calculating image features may be used to calculate image features, and scoring may be performed based on the calculated features. Examples of image quality features include image brightness (grayscale value), hue, saturation, resolution, presence or absence of noise, contrast, focus, and the degree of camera shake and blur.
[0089] The third evaluation criterion concerns subjects other than people in the first image being evaluated (hereinafter referred to as "objects"). Specifically, known subject detection techniques may be applied to the first image to detect objects within the first image, and scoring may be performed according to the detected objects. Objects may include places and buildings, and may also include landscapes (scenes) such as night views, the sea, beaches, and the sky. Furthermore, points may be added if, for example, a famous landmark or popular spot is depicted in the first image. This makes it possible to evaluate images in line with the times, for example, to evaluate the first image in a way that reflects current trends. When identifying what a building or landscape is as an object, it may be done by analyzing information from a website (specifically, a website related to the region to which the first image was taken).
[0090] Furthermore, in the third evaluation criterion, if an object characteristic of an event, such as a sports day or a wedding, is detected in the first image, points may be added. Specifically, the shooting date and location are identified from the metadata of the video containing the first image to be evaluated, and the event the user is participating in is estimated from this information. At this time, the event the user is participating in may be identified by referring to information on the website or the user's schedule information registered in the user terminal 30 (information on the events the user is participating in and the date and time of participation). To explain with a specific case as an example, suppose it is determined from the shooting date and location that the user is near the entrance of a certain theme park in the morning, and it is estimated that the user is participating in an event at that theme park. In such a case, if a character from the theme park is visible in the first image to be evaluated, points may be added. This makes it possible to appropriately evaluate the first image according to the evaluation criteria corresponding to the user's actions, specifically the event the user is participating in. When the event the user is participating in is estimated, a weighting coefficient α for each evaluation criterion is applied according to the content of the estimated event. k You may adjust it.
[0091] Furthermore, in the third evaluation criterion, if the first image contains objects that represent the user's tastes or preferences, such as food items or animals, points may be added.
[0092] The fourth evaluation criterion concerns the relationship between the first image being evaluated and the composite video. Specifically, the evaluation may involve identifying which of the multiple unit videos constituting the composite video the first image belongs to, and scoring based on that identification. More specifically, the score may be calculated based on the position in the timeline when the unit video containing the first image (hereinafter referred to as the target video) is played during composite video playback (in other words, the position of the video in the composite video in the time series). Alternatively, the evaluation may involve identifying the time point in the target video during composite video playback when the first image exists (in other words, how much time has passed since the start of the target video), and scoring based on that identification.
[0093] The fifth evaluation criterion concerns the audio output at the display timing of the first image in the unit video containing the first image being evaluated. Specifically, known audio analysis techniques may be applied to the audio to identify its content, volume, and the strength, intonation, and pitch of the voice, and points may be assigned according to the identified information about the audio. For example, if the audio analysis results reveal that the first image being evaluated represents an exciting scene, points may be added.
[0094] The sixth evaluation criterion concerns the user's actions when shooting a unit video including the first image to be evaluated, particularly the operation of the video recorder 10. Specifically, if the video recorder 10 is equipped with an acceleration sensor and a gyro sensor, the orientation and posture of the video recorder 10 during video recording may be determined based on the output signals from these sensors, and points may be assigned according to the determined orientation and posture of the video recorder 10. For example, during video recording, while the user is shooting their favorite scene, the user may press down a specific button (specifically, a button for high ratings) provided on the video recorder 10. In such a configuration, points may be added for the first image captured while the specific button is being pressed. Furthermore, suppose that during video recording, the user operates the operation panel (not shown) of the video recorder 10 to make detailed settings regarding shooting conditions, etc. Making detailed settings regarding shooting conditions etc. indicates that it is a scene of high interest to the user, and points may be added for images captured under those detailed settings.
[0095] Furthermore, the number and types of evaluation criteria used when evaluating the first image are not particularly limited, and evaluation criteria other than the six criteria mentioned above may be added. Alternatively, the first image may be evaluated using only some of the six evaluation criteria mentioned above.
[0096] Furthermore, while the evaluation value is typically calculated based on the aforementioned evaluation criteria for the first image being evaluated, it may also be calculated indirectly based on those criteria in some cases. For example, for each of several unit videos, the first image (frame image) may be evaluated, and if the evaluation value is higher than the standard value, a predetermined evaluation value may be assigned to a certain number of first videos among the constituent videos. For example, if the evaluation value of the first image of a certain unit video is higher than the standard value, that unit video may be divided into numerical values (a fixed value) corresponding to the playback time, and a predetermined evaluation value may be assigned to the first image (frame image) at the beginning of each divided video without performing any analysis of what is depicted in those first images. By repeating this process for all unit videos, an indirect evaluation value can be calculated for the entire composite video. It should be noted that this indirect assignment of evaluation values may be combined with the calculation of evaluation values based on the typical evaluation criteria described above.
[0097] [Regarding the extraction of the second image] The extraction of the second image is performed after the evaluation of the first image using the procedure described above. The second image is extracted from the composite video, and more specifically, it is extracted from two or more first images for which an evaluation value has been calculated.
[0098] Furthermore, in this embodiment, the processor 30A extracts a first image as a second image whose evaluation value satisfies the extraction conditions. The extraction conditions are not particularly limited, but for example, the second image may be extracted based on any of the following extraction conditions j1 to j3. Extraction condition j1: The evaluation value is equal to or greater than a threshold. Extraction condition j2: Among the first images for which evaluation values have been calculated, the second image is among the top X (where X is a natural number) with the highest evaluation value. Extraction condition j3: When the extraction ratio to the number g of first images for which evaluation values have been calculated is Y%, the second image is among the top g × Y / 100 with the highest evaluation value.
[0099] The threshold value in extraction condition j1 may be a constant value or a variable value that changes depending on the situation. Furthermore, a combination of the three extraction conditions j1 to j3 described above may be used as the extraction condition.
[0100] An example of a specific procedure for extracting a second image (hereinafter referred to as the image extraction flow) will be described with reference to Figure 10. The image extraction flow according to this embodiment employs the image processing method of the present invention. Each step in the image extraction flow corresponds to each process included in the image processing method of the present invention and is performed by the processor 30A. Note that the image extraction flow illustrated in Figure 10 is merely an example, and within the scope of the present invention, some steps in the flow may be deleted, new steps may be added to the flow, or the execution order of two steps in the flow may be changed.
[0101] Before the image extraction flow starts, the processor 30A acquires a composite video and analyzes each of the two or more first images included in the composite video (for example, all the first images included in the composite video) in the manner described above to calculate an evaluation value for each first image. After the evaluation value for each first image has been calculated, the image extraction flow starts. In the image extraction flow, first, the processor 30A acquires a number g of the first images for which an evaluation value has been calculated (S001).
[0102] Next, the processor 30A sets the processing variable i to 1 (S002), refers to the evaluation value of the i-th first image among the two or more first images for which evaluation values have been calculated (S003), and determines whether the i-th first image is an image that satisfies the extraction conditions (S004). In step S004, for example, it is determined whether the evaluation value of the i-th first image is equal to or greater than the threshold (i.e., whether the above-mentioned extraction condition j1 is satisfied). Here, the i-th first image is the first image that is displayed i-th when the two or more first images for which evaluation values have been calculated are arranged on the time axis during playback of the composite video.
[0103] If the evaluation value of the first image is determined to be above the threshold, it is extracted as a provisional second image (S005). Subsequently, the processor 30A determines whether the number of extracted second images is equal to the set value X, which is the required number of second images to be extracted (S006). If the number of extracted second images is less than X, the processor 30A increments the processing variable i by +1 (S007) and repeats steps S003 to S006 described above until the number of extracted second images reaches X.
[0104] On the other hand, if the number of extracted second images is X, the processor 30A identifies the second image with the lowest evaluation value and updates the threshold used in step S004 to the evaluation value of the identified second image (S008). Subsequently, the processor 30A increments the processing variable i by +1 (S007) and repeats steps S003 to S008 described above until the processing variable i reaches the number g of first images acquired in step S001 (S009). During this process, step S005 is performed, and a new second image is provisionally extracted in step S005. In this case, the second image with the lowest evaluation value among the X second images (second image group) provisionally extracted up to that point is removed from the second image group, and the newly extracted second image is added to the second image group instead. In addition, the evaluation value of the second image with the lowest evaluation value among the X second images provisionally extracted is adopted as the new threshold, thereby updating the threshold.
[0105] Then, when the processing variable i reaches the number g of the first images acquired in step S001, the image extraction flow ends at that point. By performing the image extraction flow in the above procedure, X second images are extracted from the two or more first images for which an evaluation value has been calculated. Here, X may be a natural number greater than or equal to 2. In that case, the processor 30A draws the selection screen shown in Figure 11 and displays the X second images (second image group) extracted by the image extraction flow on the selection screen and presents them to the user.
[0106] The user selects one or more second images from the group of second images displayed on the selection screen. For the sake of clarity, the following explanation will assume that the user selects only one image from the group of second images displayed on the selection screen.
[0107] The second image selected by the user is set as the image to be printed by the printer 20, i.e., the third image. Subsequently, the processor 30A controls the printer 20 to print the third image, and the printer 20 starts printing the third image under the control of the processor 30A. When the printer 20 notifies the user terminal 30 that it has started printing the third image, this triggers the processor 30A to send (upload) the composite video to the server 50 and store it in the server 50's first memory area. Furthermore, the processor 30A sends (uploads) the printed third image to the server 50 and stores it in a predetermined memory area on the server 50.
[0108] As explained above, the processor 30A analyzes each first image included in the composite video and calculates an evaluation value. Based on the evaluation value of each first image, it extracts a second image from the composite video that is a candidate for a still image to be printed. This allows the user to appropriately and easily select a still image to be printed from the composite video.
[0109] [Determination of Printer Conditions] As described above, after printing by the printer 20, the processor 30A can determine the printing conditions for subsequent prints based on the user's answers to questions about the image quality of the printed image. More specifically, after printing the second image (more precisely, the third image), the processor 30A displays questions on the question screen shown in Figure 7 regarding the image output by the printer 20, or more specifically, the image quality of the printed image Pa in the printed item P. The user answers the above questions by operating the user terminal 30, and the processor 30A obtains the user's answers to the above questions through communication with the user terminal 30.
[0110] The acquired response information should be stored in the user terminal 30, video recorder 10, or server 50. In this case, the response information should be stored linked to the identification information of the user terminal 30 (specifically, the serial number or product name, etc.). This is because the image to be printed is displayed on the user terminal 30's display before printing, and the user answers the above question by comparing the image quality of the image displayed on the display with the image quality of the printed image.
[0111] Then, the processor 30A determines the print conditions for subsequent prints in accordance with the user's answers to the above questions. In this embodiment, the processor 30A derives derived information from video generation information related to the generation of a composite video including the printed second image (third image), the feature quantities of the printed second image, and the acquired user answers. Then, the processor 30A determines the print conditions for subsequent prints based on this derived information.
[0112] Specifically, the shooting conditions adopted during video recording (e.g., sensitivity and shutter speed), and the selected effect information chosen when generating the composite video, correspond to the aforementioned video generation information. The features of the second image include, for example, composition, color, contrast, degree of camera shake and blur, resolution, and the presence or absence of noise. In this embodiment, the processor 30A performs principal component analysis using the video generation information, the features of the second image, and the user's response to identify principal components that have a significant influence on the user's response.
[0113] Furthermore, based on the results of the principal component analysis shown in Figure 12, the processor 30A identifies content (correlation factors) that correlate with the identified principal components from the video generation information and the features of the second image. Here, the identified principal components and correlation factors correspond to the derived information mentioned above, and these reflect the user's preferences regarding the quality of the printed image. Figure 12 shows an example of the results of the principal component analysis. By referring to the results of the principal component analysis, it is possible to predict the user's preferences regarding the quality of the printed image. For example, if the correlation factors are effect information, composition, and the color tone of the printed image, it can be predicted that the user is particular about "skin tone". Note that since the prediction of user preferences is performed by a person, it may not be necessary for the image processing system 100 to perform this.
[0114] The processor 30A then determines the print conditions for subsequent prints based on the principal components and correlation factors. For example, if the correlation factor is "color tone," the processor 30A may determine the print conditions for subsequent prints to be lower in saturation than usual if the image to be printed has high saturation. In this embodiment, the user's requests and preferences regarding image printing (i.e., the tendency of the print conditions to be adopted) are fed back, and the system can automatically set print conditions that the user prefers.
[0115] The method for determining print conditions according to the correlation factors is not particularly limited, but for example, print conditions according to the correlation factors may be determined by referring to a Look Up Table (LUT) that defines the correspondence between the correlation factors and the print conditions.
[0116] Furthermore, the derived information described above is not limited to the principal components and correlation factors identified by principal component analysis. For example, machine learning may be performed using video generation information when a composite video including the second image is generated, print conditions when the second image is printed, and answers to questions regarding the image quality of the second image as training data, and the mathematical model obtained as a result of this learning (hereinafter referred to as the machine learning model) may be used as the derived information. In this case, by inputting the video generation information and user answers for the second image that has already been printed into the machine learning model, the machine learning model can output the print conditions for the next time and beyond. As a result, the processor 30A can determine the print conditions for the next time and beyond based on the machine learning model as the derived information.
[0117] <<Other Embodiments>> Although specific embodiments of the present invention have been described above, the embodiments described above are merely examples given to facilitate understanding of the present invention and do not limit it. That is, the present invention can be modified or improved from the embodiments described below, without departing from its spirit. Furthermore, the present invention includes equivalents thereof. Moreover, embodiments of the present invention may include forms that combine the embodiments described above with one or more of the following modifications.
[0118] (Modifications of the device for image evaluation, extraction, and printer control) In the above embodiment, the processor 30A of the user terminal 30 performs image evaluation and extraction, and printer control. However, it is not limited to this, and for example, the video recorder 10 (more specifically, the processor installed in the video recorder 10) may perform at least one of the image evaluation and extraction, and printer control. Alternatively, the image processing system 100 may have a cloud server (not shown), and this cloud server may perform image evaluation and extraction, and printer control on behalf of the user terminal 30, based on information obtained through communication with the user terminal 30. In this case, the user terminal 30 only needs to receive information from the cloud server indicating the results of the processing performed by the cloud server and display the received information. The above cloud server may be an ASP (Application Service Provider), SaaS (Software as a Service), PaaS (Platform as a Service), or IaaS (Infrastructure as a Service) server computer.
[0119] (Modified version of the device for classifying composite videos and generating integrated videos) In the above embodiment, the video recorder 10 performs the process of classifying composite videos, the process of generating integrated videos, and the process of analyzing the history of effect information to determine recommended effect information. However, it is not limited to this, and at least one of these processes may be performed by the processor 30A of the user terminal 30. Also, in the case where the image processing system 100 has a cloud server (not shown), at least one of the above processes may be performed by the cloud server.
[0120] (Modifications regarding the transmission of the synthesized video) In the above embodiment, when the printer 20 prints the second image (more specifically, the third image which is the second image selected by the user), and the printer 20 notifies the user terminal 30 that printing has started, this triggers the processor 30A of the user terminal 30 to send (upload) the synthesized video to the server 50. However, the invention is not limited to this, and when the user terminal 30 receives the synthesized video from the video recorder 10, the processor 30A may immediately send (upload) the synthesized video to the server 50.
[0121] (Modification regarding the evaluation of the first image) In the above embodiment, after acquiring the composite video, the acquired composite video is analyzed to calculate an evaluation value for the first image included in the composite video. However, the method is not limited to this, and before the composite video is generated, specifically, after each of the multiple unit videos constituting the composite video has been filmed, and before the multiple unit videos are combined to generate a single composite video, each of the multiple unit videos may be analyzed. Then, for each of the multiple unit videos, an evaluation value may be calculated for the frame image (first image) included in each unit video.
[0122] (Modification regarding the number of extracted second images) In the above embodiment, two or more second images were extracted from two or more first images for which an evaluation value was calculated. In other words, in the above embodiment, two or more second images were extracted, but the embodiment is not limited to this. Only one second image may be extracted from two or more first images for which an evaluation value was calculated. In this case, the extracted single second image may be displayed on the display of the user terminal 30 and presented to the user, and the user may be asked whether or not to print that second image as a third image.
[0123] (First Modified Example Regarding the Method for Extracting the Second Image) In the above embodiment, the second image is extracted from two or more first images for which an evaluation value has been calculated, but the method is not limited to this. For example, for each of the multiple unit videos constituting the composite video, an evaluation value may be calculated for the first first image (frame image) in each unit video, and if the evaluation value is higher than the reference value, a predetermined number d (d is a natural number of 2 or more) of first images may be extracted from each unit video as the second image. Alternatively, multiple videos may be divided into the above predetermined number d, and the first first image (frame image) in each of the divided d videos may be extracted as the second image.
[0124] (Second Modification Regarding the Method for Extracting the Second Image) In the above embodiment, from the first images for which an evaluation value has been calculated, the first image whose evaluation value satisfies the extraction conditions is extracted as the second image. For example, the first image that is the Xth highest in evaluation value is extracted as the second image. However, with such an extraction method, in one scene of the composite video, there may be a concentration of first images with high evaluation values, and multiple second images may be extracted from the same scene in the composite video. In other words, there may be cases in which the second image is extracted from only one of the multiple unit videos that make up the composite video.
[0125] To avoid the above situation, as shown in Figure 13, a second image may be extracted from each of two or more unit videos among the multiple unit videos that make up the composite video. More specifically, the processor 30A may analyze two or more unit videos among the multiple unit videos and calculate an evaluation value for each of the two or more unit videos containing the first image. Then, for each of the two or more unit videos, the first image that satisfies the extraction conditions can be extracted as the second image from the first images for which an evaluation value has been calculated. This makes it possible to avoid the situation of extracting a second image only from the same scene.
[0126] Furthermore, when extracting a second image from each of two or more unit videos, for example, in each unit video, the first image with the highest evaluation value, or the top X first images counting from the highest evaluation value, may be extracted as the second image. Alternatively, the extraction ratio of the number of first images g included in each video may be set to Y%, and the first images that fall within g × Y / 100, counting from the highest evaluation value, may be extracted as the second image. In addition, if there are two or more first images in each unit video that are similar in composition and subject matter, the evaluation value of one of these first images (for example, the first image with the higher evaluation value) may be maintained while the evaluation values of the remaining first images may be reduced to adjust the result.
[0127] Other methods can be considered to avoid the situation of extracting a second image only from the same scene. For example, when extracting two or more second images from a composite video, the processor 30A of the user terminal 30 may choose not to extract as a second image any first image that is displayed within a predetermined time frame from the time the first image extracted as a second image is displayed, within the playback time of the composite video (i.e., the time axis during playback). More specifically, the evaluation value of the first image included in the composite video tends to change as the playback time of the composite video progresses. Taking this into account, let's assume that the first image with the maximum (peak) evaluation value is extracted as the second image. In this case, first images within a range of t seconds before and after the display time of that second image (t is a real number) will not be extracted as second images, even if their evaluation value is at its peak. This makes it possible to avoid the situation of extracting a second image only from the same scene.
[0128] (Regarding the second image stored on the server) In the above embodiment, of the two or more second images extracted from the composite video by the image extraction function of the user terminal 30, the second image printed by the printer 20, i.e., the third image, is stored on the server 50. However, the server 50 may also store a second image different from the third image, that is, a second image that was not selected as the image to be printed, along with the third image, or in place of the third image.
[0129] (Regarding the processor in the image processing system) The image processing system of the present invention comprises at least one processor, which may include various types of processors. These various types of processors include, for example, a CPU, which is a general-purpose processor that executes software (programs) and functions as various processing units. Furthermore, each process in each embodiment of the present invention may be executed on any computer. Furthermore, any computer may execute these processes by a processor, a program, or a combination thereof. Any computer may be a general-purpose computer, a computer for a specific purpose, a workstation or other system, or other hardware elements capable of executing programs.
[0130] A processor may consist of one or more hardware components, and the type of hardware is not limited. For example, a processor may consist of programmable logic devices such as a CPU (Central Processing Unit), MPU (Micro Processing Unit), FPGA (Field Programmable Gate Array), dedicated circuits for performing specific processing such as an ASIC (Application Specific Integrated Circuit), a GPU (Graphic Processing Unit), or an NPU (Neural Processing Unit).
[0131] Furthermore, the processor has units or means that perform various processes in each embodiment of the present invention. The type of hardware may also be a combination of different types of hardware. When multiple pieces of hardware are configured to perform one or more processes of a processor, these multiple pieces of hardware may be located in physically separate devices or in the same device. In addition, in any embodiment of the present invention, the order of the processes performed by the processor is not limited to the order described above and may be changed as appropriate. The hardware is composed of an electrical circuit (circuitry) or the like, which is a combination of circuit elements such as semiconductor elements.
[0132] Furthermore, each embodiment of the present invention may be implemented by hardware, software, firmware, microcode, or a combination thereof. The software, firmware, and microcode are composed of a program. The program may also be, for example, a group of program modules, and each function may be implemented by a processor configured to perform its respective function. The program may be program code and multiple code segments stored on one or more non-temporary computer-readable media (e.g., storage media or other storage). The program may be divided and stored on multiple non-temporary computer-readable media located in physically separate devices. The program code or code segments may represent any combination of procedures, functions, subprograms, routines, subroutines, modules, software packages, classes, or instructions, data structures, or program statements. The program code or code segments may be connected to other code segments or hardware circuits by sending and receiving information, data, arguments, parameters, or memory contents.
[0133] 10 Video recorder (video generation device) 20 Printer 30 User terminal 30A Processor 30B Memory 30C Communication interface 30D Storage 30E Input device 30F Output device 50 Server (source, storage device) 100 Image processing system P Printed material Pa Printed image Pb Encoded image
Claims
1. An image processing system comprising a processor and a printer, wherein the processor acquires a composite video composed of multiple videos, analyzes the composite video or the multiple videos to calculate an evaluation value for a first image which is a still image included in the composite video, extracts a second image from the first image for which the evaluation value has been calculated, the second image whose evaluation value satisfies the extraction conditions, and prints the second image to the printer.
2. The image processing system according to claim 1, wherein the processor transmits the synthesized video to the destination.
3. The image processing system according to claim 2, wherein the processor transmits the synthesized video to the destination when the second image is printed by the printer.
4. The image processing system according to claim 2, wherein the processor transmits a derived image generated based on the synthesized video to the destination.
5. The image processing system according to claim 2, wherein the processor causes the printer to print the second image on which information regarding the storage area of the synthesized video at the destination is superimposed.
6. The image processing system according to claim 1, wherein the processor calculates the evaluation value for two or more first images, extracts one or more second images from the two or more first images for which the evaluation value has been calculated, and prints a third image selected by the user from among the one or more second images to the printer.
7. The image processing system according to claim 6, comprising a storage device for storing images, wherein the processor causes the printer to print one of the images stored in the storage device.
8. The image processing system according to claim 7, wherein the third image is stored in the storage device.
9. The image processing system according to claim 7 or 8, wherein the storage device stores a second image different from the third image among two or more second images.
10. The image processing system according to claim 1, comprising a video generation device that shoots the multiple videos and generates the composite video based on effect information selected from among multiple effect information, wherein the video generation device classifies the composite video based on the effect information selected when generating the composite video.
11. The image processing system according to claim 10, wherein the printer and the video generation device are integrated.
12. The image processing system according to claim 10, wherein the video generation device generates a generated image based on two or more composite videos classified into the same classification category.
13. The image processing system according to claim 10, wherein the video generation device stores the effect information selected when generating the composite video in association with the time when the composite video was generated based on the effect information, and determines the effect information to be used when generating the composite video next from among the plurality of effect information based on the relationship between the effect information and the time.
14. The image processing system according to claim 1, wherein the processor displays a question on the screen regarding the image quality of the image output by the printer after the printer has printed the second image, and determines the print conditions for the next image to be printed by the printer in accordance with the user's answer to the question.
15. The image processing system according to claim 14, wherein the processor determines the print conditions based on video generation information relating to the generation of the composite video including the printed second image and derived information derived from the answer.
16. The image processing system according to claim 1, wherein, upon playback of the composite video, the plurality of videos are displayed in order, the processor extracts two or more second images from the composite video, and the processor does not extract as a second image a first image that is displayed within a predetermined time from the time the first image extracted as a second image is displayed during the playback time of the composite video.
17. The image processing system according to claim 1, wherein the processor analyzes two or more of the plurality of videos, calculates an evaluation value for the first image for each of the two or more videos, and extracts the second image from the first image for which the evaluation value has been calculated for each of the two or more videos.
18. An image processing method comprising: a step of a processor acquiring a composite video composed of multiple videos; a step of a processor analyzing the composite video or the multiple videos to calculate an evaluation value for a first image which is a still image included in the composite video; a step of a processor extracting a second image from the first image for which the evaluation value has been calculated, the second image whose evaluation value satisfies the extraction conditions; and a step of a processor printing the second image to a printer.
19. The image processing method according to claim 18, wherein the processor transmits the synthesized video to the destination.
20. The image processing method according to claim 18, wherein in the step of calculating the evaluation value, the evaluation value is calculated for two or more first images; in the step of extracting the second image, one or more second images are extracted from the two or more first images for which the evaluation value has been calculated; and in the step of printing the second image to the printer, a third image selected by the user from among the one or more second images is printed to the printer.
21. A program for causing a computer to perform each step included in the image processing method according to any one of claims 18 to 20.
22. A computer-readable recording medium on which a program is recorded causing a computer to perform each step included in the image processing method described in any one of claims 18 to 20.