Image generating device, image generating method, program, and recording medium

WO2026181749A1PCT designated stage Publication Date: 2026-09-03FUJIFILM CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/005226
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-28
Filing Date
2026-02-13
Publication Date
2026-09-03

Smart Images

  • Figure JP2026005226_03092026_PF_FP_ABST
    Figure JP2026005226_03092026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention makes it possible to generate a suitable interpolation image efficiently and without hassle. Provided is an image generating device 10 according to one embodiment of the present invention, wherein a processor: generates text information Tx, which is first feature information common to an input image group G1 constituted of two or more images; and generates an interpolation image G2 on the basis of the text information Tx, which is the first feature information, and second feature information, which is related to at least one individual image G1A from among the input image group G1 constituted of two or more images.
Need to check novelty before this filing date? Find Prior Art

Description

Image generation apparatus, image generation method, program, and recording medium

[0001] One embodiment of the present invention relates to an image generation apparatus, an image generation method, a program, and a recording medium that can efficiently generate a suitable interpolated image without labor.

[0002] With the widespread popularization of digital cameras and smartphones and the evolution of imaging functions, users can easily capture high-quality images in various situations. Accordingly, opportunities for such users to create various types of content such as electronic event albums from their accumulated images and enjoy looking back on memories with themselves and their families are also increasing.

[0003] However, the quality of content created in this way can vary depending on various factors, such as the number and variety of images used as creation materials, and the comprehensiveness of shooting opportunities, shooting subjects, and the like. On the other hand, from the user's perspective, it is difficult to reliably capture every opportunity and always perform complete shooting without omission.

[0004] Accordingly, conventional techniques have been proposed to compensate for missing images that do not exist in existing stored images as images relating to specific scenes, purposes, and the like. An example of such a technique is the technique described in Patent Document 1. Patent Document 1 discloses a technique for providing a moving image display apparatus and method capable of automatically and rapidly generating and displaying a series of connected still image sequences required for moving image display from a base image by providing parameters representing the movement of feature points.

[0005] This technology relates to a video display device characterized by comprising: an image storage means for storing multiple basic image data including feature points; a buffer means for selectively reading and temporarily storing basic image data necessary for generating and displaying a target video from the image storage means; a control data assignment means for providing a parameter sequence representing the movement of feature points during video display as control data; an image generation means for continuously generating new image data as interpolated images of basic images based on the basic image data read into the buffer means, according to the control data provided by the control data assignment means; and a display means for sequentially displaying the images represented by the image data generated by the image generation means on the screen.

[0006] Furthermore, Patent Document 2 discloses a technology that automatically creates a video summary by integrating important words from text data obtained by converting audio data into text using speech recognition, and representative images extracted based on the same concept as extracting important words from text data.

[0007] This technology relates to an automatic video summarization device that automatically creates a video summary from video data composed of multiple image data, for each scene composed of a block of the image data, and comprises: a scene extraction unit that divides the video data into scenes and extracts scene video data for each divided scene; a speech recognition unit that recognizes audio data contained in the scene video data and generates text data from the audio data; a key word extraction unit that extracts important words that are keywords from the text data; a representative image extraction unit that extracts representative image data that is representative of the scene from the scene video data; and a scene integration unit that integrates the key words and the representative image data for each scene to create video summary data.

[0008] Japanese Patent Publication No. 9-97348 Japanese Patent Publication No. 2008-148121

[0009] On the other hand, it is conceivable to use Artificial Intelligence (AI) to generate so-called interpolated images to compensate for the missing images mentioned above. However, when generating images with AI, it is difficult to obtain the desired image unless appropriate prompts are input into the user interface (for the AI). Moreover, for an average user who is not an expert in AI, creating prompts for generating interpolated images that do not contradict existing images and meet the user's desired conditions is extremely difficult and impractical.

[0010] One embodiment of the present invention has been made in view of the above circumstances, and aims to provide an image generation device, an image generation method, a program, and a recording medium that can generate suitable interpolated images efficiently and without effort.

[0011] The above objective is achieved by an image generation device described in any of the following [1] to [9]. [1] An image generation device equipped with a processor, wherein the processor generates first feature information common to two or more images, and generates an image based on second feature information relating to at least one of the two or more images and the first feature information.

[0012] [2] The image generation device described in [1], wherein the processor generates text information as first feature information by inputting two or more images to a trained model that outputs text information for an input image.

[0013] [3] The image generation device according to either [1] or [2], wherein the processor generates second feature information by inputting at least one image out of two or more images to a trained model that outputs text information for an input image.

[0014] [4] An image generation apparatus according to any one of [1] to [3], wherein the processor extracts second feature information based on the supplementary information of at least one of two or more images.

[0015] [5] An image generation device according to any one of [1] to [4], wherein the processor accepts editing by the user of the first feature information and generates an image based on the second feature information and the edited first feature information.

[0016] [6] The image generation apparatus according to any one of [1] to [5], wherein the processor determines the output form of the generated image based on at least one of the first feature information and the second feature information.

[0017] [7] An image generating device according to any one of [1] to [6], wherein the processor accepts user specifications for the output format of the generated image.

[0018] [8] The image generation apparatus according to any one of [1] to [7], wherein the processor performs the generation of first feature information and the generation of images based on a video containing two or more images.

[0019] [9] The image generation apparatus according to any one of [1] to [8], wherein the processor selects an image from two or more images according to a predetermined rule, and generates first feature information and an image based on the selected image.

[0020] Furthermore, the above objective can also be achieved by the image generation method described in

[10] below.

[10] An image generation method using a processor, comprising the steps of: generating first feature information common to two or more images using a processor; and generating an image using a processor from second feature information relating to at least one of the two or more images and the first feature information.

[0021] Furthermore, the above objectives can also be achieved by the program described in

[11] below.

[11] A program for causing a computer to perform each step included in the image generation method described in

[10] .

[0022] Furthermore, the above objectives can also be achieved by the recording medium described in

[12] below.

[12] A computer-readable recording medium on which a program is recorded causing a computer to perform each step of the image generation method described in

[10] .

[0023] According to one embodiment of the present invention, an image generation apparatus, an image generation method, a program, and a recording medium are provided that enable the efficient and effortless generation of suitable interpolated images.

[0024] This figure shows an example of a system configuration including the image generation device according to this embodiment. This figure shows an example of the hardware configuration of the image generation device according to this embodiment. This figure shows the functional part of the image generation device according to this embodiment. This figure shows an example of an input image group according to this embodiment. This figure shows an example of an input image table according to this embodiment. This figure shows an example of a generated text table according to this embodiment. This figure shows an example of a missing image judgment rule in this embodiment. This figure shows an example of an interpolated image table according to this embodiment. This figure shows an example of an interpolated image according to this embodiment. This figure shows an example of a generated content table according to this embodiment. This figure shows an example of an editing information table according to this embodiment. This figure shows an example of a generated form table according to this embodiment. This figure shows an example of a selection rule table according to this embodiment. This figure shows an example of the image generation method according to this embodiment. This figure shows an example of a screen according to this embodiment. This figure shows an example of the image generation method according to this embodiment. This figure shows an example of the generation of an interpolated image and content according to this embodiment.

[0025] The following describes specific embodiments of the present invention. For convenience of explanation, the following descriptions may sometimes be based on the perspective of a GUI (Graphical User Interface). Furthermore, since the fundamental data processing technologies for realizing the present invention (communication / transmission technologies, data acquisition technologies, data recording technologies, data processing / analysis technologies, machine learning technologies, image processing technologies, and visualization technologies, etc.) are known technologies, their descriptions will be omitted.

[0026] Furthermore, in this specification, the concept of "device" includes not only a single device that performs a specific function, but also a combination of multiple devices that exist independently and in a distributed manner while cooperating (linking) to perform a specific function.

[0027] Furthermore, in this invention, "user" refers to the user of the image generation device of the present invention, and specifically refers to a person who, for example, uses the functions of the image generation device of the present invention to confirm the text information (first feature information: an interpretation of the input image group expressed in natural language) generated by the generation AI and take appropriate action based on the results (e.g., editing the first feature information, such as correcting the interpretation result to match the user's shooting intention, specifying the shooting location desired by the user, and the facial expression and movement state of the subject).

[0028] Furthermore, in this specification, "person" means an entity that performs a specific action, and includes individuals, groups, corporations and other legal entities, and organizations, and may also include computers and devices that constitute artificial intelligence (AI). Artificial intelligence (AI) realizes intelligent functions such as reasoning, prediction, and judgment using hardware and software resources. The algorithm of artificial intelligence is arbitrary and includes, for example, expert systems, case-based reasoning (CBR), convolutional neural networks (CNN), deep neural networks (DNN), Bayesian networks, or inclusion architectures.

[0029] <<About one embodiment of the present invention>> [Configuration of the image generation system] In one embodiment of the present invention (hereinafter, this embodiment), an example is given in which an image generation system 5 is configured by an image generation device 10, a user terminal 30, a shooting device 40, and an external server 50 connected to a network 1 as shown in Figure 1. The user terminal 30 is an information processing device that uploads a group of images taken by the user as an input group of images to the image generation device 10 and edits the text information (first feature information) generated by the generation AI for the input group of images. If the user terminal 30 itself is configured to function as the image generation device 10, the above upload is unnecessary, and each process in the image generation device 10 can be performed by the user terminal 30 itself. The same applies to the shooting device 40; if the shooting device 40 itself is configured to function as the image generation device 10, the shooting device 40 may perform each process in the image generation device 10.

[0030] Furthermore, the user terminal 30 and the imaging device 40 capture the above-mentioned input image group and input images from the UI (User Interface), which serve as the basis for generating text information (first feature information) and interpolated image generation by the generating AI, and upload this to the image generation device 10.

[0031] Furthermore, the user may perform input operations on the input image group and editing operations on the text information using the UI (User Interface) of the image generation device 10, without using the user terminal 30 or the imaging device 40. In this case, the UI of the image generation device 10 will read the data of such input image group from the recording medium presented by the user. Alternatively, the user in this case may directly perform editing operations on the text information using the UI, such as a keyboard, mouse, or touch panel.

[0032] The image generation device 10 stores the input image group uploaded from the user terminal 30 and the imaging device 40 in the input image table 211 (described later in Figures 3 and 5) and prepares for the generation of text information (first feature information) by the generation AI 221. The image generation device 10 also stores the text information generated by the generation AI 221 from the input image group held in the input image table 211 in the generation text table 222 (described later in Figures 3 and 6) and prepares for editing by the user.

[0033] The user's edited text information is stored in the edited information table 212, preparing for the generation of interpolated images by the generation AI 221 based on the edited text information and the input image group. The interpolated images generated by the generation AI 221 are stored in the interpolated image table 223.

[0034] The image generation device 10 is an information processing device that uses a generation AI 221 (see Figure 3), which is either built into the device itself or provided by an external server 50, to generate and output text information interpreting the input image group, and to generate an interpolated image based on the edited text information and the input image group after user editing operations (which may include the concepts of modifying or adding to text, or giving instructions for regeneration; the same applies hereinafter). Such processing is performed autonomously by the image generation device 10 in accordance with the image generation method of the present invention, or in response to user instructions received through a predetermined UI.

[0035] Furthermore, as described above, the external server 50 is a server device that provides the function of generation AI 221 to the image generation device 10 via the network 1. Of course, if the image generation device 10 is configured in a way that does not require the provision of the function of generation AI 221, the external server 50 does not need to be included in the image generation system 5.

[0036] In addition to the configuration in which the image generation device 10, user terminal 30, and imaging device 40 are connected via network 1, the internal bus wiring of the image generation device 10 may be directly connected to the interfaces of the user terminal 30 and imaging device 40.

[0037] [Example Configuration of Image Generation Device] Next, an example configuration of the image generation device 10 according to this embodiment will be described with reference to Figures 2 and 3. Specifically, the image generation device 10 is composed of a PC (Personal Computer), a smartphone, a tablet terminal, a digital camera, a notebook PC, or a server device, etc. Note that the image generation device 10 is not limited to a computer owned by the user or accessible via network 1. For example, it may be composed of a terminal that is not owned by the user but is available when visiting a store or facility, such as a store-installed terminal. In the following, we will explain using the case in which the image generation device 10 is composed of a computer owned by the user, specifically a PC, as an example.

[0038] As shown in Figure 2, the computer comprising the image generation device 10 includes a processor 11, an auxiliary storage device 12, a main storage device 13, an input device 14, an output device 15, and a communication device 16.

[0039] The processor 11 is composed of, for example, a CPU (Central Processing Unit), an MPU (Micro-Processing Unit), an MCU (Micro Controller Unit), a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), a TPU (Tensor Processing Unit), or an ASIC (Application Specific Integrated Circuit).

[0040] The auxiliary storage device 12 is composed of, for example, flash memory, HDD (Hard Disk Drive), SSD (Solid State Drive), FD (Flexible Disk), MO disk (Magneto-Optical Disk), CD (Compact Disk), DVD (Digital Versatile Disk), SD card (Secure Digital card), or USB memory (Universal Serial Bus memory).

[0041] Such an auxiliary storage device 12 may be built into the computer main body constituting the image generation apparatus 10, or may be attached to the computer main body in an external format. Alternatively, the auxiliary storage device 12 may be configured by NAS (Network Attached Storage) or the like. Further, the auxiliary storage device 12 may be an external device that can communicate with one computer constituting the image generation apparatus 10 via a communication network, for example, an online storage or a database server.

[0042] Further, the auxiliary storage device 12 stores an operating system (OS) and a program 121 such as an application for executing generation processing of text information and interpolated images. When the program 121 is read and executed by the processor 11, the computer constituting the image generation apparatus 10 implements the functions of the receiving unit 21, the generating unit 22, and the output unit 23 shown in FIG. 3. Accordingly, this implementation executes a series of processes related to the generation and output of text information based on an input image group, the generation of an interpolated image based on the input image group and the text information, and the like.

[0043] The main storage device 13 is configured by a semiconductor memory such as ROM (Read Only Memory) and RAM (Random Access Memory), for example. The processor 11 loads the program 121 onto the main storage device 13 and executes it therein.

[0044] The input device 14 is a device that receives input operations by an administrator of the image generation apparatus 10 or the image generation system 5 or the user described above, and is configured by, for example, a keyboard, a mouse, a touch panel, a camera unit, or the like. Further, the input device 14 may include an imaging device implemented in a digital camera, a microphone for sound collection, and the like. The output device 15 is configured by, for example, a display, a speaker, or the like.

[0045] The communication device 16 may be constituted by, for example, a network interface card, a communication interface board, or the like. The computer constituting the image generation apparatus 10 can communicate with other devices connected to a network 1 such as the Internet and a mobile communication line via the communication device 16.

[0046] As shown in FIG. 3, the image generation apparatus 10 includes a reception unit 21, a generation unit 22, and an output unit 23. These functional units are implemented when the processor 11 of the image generation apparatus 10 executes the program 121 and cooperates with various types of hardware in the image generation apparatus 10 and other devices (the user terminal 30, the imaging apparatus 40, and the external server 50). Further, the generation AI 221 in the generation unit 22 may be called from the external server 50 and executed.

[0047] [Regarding Input Image Group] In the present embodiment, the image generation apparatus 10 according to the present invention acquires an input image group G1 as illustrated in FIG. 4 from the user terminal 30 or the imaging apparatus 40, and generates text information Tx by providing the input image group G1 to the generation AI 221.

[0048] The input image group G1 illustrated in FIG. 4 is composed of individual images G1A to G1F (of course, the number of such individual images is just an example). Note that when a common description is given for such individual images G1A to G1F, the description will be given with the image G1A as a representative. The individual image G1A is, for example, an image in which a landscape, a building, or the like at a shooting location serves as a background image Pb, and includes one or more subject images Ps reflected in front of the background image Pb.

[0049] The individual image G1A illustrated in FIG. 4 is an image of a single woman smiling in front of the Eiffel Tower in Paris. In this case, the image G1A can be said to be an image having a composition in which the background image Pb is a cityscape of Paris including the Eiffel Tower, and the subject image Ps is a single woman. This composition is the same for the other individual images G1B to G1E. Further, the other individual image G1F is an image having a composition in which the background image Pb is the European Alps, and the subject image Ps includes a total of four women including the aforementioned woman.

[0050] The image generation device 10 inputs the input image group G1 acquired from the user terminal 30 and the shooting device 40 into the generation AI 221, thereby acquiring text information Tx that represents a common concept (interpretation result) for the individual images G1A to G1F that make up the input image group G1. In the case of the input image group G1 in Figure 4, the generation AI 221 of the image generation device 10 generates and outputs keywords such as "a woman, various expressions, a trip to France."

[0051] The user views this text information Tx "A woman, various expressions, a trip to France" on the user terminal 30, and judges whether it meets various conditions, such as the purpose and intention of taking the input image group G1, or expressions that suggest content (e.g., an electronic album) using the input image group G1 as material, and edits the text as necessary, such as correcting, adding, or deleting words.

[0052] For example, a user who wants to create content that focuses on movement rather than facial expressions might edit the text information Tx in Figure 4, changing "various facial expressions" to "various movements." Similarly, a user who wants to create content that focuses on many people rather than just one might edit the text information Tx in Figure 4, changing "one woman" to "various people." Furthermore, a user who wants to create content that focuses on images in various situations without narrowing down the shooting location might edit the text information Tx in Figure 4, deleting "trip to France."

[0053] In this invention, "image" is composed of multiple pixels, represented by the grayscale values ​​of each of the multiple pixels, and includes at least one subject image Ps. Furthermore, digital image data (hereinafter referred to as "image data") that defines an image at a set resolution is generated by compressing data in which the grayscale values ​​of each pixel are recorded using a predetermined compression method. Examples of image data include lossy compressed image data such as JPEG (Joint Photographic Experts Group) format, and lossless compressed image data such as GIF (Graphics Interchange Format) or PNG (Portable Network Graphics) format.

[0054] [Functional Units in the Image Generation Device] Next, we will explain the processing performed by each of the functional units, namely the receiving unit 21, the generation unit 22, and the output unit 23, that are provided in the processor 11 of the image generation device 10.

[0055] (Reception Unit) The reception unit 21's processing includes receiving the input image group G1 from the user terminal 30 or the imaging device 40 via the input device 14 or the communication device 16, and storing it in, for example, the auxiliary storage device 12. Furthermore, the storage process in the reception unit 21 includes storing the input image group G1 received as described above in the input image table 211.

[0056] Figure 5 shows an example of the configuration of the input image table 211 in this embodiment. The input image table 211 is a collection of records containing the values ​​of "User ID," "Date and Time," and "Image Data." The User ID is the identification information of the user who input the input image group G1, and may also be the identification information of the user terminal 30. Furthermore, if the image generation device 10 is operated by a single user, this User ID is not necessary. The Date and Time is the date and time when the input image group G1 was input.

[0057] The method for receiving the input image group G1 at the reception unit 21 is not particularly limited, but includes methods of acquiring images by reading photographic prints or playback videos taken by the user terminal 30 or the shooting device 40 using the camera unit. The reception unit 21 may also acquire images by downloading image data from an external device or a web server via the network 1. The input image group G1 may be a still image or a video.

[0058] (Generation Unit) The processing of the generation unit 22 involves inputting the input image group G1 obtained by the processing of the reception unit 21 into the generation AI 221, and includes the process of generating, for example, keywords or sentences of events common to the individual images that make up the input image group G1, i.e., text information Tx (see Figure 4). The generation AI 221 that generates text information Tx has the function of a caption generation AI that outputs keywords that appropriately represent the input image for the input image. Therefore, the generation AI 221 has a model that has been machine-learned using sets of various input images and keywords that appropriately represent the input image as training data.

[0059] The text information Tx generated by the generation unit 22 using the generation AI 221 is stored in the generation text table 222. As shown in Figure 6, this generation text table 222 is a collection of records consisting of "User ID," "Date and Time," "Input Image Group," and "Text Information." The User ID is the ID of the photographer (provider) of the input image group G1. The Date and Time is the date and time the input image group G1 was taken or the date and time the text information Tx was generated. The input image group is the file name of the individual image G1A that make up the input image group G1 used to generate the text information Tx. The text information is the text information Tx generated based on the above input image group G1. The input image group G1 may also include images generated by the generation unit 22 by inputting the text information Tx and other content into the generation AI 221.

[0060] Furthermore, the processing of the generation unit 22 also includes the process of inputting the above-mentioned text information Tx (which may include information edited by the user) and information (second feature information) of individual images G1A (at least one of which) that constitute the above-mentioned input image group G1 to the generation AI 221, thereby generating interpolated images G2 for any images missing in the input image group G1.

[0061] In the input image group G1, missing images are those that the generation unit 22 determines to be missing based on predetermined rules or algorithms, or based on user specifications, regarding the event indicated by the text information Tx. For example, if the text information Tx is "various facial expressions," the generation AI 221 analyzes the facial expressions of each subject in the individual images G1A in the input image group G1.

[0062] Based on the results of this analysis, the generation unit 22 classifies the individual images G1A according to the facial expressions of the subjects. For example, suppose it obtains a classification result (second feature information) in which 50% of the individual images G1A in the input image group G1 are smiling, 10% are troubled, 10% are angry, and 10% are expressionless. In this case, the generation unit 22 determines that there are not enough images of laughing faces in the input image group G1, based on, for example, essential facial expression patterns defined in advance for the event "various facial expressions" (e.g., laughing face, smiling face, angry face, troubled face, and expressionless face).

[0063] Furthermore, for example, if the text information Tx is "various movements", the generating AI 221 analyzes the movement of each individual subject in the input image group G1A.

[0064] Based on the results of this analysis, the generation unit 22 classifies the individual images G1A according to the movement of the subject. For example, suppose it obtains a classification result (second feature information) in which 40% of the individual images G1A in the input image group G1 are images of walking, 20% are images of standing still, 10% are images of sitting, and 10% are images of lying down. In this case, the generation unit 22 determines, for example, that images of running and jumping are lacking in the input image group G1, based on essential motion patterns defined in advance for the event "various movements" (e.g., walking, standing, sitting, running, jumping, and lying down).

[0065] Furthermore, if, for example, the text information Tx is "various shooting locations" or "various people," the generating AI 221 analyzes the shooting locations and the people who are subjects of each individual image G1A in the input image group G1.

[0066] Based on the results of this analysis, the generation unit 22 classifies individual images G1A by the location where the subject was photographed and by the person in the subject. For example, it may obtain a classification result that, among the individual images G1A in the input image group G1, 70% are images where area A is the shooting location and 30% are images where area B is the shooting location, or that 60% are images where person A is the subject, 30% are images where person B is the subject and 10% are images where person C is the subject.

[0067] In this case, the generation unit 22 determines, for example, that the input image group G1 is lacking images taken in an area different from areas A and B (e.g., a well-known spot that belongs to the same country or region as the shooting location of each individual image G1A that makes up the input image group G1, but is different from the shooting location of each individual image G1A), based on a predetermined essential pattern for the event "various shooting locations" (e.g., images of at least four people, each as the main subject).

[0068] Furthermore, the method for determining missing images is not limited to the method of determining that there are missing images for a requirement defined by the essential patterns (e.g., laughing face, smiling face, angry face, troubled face, and expressionless face) for which there is no corresponding image (e.g., laughing face). For example, a method can also be adopted in which images are determined to be missing if the proportion or number of individual images G1A classified for each event, such as facial expression, movement, shooting location, and subject person, is less than the standard.

[0069] In this case, for example, based on the classification results (second feature information) of the individual images G1A in the input image group G1, such as four images of smiling faces, four images of troubled faces, and two images of laughing faces, the generation unit 22 determines that at least one image of a laughing face is missing, based on the rules for determining missing images that have been predetermined for the event "various facial expressions" (e.g., the difference in the number of individual images G1A for each facial expression is two or less).

[0070] The essential patterns and judgment rules for each of the above events are exemplified in Figure 7 as the missing image judgment rule 225. As shown in the figure, this missing image judgment rule 225 is a table that defines the missing image judgment rule for each event.

[0071] Having determined the missing images as described above, the generation unit 22 prompts the generation AI 221 to generate an image of the subject with a "laughing face," or an image of the subject "running" or "jumping," and generates an interpolated image. Upon receiving the above prompt input, the generation AI 221 either changes the facial expression of the subject in the individual image G1A to "laughing" and generates an interpolated image, or changes the movement to a "running" or "jumping" state based on the images of each body part of the subject in the individual image G1A and the skeletal estimation results, and generates an interpolated image. The generation AI 221 is a trained model that automatically generates images according to the prompts based on the above prompts and reference images that serve as the basis for the generation process (e.g., individual image G1A).

[0072] In the processing of the generation unit 22, the generated interpolated image is stored in the interpolated image table 223. As shown in Figure 8, this interpolated image table 223 is a collection of records consisting of "User ID", "Date and Time", "Text Information" used to generate the interpolated image, and "Generated Interpolated Image". The User ID is the ID of the user who took and provided the input image group G1. The Date and Time is the date and time the input image group G1 was taken or the date and time the interpolated image was generated. The Text Information is the value of the text information Tx generated by the generation AI 221 based on the input image group G1. The generated interpolated image is a data file of the interpolated image.

[0073] Figure 9 shows an example of interpolated image G2. This interpolated image G2 is an interpolated image generated by the generating AI 221 based on the text information Tx "A woman, various movements, traveling in France" from the input image group G1 shown in Figure 4, and the information of the individual images G1A to G1F in the input image group G1 (second feature information). Of these, the interpolated image "Move01.JPG" is an image of "running" that was not included in the input image group G1. Also, the interpolated image "Move02.JPG" is an image of "jumping" that was not included in the input image group G1. In the above case, information indicating that it was "not included in the input image group G1" (or was included but in less than the standard) may correspond to the second feature information.

[0074] As already mentioned, the generation unit 22 can display the text information Tx on a user terminal 30 or the like, and accept confirmation and modification by the user. For example, the user can view the text information Tx "A woman, various expressions, a trip to France" on the user terminal 30, judge whether it is suitable for various conditions such as the purpose and intention of taking the input image group G1, the situation, or expressions that suggest content using the input image group G1 as material (e.g., an electronic album, etc.) (e.g., a collection of children's smiles), and edit the text by modifying, adding, or deleting as needed.

[0075] For example, a user who wants to create content that focuses on movement rather than facial expressions might change the phrase "various expressions" to "various movements" in the text information Tx in Figure 4. Similarly, a user who wants to create content that focuses on many people rather than just one person might change the phrase "one woman" to "various people" in the text information Tx in Figure 4. Furthermore, a user who wants to create content that focuses on the fact that the photos were taken in various situations rather than narrowing down the shooting location might delete the phrase "trip to France" in the text information Tx in Figure 4.

[0076] The generation unit 22 receives the edited text information Tx as described above and stores the edited text information Tx, which is the result of the editing, in the edited text information table 212. Figure 11 shows an example of the configuration of the edited text information table 212 in this embodiment. The edited text information table 212 is a collection of records that include the values ​​of "User ID", "Date and Time", "Text Information", and "Edited Text Information".

[0077] Of these, the User ID is the identification information of the user who edited the text information Tx, and may also be the identification information of the user terminal 30 (the same applies hereinafter). Furthermore, if the image generation device 10 is operated by a single user, this User ID is not required. The date and time is the date and time when the text information Tx was edited.

[0078] Furthermore, "text information" is a caption generated by the generation unit 22 for the input image group G1, and is a single sentence or a group of words. Furthermore, "edited text information" is a single sentence or a group of words resulting from editing performed by the user on the above text information.

[0079] Furthermore, the processing of the generation unit 22 may also include the process of selecting the form of the content to be generated (output form) according to the conditions of the text information Tx (hereinafter referred to as text conditions). The text conditions, for example, indicate what kind of vocabulary is included in the text information Tx. The form of the content, for example, indicates the type of content composed of the interpolated image G2 and the input image group G1.

[0080] The generation unit 22 has a generation format table 226, illustrated in Figure 12, for selecting the format of such content. In the example of the generation format table 226 shown in Figure 12, values ​​such as "continuous shooting," "group of images," and "children, group of images" are set as "text conditions," and "output content" such as "continuous GIF," "collage," and "memory video" are set for these "text conditions."

[0081] If the text information Tx generated based on the input image group G1 includes, for example, "continuous shooting," the generation unit 22 will match this with the above-mentioned generation format table 226 and select "continuous GIF" as the "output content." In this case, the generation unit 22 determines that the input image group G1 consists of "a series of images taken in a short time," i.e., continuous shooting images, and generates and outputs the vocabulary "continuous shooting" as text information Tx. Alternatively, the generation unit 22 identifies that the shooting mode is continuous shooting mode from the EXIF ​​tag information of at least one of the individual images G1A included in the input image group G1, and generates and outputs text information Tx containing the vocabulary "continuous shooting" for that input image group G1.

[0082] These burst shots are a series of images capturing a single moment in a subject's movement. By combining these images into a data format that displays them sequentially (e.g., a continuous GIF), a video can be generated. However, if the interval between shots in the burst shots is not sufficiently short, the subject's movement in the generated video tends to be jerky, making it difficult to create a smooth video that is watchable.

[0083] The generation unit 22 reads the EXIF ​​tag information value of at least one image from the continuous shooting images and determines that the shooting interval is "1 fps" based on the continuous shooting mode conditions indicated by that value. Alternatively, the generation unit 22 calculates the shooting interval between images from the shooting date and time information of each image that makes up the continuous shooting images.

[0084] Furthermore, the generation unit 22 estimates the actual speed of the subject's movement from the subject and situation of each image that makes up the continuous shooting sequence, and determines a suitable "shooting interval" based on the estimation result. For example, if the subject is an elementary school student and the situation is a foot race at a sports day, the generation unit 22 determines that "15 fps" is a suitable shooting interval. The generation unit 22 then generates 14 interpolated images G2 for each of the already obtained continuous shooting images (input image group G1).

[0085] Furthermore, the determination of such a suitable shooting interval can be made, for example, by using a table that defines suitable shooting interval values ​​for each combination of subject and situation, or by inputting individual images G1A into the generating AI 221 to identify a suitable shooting interval. In this case, the generating AI 221 is assumed to have a trained model that has been trained on training data combining the above-mentioned combinations of subject and situation and suitable shooting interval values.

[0086] When generating the interpolated image G2, the generation unit 22 generates an image in which the position, orientation, and size of each part of the subject between each preceding and succeeding image (the individual images G1A that make up the continuous shooting image) are gradually changed at unit time intervals as needed, so that the movement of the subject in each preceding and succeeding image (the individual images G1A that make up the continuous shooting image) becomes continuous by inserting the interpolated image G2 in between. For example, the position, orientation, and size of each part of the subject between each preceding and succeeding image and the interpolated image G2, and the position, orientation, and size of each part of the subject between each interpolated image G2, are gradually changed at unit time intervals as needed. The interpolated image G2 thus generated, when added between the individual images G1A, becomes content. Therefore, the generation unit 22 stores such content in the generated content table 224 (see Figure 10). Alternatively, instead of making the above selection using the generated form table 226, the generation unit 22 may accept user specifications regarding the form of content to be generated from the user terminal 30 and generate the interpolated image G2 corresponding to this user specification.

[0087] Furthermore, it is preferable that the generation unit 22 or the receiving unit 21 perform a selection process on the individual images G1A that make up the input image group G1 when generating text information Tx (first feature information) and the classification results (second feature information) of individual images G1A, and narrow down the number of such images. In this case, for example, the generation unit 22 maintains a selection rule table 227 and uses it for the above selection.

[0088] Figure 13 shows an example of the selection rule table 227 in this embodiment. This selection rule table 227 is a table that defines the images to be removed according to the number of individual images G1A that make up the input image group G1. In the example in Figure 13, for example, if the number of images is 1,000 or more but less than 2,000, the selection methods specified are 1) remove blurry images and 2) remove images with faces cut off. If the number of images is 2,000 or more but less than 5,000, in addition to the selection methods 1) and 2) above, the selection method specified is 3) remove duplicate images with similar compositions. If the number of images is 5,000 or more, in addition to the selection methods 1) to 3), the selection methods specified are 4) remove duplicate images of the same subject with the same expression and 5) remove duplicate images of the same shooting location.

[0089] In the above selection process, the generation unit 22 (or reception unit 21) counts the number of individual images G1A in the input image group G1 and compares this to the "criteria" in the selection rule table 227. As a result, it identifies the selection rule corresponding to the number of images and recognizes the selection method. For example, if the number of individual images G1A is "1500," the selection rule "R1" is identified, and as its "selection method," two processes are executed on the input image group G1: 1) remove blurry images, and 2) remove images with faces cut off. The generation unit 22 (or reception unit 21) either has known judgment algorithms for removal of blurry images, images with faces cut off, etc., as defined in the "selection method" column, or it can call and use the corresponding functions from the external server 50.

[0090] (Output Unit) The output unit 23 processes the text information Tx and interpolated image G2 generated by the generation unit 22, or content composed of the input image group G1 and the interpolated image G2, to display on an output device 15 such as a display, a user terminal 30, or a shooting device 40. For the display of such images and content, the output unit 23 retrieves data from the generated text table 222, the interpolated image table 223, the generated content table 224, etc., as appropriate, following the processing in the generation unit 22, or in response to a request from the user terminal 30 or the arrival of a predetermined timing, and outputs it.

[0091] The method of outputting text information Tx, interpolated images G2, etc., is not particularly limited, but for example, the output unit 23 may display the text information Tx, interpolated images G2, and content on the output device 15, user terminal 30, or the display or monitor of the shooting device 40, print the text information Tx, interpolated images G2, and content, transmit the text information Tx, interpolated images G2, and content to other users, and provide the text information Tx, interpolated images G2, and content as commercial goods. Commercial goods may include media consisting of one or more pages or cards on which images are posted, such as albums, photobooks, postcards, message cards, electronic albums, and bromide prints.

[0092] [Example of Image Generation Method Flow] Next, as an example of the operation of the image generation device 10 in this embodiment, an image processing flow using the device will be described. The image processing flow described below uses the image generation method of the present invention. In other words, each step in the image processing flow described below corresponds to a component of the image generation method of the present invention. Note that the following flow is merely an example, and some steps in the flow may be deleted, new steps added to the flow, or the execution order of two steps in the flow may be changed without departing from the spirit of this embodiment.

[0093] Each step in the image generation flow according to this embodiment is performed by the processor 11 of the image generation device 10 in the order shown in Figure 14. In other words, in each step of the image generation flow, the processor 11 executes the data processing specified in the image generation application program, which corresponds to each step in Figure 14.

[0094] To explain in more detail, in the image generation flow according to this embodiment, first, the reception unit 21 receives input of the input image group G1 to be processed from, for example, the user terminal 30 (S1). During this process, the reception unit 21 delivers a reception screen G10 including U1 shown in Figure 15 to the user terminal 30. The reception unit 21 also acquires the input image group G1 via the file specification field G11 of U1 on the reception screen G10. The reception unit 21 also stores the acquired input image group G1 in the input image table 211.

[0095] Next, the generation unit 22 inputs the input image group G1 obtained in S1 to the generation AI 221 to generate text information Tx, which is the caption for the input image group G1 (S2). An example of the text information Tx generated here is shown in Figure 4. The generation unit 22 links the text information Tx generated here with the information of the input image group G1 and stores it in the generation text table 222.

[0096] Prior to the processing in S2, it is preferable to perform a filtering process for individual images G1A in the input image group G1 to improve the efficiency of subsequent processing. In that case, the generation unit 22 (or reception unit 21) executes the flow shown in Figure 16. The generation unit 22 counts the number of individual images G1A included in the input image group G1 obtained in S1 (S20).

[0097] Next, the generation unit 22 compares the number of images counted in the process of S21 with the value in the "perspective" column of each record in the selection rule table 227 (see Figure 13) (S21). If the result of this comparison is that there is no rule in which the number of images matches the value in the "perspective" column (S22:N), the generation unit 22 terminates this flow. On the other hand, if the result of the above comparison is that there is a rule in which the number of images matches the value in the "perspective" column (S22:Y), the generation unit 22 identifies the "selection method" defined in that rule (S23). The generation unit 22 also applies the processing defined by the identified selection method to the input image group G1 and performs filtering of individual images G1A in the input image group G1 (S24).

[0098] Now, let's return to the explanation based on the flow chart in Figure 14. The generation unit 22 performs classification processing on each individual image G1A in the input image group G1 obtained in S1, from the perspective of the event indicated by the text information Tx obtained in the processing of S2. For example, if the text information Tx is "various facial expressions", the generation unit 22 analyzes each individual image G1A in the input image group G1 with respect to the event "(subject's) facial expression". This analysis can be performed using the generation AI 221 which has been trained on training data consisting of pairs of images of people and the facial expressions (correct answers) of those people.

[0099] Suppose the above-mentioned generating AI 221 performs the facial expression analysis and outputs a classification result (second feature information) such as, for example, that among the individual images G1A in the input image group G1, there are 4 images with smiling expressions, 4 images with troubled expressions, 2 images with angry expressions, and 2 images with neutral expressions (S3). The generating unit 22 compares this classification result with the text information Tx obtained in S2 against the missing image judgment rule 225 (see Figure 7) (S4).

[0100] As a result of the above matching, the generation unit 22 identifies the values ​​"laughing face, smiling face, angry face, troubled face, and expressionless face" as "essential patterns / judgment rules" corresponding to the event "various facial expressions" indicated by the text information Tx. Therefore, the generation unit 22 determines that there is a lack of images of laughing faces in the input image group G1 (S5).

[0101] Next, the generation unit 22 prompts the generation AI 221 to generate an image in which the subject is making a "laughing face," and generates the interpolated image G2 (S6). Upon receiving the prompt, the generation AI 221 uses the face image of the subject in the individual image G1A as a base and changes the expression to "laughing face" to generate the interpolated image G2. The generation AI 221 is a trained model that automatically generates images in response to prompts based on the prompts and reference images that serve as the basis for the generation process (e.g., individual image G1A). In the processing of the generation unit 22, the generated interpolated image is stored in the interpolated image table 223 (see Figure 8).

[0102] Furthermore, it is preferable that the generation unit 22 determines what kind of content the interpolated image G2 will constitute, that is, the form of the content, and generates the interpolated image G2 accordingly. An example of the flow for performing such a procedure is shown in Figure 17.

[0103] In this case, the generation unit 22 compares the text information Tx generated based on the input image group G1 with the generated morphology table 226 (S30). At this time, morphological analysis is performed on the text information Tx as needed to identify words, and then those words are compared with the generated morphology table 226. In this case, the generation unit 22 either has a known morphological analysis tool on hand or can call a function from an external server 50 or the like and use it as appropriate.

[0104] Note that the text information Tx is not limited to what the generating AI 221 generates as a caption by interpreting the input image group G1. For example, it may also include information obtained by the generating unit 22 by extracting shooting mode information from the EXIF ​​tag information (ancillary information) of at least one of the individual images G1A included in the input image group G1. A specific example of such information is information that indicates "continuous shooting mode".

[0105] If the above text information Tx includes, for example, "continuous shooting," then when this is matched against the generation format table 226, "Continuous GIF" can be identified as the value of "Output Content" from the record "No. 2" where the "Text Condition" column is "continuous shooting" (S31). Such a series of images produced by "continuous shooting," i.e., continuous shooting images, are a sequence of images that capture a single moment in the movement of the subject. Therefore, it is also possible to generate a video by connecting these images and displaying them in a data format (e.g., continuous GIF).

[0106] However, if the shooting interval between images in the above-mentioned burst images is not short enough to be processed into a video, it becomes necessary to generate an interpolation image G2 to fill in the gaps between those images. For this reason, the generation unit 22 reads the EXIF ​​tag information value of at least one image in the burst images and determines the shooting interval to be "1 fps" or the like based on the burst mode conditions indicated by that value. Alternatively, the generation unit 22 calculates the shooting interval between images from the shooting date and time information of each image that makes up the burst images.

[0107] Furthermore, the generation unit 22 estimates the actual speed of the subject's movement from the subject and situation of each image that makes up the continuous shooting image, and determines a suitable "shooting interval" based on the estimation result (S32). For example, if the subject is an elementary school student and the situation is a foot race at a sports day, the generation unit 22 determines that "15 fps" is a suitable shooting interval. The generation unit 22 then generates 14 interpolated images G2 for each of the already obtained continuous shooting images (input image group G1).

[0108] Furthermore, the determination of such a suitable shooting interval can be made, for example, by using a table that defines suitable shooting interval values ​​for each combination of subject and situation, or by inputting individual images G1A into the generating AI 221 to identify a suitable shooting interval. In this case, the generating AI 221 is assumed to have a trained model that has been trained on training data combining the above-mentioned combinations of subject and situation and suitable shooting interval values.

[0109] When generating the interpolated image G2, the generation unit 22 generates an image in which, for example, the position, orientation, and size of each part of the subject between each of the preceding and succeeding images (the individual images G1A that make up the continuous shooting image) are gradually changed at unit time intervals as necessary, by inserting the interpolated image G2 in between (S33).

[0110] Figure 18 shows an example where interpolated images G2A to G2D, in which the position, orientation, and size of the subject are gradually changed per unit time, are added between the individual images G2X and G2Y that made up the input image group G1, resulting in content G50 which captures elementary school children in a foot race. In other words, content G50 is created when the generated interpolated images G2 are added between the individual images G1A. The generation unit 22 stores this content G50 in the generated content table 224 (see Figure 10) and terminates this flow. Alternatively, instead of making the above selection using the generated form table 226, the generation unit 22 may accept user specifications regarding the form of the content G50 to be generated from the user terminal 30 and generate the interpolated images G2 corresponding to these user specifications.

[0111] Although specific embodiments of the present invention have been described above, these embodiments are merely examples given to facilitate understanding of the present invention and do not limit it. That is, the present invention can be modified or improved from the embodiments described below, without departing from its spirit. Furthermore, the present invention includes equivalents thereof. Moreover, embodiments of the present invention may include forms that combine the above embodiments with one or more of the following modifications.

[0112] (Regarding the computer constituting the image generation device) In the above embodiment, the image generation device 10 of the present invention is configured with a computer directly used by the user, such as a computer or server owned and operated by a content generation service provider or a digital camera user. However, it is not limited to this, and the image generation device 10 of the present invention may be configured with a computer that can be used indirectly by the user, for example, an external server 50. Here, the external server 50 may be, for example, an external server for cloud services, specifically an external server for ASP (Application Service Provider), SaaS (Software as a Service), PaaS (Platform as a Service), or IaaS (Infrastructure as a Service).

[0113] In this case, when the user inputs the necessary information on the user terminal 30 owned by the user, the external server 50 performs various processes (calculations) based on the input information, including the generation of text information Tx and interpolated image G2, and the calculation results are output on the user terminal 30. In other words, the functions of the external server 50 that constitute the image generation device 10 of the present invention can be used on the user terminal 30. Alternatively, the image generation device 10 may be configured with a computer or external server 50, as well as a camera 40 or smartphone (a type of user terminal 30) used by the user.

[0114] (Regarding the processor configuration) In this embodiment, each process is executed on any computer. Furthermore, any computer may execute these processes using a processor, a program, or a combination thereof. Any computer may be a general-purpose computer, a computer designed for a specific purpose, a workstation, or any other hardware element capable of executing a program.

[0115] A processor may consist of one or more hardware components, and the type of hardware is not limited. For example, a processor may consist of programmable logic devices such as a CPU (Central Processing Unit), MPU (Micro Processing Unit), FPGA (Field Programmable Gate Array), dedicated circuits for performing specific processing such as an ASIC (Application Specific Integrated Circuit), a GPU (Graphic Processing Unit), or an NPU (Neural Processing Unit).

[0116] Furthermore, the processor has various units or means that execute the various processes in this embodiment. The type of hardware may also be a combination of different types of hardware. When multiple hardware components are configured to execute one or more processes of a given processor, these multiple hardware components may reside in physically separate devices or in the same device. In any embodiment, the order of the processes performed by the processor is not limited to the order described above and may be changed as appropriate. The hardware is composed of an electrical circuit (circuitry) or the like, which is a combination of circuit elements such as semiconductor elements.

[0117] Furthermore, this embodiment may be implemented by hardware, software, firmware, microcode, or a combination thereof. The software, firmware, and microcode are composed of a program. The program may also be, for example, a group of program modules, and each of its functions may be implemented by a processor configured to perform its respective function. The program may be program code or multiple code segments stored in one or more non-temporary computer-readable media (e.g., storage media or other storage).

[0118] A program may be divided and stored on multiple non-temporary computer-readable media located on devices that are physically separated from each other. Program code, or code segments, may represent any combination of procedures, functions, subprograms, routines, subroutines, modules, software packages, classes, or instructions, data structures, or program statements. Program code, or code segments, may be connected to other code segments or hardware circuits by sending and receiving information, data, arguments, parameters, or memory contents.

[0119] 1 Network 5 Image Generation System 10 Image Generation Device 11 Processor 12 Auxiliary Storage Device 121 Program 13 Main Memory Device 14 Input Device 15 Output Device 16 Communication Device 21 Reception Unit 211 Input Image Table 212 Editing Information Table 22 Generation Unit 221 Generation AI (Trained Model) 222 Generation Text Table 223 Interpolated Image Table 224 Generation Content Table 225 Missing Image Judgment Rule 226 Generation Format Table 227 Selection Rule Table 23 Output Unit 30 User Terminal 40 Shooting Device 50 External Server G1 Input Image Group G1A Individual Image G2 Interpolated Image G50 Content

Claims

1. An image generation device equipped with a processor, wherein the processor generates first feature information common to two or more images, and generates an image based on second feature information relating to at least one of the two or more images and the first feature information.

2. The image generation apparatus according to claim 1, wherein the processor generates text information as the first feature information by inputting the two or more images into a trained model that outputs text information for an input image.

3. The image generation apparatus according to claim 1, wherein the processor generates the second feature information by inputting at least one of the two or more images into a trained model that outputs text information for an input image.

4. The image generation apparatus according to claim 1, wherein the processor extracts the second feature information based on the supplementary information of at least one of the two or more images.

5. The image generation apparatus according to claim 1, wherein the processor accepts editing by a user of the first feature information and generates an image based on the second feature information and the edited first feature information.

6. The image generation apparatus according to claim 1, wherein the processor determines the output form of the generated image based on at least one of the first feature information and the second feature information.

7. The image generation apparatus according to claim 1, wherein the processor accepts a user's specification regarding the output format of the generated image.

8. The image generation apparatus according to claim 1, wherein the processor generates the first feature information and generates the images based on a video including the two or more images.

9. The image generation apparatus according to claim 1, wherein the processor selects an image from the two or more images according to a predetermined rule, and generates the first feature information and the image based on the selected image.

10. An image generation method comprising: a step of generating first feature information common to two or more images using a processor; and a step of generating an image using a processor from second feature information relating to at least one of the two or more images and the first feature information.

11. A program for causing a computer to perform each step included in the image generation method described in claim 10.

12. A computer-readable recording medium on which a program is recorded causing a computer to perform each step included in the image generation method described in claim 10.