Image generating device, image generating method, program, and recording medium

The image generating device and method address the issue of unsuitable images by classifying and generating new images based on existing ones, ensuring they meet user-defined criteria, thereby improving image quality and relevance.

WO2025197346A1PCT designated stage Publication Date: 2025-09-25FUJIFILM CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/004165
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-22
Filing Date
2025-02-07
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Existing image generation technologies fail to produce images appropriate for a specific class when no suitable images are available within that class, leading to unsatisfactory user experience.

Method used

An image generating device and method that classify images into classes, generate new images based on existing images and their information, and arrange them in layouts to meet user-defined criteria, using machine learning models and generative techniques.

Benefits of technology

Enables the creation of images suitable for a class even when no appropriate images are initially available, enhancing user satisfaction by improving image quality and relevance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025004165_25092025_PF_FP_ABST
    Figure JP2025004165_25092025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are an image generating device, an image generating method, a program, and a recording medium capable of generating an image that is appropriate for a class even when there is no image that is appropriate as an image to be classified into said class. An image generating device according to one embodiment of the present invention is provided with a processor which classifies a plurality of first images into one or more classes, and if a class including a first image that satisfies a criterion, among the one or more classes, is deemed to be a target class, then, for a non-target class, among the one or more classes, that does not include a first image satisfying the criterion, the processor generates a second image different from the first image on the basis of at least one of the first image and information related to the first image.
Need to check novelty before this filing date? Find Prior Art

Description

Image generating device, image generating method, program, and recording medium

[0001] An embodiment of the present invention relates to an image generating device, an image generating method, a program, and a recording medium.

[0002] Techniques for generating images are already known, and one example is the technique described in Patent Document 1. The image processing device described in Patent Document 1 uses a trained learning model to generate, from a first image content, a second image content in which the fluctuation elements of the image content have different degrees of fluctuation, thereby making it possible to obtain image content that more appropriately reflects the intention of acquiring the content.

[0003] JP 2023-113444 A

[0004] For example, when selecting images from a plurality of images captured by a photographic device such as a camera, it is preferable to select images classified into different classes rather than selecting only images classified into the same class (group). On the other hand, depending on the class, the images classified into that class may not be satisfactory to the user, in other words, may not be appropriate. In this regard, the technology described in Patent Document 1 can generate new images, but does not take into consideration generating appropriate images to be classified into a class.

[0005] One embodiment of the present invention has been made in consideration of the above circumstances, and aims to provide an image generation device, an image generation method, a program, and a recording medium that are capable of generating an image appropriate for a class even when there is no appropriate image to be classified into that class.

[0006] The above object is achieved by an image generating device described in any of [1] to

[18] below. [1] An image generating device including a processor, wherein the processor classifies a plurality of first images into one or more classes, and, when a class among the one or more classes that includes a first image that satisfies a criterion is designated as a target class, generates a second image different from the first image for a non-target class among the one or more classes that does not include a first image that satisfies the criterion, based on at least one of the first images and information related to the first images. [2] An image generating device including a processor, wherein the processor classifies a plurality of first images into one or more classes, and, when the number of one or more classes determined by the classification of the plurality of first images is less than a set number, generates a new class based on the one or more classes, and generates a second image for the new class that is different from the first image based on at least one of the first images and information related to the first images. [3] The image generating device described in [1] or [2], wherein the processor determines the one or more classes based on one or more scenes obtained by analyzing the plurality of first images. [4] The image generating device according to [1] or [2], wherein the processor determines one or more classes based on one or more predetermined scenes. [5] The image generating device according to any of [1] to [4], wherein the processor, for each of the one or more classes, determines whether at least one first image in the class is a first image that satisfies a criterion. [6] The image generating device according to any of [1] to [5], wherein, for a target class among the one or more classes including a first image that satisfies the criterion, the processor determines a candidate image to be placed in an arrangement area in the arrangement image from among the first images that satisfy the criterion, and for a non-target class, determines a candidate image to be placed in an arrangement area from among the generated second images. [7] The image generating device according to any of [2] to [6], wherein the processor, for each of the one or more classes, determines a candidate image to be placed in an arrangement area in the arrangement image from among the first images included in the class, and for a new class, determines a candidate image to be placed in an arrangement area from among the generated second images.[8] The image generation device according to any one of [1] to [7], wherein, in the process of generating the second image, when a class including a first image that satisfies a criterion is set as a target class among one or more classes, the processor generates the second image based on at least one of the first image in the target class and information related to the first image in the target class. [9] The image generation device according to any one of [1] to [8], wherein, in the process of generating the second image, the processor generates the second image based on at least one of the first image in a non-target class for which the second image is to be generated and information related to the first image in the non-target class for which the second image is to be generated.

[10] The image generation device according to any one of [1] to [9], wherein, in the process of generating the second image, the processor generates a second image in which a plurality of common objects are arranged in a single image area based on two or more first images including a common object.

[11] The image generation device according to any one of [1] to

[10] , wherein the processor assigns display information to at least one of the first image and the second image, indicating whether it is an original image or a generated image.

[12] The image generating device according to any one of [1] to

[11] , wherein the processor, when arranging the first image and the second image in the arrangement area of ​​the arrangement image, respectively, sets an arrangement priority for the first image and the second image.

[13] The image generating device according to any one of [1] to

[12] , wherein the processor, when the generated second image does not satisfy the conditions for being a candidate image, selects a layout corresponding to the total number of classes having candidate images.

[14] The image generating device according to

[13] , wherein, when there is no selectable layout, the processor generates a new layout corresponding to the total number.

[15] The image generating device according to any one of [1] to

[14] , wherein, in the process of generating the second image, the processor generates the second image based on sound information included in the first image.

[16] The image generating device according to any one of [1] to

[15] , wherein, in the process of generating the second image, the processor generates the second image based on information common to two or more first images.

[17] The image generating device according to

[16] , wherein the processor generates the second image by transforming the first image based on the common information.

[18] The image generation device according to any one of [1] to

[17] , wherein in the process of generating the second image, the processor generates the second image using a generative model that has been trained to output the second image in response to an input of the first image.

[0007] The above object can also be achieved by an image generation method described in

[19] or

[20] below.

[19] An image generation method in which a processor executes the following processes: classifying a plurality of first images into one or more classes; and, when a class among the one or more classes that includes a first image that satisfies a criterion is set as a target class, generating a second image different from the first image based on at least one of the first images and information related to the first images for a non-target class among the one or more classes that does not include a first image that satisfies the criterion.

[20] An image generation method in which a processor executes the following processes: classifying a plurality of first images into one or more classes; and, when the number of one or more classes determined by the classification of the plurality of first images is less than a set number, generating a new class based on the one or more classes; and, for the new class, generating a second image different from the first image based on at least one of the first images and information related to the first images.

[0008] A program according to one embodiment of the present invention is a program for causing a computer to execute each step included in the image generation method described in the above

[19] or

[20] . A recording medium according to one embodiment of the present invention is a computer-readable recording medium on which a program for causing a computer to execute each step included in the image generation method described in the above

[19] or

[20] is recorded.

[0009] According to one embodiment of the present invention, an image generation device, an image generation method, a program, and a recording medium are provided that are capable of generating an image appropriate for a class even if there is no appropriate image to be classified into that class.

[0010] FIG. 1 is a diagram showing an example of use of an image generation device according to a first embodiment of the present invention. FIG. 2 is a diagram showing a hardware configuration of an image generation device according to a first embodiment of the present invention. FIG. 3 is an explanatory diagram showing functions of an image generation device according to a first embodiment of the present invention. FIG. 4 is a diagram showing an example of a second image generated by the image generation device according to the first embodiment of the present invention. FIG. 5 is a diagram showing an example of an image generation flow according to the first embodiment of the present invention. FIG. 6 is a diagram showing an example of use of an image generation device according to a second embodiment of the present invention. FIG. 7 is an explanatory diagram showing functions of an image generation device according to a second embodiment of the present invention. FIG. 8 is an explanatory diagram showing an example of an image generation flow according to the second embodiment of the present invention.

[0011] Specific embodiments of the present invention will be described below. For ease of explanation, the following description may be given in terms of a GUI (Graphic User Interface). Furthermore, since the basic data processing technologies (communication / transmission technologies, data acquisition technologies, data recording technologies, data processing / analysis technologies, machine learning technologies, image processing technologies, image display technologies, visualization technologies, etc.) required to realize the present invention are well-known technologies, a description of these technologies will be omitted.

[0012] In addition, in this specification, the concept of "device" includes not only a single device that performs a specific function, but also a combination of multiple devices that exist independently and in a distributed manner but cooperate (link) to perform a specific function.

[0013] In addition, in this specification, the term "user" refers to a user of the image generating device of the present invention. Specifically, a user is a person who uses information obtained by the functions of the image generating device of the present invention (more specifically, a second image as a generated image, etc., as described below).

[0014] In addition, in this specification, the term "person" refers to an entity that performs a specific action, and includes individuals, groups, corporations such as companies, and organizations, as well as computers and devices that constitute artificial intelligence (AI). Artificial intelligence (AI) is a technology that realizes intelligent functions such as inference, prediction, and judgment using hardware and software resources.

[0015] Additionally, in this specification, machine learning algorithms may include neural networks, convolutional neural networks, recurrent neural networks, attention, transformers, variational autoencoders, generative adversarial networks, deep learning neural networks, Boltzmann machines, matrix factorization, factorization machines, m-way factorization machines, field-aware factorization machines, field-aware neural factorization machines, support vector machines, Bayesian networks, decision trees, and random forests, as well as other machine learning algorithms.

[0016] <<First embodiment of the present invention>>

[0017] Image generation (hereinafter, this image generation) performed using an image generation device (hereinafter, image generation device 10) and an image generation method according to a first embodiment of the present invention (hereinafter, first embodiment) will be described with reference to FIGS. 1 to 5.

[0018] <Outline of Image Generation According to First Embodiment> This image generation is used, for example, in creating a photo book, a collage image, etc. (hereinafter also referred to as a photo book, etc.). A "photo book" is made up of multiple pages, and each page is made up of an arrangement image (composite image) obtained by arranging multiple images. A "collage image" corresponds to the aforementioned arrangement image, and for example, a collage image may make up each page of a photo book. Furthermore, a "photo book, etc." may be a photo book, etc. made up of digital data, or a photo book, etc. created by printing images.

[0019] This image generation generates a new image (corresponding to the second image) different from the original image (corresponding to the first image) when, for example, the original image is insufficient as a plurality of images to be arranged in the layout image. A case in which the original image is insufficient may be, for example, when a photographing device such as a camera fails to capture the intended photographic image (original image). More specifically, for example, a case in which a photograph is taken with a slow shutter speed set in a situation where quick movement of a person must be captured. Furthermore, when photographing in the rain, a photograph may be taken without setting the photographing mode appropriate for the rainy environment, resulting in a dark overall image with raindrops visible in the image.

[0020] Unless otherwise specified, an "image" refers to digital image data (hereinafter referred to as "image data") that specifies the gradation values ​​of each of the multiple pixels that make up an image. Image data includes raw image data before compression and image data after compression. Compressed image data includes image data that has undergone lossy compression, such as the JPEG format, and image data that has undergone lossless compression, such as the GIF (Graphics Interchange Format) or PNG (Portable Network Graphics) format. An "image" may also include additional information indicating information such as the file name, date and time of capture, and location of capture.

[0021] The "first image" is an original image used to generate the second image, and specifically includes an image captured by a photographic device such as a camera (a captured image), an image obtained by scanning an existing photograph with a scanner or the like (a scanned image), an image drawn using drawing software and CG (Computer Graphics), and an image obtained by editing these images (an edited image). An edited image may be, for example, a still image extracted from a video. The first image may be either a still image or a video, and the video may be a video with sound or a silent video. In the following, unless otherwise specified, the first image will be described as an example of a captured still image.

[0022] The "second image" is an image different from the first image, and is a generated image generated based on at least one of the first image and information related to the first image. The method for generating the second image will be described later. The second image may be either a still image or a video, and the video may be a video with sound or a silent video. In the following, unless otherwise specified, the second image will be described using as an example a case where the second image is a generated still image.

[0023] The image generation will be described in more detail below with reference to Fig. 1. Note that the example shown in Fig. 1 shows both a case where the image generation is not used (upper illustration in Fig. 1) and a case where the image generation is used (lower illustration in Fig. 1) in order to clearly explain the effect of the image generation.

[0024] Specifically, the upper illustration in Fig. 1 shows a state in which a plurality of first images are classified into three classes, a first image is selected for each class, and the three selected first images c1, f1, and m1 are arranged in an arrangement area L1. The arrangement area L1 is an area included in the arrangement images and is an area for arranging a plurality of images. For example, by selecting a layout, the user can set in advance the arrangement position of each image in the arrangement area L1 and the number of frames (number of images) to be arranged.

[0025] The images arranged in the arrangement area L1 are common to each other in higher-level categories and different from each other in lower-level categories. More specifically, in the example shown in FIG. 1 , the three first images c1, f1, and m1 arranged in the arrangement area L1 are images showing the same theme, "amusement park," but are images showing different scenes, "Ferris wheel," "merry-go-round," and "roller coaster." Note that "theme" and "scene" here are merely an example of a combination of a higher-level category and a lower-level category, and may be, for example, a combination of "concept" and "theme," or other combinations. For convenience of explanation, the following description will be given using the combination of "theme" and "scene."

[0026] However, when looking at the three first images c1, f1, and m1 arranged in the arrangement area L1 in the upper picture of Figure 1, the first image f1 showing the "Ferris wheel" scene is an image in which the child's face is not visible, and the user would prefer an image that shows the child's face. Furthermore, the first image c1 showing the "roller coaster" scene is an image in which the entire image is blurred, and the user would prefer an image that is in focus. In response to this, by using this image generation, as shown in the lower picture of Figure 1, it is possible to newly generate a second image f2 showing the "Ferris wheel" scene in which the child's face is visible, and a second image c2 showing the "roller coaster" scene in which the entire image is in focus.

[0027] More specifically, in this image generation, a plurality of first images are first classified into one or more classes ("Ferris wheel," "merry-go-round," and "roller coaster"). Then, for the classes ("Ferris wheel" and "roller coaster") that do not contain a first image that satisfies the criteria, second images (second images c2, f2) different from the first images are generated. Details of the "criteria" will be described later. Examples of the "criteria" include that the face of a person as a subject in the first image faces forward and that the degree of focus (blurring) of the first image is equal to or greater than a threshold. In this way, by using this image generation, even if there is no suitable image to be classified into a class, it is possible to generate an image appropriate for that class.

[0028] <Configuration Example of Image Generation Device According to First Embodiment> Next, a configuration example of the image generation device 10 according to this embodiment will be described with reference to FIG. 2 . The image generation device 10 is composed of a computer used by a user, specifically a client terminal, and is composed of, for example, a smartphone, a tablet terminal, or a personal computer (PC) such as a laptop or desktop PC. Note that the image generation device 10 is not limited to a computer owned by the user, and may be composed of a terminal installed in a store, etc., that is not owned by the user but can be used by entering a PIN number, password, or making a deposit when visiting a store, etc. Note that the following description will be given taking as an example a case where the image generation device 10 is composed of a user-owned computer, specifically a PC.

[0029] As shown in FIG. 2, the computer that constitutes the image generating device 10 includes a processor 10a, a memory 10b, a communication interface 10c, a storage 10d, an input device 10e, and an output device 10f.

[0030] The processor 10a is configured by, for example, a central processing unit (CPU), a micro-processing unit (MPU), a microcontroller unit (MCU), a graphics processing unit (GPU), a digital signal processor (DSP), a tensor processing unit (TPU), an application specific integrated circuit (ASIC), etc. The memory 10b is configured by, for example, semiconductor memory such as a read only memory (ROM) and a random access memory (RAM).

[0031] The communication interface 10c may be configured, for example, by a network interface card, a communication interface board, etc. The computer constituting the image generating device 10 can communicate with other devices connected to a communication network such as the Internet or a mobile communication line via the communication interface 10c.

[0032] The storage 10d may be configured, for example, by a flash memory, a hard disc drive (HDD), a solid state drive (SSD), a flexible disc (FD), a magneto-optical disc (MO disc), a compact disc (CD), a digital versatile disc (DVD), a secure digital card (SD card), or a universal serial bus memory (USB memory). The storage 10d may be built into a computer constituting the image generation device 10, or may be attached to the computer in an external format. Alternatively, the storage 10d may be configured by a network attached storage (NAS) or the like. The storage 10d may also be an external device, such as an online storage or a database server, that can communicate with one of the computers constituting the image generation device 10 via a communication network.

[0033] The input device 10e is a device that accepts user input operations and is configured, for example, by a touch panel, a keyboard, etc. The input device 10e may also include a photographing device such as a camera built into a PC or smartphone, a microphone for collecting sound, etc. The output device 10f is configured, for example, by a display, a speaker, etc.

[0034] Furthermore, the computer constituting image generation device 10 has installed therein, as software, a program for an operating system (OS) and an application program for executing image generation (hereinafter referred to as an image generation app). These programs are read and executed by processor 10a, causing the computer constituting image generation device 10 to perform its functions, specifically, to execute a series of processes related to image generation. The image generation app may be acquired by reading it from a computer-readable recording medium, or by downloading it via a communication network such as the Internet or an intranet.

[0035] <Functions of Image Generation Device According to First Embodiment> The configuration of the image generation device 10 will be described again from the functional perspective with reference to FIG. 3. As shown in FIG. 3, the image generation device 10 has an acquisition unit 21, a classification unit 22, a determination unit 23, a generation unit 24, and an arrangement unit 25. These functional units are realized by the processor 10a of the image generation device 10 executing the image generation application described above and working in cooperation with other hardware devices of the image generation device 10. In addition, some functions may be realized using artificial intelligence (AI). Each functional unit will be described below.

[0036] [Acquisition Unit] The acquisition unit 21 acquires a plurality of first images to be used in generating the main image. The acquisition unit 21 may acquire, for example, images stored in the storage 10d as the first images. Alternatively, the acquisition unit 21 may acquire, for example, images stored in an external device that can communicate via the communication interface 10c, such as an online storage or a database server, as the first images.

[0037] [Classification Unit] The classification unit 22 classifies the multiple first images acquired by the acquisition unit 21 into one or more classes. Note that "classifying" here also includes, for example, when there is only one scene, classifying into two classes: one class that corresponds to the scene and one class that does not correspond to the scene.

[0038] Specifically, the classification unit 22 first determines one or more classes based on one or more predetermined scenes, for example, one or more scenes input (selected) by a user via an image generation app. Note that "determining one or more classes based on one or more scenes" means, for example, determining one class for one scene. The classification unit 22 extracts one or more image features from each of the multiple first images, determines the similarity between the extracted image features and the image features representing each scene, and classifies each of the multiple first images into a class corresponding to the scene with the highest similarity.

[0039] "Image features" include information about the image quality of each region of the image, the gradation values ​​of the pixels contained in each region, and information about the subject estimated from this information, such as the size of the subject region relative to the image and the sharpness of the image. "Subject" refers to people, animals, objects, and backgrounds contained in the image. "Subject information" may include the type of subject, the state of the subject, the position of the subject in the image, and, if the subject is a person, facial expression, etc. "Image features" are preferably features that can be quantified, vectorized, or tensorized. In this case, the results of image analysis are quantified, vectorized, or tensorized image features, i.e., feature quantities.

[0040] To explain the classification process in more detail, for example, assume that the user has predetermined four scenes representing the emotions of joy, anger, sadness, and happiness. In this case, the classification unit 22 determines the four scenes as four classes. Then, the classification unit 22 identifies the facial expression of the person as the subject for each of the multiple images, and classifies the multiple images into each of the four classes based on information about the identified facial expressions of the subjects.

[0041] In the above description, the classification unit 22 determines one or more classes based on one or more scenes input by the user through the image generation app. However, this is not limiting, and for example, the classification unit 22 may determine one or more classes based on one or more scenes that are predetermined as initial settings of the image generation app.

[0042] In this way, in the image generating device 10, by determining one or more classes based on one or more predetermined scenes, it is possible to appropriately reflect, for example, the intention of a user or the like in the classification of a plurality of first images. Note that in the above description, with regard to the correspondence between scenes and classes, the classification unit 22 determines one class for one scene, but this is not limited to this, and, for example, one class may be determined for multiple scenes, or multiple classes may be determined for one scene.

[0043] In the above description, the classification unit 22 determines one or more classes based on one or more predetermined scenes. However, this is not limiting. For example, the classification unit 22 may determine one or more classes based on one or more scenes obtained by analyzing a plurality of first images. Specifically, the classification unit 22 may identify one or more scenes common to the plurality of first images using a known scene recognition technique, and determine one or more classes based on each of the identified one or more scenes.

[0044] 1 , the classification unit 22 identifies scenes of a Ferris wheel, a merry-go-round, and a roller coaster from the plurality of first images using a known scene recognition technique, and determines three classes based on the three identified scenes. The classification unit 22 may then classify the first image f1 related to a Ferris wheel into a class corresponding to a Ferris wheel scene, the first image m1 related to a merry-go-round into a class corresponding to a merry-go-round scene, and the first image c1 related to a roller coaster into a class corresponding to a roller coaster scene. In this way, the image generating device 10 can easily classify the plurality of first images by determining one or more classes based on one or more scenes obtained by analyzing the plurality of first images.

[0045] Furthermore, the classification unit 22 may determine one or more classes based on, for example, incidental information of the first image (for example, EXIF ​​(Exchangeable Image File Format) etc.) For example, the classification unit 22 may determine one or more classes based on information on the shooting date and time included in the incidental information of each of the multiple first images, using periods when no shooting was performed as class boundaries (classification criteria).

[0046] Furthermore, the classification unit 22 may use a classification model that has been trained to classify a plurality of first images into one or more classes by inputting the plurality of first images. The classification model may be constructed, for example, by performing machine learning using the training images and the classes into which the training images are classified as training data.

[0047] [Determination Unit] The determination unit 23 determines, for each of one or more classes, whether at least one first image in the class is a first image that satisfies the criteria. In the following description, it is assumed that the determination unit 23 determines whether all first images in the class are first images that satisfy the criteria. However, this is not limited to this. For example, if the determination unit 23 finds one first image that satisfies the criteria in a class, it may not be necessary to determine whether the remaining first images in that class are first images that satisfy the criteria. In the following description, a class among the one or more classes that includes a first image that satisfies the criteria will be referred to as a "target class," and a class among the one or more classes that does not include a first image that satisfies the criteria will be referred to as a "non-target class."

[0048] Specifically, the determination unit 23 calculates an evaluation score for each first image in the class. If the calculated evaluation score satisfies a criterion, the determination unit 23 determines that the first image to be determined is a first image that satisfies the criterion. The "criterion" may be, for example, that the facial expression of the person as the subject is a smile, that the person's face is facing forward, that the size of the person's person area is a certain proportion or more of the image area, that the degree of match between the background in the image and the scene in the class is a threshold or more, that the brightness (luminosity, etc.) of the entire image is a threshold or more, and that the degree of match between the focus of the image (degree of blur) is a threshold or more.

[0049] To explain an example of a method for calculating the evaluation score, the determination unit 23 first calculates image features for each first image in one or more classes. It is assumed that multiple types of features are extracted from each first image, and feature values ​​are calculated for each feature. The determination unit 23 then calculates an evaluation value based on the calculated feature values ​​and a predetermined threshold. The "evaluation value" may be, for example, the difference between the threshold and the feature value. For example, the higher the evaluation, the larger the "+" value, and the lower the evaluation, the larger the "-" value. The determination unit 23 then calculates the evaluation score of the first image by, for example, summing the evaluation values ​​calculated for each image feature. In this way, the image generating device 10 can appropriately separate target classes and non-target classes by determining, for each of one or more classes, whether at least one first image in the class is a first image that satisfies a criterion.

[0050] Furthermore, the determination unit 23 may use an evaluation model that has been trained to calculate an evaluation score by inputting the first image. The evaluation model may be constructed by performing machine learning using training images and the evaluation scores of the training images as training data.

[0051] [Generation Unit] The generation unit 24 generates a second image that is different from the first image for a non-target class among one or more classes, based on at least one of the first image and information related to the first image. Hereinafter, the generation of the second image based on the first image and the generation of the second image based on information related to the first image will be described separately.

[0052] (Generation of Second Image Based on First Image) The generation unit 24 generates the second image based on the first image. Specifically, the generation unit 24 may generate the second image by combining a person in one of the two first images with a background in the other of the first images.

[0053] For example, when generating the second image f2 shown in the lower part of Figure 1, the generation unit 24 extracts a background region including a Ferris wheel from the first image f1 in a class corresponding to a Ferris wheel scene. Meanwhile, the generation unit 24 identifies a first image from one or more classes that displays a face of a person presumably identical to the person in the first image f1, and extracts a person region of the person from the identified first image. The generation unit 24 then generates the second image by combining the extracted background region and person region.

[0054] Furthermore, when generating the second image c2 shown in the lower part of FIG. 1 , the generation unit 24 extracts a background region including a roller coaster from the first image c1 in a class corresponding to a roller coaster scene, and performs, for example, sharpening on the background region. Meanwhile, the generation unit 24 identifies a first image from one or more classes in which a person presumably identical to the person in the first image c1 is in focus, and extracts a person region of the person from the identified first image. The generation unit 24 then generates the second image by combining the background region subjected to the sharpening process with the extracted person region. Note that if the extracted person's posture does not match the posture appropriate for the scene, the generation unit 24 may, for example, perform a correction process to correct the extracted person's posture (skeleton) to a posture appropriate for the scene.

[0055] The generation unit 24 may also generate the second image using a generative model trained to output a second image in response to input of multiple first images. The generative model may be constructed, for example, by performing machine learning using multiple training images and a generated image generated from the training images as training data. The generation unit 24 may automatically input all of the multiple first images acquired by the acquisition unit 21 into the generative model, or may, for example, present the determination result by the determination unit 23 to the user, have the user select a first image from the multiple first images based on the determination result, and input the selected first image into the generative model. In this way, the image generation device 10 can appropriately generate a second image to be placed in the placement area by using the generative model.

[0056] The generation unit 24 may also generate a second image in which multiple common subjects are arranged within a single image region based on two or more first images containing the common subject. In other words, the generation unit 24 may generate a second image expressed using a so-called multiple exposure technique. For example, consider a case in which, for a non-target class, two or more first images are taken within a predetermined period, and each image contains a large amount of movement of a common person and little movement of the background. In this case, the generation unit 24 may synthesize the person regions of the common person included in each first image into the image region of one of the two or more first images (e.g., the last first image taken). The example shown in FIG. 4 illustrates an example of a second image e2 generated by the generation unit 24. The second image e2 contains multiple common people arranged within the image region, and the time-series movements of the people are expressed.

[0057] Furthermore, as the second image expressed using the multiple exposure technique, a second image taken from a shooting location different from the shooting location shown in Figure 4 may be generated. For example, the generation unit 24 may extract a first image taken from a shooting location different from the shooting location shown in Figure 4 and generate a new second image in which multiple common subjects are arranged within the image area of ​​the extracted first image. This makes it possible to generate a second image expressed using the multiple exposure technique that appears to have been taken from a shooting location that was not actually taken. In this way, the image generation device 10 can generate a second image expressed using the so-called multiple exposure technique.

[0058] In addition, the first image used when generating the second image is not particularly limited, and specifically, the generation unit 24 may generate the second image based on a first image within any of one or more classes.

[0059] However, without being limited thereto, for example, the generation unit 24 may limit the first image used when generating the second image to, for example, a first image in a non-target class. In this case, the generation unit 24 may generate the second image based on a first image in the same non-target class, i.e., a first image in the non-target class for which the second image is to be generated. In other words, when there are two or more non-target classes, the second images belonging to each non-target class are generated based on the first images belonging to the respective non-target classes.

[0060] Similarly, the generation unit 24 may generate the second image based on information related to a first image (described later) in a non-target class for which the second image is to be generated. That is, the generation unit 24 may generate the second image based on at least one of the first image in the non-target class and information related to the first image in the non-target class. This allows the image generation device 10 to place a second image that is more suitable for a scene of the non-target class in the placement area.

[0061] Furthermore, the generation unit 24 may limit the first images used when generating the second images to, for example, first images within the target class. That is, the generation unit 24 may generate the second images based on the first images within the target class. Similarly, the generation unit 24 may generate the second images based on information related to the first images within the target class (described below). That is, the generation unit 24 may generate the second images based on at least one of the first images within the target class and information related to the first images within the target class. This allows the image generation device 10 to generate second images that share a common category (theme, etc.) with the multiple first images, thereby providing a sense of unity among the multiple images placed in the placement area.

[0062] (Generation of second image based on information related to first image) The generator 24 generates the second image based on information related to the first image. Examples of "information related to the first image" include information indicating the content of the first image and sound information included in the first image.

[0063] The "information indicating the content of the first image" includes, for example, information indicating the content of the subject in the first image, and may be any information such as text information, symbol information, etc. Furthermore, the "information indicating the content of the first image" may be information obtained by analyzing the first image, or may be information obtained by input from the user.

[0064] More specifically, the generation unit 24 may generate a second image based on, for example, information common to two or more first images as information related to the first image. For example, if two or more first images in a non-target class each contain a "baby," the generation unit 24 may identify "baby" as the subject from each of the two or more first images in the non-target class and extract the text information "baby" as common information indicating the content of the first image. The generation unit 24 may then generate a second image using an output model trained to output a second image related to "baby" by inputting, for example, "baby" as a keyword. The output model may be constructed, for example, by performing machine learning using training keywords and generated images based on the training keywords as training data. In this way, the image generation device 10 can generate a second image to be placed in the placement area by using information common to two or more first images.

[0065] Furthermore, the generation unit 24 may generate a second image by converting the first image based on the common information. Specifically, assume that two or more first images in the non-target class each display a deformed person. Note that, here, it is assumed that the first images displaying the deformed person are not photographed images but rather hand-drawn images created using drawing software, computer graphics (CG), or the like. For example, when the generation unit 24 identifies a deformed person as a subject from each of two or more first images in the non-target class, it extracts text information such as "deformed" as common information. Then, the generation unit 24 may generate a second image using a conversion model trained to output a second image in which the person in the first image (e.g., a photographed image) is deformed by inputting the first image to be converted and "deformed" as a conversion keyword. The conversion model may be constructed, for example, by performing machine learning using training conversion keywords, training images, and the converted training images as training data. In this way, the image generating device 10 can generate a second image to be placed in the placement area by converting an existing first image.

[0066] The generation unit 24 may also generate the second image based on sound information contained in the first image, such as people's conversations, machine operating sounds (such as car engine sounds), and animal cries. For example, the generation unit 24 may estimate the gender and age of a person contained in the first image based on the frequency and waveform pattern of the person's voice, and generate a second image containing a person with attributes corresponding to the estimated result. In this case, the generation unit 24 may generate the second image using a sound conversion model trained to input a first image containing sound information and output a second image related to the sound information. The sound conversion model may be constructed by performing machine learning using, for example, training images containing sound information and generated images based on the sound information as training data. In this way, the image generation device 10 can generate a second image to be placed in the placement area by using sound information.

[0067] [Placement Unit] When one or more classes include both a target class and a non-target class, the placement unit 25 places a first image included in the target class and a second image included in the non-target class in a placement area where multiple images are placed. On the other hand, when all of the one or more classes are non-target classes, the placement unit 25 places only the second image in the placement area. Note that the following description is based on the assumption that one or more classes include both a target class and a non-target class. Specifically, the placement unit 25 selects a layout according to the number of one or more classes and places the first image and the second image included in each class in the placement area. Note that multiple types of layouts are stored in the storage 10d, and the placement unit 25 acquires layouts corresponding to the number of one or more classes from the storage 10d.

[0068] Furthermore, when placing the first image and the second image in the placement area, the placement unit 25 may set placement priorities for the first image and the second image, respectively. "Setting placement priorities" specifically corresponds to determining the placement position and size of the image in the placement area. "Placement priority" refers to the priority for the placement position and size of the image in the placement area. An image with a higher priority is positioned in a more prominent position within the placement area, specifically, closer to the center of the placement area, while an image with a lower priority is positioned in a less prominent position within the placement area, specifically, closer to the edge of the placement area. Furthermore, an image with a higher priority is displayed at a more prominent size within the placement area, specifically, at a size larger than the reference size, while an image with a lower priority is displayed at a less prominent size within the placement area, specifically, at a size smaller than the reference size.

[0069] For example, the placement unit 25 calculates an evaluation score for the generated second image using the same procedure as when the determination unit 23 calculates the evaluation score for the first image. The placement unit 25 then compares the evaluation scores of the first image and the second image, and if the evaluation score of the second image is lower than the evaluation score of the first image, the placement unit 25 reduces the size of the second image when placed in the placement area compared to the first image. This makes it possible to make the second image, which has a lower evaluation score than the first image, less noticeable than the first image in the placement area. Alternatively, if the evaluation score of the second image is lower than the evaluation score of the first image, the placement unit 25 may place the second image closer to the edge of the placement area than the first image. This makes it possible to make the second image, which has a lower evaluation score than the first image, less noticeable than the first image in the placement area.

[0070] Conversely, if the evaluation score of the second image is higher than the evaluation score of the first image, the placement unit 25 may make the size of the second image larger than that of the first image when it is placed in the placement area, or may place the second image closer to the center of the placement area than the first image. This makes it possible to make the second image, which has a higher evaluation score than the first image, more noticeable than the first image in the placement area. In this way, the image generation device 10 can change the placement priority of the second image in the placement area depending on the evaluation result of the second image, i.e., the degree of completion of the second image as a generated image.

[0071] The arrangement unit 25 may also assign display information to at least one of the first image and the second image, indicating whether the image is an original image or a generated image. For example, the arrangement unit 25 may assign display information indicating that the first image is an original image, or may assign display information indicating that the second image is a generated image, or may assign corresponding display information to both the first image and the second image. The "display information" may be, for example, a mark composed of letters, symbols, pictures, etc., and may be assigned to, for example, the edge of the image area of ​​at least one of the first image and the second image. This allows the user to distinguish between original images and generated images within the arrangement area where multiple images are arranged.

[0072] <Example of Image Generation Method According to First Embodiment> Next, as an operation example of the image generation device 10 according to the first embodiment, an image generation flow using the same device will be described. In the image generation flow described below, the image generation method of the present invention is used. In other words, each step in the image generation flow described below corresponds to a component of the image generation method of the present invention. Note that the flow below is merely an example, and new steps may be added to the flow as long as they do not deviate from the spirit of this embodiment.

[0073] The steps in the image generation flow according to the first embodiment are performed in the order shown in Fig. 5 by the processor 10a included in the image generation device 10. That is, in each process in the image generation flow, the processor 10a executes processing corresponding to each step in Fig. 5 among the data processing defined in the image generation application.

[0074] First, the process is started when a user launches an image generation app installed on the image generation device 10, which generates a signal in response to the launch. After the image generation app is launched, the processor 10a transitions the screen of the display (output device 10f) to a selection screen (not shown) that allows the user to select multiple first images, based on a predetermined operation by the user. Then, when the multiple first images are selected by the user, the processor 10a acquires the multiple first images (S001).

[0075] Next, the processor 10a classifies the acquired multiple first images into one or more classes (S002). More specifically, the processor 10a determines one or more classes based on one or more predetermined scenes, and classifies the multiple first images into the determined one or more classes. Alternatively, the processor 10a determines one or more classes based on one or more scenes obtained by analyzing the multiple first images, and classifies the multiple first images into the determined one or more classes.

[0076] Next, the processor 10a determines whether at least one first image in each of one or more classes satisfies the criteria (S003). For a target class that includes a first image that satisfies the criteria, the processor 10a skips step S004. On the other hand, for a non-target class that does not include a first image that satisfies the criteria, the processor 10a generates a second image based on at least one of the first image and information related to the first image (S004). Note that in this example of the image generation flow, it is assumed that both the target class and the non-target class are included in one or more classes. Then, the processor 10a arranges the first image included in the target class and the second image included in the non-target class in the arrangement area of ​​the arrangement image (S005).

[0077] When the series of processes described above is completed, the image generation flow according to the first embodiment is completed. In the image generation flow, the series of image generation processes shown in FIG. 5 are repeatedly executed by the processor 10 a each time a plurality of first images are selected by the user.

[0078] Effectiveness of the First Embodiment As described above, the image generating device 10 according to the first embodiment generates a second image, which is different from the first image, for a non-target class among one or more classes, based on at least one of the first image and information related to the first image. This makes it possible to generate an image appropriate for a class even when there are no images appropriate for that class. Therefore, for example, even when there are not enough images appropriate for placement in layout images such as a photo book, it is possible to place an appropriate image.

[0079] Second Embodiment of the Present Invention In the first embodiment, the number of one or more classes is determined by classifying the plurality of first images, and then a layout is selected according to the number of the one or more classes. However, this is not limited to this. For example, a layout may be selected in advance by a user or the like, and if the number of one or more classes determined by classifying the plurality of first images is less than the number of classes set for the layout selected by the user or the like, new classes may be generated. This embodiment is referred to as the second embodiment, and will be described below. Note that the second embodiment is substantially the same as the first embodiment, except that the image generating device 10 further includes a class number determination unit 31 and a class generation unit 32 (see FIG. 7 ), which will be described later, and does not include the determination unit 23 of the first embodiment. The following description of the second embodiment will focus on the differences from the first embodiment.

[0080] <Outline of Image Generation According to Second Embodiment> First, an outline of image generation according to the second embodiment will be described with reference to Fig. 6. Note that the example shown in Fig. 6 shows both a case where image generation according to the second embodiment is not used (upper picture in Fig. 6) and a case where image generation according to the second embodiment is used (lower picture in Fig. 6) in order to clearly explain the effect of this image generation according to the second embodiment.

[0081] Specifically, the upper illustration of Fig. 6 shows a state in which a plurality of first images are classified into three classes, a first image is selected for each class, and the three selected first images c3, f3, and m3 are arranged in the arrangement area L2. The layout of the arrangement area L2 is a layout selected in advance by the user, and specifically, a layout corresponding to four classes. However, since the number of classes in the layout selected by the user is "4" while the number of classes determined by classifying the plurality of first images is "3," in the upper illustration of Fig. 6, an area corresponding to one class in the arrangement area L2 is blank.

[0082] In this regard, by utilizing the image generation according to the second embodiment, it is possible to generate a new class ("Parade") that is different from the three existing classes ("Ferris Wheel", "Merry-Go-Round", and "Roller Coaster"), as shown in the lower image of Figure 6, and further to generate a new second image p3 for the new class that is different from the first image.

[0083] More specifically, first, a plurality of first images are classified into one or more classes ("Ferris wheel," "merry-go-round," and "roller coaster"). If the number of one or more classes determined by the classification of the plurality of first images does not reach a set number, a new class ("parade") is generated based on the one or more classes. Furthermore, a second image different from the first image is generated for the new class. In this way, by using the image generation according to the second embodiment, it is possible to generate an image appropriate for a class even if there is no image appropriate for that class.

[0084] 7, the image generating device 10 according to the second embodiment includes an acquisition unit 21, a classification unit 22, a number-of-class determination unit 31, a class generation unit 32, a generation unit 24, and an arrangement unit 25. As described above, the image generating device 10 according to the second embodiment differs from the image generating device 10 according to the first embodiment in that it further includes the number-of-class determination unit 31 and the class generation unit 32, and does not include the determination unit 23 of the first embodiment. The acquisition unit 21 and the classification unit 22 are the same as those in the first embodiment, and therefore description thereof will be omitted.

[0085] [Number of Classes Determination Unit] The number of classes determination unit 31 determines whether the number of one or more classes determined by classifying the multiple first images satisfies a set number. Specifically, the number of classes determination unit 31 first identifies the number of one or more classes determined by the classification unit 22. Meanwhile, the number of classes determination unit 31 determines the number of classes according to, for example, a layout selected by a user as the set number. Then, the number of classes determination unit 31 determines whether the number of one or more classes determined by the classification unit 22 satisfies the set number.

[0086] [Class Generation Unit] When the number of one or more classes determined by categorizing the plurality of first images is less than a set number, the class generation unit 32 generates a new class based on the one or more classes. Specifically, the class generation unit 32 analyzes the one or more classes to generate a new class related to the one or more classes.

[0087] More specifically, the class generation unit 32 extracts a theme (higher-level category) common to one or more classes, identifies a new scene different from the scenes in one or more classes based on the higher-level category, and generates a new class corresponding to the new scene. For example, in the example shown in the lower part of FIG. 6 , the class generation unit 32 extracts "amusement park" as a theme common to classes corresponding to three scenes ("Ferris wheel," "merry-go-round," and "roller coaster"). The class generation unit 32 then identifies a new scene, "parade," based on the common theme and generates a new class corresponding to the new scene. For example, the storage 10d stores a table listing multiple types of scenes for each theme. After extracting the common theme (amusement park), the class generation unit 32 randomly selects a new scene (parade) different from the scenes in one or more existing classes based on the table. As a result, the class generation unit 32 generates a new class corresponding to the new scene.

[0088] Furthermore, the class generation unit 32 may, for example, subdivide a scene into lower categories and generate one or more new classes according to the subdivided categories. The class generation unit 32 may, for example, subdivide a scene based on a classification rule associated with a subject, specifically, the position of a person in an image. For example, when a moving subject is included in the first image, the class generation unit 32 may divide the first image according to the direction in which the subject is moving and set one or more new scenes according to the content of each divided image (image fragment).

[0089] [Generation Unit] The generation unit 24 generates a second image for a new class, which is different from the first image, based on at least one of the first image and information related to the first image. Note that the details of the method for generating the second image are the same as those in the first embodiment, and therefore will not be described here.

[0090] [Placement Unit] When the class generation unit 32 generates new classes in numbers less than the set number, the placement unit 25 places the first images included in each of the one or more classes and the second images included in the new classes in the placement areas of the placement images. On the other hand, when the class generation unit 32 generates one or more new classes in numbers equal to the set number, the placement unit 25 may place only the second images included in each of the one or more new classes in the placement areas. Note that the following description assumes a case where the number of new classes generated is less than the set number, that is, a case where both the first images and the second images are placed in the placement areas.

[0091] <Example of Image Generation Method According to Second Embodiment> Next, referring to FIG. 8 , an image generation flow using the image generation device 10 according to the second embodiment will be described as an example of the operation of the device. Note that a description of points common to the first embodiment will be omitted. After performing preprocessing common to the first embodiment, a user selects multiple first images, and the processor 10a acquires the multiple first images (S101). Next, the processor 10a classifies the acquired multiple first images into one or more classes (S102).

[0092] Next, the processor 10a determines whether the number of one or more classes determined by classifying the multiple first images satisfies a predetermined number (S103). If the number of one or more classes satisfies the predetermined number, the processor 10a skips steps S104 and S105. On the other hand, if the number of one or more classes does not satisfy the predetermined number, the processor 10a generates a new class based on one or more classes (S104). Next, the processor 10a generates a second image for the new class that is different from the first image based on at least one of the first image and information related to the first image (S105). Thereafter, the processor 10a arranges the first image included in each of the one or more classes and the second image included in the new class in the arrangement area of ​​the arrangement image (S106). When the above series of processes are completed, the image generation flow according to the second embodiment ends.

[0093] In the above description, the image generating device 10 according to the second embodiment does not include the determination unit 23 of the first embodiment. However, this is not limited thereto, and the image generating device 10 according to the second embodiment may include the determination unit 23. For example, in the second embodiment, the determination unit 23 may determine, for each of one or more classes, whether a first image in the class is a first image that satisfies a criterion. As a result, when one or more classes include both a target class and a non-target class, the placement unit 25 according to the second embodiment may place, in the placement area, a first image included in the target class, a second image included in the non-target class, and a second image included in a new class.

[0094] <Effectiveness of the Second Embodiment> As described above, in the image generating device 10 according to the second embodiment, if the number of one or more classes does not reach a set number, a new class is generated based on one or more classes. Then, for the new class, the image generating device 10 generates a second image that is different from the first image based on at least one of the first image and information related to the first image. This makes it possible to generate an image appropriate for a class even if there is no image appropriate for classification into that class. In particular, even if the number of one or more classes does not reach a set number, a new class can be generated and an image appropriate for that class can be generated. Therefore, for example, even if there are not enough images appropriate for placement in layout images such as a photo book, appropriate images can be generated.

[0095] <<Third Embodiment of the Present Invention>> In the first embodiment, the first image and the second image included in each of one or more classes are respectively arranged in the arrangement area of ​​the arrangement image. However, this is not limited to this. For example, a candidate image to be arranged in the arrangement area may be determined from the first image and the second image included in each of one or more classes. This embodiment is referred to as the third embodiment, and will be described below. Note that the third embodiment is substantially the same as the first embodiment, except that the image generating device 10 further includes a determination unit 41, which will be described later. The following description will focus on the differences between the first embodiment and the third embodiment.

[0096] 9, the image generating device 10 according to the third embodiment has an acquisition unit 21, a classification unit 22, a determination unit 23, a generation unit 24, a decision unit 41, and an arrangement unit 25. The acquisition unit 21, the classification unit 22, the determination unit 23, and the generation unit 24 are the same as those in the first embodiment, and therefore description thereof will be omitted.

[0097] (Determination Unit) For a target class that includes a first image that satisfies a criterion among one or more classes, the determination unit 41 determines candidate images to be placed in the placement area of ​​the placement image from among the first images that satisfy the criterion. For example, the determination unit 41 may determine M (M is a natural number greater than or equal to 1) first images as candidate images for each target class in descending order of evaluation score. Alternatively, the determination unit 41 may present the first images for each target class to the user in descending order of evaluation score, and allow the user to select a candidate image.

[0098] Furthermore, for non-target classes, the determination unit 41 determines, from the generated second images, candidate images to be placed in the placement area of ​​the placement image. For example, the determination unit 41 calculates the evaluation score of the second image using the same procedure as when the determination unit 23 calculates the evaluation score of the first image. Then, the determination unit 41 may determine, for each non-target class, N second images (N is a natural number greater than or equal to 1) in descending order of evaluation score as candidate images. Alternatively, the determination unit 41 may present, to the user, second images for each non-target class in descending order of evaluation score, and allow the user to select a candidate image. In this way, the image generating device 10 according to the third embodiment determines candidate images for each of one or more classes, thereby generating an image more appropriate for that class.

[0099] (Arrangement Unit) The arrangement unit 25 arranges candidate images for the target class and candidate images for the non-target class in the arrangement area of ​​the arrangement image. If the generated second image does not satisfy the conditions for a candidate image, the arrangement unit 25 may select a layout corresponding to the total number of classes having the candidate images. Specifically, referring to the example shown in FIG. 1 , layouts corresponding to three classes are selected in the lower illustration of FIG. 1 . However, if, for example, the second image c2 does not satisfy the conditions for a candidate image, the arrangement unit 25 may select a layout corresponding to two classes instead of the layout shown in FIG. 1 . As described above, for example, multiple types of layouts are stored in the storage 10d, and the arrangement unit 25 selects a predetermined layout from among these layouts. In this way, the image generating device 10 according to the third embodiment can appropriately arrange candidate images within the arrangement area by selecting a layout corresponding to the total number of classes having candidate images.

[0100] Furthermore, for example, if there is no selectable layout corresponding to the total number of classes having candidate images among the multiple types of layouts stored in the storage 10d, the arrangement unit 25 may generate a new layout corresponding to the total number of classes having candidate images. For example, the arrangement unit 25 may generate a matrix layout according to the total number of classes having candidate images. In this way, the image generating device 10 according to the third embodiment can appropriately arrange candidate images in an arrangement area of ​​a photo book or the like even if there is no selectable layout.

[0101] While the above description has been based on the first embodiment, the present invention is not limited thereto. For example, the image generating device 10 according to the second embodiment may further include a determination unit 41. In this case, the determination unit 41 may determine, for each of one or more classes, a candidate image to be placed in the placement area of ​​the placement image from among the first images included in the class. Furthermore, the determination unit 41 may determine, for a new class, a candidate image to be placed in the placement area of ​​the placement image from among the generated second images. The placement unit 25 may then place, in the placement area, the candidate images for each of the one or more classes and the candidate image of the new class. In this way, the image generating device 10 according to the third embodiment can place more appropriate images by determining the candidate images to be placed in the placement area.

[0102] Other Embodiments (Regarding the Computer Constituting the Image Generation Apparatus) In the above-described embodiment, the image generation apparatus of the present invention is configured by a computer directly used by a user, such as a user-owned terminal (client terminal). However, this is not limited to this, and the image generation apparatus of the present invention may also be configured by a computer indirectly available to a user, such as a server computer. Here, the server computer may be, for example, a server computer for a cloud service, specifically, a server computer for an ASP (Application Service Provider), SaaS (Software as a Service), PaaS (Platform as a Service), or IaaS (Infrastructure as a Service). In this case, when necessary information is input into the client terminal, the server computer performs various processes (calculations) based on the input information, and the calculation results are output on the client terminal. In other words, the functions of the server computer constituting the image generation apparatus of the present invention can be used on the client terminal.

[0103] (Regarding the Processor Configuration) In each embodiment of the present invention, each process is executed by an arbitrary computer. Furthermore, the arbitrary computer may execute these processes using a processor as hardware, a program as software, or a combination thereof. In this case, the processor is configured to execute various processes in each embodiment in cooperation with the program, and may function as each unit or means in each embodiment. Furthermore, the order in which the processes are executed by the processor is not limited to the order described and may be changed as appropriate. The arbitrary computer may be a general-purpose computer, a computer for specific applications, a workstation, or any other system capable of executing each process.

[0104] A processor may be configured with one or more pieces of hardware, and the type of hardware is not limited. For example, a processor may be configured with hardware such as a central processing unit (CPU), a micro processing unit (MPU), a programmable logic device such as a field programmable gate array (FPGA), a dedicated circuit for executing specific processing such as an application specific integrated circuit (ASIC), a graphics processing unit (GPU), or a neural processing unit (NPU). Furthermore, the type of hardware may be a combination of different types of hardware. When multiple pieces of hardware are configured to execute one or more processes of a certain processor, the multiple pieces of hardware may be located in devices physically separated from each other or may be located in the same device. In any embodiment, the order of the processes performed by the processor is not limited to the order described above and may be changed as appropriate. Note that the hardware is configured by an electric circuit (circuitry) that combines circuit elements such as semiconductor elements.

[0105] Furthermore, a program may be software, such as firmware or microcode. Alternatively, a program may be, for example, a group of program modules, each function of which may be implemented by a processor configured to perform the respective function. A program may be program code or multiple code segments stored in one or more non-transitory computer-readable media (e.g., storage media or other storages). The program may be stored across multiple non-transitory computer-readable media that reside in physically separate devices. A program code or code segment may represent a procedure, a function, a subprogram, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A program code or code segment may be connected to another code segment or a hardware circuit by sending or receiving information, data, arguments, parameters, or memory contents.

[0106] REFERENCE SIGNS LIST 10 Image generating device 10a Processor 10b Memory 10c Communication interface 10d Storage 10e Input device 10f Output device 21 Acquisition unit 22 Classification unit 23 Determination unit 24 Generation unit 25 Placement unit 31 Number of classes determination unit 32 Class generation unit 41 Determination unit L1, L2 Placement area c1, c3, f1, f3, m1, m3 First image c2, e2, f2, p3 Second image

Claims

1. An image generating device having a processor, wherein the processor classifies a plurality of first images into one or more classes, and when a class among the one or more classes that includes the first image that satisfies a criterion is designated as a target class, generates a second image different from the first image for a non-target class among the one or more classes that does not include the first image that satisfies the criterion, based on at least one of the first images and information related to the first images.

2. An image generating device having a processor, wherein the processor classifies a plurality of first images into one or more classes, and if the number of the one or more classes determined by the classification of the plurality of first images does not reach a set number, generates a new class based on the one or more classes, and generates a second image different from the first image for the new class based on at least one of the first image and information related to the first image.

3. The image generating device according to claim 1 or 2, wherein the processor determines the one or more classes based on one or more scenes obtained by analyzing the plurality of first images.

4. The image generating device according to claim 1 or 2, wherein the processor determines the one or more classes based on one or more predetermined scenes.

5. The image generating device of claim 1, wherein the processor determines, for each of the one or more classes, whether at least one of the first images in the class is a first image that satisfies the criteria.

6. The image generating device of claim 1, wherein the processor, for a target class among the one or more classes that includes a first image that satisfies the criteria, determines a candidate image to be placed in a placement area in the placement image from among the first images that satisfy the criteria, and for a non-target class, determines a candidate image to be placed in the placement area from among the generated second images.

7. The image generating device of claim 2, wherein the processor determines, for each of the one or more classes, a candidate image to be placed in a placement area in the placement image from among the first images included in the class, and, for the new class, determines the candidate image to be placed in the placement area from among the generated second images.

8. The image generating device of claim 1, wherein, in the process of generating the second image, when the class among the one or more classes that includes the first image that satisfies the criteria is set as a target class, the processor generates the second image based on at least one of the first image in the target class and information related to the first image in the target class.

9. The image generating device of claim 1, wherein the processor, in the process of generating the second image, generates the second image based on at least one of information related to the first image in the non-target class for which the second image is to be generated and information related to the first image in the non-target class for which the second image is to be generated.

10. An image generating device as described in claim 1 or 2, wherein the processor, in the process of generating the second image, generates the second image in which multiple common subjects are arranged within a single image area based on two or more first images that include the common subject.

11. The image generating device according to claim 1 or 2, wherein the processor provides display information to at least one of the first image and the second image, indicating whether the image is an original image or a generated image.

12. An image generating device as described in claim 1 or 2, wherein the processor, when placing the first image and the second image in the placement area of ​​the placement image, sets placement priorities for the first image and the second image, respectively.

13. An image generating device according to claim 1 or 2, wherein the processor, if the generated second image does not satisfy the conditions for being a candidate image, selects a layout corresponding to the total number of classes having the candidate image.

14. The image generating device according to claim 13, wherein the processor generates a new layout corresponding to the total number when there is no selectable layout.

15. An image generating device according to claim 1 or 2, wherein the processor, in the process of generating the second image, generates the second image based on sound information contained in the first image.

16. An image generating device according to claim 1 or 2, wherein the processor, in the process of generating the second image, generates the second image based on information common to two or more of the first images.

17. The image generating device of claim 16, wherein the processor generates the second image by transforming the first image based on the common information.

18. An image generation device as described in claim 1 or 2, wherein in the process of generating the second image, the processor generates the second image using a generative model that has been trained to output the second image by inputting the first image.

19. An image generation method in which a processor performs the following steps: classifying a plurality of first images into one or more classes; and, when a class among the one or more classes that includes the first image that satisfies a criterion is defined as a target class, generating a second image different from the first image for a non-target class among the one or more classes that does not include the first image that satisfies the criterion, based on at least one of the first images and information related to the first images.

20. An image generation method in which a processor performs the following processes: classifying a plurality of first images into one or more classes; if the number of the one or more classes determined by classifying the plurality of first images is less than a set number, generating a new class based on the one or more classes; and generating a second image, different from the first image, for the new class based on at least one of the first images and information related to the first images.

21. A program for causing a computer to execute each process included in the image generating method according to claim 19 or 20.

22. A computer-readable recording medium having recorded thereon a program for causing a computer to execute each process included in the image generating method according to claim 19 or 20.

Citation Information

Patent Citations

  • Image processing method, image processing device and program

    JP2018097483A

  • Image processing apparatus, image processing method, and program

    JP2023113444A

  • Intelligent identification of replacement regions for mixing and replacing of persons in group portraits

    US20200151860A1