Image generation apparatus, image generation method, program, and recording medium
Patent Information
- Application Number
- PCT/JP2026/005251
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-28
- Filing Date
- 2026-02-13
- Publication Date
- 2026-09-03
Smart Images

Figure JP2026005251_03092026_PF_FP_ABST
Abstract
Description
Image generation apparatus, image generation method, program, and recording medium
[0001] One embodiment of the present invention relates to an image generation apparatus, an image generation method, a program, and a recording medium that enable generation support accurately reflecting a user's intention when generating content based on an image.
[0002] Various services and technologies have been proposed and operated for organizing images provided by a user into appropriate content in a form desired by the user. Specifically, this includes services that arrange images captured at various events such as travel or school events, or images captured as growth records of children, on a predetermined background image or frame, appropriately decorate them, and finish them as an electronic album. Alternatively, technologies have also been proposed for automatically shaping various images and descriptions associated with academic research or diagnosis in medical institutions into papers or medical data.
[0003] As described above, as a conventional technology supporting the generation of various content using images, for example, the technology described in Patent Document 1 can be cited. Patent Document 1 discloses technologies such as an image processing apparatus and an image processing method that can obtain a result that more reflects the user's intention.
[0004] This technology comprises: image feature information detection means that analyzes an image, detects predetermined features in the image, and creates image feature information; ranking means that ranks the image feature information; image processing condition determination means that determines image processing conditions based on the highest-ranked image feature information in the ranking; image processing means that performs image processing on the image using the image processing conditions; image display means that displays the image after image processing by the image processing means; input means that receives an instruction regarding a change to the image processing conditions from the outside; and image processing condition changing means that changes the image processing conditions based on the instruction regarding the change to the image processing conditions. When the image processing conditions are changed, the image processing means performs image processing on the image using the changed image processing conditions, and the image display means displays the image after image processing performed based on the changed image processing conditions. The present invention relates to an image processing apparatus characterized by the above.
[0005] Furthermore, Patent Document 2 discloses a program that updates text generated in relation to an image based on another image, and technology relating to an electronic device that executes the program.
[0006] This technology relates to a program that causes a computer to perform an acquisition process to acquire an image, a generation process to generate text related to the image based on the image, and an update process to update the text related to the first image generated by the generation process by recognizing a second image acquired after the first image.
[0007] Japanese Patent Publication No. 2010-9420 Japanese Patent Publication No. 2020-53064
[0008] However, none of the conventional technologies provide accurate and easily understandable information about the impression and content of the content to be generated by the system. As a result, there was a possibility that content generation would proceed with discrepancies or inconsistencies between the user's intentions and the generated content. Furthermore, no function is provided to appropriately correct the generated content based on user feedback or reactions regarding such discrepancies or inconsistencies. On the other hand, when generating images or content containing images using so-called generation AI, even if appropriate prompts are entered to make the above corrections, the results often do not meet the user's intentions.
[0009] Therefore, one embodiment of the present invention has been made in view of the above circumstances, and aims to provide an image generation device, an image generation method, a program, and a recording medium that enable generation support that accurately reflects the user's intent when generating image-based content.
[0010] The above objective is achieved by an image generation device described in any of the following [1] to [9]. [1] An image generation device comprising a processor that generates an image, wherein the processor receives input information for image generation, generates a first image based on the information, generates text information relating to a completed second image based on the first image and at least one of the information, and presents the text information to the user.
[0011] [2] The image generating apparatus according to [1], wherein the second image is identical to or related to the first image.
[0012] [3] The image generating apparatus according to either [1] or [2], wherein the processor modifies the first image based on the modified text information when it receives a user's modification of the text information.
[0013] [4] An image generating apparatus according to any one of [1] to [3], wherein the processor, when it receives a user modification to the first image, regenerates text relating to the second image based on the information and at least one of the modified first images.
[0014] [5] An image generation device according to any one of [1] to [4], wherein the processor generates text information with respect to a second image in which one or more images, which are information for image generation, are arranged based on pre-held layout information.
[0015] [6] An image generating apparatus according to any one of [1] to [5], wherein the processor generates text with respect to at least one of the entire electronic album as a second image and the images constituting the electronic album.
[0016] [7] An image generation apparatus according to any one of [1] to [6], wherein the processor generates multiple pieces of text information relating to the second image.
[0017] [8] An image generating apparatus according to any one of [1] to [7], wherein the processor generates two or more texts with different focus events for the information as text.
[0018] [9] An image generating apparatus according to any one of [1] to [8], wherein the processor generates multiple patterns of text as text, each representing a single event of interest with respect to the information in a different way.
[0019] Furthermore, the above objective can also be achieved by the image generation method described in
[10] below.
[10] An image generation method for generating an image using a processor, comprising: a step of receiving information for image generation input using the processor; a step of generating a first image based on the information using the processor; a step of generating text information relating to a completed second image based on at least one of the first image and the information using the processor; and a step of presenting the text information to a user using the processor.
[0020] Furthermore, the above objectives can also be achieved by the program described in
[11] below.
[11] A program for causing a computer to perform each step included in the image generation method described in
[10] .
[0021] Furthermore, the above objectives can also be achieved by the recording medium described in
[12] below.
[12] A computer-readable recording medium on which a program is recorded causing a computer to perform each step of the image generation method described in
[10] .
[0022] According to one embodiment of the present invention, an image generation apparatus, an image generation method, a program, and a recording medium are provided that enable the generation of image-based content with high accuracy in reflecting the user's intent.
[0023] This figure shows an example system configuration including the image generation device according to this embodiment. This figure shows an example hardware configuration of the image generation device according to this embodiment. This figure shows the functional part of the image generation device according to this embodiment. This figure shows an example of an input image according to this embodiment. This figure shows an example of an input image table according to this embodiment. This figure shows an example of a content table according to this embodiment. This figure shows an example of a text information table in this embodiment. This figure shows an example of the flow of the image generation method according to this embodiment. This figure shows an example of a screen flow of the image generation method according to this embodiment. This figure shows an example of a screen according to this embodiment. This figure shows an example of a screen according to this embodiment.
[0024] The following describes specific embodiments of the present invention. For convenience of explanation, the following descriptions may sometimes be based on the perspective of a GUI (Graphical User Interface). Furthermore, since the fundamental data processing technologies for realizing the present invention (communication / transmission technologies, data acquisition technologies, data recording technologies, data processing / analysis technologies, machine learning technologies, image processing technologies, and visualization technologies, etc.) are known technologies, their descriptions will be omitted.
[0025] Furthermore, in this specification, the concept of "device" includes not only a single device that performs a specific function, but also a combination of multiple devices that exist independently and in a distributed manner while cooperating (linking) to perform a specific function.
[0026] Furthermore, in this invention, "user" refers to the user of the image generation device 10 of the present invention, and specifically refers to a person who, for example, uses the functions of the image generation device of the present invention to confirm content based on input images (e.g., electronic albums, academic papers, etc.) and text information generated regarding said content (e.g., natural language representation of the results of interpreting electronic albums and their constituent images), and takes appropriate action based on the results (e.g., modifying text information or content to correspond to the intent of content creation).
[0027] Furthermore, in this specification, "person" means an entity that performs a specific action, and includes individuals, groups, corporations and other legal entities, and organizations, and may also include computers and devices that constitute artificial intelligence (AI). Artificial intelligence (AI) realizes intelligent functions such as reasoning, prediction, and judgment using hardware and software resources. The algorithm of artificial intelligence is arbitrary and includes, for example, expert systems, case-based reasoning (CBR), convolutional neural networks (CNN), deep neural networks (DNN), Bayesian networks, or inclusion architectures.
[0028] <<About one embodiment of the present invention>> [Configuration of the image generation system] In one embodiment of the present invention (hereinafter, this embodiment), an example is given in which an image generation system 5 is configured by an image generation device 10, a user terminal 30, a shooting device 40, and an external server 50 connected to a network 1 as shown in Figure 1. The user terminal 30 is an information processing device that uploads images taken by the user as input images to the image generation device 10, and allows the user to view and modify content generated by the generation AI for the input images, as well as text information interpreting the content, as needed.
[0029] Furthermore, if the user terminal 30 itself is configured to function as the image generation device 10, the above upload is unnecessary, and the user terminal 30 can perform each process in the image generation device 10 itself. The same applies to the shooting device 40; if the shooting device 40 itself is configured to function as the image generation device 10, the shooting device 40 may perform each process in the image generation device 10.
[0030] Furthermore, the user terminal 30 and the shooting device 40 also capture the above-mentioned input images and input images from the UI (User Interface), which serve as the basis for generating content and text information by the generating AI, and upload them to the image generation device 10.
[0031] Furthermore, the user may perform operations such as inputting the input image and modifying the text information using the UI (User Interface) of the image generation device 10, without using the user terminal 30 or the imaging device 40. In this case, the UI of the image generation device 10 will read the data of such input images from the recording medium presented by the user. Alternatively, the user in this case may directly perform operations to modify the text information using the UI, such as a keyboard, mouse, or touch panel.
[0032] The image generation device 10 stores the input image G1 uploaded from the user terminal 30 or the shooting device 40 in the input image table 211 (described later in Figures 3 and 5) and prepares for content generation by the generation AI 221. The image generation device 10 also stores the temporary content G2 (first image) generated by the generation AI 221 from the input image G1 held in the input image table 211 in the "temporary content" column of the content table 222 (described later in Figures 3 and 6) and prepares for user modification.
[0033] If the user modifies the temporary content G2, the modified content, i.e., the modified content G3, is stored in the "Modified Content" column of the content table 222. If the temporary content G2 is not modified by the user and is accepted as is, the temporary content is stored in the "Completed Content" column of the content table 222 as the completed content G50 (second image). Furthermore, if the temporary content G2 is modified by the user to become the modified content G3, and the modified content G3 is accepted by the user, the modified content G3 is stored in the "Completed Content" column of the content table 222 as the completed content G50.
[0034] Furthermore, the image generation device 10 stores the text information Tx generated by the generation AI 221 from the above content in the "Text Information" column of the text information table 223 (described later in Figures 3 and 7) in preparation for user modification. Subsequently, if the user modifies the text information Tx, the modified text information, i.e., the modified text information Tm, is stored in the "Modified Text Information" column of the text information table 223.
[0035] The image generation device 10 is an information processing device that uses a generation AI 221 (see Figure 3), which is either provided by the device itself or provided by an external server 50, to perform the following processes: generating and outputting temporary content G2 based on the input image G1; generating and outputting text information Tx interpreting the temporary content G2; generating modified content G3 after user modification operations on the temporary content G2 and text information Tx (for example, this may include concepts such as changing, deleting, or adding images or parts that constitute the temporary content G2, or modifying, adding, or deleting text in the text information Tx; the same applies hereinafter); and generating modified content G3 based on the modified text information Tm and the temporary content G2 (or input image G1). Such processes are performed autonomously by the image generation device 10 in accordance with the image generation method of the present invention, or in response to user instructions received through a predetermined UI.
[0036] Furthermore, as described above, the external server 50 is a server device that provides the function of generation AI 221 to the image generation device 10 via the network 1. Of course, if the image generation device 10 is configured in a way that does not require the provision of the function of generation AI 221, the external server 50 does not need to be included in the image generation system 5.
[0037] In addition to the configuration in which the image generation device 10, user terminal 30, and imaging device 40 are connected via a network 1, the internal bus wiring of the image generation device 10 may be directly connected to the interfaces of the user terminal 30 and imaging device 40.
[0038] [Example Configuration of Image Generation Device] Next, an example configuration of the image generation device 10 according to this embodiment will be described with reference to Figures 2 and 3. Specifically, the image generation device 10 is composed of a server device, a PC (Personal Computer), a smartphone, a tablet terminal, a digital camera, or a notebook PC, etc. Note that the image generation device 10 is not limited to a computer owned by the user or accessible via network 1. For example, it may be composed of a terminal that is not owned by the user but is available when visiting a store or facility, such as a store-installed terminal. In the following, we will explain using the case in which the image generation device 10 is composed of a computer owned by the user, specifically a PC, as an example.
[0039] As shown in Figure 2, the computer comprising the image generation device 10 includes a processor 11, an auxiliary storage device 12, a main storage device 13, an input device 14, an output device 15, and a communication device 16.
[0040] The processor 11 is composed of, for example, a CPU (Central Processing Unit), an MPU (Micro-Processing Unit), an MCU (Micro Controller Unit), a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), a TPU (Tensor Processing Unit), or an ASIC (Application Specific Integrated Circuit).
[0041] The auxiliary storage device 12 is composed of, for example, flash memory, HDD (Hard Disk Drive), SSD (Solid State Drive), FD (Flexible Disk), MO disk (Magneto-Optical Disk), CD (Compact Disk), DVD (Digital Versatile Disk), SD card (Secure Digital card), or USB memory (Universal Serial Bus memory).
[0042] Such auxiliary storage devices 12 may be built into the computer main body that constitutes the image generation device 10, or they may be attached to the computer main body as an external device. Alternatively, the auxiliary storage device 12 may be configured as a NAS (Network Attached Storage) or the like. Furthermore, the auxiliary storage device 12 may be an external device that can communicate with one of the computers that constitutes the image generation device 10 via a communication network, such as an online storage or database server.
[0043] Further, the auxiliary storage device 12 stores a program 121 including an operating system (OS) and applications for executing processing such as generation and modification of content and text information. When the program 121 is read and executed by the processor 11, the computer constituting the image generation apparatus 10 implements each function of the reception unit 21, the generation unit 22, and the output unit 23 shown in FIG. 3. Accordingly, through this implementation, a series of processing related to generation, modification, output and the like of content and text information based on an input image group is executed.
[0044] The main storage device 13 is constituted by a semiconductor memory such as a ROM (Read Only Memory) and a RAM (Random Access Memory), for example. The processor 11 loads the program 121 onto the main storage device 13 and executes the program herein.
[0045] The input device 14 is a device that receives input operations performed by an administrator of the image generation apparatus 10 or the image generation system 5 or the aforementioned user, and is constituted by, for example, a keyboard, a mouse, a touch panel, a camera unit, or the like. Further, the input device 14 may include an imaging device implemented in a digital camera, a microphone for sound collection, and the like. The output device 15 is constituted by, for example, a display, a speaker, or the like.
[0046] The communication device 16 may be preferably constituted by, for example, a network interface card, a communication interface board, or the like. The computer constituting the image generation apparatus 10 can communicate with other devices connected to a network 1 such as the Internet and a mobile communication line via the communication device 16.
[0047] As shown in FIG. 3, the image generation apparatus 10 includes the reception unit 21, the generation unit 22, and the output unit 23. These functional units are implemented by the processor 11 of the image generation apparatus 10 executing the aforementioned program 121 and cooperating with various types of hardware in the image generation apparatus 10 and other devices (the user terminal 30, the imaging apparatus 40, and the external server 50). Further, the generation AI 221 in the generation unit 22 may be called from the external server 50 for execution.
[0048] [Regarding Input Image] In the present embodiment, the image generation apparatus 10 of the present invention acquires an input image G1, for example as shown in FIG. 4, via the screen G0 of a user terminal 30 or an imaging apparatus 40, and provides the input image G1 to generation AI 221, thereby generating provisional content G2 and text information Tx based on the provisional content G2.
[0049] The input image G1 illustrated in FIG. 4 is composed of individual images G1A to G1F (of course, the number of such individual images is only an example). Note that when giving a common description to these individual images G1A to G1F, the description will be given with the image G1A as a representative. The individual image G1A is an image in which, for example, a landscape, a building, or the like of the shooting location serves as a background image Pb, and the image includes one or a plurality of subject images Ps captured on the front side of the background image Pb.
[0050] The individual image G1A illustrated in FIG. 4 is an image of a family happily preparing for a trip. In this case, the image G1A can be described as an image with a composition in which the background image Pb is a living room of a home, and the subject image Ps is the aforementioned family. For other individual images G1B to G1E, they are images with a composition in which the background image Pb is a magnificent scenery of Hokkaido, and the subject image Ps is the aforementioned family.
[0051] Note that the image generation apparatus 10 generates provisional content G2 such as an electronic album, for example, by inputting the input image G1 acquired from the user terminal 30, the imaging apparatus 40, or the like into the generation AI 221. In this case, the generation AI 221, for example, estimates a common concept (interpretation result) for the individual images G1A to G1F constituting the input image G1, estimates the subject, shooting location, shooting time, and the like in the individual images G1A to G1F, identifies content corresponding to these estimation results, and generates content by setting the appropriately processed and edited input image G1 in a template (layout information) held in advance for the content.
[0052] Therefore, the learning model in the generating AI221 is assumed to perform machine learning using, for example, a common concept for each input image G1, the subject, location, and time of shooting for each input image G1, and suitable content as training data. For example, if the common concept for each input image G1 is "family, Hokkaido," and the subject, location, and time of shooting for each input image G1 are "subject: family of three including an elementary school child, location: home, time of shooting: March 8th," "subject: family of three including an elementary school child, location: on a road in the suburbs of Kushiro City, time of shooting: March 9th," "subject: family of four including an elementary school child and a kindergarten child, location: near the foot of Mt. Akan Fuji, time of shooting: March 9th," etc., then for example, "an electronic album for family trips including outdoor activities" is identified as suitable content, and the input image G1 is set in the template (layout information) of the "electronic album for family trips including outdoor activities" to generate an electronic album (temporary content G2). Furthermore, information such as the shooting location and shooting date can be obtained from appropriate metadata, such as EXIF tag information, of each input image G1.
[0053] Furthermore, the generating AI 221 interprets the provisional content G2 and generates text information Tx that represents the provisional content G2. Therefore, the generating AI 221 is also equipped with a learning model as a caption generating AI. This learning model is a learning model that has undergone machine learning using various content and captions that accurately represent that content as training data. In the case of the input image G1 in Figure 4, the generating AI 221 of the image generation device 10 can generate text information Tx such as, for example, "A family trip during spring break. A family of four drove around Hokkaido. We toured eastern Hokkaido for 4 nights and 5 days."
[0054] The user views this text information Tx "A family trip during spring break. A family of four drove around Hokkaido. We toured eastern Hokkaido for 4 nights and 5 days." on the user terminal 30, and judges whether it meets various conditions such as the purpose and intent of taking the input image G1, or expressions that suggest content (e.g., an electronic album) using the input image G1 as material, and makes corrections such as changing, adding, or deleting the wording as necessary.
[0055] For example, if the event the user photographed was indeed a "family trip," but the purpose of the trip was to "celebrate their eldest daughter's elementary school graduation," the user might change the above text information Tx "A family trip during spring break. The four of us drove around Hokkaido. We toured eastern Hokkaido for 4 nights and 5 days." to "A trip to celebrate our eldest daughter's elementary school graduation." In this case, "A trip to celebrate our eldest daughter's elementary school graduation" would become the corrected text information Tm.
[0056] In this invention, "image" is composed of multiple pixels, represented by the grayscale values of each of the multiple pixels, and includes at least one subject image Ps. Furthermore, digital image data (hereinafter referred to as "image data") that defines an image at a set resolution is generated by compressing data in which the grayscale values of each pixel are recorded using a predetermined compression method. Examples of image data include lossy compressed image data such as JPEG (Joint Photographic Experts Group) format, and lossless compressed image data such as GIF (Graphics Interchange Format) or PNG (Portable Network Graphics) format.
[0057] [Functional Units in the Image Generation Device] Next, we will explain the processing performed by each of the functional units, namely the receiving unit 21, the generation unit 22, and the output unit 23, that are provided in the processor 11 of the image generation device 10.
[0058] (Reception Unit) The reception unit 21's processing includes receiving the input image G1 from the user terminal 30 or the imaging device 40 via the input device 14 or the communication device 16, and storing it in, for example, the auxiliary storage device 12. Furthermore, the storage process in the reception unit 21 includes storing the input image G1 received as described above in the input image table 211.
[0059] Figure 5 shows an example of the configuration of the input image table 211 in this embodiment. The input image table 211 is a collection of records containing the values of "User ID," "Date and Time," and "Image Data." The User ID is the identification information of the user who input the input image G1, and may also be the identification information of the user terminal 30. Furthermore, if the image generation device 10 is operated by a single user, this User ID is not required. The Date and Time is the date and time when the input image G1 was input or when it was taken (this can be obtained from metadata such as EXIF tag information held in the data of the target input image G1).
[0060] The method for receiving the input image G1 at the reception unit 21 is not particularly limited, but includes methods of acquiring the image by reading a photographic print or playback video of an image taken by the user terminal 30 or the shooting device 40 using the camera unit. The reception unit 21 may also acquire the input image G1 by downloading image data from an external device or a web server via the network 1. The input image G1 may be a still image or a video. The input image G1 may also include images generated by the generation unit 22 by adding text information Tx, modified content G3, and other content to the generation AI 221.
[0061] (Generation Unit) The processing of the generation unit 22 includes the process of generating temporary content G2, such as an electronic album, by inputting the input image G1 obtained by the processing of the reception unit 21 into the generation AI 221. The temporary content G2 generated by the generation unit 22 and the generation AI 221 is stored in the content table 222. As shown in Figure 6, this content table 222 is a collection of records consisting of "Content ID", "User ID", "Date and Time", "Temporary Content", "Modified Content", and "Completed Content".
[0062] Of these, the Content ID is an ID that uniquely identifies the content (at least one of the temporary content G2, the modified content G3, and the completed content G50; the same applies hereinafter). The User ID is the ID of the photographer (provider) of the input image G1 used to generate the content. The Date and Time is the date and time the input image G1 was taken or the date and time the content was generated. Temporary content G2 is the content generated based on the input image G1. Modified content G3 is the content after the user has made modifications to temporary content G2, or the content regenerated by the generation AI 221 in accordance with the modified text information Tm, which is the modified text information, after the user has made modifications to the text information Tx generated by the generation AI 221 based on temporary content G2. Furthermore, completed content G50 is the content that becomes completed for the user when the user has not made any modifications to the temporary content G2 or modified content G3, i.e., when the user has accepted that content.
[0063] As described above, the processing of the generation unit 22 includes processing to accept user modification operations on the temporary content G2 or the modified content G3, or user modification operations on the text information Tx generated based on the temporary content G2 or the modified content G3, and processing to store the modified content G3 or modified text information Tm after the user modification operations.
[0064] Furthermore, the processing of the generation unit 22 may also include the generation and storage of the modified content G3 and new modified content based on the modified text information Tm (i.e., the regeneration and re-storage of the modified content G3). Therefore, the generation AI 221 of the generation unit 22 has a learning model that performs machine learning using training data, which consists of a pair of text specified for a certain content (i.e., the sentence indicated by the modified text information Tm) and content obtained by modifying the content according to the above text (or newly generated based on the content) (i.e., new modified content G3).
[0065] Furthermore, the processing of the generation unit 22 includes the process of generating keywords and sentences, i.e., text information Tx, when interpreting the provisional content G2 and the modified content G3. The generation AI 221 that generates text information Tx has the function of a caption generation AI that outputs sentences and keywords that appropriately represent the content. Therefore, the generation AI 221 has a learning model that has been trained using machine learning with various content and sets of sentences and keywords that appropriately represent the content as training data.
[0066] The text information Tx generated by the generation unit 22 using the generation AI 221 is stored in the text information table 223. As shown in Figure 7, this text information table 223 is a collection of records consisting of "text ID," "content ID," "text information," and "modified text information." The text ID is an ID that uniquely identifies the text information. The content ID is an ID that uniquely identifies the content (temporary content G2 or modified content G3) used to generate the text information. The text information is the text information Tx generated based on the temporary content G2 or modified content G3. The modified text information is the modified text information Tm generated by the user making modifications to the text information Tx.
[0067] Furthermore, the processing of the generation unit 22 may also include the process of generating multiple copies of the above-mentioned text information Tx and modified text information Tm with respect to the target provisional content G2 and modified content G3. This processing may be performed using at least one of two methods. The first method is to generate multiple patterns of text information Tx and modified text information Tm with different representations for a single event of interest in the provisional content G2 and modified content G3. The second method is to generate text information Tx and modified text information Tm for each of the multiple events of interest in the provisional content G2 and modified content G3.
[0068] The first method involves generating sentences with different expressions, for example, when focusing on the event of "a family trip to Hokkaido." Pattern 1: "A family trip during spring break. The four of us drove around Hokkaido. We toured eastern Hokkaido in 4 nights and 5 days." Pattern 2: "A road trip to Hokkaido with the four of us during spring break♪ We toured eastern Hokkaido and made lots of memories★" Pattern 3: "A road trip to Hokkaido with the four of us during spring break. We conquered eastern Hokkaido in 4 nights and 5 days." Pattern 4: "A journey for the four of us where we felt the arrival of spring. The memories of our 4 nights and 5 days exploring the northern land will last forever."
[0069] Regarding the second method, for example, Pattern 1: Focusing on the event "Family trip to Hokkaido," it generates the sentence "A family trip during spring break. The four of us drove around Hokkaido. We toured eastern Hokkaido for 4 nights and 5 days." Pattern 2: Focusing on the event "Nature of Hokkaido," it generates the sentence "The great nature of Hokkaido in spring." Pattern 3: Focusing on the event "Outdoor travel," it generates the sentence "Early spring trekking tour." Pattern 4: Focusing on the event "Child's school graduation," it generates the sentence "A trip to commemorate our eldest daughter's elementary school graduation."
[0070] In order to execute the above method, the generation unit 22 performs two processes in the generation AI 221 based on the content or images contained in the content: 1) estimation of the events represented by the content or images, and 2) generation of text describing the content or images for each of the events that is most likely among the estimated events, or for each of all events.
[0071] Therefore, as a function to perform the processing in 1) above, the generating AI 221 has, as already stated, functions such as estimating a common concept (interpretation result) for the input image G1 (which may also include each of the individual images G1A) contained in the content, and estimating the subject, shooting location, shooting time, etc. for each of the individual images G1A.
[0072] The Generative AI 221 combines these values—the subject (e.g., age, gender, family relationships between subjects estimated from their ages, and school age based on the subject's age if they are children), the shooting location (e.g., Hokkaido, eastern Hokkaido), and the shooting period (e.g., mid-March)—to estimate that it is a "family of four including elementary school children," a "trip during the elementary school spring break period," and a "trip destination in Hokkaido, eastern Hokkaido," thus deriving the event "family trip to Hokkaido" (Generative AI 221 has already learned the process and results of this derivation).
[0073] Furthermore, the Generating AI 221 has a function to perform the process described in 2) above, for example, by inputting information about the event derived in 1), it generates text that explains the event. Therefore, the Generating AI 221 has already performed machine learning using pairs of such events and appropriate texts that explain the event as training data. Thus, the Generating AI 221 takes information about the event with the highest probability among the events estimated in 1), or information about each of the events, as input, and generates text that explains the content or image. Regarding the determination of probability, known determination algorithms and tools will be used.
[0074] (Output Unit) The output unit 23 processes the temporary content G2, modified content G3, completed content G50, text information Tx, and modified text information Tm generated by the generation unit 22, and displays them on an output device 15 such as a display, a user terminal 30, or a shooting device 40. For this display, the output unit 23 retrieves data from the content table 222, the text information table 223, or the main memory 13 as appropriate, following the processing in the generation unit 22, or in response to a request from the user terminal 30 or the arrival of a predetermined timing, and outputs it.
[0075] The output methods for the temporary content G2, revised content G3, completed content G50, text information Tx, and revised text information Tm are not particularly limited, but include, for example, the output unit 23 displaying the temporary content G2, revised content G3, completed content G50, text information Tx, and revised text information Tm on the display or monitor of the output device 15, user terminal 30, or shooting device 40; printing the temporary content G2, revised content G3, completed content G50, text information Tx, and revised text information Tm; transmitting the temporary content G2, revised content G3, completed content G50, text information Tx, and revised text information Tm to other users; and providing the temporary content G2, revised content G3, completed content G50, text information Tx, and revised text information Tm as commercial products. Products may include media consisting of one or more pages or cards containing images, such as albums, photobooks, postcards, message cards, digital albums, and bromide prints.
[0076] [Example of Image Generation Method Flow] Next, as an example of the operation of the image generation device 10 in this embodiment, an image processing flow using the device will be described. The image processing flow described below uses the image generation method of the present invention. In other words, each step in the image processing flow described below corresponds to a component of the image generation method of the present invention. Note that the following flow is merely an example, and some steps in the flow may be deleted, new steps added to the flow, or the execution order of two steps in the flow may be changed without departing from the spirit of this embodiment.
[0077] Each step in the image generation flow according to this embodiment is performed by the processor 11 of the image generation device 10 in the order shown in Figure 8. In other words, in each step of the image generation flow, the processor 11 executes the data processing defined in the application program for image generation that corresponds to each step in Figure 8.
[0078] To explain in more detail, in the image generation flow according to this embodiment, first, the reception unit 21 receives the input image G1 to be processed from, for example, the user terminal 30 (S1). During this process, the reception unit 21 delivers a reception screen G10, which includes U1 as shown in Figure 9, to the user terminal 30. The reception unit 21 also acquires the input image G1 via the file specification field G11 of U1 on the reception screen G10. The reception unit 21 also stores the acquired input image G1 in the input image table 211.
[0079] Next, the generation unit 22 inputs the input image G1 obtained in S1 into the generation AI 221 to generate temporary content G2 based on the input image G1 (S2). The generation unit 22 also links the temporary content G2 generated here with the user's ID, the date and time the input image G1 was taken, etc., and stores it in the "Temporary Content" column of the content table 222. When storing it, a content ID that identifies the temporary content G2 is assigned to the target record.
[0080] An example of the temporary content G2 generated here is illustrated in Figure 10. In the example in Figure 10, a total of six frames G201, each containing an input image G1, are arranged on a wall, with a model of a car representing a drive and a deer figurine representing Hokkaido placed in front of them. This image constitutes either one page of a multi-page electronic album, or the single page if the electronic album consists of only one page.
[0081] Furthermore, the generation unit 22 inputs the temporary content G2 obtained in S2 into the generation AI 221 to generate text information Tx, which is the caption for the temporary content G2 (S3). The generation unit 22 also links the generated text information Tx with the content ID and stores it in the "Text Information" column of the text information table 223. When storing it, a text ID that identifies the text information is assigned to the target record. An example of the generated text information Tx is shown in screen G100 of Figure 10.
[0082] Furthermore, the output unit 23 distributes the temporary content G2 and text information Tx obtained through the processing up to this point to, for example, the user's user terminal 30 for display (S4). The display status is illustrated in screen G100 of Figure 10. In addition to the temporary content G2 and text information Tx, screen G100 is equipped with a content modification button G101 and a text information modification button G102. Therefore, if a user viewing screen G100 intends to modify the temporary content G2 or text information Tx, they can click the content modification button G101 or the text information modification button G102 to perform the modification operation.
[0083] If the user has no objections to the temporary content G2 or text information Tx and has no intention to make any modifications, they click the OK button G103. At this time, the reception unit 21, upon receiving the click, recognizes the temporary content G2 displayed on screen G100 as the completed content G50 (S5:N), renames the data of the temporary content G2 as the data of the completed content G50, stores it in the "Completed Content" column of the content table 222 (S10), and terminates this flow.
[0084] On the other hand, if the user feels something is wrong with the temporary content G2 and intends to make corrections, they click the content correction button G101. At this time, the reception unit 21 receives the click and recognizes the user's intention to correct the temporary content G2 displayed on screen G100 (S5: Y, S6: Content), and for example, distributes and displays a UI (User Interface) for content correction to the user terminal 30.
[0085] The example screen G150 shown in Figure 11 is a screen that displays the corrected temporary content G2, i.e., the corrected content G3, after receiving a correction operation by the user on screen G100. In this screen G150, the generation unit 22 displays the correction menu G1011 in response to the click of the content correction button G101 and indicates that the user has specified "change background image". The generation unit 22 also displays the background image selection menu G1012 in response to this "change background image" specification operation and indicates that the user has specified "luxury".
[0086] Furthermore, the generation unit 22 has pre-configured or recallable image processing algorithms and materials for each modification event of the modification menu G1011 (e.g., changing the background image, replacing the image, changing the color tone, etc.) and for each option within that modification event (e.g., lively, dazzling, sporty, luxurious, etc.). In addition, the image processing algorithms may include learning models that perform image processing according to the algorithm. Naturally, such learning models have been trained using training data related to the image processing in question.
[0087] The generation unit 22 generates the modified content G3 by changing the background image in the temporary content G2 in response to the modification operations received from the modification menu G1011 and the selection menu G1012, and displays this on the screen G150 (S7). The generation unit 22 also stores the modified content G3 in the "Modified Content" column of the content table 222 and transitions the process to S3.
[0088] Furthermore, if the user feels something is wrong with the text information Tx and intends to correct it, they click the text information correction button G102. At this time, the reception unit 21 receives the click and recognizes the intention to correct the text information Tx displayed on the screen G100 (S5: Y, S6: text information), and for example, activates the text frame G105 for text information correction (which is also the display frame for text information Tx) to display a cursor, and distributes it to the user terminal 30 for display.
[0089] The example of screen G155 shown in Figure 12 is a screen that displays the modified text information Tm that is in the process of being entered by the user, after the user has made a correction operation on screen G100. In this example of screen G155, the input has reached the point of "My eldest daughter's elementary school graduation".
[0090] The generation unit 22 displays the modified text information Tm, which has been entered by the user in the text box G105, on the screen G160 (S8). In the screen G160 illustrated in Figure 13, the modified text information Tm is displayed as the sentence "A trip to commemorate my eldest daughter's elementary school graduation". The generation unit 22 then stores the modified text information Tm in the "Modified Text Information" column of the text information table 223 and proceeds to S9.
[0091] Next, the generation unit 22 generates the modified content G3 based on the modified text information Tm obtained in S8 (S9). The modified content G3 shown in Figure 13 is an image with a composition in which a teddy bear that matches the general impression of elementary school students is placed in front of the frame G201 on which the input image G1 is set. As already mentioned, the generation of this modified content G3 based on the modified text information Tm is a process that the generation AI 221 executes using the modified text information Tm and the temporary content G2 (or input image G1) as input. The generation unit 22 then transitions the process to S4 in order to distribute and display the generated modified content G3 and the modified text information Tm to the user terminal 30.
[0092] If a user who has viewed the modified content G3 and modified text information Tm on the user terminal 30 has no objections to the modified content G3 and modified text information Tm and has no intention to modify them, they click the OK button G103. At this time, the reception unit 21 receives the click and recognizes the modified content G3 displayed on screen G150 or screen 160 as the completed content G50 (S5:N), renames the data of the modified content G3 as the data of the completed content G50, stores it in the "completed content" column of the content table 222 (S10), and terminates this flow.
[0093] [Generating Multiple Text Information] Next, the case in which the generation unit 22 generates multiple text information Tx will be explained based on Figures 14 to 17. The generation of multiple text information Tx in this way is carried out in S3 and S8 in the flow of Figure 8. In this case, the user will view these multiple text information Tx on the user terminal 30 and select the one that they think is suitable for the purpose and intent of the shooting or content creation. Therefore, the text information Tx selected here will be the subject of subsequent processing.
[0094] The pattern for generating multiple text information Tx, as shown in the flowchart of Figure 14, generates text information Tx and modified text information Tm for each of the multiple viewpoints (points of interest) in the provisional content G2 and the modified content G3. In this case, the generation unit 22 inputs the provisional content G2, the modified content G3, and the images contained therein (mainly the input image G1) to the generation AI 221, and performs estimation of the viewpoints (events) shown in the provisional content G2 and the modified content G3, and generates text that explains the provisional content G2 and the modified content G3 for each of the viewpoints estimated here whose plausibility is above a certain standard, or for each of all viewpoints (S20).
[0095] Furthermore, the generation unit 22 distributes and displays the multiple sentences, i.e., text information Tx, generated in S20 to the user terminal 30 (S21). In the example of screen G170 in Figure 15, the text information Tx "A family trip during spring break. A family of four drove around Hokkaido. We toured eastern Hokkaido for 4 nights and 5 days.", the text information Tx "The great nature of Hokkaido in spring", the text information Tx "Early spring trekking tour", and the text information Tx "A trip to commemorate my eldest daughter's elementary school graduation" are listed and displayed along with selection UI such as radio buttons G171.
[0096] Next, the generation unit 22 accepts the user's selection of text information Tx via the radio button G171 on the screen G170 (S22). In this case, the user operates the user terminal 30 and clicks the radio button G171 corresponding to the text information Tx they desire. The user then clicks the text information selection button G172 to confirm the selection result.
[0097] The generation unit 22, upon receiving a click on the text information selection button G172, obtains the selected text information Tx (in the example in Figure 15, "A trip to commemorate my eldest daughter's elementary school graduation") and generates, for example, temporary content G2 based on this (S23). From S23 onward, the process from S4 onward, following S9 in the flow of Figure 8, will be executed.
[0098] Furthermore, the operation of selecting the desired text information Tx from multiple options can also be considered as an operation of modifying the text information Tx and generating the modified text information Tm. Therefore, although the flow in Figure 14 describes the generation of temporary content G2 based on the text information Tx, it can also be considered as the generation of modified content G3 based on the modified text information Tm obtained through the above selection.
[0099] The pattern for generating multiple text information Tx, as shown in the flowchart of Figure 16, generates multiple text information Tx and modified text information Tm with different writing styles for each of the perspectives (focus events) in the provisional content G2 and modified content G3. In this case, the generation unit 22 inputs the provisional content G2, modified content G3, and the images contained therein (mainly the input image G1) to the generation AI 221, and performs estimation of the perspective (event) indicated by the provisional content G2 and modified content G3, and generates a text explaining the provisional content G2 and modified content G3 for the perspective that is most likely among the one or more perspectives estimated here (S30).
[0100] Furthermore, the generation unit 22 distributes and displays the multiple sentences, i.e., text information Tx, generated in S30 to the user terminal 30 (S31). In the example of screen G175 in Figure 17, the following text information Tx are displayed together with selection UI such as radio buttons G171: "A family trip during spring break. A family of four drove around Hokkaido. We toured eastern Hokkaido in 4 nights and 5 days.", "A road trip to Hokkaido during spring break with the family of four♪ We toured eastern Hokkaido and made lots of memories★", "A road trip to Hokkaido during spring break with the family of four. We conquered eastern Hokkaido in 4 nights and 5 days.", and "A journey for a family of four where we felt the arrival of spring. Memories of 4 nights and 5 days touring the northern land will last forever."
[0101] Next, the generation unit 22 accepts the user's selection of text information Tx via the radio button G171 on the screen G175 (S32). In this case, the user operates the user terminal 30 and clicks the radio button G171 corresponding to the text information Tx they desire. The user then clicks the text information selection button G172 to confirm the selection result.
[0102] The generation unit 22, upon receiving a click on the text information selection button G172, acquires the selected text information Tx (in the example in Figure 17, "A family of four's journey as they felt the arrival of spring. Memories of a 4-night, 5-day trip around the northern land will last forever") and generates, for example, temporary content G2 based on this (S33). From S33 onwards, the process from S4 onwards, following S9 in the flow of Figure 8, will be executed.
[0103] Furthermore, the operation of selecting the desired text information Tx from multiple options can also be considered as an operation of modifying the text information Tx and generating the modified text information Tm. Therefore, although the flow in Figure 16 describes the generation of temporary content G2 based on the text information Tx, it can also be considered as generating modified content G3 based on the modified text information Tm obtained through the above selection.
[0104] Although specific embodiments of the present invention have been described above, these embodiments are merely examples given to facilitate understanding of the present invention and do not limit it. That is, the present invention can be modified or improved from the embodiments described below, without departing from its spirit. Furthermore, the present invention includes equivalents thereof. Moreover, embodiments of the present invention may include forms that combine the above embodiments with one or more of the following modifications.
[0105] (Regarding the computer constituting the image generation device) In the above embodiment, the image generation device 10 of the present invention is configured with a computer directly used by the user, such as a computer or server owned and operated by a content generation service provider or a digital camera user. However, it is not limited to this, and the image generation device 10 of the present invention may be configured with a computer that can be used indirectly by the user, for example, an external server 50. Here, the external server 50 may be, for example, an external server for cloud services, specifically an external server for ASP (Application Service Provider), SaaS (Software as a Service), PaaS (Platform as a Service), or IaaS (Infrastructure as a Service).
[0106] In this case, when the user inputs necessary information (e.g., captured images) on the user terminal 30 owned by the user, the external server 50 performs various processes (calculations) based on that information, including the generation of temporary content and text information Tx interpreted therefrom, and the calculation results are output on the user terminal 30. In other words, the functions of the external server 50 that constitute the image generation device 10 of the present invention can be used on the user terminal 30. Alternatively, the image generation device 10 may be configured with a computer or external server 50, as well as a camera 40 or smartphone (a type of user terminal 30) used by the user.
[0107] (Regarding the processor configuration) In this embodiment, each process is executed on any computer. Furthermore, any computer may execute these processes using a processor, a program, or a combination thereof. Any computer may be a general-purpose computer, a computer designed for a specific purpose, a workstation, or any other hardware element capable of executing a program.
[0108] A processor may consist of one or more hardware components, and the type of hardware is not limited. For example, a processor may consist of programmable logic devices such as a CPU (Central Processing Unit), MPU (Micro Processing Unit), FPGA (Field Programmable Gate Array), dedicated circuits for performing specific processing such as an ASIC (Application Specific Integrated Circuit), a GPU (Graphic Processing Unit), or an NPU (Neural Processing Unit).
[0109] Furthermore, the processor has various units or means that execute the various processes in this embodiment. The type of hardware may also be a combination of different types of hardware. When multiple hardware components are configured to execute one or more processes of a given processor, these multiple hardware components may reside in physically separate devices or in the same device. In any embodiment, the order of the processes performed by the processor is not limited to the order described above and may be changed as appropriate. The hardware is composed of an electrical circuit (circuitry) or the like, which is a combination of circuit elements such as semiconductor elements.
[0110] Furthermore, this embodiment may be implemented by hardware, software, firmware, microcode, or a combination thereof. The software, firmware, and microcode are composed of a program. The program may also be, for example, a group of program modules, and each of its functions may be implemented by a processor configured to perform its respective function. The program may be program code or multiple code segments stored in one or more non-temporary computer-readable media (e.g., storage media or other storage).
[0111] A program may be divided and stored on multiple non-temporary computer-readable media located on devices that are physically separated from each other. Program code, or code segments, may represent any combination of procedures, functions, subprograms, routines, subroutines, modules, software packages, classes, or instructions, data structures, or program statements. Program code, or code segments, may be connected to other code segments or hardware circuits by sending and receiving information, data, arguments, parameters, or memory contents.
[0112] 1 Network 5 Image generation system 10 Image generation device 11 Processor 12 Auxiliary storage device 121 Program 13 Main memory device 14 Input device 15 Output device 16 Communication device 21 Reception unit 211 Input image table 22 Generation unit 221 Generation AI (trained model) 222 Content table 223 Text information table 23 Output unit 30 User terminal 40 Shooting device 50 External server G1 Input image G2 Provisional content G3 Modified content Tx Text information Tm Modified text information G50 Final content
Claims
1. An image generating apparatus comprising a processor that generates an image, wherein the processor receives input information for image generation, generates a first image based on the information, generates text information relating to a completed second image based on the first image and at least one of the information, and presents the text information to the user.
2. The image generating apparatus according to claim 1, wherein the second image is identical to or related to the first image.
3. The image generation apparatus according to claim 1, wherein when the processor receives a user's request for modification of the text information, it modifies the first image based on the modified text information.
4. The image generation apparatus according to claim 1, wherein when the processor receives a user modification to the first image, it regenerates text relating to the second image based on the information and at least one of the modified first images.
5. The image generation apparatus according to claim 1, wherein the processor generates the text information with respect to the second image in which one or more images, which are information for image generation, are arranged based on layout information held in advance.
6. The image generation apparatus according to claim 1, wherein the processor generates the text with respect to at least one of the entire electronic album as the second image and the images constituting the electronic album.
7. The image generation apparatus according to claim 1, wherein the processor generates a plurality of text information relating to the second image.
8. The image generation apparatus according to claim 1, wherein the processor generates two or more texts as text, each focusing on a different event related to the information.
9. The image generation apparatus according to claim 1, wherein the processor generates, as text, multiple patterns of text that express a single event of interest with respect to the information in a manner different from each other.
10. An image generation method comprising: a step of receiving information for image generation input by a processor; a step of generating a first image based on the information by the processor; a step of generating text information relating to a completed second image based on at least one of the first image and the information by the processor; and a step of presenting the text information to a user by the processor.
11. A program for causing a computer to perform each step included in the image generation method described in claim 10.
12. A computer-readable recording medium on which a program is recorded causing a computer to perform each step included in the image generation method described in claim 10.