Image processing device, image processing method, program, and recording medium

The image processing device and method facilitate precise corrections to generative AI images by allowing users to input text and images, analyze characteristics, and adjust parameters intuitively, addressing the challenge of efficiently achieving desired image modifications.

WO2025204438A1PCT designated stage Publication Date: 2025-10-02FUJIFILM CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/006636
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-29
Filing Date
2025-02-26
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Users find it difficult to efficiently and accurately correct images generated by generative AI, as they lack understanding of the underlying algorithms and struggle to express subtle modifications affecting image atmosphere or impression.

Method used

An image processing device and method that allows users to input images and text information, analyze image characteristics, and provide intuitive interfaces for correcting image parameters through sliders and models, enabling precise adjustments based on user input.

Benefits of technology

Enables users to efficiently and accurately modify images generated by generative AI, ensuring the desired image quality and atmosphere is achieved through user-friendly correction tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025006636_02102025_PF_FP_ABST
    Figure JP2025006636_02102025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a device, a method, a program, and a recording medium that enable a user to accurately and efficiently correct an image generated by generative AI. An image processing device 10 according to one embodiment of the present invention comprises a processor 11. The processor 11 receives an image and information related to the image, and determines an item for processing pertaining to the image on the basis of at least one of the image or the information related to the image.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing device, image processing method, program, and recording medium

[0001] One embodiment of the present invention relates to an image processing device, an image processing method, a program, and a recording medium that enable a user to efficiently perform accurate corrections on an image generated by a generation AI.

[0002] The need to draw images on a computer and edit them according to a purpose exists not only in fields such as photography and graphic design, but also in a wide range of fields such as medicine and manufacturing. Therefore, techniques for freely editing the image quality, color tone, exposure, etc. of an image are already known. Patent Document 1 discloses an example of such a technique. Patent Document 1 discloses an interface between an image processing device and an imaging device, and an image processing program technique that allows users to easily operate the device, even when using a gradation conversion process that offers a high degree of freedom but whose parameters are difficult to intuitively grasp.

[0003] This technology relates to an image processing device comprising: a display means for displaying image data; a sensitivity word designation means for designating sensitivity words that sensorily express image processing effects; a target value setting means for setting target values ​​corresponding to the sensitivity words; an area designation means for designating a specific area on an image displayed on the display means; an image quality evaluation amount calculation means for calculating an image quality evaluation amount from the specific area; an image processing means for performing image processing on the image data based on predetermined image processing parameters corresponding to the sensitivity words designated by the sensitivity word designation means to generate a processed result image; and a control means for comparing the image quality evaluation amount calculated using the image quality evaluation amount calculation means from the area on the processed result image designated by the area designation means with the target value set by the target value setting means, modifying the image processing parameters based on the comparison result, and controlling the device to perform image processing again.

[0004] In addition, Patent Document 2 discloses a technology for generating a quantitative image or an enhanced image of a different type from a quantitative image or an enhanced image, even when the theoretical or mathematical formula representing the relationship between the quantitative image or the enhanced image and another type of quantitative image or enhanced image is not known.

[0005] This technology relates to a medical image diagnosis support device that includes: a training image accepting unit that accepts two or more training images and one or more gold standard images; a synthesis parameter determining unit that determines synthesis parameter values ​​such that, when pixel values ​​of corresponding pixels in the two or more training images are synthesized using the synthesis parameter values, post-synthesis pixel values ​​approach the pixel values ​​of corresponding pixels in the gold standard image; an inspection image accepting unit that accepts, as an inspection image, an image generated for a first subject, which is an image of the same type as the two or more training images; and an image synthesis executing unit that generates a synthetic image desired by a user by synthesizing pixel values ​​of corresponding pixels in the two or more inspection images using the synthesis parameter values ​​determined by the synthesis parameter determining unit, wherein the two or more training images are two or more types of images generated for the first subject or a second subject different from the first subject, and the one or more gold standard images are images generated for the second subject or a third subject different from the first subject and the second subject, which are images of a different type from the training images and the same type as the synthetic image desired by the user.

[0006] JP 2006-093757 A JP 2019-180637 A

[0007] In recent years, so-called generative AI (Artificial Intelligence) has become available in internet services and certain applications, and has become easily used in a variety of situations. In such generative AI, simply inputting text information corresponding to a desired image into the user interface will output a corresponding image.

[0008] On the other hand, there are many cases where the images generated by generative AI differ from the image the user had in mind. While users can easily obtain an image by simply entering text information, it can also be difficult to efficiently obtain the desired image. In such a situation, users are forced to repeatedly change, add, or delete the text information they enter in order to make the image generated by generative AI closer to the desired one.

[0009] On the other hand, it is difficult for humans to understand the algorithms used by the AI ​​to generate images. Therefore, it is uncertain whether the user's changes to text information, such as those mentioned above, will lead to the generation of the desired image. It is difficult for users to input and express in text subtle image modifications that may affect human emotions, such as changing the atmosphere or impression of objects in the background of an image, rather than simply changing the subject of a drawing from a male to a female.

[0010] One embodiment of the present invention has been made in consideration of the above circumstances, and aims to provide an apparatus, method, program, and recording medium that enable users to efficiently perform accurate corrections to images generated by generation AI.

[0011] The above object is achieved by an image processing device according to any one of [1] to [9] below. [1] An image processing device including a processor, wherein the processor accepts an image and information related to the image, and determines a processing item related to the image based on at least one of the image and the information related to the image. [2] The image processing device according to [1], wherein the image is an image generated based on information related to the image. [3] The image processing device according to [1], wherein the information related to the image is text generated from the image. [4] The image processing device according to [1], wherein the processor accepts information related to an object that a user wishes to draw in the image, and determines the processing item based on at least one of the image and the information related to the object. [5] The image processing device according to [4], wherein the processor accepts information related to a drawing area of ​​the object in the image, and determines the processing item using the information related to the drawing area as well. [6] The image processing device according to [1], wherein the processor decides a correction amount for the processing item based on at least one of the image and information related to the image. [7] The image processing device according to [6], wherein the processor determines the amount of correction for the processing item as information described in text. [8] The image processing device according to [6], wherein the amount of correction for the processing item represents at least one of the form or state of a concept. [9] The image processing device according to [6], wherein the processor presents to a user an interface that accepts a user's specification of the amount of correction for the processing item, and executes correction processing for the image based on the amount of correction accepted from the user.

[0012] The above object can also be achieved by an image processing method described in

[10] below:

[10] An image processing method including the steps of: receiving, by a processor, an image and information related to the image; and determining, by the processor, an item of processing related to the image based on at least one of the image and the information related to the image.

[0013] The above object can also be achieved by the program described in

[11] below:

[11] A program for causing a computer to execute each step included in the image processing method described in

[10] .

[0014] The above object can also be achieved by the recording medium described in

[12] below:

[12] A computer-readable recording medium having recorded thereon a program for causing a computer to execute each step included in the image processing method described in

[10] .

[0015] According to one embodiment of the present invention, an apparatus, method, program, and recording medium are provided that enable a user to efficiently perform accurate corrections on images generated by a generation AI.

[0016] 1 is a diagram illustrating an example of the configuration of an image processing system according to the present embodiment; FIG. 2 is a diagram illustrating an example of the hardware configuration of an image processing device according to the present embodiment; FIG. 3 is a diagram illustrating functional units of an image processing device according to the present embodiment; FIG. 4 is a diagram illustrating an example of a text information table according to the present embodiment; FIG. 5 is a diagram illustrating an example of an input image table according to the present embodiment; FIG. 6 is a diagram illustrating an example of a correction item table according to the present embodiment; FIG. 7 is a diagram illustrating an example of a generated image table according to the present embodiment; FIG. 8 is a diagram illustrating an example of a flow (part 1) of an image processing method according to the present embodiment; FIG. 9 is a diagram illustrating an example of a processing rule table according to the present embodiment; FIG. 10 is a diagram illustrating an example of a flow (part 3) of an image processing method according to the present embodiment; FIG. 11 is a diagram illustrating an example of an output screen according to the present embodiment; FIG. 12 is a diagram illustrating an example of an output screen according to the present embodiment; FIG. 13 is a diagram illustrating an example of an output screen according to the present embodiment; FIG. 14 is a diagram illustrating an example of an output screen according to the present embodiment;

[0017] Specific embodiments of the present invention will be described below. For ease of explanation, the following description may be given in terms of a GUI (Graphic User Interface). Furthermore, since the basic data processing technologies (communication / transmission technologies, data acquisition technologies, data recording technologies, data processing / analysis technologies, machine learning technologies, image processing technologies, visualization technologies, etc.) required to realize the present invention are well-known technologies, a description thereof will be omitted.

[0018] In addition, in this specification, the concept of "device" includes not only a single device that performs a specific function, but also a combination of multiple devices that exist independently and in a distributed manner but cooperate (link) to perform a specific function.

[0019] Furthermore, in this specification, a "user" refers to a user of the image processing device of the present invention, specifically, for example, a person who performs image correction processing using image generation AI through the functions of the image processing device of the present invention.

[0020] In addition, in this specification, the term "person" refers to an entity that performs a specific action, and includes individuals, groups, corporations such as companies, and organizations, as well as computers and devices that constitute artificial intelligence (AI). Artificial intelligence (AI) realizes intelligent functions such as inference, prediction, and judgment using hardware and software resources. The algorithm of the artificial intelligence is arbitrary, and examples include expert systems, case-based reasoning (CBR), convolutional neural networks (CNN), deep neural networks (DNN), Bayesian networks, and subsumption architectures.

[0021] <<One Embodiment of the Present Invention>> [Configuration of Image Processing System] In one embodiment of the present invention (hereinafter, this embodiment), an image processing system 5 is assumed to be configured with an image processing device 10, a user terminal 30, an image capturing device 40, and a server computer 50, all connected to a network 1 shown in Fig. 1. Of these, the user terminal 30 and the image capturing device 40 are information processing devices that allow a user to capture an image that will serve as a base image for image generation as needed, and upload this to the image processing device 10. In addition to uploading the base image, the user also uploads text information required for image generation by the image generation AI to the image processing device 10.

[0022] The user may input text information directly into the UI (User Interface) of the image processing device 10 without using the user terminal 30 or the photographing device 40, for example, by scanning or reading a photo print they have on hand, the screen of their own user terminal 30 (on which a photo image is being displayed), or the data as the base image, or by using a UI such as a keyboard, mouse, or touch panel. When such a configuration and operation is adopted, the image processing system 5 can be composed of only the image processing device 10.

[0023] The image processing device 10 stores the base images uploaded from the user terminal 30 or the photographing device 40 in an input image table 212 (described later in Figures 3 and 5), and stores the text information in a text information table 211 (described later in Figures 3 and 4), in preparation for image generation and caption generation.

[0024] The image processing device 10 is an information processing device that executes image generation and correction processes based on the base image and text information, using an image generation AI and a caption generation AI that are included in the image processing device 10 itself or that are provided by the server computer 50. However, such processes are executed in accordance with a user instruction received through a predetermined UI in accordance with the image processing method of the present invention.

[0025] Furthermore, as described above, the server computer 50 is a server device that provides the image generation AI and caption generation AI functions to the image processing device 10 via the network 1. Of course, if the image processing device 10 is configured such that it does not require the provision of the image generation AI and caption generation AI functions, the server computer 50 does not need to be included in the image processing system 5.

[0026] In addition to the configuration in which the image processing device 10 and the photographing device 40 are connected via the network 1, the internal bus wiring of the image processing device 10 and the interface of the photographing device 40 may be directly connected.

[0027] In addition, various situations can be imagined, such as the photographer who performs various operations related to taking photos with the photographing device 40 and outputting photo prints, the registrant who inputs the image (base image) obtained from the photo shoot and its photo prints into the image processing device 10 (which may include scanning operations or uploading from the user terminal 30), and the viewer who views the generated image that is the processing result of the image processing device 10 and its modified results on the user terminal 30, all being different people or all being the same person.

[0028] [Configuration Example of Image Processing Device] Next, a configuration example of the image processing device 10 according to this embodiment will be described with reference to FIGS. 2 and 3. The image processing device 10 is composed of a computer used by a user, and specifically, is composed of a PC (Personal Computer), a smartphone, a tablet terminal, a notebook PC, or the like. Note that the image processing device 10 is not limited to a computer owned by the user, and may be composed of a terminal not owned by the user but available when visiting a store, facility, or the like, such as a store-installed terminal. Note that the following description will be given taking as an example a case where the image processing device 10 is composed of a user-owned computer, specifically, a PC.

[0029] As shown in FIG. 2, the computer that constitutes the image processing apparatus 10 includes a processor 11, an auxiliary storage device 12, a main storage device 13, an input device 14, an output device 15, and a communication device 16.

[0030] The processor 11 is composed of, for example, a CPU (Central Processing Unit), an MPU (Micro-Processing Unit), an MCU (Micro Controller Unit), a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), a TPU (Tensor Processing Unit), or an ASIC (Application Specific Integrated Circuit).

[0031] The auxiliary storage device 12 may be composed of, for example, a flash memory, a hard disc drive (HDD), a solid state drive (SSD), a flexible disc (FD), a magneto-optical disc (MO disc), a compact disc (CD), a digital versatile disc (DVD), a secure digital card (SD card), or a universal serial bus memory (USB memory).

[0032] The auxiliary storage device 12 may be built into the computer that constitutes the image processing device 10, or may be externally attached to the computer. Alternatively, the auxiliary storage device 12 may be configured as a NAS (Network Attached Storage) or the like. The auxiliary storage device 12 may also be an external device, such as an online storage or a database server, that can communicate with one of the computers that constitute the image processing device 10 via a communication network.

[0033] The auxiliary storage device 12 also stores an operating system (OS) and a program 121 such as an application for executing image processing. When the program 121 is read and executed by the processor 11, the computer constituting the image processing device 10 performs the functions of the reception unit 21, generation unit 22, determination unit 23, and output unit 24 shown in Fig. 3, and specifically, executes a series of processes related to image correction based on a base image and text information.

[0034] The main storage device 13 is configured by semiconductor memories such as a ROM (Read Only Memory) and a RAM (Random Access Memory). The processor 11 loads the program 121 onto the main storage device 13 and executes it there.

[0035] The input device 14 is a device that accepts user input operations and is configured, for example, by a keyboard, a mouse, a touch panel, a camera unit, etc. The input device 14 may also include a photographing device implemented in a digital camera, a microphone for collecting sound, etc. The output device 15 is configured, for example, by a display, a speaker, etc.

[0036] The communication device 16 may be configured by, for example, a network interface card or a communication interface board, etc. The computer configuring the image processing device 10 can communicate with other devices connected to a network 1 such as the Internet or a mobile communication line via the communication device 16.

[0037] 3, the image processing device 10 has a receiving unit 21, a generating unit 22, a determining unit 23, and an output unit 24. These functional units are realized by the processor 11 of the image processing device 10 executing the above-mentioned program 121 and cooperating with other hardware devices of the image processing device 10. In addition, the image generation AI 221 and the caption generation AI 222 in the generating unit 22 may be called from the server computer 50 and executed.

[0038] 12, 14, 16, etc., the image processing device 10 of the present invention modifies an additional placement image Pa desired by the user in accordance with a user operation on a slide Us in a slider bar Ub in a base image such as an input image Pi provided by the user or P1 generated by the image processing device 10. The base image is an image including a background image Pb such as a distant view of the sky or a landscape, and one or more placement images Ph arranged on this background image Pb.

[0039] The background image Pb is an image on which the layout image Ph is placed, and functions as a mount for the layout image Ph. The background image Pb is an image that is the same size as or larger than the layout image Ph, and may be a landscape or a specific monochrome image such as white or black. The background image Pb may also be patterned, or an illustration of a character or the like may be drawn as one aspect of the pattern. The background image Pb may also be the input image Pi itself, which will be described later.

[0040] Furthermore, the image generated as a result of corrections made by the image processing device 10 is a first candidate image P1, and the image generated as a result of further corrections made in response to user operations is a second candidate image P2. For ease of explanation, the first and second candidate images have been described, but it is assumed that a third candidate image, a fourth candidate image, and so on through an n-th candidate image are generated each time a correction is made by user operation. Note that an image for which the user has completed corrections, i.e., an image for which the user is satisfied, is output as an output image Po.

[0041] In this specification, an "image" is an image composed of multiple pixels, represented by the gradation values ​​of each pixel, and includes at least one layout image Ph. Digital image data (hereinafter referred to as "image data") that defines an image at a set resolution is generated by compressing data in which the gradation values ​​for each pixel are recorded using a predetermined compression method. Examples of types of image data include lossy compressed image data such as JPEG (Joint Photographic Experts Group) format, and lossless compressed image data such as GIF (Graphics Interchange Format) or PNG (Portable Network Graphics) format.

[0042] (Reception Unit) The reception unit 21 is a functional unit that receives an input image Pi and text Ti, which is input text information (hereinafter, referred to as input text Ti), from the input device 14 or from the user terminal 30 or the imaging device 40 via the communication device 16, and stores these in, for example, the auxiliary storage device 12. The reception unit 21 stores the input text Ti received as described above in a text information table 211, and stores the input image Pi in an input image table 212.

[0043] 4 shows an example of the configuration of the text information table 211 of this embodiment. The text information table 211 is a collection of records including values ​​of a user ID, a date and time, and text information. Among these, the user ID is identification information of the user who input the input text Ti, and may be identification information of the user terminal 30. Furthermore, if the image processing device 10 is operated for use by a single user, this user ID is unnecessary. The date and time is the date and time when the input text Ti was input.

[0044] The text information is text information input by the user, stored for example by keyword. When the user inputs text using the input device 14 or the user terminal 30, the text information is stored in each field with words and phrases spaced apart on the UI. Alternatively, the reception unit 21 may apply the input sentence to the morphological analysis engine, extracting one or more words and phrases and storing them in each field. The morphological analysis engine may be held by the reception unit 21 or may be callable via the network 1.

[0045] 5 shows an example of the configuration of the input image table 212 in this embodiment. The input image table 212 is a collection of records including values ​​of a user ID, date and time, image data, and caption. Among these, the user ID is identification information of the user who input the input image Pi, and may also be identification information of the user terminal 30. Furthermore, if the image processing device 10 is operated so that it is used by a single user, this user ID is unnecessary. The date and time is the date and time when the input image Pi was input.

[0046] A caption is a keyword obtained by assigning an input image entered by a user to the caption generation AI 222. Such a caption is one or more words obtained by the reception unit 21 assigning an input image obtained from the input device 14 or the user terminal 30 to the caption generation AI 222. The caption generation AI 222 may be held by the reception unit 21 or may be callable via the network 1.

[0047] The caption generation AI 222 is a learning model for caption generation that is created in advance by machine learning. This learning model, for example, identifies the features of an input image Pi and generates a caption according to the features, and is constructed by performing machine learning using previously acquired input images Pi and captions assigned to those images (for example, appropriate captions assigned by knowledgeable individuals) as training data.

[0048] The technique for obtaining a caption from an input image Pi is not limited to the above-described technique using the caption generation AI 222. For example, the correspondence between the image analysis results of the input image Pi and the captions may be stored as data, and the reception unit 21 may select (generate) a caption by referring to the correspondence. The correspondence between the image analysis results and the captions may be, for example, a correspondence between the image analysis results, such as the color and pattern of the background image Pb and the type, shape, size, and theme of the object in the layout image Ph, and captions such as "sea," "lake," "castle," "high mountain," "sea at sunset," "rainy grassland," and "lion family."

[0049] The method of receiving the input image Pi and the input text Ti in the receiving unit 21 is not particularly limited, but includes acquiring the input image Pi and the input text Ti by reading, using a camera unit, a photo print of an image captured by the user terminal 30 or the photographing device 40. The receiving unit 21 may also acquire the image by downloading image data from an external device, a web server, or the like via the network 1.

[0050] (Generation Unit) The generation unit 22 is a functional unit that generates first to n-th candidate images P1 to PN as necessary based on the input image Pi and input text Ti obtained by the reception unit 21. The input image Pi may also include an image generated by the generation unit 22 by assigning the input text Ti to the image generation AI 221. The input text Ti may also include an image generated by the generation unit 22 by assigning the input image Pi to the caption generation AI 222. The image generation AI 221 may be used through an image generation service currently provided on the network 1, but may also be stored by the image processing device 10 itself. The image generation AI 221 is an artificial intelligence that automatically generates an image desired by a user based on text information input by the user.

[0051] In this embodiment, when generating the first to nth candidate images P1 to PN, the generation unit 22 generates an additional placement image Pa corresponding to the input text Ti using the image generation AI 221 and combines it with the input image Pi. This combination is performed by placing the additional placement image Pa in a predetermined area of ​​the background image Pb or the layout image Ph of the input image Pi. During the above-described placement, the generation unit 22 has a function of, for example, analyzing a blank area in the background image Pb of the input image Pi that does not overlap with the layout image Ph by a certain amount or more, and determining the placement area of ​​the additional placement image Pa. Alternatively, the generation unit 22 may have a function of analyzing an area in the input image Pi that has a certain positional relationship with the layout image Ph (e.g., a horizontal row, vertically connected, or located at the center of gravity of the layout image Ph) and determining the placement area of ​​the additional placement image Pa. Furthermore, the AI ​​may determine the placement of the layout image Ph in the input image Pi, or the user may specify the position.

[0052] The above analysis involves, for example, identifying the characteristics of various phenomena related to the image, such as information about the hue and tone of each region of the image, information about image quality, pixel gradation values, the size and position of blank regions, the layout of the arranged images Ph (e.g., the number of images arranged, the arrangement interval, the arrangement shape, size, size variation, etc.), and information about the main object (e.g., the subject of the arranged images Ph, such as a building or a person) estimated from this information.

[0053] The object information may include the type and state of the object, the position of the object in the image, and, if the object is a living thing, the pose, gesture, age, gender, etc. Furthermore, it is desirable that the image features be quantifiable, vectorizable, or tensorizable. In this case, the image analysis results will be the quantifiable, vectorizable, or tensorizable image features, i.e., feature quantities. The arrangement image Ph may also be a landscape image, i.e., the object contained in the arrangement image Ph may be only a landscape, with the entire image representing the landscape as an object.

[0054] (Determination Unit) The determination unit 23 is a functional unit that presents a UI for the first to nth candidate images P1 to PN generated by the generation unit 22 to the user, indicating the type and amount of correction for each image, and determines the content (item and amount of correction) selected by the user via the UI as the correction content. Examples of the correction items include object types, attributes, and states, as well as values ​​indicating these, such as the index values ​​of the scales of the slider bar Ub shown in Figures 12, 14, 16, and 17 (e.g., grassland, pond, lake, Western-style, Japanese-style, calm, egg, chick, chicken, skeleton). The user can freely adjust, for example, the type and attributes of the object corresponding to the additional placement image Pa by operating the slide Us on the slider bar Ub using the user terminal 30 or the like. In response to the user's operation of the slide Us, the determination unit 23 acquires the type and attributes of the additional placement image Pa desired by the user and determines them as generation parameters for the second candidate image P2 and subsequent candidate images. The generation parameters are passed from the determination unit 23 to the generation unit 22, and candidate image generation is executed.

[0055] When setting the correction items and correction amounts on the slider bar Ub, the determination unit 23, for example, compares the input text Ti and the caption (generated by the caption generation AI 222) with the values ​​of each item in the records of the correction item table 231 shown in FIG. 6 , identifies a record containing a corresponding value in the item, and acquires the value indicated by the record as the above-mentioned item and correction amount. The correction item table 231 may be replaced with a model obtained by machine learning using training data corresponding to the same content. In other words, the correction items and correction amounts may be determined using the model. Alternatively, both the correction item method and the model method may be implemented, and the results obtained from each may be integrated. Furthermore, by inputting text such as the captions generated from various images such as candidate images and input images Pi, or input text Ti entered by the user, or a selection of such texts from the user, into an automatic correction item generation model, the input text (e.g., "growth" or "chicken") can be treated as one of the correction items, and related intermediate concepts (e.g., egg, chick, etc.) can be automatically generated. The correction items obtained in this manner become the index values ​​of the scale on the slider bar Ub. The automatic generation model is a model that has undergone machine learning using a set of information indicating a certain concept and its intermediate concepts as training data. The automatic generation model can be provided by the image processing device 10 itself, or can be called up and used from an external device via the network 1.

[0056] The correction item table 231 shown in FIG. 6 is a collection of records including values ​​such as a broad concept, a medium concept, an event, and items 1 to N. Among these, a broad concept is the highest-level concept of a thing corresponding to an item, and a medium concept is a concept subordinate to the broad concept. For example, the medium concepts "body of water" and "mountain" are defined for the broad concept "nature." Events are definitions of events that may be included in the medium concept. For example, for the medium concept "body of water," events include "size," "water surface / waves," and "color." Items 1 to N represent the item itself corresponding to the event or its correction value. For example, for the event "size" of the medium concept "body of water," items 1 to N are defined as item 1 "ocean," item 2 "lake," item 3 "pond," and item N "puddle." The correction item table 231 may be set and edited by the user.

[0057] In the above example, the amount of correction for an item is defined using text such as "lake" or "pond," but it may also be defined using specific numerical values ​​that indicate the size, such as "area of ​​100,000 square kilometers or more," "50 to less than 100,000 square kilometers," or "0.5 to less than 50 square kilometers."

[0058] In addition to defining items and their correction amounts based on physical aspects such as the type and size of the event, as described above, the items and correction amounts may also be defined based on the form or state of the concept. For example, as shown in the correction item table 231 in FIG. 6 , for the event "Water Surface / Waves" of the intermediate concept "Water Area," the items and correction amounts may be defined based on the state of the event, such as Item 1 "Calm," Item 2 "Small Waves," Item 3 "Medium Waves," and Item N "Stormy Weather." Alternatively, for the event "Atmosphere" of the intermediate concept "Building," the items and correction amounts may be defined in the form of the "Atmosphere" of the "Building," such as Item 1 "Western," Item 2 "Greek," Item 3 "Southeast Asia," and Item N "Japanese." Furthermore, for the event "Growth" of the intermediate concept "Living Things," the items and correction amounts may be defined in the form of the "Growth" of the "Living Things," such as Item 1 "Egg," Item 2 "Chick / Infant," Item 3 "Adult," and Item N "Skeleton." In this case, the morphological changes that occur over time, from the growth of the "living thing" to its death, are defined as items and their correction amounts. This concept of the passage of time can also be applied to other major and intermediate concepts.

[0059] As shown in screen G7 of FIG. 16 , the determination unit 23 may accept a user's designation of a location for placing the additional placement image Pa in the input image Pi or each candidate image, i.e., a drawing area, by moving a location designation object G72. In the example of screen G7 of FIG. 16 , a dashed ellipse object is shown as the object G72, but this is merely an example, and any object that allows the user to move the object and designate a specific area may be used. Alternatively, the user may set the drawing area by drawing an outline (so-called freehand) without using an object. The determination unit 23 notifies the generation unit 22 of the information on the drawing area accepted by the object G72, along with the information on the above-mentioned correction items and correction amounts, and uses the information to generate candidate images.

[0060] (Output Unit) The output unit 24 is a functional unit that displays the first candidate image P1 and subsequent candidate images generated by the generation unit 22, as well as the output image Po, which is a candidate image that the user is satisfied with and whose drawing content, etc. has been finalized, on the output device 15 such as a display or the user terminal 30. For this reason, the output unit 24 stores and holds the images generated by the generation unit 22 in a generated image table 232 (see FIG. 7 ), and calls and outputs them as appropriate in response to a request from the user, the arrival of a predetermined timing, etc.

[0061] The method of outputting the output image Po etc. is not particularly limited, but includes, for example, the output unit 24 displaying the output image Po on a display or monitor of the output device 15 or the user terminal 30, printing the output image Po, sending the output image Po to another user, and providing the output image Po as a commercial material. The output image Po as a commercial material may include media consisting of one or more pages or cards on which images are published, such as an album, a photo book, a postcard, a message card, an electronic album, a photo print, etc.

[0062] [Image Processing Method Flow Example] Next, as an operation example of the image processing device 10 in this embodiment, an image processing flow using this device will be described. The image processing flow described below uses the image processing method of the present invention. In other words, each step in the image processing flow described below corresponds to a component of the image processing method of the present invention. Note that the flow below is merely an example, and some steps in the flow may be deleted, new steps may be added, or the execution order of two steps in the flow may be reversed, as long as it does not deviate from the spirit of this embodiment.

[0063] The steps in the image processing flow according to this embodiment are performed by the processor 11 included in the image processing device 10 in the order shown in Fig. 8. That is, in each step in the image processing flow, the processor 11 executes processing corresponding to each step in Fig. 8 among the data processing defined in the image processing application program.

[0064] Specifically, in the image processing flow according to this embodiment, first, the reception unit 21 receives at least one of an input image Pi and text information (text information related to an object to be drawn; hereinafter, referred to as input text Ti) from, for example, the user terminal 30 (S1). In this process, the reception unit 21 displays a reception screen G1 shown in Fig. 11 on the user terminal 30 and acquires input text Ti, which is an input value, via an input field G11 on the reception screen G1. The reception unit 21 also acquires the file of the input image Pi specified by the user in a file search field G12 on the reception screen G1.

[0065] Acquisition of the input text Ti and input image Pi is necessary for generating a first candidate image, which serves as a base image to be modified by the user. When the receiving unit 21 receives only the input text Ti, the value of the input text Ti is transmitted to the generating unit 22, and the image generation AI 221 generates a first candidate image P1 (S2). Screen G2 in FIG. 12 shows an example of a first candidate image P1 generated based on the input text "mountain grassland." This first candidate image P1 is an image in which a grassland layout image Ph is arranged in the foreground and a mountain layout image Ph is arranged in the background of a background image Pb depicting the sky.

[0066] Note that when only the input image Pi is accepted, the input image Pi is transmitted to the generation unit 22 and becomes the first candidate image P1. On the other hand, when both the input text Ti and the input image Pi are accepted on the reception screen G1, they are transmitted to the generation unit 22, and the image generation AI 221 of the generation unit 22 executes a process of generating an additional placement image Pa based on the input text Ti and a process of arranging the additional placement image Pa on the input image Pi to generate the first candidate image P1.

[0067] When the receiving unit 21 receives input of input text Ti in the additional information field G21 of the screen G2 and a click on the generate button G22 (S3), it notifies the determination unit 23 of this as a correction request. The determination unit 23 identifies the item to be corrected by comparing the input text Ti notified by the receiving unit 21 with the correction item table 231 (S4). In this identification, for example, if the input text Ti is "lake," the determination unit 23 identifies a record in the correction item table 231 related to the phenomenon "size" of the intermediate concept "body of water." This is because "lake" is included in "item 2" among "item 1" to "item N" indicated by the record. Note that, in addition to the mode of presenting the first candidate image P1 and receiving a correction instruction from the user, the following slider bar Ub may be generated and displayed without generating and presenting the first candidate image P1.

[0068] The determination unit 23 generates a slider bar Ub based on the correction amount (in the above case, ocean, lake, pond, and puddle) of the correction item identified in the correction item table 231 (in the above case, the "size" of the "water area") and responds to the user terminal 30 (S5). In the case of the "lake" above, a slider bar Ub is generated with scales indicating the correction amounts (ocean, lake, pond, marsh, and puddle) for the "size" of the correction item "water area." In the case of screen G3A of FIG. 12, "grassland" is set at the left end of the scale of the slider bar Ub. Furthermore, no correction amount value is set for the scale immediately to the right of the scale "grassland." A scale without a correction amount value corresponds to the default value for the correction item, i.e., the additional placement image Pa before receiving a user designation is displayed in the second candidate image P2. On the other hand, the scale "grassland" corresponds to the situation before the additional placement image Pa is placed, i.e., the situation when the first candidate image P1 is displayed.

[0069] In the above, the slider bar Ub is exemplified as an interface for the user to select the correction item and correction amount, but instead of such a slider bar Ub, for example, radio buttons corresponding to each correction amount may be provided, and the user may select the correction amount using the radio button.

[0070] The determination unit 23 accepts a user operation of the slide Us on the slider bar Ub (S6). In this case, the determination unit 23 acquires the scale value at which the slide Us is placed on the slider bar Ub. The determination unit 23 notifies the generation unit 22 of the accepted scale value. In response to this notification, the generation unit 22 changes the type and attributes of the additional placement image Pa displayed on the screen G3A to generate or modify a second candidate image P2, and returns this to the user terminal 30 (S7). An example of an output screen for this second candidate image P2 is shown in screen G3A in FIG. 12. As long as the OK button is not clicked on this screen G3A and the slide Us is operated on the slider bar Ub or the back button is clicked (S8: N), the generation unit 22 repeats steps S3 to S7. On the other hand, if the OK button is clicked on screen G3A (S8: Y), the generation unit 22 responds to the user terminal 30 with the candidate image currently displayed on screen G3A (in the case of Figure 12, the second candidate image P2) as the output image Po (S9), and ends this flow.

[0071] The type and attributes of the additional placement image Pa may be changed based on a rule table stored in advance by the generation unit 22. In this case, the rule table stores rules for image processing, determined according to the scale, i.e., the amount of correction for each correction item (see the processing rule table in FIG. 9). The processing rule defined in this rule table is a rule that, if the correction item is the event "size," enlarges or reduces the size of the original image by a predetermined percentage depending on the concept of each item, i.e., the amount of correction. Also, if the correction item is the event "water surface / waves," the rule is a rule that increases or decreases the height, width, and appearance frequency of the wave portion of the image in the original image by a predetermined percentage depending on the concept of each item, i.e., the amount of correction. The numerical value assigned to each scale may be the applicability when the input text Ti (corresponding event, state, etc.) entered by the user is set to a maximum of 100. In this case, the numerical values ​​may be determined, for example, by the ratio of the sizes between the corresponding objects, or by dividing 100 by the number of intermediate concepts of the event (for example, if there are three intermediate concepts between 0 and 100, resulting in a total of four levels between concepts, 25 obtained by dividing 100 by 4 is used as one scale, and the values ​​are determined as 0, 25, 50, 75, 100, etc.). In this case, it is advisable to also include information about the input text, i.e., the name of the object, etc. Also, even for correction axes determined automatically or using a table, numerical values ​​on the scale with 100 as the maximum value may be entered, as in the above case. In this case, it is preferable to also include information about what the axis to which the scale is attached (for example, the size or atmosphere of a certain object, etc.) represents.

[0072] Furthermore, when the correction item is the phenomenon "atmosphere," the rule may be one that changes, for example, the color or shape of an object in the original image according to the style of each item, i.e., the concept of the correction amount. For example, when the correction amount for "atmosphere" is "Western," the surface of the object is configured with the shape and color of stone, and when the correction amount is "Japanese," the surface of the object is configured with the shape and color of wood and a tiled roof. When the correction item is the phenomenon "growth," the rule may be one that changes the image itself or its shape of the object in the original image according to the growth of each item, i.e., the concept of the correction amount. For example, when the correction amount for "growth" is "egg," the object is changed to an image of an egg; when the correction amount is "chick," the object is changed to an image of a chick; when the correction amount is "adult bird," the object is changed to an image of an adult bird; and when the correction amount is "skeleton," the object is changed to an image of a bird skeleton.

[0073] Furthermore, if the correction item is the "painter" phenomenon, rules can be adopted that change the color and shape of the object in the original image according to the artist, depending on each item, i.e., the concept of the correction amount. For example, for the "painter" correction amount "Van Gogh," the object's surface is processed into a spiral shape to emphasize yellow and blue, while for the "painter" correction amount "Gauguin," the object's color is changed to warm colors, emphasizing the roundness and thickening of the drawn lines. In such a case, for example, as shown in the flow chart of FIG. 10 , the receiving unit 21 receives an image of a painting by a specific artist as an input image Pi from the user as a base image (S30). The determining unit 23 then inputs the base image obtained by the receiving unit 21 into, for example, an AI for style determination to identify the style (S31). This AI may use a determination model that has been machine-learned using training data that associates each artist's representative paintings with descriptions of their style. The receiving unit 21 also receives input text Ti related to the object to be drawn (S32). The generation unit 22 provides the input text Ti and drawing instructions based on the above rules to the image generation AI 221, and generates a first candidate image P1 in the style of the corresponding artist.

[0074] [Modifications of the Present Embodiment] The present embodiment is not limited to the above embodiment, and for example, the following modifications are possible. These modifications will be described below. Note that the following description will focus on the differences between the modifications and the above embodiment.

[0075] (Regarding First Modification) In the above-described embodiment, the image processing device 10 receives input text Ti related to the additional placement image Pa on the screen G2, identifies the items to be corrected by comparing the input text Ti with the correction item table 231, generates the slider bar Ub, and recognizes the user's intention to correct the items. However, without being limited to this, as a modification of the present embodiment, when a request to correct the first candidate image P1 is made by the user (for example, when a correction button is clicked on the screen of the user terminal 30), for example, the determination unit 23 may analyze the input text Ti used to generate the first candidate image P1 and estimate candidates for the items to be corrected that the user desires.

[0076] For example, suppose that the input field G41 on the screen G4 of FIG. 13 accepts two words, "sea" and "castle," as input text Ti for generating a first candidate image. The generation unit 22 assigns these two words to the image generation AI 221, resulting in the first candidate image P1 shown on the screen G5. In this case, the determination unit 23 analyzes the input text Ti for the first candidate image, "castle, sea," and determines potential correction items that the user may want to correct. The correction item candidates are determined by comparing the words with the correction item table 231 to determine corresponding phenomenon items (e.g., castle → Western style, Japanese style, sea → water surface / waves, castle × sea → sense of distance between castle and sea). As already mentioned, instead of the correction item table 231, correction item candidates may be determined using a trained model that has undergone machine learning using training data that is a set of input text Ti and correct correction items.

[0077] An example of a screen according to the first modified example is shown in Fig. 14. Screen G5 in Fig. 14 shows a configuration in which a first candidate image P1 generated from the input text "ocean castle" has a layout image Ph, which is an image of a castle and the sea, placed on a background image Pb depicting the sky, and a slider bar Ub and slide Us, which accept the user's specification of the amount of correction for the correction item, are displayed. Of these, the slider bar Ub is generated and displayed for each of "castle" and "ocean."

[0078] 14 shows a screen G6A in which a second candidate image P2 is displayed in which the castle is changed to a Japanese-style one and the sea is changed to a rough one in response to a user operation of sliding Us on the slider bars Ub for "castle" and "sea" for the first candidate image P1. Also, a screen G6B in FIG. 14 shows a configuration in which the user clicks the OK button on the second candidate image P2 to complete the modification, and an output image Po is displayed.

[0079] (Regarding the Second Modification) Alternatively, instead of identifying the item to be corrected using the input text Ti obtained from the user as a key, it is also possible to extract words by assigning the first candidate image P1 to the caption generation AI 222, and then identify the item to be corrected based on the extracted words. In this case, the determination unit 23 inputs the input image Pi (including the concept of the first candidate image P1) received from the user by the reception unit 21 to the caption generation AI 222. The caption generation AI 222 analyzes the input image Pi and generates words that are captions. This caption generation AI 222 is an AI having a learning model that has undergone machine learning using a set of the input image Pi and the correct words to be extracted from the input image Pi as training data.

[0080] The determination unit 23 acquires the above words and phrases (for example, the words "wave" and "castle" from the first candidate image P1) from the caption generation AI 222, and identifies the items to be corrected and the amount of correction by comparing these words and phrases with the correction item table 231, and generates a slider bar Ub corresponding to each word and phrase. Note that, among the words and phrases extracted by the caption generation AI 222, control may be performed so that "castle sea" and words similar thereto, which were originally input for generating the first candidate image, are not presented as correction items.

[0081] Furthermore, as described above, a set of correction items and correction amounts is determined for each word or phrase, but a predetermined number of slider bars Ub may be generated and sent to the user terminal 30. In this case, the determination unit 23 may evaluate the difference in priority between correction items and send only slider bars Ub related to items with priorities equal to or higher than a threshold.

[0082] Meanwhile, the user searches for the item they wish to modify using the slider bar Ub and executes a modification instruction using the slide Us on the target slider bar Ub (for example, sliding the slide Us toward "stormy weather" for "waves"). The determination unit 23 notifies the generation unit 22 of the modification amount obtained using the slider bar Ub and the phrase (waves, castle) obtained by the caption generation AI 222. The generation unit 22 assigns text information "castle, sea, stormy weather" to the image generation AI 221 and generates a second candidate image P2 from the text information. At this time, the first candidate image P1 may also be input to the image generation AI 221 as a reference image.

[0083] In addition, when the caption generation AI 222 extracts, for example, the word "waves," the word may be input to another AI (rather than the correction item table 231) to extract keywords indicating the degree of adjustment, such as "rough" or "gentle."

[0084] (Regarding the third modified example) Note that the first candidate image P1 itself may be an image input by the user. An example flow in this case is shown in FIG. 15 . Prior to this flow, it is assumed that the user inputs an image (real-life image) of a castle and the sea, taken by the user himself / herself using the camera device 40 or the like, as the first candidate image P1 to the reception unit 21. The reception unit 21 passes this input image Pi to the generation unit 22 as the first candidate image P1. The generation unit 22 inputs the first candidate image P1 to the caption generation AI 222 (S15). The generation unit 22 acquires the captions "castle" and "sea" generated by the caption generation AI 222 for the first candidate image P1 (S16).

[0085] The accepting unit 21 also accepts input text Ti from the user regarding the additional placement image Pa (S17). The determining unit 23 compares the words obtained in S16 and the input text Ti obtained in S17 with the correction item table 231 to identify the correction item and the amount of correction (S18). Thereafter, steps S19 to S22 are executed in the same manner as steps S5 to S9 in the flow of FIG. 8.

[0086] Examples of screens corresponding to the above-described modified example are shown in FIGS. 16 and 17 . Screens G7 and G8 shown in FIG. 16 analyze the input image Pi to acquire the phrase "grassland," acquire "lake" as input text Ti related to the additional placement image Pa, and display a scale of correction amounts ranging from "grassland" to "lake" on a slider bar Ub. Also shown is a control example in which first to fourth candidate images P1 to P4 generated according to each scale are displayed alongside the scale. In this case, the generation unit 22 generates candidate images corresponding to the scale and displays them on screen G8. At this time, an object G72 may be displayed, which allows the user to specify which region of the first candidate image P1 the additional placement image Pa should be placed in, and the object G72 may be moved by the user. In this case, the placement position of the object G72 becomes the placement position of the additional placement image Pa.

[0087] 17 also shows a screen G9 and a screen G10 in which the input image Pi is not analyzed to extract words or phrases, and the input text Ti for the additional placement image Pa is "chicken." A slider bar Ub is used to display a scale of corrections ranging from "none" to "skeleton." Also shown is a control example in which the first through fifth candidate images P1 through P5, generated according to the scales, are displayed alongside the scales. In this case, the generation unit 22 generates candidate images corresponding to the scales and displays them on the screen G10. At this time, an object G92 may be displayed, which allows the user to specify in which region of the first candidate image P1 the additional placement image Pa "chicken" should be placed, and the object G92 may be moved by the user. In this case, the placement position of the object G92 becomes the placement position of the additional placement image Pa. Moreover, each value of the scale on the slider bar Ub corresponds to the value of each item related to the "growth" phenomenon of "animal" in the correction item table 231, from "egg to skeleton."

[0088] Although specific embodiments of the present invention have been described above, the above embodiments are merely examples given to facilitate understanding of the present invention and are not intended to limit the present invention. That is, the present invention may be modified or improved from the embodiments described below without departing from the spirit of the present invention. The present invention also includes equivalents thereof. Furthermore, embodiments of the present invention may include a combination of the above embodiments with one or more of the following modifications.

[0089] (Regarding the Computer Constituting the Image Processing Device) In the above embodiment, the image processing device 10 of the present invention is configured by a computer directly used by the user, such as a user-owned PC (Personal Computer). However, this is not limited thereto, and the image processing device 10 of the present invention may also be configured by a computer indirectly available to the user, such as a server computer 50. Here, the server computer 50 may be, for example, a server computer for a cloud service, specifically, a server computer for an ASP (Application Service Provider), SaaS (Software as a Service), PaaS (Platform as a Service), or IaaS (Infrastructure as a Service). In this case, when necessary information is input into the user's user terminal 30, the server computer 50 performs various processes (calculations) including image generation based on the input information, and the calculation results are output on the user terminal 30. In other words, the functions of the server computer 50 constituting the image processing device of the present invention can be used on the user terminal 30. Alternatively, the image processing device 10 may be configured by a photographing device 40 or a smartphone (a type of user terminal 30) used by a user, in addition to the above-mentioned PC or server computer.

[0090] (Regarding Processor Configuration) The processor 11 included in the image processing device 10 of the present invention includes various types of processors. The various types of processors include, for example, a CPU, which is a general-purpose processor that executes software (programs) and functions as various processing units. The various types of processors also include a PLD (Programmable Logic Device), which is a processor whose circuit configuration can be changed after manufacture, such as an FPGA (Field Programmable Gate Array). Furthermore, the various types of processors also include dedicated electrical circuits, such as an ASIC (Application Specific Integrated Circuit), which is a processor having a circuit configuration designed specifically for performing specific processing.

[0091] Furthermore, one functional unit of the image processing device 10 of the present invention may be configured by one of the various processors described above. Alternatively, one functional unit of the image processing device 10 of the present invention may be configured by a combination of two or more processors of the same or different types, such as a combination of multiple FPGAs, or a combination of an FPGA and a CPU. Furthermore, the multiple functional units of the image processing device 10 of the present invention may be configured by one of the various processors, or two or more of the multiple functional units may be combined into a single processor. Furthermore, as in the above embodiment, one processor may be configured by a combination of one or more CPUs and software, and this processor may function as multiple functional units.

[0092] Furthermore, a processor may be used that realizes the functions of the entire system including multiple functional units in the image processing device 10 of the present invention on a single IC (Integrated Circuit) chip, as typified by, for example, an SoC (System on Chip).The hardware configuration of the various processors described above may also be an electric circuit that combines circuit elements such as semiconductor elements.

[0093] 1 Network 5 Image processing system 10 Image processing device 11 Processor 12 Auxiliary storage device 121 Program 13 Main storage device 14 Input device 15 Output device 16 Communication device 21 Reception unit 211 Text information table 212 Input image table 22 Generation unit 221 Image generation AI 222 Caption generation AI 23 Determination unit 231 Correction item table 232 Generated image table 24 Output unit 30 User terminal 40 Photography device 50 Server computer Ti Input text Tx Output text Pi Input image Po Generated image P1 First candidate image P2 Second candidate image Pb Background image Ph Placed image Pa Additional placed image Ub Slider bar Us Slide Uc Item

Claims

1. An image processing device having a processor, wherein the processor receives an image and information related to the image, and determines an item of processing related to the image based on at least one of the image and the information related to the image.

2. The image processing device according to claim 1, wherein the image is an image generated based on information related to the image.

3. The image processing device according to claim 1, wherein the information about the image is text generated from the image.

4. An image processing device as described in claim 1, wherein the processor receives information relating to an object that the user wishes to draw on the image as information related to the image, and determines the processing item based on at least one of the information about the image and the received information about the object.

5. The image processing device according to claim 4, wherein the processor receives information relating to a drawing area of ​​the object in the image, and determines the processing item using the information relating to the drawing area.

6. The image processing device according to claim 1, wherein the processor determines, for each of the processing items, the amount of correction for the processing item based on at least one of the image and information related to the image.

7. The image processing device according to claim 6, wherein said processor determines the amount of correction for said processing item as information described in text.

8. The image processing device according to claim 6, wherein the amount of correction of the processing item represents at least one of the form and state of a concept.

9. The image processing device according to claim 6, wherein the processor presents to the user an interface that accepts a user's designation regarding the amount of correction for the processing item, and executes correction processing for the image based on the amount of correction accepted from the user.

10. An image processing method comprising: a step of receiving, by a processor, an image and information related to the image; and a step of determining, by the processor, an item of processing related to the image based on at least one of the image and the information related to the image.

11. A program for causing a computer to execute each step included in the image processing method according to claim 10.

12. A computer-readable recording medium having recorded thereon a program for causing a computer to execute each step included in the image processing method according to claim 10.

Citation Information

Patent Citations

  • Query correction support system, search system, and program

    JP2021068063A