Image creation system
The image generation system addresses the challenge of adding new image objects by using a system with detection and insertion units to automatically generate and add user-desired images, ensuring ethical and legal compliance.
Patent Information
- Application Number
- JP2024082127
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-20
- Publication Date
- 2025-12-03
AI Technical Summary
Existing automatic coloring devices struggle to allow users to easily add new image objects to a target image.
An image generation system that includes a target image acquisition unit, command detection unit, user prompt conversion unit, and processing execution unit to automatically detect and insert new image objects based on user prompts and keywords using machine-learned models.
Enables the automatic generation and insertion of new image objects desired by the user into a target image, ensuring ethical acceptability and compliance with copyright and rights clearance.
Smart Images

Figure 2025175835000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image generation system. [Background technology]
[0002] An automatic coloring device colors a line drawing based on hint information, and the hint information is information that specifies colors using dots, line segments, etc. (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2018 / 203374 Summary of the Invention [Problem to be solved by the invention]
[0004] However, although the above-mentioned automatic coloring device can color objects in a target image, it is difficult for the user to add a new image object desired by the user.
[0005] The present invention has been made in view of the above-mentioned problems, and has as its object to provide an image generation system that can automatically generate and add a new image object desired by a user to a target image. [Means for solving the problem]
[0006] The image generation system of the present invention comprises a target image acquisition unit that acquires an original image of a document as a target image, a command detection unit that (a) detects additional edit commands written on the document in the target image and (b) identifies a user prompt and a processing target area in the target image corresponding to the edit command, a user prompt conversion unit that acquires keywords related to the user prompt and adds the keywords to the user prompt, and a processing execution unit that, for the detected edit command, (a) acquires, in an image generation model, a generated image corresponding to the user prompt with the keywords added and (b) inserts the generated image into the processing target area. [Effects of the Invention]
[0007] The present invention has been made in view of the above-mentioned problems, and has as its object to provide an image generation system that can automatically generate and add a new image object desired by a user to a target image.
[0008] The above and other objects, features and advantages of the present invention will become more apparent from the following detailed description taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a block diagram showing the configuration of an image generation system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram showing an example of a target image. [Figure 3] FIG. 3 is a diagram showing an example of a generated image. [Figure 4] FIG. 4 is a diagram showing an example of a target image into which the generated image shown in FIG. 3 has been inserted. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0011] Fig. 1 is a block diagram showing the configuration of an image generation system according to an embodiment of the present invention. The image generation system shown in Fig. 1 is an information processing device such as a personal computer, or an electronic device such as a digital camera or an image forming device (scanner, multifunction peripheral, etc.), and includes an arithmetic processing device 1, a storage device 2, a communication device 3, a display device 4, an input device 5, an internal device 6, etc.
[0012] The arithmetic processing device 1 includes a computer, which executes programs to function as various processing units. Specifically, the computer includes a CPU (Central Processing Unit), ROM (Read Only Memory), RAM (Random Access Memory), etc., and functions as a predetermined processing unit by loading a program stored in the ROM or storage device 2 into the RAM and executing it on the CPU. The arithmetic processing device 1 may also include an ASIC (Application Specific Integrated Circuit) that functions as a specific processing unit.
[0013] The storage device 2 is a non-volatile storage device such as a flash memory, and stores programs and data necessary for the processes described below. The storage device 2 stores setting data 2a. The setting data 2a includes data on the correspondence between revision commands and the processes to be executed.
[0014] The communication device 3 is a device that performs data communication with external devices, such as a network interface or a peripheral device interface. The display device 4 is a device that displays various information to the user, such as a display panel such as a liquid crystal display. The input device 5 is a device that detects user operations, such as a keyboard or a touch panel.
[0015] The internal device 6 is a device that executes a predetermined function. For example, if the image generation system is an image forming device such as a multifunction peripheral, the internal device 6 includes an image reading device that optically reads an original image from an original, a printing device that prints an image on printing paper, and the like.
[0016] Here, the processing device 1 operates as the above-mentioned processing units, including a target image acquisition unit 11, a command detection unit 12, a user prompt conversion unit 13, a processing execution unit 14, a generated image approval / disapproval determination unit 15, a rights clearance determination unit 16, and an output processing unit 17.
[0017] The target image acquisition unit 11 acquires (image data of) a document image of a certain document as a target image from the storage device 2, communication device 3, internal device 6, etc., and stores it in RAM, etc. For example, this document is a printout output from a printing device, and this document image is read from the document by an image reading device. This document is, for example, a slip.
[0018] The command detection unit 12 (a) detects a revision command written additionally to the document in the target image, and (b) identifies a user prompt and a processing target area corresponding to the revision command in the target image.
[0019] For example, the command detection unit 12 performs character recognition processing on the target image, detects each character string (text data) in the target image and identifies its position, identifies a character string from the detected character strings that is registered as an add command, identifies a figure of a predetermined shape (here, a rectangular frame) adjacent to the identified character string as an area designation area, and identifies a character string adjacent to the identified character string as a user prompt.
[0020] 2 is a diagram showing an example of a target image. For example, in the case of target image 101 shown in FIG. 2, character recognition processing is performed on target image 101, and character strings such as "Christmas," "SALE," "upto," "50%," "off," "GENW," "Anime Santa clause," "HURRY UP! ONLY," and "15-24 DEC" are detected, and the positions of these character strings are identified. Of these, a character string 111 called "GENW," which is registered as an add command, is identified, a character string 112 called "Anime Santa clause" adjacent to character string 111 is identified as a user prompt, and a rectangular frame 113 adjacent to character string 111 is identified as a processing target area.
[0021] Here, the retouching command "GENW" specifies the process of generating an image using an image generation model based on a user prompt and inserting the generated image (after scaling or cropping as necessary) into the area to be processed.
[0022] The user prompt conversion unit 13 acquires keywords related to the identified user prompt and adds the acquired keywords to the user prompt. The user prompt conversion unit 13 may add all of the acquired keywords to the user prompt, or may add only those keywords selected by the user from among the acquired keywords to the user prompt. In this case, for example, the acquired keywords are displayed on the display device 4, and the user performs a user operation on the input device 5 to select a keyword, and the user prompt conversion unit 13 characterizes the keyword selected by the user based on the user operation.
[0023] Specifically, the user prompt conversion unit 13 acquires keywords related to the user prompt (hereinafter referred to as related keywords) using a large-scale language model that has been machine-learned, such as PALM or Chat GPT. For example, the user prompt conversion unit 13 uses the communication device 3 to access a server of a large-scale language model, such as PALM or Chat GPT, inputs the user prompt and a prompt with an instructional phrase (e.g., "Tell me keywords related to") added to the large-scale language model, and acquires the related keywords from the large-scale language model.
[0024] For example, from the user prompt "Anime Santa claus" obtained from the target image 101 shown in Figure 2, related keywords "Christmas, present, snow" are obtained using a large-scale language model, and the user prompt is converted to "Anime, Santa claus, Christmas, present, snow."
[0025] For the identified retouching command, the processing execution unit 14 (a) obtains a generated image corresponding to the user prompt with the related keywords added using a machine-learned image generation model such as Stable Diffusion, and (b) inserts the generated image into the processing target area.
[0026] The image generation model may be built into the processing execution unit 14 or may be implemented on an external server. When the image generation model is implemented on an external server, the processing execution unit 14 uses the communication device 3 to access the server of the image generation model and acquire the generated image.
[0027] Fig. 3 is a diagram showing an example of a generated image. Fig. 4 is a diagram showing an example of a target image into which the generated image shown in Fig. 3 has been inserted. When a generated image 121, for example, as shown in Fig. 3 is acquired in response to a user prompt, the processing execution unit 14 erases the retouch command and user prompt (character strings 111, 112) and the processing target area (rectangular frame 113) in the target image 101, as shown in Fig. 4, for example, and then pastes the generated image 121 at the position of the processing target area.
[0028] In addition, if the image generation model specifies the language of the prompt as a predetermined language (e.g., English) and the user prompt is not in that predetermined language, the user prompt conversion unit 13 may translate the user prompt into that predetermined language, and the processing execution unit 14 may obtain a generated image corresponding to the translated user prompt.
[0029] Generated image acceptability determination unit 15 determines whether the generated image is ethically acceptable or not.
[0030] For example, generated image acceptability determination unit 15 uses communication device 3 to obtain the degree of likelihood that the generated image contains inappropriate content in a specific category (adult, spoof, medical, violence, and sexual depiction) using Google SafeSearch, and determines whether the generated image is ethically acceptable based on that degree. If it is determined that the generated image is ethically unacceptable, processing execution unit 14 discards the generated image.
[0031] The generated image determination unit 15 may be provided as needed, or may not be provided.
[0032] The rights clearance determination unit 16 determines whether or not a generated image generated based on a user prompt (text) to which a keyword has been added may infringe at least one of copyright, trademark right, and portrait right.
[0033] For example, rights clearance determination unit 16 uses communication device 3 to perform a web image search for the user prompt using Google's Webdetection, obtains an image as a result of the image search, and if the similarity between the obtained image and the generated image exceeds a predetermined threshold, determines that the generated image may infringe at least one of copyright, trademark, and portrait rights. If it is determined that the generated image is not likely to infringe any of copyright, trademark, and portrait rights, processing execution unit 14 discards the generated image.
[0034] The rights clearance determining unit 16 may be provided as needed, but may not be provided.
[0035] The output processing unit 17 outputs the target image after the above-mentioned processing (printing, data transmission, saving in the storage device 2, etc.).
[0036] Next, the operation of the image generation system will be described.
[0037] When the target image acquisition unit 11 acquires a target image in accordance with user operations, etc., the command detection unit 12 (a) detects any additional commands written in the target image in addition to the original, and (b) identifies the user prompt and processing target area in the target image corresponding to the additional commands.
[0038] Next, the user prompt conversion unit 13 acquires keywords related to the identified user prompt, and adds the acquired keywords to the user prompt.
[0039] Then, for the identified revision command, the process execution unit 14 acquires, in the image generation model, a generated image corresponding to the user prompt to which the related keyword has been added.
[0040] Here, the generated image acceptability determination unit 15 determines whether the generated image is ethically acceptable or not, The rights clearance determination unit 16 determines whether or not a generated image generated based on a user prompt (text) to which a keyword has been added may infringe at least one of copyright, trademark right, and portrait right.
[0041] If the generated image is determined to be ethically unacceptable or if the generated image is determined to be likely to infringe on at least one of copyright, trademark right, and portrait right, the processing execution unit 14 discards the generated image and displays an error message on the display device 4.
[0042] On the other hand, if the generated image is determined to be ethically acceptable and not likely to infringe on at least one of copyright, trademark, and portrait rights, the processing execution unit 14 inserts the generated image into the processing target area. Thereafter, the output processing unit 17 outputs the target image into which the generated image has been inserted.
[0043] As described above, according to the embodiment, the command detection unit 12 (a) detects a retouch command added to the document in the target image, and (b) identifies a user prompt and a processing target area in the target image corresponding to the retouch command. The user prompt conversion unit 13 acquires a keyword related to the user prompt and adds the keyword to the user prompt. The processing execution unit 14 includes a processing execution unit that (a) acquires, in an image generation model, a generated image corresponding to the user prompt to which the keyword has been added, and (b) inserts the generated image into the processing target area, for the detected retouch command.
[0044] As a result, a new image object desired by the user (that is, the generated image described above) is automatically and appropriately generated and added to the target image.
[0045] It should be noted that various changes and modifications to the above-described embodiments will be apparent to those skilled in the art. Such changes and modifications may be made without departing from the spirit and scope of the subject matter and without diminishing its intended advantages. In other words, it is intended that such changes and modifications be included within the scope of the claims.
[0046] For example, in FIG. 2, one retouching command is present in the target image, but multiple retouching commands (and corresponding multiple user prompts and area designations) may be present in the target image. [Industrial Applicability]
[0047] The present invention is applicable, for example, to a system for inserting a generated image into a target image. [Explanation of symbols]
[0048] 11 Target image acquisition unit 12 Command detection section 13 User Prompt Conversion Section 14 Processing execution unit 15. Generated image validity determination unit 16 Rights Clearance Determination Division
Claims
1. a target image acquisition unit that acquires a document image of a document as a target image; (a) detecting a touch-up command written in the target image in addition to the document; and (b) identifying a user prompt and a processing target area corresponding to the touch-up command in the target image. a user prompt conversion unit that obtains keywords related to the user prompt and adds the keywords to the user prompt; a processing execution unit that, for the detected touch-up command, (a) obtains, in an image generation model, a generated image corresponding to the user prompt to which the keyword has been added, and (b) inserts the generated image into the processing target area; An image generation system comprising:
2. 2. The image generation system according to claim 1, wherein the user prompt conversion unit acquires the keywords from the user prompt using a large-scale language model trained by machine learning.
3. a rights clearance determination unit that determines whether or not the generated image may infringe on at least one of copyright, trademark, and portrait rights based on the user prompt to which the keyword has been added; the processing execution unit discards the generated image when it is determined that the generated image may infringe at least one of copyright, trademark right, and portrait right; 2. The image generating system according to claim 1, wherein:
4. a generated image acceptability determination unit that determines whether the generated image is ethically acceptable; the processing execution unit discards the generated image if it is determined that the generated image is ethically unacceptable; 2. The image generating system according to claim 1, wherein:
Citation Information
Patent Citations
Line drawing automatic coloring program, line drawing automatic coloring device, and program for graphical user interface
WO2018203374A1