Information processing device and image generation method
The information processing device simplifies image generation by allowing separate input and editing of sketch and text layers, addressing the inconvenience of manual position explanations in existing AIs, enhancing user experience.
Patent Information
- Application Number
- JP2024199139
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-11-14
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2044-11-14
AI Technical Summary
Existing image generation AIs require users to accurately explain the correspondence between the position of each part of a sketch and a prompt in text, which is cumbersome and inconvenient.
An information processing device that includes a data acquisition unit for acquiring sketch images and text information, a prompt generation unit for generating prompts using a text generation model, and an image generation unit for generating images based on the sketch and prompts, allowing separate input and editing of sketch and text layers.
Improves user convenience by enabling easy addition of annotations directly on sketches without needing detailed written explanations, simplifying the image generation process.
Smart Images

Figure 0007739575000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device and an image generating method. [Background technology]
[0002] In recent years, image generation AI (Artificial Intelligence) that generates images from images such as hand-drawn sketches (Image to Image) has become known. There is also image generation AI that generates images by inputting prompts created in text (Text to Image) (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2024-120131 Summary of the Invention [Problem to be solved by the invention]
[0004] Some of the image generation AIs mentioned above can generate images by inputting both an image such as a hand-drawn sketch and a prompt, and can specify additional information for each part of the sketch in the prompt so that the image the user wants to create is generated. However, the correspondence between the position of each part of the sketch and the prompt needs to be explained in the text written in the prompt, and it is not easy and cumbersome for the user to express it accurately (and in a way that can be conveyed to the image generation AI).
[0005] The present invention has been made in consideration of the above-mentioned circumstances, and one of its objects is to provide an information processing device and an image generation method that improve convenience when using image generation AI. [Means for solving the problem]
[0006] The present invention has been made to solve the above-mentioned problems, and an information processing device according to a first aspect of the present invention comprises a data acquisition unit that acquires data for text information input and superimposed on a sketch image and the sketch image, and a generated image acquisition unit that acquires a generated image generated by the image generation model by inputting input data for image generation based on the sketch image and the text information into the image generation model.
[0007] The information processing device may further include a prompt generation instruction unit that instructs the generation of a prompt as an input sentence to an image generation model by inputting input data for prompt generation based on the sketch image and the text information acquired by the data acquisition unit into a text generation model, and an image generation instruction unit that inputs the prompt generated by the text generation model and the sketch image into the image generation model as input data for image generation.
[0008] The information processing device may further include a sketch input unit that accepts input of the sketch image, a text input unit that accepts text input of the text information by overlaying it on the sketch image accepted by the sketch input unit, and an output unit that outputs the sketch image and the text information as separate data.
[0009] In the information processing device, the sketch input unit may accept input for a first layer, and the text input unit may accept text input for a second layer that is overlaid on the first layer.
[0010] The information processing device may further include a prompt editing unit that accepts text input as a prompt to be input to the image generation model and outputs the prompt input by the user as a prompt to be added to the prompt generated by the text generation model, and the image generation instruction unit may input the prompt output by the prompt editing unit to the image generation model as a prompt to be input to the image generation model.
[0011] The information processing device may further include a prompt editing unit that acquires a prompt generated by the text generation model and presents it in an editable manner, accepts correction input for the presented prompt, and modifies and outputs the prompt generated by the text generation model based on the correction input, and the image generation instruction unit may input the prompt output by the prompt editing unit to the image generation model as a prompt to be input to the image generation model.
[0012] In the information processing device, the text generation model may be a VLM (Vision and Language Model).
[0013] In addition, an image generation method in an information processing device according to a second aspect of the present invention includes a step in which a data acquisition unit acquires data for text information input and superimposed on a sketch image and the sketch image, and a step in which a generated image acquisition unit acquires a generated image generated by the image generation model by inputting input data for image generation based on the sketch image and the text information into the image generation model. [Effects of the Invention]
[0014] According to the above-described aspects of the present invention, it is possible to improve the convenience when using an image generation AI. [Brief explanation of the drawings]
[0015] [Figure 1] FIG. 1 is a diagram showing an example of the configuration of an information processing system according to an embodiment. [Figure 2] FIG. 1 is a block diagram showing an example of a hardware configuration of an information processing apparatus according to an embodiment. [Figure 3] FIG. 10 is a diagram showing an example of an image generation UI according to the embodiment. [Figure 4] 10A and 10B are diagrams showing an input example in an input area and an output example in an output area according to the embodiment. [Figure 5] FIG. 2 is a block diagram showing an example of a functional configuration related to image generation processing according to the embodiment. [Figure 6] FIG. 2 is a block diagram showing a detailed example of the functional configuration of a UI processing unit according to the embodiment. [Figure 7] FIG. 2 is a block diagram showing a detailed example of the functional configuration of an image generation processing unit according to the embodiment. [Figure 8] 10 is a flowchart showing an example of an image generation process according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0016] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. [System Configuration] 1 is a diagram showing an example of the configuration of an information processing system according to this embodiment. The information processing system SYS includes an information processing device 10, a text generation server 20, and an image generation server 30.
[0017] The information processing device 10 is, for example, a clamshell (notebook) type PC (personal computer). The information processing device 10 has a first housing 10A and a second housing 10B, each of which is substantially rectangular and plate-shaped (for example, flat), joined (connected) via a hinge mechanism so that the first housing 10A and the second housing 10B can be opened and closed. The information processing device 10 also has a display 150 extending from the first housing 10A to the second housing 10B.
[0018] The display 150 is a flexible display that can be bent when the first housing 10A and the second housing 10B are opened or closed. An organic electroluminescence (EL) display or the like is used as the flexible display. For example, the display 150 can be used not only in a single-screen mode in which the entire screen area is a single screen, but also in a dual-screen mode in which the screen is divided into two screens: a first screen area 150A on the first housing 10A side and a second screen area 150B on the second housing 10B side. For example, when the information processing device 10 is used while placed on a desk, the second screen area 150B on the second housing 10B side becomes a substantially horizontal screen parallel to the surface of the desk.
[0019] Furthermore, a touch sensor is provided on the top (surface) of the screen of the display 150. The information processing device 10 can detect a touch operation on the screen area of the display 150. By opening the information processing device 10, the user can view the displays 150 provided on the inner surfaces of the first housing 10A and the second housing 10B and perform touch operations on the displays 150, thereby enabling the user to use the information processing device 10.
[0020] The information processing device 10 can also be connected to an external keyboard 50 via a wired or wireless connection to receive keyboard input.
[0021] The text generation server 20 includes a text generation model 21, and generates and outputs text from an input image (or image and text). The text generation model 21 is a large-scale visual language model known as a VLM (Vision and Language Model). In other words, the text generation server 20 functions as a text generation AI (Artificial Intelligence) that generates and outputs text from an image (or image and text, etc.) using the text generation model 21 as a large-scale visual language model.
[0022] The image generation server 30 includes an image generation model 31. The image generation model 31 is an image generation model trained using, for example, a diffusion model, and generates an image from an input image or prompt (text). In other words, the image generation server 30 functions as an image generation AI that uses the image generation model 31 to generate and output an image from an image and a prompt (text).
[0023] The text generation server 20 and the image generation server 30 may be configured as a single server, or may be configured as a plurality of distributed servers.
[0024] In this embodiment, a text generation AI that generates text using the text generation model 21 and an image generation AI that generates images using the image generation model 31 are provided on a server, and an example configuration is described in which the information processing device 10 uses the functions of the text generation AI and the image generation AI by communicating with the server, but the functions of the text generation AI and the image generation AI may also be downloaded to the information processing device 10 and used locally.
[0025] [Hardware configuration of information processing device 10] The specific configuration of the information processing device 10 will be described below. 2 is a block diagram showing an example of the hardware configuration of an information processing device 10 according to this embodiment. The information processing device 10 includes a communication unit 11, a RAM (Random Access Memory) 12, a storage unit 13, a speaker 14, a display unit 15, a camera 16, and a control unit 18. These units are connected to each other so as to be able to communicate with each other via a bus or the like.
[0026] The communication unit 11 includes, for example, a plurality of Ethernet (registered trademark) ports, a plurality of digital input / output ports such as USB (Universal Serial Bus), and a communication device for wireless communication such as Bluetooth (registered trademark) or Wi-Fi (registered trademark).
[0027] For example, the communication unit 11 can communicate with an external keyboard 50, a touch pen, a mouse, a touch pad, or the like shown in Fig. 1 using Bluetooth (registered trademark). In addition, the communication unit 11 can communicate with the text generation server 20 and the image generation server 30 shown in Fig. 1 by connecting to the Internet via wireless communication such as Wi-Fi (registered trademark) or wired communication such as Ethernet (registered trademark).
[0028] The RAM 12 is a volatile memory into which programs and data for processing executed by the control unit 18 are expanded and into which various data are stored or erased as appropriate. Note that, because the RAM 12 is a volatile memory, it does not retain data when power supply to the RAM 12 is stopped. Data that needs to be retained when power supply to the RAM 12 is stopped is transferred to the storage unit 13.
[0029] The storage unit 13 is a storage device configured to include one or more of an SSD (Solid State Drive), an HDD (Hard Disk Drive), a ROM (Residual Only Memory), a Flash-ROM, etc. For example, the storage unit 13 stores programs and setting data for a BIOS (Basic Input Output System), an OS (Operating System) and programs for applications that run on the OS, and various data used by applications. The speaker 14 outputs electronic sounds, voices, and the like.
[0030] The display unit 15 includes a display 150 and a touch sensor 155. As described above, the display 150 is a flexible display that can be bent in accordance with the opening and closing of the first housing 10A and the second housing 10B. Under the control of the control unit 18, the display 150 displays the desktop screen of the OS, windows of running applications, and the like.
[0031] Touch sensor 155 is provided on the screen of display 150 and detects touch operations on the screen. Touch operations include, for example, tap operations, slide operations, flick operations, swipe operations, pinch operations, etc. Furthermore, a finger or a touch pen can be used as an operating means for performing the touch operations.
[0032] The camera 16 includes a lens, an imaging element, etc. The camera 16 captures an image (a still image or a moving image) according to the control of the control unit 18, and outputs data of the captured image.
[0033] Control unit 18 includes processors such as a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), and a microcomputer, and realizes various functions by executing programs (various programs such as a BIOS, an OS, and applications that run on the OS) stored in storage unit 13. For example, control unit 18 controls the screen mode of display 150, controls display in the screen area, and detects touch operations on touch sensor 155.
[0034] Next, a description will be given of image generation processing using the image generation model 31 (image generation AI) in the information processing device 10. First, a UI (User Interface) will be described with reference to FIG.
[0035] [Image generation UI] 3 is a diagram showing an example of an image generation UI according to this embodiment. In the example shown, information processing device 10 controls display 150 in a dual-screen mode with first screen area 150A and second screen area 150B. In second screen area 150B, a window of application AP1 (hereinafter referred to as "image generation application AP1") having an image generation function that generates an image using image generation model 31 (image generation AI) is displayed. The window of image generation application AP1 has, as UI, an input area R1 for inputting information about the image to be generated and an output area R2 for displaying the generated image.
[0036] The user can hand-draw a sketch in the input area R1 and can also input annotations as text overlaid on the sketch as additional information. Additional information can be information that supplements each part of the sketch or information about objects not included in the sketch. Annotations can be input anywhere on the sketch.
[0037] When the information processing device 10 is placed on a desk and used, the second screen area 150B is a nearly horizontal screen, and therefore, the presence of the input area R1 on the second screen area 150B side makes it easy for the user to input data.
[0038] When inputting a rough sketch by freehand, the user can mainly use a touch pen, but can also use a finger without a touch pen, or can use an external mouse (not shown), or can use an external touch pad (not shown). When inputting an annotation as text, the user can use, for example, a keyboard (see FIG. 1), but can also use an on-screen software keyboard or handwritten input using a touch pen or the like.
[0039] The image generation UI shown in FIG. 3 is an example, and the layout and size of the input area R1 and output area R2, the screen area in which they are displayed, and the like are not limited to this example.
[0040] 4A and 4B are diagrams showing an example of input in the input area R1 and an example of output in the output area R2 according to this embodiment. (A) of Fig. 4 shows an example of input to the input area R1. As mentioned above, the input area R1 allows for freehand input of a sketch and text input of annotations.
[0041] For example, the image generation application AP1 has multiple overlapping layers that use the range of input area R1 as an input area, and switches between layer 1, which accepts hand-drawn input of a sketch, and layer 2, which accepts text input of an annotation, to accept input for each. That is, the user selects layer 1 to hand-draw a sketch and inputs it, and selects layer 2 to input an annotation as text. This allows the user to, for example, select layer 1 to hand-draw a sketch, and then select layer 2 to input an annotation overlaid on the sketch image input in layer 1.
[0042] In the input example shown in (A) of Figure 4, a "tree," a "person," a "furniture with legs," and a "diagonal line" extending from the upper right to the lower left within input area R1 have been hand-drawn as sketches in layer 1. In layer 2, the following text has been input as annotations, superimposed on the sketch image input in layer 1: "palm tree" in the position of the "tree" in the sketch, "woman" in the position of the "person" in the sketch, and "reclining cot" in the position of the "furniture with legs" in the sketch, as well as "ocean" above and "beach" below the sketch separated by the "diagonal line."
[0043] The image generation application AP1 generates the sketch image data input to Layer 1 and the annotation data input to Layer 2 as separate data. Note that the image generation application AP1 may obtain the sketch image data from an image data file, rather than inputting it to Layer 1. In other words, the image generation application AP1 may read image data from a file and display part or all of the read image data as a sketch image on Layer 1. This allows the user to input annotations, for example, by superimposing them on the sketch image displayed on Layer 1.
[0044] The image generation application AP1 inputs input data based on the sketch image and annotations into the image generation model 31, thereby acquiring a generated image generated by the image generation model 31. The image generation application AP1 then displays the acquired generated image in the output area R2. FIG. 4B shows an example of a generated image displayed in the output area R2 as an output example of the output area R2. Note that the generated image shown in this figure is an illustration that imitates a generated image generated by the image generation model 31. In the example of the generated image shown in FIG. 4B, an image of a palm tree, a woman, and a reclining bed arranged on a sandy beach is generated in accordance with the input example shown in FIG. 4A.
[0045] 3, a window of an application AP2 different from the image generation application AP1 is displayed in the first screen area 150A. The generated image generated by the image generation application AP1 can be used in other applications, and can be displayed, for example, by copying and pasting it into the window of the other application AP2. The other application AP2 is an application that can insert or paste images, such as an application with a presentation creation function, an application with a document creation function, or an application with an image editing function.
[0046] [Image generation processing functional configuration] Next, the functional configuration of the image generation process realized by the control unit 18 executing the image generation application AP1 will be described. Fig. 5 is a block diagram showing an example of a functional configuration related to the image generation process according to this embodiment. This diagram shows a front-end configuration related to the UI of the image generation application AP1 described with reference to Fig. 4, and a back-end configuration that generates images in accordance with the process.
[0047] The front-end configuration includes a UI processing unit 100 as a functional configuration based on the UI described with reference to Fig. 4. The UI processing unit 100 outputs the data of the sketch image and annotations input by the user to the input area R1 (see Fig. 3) to the back-end. The UI processing unit 100 also outputs the generated image acquired from the back-end to the output area R2.
[0048] The back-end configuration includes an image generation processing unit 110, a prompt generation unit 120, and an image generation unit .
[0049] The prompt generation unit 120 corresponds to the functional configuration of the text generation server 20 shown in Fig. 1, and generates and outputs a prompt (text) from an input image (or image and text) using a text generation model 21. The image generation unit 130 corresponds to the functional configuration of the image generation server 30 shown in Fig. 1, and uses an image generation model 31 to generate and output an image from an input image and prompt (text).
[0050] The image generation processing unit 110 connects the front-end UI processing unit 100 with the back-end prompt generation unit 120 and image generation unit 130, and performs image generation processing using the prompt generation unit 120 and image generation unit 130 based on data obtained from the UI processing unit 100.
[0051] The image generation processing unit 110 may be a functional configuration that is realized by the control unit 18 when the control unit 18 executes the image generation application AP1, or may be a functional configuration on the server side.
[0052] The configuration of the image generation processing according to this embodiment will be described in detail below with reference to Figures 5 to 7. Figure 6 is a block diagram showing a detailed example of the functional configuration of the UI processing unit 100. The UI processing unit 100 includes a sketch input unit 101, an annotation input unit 102 (an example of a text input unit), a data output unit 103, a generated image output unit 104, and a prompt editing unit 105.
[0053] The sketch input unit 101 accepts a freehand input of a sketch (sketch image) by a user. For example, the sketch input unit 101 accepts a freehand input of a sketch image by freehand input on layer 1 (see FIGS. 3 and 4) of input area R1.
[0054] The annotation input unit 102 receives annotation text input by the user, superimposed on the sketch image received by the sketch input unit 101. For example, the annotation input unit 102 receives annotation text input from text input to layer 2 (see FIGS. 3 and 4) superimposed on layer 1 of input area R1.
[0055] The data output unit 103 outputs the sketch image received by the sketch input unit 101 and the annotation received by the annotation input unit 102 as separate data.
[0056] Generated image output unit 104 acquires and outputs the generated image output from image generation unit 130 via image generation processing unit 110. For example, generated image output unit 104 displays the acquired generated image in output area R2 (see FIG. 3).
[0057] The prompt editing unit 105 accepts text input as a prompt from the user, and outputs the prompt input by the user as a prompt to be added to the prompt generated by the prompt generating unit 120 .
[0058] Furthermore, prompt editing unit 105 acquires the prompt generated by prompt generation unit 120, presents (displays) it in a modifiable manner, and accepts input for modifying the prompt. Then, prompt editing unit 105 modifies the prompt generated by prompt generation unit 120 based on the input for modification, and outputs the modified prompt.
[0059] 7 is a block diagram showing a detailed example of the functional configuration of image generation processing unit 110. Image generation processing unit 110 includes data acquisition unit 111, prompt generation instruction unit 112, prompt acquisition unit 113, image generation instruction unit 114, and generated image acquisition unit 115.
[0060] The data acquisition unit 111 acquires, from the UI processing unit 100, data of the sketch image and data of the annotation that has been input and superimposed on the sketch image.
[0061] The prompt generation instruction unit 112 instructs the prompt generation unit 120 to generate a prompt as an input sentence to the image generation unit 130 (image generation model 31) by inputting input data for prompt generation based on the sketch image and annotation acquired by the data acquisition unit 111 to the prompt generation unit 120. For example, the input data for prompt generation is image data obtained by integrating a sketch image layer and an annotation layer.
[0062] Here, the prompt for text generation inputted to the prompt generation unit 120 (text generation model 21) by the prompt generation instruction unit 112 when inputting the input data for prompt generation to the prompt generation unit 120 is an instruction for causing the text generation model 21 to generate a prompt (prompt for image generation) to be inputted to the image generation unit 130 (image generation model 31). For example, the prompt generation instruction unit 112 inputs a prompt for text generation, such as "generate a prompt including position information of an object in a predetermined format," to the prompt generation unit 120 (text generation model 21).
[0063] Note that a plurality of prompts for generating text may be prepared in advance according to the image to be generated using the image generating unit 130 (image generation model 31), and a selection may be made from among them.
[0064] When the prompt generation input data is input, the prompt generation unit 120 uses the text generation model 21 to generate and output a prompt (prompt for image generation) to be input to the image generation unit 130.
[0065] The prompt acquisition unit 113 acquires the prompt generated by the prompt generation unit 120 from the prompt generation unit 120. The prompt acquisition unit 113 also outputs the acquired prompt to the UI processing unit 100. When the user adds or modifies a prompt generated by the prompt generation unit 120 in the UI processing unit 100, the prompt acquisition unit 113 acquires the prompt output by the prompt editing unit 105 of the UI processing unit 100.
[0066] If the prompt generated by the prompt generation unit 120 has not been added to or modified, the prompt acquisition unit 113 outputs the prompt acquired from the prompt generation unit 120 to the image generation instruction unit 114. On the other hand, if the prompt generated by the prompt generation unit 120 has been added to or modified, the prompt acquisition unit 113 outputs the prompt acquired from the prompt editing unit 105 of the UI processing unit 100 to the image generation instruction unit 114.
[0067] The image generation instruction unit 114 instructs the image generation unit 130 to generate an image by inputting the prompt acquired from the prompt acquisition unit 113 and the sketch image acquired by the data acquisition unit 111 to the image generation unit 130 as input data for image generation.
[0068] That is, if there is no addition or modification to the prompt generated by the prompt generation unit 120, the image generation instruction unit 114 inputs the prompt generated by the prompt generation unit 120 and the sketch image acquired by the data acquisition unit 111 to the image generation unit 130. On the other hand, if there is an addition or modification to the prompt generated by the prompt generation unit 120, the image generation instruction unit 114 inputs the prompt added to or modified in the prompt generated by the prompt generation unit 120 and the sketch image acquired by the data acquisition unit 111 to the image generation unit 130.
[0069] When the prompt and the sketch image are input, the image generation unit 130 generates an image using the image generation model 31 and outputs the generated image.
[0070] The generated image acquisition unit 115 acquires the generated image generated by the image generation unit 130. The generated image acquisition unit 115 also outputs the acquired generated image to the UI processing unit 100.
[0071] [Image generation process] Next, the operation of the image generation processing executed by the image generation processing unit 110 will be described. FIG. 8 is a flowchart showing an example of image generation processing according to this embodiment.
[0072] (Step S101) The image generation processing unit 110 acquires data of a sketch image and data of annotations input and superimposed on the sketch image from the UI processing unit 100 based on the user's freehand input and text input in the input area R1 shown in Fig. 3. Then, the process proceeds to step S103.
[0073] (Step S103) The image generation processing unit 110 inputs input data for prompt generation based on the sketch image and annotation acquired in step S101 to the prompt generation unit 120, thereby instructing the prompt generation unit 120 to generate a prompt to be input to the image generation unit 130 (image generation model). For example, the image generation processing unit 110 inputs image data obtained by integrating the sketch image layer and the annotation layer to the prompt generation unit 120 as input data for prompt generation. Then, the process proceeds to step S105.
[0074] (Step S105) Image generation processing unit 110 acquires the prompt generated by prompt generation unit 120. Furthermore, image generation processing unit 110 outputs the acquired prompt to UI processing unit 100, and when the user adds or modifies a prompt in UI processing unit 100, acquires the added or modified prompt from UI processing unit 100.
[0075] (Step S107) The image generation processing unit 110 inputs the sketch image acquired in step S101 and the prompt acquired in step S105 to the image generation unit 130 as input data for image generation, thereby instructing the image generation unit 130 to generate an image.
[0076] (Step S109) Image generation processing unit 110 acquires the generated image generated by image generation unit 130, and outputs the acquired generated image to UI processing unit 100. UI processing unit 100 displays the acquired generated image in output area R2 (see FIG. 3).
[0077] As described above, the information processing device 10 according to this embodiment acquires data on the sketch image and the annotation (an example of text information) input and superimposed on the sketch image, and inputs input data for image generation based on the sketch image and the annotation to the image generation model 31. Then, the information processing device 10 acquires the generated image generated by the image generation model 31.
[0078] As a result, the information processing device 10 can easily add information that supplements each part of the sketch or information about objects that are not in the sketch, and generate an image by simply having the user input annotations by overlaying them on the sketch image. This eliminates the need to explain in writing the correspondence between the annotations and the positions of each part of the sketch, thereby improving convenience when using image generation AI.
[0079] For example, the information processing device 10 inputs input data for prompt generation based on the acquired sketch image and annotation to the text generation model 21, thereby instructing the image generation model 31 to generate a prompt as an input sentence. Then, the information processing device 10 inputs the prompt and sketch image generated by the text generation model 21 to the image generation model 31 as input data for image generation. Here, the input data for prompt generation is, for example, image data obtained by integrating a sketch image layer and an annotation layer.
[0080] As a result, the information processing device 10 can generate a prompt that associates the annotation with the position of each part of the sketch, and input it into the image generation model 31, simply by the user inputting an annotation by overlaying it on the sketch image. This eliminates the need to explain in writing the association between the annotation and the position of each part of the sketch, thereby improving convenience when using image generation AI.
[0081] The information processing device 10 also accepts input of a sketch image and input of annotation text to be superimposed on the accepted sketch image, and then outputs the sketch image and the annotation as separate data.
[0082] As a result, the information processing device 10 treats the sketch image and the annotation as independent data, so there is no need to separate the sketch image and the annotation later. Also, because the information processing device 10 treats the sketch image and the annotation as independent data, it is possible to save only the sketch image or only the annotation. Also, because the information processing device 10 treats the sketch image and the annotation as independent data, it is possible to independently perform the functions of "Undo," which returns to the state up to the immediately previous operation, and "Redo," which returns from the undone state to the state before the undone.
[0083] For example, the information processing device 10 receives an input for Layer 1 (first layer) and a text input for Layer 2 (second layer) that is overlaid on Layer 1.
[0084] As a result, the information processing device 10 can manage the sketch image and the annotations on separate layers, thereby handling the sketch image and the annotations as independent data.
[0085] Furthermore, the information processing device 10 accepts text input as a prompt to be input to the image generation model 31, and outputs the prompt input by the user as a prompt to be added to the prompt generated by the text generation model 21. Then, the information processing device 10 inputs the output prompt to the image generation model 31 as a prompt to be input to the image generation model 31.
[0086] This allows the information processing device 10 to allow the user to add a prompt to the prompt generated from the sketch image and annotation and input it to the image generation model 31, thereby allowing the user to adjust the content of the prompt.
[0087] Furthermore, the information processing device 10 acquires a prompt generated by the text generation model 21, presents it in a modifiable manner, accepts correction input for the presented prompt, and based on the correction input, corrects and outputs the prompt generated by the text generation model 21. Then, the information processing device 10 inputs the output prompt to the image generation model 31 as a prompt to be input to the image generation model 31.
[0088] This allows the information processing device 10 to allow the user to modify the prompt generated from the sketch image and annotation and input it to the image generation model 31, thereby allowing the user to adjust the content of the prompt.
[0089] Furthermore, the information processing device 10 generates a prompt using a VLM (Vision and Language Model) as the text generation model 21.
[0090] This allows the information processing device 10 to automatically generate a prompt using text generation AI in which the input annotations are associated with the positions of each part of the sketch, simply by the user inputting annotations by overlaying them on the sketch image.
[0091] In addition, the image generation method in the information processing device 10 includes a step in which the control unit 18 (image generation processing unit 110) acquires data for the sketch image and an annotation (an example of text information) that has been input and superimposed on the sketch image, and a step in which the control unit 18 acquires a generated image generated by the image generation model 31 by inputting input data for image generation based on the sketch image and the annotation into the image generation model 31.
[0092] As a result, the image generation method in the information processing device 10 allows the user to easily add information that supplements each part of the sketch or information about objects that are not in the sketch by simply inputting annotations overlaid on the sketch image, thereby generating an image. This eliminates the need to explain in writing the correspondence between the annotations and the positions of each part of the sketch, thereby improving convenience when using image generation AI.
[0093] Although the embodiments of the present invention have been described above in detail with reference to the drawings, the specific configurations are not limited to the above-described embodiments, and the present invention also includes designs that do not deviate from the gist of the present invention. For example, the configurations described in the above-described embodiments can be combined in any manner.
[0094] In the above-described embodiment, an example was described in which image data obtained by integrating a sketch image layer and an annotation layer was input to the prompt generation unit 120 (text generation model 21) as input data for generating a prompt. However, a prompt (text information) may also be generated by inputting sketch image data and annotation (annotation image) data without integrating the sketch image layer and the annotation layer.
[0095] In the above-described embodiment, an example was described in which the user can edit the prompt (prompt for image generation) generated by the text generation model 21, but the text generation prompt input to the text generation model 21 may also be configured to be editable by the user.
[0096] In the above-described embodiment, the display 150 is a single display provided across the first housing 10A and the second housing 10B, but this is not limiting. For example, the information processing device 10 may be provided with two displays in total, one on each of the first housing 10A and the second housing 10B. Furthermore, the information processing device 10 may be provided with a display only on one of the first housing 10A and the second housing 10B (for example, only on the first housing 10A).
[0097] Furthermore, the information processing device 10 is not limited to a clamshell type (notebook type) PC, but may be, for example, a tablet type PC or a desktop type PC.
[0098] In the above-described embodiment, an example of a touch panel display in which an input unit (touch sensor) and a display unit (display) are integrated has been described, but a non-touch panel display that does not have an input unit (touch sensor) may also be used. In this case, a touch pad, a mouse, or the like may be used as an input device that accepts operations instead of the input unit (touch sensor).
[0099] Furthermore, for example, the image generation model 31 may also include a text generation function using the text generation model 21, and an image may be generated by inputting a sketch image and annotations to the image generation model 31.
[0100] The information processing device 10 described above includes an internal computer system. A program for implementing the functions of each component of the information processing device 10 may be recorded on a computer-readable recording medium, and the program may be loaded into a computer system and executed to perform processing in each component of the information processing device 10. Here, "loading a program recorded on a recording medium into a computer system and executing it" includes installing the program into a computer system. The term "computer system" here includes hardware such as an OS and peripheral devices. The term "computer system" may also include multiple computers connected via a network, including the Internet, a WAN, a LAN, a dedicated line, or other communication lines. The term "computer-readable recording medium" refers to portable media such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, or a storage device such as a hard disk built into a computer system. The recording medium storing the program may also be a non-transitory recording medium such as a CD-ROM.
[0101] The recording medium also includes internal or external recording media accessible from a distribution server for distributing the program. The program may be divided into multiple parts, downloaded at different times, and then combined by each component of the information processing device 10, or each divided program may be distributed by a different distribution server. Furthermore, the term "computer-readable recording medium" also includes a medium that stores a program for a certain period of time, such as volatile memory (RAM) within a computer system that serves as a server or client when a program is transmitted over a network. The program may also be a medium that realizes part of the above-described functions. Furthermore, the program may be a so-called differential file (differential program) that can realize the above-described functions in combination with a program already stored in the computer system.
[0102] Furthermore, some or all of the functions of the information processing device 10 in the above-described embodiment may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each function may be individually implemented as a processor, or some or all of the functions may be integrated into a processor. The integrated circuit method is not limited to LSI, and may be implemented using a dedicated circuit or a general-purpose processor. Furthermore, if an integrated circuit technology that can replace LSI emerges due to advances in semiconductor technology, an integrated circuit based on that technology may be used. [Explanation of symbols]
[0103] 10 Information processing device, 10A First housing, 10B Second housing, 20 Text generation server, 21 Text generation model, 30 Image generation server, 31 Image generation model, 11 Communication unit, 12 RAM, 13 Memory unit, 14 Speaker, 15 Display unit, 16 Camera, 150 Display, 150A First screen area, 150B Second screen area, 155 Touch sensor, 18 Control unit, 100 UI processing unit, 101 Sketch input unit, 102 Annotation input unit, 103 Data output unit, 104 Generated image output unit, 105 Prompt editing unit, 110 Image generation processing unit, 111 Data acquisition unit, 112 Prompt generation instruction unit, 113 Prompt acquisition unit, 114 Image generation instruction unit, 115 Generated image acquisition unit, 120 Prompt generation unit, 130 Image generation unit, SYS Information processing system
Claims
1. a data acquisition unit that acquires data of the text information inputted and superimposed on the sketch image and the sketch image; a generated image acquisition unit that acquires a generated image generated by an image generation model by inputting input data for image generation based on the sketch image and the text information into the image generation model; An information processing device comprising:
2. a prompt generation instruction unit that instructs the generation of a prompt as an input sentence to an image generation model by inputting input data for prompt generation based on the sketch image and the text information acquired by the data acquisition unit to a text generation model; an image generation instruction unit that inputs the prompt generated by the text generation model and the sketch image to the image generation model as input data for image generation; The information processing device according to claim 1 , comprising:
3. a sketch input unit that accepts input of the sketch image; a text input unit that receives text input of the text information superimposed on the sketch image received by the sketch input unit; an output unit that outputs the sketch image and the text information as separate data; The information processing device according to claim 1 , comprising:
4. the sketch input unit accepts input for a first layer; the text input unit accepts text input for a second layer overlaid on the first layer; The information processing device according to claim 3 .
5. a prompt editing unit that accepts a text input as a prompt to be input to the image generation model, and outputs the prompt input by the user as a prompt to be added to the prompt generated by the text generation model; Equipped with The image generation instruction unit inputting the prompt output by the prompt editing unit into the image generation model as a prompt to be input into the image generation model; The information processing device according to claim 2 .
6. a prompt editing unit that acquires a prompt generated by the text generation model, presents the prompt in an editable manner, accepts input for correction of the presented prompt, and corrects and outputs the prompt generated by the text generation model based on the input for correction; Equipped with The image generation instruction unit inputting the prompt output by the prompt editing unit into the image generation model as a prompt to be input into the image generation model; The information processing device according to claim 2 .
7. The text generation model generates prompts using a Vision and Language Model (VLM). The information processing device according to claim 2 .
8. An image generation method in an information processing device, comprising: a step in which a data acquisition unit acquires data of the text information inputted and superimposed on the sketch image and the sketch image; a generated image acquisition unit inputting input data for image generation based on the sketch image and the text information into an image generation model, thereby acquiring a generated image generated by the image generation model; An image generation method comprising:
Citation Information
Patent Citations
Image generation device, prompt creation support device, program and application program
JP2024120131A