Information processing device and image generation method
The information processing device enhances image generation AI by allowing users to input annotations directly on a background image, using a Vision and Language Model to generate images, thereby simplifying the image creation process and improving user convenience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- LENOVO (SINGAPORE) PTE LTD
- Filing Date
- 2024-11-14
- Publication Date
- 2026-05-26
AI Technical Summary
Existing image generation AI systems require users to accurately describe the association between the position of each part of a sketch and a prompt in text, which is cumbersome and difficult for users to convey effectively.
An information processing device that acquires data of text information superimposed on a background image and uses a Vision and Language Model to generate an image by integrating background image and annotation data, allowing users to input annotations directly on the sketch without needing to explicitly specify positional associations in text.
Improves the convenience of using image generation AI by enabling users to easily generate images by overlaying annotations on a background image, eliminating the need to explain positional associations in text, thus simplifying the image creation process.
Smart Images

Figure 2026086179000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus and an image generation method.
Background Art
[0002] In recent years, image generation AI (Artificial Intelligence) that generates an image from an image such as a hand-drawn sketch (Image to Image) is known. There is also image generation AI that generates an image by inputting a prompt created in text (Text to Image) (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Among the above-described image generation AI, there are those that can generate an image by inputting both an image such as a hand-drawn sketch and a prompt. For example, in order to generate an image that a user wants to create, additional information can be described and specified for each part of the sketch with a prompt. However, it is necessary to describe the association between the position of each part of the sketch and the prompt in the text written in the prompt, and it is not easy for the user to accurately express (and convey to the image generation AI), which is troublesome.
[0005] The present invention has been made in view of the above circumstances, and one of the objectives is to provide an information processing apparatus and an image generation method that improve the convenience when using image generation AI.
Means for Solving the Problems
[0006] The present invention has been made to solve the above problems, and an information processing device according to a first aspect of the present invention comprises a data acquisition unit that acquires data of text information superimposed on a background image and data of the background image, and a generated image acquisition unit that acquires a generated image generated by an image generation model by inputting image generation input data based on the background image and the text information to the image generation model.
[0007] The above-described information processing device may further include a prompt generation instruction unit that instructs the generation of a prompt as an input sentence to an image generation model by inputting prompt generation input data based on the background image and text information acquired by the data acquisition unit to a text generation model, and an image generation instruction unit that inputs the prompt generated by the text generation model and the background image as image generation input data to the image generation model.
[0008] The above-described information processing device may further include: a background image input unit that receives the input of the background image; a text input unit that receives text input of the text information superimposed on the background image received by the background image input unit; and an output unit that outputs the background image and the text information as separate data.
[0009] In the above-described information processing device, the background image input unit may accept input for the first layer, and the text input unit may accept text input for the second layer which is superimposed on the first layer.
[0010] The above-described information processing device further includes a prompt editing unit that receives text input as a prompt to be input to the image generation model, and outputs the prompt input by the user as a prompt to be added to the prompt generated by the text generation model, and the image generation instruction unit may input the prompt output by the prompt editing unit to the image generation model as a prompt to be input to the image generation model.
[0011] The above-described information processing device further includes a prompt editing unit that acquires and presents a prompt generated by the text generation model in a modifiable manner, accepts modification input for the presented prompt, and modifies and outputs the prompt generated by the text generation model based on the modification input, and the image generation instruction unit may input the prompt output by the prompt editing unit to the image generation model as a prompt to be input to the image generation model.
[0012] In the above-described information processing device, the text generation model may be a VLM (Vision and Language Model).
[0013] Furthermore, an image generation method in an information processing apparatus according to a second aspect of the present invention includes the steps of: a data acquisition unit acquiring data of text information superimposed on a background image and data of the background image; and a generated image acquisition unit acquiring a generated image generated by an image generation model by inputting image generation input data based on the background image and the text information to the image generation model. [Effects of the Invention]
[0014] According to the above-described embodiment of the present invention, it is possible to improve the convenience of using image generation AI. [Brief explanation of the drawing]
[0015] [Figure 1] A diagram showing an example of the configuration of an information processing system according to the embodiment. [Figure 2] A block diagram showing an example of the hardware configuration of the information processing device according to the embodiment. [Figure 3] A diagram showing an example of an image generation UI according to the embodiment. [Figure 4] A diagram showing an example of input to the input area and an example of output to the output area according to the embodiment. [Figure 5] A block diagram showing an example of the functional configuration related to the image generation process according to the embodiment. [Figure 6] Block diagram showing a detailed example of the functional configuration of the UI processing unit according to the embodiment. [Figure 7] Block diagram showing a detailed example of the functional configuration of the image generation processing unit according to the embodiment. [Figure 8] Flowchart showing an example of the image generation processing according to the embodiment.
Mode for Carrying Out the Invention
[0016] Hereinafter, embodiments of the present invention will be described with reference to the drawings. [System Configuration] FIG. 1 is a diagram showing an example of the configuration of an information processing system according to the present embodiment. The information processing system SYS includes an information processing device 10, a text generation server 20, and an image generation server 30.
[0017] The information processing device 10 is, for example, a clamshell type (notebook type) PC (personal computer). The information processing device 10 has a substantially rectangular plate-shaped (for example, flat plate-shaped) first housing 10A and a second housing 10B that are coupled (connected) via a hinge mechanism so as to be openable and closable. Further, a display 150 is provided on the information processing device 10 across the first housing 10A to the second housing 10B.
[0018] The display 150 is a flexible display that can be bent when the first housing 10A and the second housing 10B are opened and closed. As the flexible display, for example, an organic EL display or the like is used. For example, the display 150 can be used not only in a single-screen mode in which the entire screen area is one screen but also in a two-screen mode in which the screen is divided into two, a first screen area 150A on the first housing 10A side and a second screen area 150B on the second housing 10B side. For example, when the information processing device 10 is placed on a desk and used, the second screen area 150B on the second housing 10B side becomes a substantially horizontal screen parallel to the upper surface of the desk.
[0019] In addition, a touch sensor is provided on the upper (front) surface of the display 150. The information processing device 10 can detect a touch operation on the screen area of the display 150. By turning on the information processing device 10, the user can view the display of the display 150 provided on the inner surfaces of the first housing 10A and the second housing 10B and perform a touch operation on the display 150, enabling the use of the information processing device 10.
[0020] In addition, the information processing device 10 can also be connected to an external keyboard 50, either wired or wirelessly, to receive keyboard input.
[0021] The text generation server 20 includes a text generation model 21 and generates and outputs text from the input image (or image and text). The text generation model 21 is a large-scale vision and language model, such as a VLM (Vision and Language Model) for example. That is, the text generation server 20 functions as a text generation AI that uses the text generation model 21 as a large-scale vision and language model to generate and output text from an image (or image and text, etc.).
[0022] The image generation server 30 includes an image generation model 31. The image generation model 31 is an image generation model learned by, for example, a Diffusion Model (diffusion model), and generates an image from the input image or prompt (text). That is, the image generation server 30 functions as an image generation AI that uses the image generation model 31 to generate and output an image from an image and a prompt (text).
[0023] The text generation server 20 and the image generation server 30 may be configured as one server or may be distributed among a plurality of servers.
[0024] In this embodiment, a server is equipped with a text generation AI that generates text using a text generation model 21 and an image generation AI that generates images using an image generation model 31. The information processing device 10 communicates with the server to utilize the functions of the text generation AI and the image generation AI. However, the functions of the text generation AI and the image generation AI may be downloaded to the information processing device 10 and used locally.
[0025] [Hardware configuration of the information processing device 10] The specific configuration of the information processing device 10 will be described below. Figure 2 is a block diagram showing an example of the hardware configuration of the information processing device 10 according to this embodiment. The information processing device 10 includes a communication unit 11, a RAM (Random Access Memory) 12, a storage unit 13, a speaker 14, a display unit 15, a camera 16, and a control unit 18. Each of these units is connected to communicate via a bus or the like.
[0026] The communication unit 11 is comprised of, for example, multiple Ethernet® ports, multiple USB (Universal Serial Bus) or other digital input / output ports, and communication devices that perform wireless communication such as Bluetooth® or Wi-Fi®.
[0027] For example, the communication unit 11 can communicate with an external keyboard 50, or a stylus, mouse, or touchpad, as shown in Figure 1, using Bluetooth®. Furthermore, the communication unit 11 can communicate with the text generation server 20 and image generation server 30 shown in Figure 1 by connecting to the internet via wireless communication such as Wi-Fi® or wired communication such as Ethernet®.
[0028] RAM12 is a volatile memory where programs and data for processing executed by the control unit 18 are stored, and various data are saved or erased as needed. Since RAM12 is a volatile memory, it will not retain data when power to RAM12 is cut off. Data that needs to be retained when power to RAM12 is cut off is transferred to the storage unit 13.
[0029] The memory unit 13 is a storage device that includes one or more of the following: SSD (Solid State Drive), HDD (Hard Disk Drive), ROM (Residual Only Memory), Flash-ROM, etc. For example, the memory unit 13 stores BIOS (Basic Input Output System) programs and configuration data, OS (Operating System) and application programs that run on the OS, and various data used by applications. Speaker 14 outputs electronic sounds, voices, etc.
[0030] The display unit 15 includes a display 150 and a touch sensor 155. As mentioned above, the display 150 is a flexible display that can be bent in accordance with the opening and closing of the first housing 10A and the second housing 10B. The display 150 displays the OS desktop screen, running application windows, etc., in accordance with the control of the control unit 18.
[0031] The touch sensor 155 is located on the screen of the display 150 and detects touch operations on the screen. Touch operations include, for example, tapping, sliding, flicking, swiping, and pinching. Touch operations can be performed using a finger or a stylus.
[0032] The camera 16 is composed of a lens, an image sensor, and the like. The camera 16 captures images (still images and videos) and outputs the data of the captured images in accordance with the control unit 18.
[0033] The control unit 18 is composed of processors such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and a microcomputer, and these processors execute programs (such as the BIOS, OS, and applications that run on the OS) stored in the memory unit 13, etc., to realize various functions. For example, the control unit 18 controls the screen mode of the display 150, controls the display on the screen area, and detects touch operations on the touch sensor 155.
[0034] Next, we will explain the image generation process using the image generation model 31 (image generation AI) in the information processing device 10. First, we will explain the UI (User Interface) with reference to Figure 3.
[0035] [Image generation UI] Figure 3 shows an example of an image generation UI according to this embodiment. In the illustrated example, the information processing device 10 controls the display 150 to a two-screen mode consisting of a first screen area 150A and a second screen area 150B. The second screen area 150B displays a window for application AP1 (hereinafter referred to as "image generation app AP1"), which has an image generation function that generates images using an image generation model 31 (image generation AI). The window for image generation app AP1 has an input area R1 for inputting information about the image to be generated and an output area R2 for displaying the generated image as a UI.
[0036] The user can hand-draw a preliminary sketch in input area R1, and also add annotations as text overlaying on the sketch. This additional information includes details that supplement each part of the sketch, or information about objects not present in the sketch. Annotations can be added anywhere on the sketch.
[0037] When the information processing device 10 is placed on a desk and used, the second screen area 150B becomes a nearly horizontal screen, and having the input area R1 on the second screen area 150B side makes it easier for the user to input data.
[0038] When users hand-draw preliminary sketches, they can primarily use a stylus, but they can also use their fingers, an external mouse (not shown), or an external touchpad (not shown). Furthermore, when users input annotations as text, they can use, for example, a keyboard (see Figure 1), but they can also use an on-screen software keyboard or hand-draw using a stylus.
[0039] Note that the image generation UI shown in Figure 3 is just one example, and the arrangement and size of the input area R1 and output area R2, as well as which screen area they are displayed in, are not limited to this example.
[0040] Figure 4 shows an example of input to input area R1 and an example of output to output area R2 according to this embodiment. Figure 4(A) shows an example of input to input area R1. As mentioned above, input area R1 allows for handwritten input of a background image and text input of annotations.
[0041] For example, the image generation application AP1 has multiple overlapping layers within the input area R1, and it switches between Layer 1, which accepts hand-drawn input for the background image, and Layer 2, which accepts text input for annotations, to accept each type of input. In other words, when the user is hand-drawing the background image, they select Layer 1 and input, and when they are inputting text for annotations, they select Layer 2 and input. This allows the user to, for example, select Layer 1 and hand-draw the background image, then select Layer 2 and input annotations overlaid on the background image entered in Layer 1.
[0042] In the input example shown in Figure 4(A), Layer 1 contains hand-drawn outlines of a "tree," a "person," "furniture with legs," and a "diagonal line" running from the upper right to the lower left within the input area R1. Layer 2, on top of the outline image in Layer 1, has annotations: "palm tree" at the location of the "tree," "woman" at the location of the "person," "reclining cot" at the location of the "furniture with legs," and "ocean" above and "beach" below the area separated by the "diagonal line."
[0043] The image generation application AP1 generates separate data for the background image input to Layer 1 and the annotation data input to Layer 2. The image generation application AP1 may obtain the background image data not only from input to Layer 1, but also from an image data file. In other words, the image generation application AP1 may read image data from a file and display part or all of the read image data as the background image on Layer 1. This allows the user to, for example, input annotations overlaid on the background image displayed on Layer 1.
[0044] The image generation application AP1 inputs input data based on the background image and annotations to the image generation model 31, and obtains the generated image generated by the image generation model 31. The image generation application AP1 then displays the obtained generated image in the output area R2. Figure 4(B) shows an example of the output in the output area R2, specifically an example of a generated image displayed in the output area R2. Note that the generated image shown in this figure is an illustrative image drawn to mimic the generated image produced by the image generation model 31. In the example of the generated image shown in Figure 4(B), an image is generated showing a beach with palm trees, a woman, and a reclining bed arranged according to the input example shown in Figure 4(A).
[0045] Returning to Figure 3, the first screen area 150A displays a window for another application, AP2, which is different from the image generation application AP1. Images generated by the image generation application AP1 can be used in other applications, for example, by copying and pasting them into the window of the other application AP2. The other application AP2 is an application that allows image insertion and pasting, such as an application with presentation creation functions, a document creation function, or an image editing function.
[0046] [Functional Configuration of Image Generation Processing] Next, we will describe the functional configuration of the image generation process realized by the control unit 18 executing the image generation application AP1. Figure 5 is a block diagram showing an example of the functional configuration related to the image generation process according to this embodiment. This figure shows the front-end configuration related to the UI of the image generation application AP1, which was explained with reference to Figure 4, and the back-end configuration that generates images according to the said process.
[0047] The front-end configuration includes a UI processing unit 100, as described in Figure 4. The UI processing unit 100 outputs the data of the background image and annotations input by the user into the input area R1 (see Figure 3) to the backend. The UI processing unit 100 also outputs the generated image obtained from the backend to the output area R2.
[0048] The backend configuration includes an image generation processing unit 110, a prompt generation unit 120, and an image generation unit 130.
[0049] The prompt generation unit 120 corresponds to the functional configuration of the text generation server 20 shown in Figure 1, and uses the text generation model 21 to generate and output a prompt (text) from the input image (or image and text). The image generation unit 130 corresponds to the functional configuration of the image generation server 30 shown in Figure 1, and uses the image generation model 31 to generate and output an image from the input image and prompt (text).
[0050] The image generation processing unit 110 connects the front-end UI processing unit 100 with the back-end prompt generation unit 120 and image generation unit 130, and executes image generation processing using the prompt generation unit 120 and image generation unit 130 based on data obtained from the UI processing unit 100.
[0051] The image generation processing unit 110 may have a functional configuration that is implemented by the control unit 18 when the control unit 18 executes the image generation application AP1, or it may have a functional configuration on the server side.
[0052] The configuration of the image generation process according to this embodiment will be described in detail below with reference to Figures 5 to 7. Figure 6 is a block diagram showing a detailed example of the functional configuration of the UI processing unit 100. The UI processing unit 100 includes a background image input unit 101, an annotation input unit 102 (an example of a text input unit), a data output unit 103, a generated image output unit 104, and a prompt editing unit 105.
[0053] The sketch input unit 101 accepts handwritten input of a sketch (sketch image) by the user. For example, the sketch input unit 101 accepts handwritten input of a sketch image by handwriting input to layer 1 of input area R1 (see Figures 3 and 4).
[0054] The annotation input unit 102 accepts annotation text input from the user, overlaid on the background image received by the background image input unit 101. For example, the annotation input unit 102 accepts annotation text input from the text input for layer 2 (see Figures 3 and 4) which is superimposed on layer 1 of input area R1.
[0055] The data output unit 103 outputs the background image received by the background image input unit 101 and the annotation received by the annotation input unit 102 as separate data.
[0056] The generated image output unit 104 acquires and outputs the generated image output from the image generation unit 130 via the image generation processing unit 110. For example, the generated image output unit 104 displays the acquired generated image in the output area R2 (see Figure 3).
[0057] The prompt editing unit 105 accepts text input from the user as a prompt, and outputs the prompt entered by the user as a prompt to be added to the prompt generated by the prompt generation unit 120.
[0058] Furthermore, the prompt editing unit 105 acquires the prompt generated by the prompt generation unit 120, presents (displays) it in a editable format, and accepts correction input for that prompt. Based on the correction input, the prompt editing unit 105 corrects the prompt generated by the prompt generation unit 120 and outputs the corrected prompt.
[0059] Figure 7 is a block diagram showing a detailed example of the functional configuration of the image generation processing unit 110. The image generation processing unit 110 includes a data acquisition unit 111, a prompt generation instruction unit 112, a prompt acquisition unit 113, an image generation instruction unit 114, and a generated image acquisition unit 115.
[0060] The data acquisition unit 111 acquires the data of the background image and the annotation data that has been overlaid on the background image from the UI processing unit 100.
[0061] The prompt generation instruction unit 112 instructs the prompt generation unit 120 to generate a prompt as an input sentence for the image generation unit 130 (image generation model 31) by inputting prompt generation input data based on the background image and annotations acquired by the data acquisition unit 111 to the prompt generation unit 120. For example, the prompt generation input data is image data that integrates the background image layer and the annotation layer.
[0062] Here, when the prompt generation instruction unit 112 inputs prompt generation input data to the prompt generation unit 120, the text generation prompt that it inputs to the prompt generation unit 120 (text generation model 21) is an instruction to cause the text generation model 21 to generate a prompt (image generation prompt) to be input to the image generation unit 130 (image generation model 31). For example, the prompt generation instruction unit 112 inputs a text generation prompt with an instruction such as "Generate a prompt containing object position information in a predetermined format" to the prompt generation unit 120 (text generation model 21).
[0063] In addition, depending on the image to be generated using the image generation unit 130 (image generation model 31), multiple text generation prompts are provided in advance, and one of them may be selected.
[0064] When prompt generation input data is received, the prompt generation unit 120 uses the text generation model 21 to generate and output a prompt (image generation prompt) to be input to the image generation unit 130.
[0065] The prompt acquisition unit 113 acquires prompts generated by the prompt generation unit 120 from the prompt generation unit 120. The prompt acquisition unit 113 also outputs the acquired prompts to the UI processing unit 100. If the user adds or modifies a prompt in the UI processing unit 100 that was generated by the prompt generation unit 120, the prompt acquisition unit 113 acquires the prompt output by the prompt editing unit 105 of the UI processing unit 100.
[0066] If there are no additions or modifications to the prompt generated by the prompt generation unit 120, the prompt acquisition unit 113 outputs the prompt acquired from the prompt generation unit 120 to the image generation instruction unit 114. On the other hand, if there are additions or modifications to the prompt generated by the prompt generation unit 120, the prompt acquisition unit 113 outputs the prompt acquired from the prompt editing unit 105 of the UI processing unit 100 to the image generation instruction unit 114.
[0067] The image generation instruction unit 114 instructs the image generation unit 130 to generate an image by inputting the prompt obtained from the prompt acquisition unit 113 and the background image obtained from the data acquisition unit 111 as input data for image generation.
[0068] Specifically, if there are no additions or modifications to the prompt generated by the prompt generation unit 120, the image generation instruction unit 114 inputs the prompt generated by the prompt generation unit 120 and the background image acquired by the data acquisition unit 111 to the image generation unit 130. On the other hand, if there are additions or modifications to the prompt generated by the prompt generation unit 120, the image generation instruction unit 114 inputs the added or modified prompt and the background image acquired by the data acquisition unit 111 to the prompt generation unit 120.
[0069] When the image generation unit 130 receives a prompt and a background image as input, it generates an image using the image generation model 31 and outputs the generated image.
[0070] The generated image acquisition unit 115 acquires the generated image generated by the image generation unit 130. The generated image acquisition unit 115 also outputs the acquired generated image to the UI processing unit 100.
[0071] [Image generation process operation] Next, the operation of the image generation process performed by the image generation processing unit 110 will be described. Figure 8 is a flowchart showing an example of the image generation process according to this embodiment.
[0072] (Step S101) The image generation processing unit 110 obtains the data of the background image and the annotation data superimposed on the background image from the UI processing unit 100 based on the user's handwritten input and text input to the input area R1 shown in Figure 3. Then, it proceeds to the process in step S103.
[0073] (Step S103) The image generation processing unit 110 instructs the prompt generation unit 120 to generate prompts to be input to the image generation unit 130 (image generation model) by inputting prompt generation input data based on the background image and annotations acquired in step S101 to the prompt generation unit 120. For example, the image generation processing unit 110 inputs image data, which is an integrated version of the background image layer and the annotation layer, to the prompt generation unit 120 as prompt generation input data. Then, the process proceeds to step S105.
[0074] (Step S105) The image generation processing unit 110 acquires the prompt generated by the prompt generation unit 120. The image generation processing unit 110 also outputs the acquired prompt to the UI processing unit 100, and if the user adds or modifies a prompt in the UI processing unit 100, it acquires the added or modified prompt from the UI processing unit 100.
[0075] (Step S107) The image generation processing unit 110 instructs the image generation unit 130 to generate an image by inputting the background image acquired in step S101 and the prompt acquired in step S105 as input data for image generation to the image generation unit 130.
[0076] (Step S109) The image generation processing unit 110 acquires the generated image generated by the image generation unit 130 and outputs the acquired generated image to the UI processing unit 100. The UI processing unit 100 displays the acquired generated image in the output area R2 (see Figure 3).
[0077] As described above, the information processing device 10 according to this embodiment acquires the data of the annotation (an example of text information) superimposed on the background image and the background image, and inputs the image generation input data based on the background image and annotation to the image generation model 31. The information processing device 10 then acquires the generated image generated by the image generation model 31.
[0078] As a result, the information processing device 10 can easily generate an image by simply having the user input annotations overlaid on the background image, adding supplementary information for each part of the background image and information for objects not present in the background image. This eliminates the need to explain the correspondence between the annotations and the positions of each part of the background image in text, thereby improving the convenience of using the image generation AI.
[0079] For example, the information processing device 10 instructs the text generation model 21 to generate a prompt as input text for the image generation model 31 by inputting prompt generation input data based on the acquired background image and annotations into the text generation model 21. The information processing device 10 then inputs the prompt generated by the text generation model 21 and the background image into the image generation model 31 as image generation input data. Here, the prompt generation input data is, for example, image data obtained by integrating the background image layer and the annotation layer.
[0080] As a result, the information processing device 10 can generate prompts that associate the annotations with the positions of each part of the background image simply by the user inputting annotations overlaid on the background image, and input them into the image generation model 31. This eliminates the need to explain the association between the annotations and the positions of each part of the background image in text, thereby improving the convenience of using the image generation AI.
[0081] Furthermore, the information processing device 10 accepts the input of a background image and also accepts the input of annotation text to be superimposed on the received background image. The information processing device 10 then outputs the background image and the annotation as separate data.
[0082] As a result, the information processing device 10 treats the background image and annotations as independent data, eliminating the need to separate them later. Furthermore, because the information processing device 10 treats the background image and annotations as independent data, it is also possible to save only the background image or only the annotations. Additionally, because the information processing device 10 treats the background image and annotations as independent data, it can independently perform "Undo" to return to the state before the previous operation and "Redo" to return to the state before the Undo operation.
[0083] For example, the information processing device 10 accepts input for layer 1 (first layer) and text input for layer 2 (second layer) which is superimposed on layer 1.
[0084] As a result, the information processing device 10 can manage the background image and annotations in separate layers, allowing it to treat each of them as independent data.
[0085] Furthermore, the information processing device 10 accepts text input as a prompt to be input to the image generation model 31, and outputs the prompt entered by the user as a prompt to be added to the prompt generated by the text generation model 21. The information processing device 10 then inputs the outputted prompt to the image generation model 31 as a prompt to be input to the image generation model 31.
[0086] As a result, the information processing device 10 allows the user to add prompts to the prompts generated from the background image and annotations and input them into the image generation model 31, thus allowing the user to adjust the content of the prompts.
[0087] Furthermore, the information processing device 10 acquires the prompt generated by the text generation model 21, presents it in a modifiable form, accepts modification input for the presented prompt, and modifies the prompt generated by the text generation model 21 based on the modification input and outputs it. The information processing device 10 then inputs the outputted prompt to the image generation model 31 as a prompt to be input to the image generation model 31.
[0088] As a result, the information processing device 10 allows the user to modify the prompt generated from the background image and annotations and input it into the image generation model 31, thus allowing the user to adjust the content of the prompt.
[0089] Furthermore, the information processing device 10 generates prompts using a VLM (Vision and Language Model) as the text generation model 21.
[0090] As a result, the information processing device 10 can automatically generate prompts using text generation AI, where the input annotations are associated with the positions of each part of the background image, simply by the user inputting annotations overlaid on the background image.
[0091] Furthermore, the image generation method in the information processing device 10 includes the steps of: the control unit 18 (image generation processing unit 110) acquiring data for annotations (an example of text information) superimposed on the background image and the background image; and acquiring a generated image generated by the image generation model 31 by inputting image generation input data based on the background image and annotations to the image generation model 31.
[0092] As a result, the image generation method in the information processing device 10 allows the user to easily generate an image by simply inputting annotations overlaid on a background image, adding supplementary information to each part of the background image and information about objects not present in the background image. This eliminates the need to explain the correspondence between the annotations and the positions of each part of the background image in text, thereby improving the convenience of using the image generation AI.
[0093] Although embodiments of this invention have been described in detail above with reference to the drawings, the specific configurations are not limited to the embodiments described above, and include designs and the like that do not depart from the spirit of this invention. For example, the configurations described in the embodiments described above can be combined in any way.
[0094] In the embodiment described above, an example was explained in which image data, which is an integrated version of the background image layer and the annotation layer, is input to the prompt generation unit 120 (text generation model 21) as prompt generation input data. However, it is also possible to configure the system so that prompts (text information) are generated by inputting the background image data and the annotation (annotation image) data without integrating the background image layer and the annotation layer.
[0095] In the embodiment described above, an example was explained in which the prompt (image generation prompt) generated by the text generation model 21 can be edited by the user. However, the text generation prompt input to the text generation model 21 may also be configured to be editable by the user.
[0096] Furthermore, although the above-described embodiment described an example in which the display 150 is a single display provided across the first housing 10A and the second housing 10B, the invention is not limited to this. For example, the information processing device 10 may have two displays, one in the first housing 10A and one in the second housing 10B. Alternatively, the information processing device 10 may have a display provided in only one of the first housing 10A and the second housing 10B (for example, only in the first housing 10A).
[0097] Furthermore, the information processing device 10 is not limited to a clamshell-type (notebook-type) PC, but may also be a tablet-type PC or a desktop PC, for example.
[0098] Furthermore, although the above-described embodiment described an example of a touch panel type display in which the input unit (touch sensor) and the display unit (display) are integrated, a non-touch panel type display without an input unit (touch sensor) may also be used. In that case, instead of an input unit (touch sensor), an input device that accepts operation may be used, such as a touchpad or a mouse.
[0099] Alternatively, for example, the image generation model 31 may also include a text generation function by the text generation model 21, and image generation may be performed by inputting a background image and annotations into the image generation model 31.
[0100] The information processing device 10 described above has a computer system inside. The processing in each configuration of the information processing device 10 may be performed by recording a program for realizing the functions of each configuration of the information processing device 10 onto a computer-readable recording medium, loading the program recorded on this recording medium into the computer system, and executing it. Here, "loading the program recorded on the recording medium into the computer system and executing it" includes installing the program into the computer system. Here, "computer system" includes hardware such as the OS and peripheral devices. Furthermore, "computer system" may include multiple computer devices connected via a network including communication lines such as the Internet, WAN, LAN, and dedicated lines. Also, "computer-readable recording medium" refers to portable media such as flexible disks, magneto-optical disks, ROMs, CD-ROMs, and storage devices such as hard disks built into the computer system. Thus, the recording medium storing the program may be a non-transient recording medium such as a CD-ROM.
[0101] Furthermore, the recording medium also includes internal or external recording media accessible from the distribution server for distributing the program. The program may be divided into multiple parts, downloaded at different times, and then combined in each configuration of the information processing device 10. The distribution servers for each of the divided programs may also be different. Moreover, "computer-readable recording media" includes volatile memory (RAM) within computer systems that act as servers or clients when a program is transmitted over a network, which retains the program for a certain period of time. The program itself may also be intended to implement some of the functions described above. Furthermore, the program may be a so-called differential file (differential program) that can implement the functions described above in combination with a program already recorded in the computer system.
[0102] Furthermore, some or all of the functions of the information processing device 10 in the above-described embodiment may be implemented as an integrated circuit such as an LSI (Large Scale Integration). Each function may be individually processorized, or some or all of them may be integrated into a single processor. In addition, the method of implementing the integrated circuit is not limited to LSIs; it may also be implemented using dedicated circuits or general-purpose processors. Furthermore, if an integrated circuit technology that can replace LSIs emerges due to advances in semiconductor technology, an integrated circuit using that technology may be used. [Explanation of Symbols]
[0103] 10 Information Processing Unit, 10A First Enclosure, 10B Second Enclosure, 20 Text Generation Server, 21 Text Generation Model, 30 Image Generation Server, 31 Image Generation Model, 11 Communication Unit, 12 RAM, 13 Storage Unit, 14 Speaker, 15 Display Unit, 16 Camera, 150 Display, 150A First Screen Area, 150B Second Screen Area, 155 Touch Sensor, 18 Control Unit, 100 UI Processing Unit, 101 Background Image Input Unit, 102 Annotation Input Unit, 103 Data Output Unit, 104 Generated Image Output Unit, 105 Prompt Editing Unit, 110 Image Generation Processing Unit, 111 Data Acquisition Unit, 112 Prompt Generation Instruction Unit, 113 Prompt Acquisition Unit, 114 Image Generation Instruction Unit, 115 Generated Image Acquisition Unit, 120 Prompt Generation Unit, 130 Image Generation Unit, SYS Information Processing System
Claims
1. A data acquisition unit that acquires data from text information superimposed on a background image and data from the background image, A generated image acquisition unit acquires a generated image generated by an image generation model by inputting image generation input data based on the aforementioned background image and the aforementioned text information into the image generation model. An information processing device equipped with the following features.
2. A prompt generation instruction unit instructs the text generation model to generate a prompt as an input sentence for the image generation model by inputting prompt generation input data based on the background image and text information acquired by the data acquisition unit to the text generation model. An image generation instruction unit inputs the prompt generated by the text generation model and the background image as input data for image generation to the image generation model. The information processing apparatus according to claim 1, comprising:
3. A background image input unit that receives the aforementioned background image input, A text input unit that receives text input of the text information to be superimposed on the background image received by the background image input unit, An output unit that outputs the aforementioned preliminary image and the aforementioned text information as separate data, The information processing apparatus according to claim 1, comprising:
4. The aforementioned background image input unit receives input for the first layer, The text input unit receives text input for the second layer which is superimposed on the first layer. The information processing apparatus according to claim 3.
5. A prompt editing unit that accepts text input as a prompt to be input to the image generation model, and outputs the prompt entered by the user as a prompt to be added to the prompt generated by the text generation model. Equipped with, The image generation instruction unit, The prompt output by the prompt editing unit is input to the image generation model as a prompt to be input to the image generation model. The information processing apparatus according to claim 2.
6. A prompt editing unit that acquires prompts generated by the text generation model, presents them in a modifiable format, accepts modification input for the presented prompts, and modifies and outputs the prompts generated by the text generation model based on said modification input. Equipped with, The image generation instruction unit, The prompt output by the prompt editing unit is input to the image generation model as a prompt to be input to the image generation model. The information processing apparatus according to claim 2.
7. The text generation model uses VLM (Vision and Language Model) to generate prompts. The information processing apparatus according to claim 2.
8. An image generation method in an information processing device, The data acquisition unit performs the step of acquiring data from both the text information superimposed on the background image and the background image itself. The generated image acquisition unit inputs image generation input data based on the background image and the text information to the image generation model, thereby acquiring the generated image generated by the image generation model. An image generation method that includes [a specific feature / method].