Information processing device, prompt generation method, and image generation method

The information processing device improves image generation AI usability by detecting characters, determining their positions, and generating prompts, addressing the challenges of users who are not skilled at drawing or describing text prompts effectively.

JP2026086164AActive Publication Date: 2026-05-26LENOVO (SINGAPORE) PTE LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
LENOVO (SINGAPORE) PTE LTD
Filing Date
2024-11-14
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Individuals who are not skilled at drawing face challenges in using image generation AI, as they struggle to accurately convey their desired image through text prompts, and there is anxiety about generating the correct image.

Method used

An information processing device that includes a character detection unit to identify characters in an image, a word cluster determination unit to group characters, a location information acquisition unit to determine the character positions, and a prompt generation unit to create prompts based on these clusters and positions, thereby generating images corresponding to user input.

Benefits of technology

Enhances the convenience of using image generation AI by automatically detecting characters, determining their positions, and generating prompts that accurately place image elements, reducing the need for users to manually draw or precisely describe image placements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026086164000001_ABST
    Figure 2026086164000001_ABST
Patent Text Reader

Abstract

To improve the convenience of using image generation AI. [Solution] The information processing device comprises: a character detection unit that detects characters in an image; a word cluster determination unit that determines a group of characters consisting of one or more words from among the characters detected by the character detection unit as a word cluster; a location information acquisition unit that acquires location information indicating the position of the word cluster determined by the word cluster determination unit within the image; and a prompt generation unit that generates a prompt as an input sentence to an image generation model based on the word cluster and the location information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, a prompt generation method, and an image generation method.

Background Art

[0002] In recent years, image generation AI (Artificial Intelligence) that generates an image from an image such as a hand-drawn sketch (Image to Image) is known. There is also image generation AI that generates an image by inputting a prompt created in text (Text to Image) (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, for those who are not good at drawing, there is anxiety about drawing the picture they imagine, and there is a concern that the desired image will not be generated. In addition, in the image generation AI as described above, some generate an image by inputting a prompt in which what the user wants to include in the generated image is described in text. However, it is necessary to describe in words where the user wants to place it, and it is not easy for the user to accurately express it (and convey it to the image generation AI), which is troublesome.

[0005] The present invention has been made in view of the above circumstances, and one of the objectives is to provide an information processing apparatus, a prompt generation method, and an image generation method that improve the convenience when using image generation AI.

Means for Solving the Problems

[0006] The present invention has been made to solve the above problems, and an information processing device according to a first aspect of the present invention comprises: a character detection unit that detects characters in an image; a word cluster determination unit that determines a group of characters consisting of one or more words from among the characters detected by the character detection unit as a word cluster; a location information acquisition unit that acquires location information indicating the location of the word cluster determined by the word cluster determination unit in the image; and a prompt generation unit that generates a prompt as an input sentence to an image generation model based on the word cluster and the location information.

[0007] In the above-described information processing device, the prompt generation unit may generate a prompt that includes the word cluster and the position information in the sentence.

[0008] In the above-described information processing device, the prompt generation unit may generate a prompt that instructs the generation of an image such that the image corresponding to the word cluster is placed at a position corresponding to the position information.

[0009] In the above-described information processing device, the prompt generation unit may generate a prompt indicating an instruction to place the word cluster as a character at a position corresponding to the position information if a specific symbol is included in the word cluster.

[0010] In the above-described information processing device, the character detection unit may detect characters by analyzing the handwritten trajectories contained in the image.

[0011] The above-described information processing device may also include a character removal unit that generates an image from which characters detected by the character detection unit have been removed, as an input image to the image generation model.

[0012] The above-described information processing device may also include a generated image acquisition unit that acquires a generated image output from the image generation model by inputting the input image generated by the character removal unit and the prompt generated by the prompt generation unit to the image generation model.

[0013] Furthermore, a prompt generation method in an information processing apparatus according to a second aspect of the present invention includes the steps of: a character detection unit detecting characters in an image; a word cluster determination unit determining a group of characters consisting of one or more words from among the characters detected by the character detection unit as a word cluster; a location information acquisition unit acquiring location information indicating the location of the word cluster determined by the word cluster determination unit within the image; and a prompt generation unit generating a prompt as an input sentence to an image generation model based on the word cluster and the location information.

[0014] Furthermore, an image generation method in an information processing apparatus according to a third aspect of the present invention includes the steps of: a character detection unit detecting characters in an image; a word cluster determination unit determining a group of characters consisting of one or more words from among the characters detected by the character detection unit as a word cluster; a position information acquisition unit acquiring position information indicating the position of the word cluster determined by the word cluster determination unit within the image; a prompt generation unit generating a prompt as an input sentence to an image generation model based on the word cluster and the position information; a character removal unit generating an image from which the characters detected by the character detection unit have been removed as an input image to the image generation model; and a generated image acquisition unit acquiring a generated image output from the image generation model by inputting the input image generated by the character removal unit and the prompt generated by the prompt generation unit to the image generation model. [Effects of the Invention]

[0015] According to the above-described embodiment of the present invention, it is possible to improve the convenience of using image generation AI. [Brief explanation of the drawing]

[0016] [Figure 1] A diagram showing an example of the configuration of an information processing system according to the embodiment. [Figure 2] A block diagram showing an example of the hardware configuration of the information processing apparatus according to the embodiment. [Figure 3] A diagram showing an example of the image generation UI according to the embodiment. [Figure 4] A diagram showing an input example of the input area and an output example of the output area according to the embodiment. [Figure 5] A block diagram showing an example of the functional configuration related to the image generation process according to the embodiment. [Figure 6] A flowchart showing an example of the prompt generation process according to the embodiment. [Figure 7] A flowchart showing an example of the image generation process according to the embodiment. [Figure 8] A diagram showing another example of the input of the input area and the output of the output area according to the embodiment. [Figure 9] A diagram showing examples of prompts with and without specific symbols according to the embodiment. [Figure 10] A diagram showing an example of an image generated by handwritten input of the characters shown in FIG. 9 according to the embodiment.

Embodiments of the Invention

[0017] Hereinafter, embodiments of the present invention will be described with reference to the drawings. [System Configuration] FIG. 1 is a diagram showing an example of the configuration of the information processing system according to the present embodiment. The information processing system SYS includes an information processing apparatus 10 and an image generation server 30.

[0018] The information processing apparatus 10 is, for example, a clamshell-type (notebook-type) PC (personal computer). The information processing apparatus 10 has a substantially rectangular plate-shaped (for example, flat plate-shaped) first housing 10A and a second housing 10B that are coupled (connected) via a hinge mechanism so as to be openable and closable. Further, a display 150 is provided on the information processing apparatus 10 across the first housing 10A to the second housing 10B.

[0019] The display 150 is a flexible display that can be bent when the first housing 10A and the second housing 10B are opened and closed. As a flexible display, for example, an organic EL display can be used. For example, the display 150 can be used not only in a single-screen mode in which the entire screen area is a single screen, but also in a two-screen mode in which the screen is divided into two screens: a first screen area 150A on the first housing 10A side and a second screen area 150B on the second housing 10B side. For example, when the information processing device 10 is used on a desk, the second screen area 150B on the second housing 10B side becomes a nearly horizontal screen parallel to the surface of the desk.

[0020] Furthermore, a touch sensor is provided on the top (surface) of the display 150. The information processing device 10 is capable of detecting touch operations on the screen area of ​​the display 150. By opening the information processing device 10, the user can view the displays on the displays 150 provided on the inner surfaces of the first housing 10A and the second housing 10B, and can also perform touch operations on the displays 150, thereby enabling the use of the information processing device 10.

[0021] The image generation server 30 is equipped with an image generation model 31. The image generation model 31 is, for example, an image generation model trained using a diffusion model, and generates images from input images or prompts (text). In other words, the image generation server 30 functions as an image generation AI (Artificial Intelligence) that generates and outputs images from images or prompts (text) using the image generation model 31. The image generation server 30 may be configured as a single server or distributed across multiple servers.

[0022] In this embodiment, an image generation AI that generates images using the image generation model 31 is provided on the server, and the information processing device 10 communicates with the server to utilize the functions of the image generation AI. However, it is also possible to download the functions of the image generation AI to the information processing device 10 and use them locally.

[0023] [Hardware configuration of the information processing device 10] The specific configuration of the information processing device 10 will be described below. Figure 2 is a block diagram showing an example of the hardware configuration of the information processing device 10 according to this embodiment. The information processing device 10 includes a communication unit 11, a RAM (Random Access Memory) 12, a storage unit 13, a speaker 14, a display unit 15, a camera 16, and a control unit 18. Each of these units is connected to communicate via a bus or the like.

[0024] The communication unit 11 is comprised of, for example, multiple Ethernet® ports, multiple USB (Universal Serial Bus) or other digital input / output ports, and communication devices that perform wireless communication such as Bluetooth® or Wi-Fi®.

[0025] For example, the communication unit 11 can communicate with a stylus, mouse, touchpad, etc., using Bluetooth®. Furthermore, the communication unit 11 can communicate with the image generation server 30 shown in Figure 1 by connecting to the internet via wireless communication such as Wi-Fi® or wired communication such as Ethernet®.

[0026] RAM12 is a volatile memory where programs and data for processing executed by the control unit 18 are stored, and various data are saved or erased as needed. Since RAM12 is a volatile memory, it will not retain data when power to RAM12 is cut off. Data that needs to be retained when power to RAM12 is cut off is transferred to the storage unit 13.

[0027] The memory unit 13 is a storage device that includes one or more of the following: SSD (Solid State Drive), HDD (Hard Disk Drive), ROM (Residual Only Memory), Flash-ROM, etc. For example, the memory unit 13 stores BIOS (Basic Input Output System) programs and configuration data, OS (Operating System) and application programs that run on the OS, and various data used by applications. Speaker 14 outputs electronic sounds, voices, etc.

[0028] The display unit 15 includes a display 150 and a touch sensor 155. As mentioned above, the display 150 is a flexible display that can be bent in accordance with the opening and closing of the first housing 10A and the second housing 10B. The display 150 displays the OS desktop screen, running application windows, etc., in accordance with the control of the control unit 18.

[0029] The touch sensor 155 is located on the screen of the display 150 and detects touch operations on the screen. Touch operations include, for example, tapping, sliding, flicking, swiping, and pinching. Touch operations can be performed using a finger or a stylus.

[0030] The camera 16 is composed of a lens, an image sensor, and the like. The camera 16 captures images (still images and videos) and outputs the data of the captured images in accordance with the control unit 18.

[0031] The control unit 18 is composed of processors such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and a microcomputer, and these processors execute programs (such as the BIOS, OS, and applications that run on the OS) stored in the memory unit 13, etc., to realize various functions. For example, the control unit 18 controls the screen mode of the display 150, controls the display on the screen area, and detects touch operations on the touch sensor 155.

[0032] Next, we will explain the image generation process using the image generation model (image generation AI) in the information processing device 10. First, we will explain the UI (User Interface) with reference to Figure 3.

[0033] [Image generation UI] Figure 3 shows an example of an image generation UI according to this embodiment. In the illustrated example, the information processing device 10 controls the display 150 to a two-screen mode consisting of a first screen area 150A and a second screen area 150B. The first screen area 150A and the second screen area 150B display windows of an application (hereinafter referred to as the "image generation app") that has an image generation function that generates images using an image generation model (image generation AI).

[0034] In the second screen area 150B, the UI displays a window from the image generation application that includes an input area R1 for entering information about the image to be generated. In the first screen area 150A, the UI displays a window from the image generation application that includes an output area R2 for displaying the generated image.

[0035] The user can hand-draw a preliminary sketch in the input area R1, and can also hand-draw text on top of the sketch. When the information processing device 10 is used on a desk, the second screen area 150B becomes a nearly horizontal screen, so having the input area R1 on the second screen area 150B side makes it easier for the user to input.

[0036] Users can primarily use a stylus to hand-draw sketches and text, but they may also use their fingers, an external mouse (not shown), or an external touchpad (not shown).

[0037] Note that the image generation UI shown in Figure 3 is just one example, and the arrangement and size of the input area R1 and output area R2, as well as which screen area they are displayed in, are not limited to this example.

[0038] Figure 4 shows an example of input to input area R1 and an example of output to output area R2 according to this embodiment. Figure 4(A) shows an example of input to input area R1. As mentioned above, background images and characters can be input to input area R1.

[0039] In the input example shown in Figure 4(A), the user has hand-drawn a "ship" and a "water surface" as a base image. Here, the user wants to generate an image where the area below the water surface is a rough sea and multiple birds are flying in the sky. However, it is difficult to represent a rough sea with a drawing, and drawing multiple birds is time-consuming. In cases like this, where it is difficult or time-consuming to represent something with a drawing, the user can input using text. In the example shown, the words "rough ocean" are hand-drawn for the area below the water surface, and "birds" are hand-drawn for the sky above the water surface.

[0040] The image generation application detects characters from the input background image and character images and generates prompts based on the characters. The image generation application then inputs the generated prompts and input data based on the background image to the image generation model 31, thereby obtaining the generated image produced by the image generation model 31. The image generation application then displays the obtained generated image in the output area R2.

[0041] Figure 4(B) shows an example of the output of output area R2, specifically an example of a generated image displayed in output area R2. The generated image shown in this figure is an illustrative image drawn to mimic the generated image produced by the image generation model 31. In the example of the generated image shown in Figure 4(B), an image of a ship sailing on a rough sea with several birds flying in the sky is generated, corresponding to the background image and text input example shown in Figure 4(A).

[0042] Furthermore, the generated images created by the image generation application (the generated images displayed in output area R2) can be used in other applications. For example, they can be copied and pasted into the windows of other applications. These other applications include applications with presentation creation functions, document creation functions, or image editing functions—any application that allows for image insertion and pasting.

[0043] [Functional Configuration of Image Generation Processing] Next, we will explain in detail the image generation process, which involves detecting characters to generate prompts and then generating an image based on the generated prompts and background images.

[0044] Figure 5 is a block diagram showing an example of the functional configuration related to the image generation process according to this embodiment. The illustrated image generation processing unit 110 has a functional configuration in which the control unit 18 performs the image generation process by executing an image generation application. The image generation processing unit 110 includes a character detection unit 111, a word cluster determination unit 112, a location information acquisition unit 113, a prompt generation unit 114, a character removal unit 115, and a generated image acquisition unit 116.

[0045] The character detection unit 111 detects characters in the image, including the background drawing and characters, that are input into the input area R1 (see Figure 3). For example, the character detection unit 111 detects characters by analyzing the trajectory (stroke) of handwritten characters contained in the image.

[0046] The word cluster determination unit 112 determines a group of characters consisting of one or more words from the characters detected by the character detection unit 111 as a word cluster. For example, the word cluster determination unit 112 determines word clusters based on the distance between characters detected by the character detection unit 111. The word cluster determination unit 112 also sets a bounding box for the area (sub-region) surrounding each group of characters in a word cluster.

[0047] In the illustrated example, "birds" and "rough ocean" are determined as word clusters, and the bounding boxes for each word cluster are shown as rectangles. The word cluster determination unit 112 outputs information about the bounding box for each word cluster to the position information acquisition unit 113.

[0048] The position information acquisition unit 113 acquires position information indicating the position of the word clusters within the image determined by the word cluster determination unit 112. For example, the position information acquisition unit 113 divides the image into 3x3 9 sections, determines the position of the bounding box of the word cluster from within the divided section, and acquires position information. The determination of the bounding box position may be based on, for example, the position of the upper left corner of the bounding box, or on the center position of the bounding box.

[0049] In the illustrated example, the location information acquisition unit 113 acquires "Middle Left" as the location information for "birds" and "Bottom Center" as the location information for "rough ocean". The location information acquisition unit 113 outputs the acquired word cluster location information to the prompt generation unit 114.

[0050] The prompt generation unit 114 generates a prompt as an input sentence to the image generation model 31 based on word clusters and location information. For example, the prompt generation unit 114 generates a prompt that includes word clusters and location information in the sentence. In other words, the prompt generation unit 114 generates a prompt that instructs the generation of an image so that the image corresponding to the word cluster is placed at a location corresponding to the location information acquired by the location information acquisition unit 113.

[0051] For example, the prompt generation unit 114 generates a prompt using a predetermined format. One example of this predetermined format is a sentence such as "Place [word cluster] at [location information]". The prompt generation unit 114 generates a prompt by substituting the characters (word clusters) detected from the image and their location information into the "word cluster" and [location information] parts of this sentence.

[0052] In the illustrated example, the prompt format is set to "[words cluster] definitely positioned at [position] of the image". The prompt generation unit 114 generates the prompt "[birds] definitely positioned at [middle left] of the image, [rough ocean] definitely positioned at [bottom center] of the image."

[0053] These prompts are just examples and can be set using any format.

[0054] In this embodiment, the image area was divided into 3x3 grids (9 sections) to determine the position of word clusters. However, this is not the only method, and the number of divisions in the image area can be set arbitrarily. Furthermore, the position of word clusters may be defined by coordinates within the image.

[0055] The character removal unit 115 generates an image from which the characters detected by the character detection unit 111 have been removed, and uses this image as the input image for the image generation model 31.

[0056] The generated image acquisition unit 116 inputs the input image (the background image with the text removed) generated by the character removal unit 115 and the prompt generated by the prompt generation unit 114 to the image generation model 31. The image generation model 31 generates and outputs an image based on the input image (the background image with the text removed) and the prompt. As a result, the generated image acquisition unit 116 acquires the generated image output from the image generation model 31. The generated image acquisition unit 116 also displays the acquired generated image in the output area R2 (see Figure 3).

[0057] [Procedure generation process behavior] Next, the operation of the prompt generation process performed by the image generation processing unit 110 will be described. Figure 6 is a flowchart showing an example of the prompt generation process according to this embodiment.

[0058] (Step S101) The image generation processing unit 110 detects characters within the image from the background image and character-containing image input to the input area R1 (see Figure 3). Then, it proceeds to the process in step S103.

[0059] (Step S103) The image generation processing unit 110 determines a group of characters consisting of one or more words from the characters detected in step S101 as a word cluster. For example, the word cluster determination unit 112 sets a bounding box (a sub-region surrounding the group of characters) for each word cluster. Then, the process proceeds to step S105.

[0060] (Step S105) The image generation processing unit 110 acquires position information indicating the position of the word clusters within the image determined in step S103. For example, the position information acquisition unit 113 divides the image into 3x3 9 sections (see Figure 5), determines the position of the bounding box of the word clusters within the divided sections, and acquires position information. Then, it proceeds to the process in step S107.

[0061] (Step S107) The image generation processing unit 110 generates a prompt as an input sentence to the image generation model 31 based on the word cluster and location information. For example, the image generation processing unit 110 generates a prompt that includes the word cluster and location information in the sentence (see Figure 5).

[0062] [Image generation process operation] Next, the operation of the image generation process performed by the image generation processing unit 110 will be described. Figure 7 is a flowchart showing an example of the image generation process according to this embodiment. In this figure, the processes corresponding to each process in Figure 6 are denoted by the same reference numerals, and their explanations are omitted. Steps S101 to S107 are prompt generation processes as shown in Figure 6.

[0063] (Step S109) When the image generation processing unit 110 detects characters in the image, which includes the background image and characters input to the input area R1 (see Figure 3) in step S101, it removes the detected characters from the image and generates it as an input image for the image generation model 31. Then, it proceeds to the process in step S111.

[0064] (Step S111) The image generation processing unit 110 inputs the input image (the background image with the text removed) generated in step S109 and the prompt generated in step S107 to the image generation model 31. Then, it proceeds to the processing in step S113.

[0065] Here, the image generation model 31 generates and outputs an image based on the input image (the background image excluding the text) and the prompt.

[0066] (Step S113) The image generation processing unit 110 acquires the generated image output from the image generation model 31. The image generation processing unit 110 also displays the acquired generated image in the output area R2 (see Figure 3).

[0067] [Another input example for image generation processing] Next, another example of input for the image generation process in this embodiment will be described. In this embodiment, an image can be generated from the input of a background image and text, as shown in the example in Figure 4, but an image can also be generated from the input of only a background image or only text. In the case of only a background image, no text is detected from the image, so no prompt is generated, and only the background image is input to the image generation model 31. In the case of only text, a prompt is generated, and the image generation model 31 receives an image without a background image and the generated prompt.

[0068] Figure 8 shows another example of the input to input area R1 and the output to output area R2 according to this embodiment. Figure 8(A) shows an example of the input to input area R1, and Figure 8(B) shows an example of the output to output area R2.

[0069] In the input example shown in Figure 8(A), there is no background image input, and the words "mountain," "green cars," "house," and "road" are handwritten. In this case, the generated prompt will be, for example, "[mountain] definitely positioned at [top center] of the image, [green cars] definitely positioned at [middle center] of the image, [house] definitely positioned at [middle right] of the image, [road] definitely positioned at [bottom left] of the image."

[0070] In the example of a generated image shown in Figure 8(B), based on the text input example shown in Figure 8(A), an image is generated showing a house on the right side of a road with multiple cars (green cars) driving on it, and a mountain towering in the distance. Note that the generated image shown in this figure is an illustrative image drawn to mimic the generated image produced by the image generation model 31.

[0071] [Example of text input] Furthermore, when inputting text, it is sometimes necessary not only to generate images corresponding to the characters (word clusters), but also to include the characters themselves in the images. Therefore, the image generation processing unit 110 (prompt generation unit 114) generates a prompt indicating that if a specific symbol is included in the word cluster, the word cluster should be positioned as text at a location corresponding to the positional information. The specific symbol can be arbitrarily set, but in this embodiment, the specific symbol is set to ().

[0072] Figure 9 shows examples of prompts with and without a specific symbol. Figure 9(A) is an example where the specific symbol is absent, the handwritten input is "Robot", and the generated prompt is, for example, "[Robot] definitely positioned at [position] of the image." [position] is where the position information is entered.

[0073] Figure 9(B) shows an example where a specific symbol is present, and the handwritten characters are "(Robot)". Because the word cluster "(Robot)" contains the specific symbol (), the generated prompt will be, for example, "[Robot]" definitely positioned at [position] of the image, as letters. In other words, if a specific symbol () is present, the prompt should state that it should be generated as letters.

[0074] Figure 10 shows an example of an image generated by handwritten input of characters as shown in Figure 9. Figure 10(A) shows an example of input in input area R1, and Figure 10(B) shows an example of output in output area R2.

[0075] In the input example shown in Figure 10 (A), there is no background image input, and the words "[Robot]" are handwritten at the top and "(Robot)" at the bottom. The prompt generated from these handwritten characters (word clusters) is as shown in the example in Figure 9, and it is noted that "(Robot)" is generated as text because it contains the specific symbol ().

[0076] In the example of a generated image shown in Figure 10(B), an image containing both a robot image and the word "ROBOT" is generated, corresponding to the text input example shown in Figure 10(A). Note that the generated image shown in this figure is an illustrative diagram that mimics the generated image produced by the image generation model 31.

[0077] As described above, the information processing device 10 according to this embodiment detects characters in an image and determines a group of characters consisting of one or more words as a word cluster. The information processing device 10 also acquires positional information indicating the location of the word cluster in the image and generates a prompt as an input sentence for the image generation model 31 based on the word cluster and the positional information.

[0078] As a result, the information processing device 10 automatically detects characters from an image and generates prompts corresponding to the positions of character clusters. Therefore, for people who are not good at drawing, or when it is difficult or time-consuming to represent something with a picture, they can generate an image using the image generation model 31 by inputting characters onto an image such as a sketch, without having to draw anything themselves. Thus, the information processing device 10 can improve the convenience of using the image generation AI.

[0079] Furthermore, if an image containing text were input directly to the image generation model 31 (with the text treated as an image), conventional technology would randomly generate either the text portion as an image or the text itself. In contrast, in this embodiment, the information processing device 10 separates the text from the image and creates prompts, thereby generating the intended image. Additionally, since the information processing device 10 generates prompts based on word clusters and their positional information, the user does not need to create a prompt explaining in text where they want the text to be placed in the image, which is convenient.

[0080] For example, the information processing device 10 generates a prompt that includes word clusters and location information within a sentence.

[0081] This makes the information processing device 10 convenient because it can generate a prompt containing word clusters and their positional information without requiring the user to create a written prompt explaining where they want to place them in the image.

[0082] More specifically, the information processing device 10 generates a prompt that instructs the generation of an image so that an image corresponding to a word cluster is placed at a location corresponding to the location information.

[0083] As a result, the information processing device 10 can generate prompts that instruct the image generated by the image generation model 31 to place the image intended by the user in the appropriate position, based on the characters entered by the user.

[0084] Furthermore, if a word cluster contains a specific symbol, the information processing device 10 generates a prompt indicating an instruction to place the word cluster as a character at a position corresponding to the positional information.

[0085] As a result, the information processing device 10 is convenient because even if the user wants to include actual characters in the generated image produced by the image generation model 31, they only need to input characters by inserting specific symbols.

[0086] For example, the information processing device 10 detects characters by analyzing the handwritten trajectories contained in the image.

[0087] As a result, the information processing device 10 can generate prompts appropriately simply by the user inputting characters by hand.

[0088] Furthermore, the information processing device 10 generates an image from which the detected characters have been removed as an input image for the image generation model 31.

[0089] As a result, the information processing device 10 can appropriately input images containing both background images and text into the image generation model 31 by removing the text from the image. Therefore, the information processing device 10 can prevent random occurrences where text is generated as an image or as text itself.

[0090] Furthermore, the information processing device 10 obtains the generated image output from the image generation model 31 by inputting the input image from which the characters have been removed and the generated prompt to the image generation model 31.

[0091] As a result, the information processing device 10 detects characters from an image and automatically generates prompts corresponding to the positions of character clusters. It then inputs the generated prompts and the image with the characters removed into the image generation model 31. Therefore, for people who are not good at drawing, or when it is difficult or time-consuming to represent something with a picture, they can generate an image using the image generation model 31 by inputting characters onto an image such as a sketch, without having to draw anything themselves. Thus, the information processing device 10 can improve the convenience of using the image generation AI.

[0092] Furthermore, the prompt generation method in the information processing device 10 according to this embodiment includes the steps of: the control unit 18 (image generation processing unit 110) detecting characters in an image; determining a group of characters consisting of one or more words from the detected characters as a word cluster; acquiring positional information indicating the position of the determined word cluster in the image; and generating a prompt as an input sentence to the image generation model 31 based on the word cluster and the positional information.

[0093] As a result, the prompt generation method in the information processing device 10 automatically generates prompts corresponding to the position of character clusters (word clusters) by detecting characters from an image. Therefore, for people who are not good at drawing, or when it is difficult or time-consuming to represent something with a picture, they can generate an image using the image generation model 31 by inputting characters onto an image such as a sketch, without having to draw a picture themselves. Thus, the prompt generation method in the information processing device 10 can improve the convenience of using the image generation AI.

[0094] Furthermore, according to the prompt generation method in the information processing device 10, since characters are separated from images and then converted into prompts, when an image containing characters is input directly to the image generation model 31 (with the characters treated as images), the characters are not randomly generated as images or as the characters themselves, ensuring that the intended image is generated. In addition, according to the prompt generation method in the information processing device 10, prompts are generated based on word clusters and their positional information, eliminating the need for the user to create a prompt explaining in text where they want the characters placed in the image, thus improving convenience.

[0095] Furthermore, the image generation method in the information processing device 10 according to this embodiment includes the steps of: the control unit 18 (image generation processing unit 110) detecting characters in an image; determining a group of characters consisting of one or more words from the detected characters as a word cluster; acquiring positional information indicating the position of the determined word cluster in the image; generating a prompt as an input sentence to the image generation model 31 based on the word cluster and the positional information; generating an image from which the detected characters have been removed as an input image to the image generation model 31; and acquiring a generated image output from the image generation model 31 by inputting the generated input image and the generated prompt to the image generation model 31.

[0096] As a result, the image generation method in the information processing device 10 automatically generates prompts corresponding to the positions of character clusters (word clusters) by detecting characters from an image, and inputs the generated prompts and the image with the characters removed into the image generation model 31. Therefore, for people who are not good at drawing, or when it is difficult or time-consuming to represent something with a picture, they can generate an image in the image generation model 31 by inputting characters onto an image such as a sketch, without having to draw a picture themselves. Thus, the image generation method in the information processing device 10 can improve the convenience of using image generation AI.

[0097] Although embodiments of this invention have been described in detail above with reference to the drawings, the specific configurations are not limited to the embodiments described above, and include designs and the like that do not depart from the spirit of this invention. For example, the configurations described in the embodiments described above can be combined in any way.

[0098] Furthermore, although the above-described embodiment explained an example of inputting characters by hand on a sketch, character input is not limited to handwriting; it is also possible to use a keyboard or other means.

[0099] Furthermore, although the above-described embodiment described an example in which the display 150 is a single display provided across the first housing 10A and the second housing 10B, the invention is not limited to this. For example, the information processing device 10 may have two displays, one in the first housing 10A and one in the second housing 10B. Alternatively, the information processing device 10 may have a display provided in only one of the first housing 10A and the second housing 10B (for example, only in the first housing 10A).

[0100] Furthermore, the information processing device 10 is not limited to a clamshell-type (notebook-type) PC, but may also be a tablet-type PC or a desktop PC, for example.

[0101] Furthermore, although the above-described embodiment described an example of a touch panel type display in which the input unit (touch sensor) and the display unit (display) are integrated, a non-touch panel type display without an input unit (touch sensor) may also be used. In that case, instead of an input unit (touch sensor), an input device that accepts operation may be used, such as a touchpad or a mouse.

[0102] The information processing device 10 described above has a computer system inside. The processing in each configuration of the information processing device 10 may be performed by recording a program for realizing the functions of each configuration of the information processing device 10 onto a computer-readable recording medium, loading the program recorded on this recording medium into the computer system, and executing it. Here, "loading the program recorded on the recording medium into the computer system and executing it" includes installing the program into the computer system. Here, "computer system" includes hardware such as the OS and peripheral devices. Furthermore, "computer system" may include multiple computer devices connected via a network including communication lines such as the Internet, WAN, LAN, and dedicated lines. Also, "computer-readable recording medium" refers to portable media such as flexible disks, magneto-optical disks, ROMs, CD-ROMs, and storage devices such as hard disks built into the computer system. Thus, the recording medium storing the program may be a non-transient recording medium such as a CD-ROM.

[0103] Furthermore, the recording medium also includes internal or external recording media accessible from the distribution server for distributing the program. The program may be divided into multiple parts, downloaded at different times, and then combined in each configuration of the information processing device 10. The distribution servers for each of the divided programs may also be different. Moreover, "computer-readable recording media" includes volatile memory (RAM) within computer systems that act as servers or clients when a program is transmitted over a network, which retains the program for a certain period of time. The program itself may also be intended to implement some of the functions described above. Furthermore, the program may be a so-called differential file (differential program) that can implement the functions described above in combination with a program already recorded in the computer system.

[0104] Furthermore, some or all of the functions of the information processing device 10 in the above-described embodiment may be implemented as an integrated circuit such as an LSI (Large Scale Integration). Each function may be individually processorized, or some or all of them may be integrated into a single processor. In addition, the method of implementing the integrated circuit is not limited to LSIs; it may also be implemented using dedicated circuits or general-purpose processors. Furthermore, if an integrated circuit technology that can replace LSIs emerges due to advances in semiconductor technology, an integrated circuit using that technology may be used. [Explanation of Symbols]

[0105] 10 Information processing device, 10A First enclosure, 10B Second enclosure, 30 Image generation server, 31 Image generation model, 11 Communication unit, 12 RAM, 13 Storage unit, 14 Speaker, 15 Display unit, 16 Camera, 150 Display, 150A First screen area, 150B Second screen area, 155 Touch sensor, 18 Control unit, 110 Image generation processing unit, 111 Character detection unit, 112 Word cluster determination unit, 113 Location information acquisition unit, 114 Prompt generation unit, 115 Character removal unit, 116 Generated image acquisition unit, SYS Information Processing System

Claims

1. A character detection unit that detects characters within an image, A word cluster determination unit determines a group of characters consisting of one or more words from among the characters detected by the character detection unit as a word cluster, A location information acquisition unit acquires location information indicating the position of the word cluster determined by the word cluster determination unit within the image, A prompt generation unit generates a prompt as an input sentence to an image generation model based on the word cluster and the location information, An information processing device equipped with the following features.

2. The prompt generation unit, A prompt is generated that includes the word cluster and the position information in the sentence. The information processing apparatus according to claim 1.

3. The prompt generation unit, The system generates a prompt that instructs the generation of an image so that the image corresponding to the word cluster is placed at a position corresponding to the aforementioned location information. The information processing apparatus according to claim 2.

4. The prompt generation unit, If a specific symbol is included in the word cluster, the prompt is generated, indicating an instruction to place the word cluster as a character at a position corresponding to the positional information. The information processing apparatus according to claim 2.

5. The aforementioned character detection unit, Characters are detected by analyzing the handwritten trajectories contained in the aforementioned image. The information processing apparatus according to any one of claims 1 to 4.

6. A character removal unit generates an image from which characters detected by the character detection unit have been removed, as an input image for the image generation model. The information processing apparatus according to claim 1, comprising:

7. A generated image acquisition unit acquires a generated image output from the image generation model by inputting the input image generated by the character removal unit and the prompt generated by the prompt generation unit to the image generation model. The information processing apparatus according to claim 6, comprising:

8. A method for generating a prompt in an information processing device, The character detection unit performs the step of detecting characters in the image, The word cluster determination unit determines a group of characters consisting of one or more words from among the characters detected by the character detection unit as a word cluster, The location information acquisition unit acquires location information indicating the position of the word cluster determined by the word cluster determination unit within the image, The prompt generation unit generates a prompt as an input sentence to the image generation model based on the word cluster and the position information, A prompt generation method that includes this.

9. An image generation method in an information processing device, The character detection unit performs the step of detecting characters in the image, The word cluster determination unit determines a group of characters consisting of one or more words from among the characters detected by the character detection unit as a word cluster, The location information acquisition unit acquires location information indicating the position of the word cluster determined by the word cluster determination unit within the image, The prompt generation unit generates a prompt as an input sentence to the image generation model based on the word cluster and the position information, The character removal unit generates an image from which the characters detected by the character detection unit have been removed, as an input image for the image generation model. The generated image acquisition unit acquires a generated image output from the image generation model by inputting the input image generated by the character removal unit and the prompt generated by the prompt generation unit to the image generation model. An image generation method that includes [a specific feature / method].