Information processing device, prompt generation method, and image generation method
The information processing device enhances user convenience by automatically detecting characters and generating prompts for image generation models, addressing the challenges of user skill and complexity in existing image generation AI systems.
Patent Information
- Application Number
- JP2024199115
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-11-14
- Publication Date
- 2025-12-25
- Estimated Expiration
- 2044-11-14
AI Technical Summary
Users who are not skilled at drawing may feel uneasy about creating images with image generation AI, and existing text-based prompts require cumbersome and inaccurate descriptions of image placement.
An information processing device that includes a character detection unit, word cluster determination unit, and prompt generation unit to automatically detect characters, determine their positions, and generate prompts for image generation models based on word clusters and position information.
Improves user convenience by allowing users to input characters without needing to draw complex images, with the device generating accurate prompts for image generation models, ensuring characters are placed correctly and avoiding random generation.
Smart Images

Figure 0007792494000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, a prompt generating method, and an image generating method. [Background technology]
[0002] In recent years, image generation AI (Artificial Intelligence) that generates images from images such as hand-drawn sketches (Image to Image) has become known. There is also image generation AI that generates images by inputting prompts created in text (Text to Image) (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2024-120131 Summary of the Invention [Problem to be solved by the invention]
[0004] However, people who are not good at drawing may feel uneasy about drawing the picture they imagine, and may be concerned that the image they want may not be generated as they wish.In addition, some image generation AIs, such as those mentioned above, generate images by inputting a prompt that describes in text what the user wants to include in the generated image, but it is necessary to explain in text where in the image they want to place it, and it is not easy and cumbersome for users to express it accurately (and in a way that can be conveyed to the image generation AI).
[0005] The present invention has been made in consideration of the above-mentioned circumstances, and one of its objects is to provide an information processing device, a prompt generation method, and an image generation method that improve convenience when using image generation AI. [Means for solving the problem]
[0006] The present invention has been made to solve the above-mentioned problems, and an information processing device according to a first aspect of the present invention comprises a character detection unit that detects characters in an image, a word cluster determination unit that determines a group of characters consisting of one or more words from among the characters detected by the character detection unit as a word cluster, a position information acquisition unit that acquires position information indicating the position of the word cluster determined by the word cluster determination unit within the image, and a prompt generation unit that generates a prompt as an input sentence to an image generation model based on the word cluster and the position information.
[0007] In the information processing device, the prompt generation unit may generate a prompt that includes the word cluster and the location information in a sentence.
[0008] In the information processing device, the prompt generation unit may generate the prompt instructing generation of an image such that the image corresponding to the word cluster is placed at a position according to the position information.
[0009] In the above information processing device, the prompt generation unit may generate the prompt indicating an instruction to place the word cluster as a character at a position corresponding to the position information when a specific symbol is included in the word cluster.
[0010] In the information processing device, the character detection unit may detect characters by analyzing a handwritten trajectory included in the image.
[0011] The information processing device may further include a character removal unit that generates an image obtained by removing the characters detected by the character detection unit from the image, as an input image to the image generation model.
[0012] The information processing device may also include a generated image acquisition unit that acquires a generated image output from the image generation model by inputting the input image generated by the character removal unit and the prompt generated by the prompt generation unit into the image generation model.
[0013] Furthermore, a prompt generation method in an information processing device according to a second aspect of the present invention includes the steps of: a character detection unit detecting characters in an image; a word cluster determination unit determining a group of characters consisting of one or more words from among the characters detected by the character detection unit as a word cluster; a position information acquisition unit acquiring position information indicating the position of the word cluster determined by the word cluster determination unit within the image; and a prompt generation unit generating a prompt as an input sentence to an image generation model based on the word cluster and the position information.
[0014] Furthermore, an image generation method in an information processing device according to a third aspect of the present invention includes the steps of: a character detection unit detecting characters in an image; a word cluster determination unit determining a group of characters consisting of one or more words from among the characters detected by the character detection unit as a word cluster; a position information acquisition unit acquiring position information indicating the position of the word cluster determined by the word cluster determination unit within the image; a prompt generation unit generating a prompt as an input sentence to an image generation model based on the word cluster and the position information; a character removal unit generating an image from the image by removing the characters detected by the character detection unit as an input image to the image generation model; and a generated image acquisition unit acquiring a generated image to be output from the image generation model by inputting the input image generated by the character removal unit and the prompt generated by the prompt generation unit into the image generation model. [Effects of the Invention]
[0015] According to the above-described aspects of the present invention, it is possible to improve the convenience when using an image generation AI. [Brief explanation of the drawings]
[0016] [Figure 1] FIG. 1 is a diagram showing an example of the configuration of an information processing system according to an embodiment. [Figure 2] FIG. 1 is a block diagram showing an example of a hardware configuration of an information processing apparatus according to an embodiment. [Figure 3] FIG. 10 is a diagram showing an example of an image generation UI according to the embodiment. [Figure 4] 10A and 10B are diagrams showing an input example in an input area and an output example in an output area according to the embodiment. [Figure 5] FIG. 2 is a block diagram showing an example of a functional configuration related to image generation processing according to the embodiment. [Figure 6] 10 is a flowchart illustrating an example of a prompt generation process according to the embodiment. [Figure 7] 10 is a flowchart showing an example of an image generation process according to the embodiment. [Figure 8] 10A and 10B are diagrams showing another example of input in the input area and output in the output area according to the embodiment. [Figure 9] 10A and 10B are diagrams illustrating examples of prompts with and without specific symbols according to an embodiment. [Figure 10] 10 is a diagram showing an example of an image generated by handwriting input of the characters shown in FIG. 9 according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0017] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. [System Configuration] 1 is a diagram showing an example of the configuration of an information processing system according to this embodiment. The information processing system SYS includes an information processing device 10 and an image generation server 30.
[0018] The information processing device 10 is, for example, a clamshell (notebook) type PC (personal computer). The information processing device 10 has a first housing 10A and a second housing 10B, each of which is substantially rectangular and plate-shaped (for example, flat), joined (connected) via a hinge mechanism so that the first housing 10A and the second housing 10B can be opened and closed. The information processing device 10 also has a display 150 extending from the first housing 10A to the second housing 10B.
[0019] The display 150 is a flexible display that can be bent when the first housing 10A and the second housing 10B are opened or closed. An organic electroluminescence (EL) display or the like is used as the flexible display. For example, the display 150 can be used not only in a single-screen mode in which the entire screen area is a single screen, but also in a dual-screen mode in which the screen is divided into two screens: a first screen area 150A on the first housing 10A side and a second screen area 150B on the second housing 10B side. For example, when the information processing device 10 is used while placed on a desk, the second screen area 150B on the second housing 10B side becomes a substantially horizontal screen parallel to the surface of the desk.
[0020] Furthermore, a touch sensor is provided on the top (surface) of the screen of the display 150. The information processing device 10 can detect a touch operation on the screen area of the display 150. By opening the information processing device 10, the user can view the displays 150 provided on the inner surfaces of the first housing 10A and the second housing 10B and perform touch operations on the displays 150, thereby enabling the user to use the information processing device 10.
[0021] The image generation server 30 includes an image generation model 31. The image generation model 31 is an image generation model trained by, for example, a diffusion model, and generates an image from an input image or prompt (text). In other words, the image generation server 30 functions as an image generation AI (Artificial Intelligence) that uses the image generation model 31 to generate and output an image from an image or prompt (text). The image generation server 30 may be configured as a single server, or may be configured as a plurality of distributed servers.
[0022] In this embodiment, an image generation AI that generates images using the image generation model 31 is provided on a server, and an example configuration is described in which the information processing device 10 uses the functions of the image generation AI by communicating with the server, but the functions of the image generation AI may also be downloaded to the information processing device 10 and used locally.
[0023] [Hardware configuration of information processing device 10] The specific configuration of the information processing device 10 will be described below. 2 is a block diagram showing an example of the hardware configuration of an information processing device 10 according to this embodiment. The information processing device 10 includes a communication unit 11, a RAM (Random Access Memory) 12, a storage unit 13, a speaker 14, a display unit 15, a camera 16, and a control unit 18. These units are connected to each other so as to be able to communicate with each other via a bus or the like.
[0024] The communication unit 11 includes, for example, a plurality of Ethernet (registered trademark) ports, a plurality of digital input / output ports such as USB (Universal Serial Bus), and a communication device for wireless communication such as Bluetooth (registered trademark) or Wi-Fi (registered trademark).
[0025] For example, the communication unit 11 can communicate with a touch pen, a mouse, a touch pad, etc. using Bluetooth (registered trademark). The communication unit 11 can also communicate with the image generation server 30 shown in Fig. 1 by connecting to the Internet via wireless communication such as Wi-Fi (registered trademark) or wired communication such as Ethernet (registered trademark).
[0026] The RAM 12 is a volatile memory into which programs and data for processing executed by the control unit 18 are expanded and into which various data are stored or erased as appropriate. Note that, because the RAM 12 is a volatile memory, it does not retain data when power supply to the RAM 12 is stopped. Data that needs to be retained when power supply to the RAM 12 is stopped is transferred to the storage unit 13.
[0027] The storage unit 13 is a storage device configured to include one or more of an SSD (Solid State Drive), an HDD (Hard Disk Drive), a ROM (Residual Only Memory), a Flash-ROM, etc. For example, the storage unit 13 stores programs and setting data for a BIOS (Basic Input Output System), an OS (Operating System) and programs for applications that run on the OS, and various data used by applications. The speaker 14 outputs electronic sounds, voices, and the like.
[0028] The display unit 15 includes a display 150 and a touch sensor 155. As described above, the display 150 is a flexible display that can be bent in accordance with the opening and closing of the first housing 10A and the second housing 10B. Under the control of the control unit 18, the display 150 displays the desktop screen of the OS, windows of running applications, and the like.
[0029] Touch sensor 155 is provided on the screen of display 150 and detects touch operations on the screen. Touch operations include, for example, tap operations, slide operations, flick operations, swipe operations, pinch operations, etc. Furthermore, a finger or a touch pen can be used as an operating means for performing the touch operations.
[0030] Camera 16 includes a lens, an imaging element, etc. Camera 16 captures an image (a still image or a video) according to the control of control unit 18, and outputs data of the captured image.
[0031] Control unit 18 includes processors such as a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), and a microcomputer, and realizes various functions by executing programs (various programs such as a BIOS, an OS, and applications that run on the OS) stored in storage unit 13. For example, control unit 18 controls the screen mode of display 150, controls display in the screen area, and detects touch operations on touch sensor 155.
[0032] Next, a description will be given of image generation processing using an image generation model (image generation AI) in the information processing device 10. First, a UI (User Interface) will be described with reference to FIG.
[0033] [Image generation UI] 3 is a diagram showing an example of an image generation UI according to the present embodiment. In the example shown, the information processing device 10 controls the display 150 to a dual-screen mode with a first screen area 150A and a second screen area 150B. In the first screen area 150A and the second screen area 150B, a window of an application (hereinafter referred to as an "image generation app") having an image generation function that generates an image using an image generation model (image generation AI) is displayed.
[0034] In the second screen area 150B, a window of the image generation application including an input area R1 for inputting information about the image to be generated is displayed as a UI. In addition, in the first screen area 150A, a window of the image generation application including an output area R2 for displaying the generated image is displayed as a UI.
[0035] The user can hand-draw a sketch in input area R1 and can also hand-draw text on top of the sketch. When information processing device 10 is placed on a desk and used, second screen area 150B is a nearly horizontal screen, so having input area R1 on the second screen area 150B side makes it easy for the user to input text.
[0036] When hand-drawing a sketch or inputting characters, the user can mainly use a touch pen, but can also use a finger without a touch pen, or can use an external mouse (not shown), or can use an external touch pad (not shown).
[0037] The image generation UI shown in FIG. 3 is an example, and the layout and size of the input area R1 and output area R2, the screen area in which they are displayed, and the like are not limited to this example.
[0038] 4A and 4B are diagrams showing an example of input in the input area R1 and an example of output in the output area R2 according to this embodiment. (A) of Fig. 4 shows an example of input in the input area R1. As mentioned above, sketches and characters can be input into the input area R1.
[0039] In the input example shown in FIG. 4A, a "ship" and a "water surface" are handwritten by the user as a rough sketch. Here, the user wants to generate an image in which the area below the water surface is a rough ocean and multiple birds are flying in the sky, but it is difficult to represent a rough ocean in a picture and it takes time and effort to draw multiple birds. In this way, when it is difficult or time-consuming to represent it in a picture, the user can input it in text. In the example shown, the text "rough ocean" is handwritten in the area below the water surface and "birds" is handwritten in the area above the water surface in sky.
[0040] The image generation application detects characters from the input sketch and character images and generates a prompt based on the characters. The image generation application then inputs input data based on the generated prompt and sketch image to the image generation model 31, thereby acquiring a generated image generated by the image generation model 31. The image generation application then displays the acquired generated image in the output area R2.
[0041] Fig. 4(B) shows an example of a generated image displayed in output area R2 as an output example of output area R2. Note that the generated image shown in this figure is an image that imitates the generated image generated by image generation model 31. In the example of a generated image shown in Fig. 4(B), an image of a ship sailing on rough seas and multiple birds flying in the sky is generated in accordance with the example of input of the sketch and text shown in Fig. 4(A).
[0042] The generated image generated by the image generation app (the generated image displayed in output area R2) can be used in other applications, for example, by copying and pasting it into the window of another application. The other applications are applications that allow you to insert or paste images, such as applications with presentation creation functions, applications with document creation functions, or applications with image editing functions.
[0043] [Image generation processing functional configuration] Next, the image generation process for detecting characters to generate a prompt and generating an image based on the generated prompt and a sketch image will be described in detail.
[0044] 5 is a block diagram showing an example of a functional configuration related to image generation processing according to this embodiment. The illustrated image generation processing unit 110 is a functional configuration that performs image generation processing by the control unit 18 executing an image generation application. The image generation processing unit 110 includes a character detection unit 111, a word cluster determination unit 112, a position information acquisition unit 113, a prompt generation unit 114, a character removal unit 115, and a generated image acquisition unit 116.
[0045] The character detection unit 111 detects characters in an image including a sketch and characters input into the input area R1 (see FIG. 3). For example, the character detection unit 111 detects characters by analyzing the trajectories (strokes) of handwritten characters included in the image.
[0046] The word cluster determination unit 112 determines, as a word cluster, a group of characters consisting of one or more words from among the characters detected by the character detection unit 111. For example, the word cluster determination unit 112 determines a word cluster based on the distance between the characters detected by the character detection unit 111. The word cluster determination unit 112 also sets a bounding box in a range (partial region) surrounding the group of characters for each word cluster.
[0047] In the illustrated example, "birds" and "rough ocean" are determined as word clusters, and the bounding boxes of each word cluster are indicated by squares. Word cluster determination unit 112 outputs information about the bounding boxes for each word cluster to location information acquisition unit 113.
[0048] The position information acquisition unit 113 acquires position information indicating the position within the image of the word cluster determined by the word cluster determination unit 112. For example, the position information acquisition unit 113 divides the image area into nine 3x3 ranges, and determines the position of the bounding box of the word cluster within each divided range to acquire the position information. The position of the bounding box may be determined, for example, based on the position of the upper left corner of the bounding box, or based on the center position of the bounding box.
[0049] In the illustrated example, the location information acquisition unit 113 acquires "Middle Left" as the location information for "birds" and "Bottom Center" as the location information for "rough ocean." The location information acquisition unit 113 outputs the acquired location information for the word clusters to the prompt generation unit 114.
[0050] Based on the word cluster and the position information, the prompt generation unit 114 generates a prompt as an input sentence to the image generation model 31. For example, the prompt generation unit 114 generates a prompt that includes the word cluster and the position information in the sentence. In other words, the prompt generation unit 114 generates a prompt that instructs the generation of an image such that the image corresponding to the word cluster is placed at a position according to the position information acquired by the position information acquisition unit 113.
[0051] For example, the prompt generation unit 114 generates a prompt using a predetermined format. As an example of the predetermined format, a sentence such as "Place [word cluster] at [location information]" is set in advance. The prompt generation unit 114 generates a prompt by substituting characters (word clusters) detected from an image and their location information for the "word cluster" and [location information] parts of this sentence.
[0052] In the illustrated example, the sentence "[word cluster] definitely positioned at [position] of the image" is set as the prompt format. Prompt generator 114 generates the prompt "[birds] definitely positioned at [middle left] of the image, [rough ocean] definitely positioned at [bottom center] of the image."
[0053] These prompt examples are only examples and can be set using any format.
[0054] In this embodiment, the image area is divided into nine 3x3 areas to determine the positions of word clusters, but this is not limiting and the number of divisions into the image area can be set arbitrarily. Furthermore, the positions of word clusters may be defined by coordinates within the image.
[0055] The character removal unit 115 removes the characters detected by the character detection unit 111 from the image, and generates the resulting image as an input image to the image generation model 31.
[0056] The generated image acquisition unit 116 inputs the input image (a sketch image excluding the characters) generated by the character removal unit 115 and the prompt generated by the prompt generation unit 114 to the image generation model 31. The image generation model 31 generates and outputs an image based on the input image (a sketch image excluding the characters) and the prompt. In this way, the generated image acquisition unit 116 acquires the generated image output from the image generation model 31. The generated image acquisition unit 116 also displays the acquired generated image in the output area R2 (see FIG. 3).
[0057] [Prompt generation process behavior] Next, the operation of the prompt generation process executed by the image generation processing unit 110 will be described. FIG. 6 is a flowchart showing an example of a prompt generation process according to this embodiment.
[0058] (Step S101) The image generation processing unit 110 detects characters in an image from an image including a sketch and characters input into the input area R1 (see FIG. 3), and then proceeds to the processing of step S103.
[0059] (Step S103) The image generation processing unit 110 determines, as a word cluster, a group of characters consisting of one or more words from among the characters detected in step S101. For example, the word cluster determination unit 112 sets a bounding box (a partial area surrounding the group of characters) within the range of the group of characters for each word cluster. Then, the process proceeds to step S105.
[0060] (Step S105) Image generation processing unit 110 acquires position information indicating the position within the image of the word cluster determined in step S103. For example, position information acquisition unit 113 divides the image area into nine 3x3 ranges (see FIG. 5), and determines the position of the bounding box of the word cluster within each divided range to acquire position information. Then, the process proceeds to step S107.
[0061] (Step S107) Based on the word cluster and the position information, the image generation processing unit 110 generates a prompt as an input sentence to the image generation model 31. For example, the image generation processing unit 110 generates a prompt that includes the word cluster and the position information in the sentence (see FIG. 5).
[0062] [Image generation process] Next, the operation of the image generation processing executed by the image generation processing unit 110 will be described. 7 is a flowchart showing an example of image generation processing according to this embodiment. In this figure, the same reference numerals are used to designate processes corresponding to those in FIG. 6, and their explanations will be omitted. The processing in steps S101 to S107 is the prompt generation processing shown in FIG.
[0063] (Step S109) When the image generation processing unit 110 detects characters in the image from the image including the sketch and characters input to the input area R1 (see FIG. 3) in step S101, the image generation processing unit 110 removes the detected characters from the image and generates an input image for the image generation model 31. Then, the process proceeds to step S111.
[0064] (Step S111) The image generation processing unit 110 inputs the input image (image of the sketch excluding the characters) generated in step S109 and the prompt generated in step S107 to the image generation model 31. Then, the process proceeds to step S113.
[0065] Here, the image generation model 31 generates and outputs an image based on the input image (a sketch image excluding characters) and the prompt.
[0066] (Step S113) The image generation processing unit 110 acquires a generated image output from the image generation model 31. The image generation processing unit 110 also displays the acquired generated image in the output area R2 (see FIG. 3).
[0067] [Another example of input for image generation processing] Next, another input example of the image generation process in this embodiment will be described. In this embodiment, an image can be generated from input of a sketch and text as in the example shown in Figure 4, but an image can also be generated from input of only a sketch or only text. In the case of only a sketch, no text is detected from the image, so no prompt is generated, and only the sketch image is input to the image generation model 31. In the case of only text, a prompt is generated, and an image without a sketch and the generated prompt are input to the image generation model 31.
[0068] 8A and 8B are diagrams showing another example of input in the input area R1 and output in the output area R2 according to this embodiment. (A) of Fig. 8 shows an example of input in the input area R1, and (B) of Fig. 8 shows an example of output in the output area R2.
[0069] In the input example shown in (A) of Figure 8, no sketch is input, and the words "mountain," "green cars," "house," and "road" are handwritten. In this case, the generated prompt is, for example, "[mountain] definitely positioned at [top center] of the image, [green cars] definitely positioned at [middle center] of the image, [house] definitely positioned at [middle right] of the image, [road] definitely positioned at [bottom left] of the image."
[0070] In the example of a generated image shown in (B) of Fig. 8, an image is generated in which a house is located on the right side of a road on which multiple cars (green cars) are driving, and mountains tower in the distance, in response to the example of character input shown in (A) of Fig. 8. Note that the generated image shown in this figure is an illustration that imitates the generated image generated by image generation model 31.
[0071] [Text input example] Furthermore, when inputting characters, there are cases where users not only want to generate an image corresponding to the characters (word cluster) but also want to include the characters themselves in the image. Therefore, when a specific symbol is included in a word cluster, the image generation processing unit 110 (prompt generation unit 114) generates a prompt that instructs the user to place the word cluster as a character at a position corresponding to the position information. While the specific symbol can be set arbitrarily, in this embodiment, the specific symbol is set to ().
[0072] Figure 9 shows examples of prompts with and without a specific symbol. Figure 9(A) shows an example without a specific symbol, where the handwritten input characters are "Robot." The generated prompt is, for example, "[Robot] definitely positioned at [position] of the image." [Position] is where position information is entered.
[0073] (B) in Figure 9 is an example of a case where a specific symbol is included, and the handwritten input character is "(Robot)". Because the word cluster "(Robot)" includes the specific symbol (), the generated prompt is, for example, ""[Robot]" definitely positioned at [position] of the image, as letters." In other words, if the specific symbol () is included, the prompt states that it will be generated as a character.
[0074] Fig. 10 is a diagram showing an example of an image generated by handwriting input of the characters shown in Fig. 9. Fig. 10(A) shows an example of input in input area R1, and Fig. 10(B) shows an example of output in output area R2.
[0075] In the input example shown in Figure 10(A), no sketch is input, and the text [Robot] is handwritten at the top and "(Robot)" at the bottom. The prompt generated from this handwritten input (word cluster) is as shown in the example in Figure 9, and it is stated that "(Robot)" will be generated as text because it contains the specific symbol ().
[0076] In the example of the generated image shown in Fig. 10(B), an image including an image of a robot and the character "ROBOT" is generated in response to the character input example shown in Fig. 10(A). Note that the generated image shown in this figure is an image that is drawn to imitate the generated image generated by the image generation model 31.
[0077] As described above, the information processing device 10 according to this embodiment detects characters in an image and determines a group of characters consisting of one or more words from among the detected characters as a word cluster. The information processing device 10 also acquires position information indicating the position of the word cluster in the image, and generates a prompt as an input sentence to the image generation model 31 based on the word cluster and the position information.
[0078] As a result, the information processing device 10 detects characters from an image and automatically generates a prompt according to the position of a group of characters (word cluster), so that even if a person is not good at drawing pictures or if it is difficult or time-consuming to express something in a picture, the image generation model 31 can generate an image by inputting characters onto an image such as a sketch without drawing a picture. Thus, the information processing device 10 can improve the convenience when using image generation AI.
[0079] If an image containing characters is input as is (the characters are also input as images) to the image generation model 31, in conventional technology, the characters are either generated as an image or as the characters themselves, which occurs randomly. In contrast, in this embodiment, the information processing device 10 separates the characters from the image and generates prompts, thereby generating the intended image. Furthermore, because the information processing device 10 generates prompts based on word clusters and their position information, the user does not need to create a prompt that describes in writing where they want to place the characters on the image, which is convenient.
[0080] For example, the information processing device 10 generates a prompt that includes a word cluster and location information in a sentence.
[0081] This is convenient because the information processing device 10 can generate a prompt that includes a word cluster and its position information without the user having to create a prompt that explains in writing where on the image they want to place it.
[0082] More specifically, the information processing device 10 generates a prompt that instructs the generation of an image so that the image corresponding to the word cluster is placed at a position according to the position information.
[0083] This allows the information processing device 10 to generate a prompt that instructs the user to place the image intended by the user in an appropriate position in the generated image generated by the image generation model 31, based on the characters entered by the user.
[0084] Furthermore, when a specific symbol is included in a word cluster, the information processing device 10 generates a prompt indicating an instruction to arrange the word cluster as a character at a position according to the position information.
[0085] This makes the information processing device 10 convenient because, even when a user wants to include text itself in a generated image generated by the image generation model 31, the user only needs to enter a specific symbol and then input the text.
[0086] For example, the information processing device 10 detects characters by analyzing the trajectory of handwriting included in the image.
[0087] This allows the information processing device 10 to generate an appropriate prompt simply by the user inputting characters by hand.
[0088] Furthermore, the information processing device 10 generates an image from which the detected characters have been removed as an input image to the image generation model 31.
[0089] As a result, even if an image includes a sketch and text, the information processing device 10 can appropriately input the image by excluding the text from the image when inputting the image to the image generation model 31. Therefore, the information processing device 10 can prevent the text from being randomly generated as an image or as the text itself.
[0090] Furthermore, the information processing device 10 inputs an input image, from which characters have been removed, and the generated prompt to the image generation model 31, and thereby obtains a generated image output from the image generation model 31.
[0091] As a result, the information processing device 10 detects characters from an image, automatically generates a prompt according to the position of a group of characters (word cluster), and inputs the generated prompt and the image excluding the characters to the image generation model 31. Therefore, if a person is not good at drawing pictures or if it is difficult or time-consuming to express something in a picture, the image generation model 31 can generate an image by inputting characters onto an image such as a sketch without having to draw a picture. Therefore, the information processing device 10 can improve the convenience when using image generation AI.
[0092] In addition, the prompt generation method in the information processing device 10 according to this embodiment includes the steps of: the control unit 18 (image generation processing unit 110) detecting characters in an image; determining a group of characters consisting of one or more words from among the detected characters as a word cluster; acquiring position information indicating the position of the determined word cluster in the image; and generating a prompt as an input sentence to the image generation model 31 based on the word cluster and the position information.
[0093] As a result, the prompt generation method in the information processing device 10 detects characters from an image and automatically generates a prompt according to the position of a group of characters (word cluster), so that people who are not good at drawing pictures or who find it difficult or time-consuming to express something in a picture can have the image generation model 31 generate an image by inputting characters onto an image such as a sketch without having to draw a picture. Therefore, the prompt generation method in the information processing device 10 can improve the convenience when using image generation AI.
[0094] Furthermore, according to the prompt generation method in the information processing device 10, characters are separated from images to generate prompts, so when an image containing characters is input as is (the characters are also input as images) to the image generation model 31, the intended image is generated without randomly generating either the character portion as an image or the characters themselves. Furthermore, according to the prompt generation method in the information processing device 10, a prompt is generated based on word clusters and their position information, so the user does not need to create a prompt that explains in writing where they want to place the image, which is convenient.
[0095] Furthermore, the image generation method in the information processing device 10 according to this embodiment includes the steps of: detecting characters in an image by the control unit 18 (image generation processing unit 110); determining a group of characters consisting of one or more words from among the detected characters as a word cluster; acquiring position information indicating the position of the determined word cluster in the image; generating a prompt as an input sentence to the image generation model 31 based on the word cluster and the position information; generating an image from which the detected characters have been removed as an input image to the image generation model 31; and acquiring a generated image output from the image generation model 31 by inputting the generated input image and the generated prompt into the image generation model 31.
[0096] As a result, the image generation method in the information processing device 10 detects characters from an image, automatically generates a prompt according to the position of a group of characters (word cluster), and inputs the generated prompt and the image excluding the characters into the image generation model 31. Therefore, for people who are not good at drawing pictures or when it is difficult or time-consuming to express something in a picture, they can input characters onto an image such as a sketch without having to draw a picture, and have the image generation model 31 generate an image. Therefore, the image generation method in the information processing device 10 can improve the convenience when using image generation AI.
[0097] Although the embodiments of the present invention have been described above in detail with reference to the drawings, the specific configurations are not limited to the above-described embodiments, and the present invention also includes designs that do not deviate from the gist of the present invention. For example, the configurations described in the above-described embodiments can be combined in any manner.
[0098] Furthermore, in the above-described embodiment, an example has been described in which characters are input by handwriting on a sketch, but the input of characters is not limited to handwriting and can also be done using a keyboard or the like.
[0099] In the above-described embodiment, the display 150 is a single display provided across the first housing 10A and the second housing 10B, but this is not limiting. For example, the information processing device 10 may be provided with two displays in total, one on each of the first housing 10A and the second housing 10B. Furthermore, the information processing device 10 may be provided with a display only on one of the first housing 10A and the second housing 10B (for example, only on the first housing 10A).
[0100] Furthermore, the information processing device 10 is not limited to a clamshell type (notebook type) PC, but may be, for example, a tablet type PC or a desktop type PC.
[0101] In the above-described embodiment, an example of a touch panel display in which an input unit (touch sensor) and a display unit (display) are integrated has been described, but a non-touch panel display that does not have an input unit (touch sensor) may also be used. In this case, a touch pad, a mouse, or the like may be used as an input device that accepts operations instead of the input unit (touch sensor).
[0102] The information processing device 10 described above includes an internal computer system. A program for implementing the functions of each component of the information processing device 10 may be recorded on a computer-readable recording medium, and the program may be loaded into a computer system and executed to perform processing in each component of the information processing device 10. Here, "loading a program recorded on a recording medium into a computer system and executing it" includes installing the program into a computer system. The term "computer system" here includes hardware such as an OS and peripheral devices. The term "computer system" may also include multiple computers connected via a network, including the Internet, a WAN, a LAN, a dedicated line, or other communication lines. The term "computer-readable recording medium" refers to portable media such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, or a storage device such as a hard disk built into a computer system. The recording medium storing the program may also be a non-transitory recording medium such as a CD-ROM.
[0103] The recording medium also includes internal or external recording media accessible from a distribution server for distributing the program. The program may be divided into multiple parts, downloaded at different times, and then combined by each component of the information processing device 10, or each divided program may be distributed by a different distribution server. Furthermore, the term "computer-readable recording medium" also includes a medium that stores a program for a certain period of time, such as volatile memory (RAM) within a computer system that serves as a server or client when a program is transmitted over a network. The program may also be a medium that realizes part of the above-described functions. Furthermore, the program may be a so-called differential file (differential program) that can realize the above-described functions in combination with a program already stored in the computer system.
[0104] Furthermore, some or all of the functions of the information processing device 10 in the above-described embodiment may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each function may be individually implemented as a processor, or some or all of the functions may be integrated into a processor. The integrated circuit method is not limited to LSI, and may be implemented using a dedicated circuit or a general-purpose processor. Furthermore, if an integrated circuit technology that can replace LSI emerges due to advances in semiconductor technology, an integrated circuit based on that technology may be used. [Explanation of symbols]
[0105] 10 Information processing device, 10A First housing, 10B Second housing, 30 Image generation server, 31 Image generation model, 11 Communication unit, 12 RAM, 13 Memory unit, 14 Speaker, 15 Display unit, 16 Camera, 150 Display, 150A First screen area, 150B Second screen area, 155 Touch sensor, 18 Control unit, 110 Image generation processing unit, 111 Character detection unit, 112 Word cluster determination unit, 113 Position information acquisition unit, 114 Prompt generation unit, 115 Character removal unit, 116 Generated image acquisition unit, SYS Information processing system
Claims
1. a character detection unit that detects characters in an image; a word cluster determination unit that determines a group of characters consisting of one or more words from among the characters detected by the character detection unit as a word cluster; a position information acquisition unit that acquires position information indicating a position within the image of the word cluster determined by the word cluster determination unit; a prompt generation unit that generates a prompt as an input sentence to an image generation model based on the word cluster and the position information; Equipped with The word cluster determination unit when determining a group of characters comprising the one or more words as the word cluster, determining a group of characters that are relatively close to each other based on a distance between the characters detected by the character detection unit as the word cluster; Information processing device.
2. The prompt generation unit generating a prompt including the word cluster and the location information in a sentence; The information processing device according to claim 1 .
3. The prompt generation unit generating the prompt to instruct generation of an image such that the image corresponding to the word cluster is placed at a position according to the position information; The information processing device according to claim 2 .
4. The prompt generation unit generating the prompt indicating an instruction to arrange the word cluster as a character at a position according to the position information when the word cluster includes a specific symbol; The information processing device according to claim 2 .
5. The character detection unit detecting characters by analyzing handwritten trajectories included in the image; The information processing device according to claim 1 .
6. a character removal unit that removes the characters detected by the character detection unit from the image and generates the image as an input image to the image generation model; The information processing device according to claim 1 , comprising:
7. a generated image acquisition unit that acquires a generated image output from the image generation model by inputting the input image generated by the character removal unit and the prompt generated by the prompt generation unit into the image generation model; The information processing device according to claim 6 , comprising:
8. A method for generating a prompt in an information processing device, comprising: a character detection unit detecting characters in the image; a step in which a word cluster determination unit determines, as a word cluster, a group of characters consisting of one or more words from among the characters detected by the character detection unit; a position information acquiring unit acquiring position information indicating a position within the image of the word cluster determined by the word cluster determining unit; a prompt generation unit generating a prompt as an input sentence to an image generation model based on the word cluster and the position information; Including, The word cluster determination unit when determining a group of characters comprising the one or more words as the word cluster, determining a group of characters that are relatively close to each other based on a distance between the characters detected by the character detection unit as the word cluster; The prompt generation method.
9. An image generation method in an information processing device, comprising: a character detection unit detecting characters in the image; a step in which a word cluster determination unit determines, as a word cluster, a group of characters consisting of one or more words from among the characters detected by the character detection unit; a position information acquiring unit acquiring position information indicating a position within the image of the word cluster determined by the word cluster determining unit; a prompt generation unit generating a prompt as an input sentence to an image generation model based on the word cluster and the position information; a character removal unit generating an image obtained by removing the characters detected by the character detection unit from the image as an input image to the image generation model; a generated image acquisition unit inputting the input image generated by the character removal unit and the prompt generated by the prompt generation unit into the image generation model, thereby acquiring a generated image output from the image generation model; Including, The word cluster determination unit when determining a group of characters comprising the one or more words as the word cluster, determining a group of characters that are relatively close to each other based on a distance between the characters detected by the character detection unit as the word cluster; Image generation method.
Citation Information
Patent Citations
Information processing device, information processing method, and information processing program
JP2024074020A
Sentence generator, sentence generation system, method for generating sentence, and program
JP2024159480A
Image generation device, prompt creation support device, program and application program
JP2024120131A