Information processing method, information processing device, and program
The information processing apparatus addresses the challenge of mismatched generative AI outputs by deriving and setting negative prompts, ensuring content generation aligns with user expectations and template styles.
Patent Information
- Application Number
- JP2024002654
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2025-07-24
- Estimated Expiration
- 2044-01-11
AI Technical Summary
Users face challenges in efficiently generating images with generative AI that match the style and atmosphere of a selected template, often requiring repetitive input of prompts due to mismatched expectations.
An information processing apparatus that derives a specific character string as a negative prompt based on user information or template characteristics to guide generative AI on what not to generate, thereby simplifying the prompt-setting process.
Facilitates easy and efficient generation of content that aligns with user expectations by automatically setting appropriate negative prompts, reducing user effort and improving content relevance.
Smart Images

Figure 2025109013000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to prompt setting for image generation AI.
Background Art
[0002] As one of the functions of software for creating posters, flyers, etc., there is a function that allows a user to select a desired template from various prepared templates and insert arbitrary characters or an image taken by the user himself / herself into the template. In this regard, Patent Document 1 discloses a technique that enables searching for an image suitable for a template by extracting words from the property information of objects included in the template to create search keywords and searching from a group of separately prepared images.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] For example, among software for creating posters and the like, there are those that have a function of generating images using generative AI (Artificial Intelligence). This image generation function is a function in which when a user inputs words or a sentence as a prompt to the generative AI, the AI automatically generates an image based on the input prompt (words or sentence). Here, for example, assume that the template of the poster to be used is in an illustration style. When inserting an image into an illustration-style template, the user will expect an illustration-style image and input an arbitrary prompt to the generative AI. However, if the prompt input by the user is not appropriate, the generative AI may generate a realistic image, for example, and it may not match the style and atmosphere of the template to be used. In such a case, the user has to repeat the image generation by the generative AI by trying to input another prompt, etc., which is time-consuming for the user. Thus, thinking of an appropriate prompt for obtaining desired content and inputting it to the generative AI has been a difficult and time-consuming task for the user.
Means for Solving the Problem
[0005] The information processing apparatus according to the present disclosure is an information processing apparatus for causing a generative AI to generate content, and includes a derivation means for deriving a specific character string based on information obtained about a user, and a setting means for setting the specific character string derived by the derivation means as a negative prompt for specifying what kind of content should not be generated by the generative AI.
Advantages of the Invention
[0006] According to the present disclosure, an appropriate prompt for obtaining desired content can be easily set.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
[0008] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the present invention according to the claims, and not all combinations of the features described in the present embodiments are essential for the solution means of the present invention.
[0009] [Embodiment 1] [System Configuration] FIG. 1 is a diagram showing a configuration example of a client-server system according to the present embodiment. This system is composed of a server 101 that provides a content generation service and a client 102 that uses the service. The server 101 and the client 102 communicate with each other via the Internet 104 from their respective networks 103. The server 101 has a backend application 105. The client 102 has an operating system 106 and a frontend application 107.
[0010] <Hardware Configuration> FIG. 2 is a diagram showing an example of a hardware configuration common to the server 101 and the client 102, which are information processing devices. The user interface 201 is an operation display unit that inputs and outputs information and signals, and is composed of a display, a keyboard, a mouse, buttons, a touch panel, etc. The network interface 202 connects to a network such as a LAN and communicates with other computers and network devices. The communication method may be either wired or wireless. The CPU 203 is a central processing unit that executes programs read from a ROM 204, a RAM 205, a secondary storage device 206, etc. The ROM 204 is a non-volatile memory in which embedded programs and data are recorded. The RAM 205 is a volatile memory that provides a temporary memory area. The secondary storage device 206 is a large-capacity storage device typified by an HDD or a flash memory. A computer not equipped with these hardware components can also be connected and operated from another computer by means of remote desktop, remote shell, etc. Each part is connected via an input / output interface 207.
[0011] <Software Configuration> FIG. 3 is a functional block diagram showing an example of the software configuration of each of the front - end application 107 and the back - end application 105. The front - end application 107 has a GUI control unit 301, a negative prompt derivation unit 302, a prompt setting unit 303, and a request processing unit 304. The back - end application 105 has a response processing unit 311 and a content generation unit 312. Note that one of the front - end application 107 and the back - end application 105 may be configured to include all of the above units 301 to 304, 311, and 312.
[0012] The GUI control unit 301 controls a GUI (Graphical User Interface) for presenting information to the user or for the user to input instructions. Specifically, it performs controls such as displaying a UI screen and accepting user input. Specific examples of user input include selecting a template, instructing a content insertion method, and accepting a character string as a prompt (instruction information consisting of words or sentences) to be input to the generation AI. Prompts include positive prompts and negative prompts. A positive prompt is a prompt that specifies the direction (elements that one wants to generate) of what kind of content should be generated for the generation AI. For example, when "reindeer" is input as a positive prompt to an AI that generates images (image generation AI), the image generation AI generates an image of a reindeer. On the other hand, a negative prompt is a prompt that specifies the direction (elements that one wants to exclude) of what kind of content should not be generated for the generation AI. FIG. 4 is a table summarizing an example of a general negative prompt input to an image generation AI. For example, when "low quality" is input as a negative prompt to an image generation AI, the image generation AI will not generate low - quality images.
[0013] The negative prompt derivation unit 302 derives the string to be used as the above-mentioned negative prompt among the prompts input to the generative AI based on the information obtained regarding the user. The derivation method will be described later.
[0014] The prompt setting unit 303 sets the string input by the user via the GUI and the string derived by the negative prompt derivation unit 302 as the positive prompt and the negative prompt, respectively.
[0015] The request processing unit 304 performs processing to request the server 101 to generate content. At the time of this request, the positive prompt and the negative prompt set by the prompt setting unit 303 are also sent together. Furthermore, the request processing unit 304 also performs processing to receive the content generated by the server 101 in response to the request.
[0016] The response processing unit 311 performs response processing to receive a request from the client 102 and to send the generated content to the requesting client 102 in response to the received request.
[0017] The content generation unit 312 is a generative AI that generates content such as images or text using the positive prompt and the negative prompt as inputs. As an image generation AI that generates images as content, for example, those such as "Stable Diffusion" and "Midjourney" are known. The generative AI is a learned model (content generation model) obtained by performing machine learning on various data by means such as deep learning so that the target content can be obtained.
[0018] Next, taking the poster creation software as the front-end application 107 as an example, the operation flows of the client 102 and the server 101 will be described. Note that the poster creation software is an example of the software installed on the client 102, and is not limited thereto. For example, the present disclosure is applicable to software for obtaining various work products such as photo album creation software and postcard creation software.
[0019] <Client-side operation> FIG. 5 is a flowchart showing the operation flow in the client 102 according to the present embodiment. This flowchart starts based on the activation of the poster creation software based on a user instruction. In the following description, the symbol "S" means step.
[0020] In S501, the GUI control unit 301 displays an editing UI screen (hereinafter referred to as "poster editing screen") corresponding to the poster creation software on the display of the user interface 201. An example of the poster editing screen is shown in FIG. 6(a). The poster editing screen 600 in FIG. 6(a) is composed of a template list pane 610, a template editing pane 620, and a content editing pane 630. The GUI control unit 301 displays a list of various templates according to the purpose and use on the template list pane 610. In the example of FIG. 6(a), three templates (611, 612, 613) prepared in advance are displayed on the template list pane 610.
[0021] In S502, the GUI control unit 301 accepts a user selection for a specific template among the templates listed in the template list pane 610. FIG. 6(b) shows a screen change when the user selects the template 612, and a highlight frame 640 indicating the selected state is added. Further, the template 612 selected by the user is displayed as the template 621 to be edited in the template editing pane 620. Now, in the template 621 to be edited, an image setting area 622 for the user to set an arbitrary image and a text setting area 623 for the user to set arbitrary text are arranged. Hereinafter, the image setting area and the text setting area will be collectively referred to as the "content setting area".
[0022] In S503, the GUI control unit 301 determines whether or not any of the content setting areas in the template to be edited displayed in the template editing pane 620 has been pressed. When a press on the content setting area is detected, the GUI control unit 301 subsequently executes S504. If no press is detected, it waits for a predetermined time to elapse and then determines whether or not a press has occurred.
[0023] In S504, the GUI control unit 301 displays a pop-up screen for allowing the user to select an addition method for the content according to the pressed content setting area, and accepts a user selection as to whether to manually add or automatically generate the target content. FIG. 7(a) shows an example in the state of FIG. 6(b) where the user operates the pointer to press the image setting area 622, and a pop-up screen for selecting whether to automatically generate or manually add the image to be added is displayed. Currently, as options for the image addition method, buttons for “Automatic Generation” and “Select from a folder” are displayed. Here, the “Automatic Generation” button is for the case of causing the generation AI to generate an image, and the “Select from a folder” button is for the case where the user designates and uploads an image from an arbitrary folder. If the text setting area 623 was pressed in S503, buttons for “Automatic Generation” and “Manual Input” will be displayed as options for the text addition method. In this case, the “Automatic Generation” button is for the case of causing the generation AI to generate text, and the “Manual Input” button is for the case where the user manually inputs an arbitrary character string. When the pressing of the button for manual addition is detected, the GUI control unit 301 subsequently executes S505, and when the pressing of the button for automatic generation is detected, the GUI control unit 301 subsequently executes S507.
[0024] In S505, the GUI control unit 301 accepts the user's specification of content. For example, when it is detected in S504 that the "Select from a folder" button for manually adding an image is pressed, it accepts the specification of desired image data from an arbitrary folder. The user can specify a desired image from among the images previously held in the folder, such as by taking the image himself / herself or acquiring it via the Internet 104, by an operation such as drag & drop. Also, for example, when it is detected in S504 that the "Manual Input" button for manually adding text is pressed, it accepts the specification of a character string via an input field (not shown) for directly inputting a desired character string, such as by displaying the input field.
[0025] In S506, the GUI control unit 301 inserts the content specified by the user accepted in S505 into the content setting area pressed in S503 in the selected template. FIG. 7(b) is a diagram showing an example of a state in which the image 710 specified by the user is inserted into the image setting area 622. When the content is inserted, an operation field 720 for the user to perform an editing operation on the inserted content is displayed in the content editing pane 630. Now, in FIG. 7(b), the inserted image 710 can be edited. In this example, a field 721 for instructing whether to automatically correct the image to be edited, a control bar 722 for adjusting the brightness of the image, and a control bar 723 for adjusting the contrast of the image are provided. If the content to be edited is text, an operation field (not shown) for adjusting the font, size, color, etc. of the characters will be further displayed. The user can perform corrections, etc. on the content by appropriately performing necessary operations in the operation field corresponding to the content to be edited.
[0026] S507 to S513 are processes for automatically generating and inserting content into the content setting area pressed by the user by the generative AI. Figures 8(a) and 8(b) are examples of the UI screen when the user selects the Christmas poster template 611 in the state of Fig. 6(a), and then presses the image setting area 622 of the template 621 to be edited to select automatic image generation. Hereinafter, while appropriately referring to the UI screen example, the process until the generative AI automatically generates content and completes the template will be described.
[0027] First, in S507, the GUI control unit 301 receives an input of a character string such as a word representing the elements of the content to be generated by the generative AI. Since English is generally used as the language for the prompt, it is also in English notation in this embodiment. However, the language used for the prompt depends on the generative AI and is not limited to English. Here, it is assumed that the words "Reindeer" and "Christmas" are input via an input field (not shown).
[0028] In S508, the prompt setting unit 303 sets the character string such as the word input by the user in S507 as a positive prompt. Now, since the two words "Reindeer" and "Christmas" are input by the user, they will be set as positive prompts.
[0029] In S509, the negative prompt derivation unit 302 derives a character string such as a word representing an element of content that the user does not want the generative AI to generate, based on the template selected by the user. As a method of derivation based on a template, for each template displayed in the template list, a character string for a negative prompt corresponding to its characteristics is given as metadata, and the metadata of the template related to the user selection is referred to. Alternatively, a table associating each template with a character string for a negative prompt may be prepared in advance and the table may be referred to. Furthermore, it may be derived using a learned model that estimates a character string inappropriate for the impression of the template selected by the user. The learned model used for this estimation can be obtained by learning a large amount of teacher data in which a template and a word or the like inappropriate for the impression of the template are paired. Note that at the time of estimation, the overall impression of the target template may be estimated, or the overall impression of the template may be estimated after estimating the impressions of each content such as an image or text included in the target template. In the example of Fig. 8(a), a template in an illustration style with a Christmas theme is selected. In this case, even if a realistic image of a reindeer such as a photographed image of a reindeer in the natural world is generated, it is not suitable for the overall illustration-style template. Therefore, words such as "Realistic" are associated in advance with the illustration-style template by the method described above. As a result, just by the user selecting a desired template, a negative prompt suitable for the template is automatically acquired and set, so the user can save the trouble of thinking about and setting the negative prompt by themselves. Note that when an illustration-style template is selected, a character string such as "Illustration style" can be set by the user himself / herself as a positive prompt to cause the generative AI to generate an illustration-style image. However, with the method of this embodiment, the user does not need to think about and set such a positive prompt one by one, so more labor can be saved.
[0030] In S510, the prompt setting unit 303 sets the character string such as the word derived in S509 as a negative prompt. In the above example, the word "Realistic" will be automatically set as the negative prompt. Note that, as shown in FIG. 11(b) described later, the user may be able to set the word to be set as the negative prompt from among the group of words derived in S509 via the GUI.
[0031] In S511, the request processing unit 304 sends a content generation request to the server 101 that provides the content generation service. This content generation request includes the positive prompt set in S508 and the information of the negative prompt set in S510. In response to this content generation request, the server 101 automatically generates content based on the positive prompt and the negative prompt. Details of the automatic content generation in the server 101 will be described later.
[0032] In S512, the request processing unit 304 receives the content generated based on the content generation request from the server 101.
[0033] In S513, the GUI control unit 301 inserts the content received in S512 into the content setting area pressed in S503 in the selected template. In the example of Fig. 8(a), an illustration-style reindeer image 800 is inserted into the image setting area 622, and an operation field 810 for the user to perform an editing operation on the image 800 is displayed in the content editing pane 630. The operation field 810 is provided with two sub-fields 820 and 830 indicating the positive prompt and negative prompt used for generating the inserted image 800, and further a sub-field 840 for displaying only the generated image. Fig. 8(b) is a diagram showing a specific example when "Realistic" is not set as the negative prompt, and a realistic reindeer image 801 like a natural photography is generated by the generation AI. Note that the user may be able to finely adjust the color and the like using the above-mentioned control bar or the like for the image displayed in the sub-field 840.
[0034] In S514, it is determined whether all the content to be set has been set for the template selected in S502. If there is unset content, the process returns to S503 and continues. On the other hand, if all the content has been set, this process ends.
[0035] The above is the description of the operation on the client side. In the flow of Fig. 5 described above, before the content generation request is sent, a step may be provided for the user to confirm on the UI screen whether to apply the negative prompt automatically set in S510 as it is. Also, the user may add or change a character string in the sub-fields 820 and 830 to cause the generation AI to generate an image again.
[0036] The above-described method is applicable to various cases. For example, in a case where there are implicit rules for a certain traditional dish (for example, ingredient X should not be used), a string such as "X as an ingredient that should not be used" can be associated with the template for the traditional dish. As a result, even for a user who is not aware of the rules regarding the traditional dish, when selecting the template for it, the string representing ingredient X is automatically set as a negative prompt. Therefore, it is possible to prevent the content generation AI from accidentally creating content that includes ingredient X.
[0037] <Server-side operations> FIG. 9 is a flowchart showing the flow of operations in which the backend application 105 of the server 101 generates content in response to a content generation request (S511) from the client 102. In the following description, the symbol "S" means step.
[0038] In S901, the response processing unit 311 receives a content generation request from the client 102. In S902, the content generation unit 312 obtains a positive prompt and a negative prompt from the content generation request received in S901. In S903, the content generation unit 312 generates content using the positive prompt and the negative prompt obtained in S902 as inputs. In S904, the response processing unit 311 transmits the data of the content generated in S903 to the client 102 that is the request source.
[0039] The above is the description of the server-side operations. In this embodiment, the backend application 105 is configured to include a generation AI, but it is not limited to this. For example, the generation AI may be external to the backend application 105, and the backend application 105 may call the external generation AI to respond to a content generation request.
[0040] <Variations of the string derivation method> In the above example, the string used as the negative prompt was derived by pre-associating a specific string for each template. However, the method for deriving the string of the negative prompt is not limited to this. Below, variations of the method for deriving the string for the negative prompt will be described.
[0041] <<Derivation Based on User Attributes>> Generally, in the case of software that realizes a function such as poster creation as a cloud-based service, users often register an account in advance and log in to use the software. When a user registers an account, attribute information such as the user's name, sex, nationality, region, language, and hobbies are also registered. If the software is provided in countries around the world, it will be used by users of various nationalities, but different nationalities have different cultures and customs, and the same gesture may be perceived differently depending on the country. For example, in the case of the peace sign, which is a type of body language, it gives a good impression in some countries and a bad impression in others. Therefore, if the nationality indicated by the attribute information of the logged-in user is a country where the peace sign gives a bad impression, the "Peace sign" is derived as a string of characters for the negative prompt. As a specific method of derivation, strings representing gestures that are taboo for each country are registered in a database, and when a user logs in, the nationality is inquired from the database, and the string registered in association with the user's country is obtained. 10(a) and 10(b) are examples of UI screens when a user selects a template 611 on the UI screen of FIG. 6(a) and presses an image setting area 622 of an edit target template 621 to select automatic generation of an image. FIG. 10(a) is an example of a case where the nationality of the logged-in user is a country where the peace sign does not give a bad impression, and FIG. 10(b) is an example of a case where the nationality of the logged-in user is a country where the peace sign gives a bad impression. In both cases, the character string "woman is laughing" is set as a positive prompt, but in FIG. 10(b), "Peace sign" is further set as a negative prompt. As a result, in the example of FIG. 10(a), the generation AI generates an image 1000 of a woman laughing with a peace sign, whereas in the example of FIG. 10(b), the generation AI generates an image 1001 of a woman laughing without making a peace sign. In this way, by setting a negative prompt according to the nationality of the user, it is possible to prevent content that is not suitable for the user from being generated.Here, nationality is used as the user's attribute information. However, for example, other user attributes such as a region including multiple countries or the language used may also be used. In this case, a string representing a gesture or the like that is taboo for each attribute such as a region or a language is registered in a database, and a negative prompt may be set by deriving the string according to these attributes for the logged-in user. Furthermore, instead of these attributes at the time of account registration, for example, the above attributes of the user may be estimated from information after account registration such as the user's behavior history or IP address, and the string of the negative prompt may be derived using the estimated attribute information. Note that even for a user of a nationality that gives a bad impression of the peace sign, if a string of an expression that intentionally insults with a positive prompt is input, the positive prompt may be prioritized so that "Peace sign" is not set as the negative prompt.
[0042] ≪Derivation Based on User Selection for Content≫ Next, a method for deriving a negative prompt string based on a user's selection of the content within the template will be described. Specifically, the impression of the image selected by the user from among the images generated by the generative AI is estimated using a pre-trained model for impression estimation (impression estimation model), and a string corresponding to the antonym of the string representing the estimated impression is obtained from a dictionary database. Figure 11(a) is a diagram for explaining how a negative prompt is set based on the image selected by the user from among a plurality of images generated by an image generation model which is a generative AI. Now, "Halloween pumpkin" is set as the positive prompt. Then, using this as an input, the image generation model 1110 has generated three images 1120, 1121, and 1122 of pumpkins deformed in a Halloween style. Among these three images, there is an image 1120 with a "cute" impression, as well as images 1121 and 1122 with a "scary" impression. Now, assume that the user has selected the image 1120 of a cute pumpkin as the image to be inserted into the template. In this case, when the image 1120 selected by the user is input into the impression estimation model, the impression "cute" is estimated. Next, the word "pretty (cute)" representing the estimated impression is input into the dictionary database. As a result, words such as "scary", "eerie", "weird", and "frightening" which represent an impression opposite to or deviating from "cute" are obtained. Figure 12(a) is a diagram showing how the impression "pretty" obtained by inputting the image 1120 selected by the user into the impression estimation model 1200 is further input into the dictionary database, and the opposite impression "scary" of "pretty" is derived. Figure 11(b) is an example of a UI screen where the user selects and sets a negative prompt from the group of words thus derived. The GUI control unit 301 of the front-end application 107 displays a candidate dialog 1130 including the derived group of words. The user selects a word from among the group of words displayed in the candidate dialog 1130 that represents an impression that they do not want to be generated. In this example, "eerie" is selected and "eerie" is set as the negative prompt.When "Halloween bat" is set as a positive prompt and input into the image generation model 1110, images 1140 and 1141 of bats with a "cute" impression will be generated. In this way, the string of the negative prompt may be derived from the image selected by the user using the impression estimation model and the dictionary database. Here, one word is selected by the user from the derived word group, but one or more of the derived words may be directly set as the negative prompt.
[0043] Also, the word "pretty" representing the impression estimated from the image selected by the user may be added to the positive prompt. Also, as shown in Fig. 12(b), a learned model (antonym estimation model) 1210 capable of estimating the antonyms of the words representing the impression of the image may be used to more directly derive the string of the negative prompt from the image selected by the user. The antonym estimation model in this case can be obtained by learning a large amount of teacher data in which an image and a word not familiar with its impression form a pair. Note that the images used for the teacher data during learning may be those generated by the generative AI or those taken by a person with a camera.
[0044] In the above example, the string of the negative prompt was derived based on the image 1120 selected by the user. However, as shown in Fig. 12(c), it may be derived based on the images 1121 and 1122 not selected by the user. Now, both of the images 1121 and 1122 not selected by the user are images with a "scary" impression. Therefore, by inputting these into the impression estimation model 1200, "scary" representing the impression common to both images can be estimated. Note that the number of images input into the impression estimation model may be three or more or one.
[0045] Also, although an example of derivation based on the images not selected by the user among the images generated by the generative AI has been described, it is not limited to this. For example, it may be derived based on the images not selected by the user among the images for inserting templates prepared in advance by the poster creation software.
[0046] Furthermore, for images not selected by the user, as bad examples, additional learning may be associated with the user using a method such as LoRA (Low-Rank Adaptation), and the generative AI may be updated. As a result, when each user uses it after the next time, by using the generative AI updated on a per-user basis, it becomes difficult to generate images that do not match the preferences of the logged-in user.
[0047] Note that as the impression estimation model and the antonym estimation model, a pre-trained model learned by, for example, a deep learning method is assumed, similar to the content generation model, but it is not limited to this.
[0048] <Other Variants> The user may automatically update the automatically set negative prompt during the process of editing the template. Figures 13(a) and (b) are diagrams for explaining the automatic update of the negative prompt. Now, assume that "realistic" and "pretty" are pre-associated as negative prompts for template 1300. And the user instructs the addition of an image by automatic generation using "Halloween pumpkin" as the positive prompt. As a result, an illustration-style and "scary" impression pumpkin image 1301 is generated by the generative AI and placed at the upper right position of template 1300. In Figure 13(a), V1 and V2 are the results of converting "Realistic" and "pretty" into vector representations using a method such as Word2Vec. Note that the method of converting words into vector representations is not limited to the Word2Vec method and other methods may also be used. Next, assume that the user inserts the pumpkin image 1302 into the lower left position of template 1300 by means of drag & drop or the like. And this pumpkin image 1302 is input into the aforementioned impression estimation model 1200, and the impression "cute" is estimated. In Figure 13(a), V is the result of converting the estimated impression "cute" into a vector representation. In this case, the cosine similarity S1 between V and V1 and the cosine similarity S2 between V and V2 are obtained and compared, and the value of S2 is larger. This is because "pretty" is closer to the meaning of "cute" than "realistic". When the impression has a high similarity to the impression of the image added by the user himself / herself, it is not suitable as a negative prompt. Therefore, as shown in Figure 12(b), "pretty" is deleted from the negative prompt. As a result, the negative prompt for template 1300 is updated to only "realistic". In this way, the negative prompt may be automatically updated during the process of the user editing the template. As a result, when the user designates "Halloween bat" as the positive prompt and instructs the automatic generation of an image, an image of a bat with a "cute" impression will also be generated.By automatically updating the negative prompt according to the user's editing operation in this way, it is possible to cause the content generation AI to generate content suitable for the impression changed by the editing.
[0049] In addition, in the above-described embodiment, although the user can confirm the automatically set negative prompt on the UI screen, it may be applied in a form that is not visible to the user without being displayed on the UI screen.
[0050] Also, the expression of the negative prompt automatically derived and set according to the above-described embodiment may be made changeable by the user, for example, changing “cute” to “pretty”.
[0051] (Other Examples) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiment to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (for example, ASIC) that realizes one or more functions.
[0052] In addition, the disclosure of this embodiment includes the following configurations and methods.
[0053] [Configuration 1] An information processing apparatus for causing a content generation AI to generate content, a derivation means for deriving a specific character string based on information obtained about a user; a setting means for setting the specific character string derived by the derivation means as a negative prompt for specifying what kind of content should not be generated for the content generation AI; An information processing apparatus characterized by comprising:
[0054] [Configuration 2] Software for obtaining a work product incorporating the content generated by the content generation AI is installed in the information processing apparatus, The deriving means derives a specific character string that is pre-associated with the selected template based on the user's selection of the template prepared for the work product. The information processing apparatus according to Configuration 1, characterized by the above.
[0055] [Configuration 3] Specific character strings corresponding to the respective characteristics are given as metadata to the plurality of templates. The deriving means refers to the metadata of the template selected by the user and derives a specific character string associated with the template. The information processing apparatus according to Configuration 2, characterized by the above.
[0056] [Configuration 4] The deriving means refers to a table in which each of the plurality of templates is associated with a character string for a negative prompt, and derives a specific character string associated with the template selected by the user. The information processing apparatus according to Configuration 2, characterized by the above.
[0057] [Configuration 5] The deriving means uses a learned model that estimates a character string inappropriate for the impression of the input template, and derives a character string inappropriate for the impression of the template selected by the user as the specific character string. The information processing apparatus according to Configuration 1, characterized by the above.
[0058] [Configuration 6] It has control means for controlling a GUI for the user. The control means automatically updates the negative prompt set by the setting means based on the user's editing operation on the GUI for the template selected by the user. The information processing apparatus according to any one of Configurations 2 to 5, characterized by the above.
[0059] [Configuration 7] The automatic update is a process of deleting the negative prompt when the impression of a specific character string set by the setting means in the negative prompt is similar to the impression of the content added by the user himself / herself to the template selected by the user. The information processing apparatus according to Configuration 6 is characterized by this.
[0060] [Configuration 8] The derivation means derives the specific character string based on the attribute information indicating the attributes of the user. The information processing apparatus according to Configuration 1 is characterized by this.
[0061] [Configuration 9] The derivation means uses a database in which the specific character strings are registered for each attribute to derive a specific character string corresponding to the attribute specified by the attribute information. The information processing apparatus according to Configuration 8 is characterized by this.
[0062] [Configuration 10] The attribute specified by the attribute information is any one of the user's gender, age, hobby, nationality, region including the country of the nationality, and language used by the user. The information processing apparatus according to Configuration 9 is characterized by this.
[0063] [Configuration 11] The derivation means estimates the attribute using the user's behavior history or information on the user's IP address and performs the derivation. The information processing apparatus according to Configuration 10 is characterized by this.
[0064] [Configuration 12] The information processing apparatus is an information processing apparatus in which software for obtaining a work incorporating the content generated by the generation AI is installed. It has control means for controlling the GUI for the user. The control means receives the user's selection for the content in the template prepared for the work via the GUI. The derivation means derives, as the specific string, a string corresponding to an antonym of a string representing an impression of the selected content based on the user's selection of the content in the template prepared for the work product, or a string representing an impression of the content not selected by the user. The information processing apparatus according to Configuration 1, characterized in that.
[0065] [Configuration 13] When the derivation means derives, as the specific string, a string corresponding to an antonym of a string representing an impression of the content selected by the user, the learning model for estimating the impression of the input content is used to estimate the impression of the content selected by the user, and a string representing an impression opposite to the estimated impression is derived. The information processing apparatus according to Configuration 12, characterized in that.
[0066] [Configuration 14] When the derivation means derives, as the specific string, a string corresponding to an antonym of a string representing the estimated impression, the dictionary database is used to derive the string. The information processing apparatus according to Configuration 13, characterized in that.
[0067] [Configuration 15] When the derivation means derives, as the specific string, a string representing an impression of the content not selected by the user, the learning model for estimating the impression of the input content is used to estimate the impression of the content not selected by the user, and a string representing the estimated impression is derived. The information processing apparatus according to Configuration 12, characterized in that.
[0068] [Configuration 16] The generation AI is updated by performing additional learning using the content not selected by the user as a bad example. The information processing apparatus according to any one of Configurations 12 to 15, characterized in that.
[0069] [Configuration 17] It has control means for controlling the GUI for the user. The control means causes the GUI to display the negative prompt set by the setting means, and accepts a selection by the user with respect to the set negative prompt. The information processing apparatus according to any one of Configurations 1, 8, and 12, characterized in that.
[0070] [Configuration 18] The information processing apparatus according to any one of Configurations 1 to 17, further comprising generation means for inputting a positive prompt for specifying what content should be generated for the generation AI and the negative prompt set by the setting means to the generation AI to generate content.
[0071] [Configuration 19] The information processing apparatus according to any one of Configurations 1 to 17, further comprising processing means for requesting the generation of content together with a positive prompt for specifying what content should be generated for the generation AI and the negative prompt set by the setting means to an external information processing apparatus having the generation AI, and receiving the content generated based on the request.
[0072] [Configuration 20] The information processing apparatus according to any one of Configurations 1 to 19, characterized in that the content is an image.
[0073] [Configuration 21] The information processing apparatus according to any one of Configurations 1 to 19, characterized in that the content is text.
[0074] [Method 1] An information processing method for causing a generation AI to generate content, A derivation step of deriving a specific string based on information obtained about a user, A setting step of setting the specific string derived in the derivation step as a negative prompt for specifying what content should not be generated for the generation AI, An information processing method characterized by including.
[0075] [Configuration 22] A program for causing a computer to function as the information processing apparatus according to any one of Configurations 1 to 21.
Claims
1. An information processing apparatus for causing a generative AI to generate content, comprising: a derivation means for deriving a specific character string based on information obtained about a user; a setting means for setting the specific character string derived by the derivation means as a negative prompt for specifying what content should not be generated for the generative AI; An information processing apparatus characterized by comprising the above.
2. Software for obtaining a work product incorporating the content generated by the generative AI is installed in the information processing apparatus, wherein the derivation means derives a specific character string pre-associated with the selected template based on the user's selection of the template prepared for the work product. The information processing apparatus according to claim 1, characterized by the above.
3. Specific character strings corresponding to the respective characteristics are given as metadata to the plurality of templates, and the derivation means refers to the metadata of the template selected by the user and derives a specific character string associated with the template. The information processing apparatus according to claim 2, characterized by the above.
4. The derivation means refers to a table in which each of the plurality of templates is associated with a character string for a negative prompt, and derives a specific character string associated with the template selected by the user. The information processing apparatus according to claim 2, characterized by the above.
5. The derivation means uses a learned model that estimates a character string inappropriate for the impression of the input template, and derives a character string inappropriate for the impression of the template selected by the user as the specific character string. The information processing apparatus according to claim 1, characterized by the above.
6. having a control means for controlling a GUI for the user, wherein the control means automatically updates the negative prompt set by the setting means based on an editing operation by the user on the GUI for the template selected by the user. The information processing apparatus according to claim 2, characterized by the above.
7. The automatic update is a process of deleting the negative prompt when the impression of a specific character string set by the setting means in the negative prompt is similar to the impression of the content added by the user himself / herself to the template selected by the user. The information processing apparatus according to claim 6 is characterized by this.
8. The deriving means derives the specific character string based on the attribute information indicating the attributes of the user. The information processing apparatus according to claim 1 is characterized by this.
9. The deriving means uses a database in which the specific character strings are registered for each attribute to derive a specific character string corresponding to the attribute specified by the attribute information. The information processing apparatus according to claim 8 is characterized by this.
10. The attribute specified by the attribute information is any one of the user's gender, age, hobby, nationality, region including the country of the nationality, and language used by the user. The information processing apparatus according to claim 9 is characterized by this.
11. The deriving means performs the derivation by estimating the attribute using the user's behavior history or information of the user's IP address. The information processing apparatus according to claim 10 is characterized by this.
12. The information processing apparatus is an information processing apparatus in which software for obtaining a result incorporating the content generated by the generation AI is installed. It has control means for controlling the GUI for the user. The control means receives, via the GUI, the user's selection of the content in the template prepared for the result. Based on the user's selection of the content in the template prepared for the result, the deriving means derives, as the specific character string, a character string corresponding to the antonym of the character string representing the impression of the selected content, or a character string representing the impression of the content not selected by the user. The information processing apparatus according to claim 1 is characterized by this.
13. When the deriving means derives, as the specific character string, a character string corresponding to the antonym of the character string representing the impression of the content selected by the user, it estimates the impression of the content selected by the user using a learned model for estimating the impression of the input content, and derives a character string representing the impression opposite to the estimated impression. The information processing apparatus according to claim 12 is characterized by this.
14. The information processing apparatus according to claim 13, wherein the derivation means derives a character string corresponding to the antonym of the character string representing the estimated impression using a dictionary database.
15. When the derivation means derives a character string representing the impression of the content not selected by the user as the specific character string, the learning model that estimates the impression of the input content is used to estimate the impression of the content not selected by the user, and a character string representing the estimated impression is derived. The information processing apparatus according to claim 12, characterized by the above.
16. The information processing apparatus according to claim 12, wherein the generation AI is updated by performing additional learning using the content not selected by the user as a bad example.
17. It has control means for controlling the GUI for the user, The control means causes the GUI to display the negative prompt set by the setting means and accepts the user's selection for the set negative prompt. The information processing apparatus according to claim 1, characterized by the above.
18. The information processing apparatus according to claim 1, further comprising generation means for inputting a positive prompt for specifying what kind of content should be generated for the generation AI and the negative prompt set by the setting means to the generation AI to generate content.
19. The information processing apparatus according to claim 1, further comprising processing means for requesting the generation of content together with a positive prompt for specifying what kind of content should be generated for the generation AI and the negative prompt set by the setting means to an external information processing apparatus having the generation AI, and receiving the content generated based on the request.
20. The information processing apparatus according to claim 1, wherein the content is an image.
21. The information processing apparatus according to claim 1, wherein the content is text.
22. An information processing method for causing a generation AI to generate content, A derivation step of deriving a specific character string based on information obtained about the user, A setting step of setting the specific character string derived in the derivation step as a negative prompt for specifying what kind of content should not be generated for the generation AI. An information processing method characterized by including
23. A program for causing a computer to execute the information processing method according to Claim 22.
Citation Information
Patent Citations
Picture book creation system, picture book creation program, and picture book creation method
JP7462991B1
Information processing device and program
JP2017037557A