Information processing method, information processing device and storage medium

By setting negative prompts in the generated AI, the problem of users' difficulty in finding appropriate prompts is solved, and efficient generation of images that conform to the template style is achieved.

CN120298541APending Publication Date: 2025-07-11CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510028930.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-11
Filing Date
2025-01-08
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When users use Generate AI to generate images, it is difficult to find appropriate prompts, resulting in the generated images not meeting the expected template style, and a waste of time and effort.

Method used

Through the information processing device, a specific character string is derived based on user-related information as a negative prompt, and a content that the generation AI should not generate is set, and an image that conforms to the template style is generated based on the positive prompt.

Benefits of technology

Reduces the time and effort of users to find appropriate prompts, improves the efficiency and quality of generated images, and ensures that the generated content matches the template style.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298541A_ABST
    Figure CN120298541A_ABST
Patent Text Reader

Abstract

The invention provides an information processing method, an information processing apparatus, and a storage medium. The information processing apparatus for causing a generation AI to generate content includes: a deriving unit configured to derive a specific character string based on information obtained in relation to a user; and a setting unit configured to set the specific character string exported by the exporting unit as a negative cue specifying which content the generation AI should not generate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to prompt settings for image generation AI. Background Art

[0002] As one of the functions of software for creating posters, flyers, etc., there is a function that allows a user to select a desired template from various pre-prepared templates and insert arbitrary characters, images taken by the user himself / herself, etc. into the template. In this regard, Japanese Patent Application Laid-Open No. 2017-037557 discloses a technique that enables retrieval of an image suitable for the template from a separately prepared group of images by extracting words from the attribute information of the objects included in the template and creating a search keyword.

[0003] For example, in software for creating posters, etc., there has emerged software having a function of generating an image by using generative AI (artificial intelligence). This image generation function is a function that, when a user inputs a word or a sentence as a prompt to the generative AI, the AI automatically generates an image based on the input prompt (word or sentence). For example, here it is assumed that the template of the poster to be used is in an illustration style. When inserting an image into the illustration-style template, the user inputs an arbitrary prompt to the generative AI in order to obtain an image of the same illustration style. However, when the prompt input by the user is not suitable, the generative AI generates, for example, a realistic image, etc., so there may be a situation where the image does not match the style and atmosphere of the template to be used. In such a case, the user needs to repeat image generation by the generative AI by trying to input other prompts, etc., so it takes time and effort for the user. As described above, it is difficult and time-consuming work for the user to find and input an appropriate prompt for obtaining the desired content to the generative AI. Summary of the Invention

[0004] An information processing apparatus for causing a generative AI to generate content according to the present disclosure includes: one or more memories that store instructions; and one or more processors that execute the instructions to perform the following steps: deriving a specific string based on information obtained related to a user; and setting the derived specific string as a negative prompt specifying what kind of content the generative AI should not generate.

[0005] Other features of the present disclosure will become apparent from the following description of exemplary embodiments with reference to the drawings. Brief Description of the Drawings

[0006] Figure 1 is a diagram showing a configuration example of a client-server system according to the present embodiment;

[0007] Figure 2 is a diagram showing an example of the hardware configuration of the information processing apparatus;

[0008] Figure 3 is a functional block diagram showing an example of the software configuration of a front - end application and a back - end application;

[0009] Figure 4 is a diagram showing an example of a negative prompt;

[0010] Figure 5 is a flowchart showing the process of operations in a client;

[0011] Figure 6A and Figure 6B each is a diagram showing an example of a poster editing screen;

[0012] Figure 7A and Figure 7B each is a diagram showing an example of a poster editing screen;

[0013] Figure 8A and Figure 8B each is a diagram showing an example of a poster editing screen;

[0014] Figure 9 is a flowchart showing the process of operations in a server;

[0015] Figure 10A and Figure 10B each is a diagram showing an example of a poster editing screen;

[0016] Figure 11A is a diagram explaining the method of exporting a string of negative prompts, and Figure 11B is a diagram showing an example of a UI screen for setting a negative prompt from an exported string group;

[0017] Figures 12A to 12C each is a diagram explaining the method of obtaining an impression of an image by using a learning model; and

[0018] Figure 13A and Figure 13B each is a diagram explaining the automatic update of negative prompts. Detailed Description of the Invention

[0019] Hereinafter, with reference to the accompanying drawings, the present disclosure will be described in detail according to preferred embodiments. The configurations shown in the following embodiments are merely exemplary, and the present disclosure is not limited to the configurations shown schematically.

[0020] [First Embodiment]

[0021] <System Configuration>

[0022] Figure 1It is a diagram showing an example of the configuration of a client-server system according to the present embodiment. The system includes: a server 101 that provides a content generation service and a client 102 that uses the service. The server 101 and the client 102 communicate with each other via the respective networks 103 through the Internet 104. The server 101 has a backend application 105. The client 102 has an operating system 106 and a frontend application 107.

[0023] <Hardware Configuration>

[0024] Figure 2 It is a diagram showing an example of the common hardware configuration of the server 101 and the client 102, and both the server 101 and the client 102 are information processing devices. The user interface 201 is an operation display unit configured to input and output information and signals, including a display, a keyboard, a mouse, buttons, a touch panel, etc. The network interface 202 is connected to a network such as a LAN and communicates with other computers and network devices. The communication method can be a wired method or a wireless method. The CPU 203 is a central processing unit configured to execute programs read from the ROM 204, the RAM 205, the secondary storage device 206, etc. The ROM 204 is a non-volatile memory that stores integrated programs and data. The RAM 205 is a volatile memory that provides a temporary storage area. The secondary storage device 206 is typically a large-capacity storage device such as an HDD and a flash memory. It is also possible to connect to and operate a computer not equipped with such hardware from another computer via remote desktop, remote shell, etc. Each unit is connected via an input / output interface 207.

[0025] <Software Configuration>

[0026] Figure 3 It is a functional block diagram showing an example of the software configuration of the frontend application 107 and the backend application 105, respectively. The frontend application 107 has a GUI control unit 301, a negative prompt export unit 302, a prompt setting unit 303, and a request processing unit 304. The backend application 105 has a response processing unit 311 and a content generation unit 312. The following configuration is also acceptable, in which one of the frontend application 107 and the backend application 105 includes each of the above units 301 to 304 and 311 and 312.

[0027] The GUI control unit 301 controls a GUI (Graphical User Interface) for presenting information to the user or for the user to input instructions. Specifically, it controls the display of the UI screen, the reception of user input, etc. As specific examples of user input, there are the selection of templates, instructions for content insertion methods, the reception of strings as prompts (including instruction information of words and sentences) input to the generation AI, etc. Prompts include positive prompts and negative prompts. A positive prompt is a prompt that designates a desirable object (an element expected to be generated) of the content that the generation AI should generate. For example, when "reindeer" is input as a positive prompt to an AI that generates an image as content (image generation AI), the image generation AI generates an image of a reindeer. In contrast, a negative prompt is a prompt that designates an undesirable object (an element expected to be excluded) of the content that the generation AI should not generate. Figure 4 is a table summarizing examples of general negative prompts input to the image generation AI. For example, when "low quality" is input as a negative prompt to the image generation AI, the image generation AI no longer generates low-quality images.

[0028] The negative prompt derivation unit 302 derives a string to be used as the above-mentioned negative prompt from the prompts input to the generation AI based on information obtained related to the user. The derivation method will be described later.

[0029] The prompt setting unit 303 sets the string input by the user via the GUI and the string derived by the negative prompt derivation unit 302 as positive prompts and negative prompts, respectively.

[0030] The request processing unit 304 performs processing to request the server 101 to generate content. At the time of the request, the positive prompt and negative prompt set by the prompt setting unit 303 are also sent together. In addition, the request processing unit 304 also performs processing to receive the content generated by the server 101 in response to the request.

[0031] The response processing unit 311 performs response processing to receive a request from the client 102, and sends the content generated in response to the received request to the client 102 that made the request, etc.

[0032] The content generation unit 312 is a generation AI that takes positive prompts and negative prompts as input to generate content such as images and texts. As an image generation AI that generates an image as content, for example, "StableDiffusion", "Midjourney", etc. are known. The generation AI is a learning model (content generation model) obtained by performing machine learning on various data by methods such as deep learning so as to obtain the target content.

[0033] Based on the above content, taking the poster creation software as the front-end application 107 as an example, the operation processes of each in the client 102 and the server 101 are described. The poster creation software is an example installed in the client 102, and the software is not limited thereto. For example, the present disclosure can be applied to software for obtaining various products, such as photo album creation software and postcard creation software.

[0034] <Operations on the client side>

[0035] Figure 5 It is a flowchart showing the operation process in the client 102 according to the present embodiment. This flowchart starts by launching the poster creation software based on a user instruction. Hereinafter, the symbol "S" represents a step.

[0036] In S501, the GUI control unit 301 displays an editing UI screen (hereinafter described as "poster editing screen") on the display of the user interface 201 according to the poster creation software. Figure 6A An example of the poster editing screen is shown. Figure 6A The poster creator screen 600 in includes a template list pane 610, a template editing pane 620, and a content editing pane 630. The GUI control unit 301 displays a list of various templates according to the purposes and uses in the template list pane 610. In Figure 6A In the example in, three pre-prepared templates (611, 612, 613) are displayed in the template list pane 610.

[0037] In S502, the GUI control unit 301 receives a user selection of a specific template among the templates displayed in the list in the template list pane 610. Figure 6B The change of the screen in the case where the template 612 is selected by the user and a highlight box 640 indicating the selection state is added is shown. In addition, the selected template 612 by the user is displayed as an editing object template 621 in the template editing pane 620. Here, in the editing object template 621, an image setting area 622 for the user to set an arbitrary image and a text setting area 623 for the user to set an arbitrary text are arranged. Hereinafter, the image setting area and the text setting area are collectively referred to as "content setting areas".

[0038] In S503, the GUI control unit 301 determines whether there is a press on one of the content setting areas in the editing object template displayed in the template editing pane 620. In the case where a press on the content setting area is detected, the GUI control unit 301 proceeds to S504 subsequently, and in the case where no press is detected, the GUI control unit 301 determines whether there is a press after waiting for a predetermined time.

[0039] In S504, the GUI control unit 301 displays a pop-up screen for allowing a user to select a method for adding content according to a pressed content setting area, and receives a user selection as to whether to manually add object content or to automatically generate it. Figure 7A Shows a state in Figure 6B where, when the user presses the image setting area 622 by operating a pointer and a pop-up screen for selecting whether to automatically generate the added image or to manually add it is displayed. Here, as an alternative to the image addition method, buttons for "Automatically Generate" and "Select from Folder" are displayed. Here, the "Automatically Generate" button is a button for generating an AI-generated image, and the "Select from Folder" button is a button for the user to specify an image from an arbitrary folder and upload the image. When the text setting area 623 is pressed in S503, as an alternative to the text addition method, buttons for "Automatically Generate" and "Manually Enter" are displayed. The "Automatically Generate" button in this case is a button for generating AI-generated text, and the "Manually Enter" button is a button for the user to manually enter an arbitrary string. When it is detected that the button for manual addition is pressed, the GUI control unit 301 proceeds to S505 that follows, and when it is detected that the button for automatic generation is pressed, the GUI control unit 301 proceeds to S507 that follows.

[0040] In S505, the GUI control unit 301 receives a user's specification of content. For example, when it is detected in S504 that the "Select from Folder" button for manually adding an image is pressed, the GUI control unit 301 receives a specification of desired image data from an arbitrary folder. The user can specify a desired image from among images that are pre-stored in a folder obtained by the user's own photography or via the Internet 104 or the like through an operation such as drag-and-drop. Further, when it is detected in S504 that the "Manually Enter" button for manually adding text is pressed, the GUI control unit 301 receives a specification of a string via an input bar (not schematically shown) that is displayed for the user to directly input a desired string.

[0041] In S506, the GUI control unit 301 inserts the content specified by the user received in S505 into the content setting area in the selected template pressed in S503. Figure 7B is a diagram showing an example of a state in which a user-specified image 710 is inserted into the image setting area 622. When inserting content, an operation bar 720 for allowing the user to perform an edit operation on the inserted content is displayed in the content editing pane 630. Here, in Figure 7BIn [the figure], the inserted image 710 can be edited. In this example, a bar 721 for the user to indicate whether to automatically correct the edited object image, a control bar 722 for adjusting the brightness of the image, and a control bar 723 for adjusting the contrast of the image are provided. When the content of the edited object is text, an operation bar (not schematically shown) for adjusting the font, size, color, etc. of the characters is displayed. The user can modify the content, etc. by appropriately performing necessary operations in the operation bar according to the content of the edited object.

[0042] The processes of S507 to S513 are processes in which the generation AI automatically generates content for the content setting area pressed by the user and inserts the content. Figure 8A and Figure 8B each is an example of a UI screen when the user selects the template 611 of the Christmas poster in the state of Figure 6A and then presses the image setting area 622 of the edited object template 621 and selects image automatic generation. Hereinafter, with reference to the UI screen example as appropriate, the process of the template until the content is automatically generated by the generation AI is described.

[0043] First, in S507, the GUI control unit 301 receives an input of a string of words or the like representing an element of the content that the user desires to be generated by the generation AI. As the language used in the prompt, English is generally often used, so in this embodiment, the description is made in English. However, it goes without saying that the language is not limited to English because the language used in the prompt depends on the generation AI. Here, it is assumed that the words "reindeer" and "Christmas" are input via an input bar (not schematically shown).

[0044] In S508, the prompt setting unit 303 sets the string of words or the like input by the user in S507 as a positive prompt. Here, since the user has input two words, "reindeer" and "Christmas", these words are set as positive prompts.

[0045] In S509, the negative prompt export unit 302 exports a string such as words representing elements of content that the user does not expect to be generated by the generative AI, based on the template selected by the user. As a template-based export method, consider a method in which, for each template displayed in the template list, a string for negative prompting according to its characteristics is pre-attached as metadata, and the metadata of the template related to the user's selection is referred to. Alternatively, a table associating each template with a string for negative prompting may be prepared in advance and referred to. In addition, a string may be exported by using a trained model that estimates a string unsuitable for the impression of the template selected by the user. The trained model for this estimation can be obtained by learning a large amount of training data in which templates and words etc. unsuitable for the impression of the template are paired. In the case of estimation, after estimating the impression of each content (such as an image and text included in the object template), the impression of the entire object board or the comprehensive impression of the entire template may be estimated. In Figure 8A the example in

[0046] Figure 8A , a template with an illustration style of the Christmas theme is selected. In this case, even if an image of a reindeer living in the natural world (such as a photographed reindeer) is generated, it is overall unsuitable for the illustration-style template. Therefore, for the illustration-style template, words such as "realistic" are pre-associated by the above method. Thus, when the user simply selects the desired template, a negative prompt suitable for the template is automatically obtained and set, so the user can save the time and effort of searching for and setting negative prompts by themselves. In the case of selecting an illustration-style template, the user can also set a string such as "illustration style" as a positive prompt and have the generative AI generate an illustration-style image. However, with the method of this embodiment, the user does not need to search for and set such positive prompts every time, so more time and effort can be saved. Figure 11B As shown later

[0047] Figure 11B , it is also possible to enable the user to set, via the GUI, the word set as the negative prompt from the group of words exported in S509.

[0048] In S512, the request processing unit 304 receives the content generated based on the content generation request from the server 101.

[0049] In S513, the GUI control unit 301 inserts the content received in S512 into the content setting area pressed in S503 in the selected template. In Figure 8A the example in, the image 800 of a reindeer in an illustration style is inserted into the image setting area 622, and in the content editing pane 630, an operation bar 810 for the user to perform editing operations on the image 800 is displayed. In the operation bar 810, two sub-bars 820 and 830 indicating positive prompts and negative prompts for generating the inserted image 800 are set. In addition, a sub-bar 840 for only displaying the generated image is set. Figure 8B FIG. is a specific example showing the case where "reality" is not set as a negative prompt and the generated AI generates a real image 801 of a reindeer (such as a photographed picture in nature). It is also possible to enable the user to fine-tune the hue and the like of the image displayed in the sub-bar 840 as an object by using the aforementioned control bar.

[0050] In S514, for the template selected in S502, it is determined whether all the content to be set has been set. In the case where there is still content not set, the process returns to S503 and continues the process. On the other hand, in the case where all the content has been set, the process is terminated.

[0051] The above is the description of the operations on the client side. In the above Figure 5 process, before sending the content generation request, it is also possible to set a step for the user to check on the UI screen whether to apply the automatically set negative prompt as it is in S510. In addition, it is also possible to make the generated AI regenerate the image by the user's own operations such as adding and changing the strings in the sub-bars 820 and 830.

[0052] The above method can be applied to various situations. For example, in the case where there are some unwritten rules for certain traditional foods (for example, ingredient X shall not be used), it is only necessary to associate a string such as "X as an ingredient that should not be used" with the template of the traditional food. Thus, even for a user who does not know the rules related to the traditional food, when the user selects the template of the traditional food, the string representing ingredient X will be automatically set as a negative prompt. Therefore, it is possible to prevent the generated AI from erroneously generating content including ingredient X.

[0053] <Operations on the server side>

[0054] Figure 9It is a flowchart showing the operation process of the backend application 105 of the server 101 when generating content in response to a content generation request (S511) from the client 102. In the following description, the symbol "S" represents a step.

[0055] In S901, the response processing unit 311 receives a content generation request from the client 102. In S902, the content generation unit 312 obtains positive prompts and negative prompts from the content generation request received in S901. In S903, the content generation unit 312 generates content using the positive prompts and negative prompts obtained in S902 as input. In S904, the response processing unit 311 sends the data of the content generated in S903 to the requesting client 102.

[0056] The above is the description of the operations on the server side. In this embodiment, it is configured such that a generation AI is included in the backend application 105, but the configuration is not limited thereto. For example, the following configuration can also be accepted: the generation AI is located outside the backend application 105, and the backend application 105 responds to the content generation request by calling the external generation AI.

[0057] <Variations in the method of exporting strings>

[0058] In the above example, the string used as the negative prompt is exported by associating specific strings with each template in advance, but the method of exporting the string for the negative prompt is not limited thereto. Hereinafter, variations in the method of exporting the string for the negative prompt will be described.

[0059] <<Export based on user attributes>>

[0060] Generally, when software implements functions such as poster creation through cloud services, users often use the software by registering an account in advance and logging in. When registering an account, users also register attribute information such as name, gender, nationality, region, language, and hobbies. When services of software are developed in many countries around the world, the software is used by users from different countries, but the cultures and customs of different countries are different. For example, the same gesture may have different interpretations in different countries. For example, in the case of the peace gesture, which is a kind of body language, in some countries, it gives a good impression, while in some countries, it gives a bad impression. Therefore, in the case where the nationality indicated by the attribute information of the logged-in user indicates that the peace gesture gives a bad impression in that country, the export method is designed such that "peace gesture" is exported as the string for the negative prompt. As a specific export method, it is only necessary to register in the database in advance the strings indicating the gestures regarded as taboos in each country, etc., and then, when the user logs in, ask the user about their nationality and obtain the string registered in association with the user's country.Figure 10A and Figure 10B are each an example of a UI screen when the user selects a template 611 on the UI screen in the foregoing Figure 6A and presses the image setting area 622 of the editing target template 621 and selects automatic image generation. Figure 10A is an example when the nationality of the logged-in user is a country where a peace gesture does not give a bad impression, while Figure 10B is an example when the nationality of the logged-in user is a country where a peace gesture gives a bad impression. In each case, the string "A woman is smiling" is set as a positive prompt, but in Figure 10B , "peace gesture" is further set as a negative prompt. As a result, in the example in Figure 10A , the AI generates an image 1000 of a smiling woman with a peace gesture, but in the example in Figure 10B , the AI generates an image 1001 of a smiling woman without a peace gesture. As described above, by setting a negative prompt according to the nationality of the user, it is possible to prevent the generation of content inappropriate for the user. Here, for example, the nationality is used as attribute information about the user, but other user attributes such as a region including multiple countries or the language used can also be used. In this case, for each attribute such as the region and the language, a string representing a gesture considered taboo, etc. is simply registered in advance in the database, and then the string is derived based on the attributes of the logged-in user and the negative prompt is set. In addition, instead of the attributes at the time of account registration, the above attributes of the user can also be estimated based on information after account registration (such as the user's behavior history and IP address), and the estimated attribute information can be used to derive the string of the negative prompt. Even when the nationality of the user indicates a country where a peace gesture gives a bad impression, under the condition that the user inputs a string with a deliberately insulting expression form as a positive prompt, the positive prompt can be prioritized so that "peace gesture" is not set as a negative prompt.

[0061] <<Derivation Based on User Selection of Content>>

[0062] Based on the above, a method for deriving a string of a negative prompt based on the user's selection of content within a template is described. Specifically, a learning model for impression estimation (impression estimation model) is used to estimate the impression of the image selected by the user from the images generated by the generative AI, and a string corresponding to the antonym of the string representing the estimated impression is obtained from the dictionary database. Figure 11AThis is a diagram illustrating the method of setting negative prompts based on an image selected by a user from multiple images generated by an image generation model (i.e., generative AI). Here, as a positive prompt, "Halloween pumpkin" is set. Then, the image generation model 1110 takes this as input and generates three images 1120, 1121, and 1122 of pumpkins transformed into a Halloween style. These three images include: an image 1120 with an impression of "nice" and images 1121 and 1122 with an impression of "scary". Here, it is assumed that the user selects the image 1120 with an impression of "nice" as the image the user desires to insert into the template. In this case, if the user inputs the selected image 1120 into the impression estimation model, the impression of "nice" is estimated. Next, the user inputs the word "nice" representing the estimated impression into the dictionary database. As a result, words such as "terrifying", "ghoulish", "weird", and "frightening" representing impressions opposite or deviating from "nice" are obtained. Figure 12A This is a diagram showing the method of deriving the impression "terrifying" opposite to "nice" by further inputting the impression "nice" obtained by inputting the image 1120 selected by the user into the impression estimation model 1200 into the dictionary database. Figure 11B This is an example of a UI screen where the user selects and sets negative prompts from the derived word group. The GUI control unit 301 of the front-end application 107 displays a candidate dialog 1130 including the derived word group. The user selects a word representing the impression the user does not want to generate from the word group displayed in the candidate dialog. In this example, "ghoulish" is selected and set as the negative prompt. Then, in the case where the user sets "Halloween bat" in the positive prompt and inputs it into the image generation model 1110, images 1140 and 1141 of bats with an impression of "nice" are generated. It is also possible to derive a string of negative prompts from the image selected by the user by using the impression estimation model and the dictionary database as described above. Here, the user selects one word from the derived word group, but it is also possible to set one or more of the derived words as they are in the negative prompt.

[0063] In addition, the word "nice" representing the impression estimated from the image selected by the user can be added to the positive prompt. In addition, as Figure 12B shown, it is also possible to more directly derive a string of negative prompts from the image selected by the user by using a learning model (antonym estimation model) 1210 that can estimate the antonyms of words representing the impression of an image. In this case, the antonym estimation model can be obtained by learning a large amount of training data as follows, in which images and words representing impressions not suitable for the images are paired. The images used as training data for learning can be images generated by generative AI or images taken by a person using a camera.

[0064] In the above example, the strings of the negative prompts are exported based on the image 1120 selected by the user, but the strings can also be exported based on the images 1121 and 1122 not selected by the user. Here, the images 1121 and 1122 not selected by the user are both images with the impression of "scary". Therefore, by inputting these images into the impression estimation model 1200, "scary" representing the common impression of these two images can be estimated. The number of images input into the impression estimation model can be one, three, or more than three.

[0065] In addition, an example of exporting a string based on an image not selected by the user among the images generated by the generative AI is illustrated, but the example is not limited thereto. For example, a string can also be exported based on an image not selected by the user among the images prepared in advance by the poster creation software for inserting into the template.

[0066] In addition, an image not selected by the user can be used as a bad example, and the generative AI can be updated by using a method such as LoRA (Low-Rank Adaptation) for additional learning in association with the user. Thus, in the case where each user uses the generative AI next time and thereafter, it is difficult to generate an image that does not suit the preferences of the logged-in user by using the generative AI updated for each user.

[0067] For example, as the impression estimation model or the antonym estimation model, a learning model (such as a content generation model) learned by the method of deep learning is assumed, but the model is not limited thereto.

[0068] <Variant Example>

[0069] The negative prompt automatically set in the process of the user editing the template can also be updated automatically. Figure 13A and Figure 13B is a diagram illustrating the automatic update of the negative prompt. Here, it is assumed that "real" and "beautiful" are associated with the template 1300 in advance as negative prompts. Then, the user issues an instruction to automatically generate and add an image by using "Halloween pumpkin" in the positive prompt. As a result, an image 1301 of a pumpkin in an illustration style with the impression of "scary" is generated by the generative AI, and this image is arranged at the upper right position of the template 1300. In Figure 13A V1 and V2 are "real" and "beautiful" respectively converted into vector representation forms by using a method such as Word2Vec. The method of converting a word into a vector representation form is not limited to the Word2Vec method, and other methods can also be used. Next, it is assumed that the user inserts the image 1302 of the pumpkin into the lower bottom position of the template 1300 by drag-and-drop or the like. Then, the image 1302 of the pumpkin is input into the aforementioned impression estimation model 1200, and the impression "lovely (beautiful / young)" is estimated. In Figure 13AAmong them, V is the estimated impression "cute" converted into a vector representation form. In this case, under the condition of obtaining the cosine similarity S1 between V and V1 and the cosine similarity S2 between V and V2 and comparing the two, the value of S2 will be larger. The reason is that the meaning of "beautiful" is closer to the meaning of "cute" than the meaning of "realistic". When the similarity between this impression and the impression of the image added by the user himself is relatively high, this impression is not suitable as a negative prompt. Therefore, as Figure 13B shown, "beautiful" is deleted from the negative prompt. Thus, the negative prompt of the template 1300 is only updated to "realistic". As described above, the negative prompt can also be automatically updated during the process of the user editing the template. Thus, for example, when the user specifies "Halloween bat" in the positive prompt and issues an instruction to automatically generate an image, an image of a bat with the impression of "beautiful" will also be generated. As described above, by automatically updating the negative prompt according to the user's editing operation, the generation AI can be made to generate content suitable for the impression changed by editing.

[0070] In addition, in the above embodiment, the user can check the automatically set negative prompt on the UI screen, but the automatically set negative prompt can also be applied in a form that the user cannot see without displaying it on the UI screen.

[0071] In addition, it can also be made possible for the user to change the expression form of the negative prompt automatically derived and set through the above embodiment. For example, change "cute" to "beautiful".

[0072] [Other Embodiments]

[0073] Embodiments of the present invention can also be implemented by a computer of a system or device that reads and executes computer-executable instructions (e.g., one or more programs) recorded on a storage medium (which may also be more fully referred to as a "non-transitory computer-readable storage medium") to perform the functions of one or more of the above-described embodiments, and / or includes one or more circuits (e.g., an application specific integrated circuit (ASIC)) for performing the functions of one or more of the above-described embodiments. Further, embodiments of the present invention can be implemented by a method of, for example, reading and executing the computer-executable instructions from the storage medium by the computer of the system or device to perform the functions of one or more of the above-described embodiments, and / or controlling the one or more circuits to perform the functions of one or more of the above-described embodiments. The computer may include one or more processors (e.g., a central processing unit (CPU), a microprocessing unit (MPU)), and may include a network of separate computers or separate processors to read and execute the computer-executable instructions. The computer-executable instructions may be provided to the computer, for example, from a network or the storage medium. The storage medium may include, for example, a hard disk, a random access memory (RAM), a read only memory (ROM), the memory of a distributed computing system, an optical disk (such as a compact disk (CD), a digital versatile disk (DVD), or a Blu-ray disk (BD) TM ), a flash device, and a memory card, among others.

[0074] According to the present disclosure, an appropriate prompt for obtaining desired content can be easily set.

[0075] Although the present invention has been described with reference to exemplary embodiments, it should be understood that the present invention is not limited to the disclosed exemplary embodiments. The scope of the appended claims should be given the broadest interpretation so as to cover all such variations and equivalent structures and functions.

Claims

1. An information processing apparatus for causing a generative AI to generate content, the information processing apparatus comprising: one or more memories that store instructions; and one or more processors that execute the instructions to perform the following steps: an export step of exporting a specific string based on information obtained related to a user; and a setting step of setting the exported specific string as a negative prompt specifying what kind of content the generative AI should not generate.

2. The information processing apparatus according to claim 1, wherein in the information processing apparatus, software for obtaining a product integrating content generated by the generative AI is installed, and in the export step, a specific string associated in advance with the selected template is exported based on the user's selection of a template prepared for the product.

3. The information processing apparatus according to claim 2, wherein specific strings according to respective features are attached as metadata to a plurality of templates, and in the export step, the specific string associated with the template is exported with reference to the metadata of the template selected by the user.

4. The information processing apparatus according to claim 2, wherein in the export step, the specific string associated with the template selected by the user is exported with reference to a table in which each of the plurality of templates is associated with a string for a negative prompt area.

5. The information processing apparatus according to claim 1, wherein in the export step, a string that is estimated not to suit the impression of the input template is exported as the specific string by using a learning model for strings.

6. The information processing apparatus according to claim 2, wherein the one or more processors further execute the instructions to perform the following steps: a control step of controlling a GUI of the user; and in the control step, the set negative prompt is automatically updated based on an editing operation of the template selected by the user on the GUI.

7. The information processing apparatus according to claim 6, wherein the automatic update is a process of deleting the negative prompt when the impression of the specific string set as the negative prompt is similar to the impression of the content that the user has added to the template selected by the user.

8. The information processing apparatus according to claim 1, wherein in the export step, the specific string is exported based on attribute information indicating an attribute of the user.

9. The information processing apparatus according to claim 8, wherein in the export step, the specific string is exported by using a database in which specific strings according to attributes identified by the attribute information are registered for each attribute.

10. The information processing apparatus according to claim 9, wherein the attribute identified by the attribute information is one of gender, age, hobby, nationality, a region of a country including the nationality of the user, and a language used by the user.

11. The information processing apparatus according to claim 10, wherein In the export step, the attribute is estimated by using information on the user's behavior history or the user's IP address, and a specific string according to the estimated attribute is exported.

12. The information processing apparatus according to claim 1, wherein the information processing apparatus is an information processing apparatus installed with software for obtaining a product integrating the content generated by the generation AI, the one or more processors further execute the instructions to perform the following steps: a control step of controlling the user's GUI, in the control step, receiving, via the GUI, a user's selection of content within a template prepared for the product, and in the export step, based on the user's selection of content within a template prepared for the product, exporting, as the specific string, a string corresponding to the antonym of a string representing the impression of the selected content or a string representing the impression of the content not selected by the user.

13. The information processing apparatus according to claim 12, wherein in the export step, when exporting, as the specific string, a string corresponding to the antonym of a string representing the impression of the content selected by the user, estimating the impression of the content selected by the user by using a learning model for estimating the impression of input content, and exporting a string representing an impression opposite to the estimated impression.

14. The information processing apparatus according to claim 13, wherein in the export step, exporting a string corresponding to the antonym of a string representing the estimated impression by using a dictionary database.

15. The information processing apparatus according to claim 12, wherein in the export step, when exporting, as the specific string, a string representing the impression of the content not selected by the user, estimating the impression of the content not selected by the user by using a learning model for estimating the impression of input content, and exporting a string representing the estimated impression.

16. The information processing apparatus according to claim 12, wherein the generation AI is updated by additionally learning with the content not selected by the user as a bad example.

17. The information processing apparatus according to claim 1, wherein the one or more processors further execute the instructions to perform the following steps: a control step of controlling the user's GUI, and in the control step, displaying the set negative prompt on the GUI and receiving the user's selection of the set negative prompt.

18. The information processing apparatus according to claim 1, wherein the one or more processors further execute the instructions to perform the following steps: generating content by inputting a positive prompt specifying what kind of content the generation AI generates and the set negative prompt into the generation AI.

19. The information processing apparatus according to claim 1, wherein the one or more processors further execute the instructions to perform the following steps: requesting an external information processing apparatus having the generation AI to generate content, a positive prompt specifying what kind of content the generation AI generates, and the set negative prompt; and receiving the content generated based on the request.

20. The information processing apparatus according to claim 1, wherein the content is an image.

21. The information processing apparatus according to claim 1, wherein the content is text.

22. An information processing method for causing a generative AI to generate content, the information processing method comprising: deriving a specific string based on information obtained related to a user; and setting the derived specific string as a negative prompt specifying what kind of content the generative AI should not generate.

23. A non-transitory computer-readable storage medium storing a program for causing a computer to perform an information processing method for causing a generative AI to generate content, the information processing method comprising: deriving a specific string based on information obtained related to a user; and setting the derived specific string as a negative prompt specifying what kind of content the generative AI should not generate.

Citation Information

Patent Citations

  • Information processing device and program

    JP2017037557A