Image generation method using generative model, and computing device for performing same

By employing a generative model to convert conversations into prompts and then generate images, the method effectively addresses the challenge of accurately representing conversation contexts in images, enhancing efficiency and user experience.

WO2025127406A1PCT designated stage expired Publication Date: 2025-06-19SAMSUNG ELECTRONICS CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/017159
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-11-04
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing methods for converting conversations into images often fail to accurately match the conversation context with the generated image, requiring users to manually enter image descriptions, which is cumbersome.

Method used

A method using a generative model that first converts a conversation into a prompt through a prompt generation model, and then uses an image generation model to produce an image that reflects the conversation context.

Benefits of technology

This approach enables the generation of images that accurately represent the conversation context, improving the efficiency and effectiveness of image creation from conversations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024017159_19062025_PF_FP_ABST
    Figure KR2024017159_19062025_PF_FP_ABST
Patent Text Reader

Abstract

This method for generating an image by using a generative model comprises the steps of: receiving a dialogue; obtaining a prompt corresponding to the dialogue by inputting the dialogue into a prompt generative model; and obtaining an image by inputting the prompt into an image generative model, wherein the prompt includes a description reflecting the context of the dialogue, and the prompt generative model can be a language model trained through a plurality of dialogue-prompt pairs.
Need to check novelty before this filing date? Find Prior Art

Description

Image generation method using a generative model and a computing device for performing the same

[0001] The present disclosure relates to a method for generating an image using a generative model, and more particularly, to a method for generating an image corresponding to a conversation using a generative model.

[0002] Generative AI technology learns the patterns and structures of massive training data and, based on that data, generates new data similar to the input data. Generative AI technology can be used to obtain images corresponding to text.

[0003] Meanwhile, with the recent surge in messenger-based online communication, some devices and programs now support features that allow users to convert their conversations into images. However, there are instances where the conversation and the converted images don't properly match. While inputting a separate image description to create an appropriate image for the conversation can be helpful, it can also be a cumbersome process for users.

[0004] A method for generating an image using a generative model according to one embodiment of the present disclosure may include a step of receiving a dialogue. The method may include a step of inputting the dialogue into a prompt generation model to obtain a prompt corresponding to the dialogue. The method may include a step of inputting the prompt into an image generation model to obtain an image. The prompt may include a description reflecting the context of the dialogue. The prompt generation model may be a language model trained through a plurality of dialogue-prompt pairs.

[0005] A computing device according to an embodiment of the present disclosure may include an input / output interface, a memory, and at least one processor. The input / output interface may receive a user input requesting image processing. The input / output interface may output a processed image according to the user input. The memory may store commands for processing the image. At least one processor may execute the commands. At least one processor may receive a dialogue. At least one processor may input the dialogue into a prompt generation model to obtain a prompt corresponding to the dialogue. At least one processor may input the prompt into an image generation model to obtain an image. The prompt may include a description reflecting the context of the dialogue. The prompt generation model may be a language model trained through a plurality of dialogue-prompt pairs.

[0006] According to one embodiment of the present disclosure, a non-transitory computer-readable recording medium may have stored thereon a program for executing at least one of the embodiments of the disclosed method on a computer.

[0007] According to one embodiment of the present disclosure, a computer program may be stored on a medium for performing at least one of the embodiments of the disclosed method on a computer.

[0008] FIG. 1 is a conceptual diagram illustrating a process of generating an image using a prompt generation model according to an embodiment of the present disclosure.

[0009] FIG. 2 is a flowchart illustrating a method for generating an image using a prompt generation model according to an embodiment of the present disclosure.

[0010] FIG. 3 is a diagram illustrating a pair between a dialogue input to a prompt generation model and a prompt output by the prompt generation model according to one embodiment of the present disclosure.

[0011] FIG. 4 is a diagram illustrating information included in a prompt according to an embodiment of the present disclosure.

[0012] FIG. 5 is a diagram illustrating information included in a prompt according to an embodiment of the present disclosure.

[0013] FIG. 6 is a conceptual diagram illustrating a method for generating a prompt using a prompt generation model according to one embodiment of the present disclosure.

[0014] FIG. 7 is a conceptual diagram illustrating a method for training a prompt generation model and generating prompts using example dialogue-prompt pairs according to one embodiment of the present disclosure.

[0015] FIG. 8 is a conceptual diagram illustrating a method for training a conversation-image generation model and generating images using conversation-image example pairs according to one embodiment of the present disclosure.

[0016] FIG. 9 is a flowchart illustrating a method for generating an image using a prompt generation model according to an embodiment of the present disclosure.

[0017] FIG. 10 is a flowchart illustrating a method for generating an image using a prompt generation model according to an embodiment of the present disclosure.

[0018] FIG. 11 is a flowchart illustrating a method for performing communication with a counterpart user based on an image generated using a prompt generation model according to an embodiment of the present disclosure.

[0019] FIG. 12 is a diagram illustrating a configuration of a computing device for performing image generation using a prompt generation model according to one embodiment of the present disclosure.

[0020] In describing this disclosure, descriptions of technical details that are well-known in the technical field to which this disclosure pertains and are not directly related to this disclosure will be omitted. This is to avoid obscuring the gist of this disclosure by omitting unnecessary explanations and to convey it more clearly. Furthermore, the terms described below are defined based on their functions in this disclosure and may vary depending on the intent or custom of the user or operator. Therefore, their definitions should be based on the contents of this specification as a whole.

[0021] For the same reason, some components in the attached drawings are exaggerated, omitted, or schematically depicted. Furthermore, the dimensions of each component do not entirely reflect its actual size. Identical or corresponding components in each drawing are assigned the same reference numbers.

[0022] The advantages and features of the present disclosure, and methods for achieving them, will become clearer with reference to the embodiments described below in detail with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. The disclosed embodiments are provided to ensure that the disclosure of the present disclosure is complete and to fully inform those skilled in the art of the present disclosure of the scope of the disclosure. An embodiment of the present disclosure may be defined according to the claims. Like reference numerals denote like elements throughout the specification. In addition, when describing an embodiment of the present disclosure, if a detailed description of a related function or configuration is determined to unnecessarily obscure the gist of the present disclosure, the detailed description thereof will be omitted. In addition, the terms described below are terms defined in consideration of the functions of the present disclosure and may vary depending on the intention or custom of the user or operator. Therefore, the definitions should be made based on the contents throughout this specification.

[0023] In one embodiment, each block of the flowchart diagrams and combinations of the flowchart diagrams can be performed by computer program instructions. The computer program instructions can be installed on a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, and the instructions, when executed by the processor of the computer or other programmable data processing apparatus, can create means for performing the functions described in the flowchart block(s). The computer program instructions can also be stored in a computer-available or computer-readable memory that can direct a computer or other programmable data processing apparatus to implement the functions in a particular manner, and the instructions stored in the computer-available or computer-readable memory can also produce an article of manufacture that includes instruction means for performing the functions described in the flowchart block(s). The computer program instructions can also be installed on a computer or other programmable data processing apparatus.

[0024] Additionally, each block in the flowchart diagram may represent a module, segment, or portion of code that includes one or more executable instructions for performing a specified logical function(s). In one embodiment, the functions described in the blocks may occur out of order. For example, two blocks depicted in succession may be executed substantially simultaneously or, depending on the function, may be executed in reverse order.

[0025] The term '~ unit' used in one embodiment of the present disclosure may represent software or a hardware component such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC), and the '~ unit' may perform a specific role. Meanwhile, the '~ unit' is not limited to software or hardware. The '~ unit' may be configured to be on an addressable storage medium and may be configured to play one or more processors. In one embodiment, the '~ unit' may include components such as software components, object-oriented software components, class components, and task components, processes, functions, properties, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functionality provided through a specific component or a specific '~ unit' may be combined to reduce the number of components or separated into additional components. In addition, in one embodiment, the '~ unit' may include one or more processors.

[0026] Embodiments of the present disclosure relate to a method for generating images using a language model. Before describing specific embodiments, the meanings of terms frequently used in this specification are defined.

[0027] In this disclosure, "generative AI" may refer to artificial intelligence technology capable of generating new text, images, etc. in response to input data (e.g., text, images, etc.). Representative examples of generative AI are described in the "Generation Model" section below.

[0028] In this disclosure, a "generative model" may refer to a neural network model that implements generative AI technology. By learning the patterns and structure of training data, a generative model can generate new data with similar characteristics to the input data or new data corresponding to the input data. For example, if the input data is text and the generative model is requested to generate an image corresponding to the text, the generative model can generate a new image that reflects the context of the original text.

[0029] In the present disclosure, a 'language model' may refer to a generative model for obtaining the most natural word sequence by assigning probabilities to word sequences. For example, a language model may obtain text as input data and output a word sequence that describes the context of the obtained text. Although the term 'language model' is used in the present disclosure, this does not limit the technical idea of ​​the present disclosure, and a language model may be expressed as a generative model, an AI model, a language generation model, a natural language processing model, a text generation model, a conversation simulator, a conversational artificial intelligence, and a natural language understanding and generation system, depending on the intention. Among language models, a large language model may be a language model composed of an artificial neural network with a larger number of parameters.

[0030] In the present disclosure, the generative model may include a "prompt generation model." The prompt generation model can acquire text as input data and transform the acquired text to generate prompts. The prompt generation model can learn the patterns and structures of training text data to generate prompts corresponding to the input data. The prompt generation model may be a language model that outputs text reflecting the context of the conversation (e.g., text containing descriptions reflecting the context of the conversation).

[0031] The term "prompt" may refer to input data that instructs an image generation model to perform a task. While a "prompt" may collectively refer to input data that instructs a generation model to perform a task, in the present disclosure, the term "prompt" is described as meaning input data that instructs an "image generation model" to perform a task to avoid confusion in meaning.

[0032] In the present disclosure, the generative model may include an "image generation model." The image generation model can acquire text as input data and transform the acquired text to generate an image. The image generation model can generate an image corresponding to the input data by learning the patterns and structures of the training text data.

[0033] Hereinafter, a method for generating an image using a language model according to embodiments of the present disclosure and a computing device for performing the same will be described with reference to drawings.

[0034] The processes described in this disclosure are assumed to be performed by a computing device supporting image processing capabilities. Therefore, in describing FIGS. 1 through 12 , the computing device is described as performing the processes. Detailed components included in a computing device according to one embodiment are illustrated in FIG. 12 , and these components will be described in detail later.

[0035] In embodiments of the present disclosure, a computing device can input a conversation into a prompt generation model to obtain a prompt corresponding to the conversation, thereby generating an appropriate prompt to be input into the image generation model. The final image generation from the conversation can be performed using two generation models, but efficiency and performance can also be improved by using a single generation model that integrates each step.

[0036] FIG. 1 is a conceptual diagram illustrating a process of generating an image using a prompt generation model according to an embodiment of the present disclosure.

[0037] Referring to FIG. 1, a computing device can generate an image using multiple generation models.

[0038] In one embodiment, the plurality of generation models may include a prompt generation model (200) and an image generation model (300).

[0039] The prompt generation model (200) is a model for obtaining a prompt (20) to be input into the image generation model (300). The prompt (20) may refer to input data that instructs the image generation model (300) to perform a task. The prompt (20) needs to be configured to be easily recognized by the image generation model (300), and the prompt generation model (200) may be a model that converts input data into a prompt (20) so that the image generation model (300) can easily recognize it.

[0040] The prompt generation model (200) may be a language model trained through multiple dialogue-prompt pairs. The prompt generation model (200) may be a language model trained through training data including multiple dialogue-prompt pairs. The prompt generation model (200) may also be a general language model (e.g., a large language model) trained through various types of text data. The prompt generation model (200) may learn patterns of transitions between dialogues and prompts from multiple dialogue-prompt pairs, thereby generating a prompt (20) from a new dialogue (10). Of course, the learning method of the prompt generation model (200) does not limit the technical idea of ​​the present disclosure.

[0041] In one embodiment, the prompt generation model (200) may acquire user conversations (10) as input data. For example, the user conversations (10) may include everyday expressions such as greetings, questions, and requests between users. Specifically, the user conversations (10) may include everyday expressions such as "hello" and "want to go get pizza?"

[0042] The user's dialogue (10) may be a text composed without considering whether it can be recognized by a generative model (e.g., an image generation model). The prompt generation model (200) may be a model that generates a prompt (20) by reconstructing the user's dialogue so that another generative model (e.g., an image generation model) can easily recognize the user's dialogue (10) as input data.

[0043] The data type of the user's conversation (10) input to the prompt generation model (200) does not limit the technical concept of the present disclosure. For example, the user's conversation (10) may be text data or voice data.

[0044] In one embodiment, the prompt generation model (200) may be a language model that converts text into text. The user's dialogue (10) may be text data obtained based on user input. The prompt generation model (200) may generate a text-based prompt (20) by converting the text-based dialogue (10).

[0045] In one embodiment, the prompt generation model (200) may be a language model that converts voice into text. The user's dialogue (10) may be voice data acquired based on user input. The prompt generation model (200) may generate a text-based prompt (20) by converting the voice-based dialogue (10).

[0046] In one embodiment, a prompt generation model (200) may input a user's conversation (10) and output a prompt (20). The prompt (20) may be text data generated to reflect the context of the conversation (10). The prompt (20) may include a description reflecting the context of the conversation (10). Specifically, the prompt (20) may be text including a description for generating an image reflecting the context of the user's conversation (10). Alternatively, the prompt (20) may be text including a description of an image corresponding to the context of the user's conversation (10). The configuration of the prompt (20) will be described in detail below using FIGS. 3 to 5 .

[0047] The image generation model (300) is a model for obtaining an image (30) from a prompt (20). The prompt (20) may be data converted to reflect the context of the user's conversation (10), and the image generation model (300) may ultimately output an image (30) from the prompt (20). The image (30) may be an image converted to reflect the context of the user's conversation (10).

[0048] For example, the image generation model (300) may be a generative model trained through multiple prompt-image pairs. The image generation model (300) may be a generative model trained through training data including multiple prompt-image pairs. The image generation model (300) may learn patterns of transitions between prompts and images from multiple prompt-image pairs, thereby generating an image (30) from a new prompt (20). Of course, the learning method of the image generation model (300) does not limit the technical idea of ​​the present disclosure.

[0049] In one embodiment, the image generation model (300) can obtain a prompt (20) as input data. The image generation model (300) can be a generation model that converts text into an image. The user's dialogue (10) can be text data obtained according to user input. The prompt generation model (200) can generate a text-type prompt (20) by converting the text-type dialogue (10).

[0050] FIG. 2 is a flowchart illustrating a method for generating an image using a prompt generation model according to an embodiment of the present disclosure. For convenience of explanation, details that overlap with those described using FIG. 1 are simplified or omitted.

[0051] In step S210, the computing device can receive a dialogue. The computing device can obtain user input for communication between users as dialogue. The computing device can obtain the dialogue through an input interface. The dialogue can be obtained in the form of text or voice.

[0052] In step S220, the computing device can input the conversation into a prompt generation model to obtain a prompt corresponding to the conversation.

[0053] In one embodiment, the prompt generation model may be a language model trained through multiple dialogue-prompt pairs. The prompt generation model may be a model that has learned transition patterns between dialogues and prompts from multiple dialogue-prompt pairs. The computing device may use the trained prompt generation model to obtain prompts from the dialogue.

[0054] In one embodiment, the acquired prompt may include a description reflecting the context of the conversation. The acquired prompt may include text data describing the context of the conversation.

[0055] In step S230, the computing device can acquire an image by inputting a prompt into the image generation model.

[0056] In one embodiment, the image generation model may be a pre-trained generative model trained on multiple prompt-image pairs. The image generation model may be a model that learns a transformation pattern between prompts and images from multiple prompt-image pairs. The computing device may use the pre-trained image generation model to obtain an image from the prompt.

[0057] In one embodiment, the acquired image may be generated in response to a prompt. Since the prompt reflects the context of the conversation, the acquired image may also reflect the context of the conversation. The acquired image may include at least one image composed to reflect the context of the conversation.

[0058] FIG. 3 is a diagram illustrating pairs between a dialogue input into a prompt generation model and a prompt output by the prompt generation model according to one embodiment of the present disclosure. For reference, FIG. 3 is a diagram illustrating a table of correspondences between dialogue, which is input data of the prompt generation model, and prompts, which is output data.

[0059] Referring to FIG. 3, a computing device can acquire a user input such as "Let's go eat pizza." as a dialogue, and the acquired dialogue can be input into a prompt generation model. In response to the user input "Let's go eat pizza.", the prompt generation model can output the prompt "Looking at a pizza place, smiling."

[0060] For example, the text input "Let's go eat pizza" may not be easy for an image generation model to extract its context. An image generation model may generate an image based on the meaning of "pizza," "eat," or "go" from the data "Let's go eat pizza." Therefore, an action may be required to convert the conversation into prompts so that the image generation model can easily determine the context of the conversation. Furthermore, since the conversation "Let's go eat pizza" does not include information about which image to generate, i.e., a description of the image, if this conversation is input directly into an image generation model, the image output from the image generation model may not properly reflect the context of the conversation or the user's intention.

[0061] The prompt "Looking at a pizza place and smiling" can include action descriptions, corresponding to the context of the conversation "Let's go eat pizza." That is, the prompt "Looking at a pizza place and smiling" can include descriptions of the actions "looking at a pizza place" and "smiling." Of course, a prompt can include multiple action descriptions, or it can include a single description.

[0062] In one embodiment, as illustrated in FIG. 3, the prompt is structured as 'looking at a pizza shop, smiling', including descriptions of the actions 'looking at a pizza shop' and 'smiling', but it is of course possible for the prompt to be structured as a sentence structure of 'looking at a pizza shop + smiling', taking into account the patterns learned by the image generation model.

[0063] As another example, a prompt may have the same meaning, but structured as "See a pizza place. + Smile." The structure of the prompt is merely designed to facilitate the image generation model in determining the context of the prompt as input data, and does not limit the technical concept of the present disclosure.

[0064] A computing device can acquire user input such as "I'm playing with a cat" as a conversation, and input the acquired conversation into a prompt generation model. The prompt generation model can then output the prompt "I'm petting a cat" in response to the user input "I'm playing with a cat."

[0065] For example, the text input "I'm playing with a cat" may not be easily extractable by an image generation model. An image generation model could generate an image based on the meaning of "cat" or "play" from the data "I'm playing with a cat." Therefore, converting the dialogue into prompts may be required to facilitate the image generation model's ability to determine the context of the conversation.

[0066] The prompt "petting a cat" could include a description of an action, corresponding to the context of the conversation "I'm playing with a cat." That is, the prompt "petting a cat" could include the action "petting a cat."

[0067] In one embodiment, as illustrated in FIG. 3, the prompt is configured to include a description of the action "petting a cat." However, it should be understood that the prompt could also be structured as "petting a cat" considering patterns learned by the image generation model. The structure of the prompt is merely configured to facilitate the image generation model in determining the context of the prompt as input data, and does not limit the technical concept of the present disclosure.

[0068] In one embodiment, the same conversation may be transformed into various prompts based on the training of the prompt generation model. For example, from the input data "I am playing with a cat," the prompt generation model may generate the prompt "petting a cat," as illustrated in FIG. 3. However, it may also generate one of various prompts, such as "feeding a cat," "playing with a toy with a cat," etc.

[0069] In one embodiment, a computing device may acquire a user input such as "I think I'll be late today" as a dialogue, and the acquired dialogue may be input into a prompt generation model. The prompt generation model may output the prompt "Looking at a clock to check the time" in response to the user input "I think I'll be late today."

[0070] In one embodiment, a computing device may acquire a user input such as "What's the problem?" as a dialogue, and the acquired dialogue may be input into a prompt generation model. The prompt generation model may output the prompt "Frowning, head down" in response to the user input "What's the problem?"

[0071] The method of obtaining each prompt is the same as that of obtaining the prompts 'Looking at a pizza shop and smiling' and 'Stroking a cat', so duplicate explanations are omitted.

[0072] FIG. 4 is a diagram illustrating information included in a prompt according to an embodiment of the present disclosure.

[0073] Referring to Figure 4, the prompt generation model can acquire the dialogue "Let's go eat pizza" as input data and output the prompt "Looking at a pizza place, smiling." In one embodiment, the prompt can include a description of an object's behavior corresponding to the context of the dialogue.

[0074] A computing device can receive a user's conversation and, using a prompt generation model, obtain a prompt corresponding to the context of the conversation. The prompt may include a description of an object's behavior corresponding to the context of the conversation.

[0075] In one embodiment, the prompt may include, as essential information, a description of the object's behavior that corresponds to the context of the conversation. In addition to the description of the object's behavior, the prompt may also include additional information. The additional information included in the prompt is described in detail below using Figure 5.

[0076] For convenience of explanation, we will use the prompt "Looking at a pizza place and smiling" as an example, as shown in Figure 4. This prompt includes the descriptions "Looking at a pizza place" and "Smiling." The descriptions "Looking at a pizza place" and "Smiling" may each be descriptions of an unseen object's behavior.

[0077] Of course, the example prompt described in Figure 4 is structured so that no object is revealed, but it is certainly possible to specify an object. For example, in response to the dialogue "Let's go eat pizza," a prompt could be generated with the structure "A smiling dog character looking at a pizza place."

[0078] FIG. 5 is a diagram illustrating information included in a prompt according to an embodiment of the present disclosure. For convenience of explanation, details that overlap with those described using FIG. 4 are simplified or omitted.

[0079] Referring to Figure 5, the prompt generation model can obtain the dialogue 'Let's go eat pizza' as input data and output the prompt "i) smiling while looking at a pizza shop ii) focus line emphasizing the pizza shop iii) insertion of 'Pizza!!' in cursive font iv) character is a dog."

[0080] In one embodiment, the prompt may include a description of an object's behavior that corresponds to the context of the conversation. The prompt may include data such as "i) looking at a pizza place and smiling" as a description of the object's behavior. The description of the object's behavior may be included as essential information in the prompt, and is omitted as it overlaps with the description using Figure 4.

[0081] In one embodiment, the prompt may include a description of an image effect. The prompt may include data such as "ii) a focus line highlighting a pizza place" as a description of the image effect. The description of the image effect may be included as additional information in the prompt, and the prompt may also be configured to exclude the description of the image effect.

[0082] Descriptions of image effects can be expressed in various ways, such as the background of the image (e.g., no background, sky background, etc.), the number of colors used in the image (e.g., black and white image, color image, etc.), and the drawing style (e.g., cartoon style, cubism style, etc.).

[0083] In one embodiment, the prompt may include a description of the inserted text. The prompt may include data such as "iii) Insert 'Pizza!!', cursive font" as a description of the inserted text. The description of the inserted text may be included as additional information in the prompt, and the prompt may also be configured to exclude the description of the inserted text.

[0084] Descriptions of the text to be inserted can include a variety of possible effects for inserting text into an image, such as the content of the text to be inserted, font, size, color, and effects for the text to be inserted (e.g., italics, bold, emphasis, 3D expression, etc.).

[0085] In one embodiment, the prompt may include a description of the selection of an object. The prompt may include data such as "iv) the character is a puppy" as a description of the selection of an object. The description of the selection of an object may be included as additional information in the prompt, and the prompt may also be configured to exclude the description of the selection of an object.

[0086] Descriptions of object selection can be expressed in various ways through descriptions that specify the object, such as its type, size, and color.

[0087] In one embodiment, the prompt may essentially include a description of an object's behavior that corresponds to the context of the conversation. The prompt may optionally further include at least one of a description of an image effect, a description of inserted text, and a description of selecting the object.

[0088] For example, the prompt in FIG. 5 is structured as "i) looking at a pizza shop, smiling ii) a focused line emphasizing the pizza shop iii) inserting 'Pizza!!', cursive font iv) the character is a puppy", but some of the prompts "ii) a focused line emphasizing the pizza shop iii) inserting 'Pizza!!', cursive font iv) the character is a puppy" may be excluded. The prompt may also be structured as "i) looking at a pizza shop, smiling ii) inserting 'Pizza!!', cursive font iii) the character is a puppy" excluding the description of the image effect.

[0089] FIG. 6 is a conceptual diagram illustrating a method for generating a prompt using a prompt generation model according to an embodiment of the present disclosure. For convenience of explanation, details that overlap with those described using FIGS. 1 and 2 are simplified or omitted.

[0090] Referring to FIG. 6, a computing device can input a conversation into a prompt generation model (200) to obtain a prompt (620) corresponding to the conversation.

[0091] In one embodiment, a conversation (611) is input to a prompt generation model (200), and the prompt generation model (200) can generate a prompt (620) corresponding to the conversation (611).

[0092] Of course, as illustrated in FIG. 6, the data input to the prompt generation model (200) may be at least one dialogue-prompt example pair and a combination of dialogues (610). The combination of at least one dialogue-prompt example pair and dialogues (610) may be data listing at least one dialogue-prompt example pair and dialogues.

[0093] In one embodiment, at least one of the dialogue-prompt example pairs may be obtained using a large language model trained over multiple dialogue-prompt pairs, or may be data input by a user.

[0094] For example, at least one dialog-prompt example pair may include a pair of dialog examples "Hello!" and a corresponding prompt example "Waving." At least one dialog-prompt example pair may include a pair of dialog examples "I adopted a puppy!" and a corresponding prompt example "Holding a puppy." At least one dialog-prompt example pair may include a pair of dialog examples "Cheer up!" and a corresponding prompt example "Clenching a fist."

[0095] In one embodiment, the combination of at least one dialogue-prompt example pair and dialogue (610) may be data including a dialogue (611). The dialogue (611) may be data input by a user. The computing device may obtain text data input by the user as the dialogue (611) using an input interface, or may obtain voice data input by the user as the dialogue (611). For example, the computing device may obtain the dialogue (611) "Aren't you hungry?" based on a user input.

[0096] In one embodiment, a computing device may input at least one dialogue-prompt example pair and a combination of dialogues (610) into a prompt generation model (200), thereby obtaining a prompt (620) corresponding to the at least one dialogue-prompt example pair and the combination of dialogues (610). The obtained prompt (620) may be a description reflecting the context of the dialogue (611). The obtained prompt (620) may be a description reflecting the context of the dialogue (611), based on at least one dialogue-prompt example pair.

[0097] In one embodiment, the prompt generation model (200) can output a prompt (620) corresponding to a conversation (611) based on a relationship between a conversation example and a prompt example included in at least one conversation-prompt example pair and a combination of conversations (610).

[0098] In one embodiment, the prompt generation model (200) may be a Large Language Model (LLM) trained using various text data. If the prompt generation model (200) is a Large Language Model, and at least one dialogue-prompt example pair and a combination of dialogues (610) as described above are input to the prompt generation model (200), the prompt generation model (200) may be fine-tuned by the at least one dialogue-prompt example pair included in the input. Accordingly, the prompt generation model (200) may output a prompt (620) corresponding to the dialogue (611) based on the relationship between the dialogue examples and the prompt examples included in the at least one dialogue-prompt example pair. In other words, the prompt generation model (200) may be additionally trained using at least one dialogue-prompt example pair, and may output a prompt (620) corresponding to the dialogue (611) by reflecting the training result.

[0099] FIG. 7 is a conceptual diagram illustrating a method for training a prompt generation model and generating prompts using example dialogue-prompt pairs according to one embodiment of the present disclosure.

[0100] For convenience of explanation, parts that overlap with those described using Figure 1 are simplified or omitted.

[0101] Referring to FIG. 7, a computing device can input a conversation (710) into a prompt generation model (200) to obtain a prompt (720) corresponding to the conversation (710).

[0102] In one embodiment, the prompt generation model (200) may be a language model trained through multiple dialogue-prompt pairs. The prompt generation model (200) may be a pre-trained language model using multiple text data.

[0103] In one embodiment, the prompt generation model (200) may not be constructed with a sufficiently large model size for lightweight purposes, and may be trained with a small amount of training data. In this case, a step of fine-tuning the prompt (720) output by the prompt generation model (200) may be required to obtain an appropriate prompt (720). Through the fine-tuning step, the weight of the prompt generation model (200) can be updated, and the fine-tuned prompt generation model (200) can output a more appropriate prompt in response to the conversation (710). The fine-tuned prompt generation model (200) can output a description as a prompt that more appropriately reflects the context of the conversation (710).

[0104] In the present disclosure, at least one dialogue-prompt example pair (715) for fine-tuning a prompt generation model (200) and a plurality of dialogue-prompt pairs for training the prompt generation model (200) may be data regarding pairs of dialogues and prompts that correspond to each other, respectively. However, at least one dialogue-prompt example pair (715) is data used for fine-tuning a prompt (720) output by the prompt generation model (200), and the plurality of dialogue-prompt pairs are training data used when the prompt generation model (200) is pre-trained.

[0105] In one embodiment, the prompt generation model (200) can fine-tune the prompt (720) based on at least one dialog-prompt example pair.

[0106] For example, a computing device can obtain a prompt (720) from a conversation (710) using a pre-trained prompt generation model (200). The computing device can retrain the pre-trained prompt generation model (200) based on at least one conversation-prompt example pair. By retraining the pre-trained prompt generation model (200) based on at least one conversation-prompt example pair, weights can be updated.

[0107] Techniques for fine-tuning do not limit the technical idea of ​​the present disclosure. For example, the prompt generation model (200) can be retrained using techniques for fine-tuning, such as Stochastic Gradient Descent (SGD), Adaptive Moment Estimation (ADAM), Adaptive Gradient (Adagrad), and Nesterov Accelerated Gradient (NAG).

[0108] In one embodiment, at least one of the dialogue-prompt example pairs (715) may be obtained using a separate large language model trained over multiple dialogue-prompt pairs, or may be data input by a user.

[0109] For example, at least one dialog-prompt example pair may include a pair of dialog examples "Hello!" and a corresponding prompt example "Waving." At least one dialog-prompt example pair may include a pair of dialog examples "I adopted a puppy!" and a corresponding prompt example "Holding a puppy." At least one dialog-prompt example pair may include a pair of dialog examples "Cheer up!" and a corresponding prompt example "Clenching a fist."

[0110] Below, an example of an operation of retraining a prompt generation model (200) to fine-tune a prompt (720) generated by a computing device is specifically described.

[0111] In one embodiment, the prompt generation model (200) can acquire a user's dialogue (710). The acquired dialogue (710) can be, for example, a text input or voice input such as 'Aren't you hungry?'

[0112] The prompt generation model (200) can generate a prompt (720) corresponding to the conversation (710) "Aren't you hungry?" The generated prompt (720) can include a description reflecting the context of the conversation (710) "Aren't you hungry?" For example, the prompt (720) can include text data such as "the image of a stomach rumbling," "the image of holding a stomach," and "the image of imagining a food one wants to eat," which describe the context of the conversation (710) "Aren't you hungry?"

[0113] In one embodiment, the computing device may use the prompt generation model (200) to generate a single prompt or a list of multiple prompts.

[0114] In one embodiment, a computing device may obtain at least one dialog-prompt example pair. The computing device may obtain a dialog example corresponding to a dialog (710) input to the prompt generation model (200) from the at least one dialog-prompt example pair (715).

[0115] For example, in response to the dialogue (710) input to the prompt generation model (200), “Aren’t you hungry?”, the computing device can extract the dialogue example “Aren’t you hungry?” from at least one dialogue-prompt example pair (715). In FIG. 7, the dialogue (710) input to the prompt generation model (200) and the dialogue example extracted from the at least one dialogue-prompt example pair (715) are illustrated as being composed of the same expressions, but the dialogue (710) and the dialogue example may not be exactly identical. The computing device can extract the dialogue example from at least one dialogue-prompt example pair (715) to the extent that they share the context of the dialogue (710). The dialogue (710) and the dialogue example may literally match, or the context of the dialogue (710) and the context of the dialogue example may be identical.

[0116] In one embodiment, the computing device can extract a prompt example corresponding to the extracted dialog example from at least one dialog-prompt example pair (715). For example, the computing device can extract a dialog example from at least one dialog-prompt example pair (715). The computing device can obtain the dialog example 'Aren't you hungry?' from the at least one dialog-prompt example pair (715). The computing device can extract a prompt example corresponding to the dialog example 'Aren't you hungry?'. The computing device can extract the prompt example 'Holding your stomach' corresponding to 'Aren't you hungry?'

[0117] In one embodiment, the computing device can fine-tune the prompt generation model (200) by comparing prompts (720) generated by the prompt generation model (200) with prompt examples extracted from at least one pair of conversational prompt examples (715). The computing device can retrain the prompt generation model (200) based on the extracted prompt examples. The computing device can update the weights of the prompt generation model (200) based on the extracted prompt examples.

[0118] In one embodiment, the computing device can fine-tune the prompt generation model (200) based on the difference between the prompt (720) generated by the prompt generation model (200) and the prompt examples extracted from the at least one pair of dialogue prompt examples (715). For example, the computing device can set the difference between the prompt (720) generated by the prompt generation model (200) and the prompt examples extracted from the at least one pair of dialogue prompt examples (715) as a loss function, and retrain the prompt generation model (200) so that the loss function is minimized.

[0119] However, the method by which the computing device fine-tunes the prompt generation model (200) is merely an example and does not limit the technical concepts of the present disclosure. For example, various fine-tuning techniques, such as Stochastic Gradient Descent (SGD) and Adaptive Moment Estimation (ADAM), can be utilized.

[0120] FIG. 8 is a conceptual diagram illustrating a method for training a conversation-image generation model and generating images using conversation-image example pairs according to one embodiment of the present disclosure.

[0121] For convenience of explanation, parts that overlap with those described using Figures 1 and 7 are simplified or omitted.

[0122] Referring to FIG. 8, a computing device can input a conversation (810) into a conversation-image generation model (800) to obtain an image (820) corresponding to the conversation (810).

[0123] In one embodiment, the conversation-to-image generation model (800) may be a single generative model that comprehensively performs the operations performed by the prompt generation model (200) and the image generation model (300) described using FIGS. 1 and 7. For example, the conversation-to-image generation model (800) may be a generative model that converts a conversation (810) into a prompt and generates an image (820) based on the prompt, or may be a generative model that generates an image (820) from the conversation (810). The network structure of the conversation-to-image generation model (800) does not limit the technical idea of ​​the present disclosure. Consequently, the conversation-to-image generation model (800) may be a generative model that inputs a conversation (810) and outputs an image (820) corresponding to the input conversation (810).

[0124] In one embodiment, the conversation-image generation model (800) may be a generative model trained through multiple conversation-image pairs. The conversation-image generation model (800) may be a pre-trained generative model that is trained to convert multiple text data into image data.

[0125] In one embodiment, the dialogue-image generation model (800) may not be constructed to be sufficiently large in size for lightweight purposes, and may be trained with a small amount of training data. In this case, a step of fine-tuning the image (820) output by the dialogue-image generation model (800) may be required to obtain an appropriate image (820). Through the fine-tuning step, the weights of the dialogue-image generation model (800) may be updated, and the fine-tuned dialogue-image generation model (800) may output a more appropriate image corresponding to the dialogue (810). The fine-tuned dialogue-image generation model (800) may output an image that more appropriately reflects the context of the dialogue (810) compared to the generated image (820).

[0126] In the present disclosure, at least one dialogue-image example pair (815) for fine-tuning a dialogue-image generation model (800) and a plurality of dialogue-image pairs for training the dialogue-image generation model (800) may be data regarding pairs of dialogues and images that correspond to each other, respectively. However, at least one dialogue-image example pair (815) is data used for fine-tuning an image (820) output by the dialogue-image generation model (800), and the plurality of dialogue-image pairs are training data used when the dialogue-image generation model (800) is pre-trained.

[0127] In one embodiment, the conversation-image generation model (800) can fine-tune an image (820) based on at least one conversation-image example pair (815).

[0128] For example, a computing device can obtain an image (820) from a conversation (810) using a pre-trained conversation-image generation model (800). The computing device can retrain the pre-trained conversation-image generation model (800) based on at least one conversation-image example pair (815). The pre-trained conversation-image generation model (800) can have its weights updated by being retrained based on at least one conversation-image example pair (815).

[0129] Techniques for fine-tuning do not limit the technical concepts of the present disclosure. For example, the dialogue-image generation model (800) can be retrained using techniques for fine-tuning, such as Stochastic Gradient Descent (SGD), Adaptive Moment Estimation (ADAM), Adaptive Gradient (Adagrad), and Nesterov Accelerated Gradient (NAG).

[0130] In one embodiment, at least one dialogue-image example pair (815) may be obtained using a separate generative model trained through multiple dialogue-image pairs, or may be data input by a user.

[0131] For example, at least one dialogue-image example pair (815) may include a dialogue example "I'm studying in the library" and a corresponding image example depicting "looking at a book with a library in the background." At least one dialogue-prompt example pair (815) may include a dialogue example "Would you like some coffee?" and a corresponding image example depicting "holding coffee." At least one dialogue-prompt example pair (815) may include a dialogue example "I'm playing with a cat" and a corresponding image example depicting "clapping while looking at a cat."

[0132] Below, the operation of retraining the conversation-image generation model (800) to fine-tune the image (820) generated by the conversation-image generation model (800) is specifically described.

[0133] In one embodiment, the conversation-image generation model (800) can acquire a user's conversation (810). The acquired conversation (810) can be, for example, a text input or voice input such as 'I'm playing with a cat.'

[0134] The conversation-image generation model (800) can generate an image (820) corresponding to the conversation (810) of "I'm playing with a cat." The generated image (820) can include an image depicting the context of the conversation (810) of "I'm playing with a cat." For example, the image (820) can include images depicting "applause while looking at a cat," "playing with a cat using a cat toy," "petting a cat," etc., corresponding to the context of the conversation (810) of "I'm playing with a cat."

[0135] In one embodiment, the computing device may generate a single image or a list of multiple images using the conversation-image generation model (800).

[0136] In one embodiment, the computing device can obtain at least one dialogue-image example pair (815). The computing device can obtain, from the at least one dialogue-image example pair (815), a dialogue example corresponding to a dialogue (810) input to the dialogue-image generation model (800).

[0137] For example, the computing device can extract a dialogue example 'I'm playing with the cat' from at least one dialogue-image example pair (815), corresponding to the dialogue (810) 'I'm playing with the cat' input into the dialogue-image generation model (800). In FIG. 8, the dialogue (810) input into the dialogue-image generation model (800) and the dialogue example extracted from at least one dialogue-prompt example pair (815) are illustrated as being composed of the same expressions, but the dialogue (810) and the dialogue example may not be exactly identical. The computing device can extract the dialogue example from at least one dialogue-image example pair (815) to the extent that they share the context of the dialogue (810). The dialogue (810) and the dialogue example may literally match, or the context of the dialogue (810) and the context of the dialogue example may be identical.

[0138] In one embodiment, the computing device can extract an image example corresponding to the extracted dialogue example from at least one dialogue-image example pair (815). For example, the computing device can extract a dialogue example from at least one dialogue-image example pair (815). The computing device can obtain the dialogue example 'I'm playing with the cat' from the at least one dialogue-image example pair (815). The computing device can extract an image example corresponding to the dialogue example 'I'm playing with the cat'. The computing device can extract an image example depicting 'an image of clapping while looking at a cat' corresponding to the context of 'I'm playing with the cat'.

[0139] In one embodiment, the computing device can fine-tune the conversation-image generation model (800) by comparing images (820) generated by the conversation-image generation model (800) with image examples (825) extracted from at least one conversation-image example pair (815). The computing device can retrain the conversation-image generation model (800) based on the extracted image examples (825). The computing device can update weights of the conversation-image generation model (800) based on the extracted image examples (825).

[0140] In one embodiment, the computing device can fine-tune the conversation-image generation model (800) based on the difference between the image (820) generated by the conversation-image generation model (800) and the image example (825) extracted from at least one conversation-image example pair (815). For example, the computing device can set the difference between the image (820) generated by the conversation-image generation model (800) and the image example (825) extracted from at least one conversation-image example pair (815) as a loss function, and retrain the conversation-image generation model (800) such that the loss function is minimized.

[0141] However, the method by which the computing device fine-tunes the conversation-image generation model (800) is merely an example and does not limit the technical concepts of the present disclosure. For example, various fine-tuning techniques, such as Stochastic Gradient Descent (SGD) and Adaptive Moment Estimation (ADAM), can be utilized.

[0142] FIG. 9 is a flowchart illustrating a method for generating an image using a prompt generation model according to an embodiment of the present disclosure.

[0143] For convenience of explanation, parts that overlap with those described using Figure 2 are simplified or omitted.

[0144] Step S230 of FIG. 2 may include steps S910 and S920.

[0145] Referring to FIG. 9, in step S910, the computing device can determine a style for generating an image.

[0146] A computing device can determine a style for generating an image based on user input. The computing device can obtain user input for selecting a style for generating an image using an input interface, and can determine a style for generating an image based on the user input.

[0147] In one embodiment, the style for generating an image may refer to a style in which the image generation model generates an image from a prompt. The style for generating an image may include settings for the implementation scope of the image, which is output data of the image generation model.

[0148] For example, a computing device may provide a user interface for selecting an image generation style and obtain user input regarding the selection of the image generation style. The image generation style may include, for example, styles such as "exact" and "creative."

[0149] In step S920, the computing device can obtain an image corresponding to the prompt based on the determined method.

[0150] When a computing device obtains user input selecting an "exact" style, the image generation model can be used to obtain an image that faithfully matches the context of the prompt. When the image generation model generates an image based on the "exact" style, the image can be implemented in a limited style. When the computing device obtains user input selecting a "creative" style, the image generation model can be used to generate an image that corresponds to the context of the prompt, but the generated image can be implemented more freely or creatively. Accordingly, the image generation model can implement images in a wider variety of forms even when receiving the same dialogue input.

[0151] In one embodiment, the computing device may acquire multiple images corresponding to the prompt based on a determined method and provide the user with a list of generated images. For example, the computing device may generate images corresponding to the context of the prompt based on user input regarding the style of "creative." In this case, the images may be implemented more freely, and multiple images may be generated in response to the prompt. The computing device may provide the user with a list of images based on the generated multiple images.

[0152] In one embodiment, the computing device may provide a service for displaying or transmitting the selected image to an external user terminal based on a user input for selecting one image from a list of images.

[0153] In one embodiment, a computing device may acquire multiple images depicting a conversation. The computing device may acquire a user input for selecting a first image from the multiple images. Based on the acquired user input, the computing device may transmit at least one of the first image and the conversation to an external user terminal.

[0154] In one embodiment, the operation of obtaining an image based on the style information described in FIG. 9 may be performed through the dialogue-image generation model (800) described in FIG. 8, replacing the prompt generation model and the image generation model.

[0155] For example, a computing device can receive a user's conversation. The computing device can input the conversation into a conversation-image generation model to obtain an image. Specifically, the computing device can obtain style information related to the style in which the image is generated. Based on the style information, the computing device can obtain an image corresponding to the conversation. The computing device inputs the conversation into the conversation-image generation model to obtain an image, and the obtained image may be an image generated based on the style information.

[0156] Duplicate descriptions regarding the operation of obtaining an image based on style information are omitted.

[0157] FIG. 10 is a flowchart illustrating a method for generating an image using a prompt generation model according to an embodiment of the present disclosure.

[0158] For convenience of explanation, parts that overlap with those described using Figure 2 are simplified or omitted.

[0159] Step S220 of FIG. 2 may include steps S1010 and S1020.

[0160] Referring to FIG. 10, in step S1010, the computing device can obtain personal information about the subject of the conversation.

[0161] In one embodiment, a computing device may obtain personal information about a user entering a conversation based on user input. The computing device may obtain personal information about the subject of the conversation using an input interface.

[0162] In one embodiment, personal information about the subject of a conversation may include personal information that can identify a specific person, status information that can confirm the status of a specific person, information related to the surrounding environment of a specific person, etc.

[0163] For example, information related to personal identification may include information that identifies a specific person, such as name, resident registration number, age, gender, place of birth, etc., or information that can be easily combined with other information to identify a specific person.

[0164] For example, status information may include information that can be used to determine the status of a specific person, such as the target person's heart rate, body temperature, stress, and sleep time.

[0165] For example, information related to the surrounding environment may include information about the surrounding environment of a specific person, such as information about the place where the specific person is located, information about the weather at the place where the specific person is located, information about the temperature at the place where the specific person is located, and information about the altitude at the place where the specific person is located.

[0166] In one embodiment, a computing device may obtain information related to the surrounding environment of a specific individual based on the individual's location. For example, the computing device may determine the individual's location and obtain information related to the surrounding environment based on the determined location.

[0167] In step S1020, the computing device may obtain a prompt corresponding to the conversation based on personal information.

[0168] For example, a computing device can obtain information related to personal information. The computing device can obtain information related to the user's gender (e.g., female) and age (e.g., 12 years old). The conversation obtained in step S210 may be "Where should we go to play today?", and the computing device can obtain a prompt corresponding to "Where should we go to play today?" based on the information related to personal information. The computing device can obtain a prompt such as "Tell me a place a teenage girl would like to go."

[0169] As another example, a computing device can obtain status information. The computing device can obtain information about the user's stress index (e.g., high). The conversation obtained in step S210 may be "Where should we go for fun today?", and the computing device can obtain a prompt corresponding to "Where should we go for fun today?" based on the status information. The computing device can obtain a prompt such as "Tell me a good place to go for a break."

[0170] As another example, a computing device can obtain information related to the surrounding environment. The computing device can obtain information about the location where the user is located (e.g., Seogwipo-si, Jeju Island) and information about the weather at the location where the user is located (e.g., rain). The conversation obtained in step S210 may be "Where should we go for fun today?", and the computing device can obtain a prompt corresponding to "Where should we go for fun today?" based on the information related to the surrounding environment. The computing device can obtain the prompt "Tell me a good place to go in Seogwipo-si, Jeju Island on a rainy day."

[0171] As another example, a computing device may obtain at least one of personal information, status information, and information related to the surrounding environment. The computing device may obtain a prompt corresponding to the conversation based on at least one of personal information (e.g., female, 12 years old), status information (e.g., high stress level), and information related to the surrounding environment (e.g., Seogwipo-si, Jeju-do, rain). In response to the conversation "Where should we go for fun today?", the computing device may obtain the prompt "Tell me a place in Seogwipo-si, Jeju-do that a teenage girl would like to go to for relaxation on a rainy day."

[0172] In one embodiment, the action of obtaining a prompt based on personal information, as described in FIG. 10, can be utilized as an action of obtaining various content other than a prompt. The action of converting a conversation into content based on personal information can also be performed through a conversation-content generation model that generates content corresponding to the conversation from the conversation.

[0173] In one embodiment, content in a conversational-content generation model may include various types of content, such as images, videos, code, and text.

[0174] For example, an operation of converting a conversation into an image based on personal information may be performed through a conversation-image generation model (800) described in FIG. 8, replacing the prompt generation model and the image generation model. For example, if personal information (e.g., female, 12 years old) is obtained, the computing device may convert the conversation into an image based on the personal information, and the converted image may include an image of a character preferred by 12-year-old girls.

[0175] As another example, an action of converting a conversation into a video based on personal information may be performed through a conversation-to-video generation model that generates a video corresponding to the conversation from the conversation.

[0176] FIG. 11 is a flowchart illustrating a method for performing communication with a counterpart user based on an image generated using a prompt generation model according to an embodiment of the present disclosure.

[0177] Referring to FIG. 11, in step S1110, the computing device can receive a dialogue. The description of step S1110 is omitted as it overlaps with the description using step S210 of FIG. 2.

[0178] In step S1120, the computing device can determine whether to transmit the conversation without an image. The computing device can determine whether to transmit the conversation between users to the terminal of the other user.

[0179] In one embodiment, a computing device may obtain user input from a user entering a conversation between users to determine whether to transmit the conversation without an image. The computing device may determine whether to transmit the conversation without an image based on the obtained user input.

[0180] If it is determined in step S1120 to convey the conversation with an image, step S1130 may be performed.

[0181] In step S1130, the computing device can determine whether to convert the dialogue into a prompt by activating the prompt generation model. In step S1140, if the prompt generation model is activated, the computing device can convert the dialogue into a prompt using the prompt generation model. The description of the operation according to step S1140 overlaps with that described using FIGS. 1 to 10 and is therefore simplified.

[0182] In one embodiment, a computing device can convert a conversation into a prompt using a prompt generation model. The generated prompt may include a description reflecting the context of the conversation. The generated prompt may include a description of an object's behavior corresponding to the context of the conversation. The generated prompt may include at least one of a description of an image effect, a description of inserted text, and a description of an object's selection.

[0183] In step S1150, the computing device can generate an image list using an image generation model.

[0184] In one embodiment, when the prompt generation model is activated, the computing device can convert the generated prompt into at least one image using the image generation model. The computing device can then generate a list of images based on the converted at least one image. The image generation model can include an image generation model (300) that generates an image from the prompt, as described with reference to FIG. 1.

[0185] In one embodiment, when the prompt generation model is not activated, the computing device may convert the input user dialogue into at least one image using the image generation model. The computing device may then generate an image list based on the converted at least one image. The image generation model of step S1150 may include the dialogue-image generation model (800) for generating images from dialogue, as described with reference to FIG. 8.

[0186] In step S1160, the computing device can display the generated image list. The computing device can display the image list through a user display.

[0187] In step S1170, the computing device may obtain a user input for selecting one image from the image list.

[0188] In step S1180, the computing device can transmit text and a selected image to the user terminal of the conversation partner based on the acquired user input. Furthermore, the computing device can display a situation in which the text and the selected image have been transmitted to the user terminal of the conversation partner based on the acquired user input.

[0189] If it is determined in step S1120 to transmit the conversation with an image, step S1190 may be performed. In step S1190, the computing device may transmit the input conversation to the user terminal of the conversation partner.

[0190] Hereinafter, with reference to FIG. 12, a configuration of a computing device for performing the image generation operations described so far will be described. FIG. 12 is a diagram illustrating the configuration of a computing device for performing image generation using a prompt generation model according to an embodiment of the present disclosure.

[0191] Referring to FIG. 12, a computing device (1000) according to an embodiment may include an input / output interface (1100), a memory (1200), and a processor (1300). However, the components of the computing device (1000) are not limited to the above-described examples, and the computing device (1000) may include more or fewer components than the above-described components. In one embodiment, some or all of the input / output interface (1100), the memory (1200), and the processor (1300) may be implemented in the form of a single chip, and the processor (1300) may include one or more processors.

[0192] The input / output interface (1100) may include an input interface (e.g., touch screen, hard button, microphone, etc.) for receiving control commands or information from a user, and an output interface (e.g., display panel, speaker, etc.) for displaying the results of execution of an operation according to the user's control or the status of the computing device (1000).

[0193] The memory (1200) is a configuration for storing various programs or data, and may be configured as a storage medium such as a ROM, a RAM, a hard disk, a CD-ROM, and a DVD, or a combination of storage media. The memory (1200) may not exist separately and may be configured to be included in the processor (1300). The memory (1200) may be configured as a volatile memory, a non-volatile memory, or a combination of volatile memory and non-volatile memory. Programs or instructions for performing operations according to the embodiments described with reference to FIGS. 1 to 11 may be stored in the memory (1200). The memory (1200) may also provide stored data to the processor (1300) upon request of the processor (1300).

[0194] The processor (1300) controls a series of processes so that the computing device (1000) operates according to the embodiments described with reference to FIGS. 1 to 11, and may be configured with one or more processors. In this case, the one or more processors may be a general-purpose processor such as a CPU, an AP, a DSP (Digital Signal Processor), a graphics-only processor such as a GPU, a VPU (Vision Processing Unit), or an artificial intelligence-only processor such as an NPU. For example, if one or more processors are artificial intelligence-only processors, the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.

[0195] The processor (1300) can record data in the memory (1200) or read data stored in the memory (1200), and in particular, process data according to predefined operation rules or artificial intelligence models by executing programs or commands stored in the memory (1200). Accordingly, the processor (1300) can perform the operations described in the embodiments described above, and the operations described as being performed by the computing device (1000) in the embodiments described above can be regarded as being performed by the processor (1300) unless otherwise specified.

[0196] A method according to one embodiment may include receiving a dialogue. The method may include inputting the dialogue into a prompt generation model to obtain a prompt corresponding to the dialogue. The method may include inputting the prompt into an image generation model to obtain an image. The prompt may include a description reflecting the context of the dialogue. The prompt generation model may be a language model trained through multiple dialogue-prompt pairs.

[0197] In one embodiment, the prompt may include a description of an object's behavior that corresponds to the context of the conversation.

[0198] In one embodiment, the prompt may further include at least one of a description of an image effect, a description of text to be inserted, and a description of selection of an object.

[0199] In one embodiment, the step of obtaining a prompt may include the step of combining at least one dialogue-prompt example pair and a dialogue. The step of obtaining a prompt may include the step of inputting the combination of at least one dialogue-prompt example pair and the dialogue into a prompt generation model to obtain a prompt corresponding to the dialogue. The prompt generation model may output a prompt corresponding to the dialogue based on a relationship between the dialogue example and the prompt example included in the at least one dialogue-prompt example pair.

[0200] In one embodiment, the step of combining at least one dialogue-prompt example pair and the dialogue may be listing at least one dialogue-prompt example pair and the dialogue.

[0201] In one embodiment, the method may further comprise the step of obtaining a plurality of dialogue-prompt pairs and at least one dialogue-prompt example pair. The method may further comprise the step of training a prompt generation model by comparing the prompts with the at least one dialogue-prompt example pair.

[0202] In one embodiment, the step of obtaining an image may include the step of determining a method for generating the image. The step of obtaining an image may include the step of obtaining an image corresponding to the prompt based on the determined method.

[0203] In one embodiment, the step of obtaining a prompt may include obtaining personal information about the subject of the conversation. The step of obtaining a prompt may include obtaining a prompt corresponding to the conversation based on the personal information.

[0204] In one embodiment, the personal information may include at least one of information about the personal identity of the subject of the conversation, information about the location where the subject of the conversation is located, and information about the weather at the location.

[0205] In one embodiment, the acquired image may include a plurality of images depicting a conversation. The method may further include obtaining a user input for selecting a first image from the plurality of images. The method may further include transmitting at least one of the first image and the conversation to an external user terminal based on the user input.

[0206] A non-transitory computer-readable recording medium having recorded thereon a program for performing any one of the methods according to one embodiment of the present disclosure on a computer may be provided.

[0207] A computing device according to one embodiment may include an input / output interface, a memory, and at least one processor. The input / output interface may receive a user input requesting image processing. The input / output interface may output a processed image according to the user input. The memory may store commands for processing the image. At least one processor may execute the commands. At least one processor may receive a dialogue. At least one processor may input the dialogue into a prompt generation model to obtain a prompt corresponding to the dialogue. At least one processor may input the prompt into an image generation model to obtain an image. The prompt may include a description reflecting the context of the dialogue. The prompt generation model may be a language model trained through a plurality of dialogue-prompt pairs.

[0208] In one embodiment, the prompt may include a description of an object's behavior that corresponds to the context of the conversation.

[0209] In one embodiment, the prompt may further include at least one of a description of an image effect, a description of text to be inserted, and a description of selection of an object.

[0210] In one embodiment, at least one processor can obtain a prompt by combining a plurality of dialogue-prompt pairs and at least one other dialogue-prompt example pair and a dialogue, and inputting the combination of at least one dialogue-prompt example pair and the dialogue into a prompt generation model, thereby obtaining a prompt corresponding to the combination of at least one dialogue-prompt example pair and the dialogue.

[0211] In one embodiment, at least one processor may list at least one dialog-prompt example pair and a dialog, in combination with at least one dialog-prompt example pair.

[0212] In one embodiment, at least one processor can train a prompt generation model by obtaining a plurality of dialogue-prompt pairs and at least one other dialogue-prompt example pair, and comparing the prompts with the at least one dialogue-prompt example pair.

[0213] In one embodiment, at least one processor can determine a method for generating an image when acquiring an image, and acquire an image corresponding to a prompt based on the determined method.

[0214] In one embodiment, at least one processor may obtain personal information about a subject of a conversation when obtaining a prompt, and obtain a prompt corresponding to the conversation based on the personal information.

[0215] In one embodiment, the acquired image may include a plurality of images depicting a conversation. At least one processor may obtain a user input for selecting a first image from the plurality of images, and, based on the user input, transmit at least one of the first image and the conversation to an external user terminal.

[0216] Various embodiments of the present disclosure may be implemented or supported by one or more computer programs, and the computer programs may be formed from computer-readable program code and embodied in a computer-readable medium. In the present disclosure, "application" and "program" may refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, associated data, or portions thereof suitable for implementation in computer-readable program code. "Computer-readable program code" may include various types of computer code, including source code, object code, and executable code. "Computer-readable medium" may include various types of media that can be accessed by a computer, such as read-only memory (ROM), random access memory (RAM), a hard disk drive (HDD), a compact disc (CD), a digital video disc (DVD), or various types of memory.

[0217] Additionally, a device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, a 'non-transitory storage medium' is a tangible device and may exclude wired, wireless, optical, or other communication links that transmit temporary electrical or other signals. Meanwhile, this 'non-transitory storage medium' does not distinguish between cases where data is permanently stored in the storage medium and cases where it is temporarily stored. For example, a 'non-transitory storage medium' may include a buffer where data is temporarily stored. A computer-readable medium may be any available medium that can be accessed by a computer, and may include both volatile and non-volatile media, and removable and non-removable media. A computer-readable medium includes a medium on which data can be permanently stored and a medium on which data can be stored and later overwritten, such as a rewritable optical disk or an erasable memory device.

[0218] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0219] The above description of the present disclosure is for illustrative purposes only, and those skilled in the art will appreciate that the present disclosure can be readily modified into other specific forms without altering the technical spirit or essential characteristics of the present disclosure. For example, suitable results can be achieved even if the described techniques are performed in a different order than the described method, and / or components of the systems, structures, devices, circuits, etc. described are combined or combined in a different form than the described method, or are replaced or substituted by other components or equivalents. Therefore, it should be understood that the embodiments described above are illustrative in all respects and not restrictive. For example, each component described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined form.

[0220] The scope of the present disclosure is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present disclosure.

Claims

1. Step of receiving a dialogue (10); A step of inputting the above conversation into a prompt generation model (200) to obtain a prompt (20) corresponding to the above conversation (10); and It includes a step of inputting the above prompt (20) into the image generation model (300) to obtain an image (30). The above prompt includes a description that reflects the context of the conversation, A method wherein the above prompt generation model is a language model learned through multiple dialogue-prompt pairs.

2. In paragraph 1, A method wherein the above prompt includes a description of an object's behavior corresponding to the context of the above conversation.

3. In paragraph 2, A method wherein the above prompt further includes at least one of a description of an image effect, a description of text to be inserted, and a description of selection of the object.

4. In any one of the clauses 1 to 3, The steps to obtain the above prompt are: at least one dialogue-prompt example pair and a step of combining said dialogue; and A step of inputting at least one of the above dialogue-prompt example pairs and the combination (610) of the above dialogues into the prompt generation model (200) to obtain the prompt (620) corresponding to the above dialogues, The above prompt generation model (200) is a method for outputting the prompt (620) corresponding to the conversation based on the relationship between the conversation example and the prompt example included in the at least one conversation-prompt example pair.

5. In any one of paragraphs 1 to 3, A step of obtaining at least one dialogue-prompt example pair different from the above plurality of dialogue-prompt pairs; A method further comprising the step of training the prompt generation model (200) by comparing the prompt (720) with at least one example dialogue-prompt pair (715).

6. In any one of paragraphs 1 to 5, The steps for obtaining the above image are: a step of determining a method for generating said image; and A method comprising the step of obtaining the image corresponding to the prompt based on the determined method.

7. In any one of paragraphs 1 to 6, The steps to obtain the above prompt are: A step of obtaining personal information about the subject of the above conversation; A method comprising the step of obtaining a prompt corresponding to the conversation based on the personal information.

8. In paragraph 7, A method wherein the personal information includes at least one of information about the personal details of the subject of the conversation, information about the location where the subject of the conversation is located, and information about the weather at the location.

9. In any one of paragraphs 1 to 8, The above acquired image includes multiple images depicting the above conversation, A step of obtaining a user input for selecting a first image from among the plurality of images; and A method further comprising the step of transmitting at least one of the first image and the conversation to an external user terminal based on the user input.

10. A non-transitory computer-readable recording medium having recorded thereon a program for performing any one of the methods of clauses 1 to 9 on a computer.

11. An input / output interface for receiving user input requesting image processing and outputting an image processed according to the user input; Memory where commands for processing images are stored; and Contains at least one processor, At least one processor executes the instructions, Receive conversations, By inputting the above conversation into the prompt generation model, a prompt corresponding to the above conversation is obtained, By inputting the above prompt into the image generation model, an image is obtained, The above prompt includes a description that reflects the context of the conversation, A computing device, wherein the above prompt generation model is a language model learned through multiple dialogue-prompt pairs.

12. In paragraph 11, A computing device, wherein the prompt includes a description of an object's behavior corresponding to the context of the conversation.

13. In any one of paragraphs 11 to 12, At least one processor, in obtaining the prompt, At least one dialogue-prompt example pair and a combination of said dialogue, By inputting at least one of the above dialogue-prompt example pairs and the combination of the above dialogues into the prompt generation model, the prompt corresponding to the above dialogue is obtained, A computing device, wherein the prompt generation model outputs the prompt corresponding to the dialogue based on a relationship between a dialogue example and a prompt example included in the at least one dialogue-prompt example pair.

14. In any one of paragraphs 11 to 13, At least one processor, in obtaining said prompt, Obtain personal information about the subject of the above conversation, A computing device that obtains a prompt corresponding to the conversation based on the personal information.

15. In any one of paragraphs 11 to 14, The above acquired image includes multiple images depicting the above conversation, At least one processor of the above, Obtaining a user input for selecting a first image from among the above multiple images, A computing device that transmits at least one of the first image and the conversation to an external user terminal based on the user input.

Citation Information

Patent Citations

  • Image generation method and device, electronic equipment and storage medium

    CN116843795A

  • Accident information providing system of autonomous vehicle, and autonomous vehicle apparatus thereof

    KR1020210005417A

  • A torque vectoring control method for vehicles

    KR1020230127149A

  • Air quality monitoring system using vehicle and air quality sensing apparatus

    KR102175460B1