Method for generating images and electronic device supporting same

WO2026168968A1PCT designated stage Publication Date: 2026-08-13SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-02-04
Publication Date
2026-08-13

Smart Images

  • Figure KR2026002074_13082026_PF_FP_ABST
    Figure KR2026002074_13082026_PF_FP_ABST
Patent Text Reader

Abstract

A method for obtaining a plurality of images, according to the present disclosure, may comprise the operations of: obtaining, on the basis of a user input, text for generating an image; generating a first prompt for generating the image on the basis of the text; obtaining a second prompt including description information for generating a second image including a plurality of first images distinguishable from each other by applying the first prompt to a first artificial intelligence model; obtaining the second image including the plurality of first images by applying the second prompt to a second artificial intelligence model; and displaying the plurality of first images.
Need to check novelty before this filing date? Find Prior Art

Description

Method for generating images and electronic devices supporting this

[0001] The present disclosure relates to a method for generating an image and an electronic device supporting the same. More specifically, the present disclosure relates to a method for generating an image using an artificial intelligence (AI) model and an electronic device supporting the same.

[0002] The development of information and communication technology has transformed the ways in which users form social relationships and communicate with one another in online environments. For example, people communicate with others using messaging services or social network services (SNS), and these digital means of communication are actively used not only for everyday conversations but also in various fields such as work, education, and leisure.

[0003] Recently, artificial intelligence systems capable of achieving human-level intelligence are being utilized in various fields. Unlike conventional rule-based smart systems, artificial intelligence systems are systems in which machines learn, make judgments, and become smarter on their own. As artificial intelligence systems improve in recognition accuracy and gain a more accurate understanding of user preferences with continued use, existing rule-based smart systems are gradually being replaced by deep learning-based artificial intelligence systems.

[0004] Image generation technology using generative AI models generates new images based on user input, primarily utilizing deep learning-based neural network models. These models are trained by learning from large amounts of image data to generate specific styles, compositions, and objects; generally, they can receive inputs such as text prompts, existing images, or sketches to generate corresponding images. Representative generative AI models include Diffusion Models, Generative Adversarial Networks (GANs), and Transformer-based models, which enable the generation of high-resolution images, realistic synthetic images, and images with artistic styles. Unlike conventional image generation methods, generative AI models offer high utility as they can produce diverse results based on user input without relying on standardized templates. However, generating high-quality images requires massive computational resources, which can lead to issues regarding computational speed, generation costs, and network load. Consequently, there is a demand for technologies that can generate images more efficiently and optimize computational resources.

[0005] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art in relation to the present disclosure.

[0006] A method for acquiring a plurality of images according to the present disclosure may include: acquiring text for generating an image based on user input; generating a first prompt for generating the image based on the text; acquiring a second prompt including description information for generating a second image including a plurality of first images that are distinguishable from one another by applying the first prompt to a first artificial intelligence model; acquiring the second image including the plurality of first images by applying the second prompt to a second artificial intelligence model; and displaying the plurality of first images.

[0007] An electronic device according to the present disclosure may include a memory for storing instructions; and at least one processor. The instructions may be executed individually or collectively by the at least one processor so that the electronic device, based on user input, obtains text for generating an image, generates a first prompt for generating the image based on the text, obtains a second prompt including descriptive information for generating a second image including a plurality of first images that are distinguishable from one another by applying the first prompt to a first artificial intelligence model, obtains the second image including the plurality of first images by applying the second prompt to a second artificial intelligence model, and displays the plurality of first images.

[0008] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.

[0009] FIG. 1 is a diagram illustrating an overview of a method in which an electronic device according to the present disclosure provides a plurality of first images.

[0010] FIG. 2 is a flowchart illustrating the process of generating and outputting a plurality of images in one embodiment.

[0011] FIG. 3 is a diagram showing an example of user input for generating a plurality of images according to one embodiment.

[0012] FIG. 4 is a flowchart illustrating the process of obtaining a first prompt based on user input for generating a plurality of images according to one embodiment.

[0013] FIG. 5 is a block diagram illustrating the input and output data of a first artificial intelligence model and a second artificial intelligence model in one embodiment.

[0014] FIG. 6 is a drawing showing an example of a second image generated according to one embodiment.

[0015] FIG. 7 is a diagram showing an example in which a second image generated according to one embodiment is displayed.

[0016] FIG. 8 is a flowchart illustrating the process of separating and displaying a plurality of first images within a second image generated in one embodiment.

[0017] FIG. 9 is a flowchart illustrating the process of a first image displayed in an area where input is received from a user being separated from a second image and displayed, in one embodiment.

[0018] FIG. 10 is an example drawing for explaining a method in which, in one embodiment, a second image is divided by region, and a first image displayed in the region where user input is received within the second image is separated from the second image.

[0019] FIG. 11 is a diagram showing an example in which, in one embodiment, a plurality of first images are separated from a second image and each first image is displayed.

[0020] FIG. 12 is a flowchart illustrating the process of acquiring and displaying a plurality of third images by increasing the resolution of a plurality of first images in one embodiment.

[0021] FIG. 13 is a block diagram showing an electronic device according to one embodiment and artificial intelligence models stored in a server.

[0022] FIG. 14 is a block diagram illustrating data transmitted and received between an electronic device, a first artificial intelligence model, and a second artificial intelligence model according to one embodiment.

[0023] FIG. 15 is a block diagram of an electronic device in a network environment according to various embodiments.

[0024] FIG. 16 is a diagram showing an example of a system including a generative artificial intelligence model according to one embodiment.

[0025] Embodiments of the present disclosure are described below in detail with reference to the attached drawings so that those skilled in the art can easily implement them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly explain the present disclosure in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification are denoted by similar reference numerals.

[0026] The terms used in this disclosure are described in their current, general form considering the functions mentioned herein; however, they may refer to various other terms depending on the intent of those skilled in the art, case law, or the emergence of new technologies. Accordingly, the terms used in this disclosure should not be interpreted solely by their names, but should be interpreted based on the meaning of the terms and the overall content of this disclosure.

[0027] Additionally, terms such as the first, second, third, ..., Nth may be used to describe various components, but the components should not be limited by these terms. These terms are used for the purpose of distinguishing one component from another.

[0028] Throughout the specification, when a part is described as being "connected" to another part, this includes not only cases where they are "directly connected," but also cases where they are "electrically connected" with other components interposed between them. Furthermore, when a part is described as "including" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.

[0029] Phrases such as "in one embodiment" appearing in various places in this disclosure do not necessarily refer to the same embodiment.

[0030] One embodiment of the present disclosure may be represented by functional block configurations and various processing steps. Some or all of these functional blocks may be implemented by various numbers of hardware and / or software configurations that execute specific functions. For example, the functional blocks of the present disclosure may be implemented by one or more microprocessors or by circuit configurations for a specific function. Additionally, for example, the functional blocks of the present disclosure may be implemented in various programming or scripting languages. The functional blocks may be implemented as algorithms executed on one or more processors. Furthermore, the present disclosure may employ prior art for electronic configuration, signal processing, and / or data processing. Terms such as "mechanism," "element," "means," and "configuration" may be used broadly and are not limited to mechanical and physical configurations.

[0031] Furthermore, the connecting lines or connecting members between the components depicted in the drawings are merely illustrative of functional connections and / or physical or circuit connections. In the actual device, connections between components may be represented by various alternative or added functional connections, physical connections, or circuit connections.

[0032] Artificial intelligence technology consists of machine learning (e.g., deep learning) and elemental technologies utilizing machine learning.

[0033] Machine learning is an algorithmic technology that classifies and learns the features of input data on its own, and the elemental technology is a technology that mimics functions such as cognition and judgment of the human brain by utilizing deep learning machine learning algorithms, and consists of the fields of linguistic understanding, visual understanding, reasoning / prediction, knowledge representation, and motion control.

[0034] The various fields where artificial intelligence technology is applied are as follows. Linguistic understanding is a technology that recognizes, applies, and processes human language and text, and includes natural language processing, machine translation, dialogue systems, question answering, and speech recognition / synthesis. Visual understanding is a technology that perceives and processes objects like human vision, and includes object recognition, object tracking, image search, person recognition, scene understanding, spatial understanding, and image enhancement. Inference and prediction is a technology that judges information to logically infer and predict, and includes knowledge / probability-based inference, optimization prediction, preference-based planning, and recommendation. Knowledge representation is a technology that automatically processes human experiential information into knowledge data, and includes knowledge construction (data generation / classification) and knowledge management (data utilization). Motion control is a technology that controls the autonomous driving of vehicles and the movement of robots, and includes motion control (navigation, collision, driving) and manipulation control (behavior control).

[0035] In the present disclosure, the first prompt may be a prompt applied to a first artificial intelligence model, comprising information for generating an image. The first prompt may be a prompt generated based on text obtained as text for generating an image is obtained by user input.

[0036] In the present disclosure, the first artificial intelligence model may be a multimodal model that learns and processes relationships between data of various types or various modalities, such as text and images. The first artificial intelligence model may be a large multimodal model (LMM) trained using text data and image data. The first artificial intelligence model may generate an image or generate text associated with an image based on a text prompt, an image prompt, or a prompt composed of text and an image input to the first artificial intelligence model. The first artificial intelligence model is not limited to an LMM, and the first artificial intelligence model may be a generative artificial intelligence model trained to generate descriptive information of an image or a prompt containing descriptive information based on a user's text input. However, it is not limited thereto.

[0037] In the present disclosure, the second prompt may be data output by the first artificial intelligence model as the first prompt is applied to the first artificial intelligence model. The second prompt may be a prompt applied to the second artificial intelligence model. The second prompt may include description information for generating a second image comprising a plurality of first images that are distinguishable from one another.

[0038] In the present disclosure, description information may refer to text information describing an image to be generated based on input received from a user. Description information may be generated by a first artificial intelligence model as a first prompt, comprising text input from a user, is applied to a first artificial intelligence model. A second prompt, comprising description information, may be generated by the first artificial intelligence model, and the second prompt may be applied to a second artificial intelligence model. The second artificial intelligence model may generate an image based on the second prompt and the description information included in the second prompt.

[0039] In the present disclosure, the second artificial intelligence model may be an artificial intelligence model trained to generate an image based on an input second prompt. The second artificial intelligence model may be a text-to-image model and may be a model in the form of a combination of a transformer-based language model and a diffusion-based image generation model, but is not limited thereto.

[0040] The first artificial intelligence model may be any artificial intelligence model capable of performing bidirectional interaction between text and image, and the second artificial intelligence model may be capable of performing unidirectional interaction that generates an image based on text, but is not limited thereto.

[0041] In the present disclosure, the second image may be a single image generated by a second artificial intelligence model, and the first image may refer to each of a plurality of distinct images included in the second image. The first images included in the second image may be images that are independently distinguishable from one another.

[0042] The images generated by the second artificial intelligence model (e.g., the first image, the second image) may be static images, but are not limited thereto, and may be dynamic images made up of multiple images, e.g., videos or animated images. The images may be graphic elements, such as stickers, for use in messaging services or social media.

[0043] In the present disclosure, the style of an image may refer to a visual representation method or technique of an image generated by an artificial intelligence model. The style may be identified by setting values ​​that determine the appearance of the image by reflecting specific aesthetic characteristics or representation techniques. The style may include elements such as artistic techniques, textures, colors, shapes, and lighting effects. For example, the style may include a doodle style that gives a hand-drawn feel, a 3D emoji style utilizing realistic 3D rendering, and an illustration style that emphasizes illustrative expressions. However, examples of styles are not limited thereto. A user may request the generation of an image that reflects a desired mood or representation method by specifying a particular style.

[0044] In the present disclosure, the visual characteristics of an image are a set of visual elements possessed by the generated image and may include characteristics related to elements such as style, color tone, composition, objects included within the image, types of objects, arrangement and form of objects, texture, lighting, or specific actions performed by objects. The visual characteristics may vary depending on user input. For example, specific attributes including a style with specific artistic techniques applied, warm or cool color tones, a composition including specific people or objects, and the movements of animals or facial expressions of people may be included in the visual characteristics. A generative AI model can generate an image that meets the user's request by reflecting these visual characteristics in the image.

[0045] The present disclosure will be described in detail below with reference to the attached drawings.

[0046] FIG. 1 is a drawing for illustrating an overview of a method in which an electronic device (100) according to the present disclosure provides a plurality of first images (10).

[0047] Referring to FIG. 1, an artificial intelligence model (e.g., a second artificial intelligence model (102)) may be utilized to generate an image according to one embodiment. A user may explain to the artificial intelligence model the image they wish to create, and the artificial intelligence model may generate the image by understanding and analyzing the user's explanation. The time required for a generative artificial intelligence model to generate an image may be longer than the time required to generate text. This is because, even if the capacity of the text and the image to be generated are the same, generating an image requires higher computational complexity and computational power. Furthermore, in a service that uses a generative artificial intelligence model to generate images, fees may be charged based on the number of images to be generated; therefore, as the number of images to be generated increases, the cost required to generate the images increases. Accordingly, when multiple images are generated by a generative artificial intelligence model, the time required and the cost required increase compared to when only one image is generated.

[0048] An electronic device (100) according to the present disclosure may receive an input from a user requesting the generation of a plurality of images. For example, the user's input may include an input specifying a plurality of styles for the images to be generated. Based on the user's input, the electronic device (100) may obtain a first prompt (110) to be applied to a first artificial intelligence model (101). By applying the first prompt (110) to the first artificial intelligence model (101), the electronic device (100) may obtain a second prompt (120) that includes description information for generating a second image (102) comprising a plurality of first images that are distinguishable from one another. The description information may include a plurality of description texts (121, 122, 123, 124) describing the plurality of first images (11, 12, 13, 14) to be generated. The electronic device (100) can obtain a second image (20) including a plurality of first images (11, 12, 13, 14) by applying a second prompt (120) to a second artificial intelligence model (102). Accordingly, the electronic device (100) can display a plurality of first images (11, 12, 13, 14).

[0049] The present disclosure may provide a method for reducing the time and cost required to generate multiple images by reducing the number of prompts applied to the second artificial intelligence model (102) from multiple to one when user input for acquiring multiple images is received.

[0050] FIG. 2 is a flowchart illustrating the process of generating and outputting a plurality of images (10) in one embodiment.

[0051] Referring to identification number 210, an electronic device (100) according to one embodiment can obtain text for generating an image based on user input.

[0052] An electronic device (100) according to one embodiment may receive input from a user for generating an image. The electronic device (100) may receive input from a user for generating a plurality of images. The input for generating a plurality of images may be included in the input for generating an image, or it may be an input separate from the input for generating an image. A detailed description regarding the operation of receiving user input for generating a plurality of images will be described later in the description section for FIG. 3.

[0053] Referring to identification number 220, an electronic device (100) according to one embodiment can generate a first prompt (110) for generating an image.

[0054] According to one embodiment, the first prompt (110) may be a prompt to be input into the first artificial intelligence model (101). In one embodiment, the first prompt (110) may include text obtained based on user input. For example, the first prompt (110) may be the text itself received from the user. For example, the first prompt (110) may include at least one of text received from the user, text containing a query to generate a plurality of images, or information regarding the number of images to be generated. For example, if the text obtained from the user includes text to generate a plurality of images and information regarding the number of images to be generated, the first prompt (110) may be the text itself received from the user. In one embodiment, the electronic device (100) may generate the first prompt (110) by reconstructing text for generating images. For example, if an input is obtained from a user to generate multiple images through an input different from the text input, separate from the text obtained from the user, the electronic device (100) can generate a first prompt (110) by adding at least one of text to generate multiple images or text regarding the number of images to be generated to the text obtained from the user. In one embodiment, the first prompt (110) can be generated by inputting the user's input and certain information into a certain template. An explanation related to this will be provided later in FIG. 4.

[0055] Referring to identification number 230, an electronic device (100) according to one embodiment can obtain a second prompt (120) including description information for generating a second image (20) comprising a plurality of first images (10) that are distinguishable from one another by applying a first prompt (110) to a first artificial intelligence model (101).

[0056] In one embodiment, the first artificial intelligence model (101) may be an artificial intelligence model trained to receive a first prompt (110) and output a second prompt (120) containing description information for generating a second image (20) that includes a plurality of first images (10) that are distinguishable from one another.

[0057] In one embodiment, a plurality of first images (10) within a second image (20) may be arranged so as to be distinguishable from one another. A plurality of first images (10) may be independently distinguished from one another within a single second image (20). For example, the area within the second image (20) may be divided by the number of a plurality of first images (10), and each first image (10) may be placed in a respective divided area within the second image (20). For example, a plurality of first images (10) may be images that reflect the user's intention for image generation. For example, a plurality of first images (10) may be images that represent similar themes and have different styles. For example, a plurality of first images (10) may have at least one identical visual characteristic. Also, for example, a plurality of first images (10) may have at least one different visual characteristic. Accordingly, each first image (10) may have characteristics corresponding to the input received from the user, and may have detailed characteristics different from different first images (10). However, the characteristics of the second image (20) and the first image (10) are not limited thereto. If the input received from the user includes information requesting the output of images having multiple different characteristics, each first image (10) may have different characteristics depending on the input received from the user. The second prompt (120) may be data output from the first artificial intelligence model (101) as the first prompt (110) is applied to the first artificial intelligence model (101). In one embodiment, the first artificial intelligence model (101) may output data having the format of a prompt to be applied to the second artificial intelligence model (102). In one embodiment, the second prompt (120) may include data output from the first artificial intelligence model (101).For example, the second prompt (120) may be the data itself output from the first artificial intelligence model (101). For example, the second prompt (120) may be generated by formally modifying the data output from the first artificial intelligence model (101) so that it can be applied to the second artificial intelligence model (102).

[0058] In one embodiment, the second prompt (120) may include a plurality of descriptive texts describing a plurality of first images to be generated. The second prompt (120) may be a single prompt including a plurality of descriptive texts (121, 122, 123, 124) describing a plurality of first images (10) to be generated. The second prompt (120) may include content for generating one second image (20). The second prompt (120) may include information that the number of images to be generated is one. The second prompt (120) may include information for outputting one image by being input to the second artificial intelligence model (102). For example, the second prompt (120) may include the text 'samplecount: 1'.

[0059] In one embodiment, the description information may include a plurality of description texts (121, 122, 123, 124) describing a plurality of first images (10) to be generated. In this case, the plurality of first images (11, 12, 13, 14) to be generated may correspond to each of the plurality of description texts (121, 122, 123, 124).

[0060] The description of the second prompt (120), the description information, and the description text will be described later in the description section for Fig. 5.

[0061] Referring to identification number 240, an electronic device (100) according to one embodiment can obtain a second image (20) including a plurality of first images (10) by applying a second prompt (120) to a second artificial intelligence model (102).

[0062] According to one embodiment, the electronic device (100) may apply a second prompt (120) to a second artificial intelligence model (102). In one embodiment, the second artificial intelligence model (102) may be an artificial intelligence model trained to generate an image based on an input prompt. The second artificial intelligence model (102) may be an artificial intelligence model trained to generate an image having characteristics included in an input prompt.

[0063] In one embodiment, the second artificial intelligence model (102) can receive the second prompt (120) and output one second image (20). For example, the second prompt (120) includes information that causes one image to be output by being input to the second artificial intelligence model (102), and the second artificial intelligence model (102) that receives the second prompt (120) can output one second image (20).

[0064] An electronic device (100) according to one embodiment can separate a plurality of first images (11, 12, 13, 14) within a second image (20) based on the acquisition of a second image (20).

[0065] For example, the electronic device (100) can obtain a plurality of first images (11, 12, 13, 14) separated from the second image (20) by applying the second image (20) to the third artificial intelligence model (103).

[0066] Referring to identification number 250, an electronic device (100) according to one embodiment can display first images (10).

[0067] In one embodiment, the electronic device (100) may display a second image (20) comprising a plurality of first images (11, 12, 13, 14). For example, the electronic device (100) may receive user input on an area where the second image (20) is displayed through the display. In this case, the electronic device (100) may distinguish areas where each first image (10) is displayed on the second image (20) and identify the first image (10) displayed in the area where the user input is received among the distinguished areas. Accordingly, the electronic device (100) may display the first image (10) displayed in the area where the user input is received. For example, the electronic device (100) may separate the first image (10) displayed in the area where the user input is received from the second image (20) and display the separated first image (10).

[0068] According to one embodiment, the electronic device (100) can separate a plurality of first images (11, 12, 13, 14) within a second image (20) and display the separated plurality of first images (11, 12, 13, 14). The details regarding the separation of a plurality of first images (10) within the second image (20) will be explained in detail in the description of FIG. 8.

[0069] FIG. 3 is a diagram showing an example of user input for generating a plurality of images according to one embodiment.

[0070] According to one embodiment, an electronic device (100) may receive input from a user for generating a plurality of images. In one embodiment, the input for generating a plurality of images may include input in text format. The electronic device (100) may receive input from a user regarding features of a first image (10) to be generated. The electronic device (100) may receive input from a user regarding visual characteristics of a first image (10) to be generated. For example, the electronic device (100) may receive input from a user that includes information describing the first image (10) to be generated. For example, the electronic device (100) may receive input regarding characteristics of an object to be included in the first image (10) to be generated. The input for generating a plurality of images may include an input for selecting a UI displayed on a display. For example, the input for generating a plurality of images may include a touch input for a UI displayed on a display.

[0071] An electronic device (100) according to one embodiment receives input text (105) for generating an image from a user, and separately from the received input text (105), may receive input from the user for generating a plurality of images. Referring to the example of FIG. 3, the electronic device (100) may receive input text (105) for generating an image from a user. In one embodiment, the input text (105) may include text indicating the characteristics of the image to be generated. The input text (105) may include text indicating the type or characteristics of an object to be included in the image to be generated. The input text (105) may include text indicating the visual characteristics of an object to be included in the image to be generated. An electronic device (100) according to one embodiment may receive input from the user for generating a plurality of images separately from the input text (105). Referring to the example of FIG. 3, the electronic device (100) may receive 'dragon chef' as the input text (105). The electronic device (100) can, together with this, receive user input to generate multiple images through a UI (305) displayed on a display. For example, the electronic device (100) can display a UI (305) to generate multiple images through a display and can receive input from a user regarding the UI (305) to generate multiple images. For example, the electronic device (100) can receive input specifying the style of the image to be generated according to input text (105). In this case, the electronic device (100) can receive input from a user specifying multiple styles of the image to be generated. For example, the electronic device (100) can receive input to generate the image to be generated according to a 'doodle' style, an 'illustration' style, and a '3D emoji' style.In this case, the electronic device (100) may determine that an input has been received to generate multiple images according to multiple different styles. For example, the electronic device (100) may receive an input specifying the style of the image to be generated according to the input text (105). In this case, the electronic device (100) may receive an input from the user specifying one of the styles of the multiple images. For example, the electronic device (100) may receive an input to generate the image to be generated according to the 'illustration' style. In this case, the electronic device (100) may determine that an input has been received to generate multiple images having different features according to the selected 'illustration' style.

[0072] In one embodiment, an electronic device (100) may receive input text (105) for generating an image from a user. The input text (105) for generating an image may include text input for generating a plurality of images. For example, the input text (105) for generating an image may include text input related to the number of first images (10) to be generated. The electronic device (100) that receives the input text (105) for generating an image may identify whether text input related to the number of first images (10) to be generated is included in the input text (105). If it is identified that text input related to the number of first images (10) to be generated is included in the received input text (105), the electronic device (100) may check whether the number of first images (10) to be generated is two or more. Upon verification, if the number of first images (10) to be generated is two or more, the electronic device (100) can determine that text input for generating multiple images has been received.

[0073] FIG. 4 is a flowchart illustrating the process of obtaining a first prompt (110) based on user input for generating a plurality of images according to one embodiment.

[0074] The description related to identification numbers 410 to 430 in FIG. 4 may correspond to the description related to identification numbers 210 to 220 in FIG. 2. In this regard, the process of obtaining the first prompt (110) will be explained in more detail below.

[0075] In identification number 410, the electronic device (100) can receive input from a user for generating multiple images. Details regarding the operation of receiving input from a user for generating multiple images have been explained in detail in identification number 210 of FIG. 2 and FIG. 3, so they will be omitted here.

[0076] In identification number 420, the electronic device (100) may input text (150) into a template, text indicating the number of first images (10) to be generated and the position where each of the first images (10) is to be placed within the second image (20). According to one embodiment, the electronic device (100) may input text into a previously stored template, text indicating the number of first images (10) to be generated and the position where each of the first images (10) is to be placed within the second image (20). For example, the template may be stored in advance in the electronic device (100).

[0077] Referring to [Table 1] below, which shows an example of a template, details regarding the information included in the template will be described later. [Table 1] below is merely an example for generating the first prompt (110), and the format of the first prompt (110) and the template for generating it is not limited to the format of [Table 1].

[0078] Your goal is to produce prompts given input text provided.The prompt must meet the following conditions:… Make NN promptsthe result should follow the result form belowCreate an ideal prompt to generate a sticker image for the input: “XX”==result form==at upper-left, prompt_1at upper-right, prompt_2at bottom-left, prompt_3at bottom-right, prompt_4...

[0079] In one embodiment, the template may include information that causes an output value to be generated under predetermined conditions. The template may include a portion where input text (150) received from a user is placed, a portion where information on the number of first images (10) to be generated is placed, and a portion where text indicating the position where each of the first images (10) is placed within the second image (20) is placed. The template may include a portion where information indicating the format of an output value to be output from the first artificial intelligence model (101) is placed.

[0080] In 'Make NN prompts' of [Table 1], information regarding the number of first images (10) to be generated may be placed in 'NN'. For example, if '4' is placed in the 'NN' section, information to generate 4 first images (10) according to the content of 'Make 4 prompts' may be included in the first prompt (110).

[0081] In 'Create an ideal prompt to generate a sticker image for the input : “XX”' of [Table 1], input text (150) received from the user may be placed in 'XX'. For example, if 'dragon chef' is placed in the 'XX' part, information that causes first images (10) containing characteristics corresponding to 'dragon chef' to be generated according to the content of 'Create an ideal prompt to generate a sticker image for the input : “dragon chef”' may be included in the first prompt (110).

[0082] Information indicating the format of the output value to be output from the first artificial intelligence model (101) may be placed on or after '==result form==' in [Table 1]. For example, the template may include information that causes descriptive texts to be generated equal to the number of first images (10) to be generated. That is, the template may include information that causes multiple texts to be included in the second prompt (120) to generate each first image (10). For example, if the number of first images (10) to be generated is 4, the first prompt (110) may include information that causes 4 descriptive texts to be included in the second prompt (120) to generate each first image (10). Text indicating the position where each of the first images (10) is to be placed within the second image (20) may be placed in the template. Text indicating the position where descriptive texts describing each of the first images (10) to be generated are to be placed in the template may be placed. For example, '==result form==' of the first prompt (110) or thereafter, 'at upper-left, prompt_1', 'at upper-right, prompt_2', 'at bottom-left, prompt_3', and 'at bottom-right, prompt_4' may be arranged in order. In this case, the first prompt (110) may include 'prompt_1', a portion where descriptive text for the first image (11) placed at the top-left of the second image (20) is to be placed; 'prompt_2', a portion where descriptive text for the first image (12) placed at the top-right of the second image (20) is to be placed; 'prompt_3', a portion where descriptive text for the first image (13) placed at the bottom-left of the second image (20) is to be placed; and 'prompt_4', a portion where descriptive text for the first image (14) placed at the bottom-right of the second image (20) is to be placed.

[0083] In identification number 430, the electronic device (100) can obtain a first prompt (110). The electronic device (100) can obtain the first prompt (110) by inputting text (150) into a template, text indicating the number of first images (10) to be generated and the position where each of the first images (10) is to be placed within the second image (20).

[0084] FIG. 5 is a block diagram illustrating the input and output data of a first artificial intelligence model (101) and a second artificial intelligence model (102) in one embodiment.

[0085] An electronic device (100) according to one embodiment can obtain a second prompt (120) output from a first artificial intelligence model (101) by applying a first prompt (110) to a first artificial intelligence model (101). The first artificial intelligence model (101) may be an artificial intelligence model trained to output a second prompt that includes descriptive information for generating a second image (20) comprising a plurality of first images (10) that are distinguishable from one another, based on a first prompt (110) for generating a plurality of images.

[0086] In one embodiment, the second prompt (120) may be output by the first artificial intelligence model (101) as the first prompt (110) is applied to the first artificial intelligence model (101). The second prompt (120) may be a prompt applied to the second artificial intelligence model (102) as it is output from the first artificial intelligence model (101).

[0087] Referring to [Table 2] below, which shows an example of the second prompt (120), details regarding the information included in the second prompt (120) will be described later. [Table 2] below is merely an example of the second prompt (120) for generating a second image (20) containing a plurality of first images (10) by inputting it into the second artificial intelligence model (102), and the content and format of the second prompt (120) are not limited to the content and format of [Table 2].

[0088] **Upper-Left:** A bold kiss-cut sticker featuring a whimsical dragon chef with a tiny chef's hat, holding a giant lollipop spoon. The dragon is vibrant teal, pink, and orange, with simple, bold outlines and cheerful, childish features. The background is pure white. / n / n**Upper-Right:** Generate a colorful, doodle-style sticker of a smiling dragon chef wearing a striped apron and oversized sunglasses. Include a playful, simple design with clear lines, using bright yellows, reds, and blues. The sticker should have a white background and a kiss-cut edge for easy peeling. / n / n**Bottom-Left:** Create a lovely, childish sticker of a dragon chef preparing a miniature cake. The dragon should be drawn with simple shapes and bold, black outlines, using a palette of pastel pinks, greens, and purples. The background is white, and the sticker has a bold kiss-cut. / n / n**Bottom-Right:** Design a bold, simple sticker featuring a dragon chef with flames for hair, happily stirring a pot of rainbow-colored soup. Use a limited color palette of primary colors with black outlines. The style should be childish and playful, with a white background and a kiss-cut edge. / n…"sampleCount": 1,….

[0089] According to one embodiment, the second prompt (120) may include description information for generating a second image (20) comprising a plurality of first images (10) that are distinguishable from one another. The description information may include a plurality of description texts (121, 122, 123, 124) describing the plurality of first images (10) to be generated. The plurality of description texts (121, 122, 123, 124) may correspond to each of the plurality of first images (11, 12, 13, 14). The number of description texts (121, 122, 123, 124) included in the description information may correspond to the number of the plurality of first images (11, 12, 13, 14) to be generated. For example, the number of descriptive texts (121, 122, 123, 124) included in the descriptive information may be equal to the number of multiple first images (11, 12, 13, 14) to be generated. Each descriptive text may include text describing the visual characteristics of each of the first images (10) to be generated. For example, each descriptive text included in the second prompt (120) may include text that causes the second artificial intelligence model (102) to generate one first image (11, 12, 13, 14) corresponding to each descriptive text as it is input to the second artificial intelligence model (102). The descriptive information may include information related to the position where the multiple first images (10) to be generated are placed within the second image (20). For example, the description information may include text containing information related to the location within the second image (20) where each first image (10) is to be placed, corresponding to the description text describing each first image (10).

[0090] Referring to [Table 2], for example, one descriptive text included in the second prompt (120) may include text that generates one first image (11, 12, 13, 14), such as ‘A bold kiss-cut sticker featuring a whimsical dragon chef with a tiny chef's hat, holding a giant lollipop spoon. The dragon is vibrant teal, pink, and orange, with simple, bold outlines and cheerful, childish features. The background is pure white.’ Additionally, in this case, the descriptive text may include ‘**Upper-Left:**’, which is text containing information regarding the position where the first image (10) to be generated by the descriptive text will be placed on the second image (20). Accordingly, the descriptive text describing the first image (11) to be placed on the upper left side of the second image (20) is ‘**Upper-Left:** A bold kiss-cut sticker featuring a whimsical dragon chef with a tiny chef's hat, holding a giant lollipop spoon. The dragon is vibrant teal, pink, and orange, with simple, bold outlines and cheerful, childish features. The background is pure white. It can be composed like / n / n'.

[0091] The second prompt (120) may be a single prompt that includes a plurality of descriptive texts describing a plurality of first images (10) to be generated. For example, the second prompt (120) may be a single prompt that includes a text in which a plurality of descriptive texts are listed.

[0092] In one embodiment, a plurality of descriptive texts (121, 122, 123, 124) may include text describing the visual characteristics of a plurality of first images to be generated. The descriptive texts (121, 122, 123, 124) may include text describing the visual characteristics of each first image (11, 12, 13, 14) corresponding to each descriptive text (121, 122, 123, 124). In one embodiment, the plurality of descriptive texts (121, 122, 123, 124) may include information related to the same characteristics. The plurality of descriptive texts (121, 122, 123, 124) may include text describing characteristics included in an input received from a user. For example, a plurality of descriptive texts (121, 122, 123, 124) may include texts of substantially the same meaning that describe characteristics included in the input received from the user. For example, a plurality of descriptive texts (121, 122, 123, 124) may include texts related to the same characteristics included in the input received from the user. For example, if the input received from the user includes information related to characteristics of an object to be included in the first image, each of the plurality of descriptive texts (121, 122, 123, 124) may include text describing the same characteristics related to the object to be included in the first image. In one embodiment, each descriptive text (121, 122, 123, 124) may have at least one difference in relation to other descriptive texts (121, 122, 123, 124). Multiple descriptive texts (121, 122, 123, 124) may include information related to different visual characteristics. That is, each descriptive text (121, 122, 123, 124) may include information for generating an image having different visual characteristics from the different descriptive texts (121, 122, 123, 124).For example, each descriptive text (121, 122, 123, 124) may be such that when applied to the second artificial intelligence model (102), it outputs an image having at least one difference from the image output when applying another descriptive text to the second artificial intelligence model (102).

[0093] In one embodiment, each descriptive text (121, 122, 123, 124) may include text describing the same characteristics related to information included in the input received from the user, but may also include text describing different characteristics related to information not included in the input received from the user. Accordingly, each descriptive text (121, 122, 123, 124) may be a text describing a first image (10) having characteristics corresponding to the input received from the user, while having detailed characteristics different from the other descriptive texts.

[0094] In one embodiment, the second prompt (120) may include text indicating the type of object to be included in the plurality of first images (10). In one embodiment, the second prompt (120) may include text indicating the position where each of the plurality of first images (10) is to be placed within the second image (20). For example, the second prompt (120) may include text indicating the position where each first image (10) is to be placed within the second image (20) so as to correspond to each description text corresponding to each first image (10).

[0095] In one embodiment, the second prompt (120) may include information that causes one second image (20) to be generated by inputting it into the second artificial intelligence model (102). The second prompt (120) may include information that the number of images to be generated by the second artificial intelligence model (102) is one. For example, referring to [Table 2], the second prompt (120) may include the text "sampleCount": 1. Accordingly, the second artificial intelligence model (102) that receives the second prompt (120) can generate one image.

[0096] For example, the second prompt (120) may further include information that causes the plurality of first images to be arranged based on the first layout within the second image (20). In this case, the second image obtained may be one in which the plurality of first images are arranged based on the first layout. For example, the electronic device (100) may display a layout list including a plurality of layouts. In this case, the electronic device (100) may receive user input selecting the first layout from the list. When user input selecting the first layout is received, the electronic device (100) may display at least some of the plurality of first images based on the selected first layout.

[0097] An electronic device (100) according to one embodiment may obtain a second image (20) output from a second artificial intelligence model (102) by applying a second prompt (120) to a second artificial intelligence model (102). The second artificial intelligence model (102) may be an artificial intelligence model trained to generate an image based on an input second prompt. The second artificial intelligence model (102) may generate a second image (20) based on the second prompt (120). The second artificial intelligence model (102) may generate a second image (20) comprising a plurality of first images (10) based on the second prompt (120). The second artificial intelligence model (102) may generate a second image (20) comprising a plurality of first images (10) that are distinguishable from one another.

[0098] FIG. 6 is a drawing showing an example of a second image generated according to one embodiment.

[0099] Identification numbers 610 and 620 illustrated in FIG. 6 represent examples of a second image (20) output by a second artificial intelligence model (102) based on a first prompt (110). The second image (20) may include a plurality of first images (11, 12, 13, 14). The second image (20) may include a number of first images (11, 12, 13, 14) as requested through user input. For example, referring to identification numbers 610 and 620 of FIG. 6, the electronic device (100) may receive input from a user to generate four images. In this case, the second image (20) output from the second artificial intelligence model (102) may include four first images (11, 12, 13, 14).

[0100] In one embodiment, a plurality of first images (10) within a second image (20) may be arranged so as to be distinguishable from one another. For example, the area within the second image (20) may be divided into a number of the plurality of first images (10), and each first image (10) may be placed in each divided area within the second image (20). For example, the plurality of first images (10) may be distinguished by a predetermined boundary within the second image (20). The boundary separating each of the first images (11, 12, 13, 14) within the second image (20) may be displayed on the second image (20) so as to be visually recognizable (e.g., example of identification number 620), or it may not be displayed on the second image (20) (e.g., example of identification number 610). For example, the second image (20) may be divided into equal parts as many times as there are multiple first images (10), and each first image (10) may be placed in each of the equally divided areas on the second image (20).

[0101] For example, with reference to identification numbers 610 and 620, the electronic device (100) may receive input from a user to generate four images related to a 'Dragon chef'. In this case, the second image (20) generated by the second artificial intelligence model (102) may include a plurality of first images (11, 12, 13, 14) that include common characteristics related to cooking dragons. Also, in this case, each of the first images (11, 12, 13, 14) may include at least some visual characteristics that are different from the other first images (11, 12, 13, 14).

[0102] Referring to the example of identification number 610, each first image (11, 12, 13, 14) can be generated according to the same style. For example, if the user input includes information to generate '4' images having an 'illustration style', the first prompt (110) may include content to generate four first images (11, 12, 13, 14) according to the same style called an 'illustration style'. Also, in this case, each descriptive text included in the second prompt (120) may include content to generate an image according to the 'illustration style'. Accordingly, the second image (20) generated from the second artificial intelligence model (102) may include four first images (11, 12, 13, 14) of the 'illustration style'.

[0103] Referring to the example of identification number 620, each first image (11, 12, 13, 14) may be generated according to different styles. For example, if the user input includes information to generate 'four' images of 'different styles', the first prompt (120) may include content to generate four first images (11, 12, 13, 14) according to different styles. Also, in this case, each descriptive text included in the second prompt (120) may include content to generate images according to different styles. For example, the first descriptive text (e.g., the first descriptive text (121) of FIG. 1) included in the second prompt (120) may include content that causes the first image (11) to be generated according to a style called 'doodle', the second descriptive text (e.g., the second descriptive text (122) of FIG. 1) may include content that causes the first image (12) to be generated according to a style called 'illustration', the third descriptive text (e.g., the third descriptive text (123) of FIG. 1) may include content that causes the first image (13) to be generated according to a style called '3D emoji', and the fourth descriptive text (e.g., the fourth descriptive text (124) of FIG. 1) may include content that causes the first image (14) to be generated according to a style called 'retro logo'. Accordingly, the second image (20) generated may include a first image (11) having a 'doodle' style, a first image (12) having an 'illustration' style, a first image (13) having a '3D emoji' style, and a first image (14) having a 'retro logo' style.

[0104] FIG. 7 is a drawing showing an example in which a second image (20) generated according to one embodiment is displayed.

[0105] For example, the second image (20) can be displayed through a display (e.g., the display module (1560) of FIG. 15). For example, the second image (20) generated by the second artificial intelligence model (102) can be displayed through a display (e.g., the display module (1560) of FIG. 15) without separate processing. In this case, the electronic device (100) can display the second image (20) on the display without separating the plurality of first images (10) included in the second image (20), and can store the second image (20) in memory (e.g., the memory (1530) of FIG. 15) without separating the plurality of first images (10) included in the second image (20).

[0106] For example, the electronic device (100) may display the second image (20) on a display without separating the plurality of first images (10) included in the second image (20), receive user input for a predetermined area on the displayed second image (20), and separate the first images (11, 12, 13, 14) displayed in the area where the user input was received from the second image (20) and provide them. Details related to this will be described later in FIGS. 8 to 11.

[0107] FIG. 8 is a flowchart illustrating the process of separating and displaying a plurality of first images (10) within a second image (20) generated in one embodiment.

[0108] The description related to identification numbers 810 to 820 of FIG. 8 may correspond to the description related to identification numbers 250 to 260 of FIG. 2. Below, in relation to this, the process of separating and displaying a plurality of first images (10) within a generated second image (20) will be described in more detail.

[0109] Referring to identification number 810, in one embodiment, the first image (10) may be separated within the second image (20). All of the plurality of first images (10) included within the second image (20) may be separated, or at least one first image (10) included within the second image (20) may be separated. An electronic device (100) according to one embodiment may divide the second image (20) into regions corresponding to each first image (10). For example, the electronic device (100) may divide regions on the second image (20) using coordinate values. In this case, the electronic device (100) may divide each region corresponding to each first image (10) on the second image (20) using coordinate values. Additionally, the electronic device (100) can identify coordinate values ​​for regions corresponding to each first image (10) or coordinate values ​​for boundaries separating regions corresponding to each first image (10). The electronic device (100) can separate regions on the second image (20) into regions corresponding to each first image (10) through coordinate values, and separate regions corresponding to each first image (10) based on boundaries between each separated region, thereby obtaining a first image (10) separated from the second image (20). An electronic device (100) according to one embodiment may store an algorithm for separating at least one image or object from an image containing a plurality of images or a plurality of objects. In this case, the electronic device (100) can apply the second image (20) to the stored algorithm to obtain a first image (10) separated within the second image (20). According to one embodiment, the electronic device (100) can obtain a first image (10) separated from the second image (20) by applying the second image (20) to a third artificial intelligence model.The third artificial intelligence model may be an artificial intelligence model trained to receive an image containing multiple images or multiple objects, and to separate and output at least one image or object from the input image. For example, the third artificial intelligence model may be an artificial intelligence model trained to perform a segmentation function, and may be stored in an electronic device (100) or may be stored in a server (e.g., the server (1508) of FIG. 15).

[0110] Referring to identification number 820, in one embodiment, separated first images (11, 12, 13, 14) may be displayed. The separated first images (11, 12, 13, 14) may be displayed separately through a display (e.g., the display module (1560) of FIG. 15) or may be displayed simultaneously on the display.

[0111] FIG. 9 is a flowchart illustrating the process of a first image displayed in an area where input is received from a user being separated from a second image and displayed, in one embodiment.

[0112] The description related to identification numbers 910 to 930 in FIG. 9 may correspond to the description related to identification numbers 810 to 820 in FIG. 8. In this regard, the process of a first image (11, 12, 13, 14) displayed in an area where input is received from a user being separated from a second image (20) and displayed will be described in more detail below.

[0113] Referring to identification number 910, the electronic device (100) can receive input from a user regarding a specific area within the second image (20). For example, referring to the drawing illustrated in FIG. 7, the electronic device (100) can output the second image (20) output by the second artificial intelligence model (120). In this case, the electronic device (100) can receive input from a user regarding a specific area within the second image (20). For example, the electronic device (100) can receive touch input from a user touching a specific area among the areas on a display (e.g., the display module (1530) of FIG. 15) where the second image (20) is displayed. In one embodiment, the electronic device (100) can divide the second image (20) into areas where each first image (10) is displayed, and identify the first image (10) displayed in the area that includes the area where the input from the user was received among each divided area. For example, the regions within the second image (20) may be divided into regions corresponding to each of the first images (10). In this case, the electronic device (100) can identify, among the divided regions, the region corresponding to the region where input from the user is received. Accordingly, the electronic device (100) can determine that input from the user has been received, which causes the first images (11, 12, 13, 14) displayed in the region identified as corresponding to the region where input from the user is received to be separated from the second image (20).

[0114] Referring to identification number 920, the electronic device (100) can separate the first image (10) displayed in the input received area within the second image (20). A detailed description of the method for separating the first image (10) within the second image (20) has been described in the description section related to identification number 810 of FIG. 8, so it will be omitted here.

[0115] Referring to identification number 930, the electronic device (100) can display a separated first image (10). A detailed description of the method for displaying the first image (10) has been described in the description section related to identification number 820 of FIG. 8, so it will be omitted here.

[0116] FIG. 10 is an example drawing for explaining a method in which, in one embodiment, the second image (20) is divided by region, and the first image (10) displayed in the region where user input is received within the second image (20) is separated from the second image (20).

[0117] FIG. 10 is an example drawing showing the process of separating a first image (11) displayed in an area where user input is received within a second image (20) from the second image (20) through a process corresponding to the flowchart of FIG. 9, and the first image (11) obtained accordingly.

[0118] Identification number 1010 represents a second image (20) in a state divided into regions corresponding to each first image (10). For example, the second image (20) may include four first images (11, 12, 13, 14). In this case, the regions on the second image (20) may be divided into four regions in which each first image (11, 12, 13, 14) is displayed. For example, the electronic device (100) may receive a user's touch input for a region in which a predetermined first image (11) is displayed among the four regions in which the first images (11, 12, 13, 14) are displayed within the second image (20). In this case, the electronic device (100) may separate the first image (11) included in the region in which the user's touch input was received within the second image (20). Accordingly, in identification number 1020, the electronic device (100) can obtain the first image (11) separated from the second image (20).

[0119] FIG. 11 is a drawing showing an example in which, in one embodiment, a plurality of first images (10) are separated from a second image (20) and each first image (10) is displayed.

[0120] The drawings corresponding to identification numbers 1110, 1120, 1130, and 1140 disclosed in FIG. 11 illustrate an example in which a plurality of first images (10) are separated from a second image (20) according to the process disclosed in FIG. 8, and then the separated plurality of first images (10) are displayed.

[0121] For example, the electronic device (100) may separate a plurality of first images (10) included in a second image (20) output from a second artificial intelligence model (102), and display the separated plurality of first images (10) separately through a display (e.g., the display module (1560) of FIG. 15). For example, when the separated plurality of first images (10) are displayed separately, they may be displayed through different pages, and the pages may be switched so that different first images (10) are displayed by user input. For example, referring to the example of FIG. 11, the electronic device (100) may provide different first images (10) among the separated plurality of first images (10) to the user by displaying the page with identification number 1120 upon receiving user input to move to the next page while displaying the page with identification number 1110. However, the method of displaying the separated first image (10) is not limited to this, and the separated first images (10) can be provided to the user through various methods.

[0122] FIG. 12 is a flowchart for explaining the process of acquiring and displaying a plurality of third images by increasing the resolution of a plurality of first images (10) in one embodiment.

[0123] The description related to identification numbers 1210 to 1220 of FIG. 12 is a description related to operations that can be performed after the second image (20) is acquired through the process of identification number 240 of FIG. 2. Specifically, after the second image (20) is acquired, I will describe a process of acquiring a third image having a higher resolution than the acquired first images (10) by increasing the resolution of the first images (10) included in the second image (20).

[0124] The operations corresponding to identification numbers 1210 to 1220 may be performed on the acquired second image (20) itself, or may be performed on the separated first images (10) after separating a plurality of first images (10) from the acquired second image (20).

[0125] In identification number 1210, the electronic device (100) can acquire a plurality of third images by increasing the resolution of a plurality of first images (10). According to one embodiment, the electronic device (100) can convert a low-resolution first image (10) generated from a generative artificial intelligence model into a high-resolution image. The electronic device (100) can acquire a third image by increasing the number of pixels of the first image (10) while maintaining or enhancing the details of the first image (10). For example, the resolution of the first image (10) can be converted through upscaling so that a third image with a higher resolution than the first image (10) can be acquired. For example, the electronic device (100) can utilize a neural network-based super-resolution technology to acquire a third image based on the first image (10). For example, a convolutional neural network (CNN), a generative adversarial network (GAN), a transformer model, etc., may be used to obtain a third image based on a first image (10). The third artificial intelligence model used to obtain a third image based on the first image (10) may be trained to predict and restore details of a high-resolution image from a low-resolution image. The third artificial intelligence model used to obtain a third image based on the first image (10) may be trained to generate the most natural and clear results based on patterns learned from training data. For example, to obtain a third image based on the first image (10), post-processing techniques including contour enhancement, noise removal, and texture enhancement may be applied. Accordingly, a third image (30) having a higher resolution may be output while maintaining the quality of the first image (10).

[0126] In identification number 1220, the electronic device (100) can display third images.

[0127] The process of FIG. 12 is not an operation that must be performed in order to carry out the method of the present disclosure, and may be performed optionally when a high-resolution image is required. However, it is not limited to such cases, and the operations of FIG. 12 may be performed in various cases to obtain an image having a higher resolution.

[0128] FIG. 13 is a block diagram showing an electronic device according to one embodiment and artificial intelligence models stored in a server.

[0129] According to one embodiment, an electronic device (100) can be connected to an AI server (1200). The electronic device (100) can be connected to the AI ​​server (1200) through a communication module (e.g., (1590) of FIG. 15) to transmit and receive data.

[0130] In one embodiment, the AI ​​server (1200) may store at least one of a first artificial intelligence model (101) or a second artificial intelligence model (102). The electronic device (100) may transmit a first prompt (110) to the first artificial intelligence model (101) deployed in the AI ​​server (1200) and receive a second prompt (120) from the first artificial intelligence model (101) deployed in the AI ​​server (1200). The electronic device (100) may transmit a second prompt (120) to the second artificial intelligence model (102) deployed in the AI ​​server (1200) and receive a second image (20) including a plurality of first images (10) from the second artificial intelligence model (102) deployed in the AI ​​server (1200).

[0131] However, unlike the block diagram shown in FIG. 13, at least one of the artificial intelligence models according to the present disclosure (e.g., first artificial intelligence model (101), second artificial intelligence model (102), third artificial intelligence model) may be stored in the memory of the electronic device (100) (e.g., memory (1530) of FIG. 15). For example, at least one of the artificial intelligence models of the present disclosure may be stored in the memory of the electronic device (100) (e.g., memory (1530) of FIG. 15) in the form of a program.

[0132] FIG. 14 is a block diagram illustrating data transmitted and received between an electronic device (100), a first artificial intelligence model (101), and a second artificial intelligence model (102) according to one embodiment.

[0133] The electronic device (100) can transmit and receive data with at least one of the first artificial intelligence model (101) or the second artificial intelligence model (102).

[0134] The electronic device (100) can receive input from a user. The electronic device (100) that receives input from a user can generate a first prompt (110) based on the input from the user. The electronic device (100) can transmit the generated first prompt (110) to a first artificial intelligence model (101) and receive a second prompt (120) from the first artificial intelligence model (101). The electronic device (100) can transmit the second prompt (120) received from the first artificial intelligence model (101) to a second artificial intelligence model (102) and receive a second image (20) including a plurality of first images (10) from the second artificial intelligence model (102).

[0135] FIG. 15 is a block diagram of an electronic device (1501) in a network environment (1500) according to various embodiments.

[0136] The electronic device (100) of the present disclosure may correspond to the electronic device (1501) of FIG. 15 and may include components that may be included in the electronic device (1501) disclosed in FIG. 15. However, the electronic device (100) of the present disclosure is not required to include all components of FIG. 15, and may include only some components of FIG. 15.

[0137] Referring to FIG. 15, in a network environment (1500), an electronic device (1501) may communicate with an electronic device (1502) through a first network (1598) (e.g., a short-range wireless communication network) or with at least one of an electronic device (1504) or a server (1508) through a second network (1599) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (1501) may communicate with the electronic device (1504) through a server (1508). According to one embodiment, the electronic device (1501) may include a processor (1520), memory (1530), input module (1550), sound output module (1555), display module (1560), audio module (1570), sensor module (1576), interface (1577), connection terminal (1578), haptic module (1579), camera module (1580), power management module (1588), battery (1589), communication module (1590), subscriber identification module (1596), or antenna module (1597). In some embodiments, at least one of these components (e.g., connection terminal (1578)) may be omitted from the electronic device (1501), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (1576), camera module (1580), or antenna module (1597)) may be integrated into a single component (e.g., display module (1560)).

[0138] The processor (1520) can, for example, execute software (e.g., program (1540)) to control at least one other component (e.g., hardware or software component) of the electronic device (1501) connected to the processor (1520) and perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (1520) can store commands or data received from other components (e.g., sensor module (1576) or communication module (1590)) in volatile memory (1532), process the commands or data stored in volatile memory (1532), and store the resulting data in non-volatile memory (1534). According to one embodiment, the processor (1520) may include a main processor (1521) (e.g., a central processing unit or an application processor) or an auxiliary processor (1523) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (1501) includes a main processor (1521) and an auxiliary processor (1523), the auxiliary processor (1523) may be configured to use lower power than the main processor (1521) or to be specialized for a specified function. The auxiliary processor (1523) may be implemented separately from the main processor (1521) or as part thereof.

[0139] The auxiliary processor (1523) may control at least some of the functions or states associated with at least one component of the electronic device (1501) (e.g., display module (1560), sensor module (1576), or communication module (1590)) on behalf of the main processor (1521) while the main processor (1521) is in an inactive (e.g., sleep) state, or together with the main processor (1521) while the main processor (1521) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (1523) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (1580) or communication module (1590)). According to one embodiment, the auxiliary processor (1523) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (1501) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (1508)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.

[0140] The memory (1530) can store various data used by at least one component of the electronic device (1501) (e.g., processor (1520) or sensor module (1576)). The data may include, for example, input data or output data for software (e.g., program (1540)) and related commands. The memory (1530) may include volatile memory (1532) or non-volatile memory (1534).

[0141] The program (1540) may be stored as software in memory (1530) and may include, for example, an operating system (1542), middleware (1544), or an application (1546).

[0142] The input module (1550) can receive commands or data to be used for a component of the electronic device (1501) (e.g., processor (1520)) from outside the electronic device (1501) (e.g., user). The input module (1550) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0143] The sound output module (1555) can output a sound signal to the outside of the electronic device (1501). The sound output module (1555) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.

[0144] The display module (1560) can visually provide information to an external (e.g., user) of the electronic device (1501). The display module (1560) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (1560) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.

[0145] The audio module (1570) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (1570) can acquire sound through an input module (1550) or output sound through an audio output module (1555) or an external electronic device (e.g., electronic device (1502)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (1501).

[0146] The sensor module (1576) can detect the operating state of the electronic device (1501) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (1576) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0147] The interface (1577) may support one or more specified protocols that can be used for the electronic device (1501) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (1502)). According to one embodiment, the interface (1577) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0148] The connection terminal (1578) may include a connector through which the electronic device (1501) can be physically connected to an external electronic device (e.g., electronic device (1502)). According to one embodiment, the connection terminal (1578) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0149] The haptic module (1579) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (1579) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.

[0150] The camera module (1580) can capture still images and video. According to one embodiment, the camera module (1580) may include one or more lenses, image sensors, image signal processors, or flashes.

[0151] The power management module (1588) can manage the power supplied to the electronic device (1501). According to one embodiment, the power management module (1588) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).

[0152] The battery (1589) can supply power to at least one component of the electronic device (1501). According to one embodiment, the battery (1589) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0153] The communication module (1590) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (1501) and an external electronic device (e.g., electronic device (1502), electronic device (1504), or server (1508)), and the performance of communication through the established communication channel. The communication module (1590) may include one or more communication processors that operate independently of the processor (1520) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1590) may include a wireless communication module (1592) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (1594) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (1504) via a first network (1598) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (1599) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1592) can identify or authenticate the electronic device (1501) within a communication network such as the first network (1598) or the second network (1599) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (1596).

[0154] The wireless communication module (1592) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (1592) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (1592) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (1592) can support various requirements specified in the electronic device (1501), external electronic device (e.g., electronic device (1504)), or network system (e.g., second network (1599)). According to one embodiment, the wireless communication module (1592) can support a Peak data rate (e.g., 20 Gbps or more) for realizing eMBB, loss coverage (e.g., 164 dB or less) for realizing mMTC, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for realizing URLLC.

[0155] An antenna module (1597) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (1597) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (1597) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (1598) or a second network (1599), may be selected from the plurality of antennas, for example, by a communication module (1590). A signal or power may be transmitted or received between the communication module (1590) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (1597).

[0156] According to various embodiments, the antenna module (1597) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.

[0157] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.

[0158] According to one embodiment, commands or data may be transmitted or received between the electronic device (1501) and an external electronic device (1504) through a server (1508) connected to a second network (1599). Each of the external electronic devices (1502, or 1504) may be the same or a different type of device as the electronic device (1501). According to one embodiment, all or part of the operations performed on the electronic device (1501) may be performed on one or more of the external electronic devices (1502, 1504, or 1508). For example, if the electronic device (1501) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (1501) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (1501). The electronic device (1501) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (1501) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (1504) may include an Internet of Things (IoT) device. The server (1508) may be an intelligent server using machine learning and / or neural networks.According to one embodiment, an external electronic device (1504) or server (1508) may be included within the second network (1599). The electronic device (1501) may be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0159] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.

[0160] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.

[0161] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0162] Various embodiments of the present document may be implemented as software (e.g., program (1540)) comprising one or more instructions stored in a storage medium (e.g., internal memory (1536) or external memory (1538)) readable by a machine (e.g., electronic device (1501)). For example, a processor (e.g., processor (1520)) of the machine (e.g., electronic device (1501)) may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0163] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TM It can be distributed online (e.g., downloaded or uploaded) through ) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0164] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

[0165] FIG. 16 is a diagram showing an example of a system including a generative artificial intelligence model according to one embodiment.

[0166] The artificial intelligence system may include a user query / response interface (1610), an AI framework (1620), an application / service component (1630), knowledge repositories (1640), and / or a generative AI model (1650).

[0167] Referring to FIG. 16, a user query / response interface (1610) may receive input. The input may include user input and / or data obtained or generated by an electronic device (e.g., electronic device (101, 201, 401, 601, 701, 901)). The data may include images, videos, and / or sensor data generated by at least one processor of the electronic device (e.g., at least one processor (120)) (e.g., illuminance data around the electronic device obtained from a sensor or sensor hub (e.g., auxiliary processor (123), attitude data (or orientation data) of the electronic device, temperature inside the electronic device (e.g., display module (160)) or temperature of at least one processor (120)), size information of the display area of ​​the display module (160), and / or images obtained through an image sensor of the electronic device (e.g., included in a camera module (180)). For example, user input may take the form of biosignals, natural language, touch data acquired through touch circuits included within the display module (e.g., used to identify input from a finger and / or stylus), images, and / or videos. Additionally, context information may be transmitted along with the user input. Context information may include various additional information at the time of user input. For example, this could include information about the application currently being used by the user or the user's location information. Furthermore, user input may take the form of a mixture of the aforementioned biosignals, natural language, images, sounds, and context information. Additionally, user input may take the form of non-natural language input, such as input for selecting a menu.

[0168] The user query / response interface (1610) can output results from a generative artificial intelligence system to the user. The output may include results (or result information) generated or obtained by the artificial intelligence system (1600) based on at least part of the input. The output may be in the form of natural language or specific content, and may also be provided in a form such as an action requested by the user. For example, the output may have a format according to the user settings of the electronic device. For example, the output may include information regarding an emergency situation in which the user is facing. The output may include detailed information about the emergency situation, for example, information indicating the type of emergency situation, information regarding the user's surrounding circumstances, personal status (health status) information, location information, information regarding the user's request, and information indicating the truthfulness of the user's request.

[0169] The AI ​​framework (1620) can receive input from the user and coordinate and control each component necessary to perform the user's intent based on the user's query.

[0170] User input received from the user query / response interface (1610) can be transmitted to a prompt design component (1621). The prompt design component (1621) can be used to generate a prompt suitable for inputting the user input into a generative AI model (1650) (e.g., a large language model (LLM), a large vision model (LVM), and / or large multimodal models (LMM)).

[0171] The prompt design component (1621) may be an AI component that uses machine learning algorithms or neural networks to develop better prompts over time. The prompt design component (1621) may generate prompts by accessing a knowledge component (e.g., knowledge repository (1640)) containing user preference data, a prompt library, and prompt examples based on user input, and may pass the generated prompts to a generative AI model (1650) (e.g., LLM, LVM, and / or LMM).

[0172] The APIs / Plugins management component (1623) can perform the role of communicating with external information when there is a request for additional information when user input is passed as input to the generative AI model (1650). The APIs / Plugins management component (1623) establishes a channel to communicate with the outside of the AI ​​interface via APIs, and can enable access to various data sources (e.g., knowledge repository (1640)) through the established channel. For example, the APIs / Plugins management component (1623) can be used to request another component (e.g., application / service component (1630)) that performs feedback (or response) according to the prompt. If the application or service needs to perform an action that ultimately executes the user's input rather than an intermediate result, the APIs / Plugins management component (1623) can request that action from the application / service component (1630) via APIs. Information obtained from the outside may be used to generate a prompt in the prompt design component (1621) along with user input, or it may be passed as input to the generative model.

[0173] A refiner component (e.g., output modification component (1625)) can fine-tune (or adjust) (or modify) the output produced by a generative AI model (1650) (e.g., LLM, LVM, and / or LMM). For example, the refiner component can verify whether the content generated by the generative AI model (1650) (e.g., LLM, LVM, and / or LMM) is irrelevant, contains biased content, or contains harmful content. Additionally, the refiner component can determine the extent to which the output matches the desired result and, if additional processing is required, proceed with that process. Furthermore, the refiner component can configure and provide hints to the user to help avoid unwanted outputs.

[0174] A generative AI model (1650) generally refers to an artificial intelligence neural network that generates new forms of data based on user input information. A generative AI model (1650) may include a model that generates images and / or a model that generates language. Models that generate images include, but are not limited to, GANs (generative adversarial networks) and VAEs (variational autoencoders), and examples include diffusion-based generative models that use VAEs and Transformer structures. Models that generate language are models trained to output the most statistically appropriate output value based on input values, and examples include models such as CHAT-GPT 3 and CHAT-GPT 4. There are also LMMs that can recognize various forms of data input, such as sensing data, biosignals, text, images, and voice, and generate new data corresponding to them.

[0175] In one embodiment, the AI ​​framework (1620) and / or generative AI model (1650) may be included within a server communicating with the electronic device, but is not limited thereto, and may be included within an AI module (e.g., including a processing circuit) within the electronic device. For example, the AI ​​module may be operatively coupled with at least one processor of the electronic device (e.g., at least one processor (120)). For example, the AI ​​module may be operatively coupled with a sensor hub of the electronic device for one or more sensors within the electronic device.

[0176] According to one embodiment, the electronic device may be configured to include at least some of the user query / response interface (1610) AI framework (1620), application / service component (1630), knowledge repository (1640), or generative AI model (1650) of FIG. 10. According to one embodiment, at least some of the user query / response interface (1610) AI framework (1620), application / service component (1630), knowledge repository (1640), or generative AI model (1650) of FIG. 10 may be included in another electronic device (e.g., a server, a wearable electronic device, or another user's electronic device) that communicates with the electronic device.

[0177] A method for acquiring a plurality of images according to the present disclosure may include: acquiring text for generating an image based on user input; generating a first prompt for generating the image based on the text; acquiring a second prompt including description information for generating a second image including a plurality of first images that are distinguishable from one another by applying the first prompt to a first artificial intelligence model; acquiring the second image including the plurality of first images by applying the second prompt to a second artificial intelligence model; and displaying the plurality of first images.

[0178] The above method may further include, based on the acquisition of the second image, the operation of separating the plurality of first images within the second image, and the operation of displaying at least one of the separated plurality of first images.

[0179] The above user input may include user input for generating multiple images.

[0180] The above description information may include a plurality of description texts describing a plurality of first images to be generated.

[0181] The above plurality of first images may each correspond to the above plurality of descriptive texts.

[0182] The above plurality of descriptive texts may include text describing the visual characteristics of the plurality of first images to be generated.

[0183] The second prompt above may be a single prompt comprising the plurality of descriptive texts describing the plurality of first images to be generated.

[0184] The above multiple descriptive texts may include at least some information related to different visual characteristics.

[0185] The first prompt above may include information related to the number of the plurality of first images to be included in the second image.

[0186] The second prompt may include information that generates one second image by being input into the second artificial intelligence model.

[0187] The second prompt may include descriptive text describing the visual characteristics of each of the plurality of first images, and text indicating the position where each of the plurality of first images is to be placed within the second image.

[0188] The operation of separating the plurality of first images within the second image may include the operation of obtaining the plurality of first images separated from the second image by applying the second image to a third artificial intelligence model.

[0189] The second prompt may further include information that causes the plurality of first images to be arranged based on the first layout within the second image. The acquired second image may be one in which the plurality of first images are arranged based on the first layout.

[0190] The above method may further include the operation of displaying a layout list including a plurality of layouts, and the operation of receiving user input selecting the first layout from the list.

[0191] The operation of displaying the plurality of first images may include the operation of displaying at least some of the plurality of first images based on the selected first layout.

[0192] The above method may further include the operation of displaying the second image, the operation of receiving input from a user regarding a predetermined area within the second image, the operation of separating a first image displayed in the predetermined area where the input was received from the second image, and the operation of displaying the separated first image.

[0193] The above method may further include the operation of acquiring a plurality of third images based on the plurality of first images. The plurality of third images may include images with increased resolution of the plurality of first images.

[0194] The above method may further include the operation of displaying at least one of the plurality of third images obtained above.

[0195] The above method may be performed by an electronic device. At least one of the first artificial intelligence model or the second artificial intelligence model may be deployed on a server communicating with the electronic device.

[0196] An electronic device according to the present disclosure may include a memory for storing instructions and at least one processor. The instructions may be executed individually or collectively by the at least one processor so that the electronic device, based on user input, obtains text for generating an image, generates a first prompt for generating the image based on the text, obtains a second prompt including descriptive information for generating a second image including a plurality of first images that are distinguishable from one another by applying the first prompt to a first artificial intelligence model, obtains the second image including the plurality of first images by applying the second prompt to a second artificial intelligence model, and displays the plurality of first images.

[0197] The above instructions may be executed individually or collectively by the at least one processor so that the electronic device separates the plurality of first images within the second image based on the acquisition of the second image, and displays at least one of the separated plurality of first images.

[0198] The second prompt may include information that generates one second image by being input into the second artificial intelligence model.

[0199] The above memory can further store the above first artificial intelligence model and the above second artificial intelligence model.

[0200] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0201] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TMIt can be distributed online (e.g., downloaded or uploaded) through ) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0202] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. A method for an electronic device to acquire a plurality of images, An operation to obtain text for generating a second image including a plurality of first images that are distinguishable from one another, based on user input; The operation of generating a first prompt to be applied to a first artificial intelligence model based on the above text; The operation of obtaining a second prompt including description information for generating the second image by applying the first prompt to the first artificial intelligence model; The operation of obtaining the second image including the plurality of first images by applying the second prompt to the second artificial intelligence model; and It includes the operation of displaying the plurality of first images above, A method in which the plurality of first images are independently distinguishable from each other within the second image.

2. In Claim 1, Based on the second image obtained above, the operation of separating the plurality of first images within the second image; and A method further comprising the operation of displaying at least one of the above-described separated plurality of first images.

3. In Claim 1, A method in which the input of the above user includes text for generating the second image including the plurality of first images.

4. In Claim 1, A method in which the above-described information includes a plurality of descriptive texts describing a plurality of first images to be generated. A method in which the plurality of first images correspond to each of the plurality of descriptive texts.

5. In Claim 4, A method in which the plurality of descriptive texts described above include text describing the visual characteristics of each of the plurality of first images to be generated.

6. In Claim 5, A method in which the second prompt is a single prompt comprising the plurality of descriptive texts describing the plurality of first images to be generated.

7. In Claim 5, A method in which the above-mentioned plurality of descriptive texts contain at least some information related to different visual characteristics.

8. In Claim 1, A method in which the first prompt includes information related to the number of the plurality of first images to be included in the second image.

9. In Claim 1, A method in which the second prompt above includes information that generates one image by being input into the second artificial intelligence model.

10. In Claim 1, A method in which the second prompt comprises descriptive texts describing the visual characteristics of each of the plurality of first images, and text indicating the position where each of the plurality of first images is to be placed within the second image.

11. In Claim 1, The second prompt further includes information that causes the plurality of first images to be arranged based on the first layout within the second image, and A method in which the second image obtained above is a plurality of first images arranged based on the first layout.

12. In Claim 1, The operation of displaying the above second image; The operation of receiving input from the user regarding a predetermined area within the second image; The operation of separating a first image displayed in the predetermined area where the above input is received from the second image; and Further including the operation of displaying the above-described separated first image, method.

13. In Claim 1, The operation of acquiring a plurality of third images based on the plurality of first images, wherein the plurality of third images include images with increased resolution of the plurality of first images; and A method further comprising the operation of displaying at least one of the plurality of third images obtained above.

14. In Claim 1, A method in which at least one of the first artificial intelligence model or the second artificial intelligence model is placed on a server communicating with the electronic device.

15. In electronic devices: Memory for storing instructions; and It includes at least one processor, The above instructions are executed individually or collectively by the above at least one processor, and the electronic device: Based on user input, obtain text for generating an image, and Based on the above text, a first prompt for generating the above image is generated, and By applying the above-mentioned first prompt to a first artificial intelligence model, a second prompt is obtained that includes description information for generating a second image comprising a plurality of first images that are distinguishable from one another, and By applying the second prompt to the second artificial intelligence model, the second image including the plurality of first images is obtained, and An electronic device for displaying the plurality of first images.