Method and electronic device for providing personalized image

By extracting keywords and feature information from user images, personalized images are generated, solving the problem of generating images specifically for a particular user in existing technologies, and realizing the generation of personalized images and privacy protection.

CN121079680APending Publication Date: 2025-12-05SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480028082.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-11
Filing Date
2024-06-11
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing generative AI image generation models struggle to generate personalized results tailored to specific users.

Method used

By extracting keywords from users' images, searching for reference images corresponding to the keywords, segmenting relevant regions, obtaining image feature information, and using a generative model to generate personalized images.

Benefits of technology

It enables the generation of personalized images based on user images, reflecting the user's individual characteristics, avoiding the need to retrain the generation model, and protecting user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121079680A_ABST
    Figure CN121079680A_ABST
Patent Text Reader

Abstract

A method by which an electronic device generates a personalized image is provided. The method may include obtaining a cue for image generation, extracting a keyword from the cue, searching for one or more reference images corresponding to the keyword from among a plurality of images stored in the electronic device, segmenting a region related to the keyword in each of the one or more reference images, and generating an image based on the segmented region. And obtaining image feature information based on the region related to the keyword, transmitting the cue and the image feature information to a server, and receiving, from the server, a personalized image generated by providing the cue and the image feature information as inputs to the generative model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The disclosure relates to a method, an electronic device, a server, and a system for generating a personalized image and providing the personalized image to a user. BACKGROUND

[0002] Generative artificial intelligence (AI) can refer to a technology for learning about structures and patterns of large-scale data and generating new synthetic data based on input data. In some embodiments, generative AI can be used to generate human-level results from various tasks related to text, images, speech, video, music, etc. For example, an image generation model can be used to generate a new image based on given text, or can be used to generate a new image based on a given image.

[0003] However, even when an image generation model can generate an image containing content matching an input based on given text or a given image, it can be difficult to obtain a result specialized for a specific user. SUMMARY

[0004] TECHNICAL SOLUTION According to an aspect of the disclosure, a method by which an electronic device provides a personalized image can be provided. The method can include obtaining a prompt for image generation. The method can include extracting a keyword from the prompt. The method can include searching for one or more reference images corresponding to the keyword from among a plurality of images stored in the electronic device. The method can include segmenting a region related to the keyword within each of the one or more reference images. The method can include obtaining image feature information based on the region related to the keyword of the reference images. The method can include transmitting the prompt and the image feature information to a server. The method can include receiving a personalized image generated by using a generative model that receives the prompt and the image feature information as input from the server.

[0005] According to an aspect of the disclosure, an electronic device for generating a personalized image is provided. The electronic device can include a communication interface, a memory storing one or more instructions, and one or more processors configured to execute the one or more instructions stored in the memory. The one or more processors can be configured to execute the one or more instructions to obtain a prompt for image generation. The one or more processors can be configured to execute the one or more instructions to extract a keyword from the prompt. The one or more processors can be configured to execute the one or more instructions to search for one or more reference images corresponding to the keyword from among a plurality of images stored in the electronic device. The one or more processors can be configured to execute the one or more instructions to segment an area related to the keyword within each of the one or more reference images. The one or more processors can be configured to execute the one or more instructions to obtain image feature information based on the area related to the keyword of the reference images. The one or more processors can be configured to execute the one or more instructions to transmit the prompt and the image feature information to a server through the communication interface. The one or more processors can be configured to execute the one or more instructions to receive, through the communication interface from the server, a personalized image generated by using a generation model that receives the received prompt and the image feature information as input.

[0006] According to an aspect of the disclosure, a method of a server through which a personalized image is provided can be provided. The method can include obtaining a prompt for image generation. The method can include extracting a keyword from the prompt. The method can include searching for one or more reference images corresponding to the keyword from among a plurality of images stored in an electronic device. The method can include segmenting an area related to the keyword within each of the one or more reference images. The method can include obtaining image feature information based on the area related to the keyword of the reference images. The method can include generating a personalized image by using a generation model. The generation model can be an artificial intelligence model configured to receive the prompt and the image feature information as input and output the personalized image.

[0007] According to an aspect of the disclosure, a server for providing a personalized image can be provided. The server can include a communication interface, a memory storing one or more instructions, and one or more processors configured to execute the one or more instructions. The one or more processors can be configured to execute the one or more instructions to obtain a prompt for image generation. The one or more processors can be configured to execute the one or more instructions to extract a keyword from the prompt. The one or more processors can be configured to execute the one or more instructions to search for one or more reference images corresponding to the keyword from among a plurality of images stored in an electronic device. The one or more processors can be configured to execute the one or more instructions to segment an area related to the keyword within each of the one or more reference images. The one or more processors can be configured to execute the one or more instructions to obtain image feature information based on the area related to the keyword of the reference images. The one or more processors can be configured to execute the one or more instructions to generate a personalized image by using a generative model. The generative model can be an artificial intelligence model configured to receive the prompt and the image feature information as input and output the personalized image.

[0008] According to an aspect of the disclosure, a computer-readable recording medium having recorded thereon a program for executing any one of the above and below-described methods by which an electronic device and / or a server provides a personalized image can be provided. BRIEF DESCRIPTION OF DRAWINGS

[0009] Figure 1 FIG. 1 is a diagram schematically illustrating an electronic device according to an embodiment of the disclosure.

[0010] Figure 2 FIG. 2 is a flowchart for describing an operation of an electronic device generating a personalized image according to an embodiment of the disclosure.

[0011] Figure 3 FIG. 3 is a diagram illustrating operations of an electronic device and a server according to an embodiment of the disclosure.

[0012] Figure 4 FIG. 4 is a diagram for describing an operation of an electronic device extracting a keyword according to an embodiment of the disclosure.

[0013] Figure 5a FIG. 5 is a diagram for describing an operation of an electronic device searching for an image according to an embodiment of the disclosure.

[0014] Figure 5b FIG. 6 is a diagram for describing an operation of an electronic device searching for an image according to an embodiment of the disclosure.

[0015] Figure 6FIG. 1 is a diagram for describing an operation of an electronic device segmenting an image according to an embodiment of the disclosure.

[0016] Figure 7 FIG. 2 is a diagram for describing an operation of an electronic device generating a vector representation according to an embodiment of the disclosure.

[0017] Figure 8 FIG. 3 is a diagram for describing an operation of a server generating a personalized image according to an embodiment of the disclosure.

[0018] Figure 9 FIG. 4 is a diagram for describing an operation of an electronic device providing an example of a personalized image according to an embodiment of the disclosure.

[0019] Figure 10a FIG. 5 is a flowchart illustrating operations of an electronic device and a server according to an embodiment of the disclosure.

[0020] Figure 10b FIG. 6 is a flowchart illustrating additional operations of an electronic device and a server according to an embodiment of the disclosure.

[0021] Figure 11 FIG. 7 is a diagram for describing an operation of an electronic device providing image feature information to a user according to an embodiment of the disclosure.

[0022] Figure 12 FIG. 8 is a diagram for describing an operation of an electronic device generating or processing a hint according to an embodiment of the disclosure.

[0023] Figure 13 FIG. 9 is a diagram for describing an operation of an electronic device providing a personalized image according to an embodiment of the disclosure.

[0024] Figure 14a FIG. 10 is a diagram for describing an operation of an electronic device providing a personalized image according to an embodiment of the disclosure.

[0025] Figure 14b FIG. 11 is a diagram for describing an operation of an electronic device providing a personalized image according to an embodiment of the disclosure.

[0026] Figure 14c FIG. 12 is a diagram for describing an operation of an electronic device providing a personalized image according to an embodiment of the disclosure.

[0027] Figure 14d FIG. 13 is a diagram for describing an operation of an electronic device providing a personalized image according to an embodiment of the disclosure.

[0028] Figure 15 FIG. 14 is a diagram for describing an operation of an electronic device providing a personalized image according to an embodiment of the disclosure.

[0029] Figure 16FIG. 1 is a diagram for describing an operation of an electronic device providing information related to a personalized image according to an embodiment of the disclosure.

[0030] Figure 17 FIG. 2 is a diagram for describing an operation of an electronic device providing a personalized image according to an embodiment of the disclosure.

[0031] Figure 18a FIG. 3 is a block diagram illustrating a configuration of an electronic device according to an embodiment of the disclosure.

[0032] Figure 18b FIG. 4 is a block diagram illustrating a configuration of an electronic device according to an embodiment of the disclosure.

[0033] Figure 19a FIG. 5 is a block diagram illustrating a configuration of a server according to an embodiment of the disclosure.

[0034] Figure 19b FIG. 6 is a block diagram illustrating a configuration of a server according to an embodiment of the disclosure. DETAILED DESCRIPTION

[0035] The terms used herein will be described briefly, and the disclosure will be described in detail. Throughout the disclosure, the expression "at least one of a, b, or c" indicates only a, only b, only c, both a and b, both a and c, both b and c, all of a, b, and c, or variations thereof.

[0036] All terms used herein including descriptive or technical terms are to be interpreted, as apparent to one of ordinary skill in the art, with the meaning consistent with the intention of the ordinary skilled in the art. However, the terms can have different meanings according to the intention of the ordinary skilled in the art, judicial precedents, or appearance of new technologies. Also, some terms used herein can be arbitrarily selected by the applicant, and in this case, the terms are defined in detail below. Therefore, the specific terms used herein are to be defined based on their unique meanings and the entire context of the disclosure.

[0037] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. The terms used herein, including technical or scientific terms, can have the same meanings as those generally understood by one of ordinary skill in the art to which the disclosure pertains. It will be understood that, although the terms "first," "second," etc. can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another.

[0038] It will be understood that when a certain part "includes" a certain component, the part does not exclude another component but can further include another component, unless the context clearly dictates otherwise. Further, the term such as "unit" or "module" used in the embodiments indicates a unit for processing at least one function or operation and can be implemented in hardware, software, or a combination of hardware and software.

[0039] The disclosure relates to a method of generating a personalized image by reflecting an image owned by a user (e.g., an image stored in a device of the user such as a smartphone) when generating an image according to a request of the user using a generative model. In an embodiment, a second image generated by reflecting a first image can refer to a second image generated with respect to, in consideration of, or based on the first image. According to an embodiment of the disclosure, a synthetic image can be obtained by reflecting an image that is not used to train a generative model during inference (e.g., image generation) performed by the generative model.

[0040] In the disclosure, the term "personalized image" can refer to synthetic data (e.g., a synthetic image) generated using a generative model according to an image generation request of a user. Further, in the disclosure, the term "personalized image" can refer to an image generated using a generative model by reflecting an image owned by a user. For example, a personalized image can be generated to include a personalized feature (e.g., a specific object or image style) of a user. The personalized image can be referred to as a "customized image."

[0041] In the disclosure, the term "generative artificial intelligence (AI)" can refer to an artificial intelligence technology capable of generating new text or images in response to input data (e.g., text or images). The term "generative model" can refer to a neural network model for implementing a generative AI technology. The generative model can be trained with respect to patterns and structures of training data to generate new data having similar features to input data or new data corresponding to input data. For example, the generative model can generate images by using an image-to-image method or a text-to-image method.

[0042] In the disclosure, the term "reference image" can refer to data to be referred to in a process of generating a personalized image. The reference image can include an element that can represent a characteristic of a user (e.g., a person of the user or around the user, an object owned by the user, a space around the user, etc.).

[0043] The disclosure will be described more fully with reference to the accompanying drawings such that one of ordinary skill in the art can perform the disclosure without difficulty. However, the disclosure can be implemented in various different forms and is not limited to the embodiments described herein. Also, in the drawings, parts unrelated to the description are omitted for the sake of clarity in describing the disclosure, and the same reference numerals denote the same elements throughout the specification.

[0044] The present disclosure will now be described with reference to the accompanying drawings.

[0045] Figure 1 FIG. 1 is a diagram schematically illustrating an electronic device according to an embodiment of the present disclosure.

[0046] In an embodiment of the present disclosure, the electronic device 2000 and / or the server 3000 can provide an image personalization service for generating an image by using a generative model to the user 100. The image personalization technology of the present disclosure allows the user 100 to generate a new image including his / her personalized features by using an image(s) owned by the user 100, which is synthetic data.

[0047] For example, when the user 100 inputs a prompt of "draw a puppy" requesting the electronic device 2000 to generate an image, the electronic device 2000 can search for and retrieve or otherwise obtain one or more images 120 related to the keyword "puppy" from a personal image set 110 stored in a media storage of the electronic device 2000. The personal image set 110 can include personal images owned by the user that are not used for training the generative model.

[0048] The electronic device 2000 can generate data for image personalization by processing the retrieved one or more images 120. The data for image personalization can refer to embedding data obtained by converting the one or more images 120 into a vector representation. The data for image personalization generated from the images owned by the user is transmitted to the server 3000, and when an inference operation using the generative model is performed, the server 3000 can generate a personalized image 130 by using the data for image personalization received from the electronic device 2000.

[0049] The server 3000 can transmit the personalized image 130 to the electronic device 2000, and the electronic device 2000 can provide the personalized image 130 to the user 100. The personalized image 130 can include features of the one or more images 120 retrieved from the electronic device 2000. For example, based on the user's request of "draw a puppy", the user can be provided with a personalized image 130 that draws a puppy (e.g., a puppy owned by the user) included in the images owned by the user 1000.

[0050] In an embodiment of the disclosure, the server 3000 can provide an image personalization service to a plurality of different users. When the server 3000 provides an image personalization to a plurality of users, the generation model does not have to be additionally trained with respect to the personal image set 110 of each of the plurality of users. In other words, the server 3000 does not provide each user with a generation model that is custom-trained with respect to the personalization features of each user. Instead, a personalized image 130 reflecting features included only in the personal image of the user can be generated using a deployed generation model that is generally available to other users. To this end, the electronic device 2000 generates data for image personalization and provides the data for image personalization to the server 3000. The phrase "model has been deployed" means that the model has been migrated to an actual operating environment so that an inference operation can be performed, not an environment in which the model is trained (e.g., pre-trained and / or fine-tuned / trained) or tested. In other words, when a model is described as "deployed", this means that the model is ready to provide a service after completing training and completing performance verification. However, because a service can be provided to individuals and / or a limited group, deployment does not necessarily mean that something should be publicly released.

[0051] In the disclosure, an image owned by a user used as data for image personalization reference can be referred to as a "reference image", and data for image personalization generated using the "reference image" is referred to as "image feature information".

[0052] In an embodiment of the disclosure, the electronic device 2000 can be any one of various types of devices for providing a personalized image 130 to a user 100. For example, the electronic device 2000 can be implemented as any one of various types and forms of electronic devices including a display. Examples of the electronic device 2000 can include devices capable of displaying a personalized image 130 through a display, such as but not limited to a smart TV, a smart phone, a tablet PC, a laptop PC, a glasses-type display, and a head-mounted display (HMD). For example, the electronic device 2000 can be implemented as any one of various types and forms of electronic devices that can be connected to a display through a wire or wirelessly. Examples of the electronic device 2000 can include devices that are connected to a display through a wire or wirelessly and are capable of displaying a personalized image 130, such as but not limited to a set-top box and a desktop PC.

[0053] In an embodiment of the disclosure, the server 3000 can be a device for generating a personalized image 130 by using a generative model. The server 3000 can be a device capable of performing complex calculations and tasks such as training, inference, management, and deployment of a generative model using large-scale data. According to an embodiment of the disclosure, the training of the generative model performed in the server 3000 can be performed by another computing device. The server 3000 can receive data for personalized image generation from the electronic device 2000 as a client device, and can return the personalized image 130 to the electronic device 2000.

[0054] In an embodiment of the disclosure, the electronic device 2000 and the server 3000 for providing an image personalization service can be referred to as an "image personalization system." Alternatively, for convenience, the electronic device 2000 and the server 3000 for providing an image personalization service can be simply referred to as a "system." In the disclosure, an embodiment in which a personalized image is generated by the image personalization system and provided to a user will be described. However, the embodiment is not limited thereto, and the operations described herein can or can not be performed by the image personalization system.

[0055] In an embodiment of the disclosure, the image personalization operation can be independently performed by the electronic device 2000. In this case, the generative model for generating the personalized image 130 can be stored in the electronic device 2000. The electronic device 2000 can generate the personalized image 130 by performing the operations of the disclosure using the generative model. Because the electronic device 2000 has lower computing performance than the server 3000, the generative model used by the electronic device 2000 can be a relatively lightweight AI model having resource requirements that can be tailored to the computing performance of the electronic device 2000.

[0056] In an embodiment of the disclosure, the image personalization operation can be independently performed by the server 3000. In this case, the server 3000 can receive only an image personalization request from the electronic device 2000, can generate the personalized image 130 by performing the operations of the disclosure, and can return the generated result to the electronic device 2000.

[0057] In an embodiment of the disclosure, the image personalization operation can be performed by a plurality of electronic devices, respectively. For example, in the operations of the disclosure, "Operation A" can be performed by "Electronic Device A," and "Operation B" can be performed by "Electronic Device B." Examples of various methods of the disclosure such as those described above are obvious, and thus, for the sake of brevity, the description thereof will be omitted in the disclosure.

[0058] Examples of the operation of the electronic device 2000 to provide the personalized image 130 to the user 100 will be described in more detail below in conjunction with the following drawings and their descriptions.

[0059] Figure 2 is a flowchart for describing an operation of an electronic device generating a personalized image according to an embodiment of the disclosure.

[0060] Referring to Figure 2 An operation performed by the electronic device 2000 to generate a personalized image will be briefly described, and the operation will be described in detail with reference to subsequent drawings.

[0061] In operation S210, the electronic device 2000 can obtain a prompt for image generation.

[0062] The term "prompt" can refer to an input for starting or initiating an interaction with a generation model for generating an image. The prompt can be a text input or a voice input including one or more words and / or one or more sentences.

[0063] In an embodiment of the disclosure, the prompt can include natural language text. The natural language text can include various information, such as context, intent, task, and constraint, which the generation model can use to generate an image. The electronic device 2000 can process the natural language using a natural language processing (NLP) model.

[0064] In an embodiment of the disclosure, the prompt can be received as a user input. For example, the electronic device 2000 can receive a text input of a user. Alternatively, the electronic device 2000 can receive a voice input of a user. The electronic device 2000 can convert the voice input of the user into text by using automatic speech recognition (ASR).

[0065] In an embodiment of the disclosure, the prompt can be generated by the electronic device 2000. For example, the electronic device 2000 can receive an image input from a user. The electronic device 2000 can extract a textual description of the image from the image. The electronic device 2000 can generate the prompt using any one of various methods for extracting text from an image (e.g., captioning an image).

[0066] The prompt can correspond to various expressions that can represent the same / similar concept. For example, the prompt can correspond to expressions such as "user input," "input phrase," "user command," "instruction," "starting sentence," "task query," "trigger sentence," and "message," but embodiments are not limited thereto.

[0067] In operation S220, the electronic device 2000 can extract a keyword from the prompt.

[0068] The keyword can include a word, a phrase, or a clause indicating a necessary part of the prompt. At least one of the extracted keywords can be a keyword for image personalization. When there is one or more keywords, the keyword for image personalization can be selected from the one or more keywords. The keyword for image personalization can correspond to a personalized element reflected in the personalized image. For example, when the keyword for image personalization is "A", the personalization can be applied to an element corresponding to "A" among elements in the personalized image. There can be one or more keywords for image personalization. In the disclosure, the keyword for image personalization can also be referred to as a "personalization keyword". Alternatively, the keyword for image personalization can be simply referred to as a "keyword".

[0069] In an embodiment of the disclosure, the electronic device 2000 can parse the prompt to decompose the text into a plurality of components and extract one or more keywords from the plurality of components.

[0070] In an embodiment of the disclosure, the electronic device 2000 can extract one or more keywords from the prompt using an NLP model. For example, the electronic device 2000 can analyze a sentence using a pre-trained language model, can determine the meaning and context of the words in the sentence, and can extract one or more keywords.

[0071] The electronic device 2000 can determine a personalization keyword from the one or more keywords extracted from the prompt.

[0072] In an embodiment of the disclosure, the electronic device 2000 can determine the one or more extracted keywords as the personalization keyword. For example, the electronic device 2000 can determine all or some of the one or more extracted keywords as the personalization keyword. In this case, the method by which the electronic device 2000 determines the personalization keyword can be a rule-based method in which the personalization keyword is pre-defined (e.g., predetermined / registered / recently used / frequently used / context analysis-based determination), but the disclosure is not limited thereto.

[0073] In an embodiment of the disclosure, the electronic device 2000 can determine the personalization keyword based on a user input. For example, the electronic device 2000 can display the one or more keywords extracted from the prompt on a screen and can receive a user input selecting at least one of the one or more keywords displayed on the screen. The electronic device 2000 can determine the keyword corresponding to the user input as the personalization keyword.

[0074] The term "keyword" can correspond to various expressions that can represent the same / similar concept. For example, the term "keyword" can correspond to expressions such as "word", "item", "main word", "search item", "core item", "main item", "relevant item", and "recommendation", but embodiments are not limited thereto.

[0075] In operation S230, the electronic device 2000 can search for one or more reference images corresponding to the keyword from among a plurality of images stored in the electronic device 2000. The operation of the electronic device 2000 to search for one or more reference images corresponding to the keyword can include a process in which the electronic device 2000 searches for one or more images corresponding to the keyword and selects reference images from among the searched and retrieved one or more images.

[0076] In an embodiment of the disclosure, the electronic device 2000 can search for one or more images in the electronic device 2000 based on the keyword. For example, the electronic device 2000 can search for one or more images matching the keyword. Alternatively, the electronic device 2000 can search for one or more images from among a plurality of images stored in the electronic device 2000 by referring to the keyword. For example, the electronic device 2000 can search for one or more images matching a word (e.g., a synonym or a similar word) related to the keyword.

[0077] The electronic device 2000 can search for one or more images corresponding to the keyword from among a plurality of images stored in a local media storage of the electronic device 2000. That is, the electronic device 2000 does not search for images corresponding to the keyword from among images posted to a public domain, but can select, retrieve, or otherwise obtain one or more images corresponding to the keyword from among a plurality of personal images of the user stored in the local media storage of the electronic device 2000. The images stored in the local media storage can include images directly captured by the user and / or images obtained by the user from an external source (e.g., a downloaded image among images posted to a public domain and an image received from another user). In an embodiment of the disclosure, the electronic device 2000 can find images matching the keyword and / or images matching a synonym or a similar word of the keyword from among various images posted to a public domain through an external database (e.g., a web).

[0078] In an embodiment of the disclosure, the electronic device 2000 can search for an image corresponding to a keyword using image metadata. The image metadata can include, but is not limited to, a title, a description, an image capture location, an image capture date and time, person information, a tag, etc. of the image. The image metadata can include information about various objects included in the image. When the electronic device 2000 classifies the image or detects objects in the image and performs object recognition, information about various objects included in the image metadata can be obtained. When the electronic device 2000 takes a photo, the image metadata can be automatically generated. Alternatively, the image metadata can be manually input by a user. For example, the electronic device 2000 can provide an interface through which a user can input additional information about an image at the time of capturing the image or after capturing the image. In this case, the interface through which additional information about an image can be input can be provided through any one of various applications, for example, a camera application, a gallery application, and a photo editing application.

[0079] In an embodiment of the disclosure, the electronic device 2000 can generate image metadata. The electronic device 2000 can generate the image metadata using various image processing / analysis algorithms. For example, the electronic device 2000 can cluster images according to a theme, and can label the theme of the images in the image cluster with metadata. Alternatively, the electronic device 2000 can extract a text description of an image from the image, can extract one or more keywords from the text description, and can label the extracted keywords with metadata. Alternatively, the electronic device 2000 can perform image classification, and can label the class of the image as a result of the classification with metadata. However, the method by which the electronic device 2000 generates image metadata is not limited to the above-described examples.

[0080] In an embodiment of the disclosure, the electronic device 2000 can select a reference image from one or more images obtained.

[0081] The term "reference image" can be an image selected from the obtained images corresponding to a keyword, and refers to an image that can be used for image personalization. For example, data indicating the reference image (for example, image feature information described below) can be used as input data for a generation model, and thus, the features of the reference image can be included in a personalized image generated by the generation model. For example, various features of the reference image, such as a background, an object, a style, a theme, and a context, can be included in the personalized image. According to an embodiment, one or more reference images can be selected.

[0082] The term "reference image" can correspond to various expressions that can represent the same / similar concept. For example, the term "reference image" can correspond to expressions such as "image," "photo," "reference photo," "standard image / photo," and "benchmark image / photo," but embodiments are not limited thereto.

[0083] In an embodiment of the disclosure, the electronic device 2000 can generate a reference image based on a pre-defined or predetermined condition. For example, the electronic device 2000 can select an image having a similarity to the keyword equal to or greater than a threshold value from among the one or more images taken corresponding to the keyword as the reference image. However, the condition for selecting the reference image is not limited to the above-described example.

[0084] In an embodiment of the disclosure, the electronic device 2000 can determine or select a reference image based on a user input. For example, the electronic device 2000 can display the one or more images taken corresponding to the keyword on a screen, and can receive a user input selecting at least one of the one or more images displayed on the screen. The electronic device 2000 can determine the image corresponding to the user input as the reference image.

[0085] In an embodiment of the disclosure, when the reference image is determined, the keyword and the reference image corresponding to the keyword can be stored as a text-image pair.

[0086] In operation S240, the electronic device 2000 can segment a region related to the keyword within the reference image. In other words, the electronic device 2000 can segment the reference image into a plurality of regions. The electronic device 2000 can determine a segment corresponding to a region related to the keyword from segments of the reference image based on the keyword. For example, the electronic device 2000 can select a segment matching the keyword from the segments of the reference image. The electronic device 2000 can determine a segment related to the keyword from the segments of the reference image by referring to the keyword. For example, the electronic device 2000 can select a segment matching a word (e.g., a synonym or a similar word) related to the keyword.

[0087] The region related to the keyword can be a region in the image indicating a visual feature of the keyword. For example, when the keyword is "object A", the region related to the keyword can refer to a region in the image corresponding to "object A". The electronic device 2000 can perform image segmentation to segment the image from the reference image into a plurality of regions based on objects, foregrounds, backgrounds, etc. included in the image. In this case, the region related to the keyword can be selected from among the plurality of segmented regions.

[0088] In an embodiment of the disclosure, the electronic device 2000 can determine a region related to the keyword by analyzing the reference image. For example, the electronic device 2000 can separate a plurality of semantic elements in the image using semantic segmentation. The electronic device 2000 can classify the plurality of semantic elements, and can select a segment indicating a region related to the keyword based on a classification result.

[0089] In an embodiment of the disclosure, the electronic device 2000 can determine a region related to a keyword based on a user input. For example, the electronic device 2000 can separate a plurality of instances (e.g., objects) in an image using instance segmentation. The electronic device 2000 can display a segmentation result on a screen, and can receive a user input selecting at least one of the segments displayed on the screen.

[0090] The image segmentation method performed by the electronic device 2000 is not limited to the above example. For example, the electronic device 2000 can use a panoptic segmentation method that is a combination of semantic segmentation and instance segmentation. Alternatively, the electronic device 2000 can select a segment indicating a region related to a keyword based on a user input (e.g., a rectangular selection or a free selection (lasso selection)) that specifies a region of an object corresponding to a keyword in a reference image.

[0091] In operation S250, the electronic device 2000 can obtain image feature information based on a region related to a keyword of a reference image.

[0092] In the disclosure, data for reflecting a personalized element when generating an image can be referred to as "reference data." For example, the reference data can include text and / or an image. In detail, examples of the text included in the reference data can include, but are not limited to, a text description of an image, a keyword, a synonym of a keyword, and a similar word of a keyword. In addition, examples of the image included in the reference data can include, but are not limited to, an original reference image and a region (e.g., a segment) of a reference image related to a keyword. The term "reference data" can correspond to various expressions that can represent the same / similar concept. In an embodiment, the reference data can be included in data that can be referred to as "data for personalization."

[0093] In an embodiment of the disclosure, the electronic device 2000 can generate image feature information by pre-processing the reference data. The image feature information can include various data related to features of an image, such as, but not limited to, a visual pattern, an attribute, a structure, or text corresponding to an image extracted from an image.

[0094] The image feature information can be used as an input of a generation model. For example, the image feature information can be obtained through certain pre-processing processes that allow an image and / or text of the reference data to be used as an input of a generation model.

[0095] In an embodiment of the disclosure, the image feature information can be obtained through a vector embedding, which is a preprocessing method of converting data into a vector representation. The electronic device 2000 can generate a reference embedding by converting the reference data into a vector representation. The reference embedding refers to data obtained by converting the reference data into a vector representation. For example, the electronic device 2000 can convert a region of a reference image related to a keyword into data in a vector representation. Also, for example, the electronic device 2000 can convert a text-image pair including a keyword and a reference image into data in a vector representation. Also, for example, the electronic device 2000 can convert a text-image pair including a keyword and a region of a reference image related to the keyword into data in a vector representation. The electronic device 2000 can generate a reference embedding having a plurality of vector values by converting the keyword and the reference image into a vector representation including normalized values. The electronic device 2000 can perform tokenization for dividing data into small unit elements, and can convert each token into a unique vector value. The reference embedding can include information about the keyword and the reference image, and can be used as input data when a model generates a personalized image.

[0096] The term "reference embedding" can correspond to various expressions that can represent the same / similar concept. For example, the term "reference embedding" can correspond to expressions such as "vector," "embedding," "vector embedding," "vector representation," or "vector encoding," but embodiments are not limited thereto. In an embodiment, the reference embedding can be included in data that can be referred to as "data for personalization."

[0097] In an embodiment of the disclosure, the electronic device 2000 can generate a reference embedding by encoding a text-image pair of a keyword and a reference image corresponding to the keyword using an encoder. The encoder can be trained to receive text and / or an image and compress information of the text and / or the image.

[0098] The preprocessing method through which the electronic device 2000 obtains the image feature information is not limited to a vector embedding that converts data into a vector representation. For example, the electronic device 2000 can perform preprocessing such as tokenization or normalization on text, and can perform preprocessing such as normalization, resizing, cropping, or image attribute change (brightness, contrast, or color) on an image. That is, the electronic device 2000 can obtain the image feature information by processing the reference data using any one of various methods.

[0099] Because there is a limitation to list all cases in which the electronic device 2000 obtains image feature information, it will be assumed that a reference embedding obtained by converting reference data into a vector is used as an example of image feature information to describe the present disclosure. However, the vector embedding is an example for describing a technical implementation method of the present disclosure. The present disclosure can also be implemented by using a similar method such as the pre-processing methods listed in the above examples, including using a vector embedding.

[0100] In operation S260, the electronic device 2000 can transmit the prompt and the image feature information to the server 3000.

[0101] The server 3000 can operate a generation model. To make the server 3000 operate the generation model, the server 3000 can be a device having relatively high computing performance compared to the electronic device 2000, which can allow the server 3000 to perform more computations than the electronic device 2000. The server 3000 can perform training and / or inference of the generation model that requires a large amount of computation. The server 3000 can generate a personalized image using the generation model, and can transmit the personalized image to the electronic device 2000.

[0102] In an embodiment of the present disclosure, the generation model can be a version deployed after completing training of the generation model. The server 3000 can generate a "personalized image" reflecting features of images owned by a user using the deployed generation model without retraining the generation model. For example, when the server 3000 generates a personalized image, existing weights of the deployed generation model are used, and only the image feature information (e.g., a reference embedding) is used as input data during an inference operation using the generation model. Therefore, even when the generation model is not trained with respect to the user's personalized features (e.g., images owned by the user), the generation model can generate a personalized image including features of a reference image that is a personal image of the user.

[0103] In an embodiment of the present disclosure, the generation model can generate a prompt embedding by encoding a prompt received from the electronic device 2000. Alternatively, when the reference embedding is generated in operation S250, the prompt embedding can be generated by the electronic device 2000. The generation model generates a personalized image using the prompt embedding and the reference embedding.

[0104] The electronic device 2000 can provide various advantages by transmitting image feature information including preprocessed data to the server 3000 without transmitting reference data to the server 3000. For example, the electronic device 2000 can reduce the amount of data processing of the server 3000 by transmitting, to the server 3000, reference embeddings obtained by converting reference images into vector representations. Also, by transmitting only the reference embeddings to the server 3000, the electronic device 2000 can prevent images stored only in a media storage of the electronic device 2000 from being shared, thereby protecting user privacy while generating a personalized image.

[0105] In some embodiments of the disclosure, the electronic device 2000 can transmit the reference data to the server 3000. In this case, the generation model can receive the reference data as input, can perform preprocessing, and can generate a personalized image. Alternatively, the generation model can receive the preprocessed reference data as input, and can generate a personalized image.

[0106] In operation S270, the electronic device 2000 can receive the personalized image from the server 3000. There can be one or more personalized images.

[0107] The electronic device 2000 can display the personalized image received from the server 3000.

[0108] Figure 3 FIG. 2 is a diagram illustrating an operation of an electronic device and a server according to an embodiment of the disclosure.

[0109] Referring to Figure 3 The electronic device 2000 can include at least a keyword extraction module 310, an image search module 320, an image segmentation module 330, and an encoder 340. The server 3000 can include an image generation module 350. Figure 3 The modules illustrated in FIG. 3 can be elements implemented when at least one processor included in the electronic device 2000 executes programs or instructions stored in a memory included in the electronic device 2000. Accordingly, operations performed by the modules included in the electronic device 2000 described below can be performed by at least one processor included in the electronic device 2000.

[0110] The keyword extraction module 310 extracts one or more keywords by processing a hint. When one or more keywords are extracted, all or some of the extracted one or more keywords can be determined or selected as a personalized keyword. There can be one or more personalized keywords.

[0111] The keywords can be transmitted to the encoder 340 to be converted into vectors, can be transmitted to the image search module 320 to be used for searching images, and can be transmitted to the image segmentation module to be used for segmenting images. Referring to Figure 4Further examples of the keyword extraction module 310 are described further.

[0112] The image search module 320 searches one or more images corresponding to the keyword from among images stored in the media storage 322. The image search module 320 can search one or more images corresponding to the keyword using image metadata. When one or more images are searched, a reference image to be reflected in the personalized image can be determined or selected from among the searched and retrieved one or more images. There can be one or more reference images.

[0113] The reference image can be selected by the electronic device 2000 based on a predetermined condition, and can be selected based on a user input. The reference image can be transmitted to the image segmentation module 330 to be segmented, and can be transmitted to the encoder 340 to be converted into a vector. Referring to Figure 5a - Figure 5b Further examples of the image search module 320 are described further.

[0114] The image segmentation module 330 segments and determines a region related to the keyword within the reference image. There can be one or more regions related to the keyword.

[0115] The region related to the keyword can be determined based on a result (e.g., semantic segmentation) obtained after the electronic device 2000 analyzes the keyword and the reference image, and the region to be segmented can be selected based on a user input. The image segmentation module 330 can obtain a segment of the reference image by extracting a region related to the keyword. The segment of the reference image can be transmitted to the encoder 340 to be converted into a vector. Referring to Figure 4 Further examples of the image segmentation module 330 are described further.

[0116] In an embodiment of the disclosure, the operation of the image segmentation module 330 can be omitted. In this case, the operation of using the segment of the reference image can be replaced with an operation of using the reference image.

[0117] The encoder 340 converts image and / or text data into a vector representation.

[0118] For example, the encoder 340 can generate a prompt embedding by converting a prompt into a vector representation. Alternatively, the prompt can be transmitted to the server 3000 as an original form, and can be encoded by the server 3000.

[0119] For example, the encoder 340 can generate a reference embedding by converting a text-image pair including a keyword and a reference image into a vector representation.

[0120] For example, the encoder 340 can generate a reference embedding by converting a text-image pair including a keyword and a segment of a reference image into a vector representation.

[0121] The prompt (or prompt embedding) and the reference embedding can be transmitted to the server 3000 so that the server 3000 generates a personalized image using the prompt and the reference embedding. Referring to Figure 7 An example of the encoder 340 is further described.

[0122] The image generation module 350 generates a personalized image. For the personalized image, a generative model can be used. The generative model can be an AI model that uses a mechanism to process prompt-image correlation (e.g., cross-attention) to generate an image based on a prompt (or prompt embedding).

[0123] In an embodiment of the disclosure, the generative model can be a version deployed after training is completed. In this case, the reference embedding received from the electronic device 2000 can be used only during an inference operation of the generative model. That is, the generative model is an AI model that receives a prompt (or prompt embedding) and a reference embedding as input and outputs a personalized image.

[0124] The server 3000 can generate a "personalized image" reflecting only the features of an image owned by a user using the image generation module 350 without retraining the generative model.

[0125] Referring to Figure 8 An example of the image generation module 350 is further described.

[0126] Figure 4 is a diagram for describing an operation of an electronic device extracting a keyword according to an embodiment of the disclosure.

[0127] In an embodiment of the disclosure, the electronic device 2000 can display the extracted one or more keywords 420. As Figure 4 As shown, a graphic interface 400 (hereinafter referred to as graphic interface 400) showing a result obtained after the electronic device 2000 extracts a keyword is shown. Various elements can be included in the graphic interface 400. For example, a prompt 410, one or more keywords 420, and various interface elements for guiding image personalization (e.g., text 430 and buttons 440) can be included in the graphic interface 400. However, some of the above elements can be omitted or other elements can be added.

[0128] In an embodiment of the disclosure, a prompt 410 for generating a personalized image can be displayed on the graphic interface 400. For example, a prompt 410 input by a user saying "draw a picture of a puppy playing in a park" can be included in the graphic interface 400. The electronic device 2000 can display the prompt 410 on the graphic interface 400 so that the user checks whether the prompt 410 is correctly input. The electronic device 2000 can determine the prompt 410 based on the user input of checking the prompt 410. Alternatively, the electronic device 2000 can change the prompt 410 based on the user input for changing the prompt 410.

[0129] In an embodiment of the disclosure, one or more keywords 420 extracted can be displayed on the graphic interface 400. For example, keywords 420 such as "puppy," "park," and "picture" extracted from the prompt 410 can be included in the graphic interface 400.

[0130] The electronic device 2000 can parse the prompt to decompose text into a plurality of components, and can extract one or more keywords 420 from the plurality of components. For example, the electronic device 2000 can perform text preprocessing for removing special characters, etc. in a sentence. The electronic device 2000 can create a list of words such as nouns, adjectives, and verbs, and can select at least some of the words. In this case, a word frequency, a predetermined word importance, etc. can be used, but embodiments are not limited thereto.

[0131] The electronic device 2000 can extract one or more keywords 420 from the prompt 410 using an NLP model. For example, the electronic device 2000 can analyze a sentence using a pre-trained language model. For example, the electronic device 2000 can tokenize a sentence, and can perform semantic analysis to determine the meaning, context, and intent of the words in the sentence. The electronic device 2000 can extract one or more keywords 420 based on the sentence analysis result.

[0132] In an embodiment of the disclosure, the electronic device 2000 can determine a personalized keyword 422 among the one or more keywords 420 extracted from the prompt 410. Since the personalized keyword 422 can be all or some of the one or more keywords 420, the personalized keyword 422 can be included in the one or more keywords 420. The personalized keyword 422 can be determined automatically by the electronic device 2000, or can be determined manually based on user input.

[0133] In an embodiment of the disclosure, the electronic device 2000 can change the one or more keywords 420 based on user input. For example, when the user selects an interface element 424 for keyword change in the graphic interface 400, the word "play" that was not previously extracted as a keyword can be added to the one or more keywords 420 and displayed.

[0134] In an embodiment of the disclosure, various interface elements (e.g., text 430 and buttons 440) for guiding image personalization can be displayed on the graphic interface 400. For example, text 430 asking what point the user wants to focus on to generate a personalized image (e.g., "What do you want to change?") can be displayed on the graphic interface 400. Alternatively, buttons 440 allowing the user to proceed to the next step for image personalization (e.g., an "image generation" button) can be displayed on the graphic interface 400.

[0135] In an embodiment of the disclosure, the electronic device 2000 can suggest keywords related to the prompt. For example, the electronic device 2000 can recommend words not included in the prompt 410 but highly related to the prompt to be included in the one or more keywords 420. Alternatively, the electronic device 2000 can recommend keywords previously used by the user based on the user's personalized image generation history to be included in the one or more keywords 420.

[0136] Figure 5a FIG. 1 is a diagram for describing an operation of an electronic device searching for an image according to an embodiment of the disclosure.

[0137] In an embodiment of the disclosure, the electronic device 2000 can search for one or more images corresponding to the keywords from the user's images stored in the local media storage.

[0138] In an embodiment of the disclosure, the images stored in the media storage can include metadata. The image metadata can include, but is not limited to, tiles, descriptions, image capture locations, image capture dates and times, person information, and tags of the image. The image metadata can include information about various objects included in the image. When the electronic device 2000 classifies images, detects objects in the images, and performs object recognition, information about various objects included in the image metadata can be obtained. When the electronic device 2000 takes a photo, the image metadata can be automatically generated. Alternatively, the image metadata can be generated by the electronic device 2000 by analyzing the stored images. Alternatively, the image metadata can be manually input by the user.

[0139] For convenience of explanation, the images stored in the media storage will be described with reference to the application's screen 510 on which the images stored in the media storage are visually output. The application's screen 510 can be displayed on the screen of the electronic device 2000.

[0140] In an embodiment of the disclosure, the images stored in the media storage of the electronic device 2000 can be pre-arranged for image search. For example, the electronic device 2000 can group the images stored in the media storage by theme through image clustering.

[0141] In more detail, for example, the images stored in the media storage can be divided into a plurality of image groups 514 (#person, #place, #my pet, #OOTD (which can refer to outfit of the day)).

[0142] Further, each of the plurality of image groups 514 can be divided into subgroups. For example, the person group 516 can include subgroups respectively corresponding to persons (e.g., Chris, Joy, Tom, and John). Images of the persons can be respectively included in the subgroups. Alternatively, the place group 518 can include subgroups respectively corresponding to places (e.g., Seoul, Jeju Island, Busan, and Paris). Images of the places can be respectively included in the subgroups.

[0143] When the electronic device 2000 searches for images corresponding to the keyword, the electronic device 2000 can use information (e.g., image group information) that is pre-arranged based on metadata. However, the disclosure is not limited thereto, and the electronic device 2000 can search for images corresponding to the keyword by using various methods.

[0144] For example, the electronic device 2000 can search for images corresponding to the keyword using original metadata. The original metadata can include metadata generated by the electronic device 2000 or added by a user. For example, category information obtained by the electronic device 2000 through image analysis (e.g., captioning an image, image classification, or object recognition) can be included in the metadata. Alternatively, information directly tagged by a user for an image can be included in the metadata.

[0145] In an embodiment of the disclosure, the electronic device 2000 can display one or more images 522 searched and retrieved. For example, the electronic device 2000 can display one or more images 522 corresponding to the keyword "puppy."

[0146] A reference image 530 can be determined or selected from the one or more retrieved images 522. For example, the electronic device 2000 can determine the reference image 530 based on a predetermined condition. For example, the electronic device 2000 can select an image having a similarity to the keyword equal to or greater than a threshold value from the one or more retrieved images 522 as the reference image 530. However, the predetermined condition of the electronic device 2000 is not limited to the above-described example. Alternatively, the electronic device 2000 can receive a user input selecting at least one of the one or more images 522 displayed on the screen. The electronic device 2000 can determine an image corresponding to the user input as the reference image 530.

[0147] In an embodiment of the disclosure, there can be one or more reference images 530. When a plurality of reference images 530 are selected, the electronic device 2000 can apply a weight to each reference image 530 based on a user input.

[0148] Figure 5b FIG. 4 is a diagram for describing an operation of an electronic device searching for an image according to an embodiment of the disclosure.

[0149] The electronic device 2000 can search for one or more images corresponding to the keyword inside and / or outside the electronic device 2000 using the image search module. In an embodiment of the disclosure, the electronic device 2000 can determine a database in which an image is to be searched. For example, the electronic device 2000 can search for an image in at least one of the media storage 540 or the external database 550 based on a user input.

[0150] In an embodiment of the disclosure, the electronic device 2000 can search for one or more images corresponding to the keyword from among images stored in the media storage 540, which can be or can include an image database. The operation has been described with reference to FIG. 3, and thus a repeated description will be omitted for the sake of brevity. Figure 5a The operation has been described, and thus a repeated description will be omitted for the sake of brevity.

[0151] In an embodiment of the disclosure, the electronic device 2000 can search for one or more images corresponding to the keyword from among images stored in the external database 550. For example, the electronic device 2000 can search for images matching the keyword and / or images matching synonyms or similar words of the keyword from among various images shared on a website using a search engine. The electronic device 2000 can search for one or more images corresponding to the keyword in a web page or an image sharing platform supporting image search. The electronic device 2000 can provide the image search result to the user. For example, the electronic device 2000 can display the image search result on a screen.

[0152] The electronic device 2000 can determine or select a reference image from among the one or more images searched for and retrieved. The operation has been described with reference to FIG. 3, and thus a repeated description will be omitted for the sake of brevity. Figure 5a The operation has been described, and thus a repeated description will be omitted for the sake of brevity.

[0153] In an embodiment of the disclosure, the electronic device 2000 can use the user information 560 when performing the image search. The user information 560 can include various items indicating information related to a user. The user information can include, for example, and without limitation, profile information, preference and interest information, history and activity record information.

[0154] The profile information can include basic information about a user. For example, the user profile can include, but is not limited to, a name, a gender, a date of birth, a residence, a current location of the user, a travel route, and a profile photo. The profile information can be obtained through a user input.

[0155] The preference and interest information can include various information related to the user's personal tastes and interests in various categories. In an embodiment of the disclosure, the preference and interest information can be obtained through user input. In an embodiment of the disclosure, the preference and interest information can be collected by the electronic device 2000. For example, the electronic device 2000 can collect and store information related to the user's preferences and interests, such as the content that the user clicks the "like" button in the SNS service, and the emoticon, mood symbol, and word that the user frequently uses in the message service.

[0156] The history and activity record information can include various information related to the user's use of the electronic device 2000. For example, the history and activity record information can include, but is not limited to, a search history, a purchase history, a service use history, a website access history, and an application history.

[0157] The electronic device 2000 can use the user information 560 when searching for one or more images related to the keyword from the media storage 540 and / or the external database 550. The electronic device 2000 can refine and process the data in the user information 560 into a form that can be used for image search.

[0158] The electronic device 2000 can search for one or more images related to the keyword and can filter images that match or are similar to the user's preferences and interests. For example, the electronic device 2000 can filter images that match or are similar to the user's preferences and interests based on the keyword, content category, color, and style related to the user's interests. Accordingly, images suitable for the user's preferences can be included in the image search result.

[0159] The electronic device 2000 can determine or select a reference image from one or more images searched and obtained. As described above with reference to FIG. 6, the reference image can be an image that is determined to be most suitable for the user's preferences and interests. Figure 5a The operation is described, and thus a detailed description will be omitted for the sake of brevity.

[0160] Figure 6 is a diagram for describing an operation of the electronic device segmenting an image according to an embodiment of the disclosure.

[0161] In an embodiment of the disclosure, the electronic device 2000 can segment a region related to the keyword within the reference image 610.

[0162] The region related to the keyword can be a region in the image indicating a visual feature of the keyword. For example, when the keyword is "object A," the region related to the keyword refers to a region in the image corresponding to "object A." For example, when the keyword is "puppy," the region refers to a region in the image corresponding to "puppy."

[0163] In an embodiment of the disclosure, the electronic device 2000 can segment the image from the reference image 610 into a plurality of regions based on objects, foreground, and background included in the image by performing image segmentation. For example, when the image segmentation is performed, the reference image 610 can be divided into a first region 611 corresponding to an object (e.g., a person), a second region 612 corresponding to an object (e.g., a puppy), and a third region 613 corresponding to a background (e.g., a forest).

[0164] A region related to the keyword can be selected from among the plurality of segmented regions. Figure 6 In an example of the disclosure, the second region 612 can be determined or selected as a region corresponding to the keyword. The electronic device 2000 can segment a region corresponding to the keyword from the reference image 610, and can isolate only the segmented region. The isolated segment can be referred to as a "segment 620 of the reference image 610."

[0165] In an embodiment of the disclosure, the electronic device 2000 can determine a region related to the keyword by analyzing the reference image 610.

[0166] For example, the electronic device 2000 can separate a plurality of semantic elements in the image using semantic segmentation. The electronic device 2000 can classify the plurality of semantic elements, and can select a segment indicating a region related to the keyword based on a classification result. For example, the electronic device 2000 can segment the reference image 610 into the first region 611, the second region 612, and the third region 613. The electronic device 2000 can classify the first region 611 as a person, can classify the second region 612 as a puppy, and can classify the third region 613 as a forest. The electronic device 2000 can determine the second region 612 corresponding to the puppy as a region related to the keyword. Accordingly, a segment 620 of the reference image 610 corresponding to the second region 612 can be extracted.

[0167] In an embodiment of the disclosure, the electronic device 2000 can determine or select a region related to the keyword based on a user input.

[0168] For example, the electronic device 2000 can separate a plurality of instances (e.g., objects) in the image using instance segmentation. The electronic device 2000 can display a segmentation result on a screen, and can receive a user input selecting at least one of the segments displayed on the screen. In detail, for example, the electronic device 2000 can segment the reference image 610 into the first region 611, the second region 612, and the third region 613, and can display the segmentation result on the screen. The electronic device 2000 can determine the second region 612 as a region related to the keyword based on a user input selecting the second region 612. Accordingly, a segment 620 of the reference image 610 corresponding to the second region 612 can be extracted.

[0169] The image segmentation method performed by the electronic device 2000 is not limited to the above examples. For example, the electronic device 2000 can use a panoramic segmentation method that is a combination of semantic segmentation and instance segmentation. Alternatively, the electronic device 2000 can select the segment 620 of the reference image 610 indicating a region related to the keyword based on a user input (e.g., a rectangular selection or a free selection (lasso selection)) that specifies a region of an object corresponding to the keyword in the reference image.

[0170] In an embodiment of the disclosure, the electronic device 2000 can store the keyword and the reference image 610 corresponding to the keyword as a text-image pair. In an embodiment of the disclosure, the electronic device 2000 can store the keyword and the segment 620 of the reference image 610 corresponding to the keyword as a text-image pair. The stored text-image pair can be encoded by the encoder.

[0171] Figure 7 is a diagram for describing an operation of the electronic device generating a vector representation according to an embodiment of the disclosure.

[0172] In an embodiment of the disclosure, the electronic device 2000 can generate the reference embedding 712 by converting the keyword and the reference image (or the segment of the reference image) into a vector representation.

[0173] In Figure 7 In order to facilitate explanation, the text-image pair of the keyword and the reference image (or the segment of the reference image) can be included in the reference data 710. In an embodiment, the text-image pair of the keyword and the reference image can be expressed as {keyword, reference image}.

[0174] The electronic device 2000 can encode the reference data 710 using the encoder 720. The encoder 720 can include a text encoder 722 and an image encoder 724. The encoder 720 can be trained to find a relationship between text and image and to generate a common vector representation between text and image. The encoder 720 can be implemented using a neural network architecture capable of processing text and image or by modification of the neural network architecture. For example, the encoder 720 can be implemented based on, but not limited to, a contrastive language-image pre-training (CLIP) architecture.

[0175] In an embodiment of the disclosure, the electronic device 2000 can convert the reference data 710 into a vector representation to generate the reference embedding 712 as data for image personalization, and can transmit the reference embedding 712 to the server 3000.

[0176] In an embodiment of the disclosure, the electronic device 2000 can convert the prompt 730 into a vector representation to generate a prompt embedding 732. The electronic device 2000 can transmit the prompt embedding 732 to the server 3000. The prompt 730 can be transmitted to the server 3000 and then can be converted into the prompt embedding 732 in the server 3000.

[0177] Figure 8 is a diagram for describing an operation of the server generating a personalized image according to an embodiment of the disclosure.

[0178] In an embodiment of the disclosure, the server 3000 can generate a personalized image. The server 3000 can generate a personalized image by applying or providing data (e.g., a reference embedding and a prompt embedding) received from the electronic device 2000 as input data to a generation model 800. The generation model 800 can include an image information generator 802 and an image decoder 804.

[0179] In an embodiment of the disclosure, the generation model 800 can process input data as various types of modalities according to various conditions. For example, the generation model 800 can receive text and can generate text-to-image from the text, or can receive an image and can generate image-to-image from the image. In another example, the generation model 800 can receive a text and an image and can generate image-to-image from the text and the image. Figure 8 In the following, the generation model 800 that processes a prompt as text as input will be described.

[0180] In an embodiment of the disclosure, the generation model 800 can be a model that applies a classifier-free guidance (CFG) method. When a condition for generating an image from text among condition options is described as an example, in a process of training the generation model 800, training of the generation model 800 can be performed when a condition given text is randomly removed. Accordingly, image generation of the generation model 800 can be guided according to a trade-off between a conditional likelihood, which is a probability of generating an image given text, and an unconditional likelihood, which is a probability of generating an image without text, adjusted. In an inference operation of the generation model 800 trained through the above process, a CFG scale s can be used as an input parameter. When a higher CFG scale value is input, a probability that the generation model 800 generates an image similar to a prompt can increase, but image quality can decrease. When a lower CFG scale value is input, the quality of an image generated by the generation model 800 can be improved, but the similarity to the prompt can decrease.

[0181] In an embodiment of the disclosure, the generation model 800 can be deployed after the training is completed. The training process of the generation model 800 can include, but is not limited to, a diffusion-backdiffusion process 820, which is a diffusion process of gradually adding noise to an original image, and a backdiffusion process of recovering the original image by denoising the noise from the noisy image.

[0182] In an embodiment of the disclosure, the diffusion-backdiffusion process 820 of the training process of the generation model 800 can be performed by the image information generator 802 of the generation model 800. In this case, the image information generator 802 can include a noise predictor for processing the backdiffusion process. The noise predictor can be implemented using a neural network architecture or by modification of a neural network architecture. For example, the noise predictor can be implemented based on, but not limited to, a U-Net architecture.

[0183] In the diffusion-backdiffusion process 820, training is performed by gradually adding noise to an image, predicting the amount of noise, and removing the noise. In this case, in the training process of the generation model 800 for generating an image from text, in order to find a text-image relationship, a prompt embedding obtained by converting a prompt (e.g., text) can be merged with an image to which noise is gradually added using an attention mechanism (e.g., cross-attention).

[0184] In an embodiment of the disclosure, the reference embedding is used only during an inference operation of the generation model 800. The inference operation of the generation model 800 can be performed by the inference code 810. When the server 3000 performs the inference operation using the inference code 810, the reference embedding and the prompt embedding can be used as input data of the generation model 800.

[0185] The image information generator 802 of the generation model 800 can process input data (e.g., embedding) to generate a multi-dimensional information array (e.g., tensor) used by the image decoder 804 to generate an image. The image decoder 804 can output a personalized image by converting the tensor output from the image information generator 802.

[0186] In an embodiment of the disclosure, the server 3000 can apply a weight to a reference image based on a user input. In this case, the input of selecting the weight can be obtained from the electronic device 2000 of the user and can be transmitted to the server, but the disclosure is not limited thereto. When the server 3000 performs the inference operation using the generation model 800, the server 3000 can receive the CFG scale s p for the prompt and the CFG scale s r for the reference as input parameters. In this case, when s p increases, a personalized image similar to the prompt can be generated, and when s rAs the value increases, a personalized image similar to a reference (e.g., a keyword or a reference image) can be generated.

[0187] In an embodiment of the disclosure, the server 3000 can receive a plurality of CFG scales s r for a reference, for example, {(keyword1, image1),..., (keywordn, imagen)}. The server can receive CFG scales {s r1 ,..., s rn} corresponding to the reference, respectively, and can perform an inference operation using the generation model 800 to adjust the degree to which the plurality of references are reflected in a personalized image.

[0188] Accordingly, the server 3000 can perform an inference operation by using the generation model 800 deployed after training is completed, and can use a reference embedding as input data only during the inference operation.

[0189] Because a reference image used to generate a reference embedding can be an image stored only in a local media storage of a user, the reference image can be data that is not used to train the deployed generation model 800. That is, an image personalization service provided by the server 3000 does not provide personalization by additionally training the deployed generation model 800 with respect to a reference image possessed by a user. Rather, because the server 3000 can use a reference embedding during an inference operation of the generation model 800, the server 3000 can generate a personalized image to include features of a reference image without retraining the generation model 800.

[0190] That is, features of a reference image that are not used to train the generation model 800 can be included in a personalized image generated by the generation model 800.

[0191] Figure 9 is a diagram for describing an example in which an electronic device according to an embodiment of the disclosure provides a personalized image.

[0192] In an embodiment of the disclosure, the electronic device 2000 can obtain a prompt. For example, the electronic device 2000 can obtain a prompt saying "draw our puppy." When "our puppy" is mentioned below, it is assumed that a user means a puppy that he / she raises.

[0193] The electronic device 2000 can extract a keyword from the obtained prompt by using the keyword extraction module 910. For example, the electronic device 2000 can extract "puppy" as a keyword. When there is one extracted keyword, the extracted keyword can be used as a personalization keyword. Alternatively, although not illustrated in FIG. 20, when there are a plurality of extracted keywords, the electronic device 2000 can select one keyword from among the plurality of extracted keywords and use the selected keyword as a personalization keyword. Figure 9"our" and "dog" as keywords. In this case, both "our" and "dog" can be determined or selected as personalized keywords, only "our" can be determined or selected as a personalized keyword, or only "dog" can be determined or selected as a personalized keyword.

[0194] The electronic device 2000 can search for one or more images corresponding to the keywords from images stored in a media storage 922 of the electronic device 2000 using an image search module 920. In this case, image metadata can be used. For example, the electronic device 2000 can provide dog images stored in the media storage 922 as search results using the image metadata. In this case, the image search results can include images of a dog owned by the user called our dog, and images of other dogs stored in the media storage 922. In this case, the user can check the image search results, and the electronic device 2000 can receive user input selecting the dog owned by the user as a reference image.

[0195] Alternatively, the image metadata can include a tag associated with "our dog." In this case, the image search results can display an image of a dog owned by the user called our dog.

[0196] When the reference image is determined, a region related to the keyword within the reference image can be separately segmented.

[0197] The electronic device 2000 can generate a reference embedding, and can transmit the reference embedding to the server 3000. The reference embedding can be obtained by converting a text-image pair into a vector representation using an encoder. The reference embedding can be obtained by converting, for example, a keyword-reference image pair into a vector representation. Alternatively, for example, the reference embedding can be obtained by converting a keyword-reference image segment pair into a vector representation.

[0198] The server 3000 can generate a personalized image using an image generation module 930. As a result of performing operations according to the above-described example, the personalized image can be a new image in which a dog owned by the user is drawn. The server 3000 can transmit the personalized image to the electronic device 2000, and the electronic device 2000 can display the personalized image.

[0199] Figure 10a is a flowchart illustrating operations of an electronic device and a server according to an embodiment of the disclosure.

[0200] In an embodiment of the disclosure, the electronic device 2000 can interact with the server 3000 to generate a personalized image. For example, a system for generating a personalized image can include the electronic device 2000 and the server 3000, and the electronic device 2000 can operate as a client.

[0201] Operation S1010 of the electronic device 2000 corresponds to operation S110 of FIG. 11 of the electronic device 1000, and thus a repeated description will be omitted for brevity. Figure 2 Operation S1020 of the electronic device 2000 corresponds to operation S120 of FIG. 11 of the electronic device 1000, and thus a repeated description will be omitted for brevity.

[0202] Figure 2 Operation S1020 of the electronic device 2000 corresponds to operation S120 of FIG. 11 of the electronic device 1000, and thus a repeated description will be omitted for brevity.

[0203] In an embodiment of the disclosure, there can be a plurality of keywords for personalization. When the electronic device 2000 extracts a plurality of keywords, in operation S1025, the electronic device 2000 can determine whether images corresponding to all of the plurality of keywords have been searched and obtained. Operations S1030 to S1036 can be repeatedly performed based on a result of operation S1025 until images corresponding to all of the keywords are searched.

[0204] Operation S1030 of the electronic device 2000 corresponds to operation S230 of FIG. 12 of the electronic device 1000, and thus a repeated description will be omitted for brevity. Figure 2 Operation S1030 of the electronic device 2000 corresponds to operation S230 of FIG. 12 of the electronic device 1000, and thus a repeated description will be omitted for brevity.

[0205] Figure 2 Operation S1032 of the electronic device 2000 corresponds to operation S232 of FIG. 12 of the electronic device 1000, and thus a repeated description will be omitted for brevity.

[0206] The electronic device 2000 can perform image segmentation on the selected image. For example, in operation S1033, the electronic device 2000 can determine whether to apply image segmentation based on a user input. When image segmentation is applied, operation S1034 can be performed. When image segmentation is not performed, operation S1036 can be performed.

[0207] Operation S1034 of the electronic device 2000 corresponds to operation S240 of FIG. 12 of the electronic device 1000, and thus a repeated description will be omitted for brevity. Figure 2 Operation S1034 of the electronic device 2000 corresponds to operation S240 of FIG. 12 of the electronic device 1000, and thus a repeated description will be omitted for brevity.

[0208] ​​In operation S1036, the electronic device 2000 can add the keyword and the selected image to the reference data. The selected image can be used as a reference image. The reference data can be a text-image pair including the keyword and the reference image, or a text-image pair including a clip of the keyword and the reference image.

[0209] When one or more reference images corresponding to each of the plurality of keywords are determined, respectively, the electronic device 2000 can perform operation S1040.

[0210] In operation S1040, the electronic device 2000 can generate a prompt embedding and a reference embedding. Operation S1040 corresponds to operation S250 of FIG. 2, and thus, a repeated description will be omitted for the sake of brevity. Figure 2

[0211] The electronic device 2000 can transmit the prompt embedding (or prompt) and the reference embedding to the server 3000.

[0212] In operation S1045, the server 3000 can generate an image. The server 3000 can generate one or more images using a generation model. The generation model can use the reference embedding as input data. Because the generated image includes features of the reference image stored in the user's media storage, the generated image can be referred to as a personalized image. The server 3000 can transmit the personalized image to the electronic device 2000.

[0213] In operation S1050, the electronic device 2000 can display one or more images received from the server 3000.

[0214] Figure 10b is a flowchart illustrating additional operations of an electronic device and a server according to an embodiment of the disclosure.

[0215] The electronic device 2000 can perform an operation of regenerating a personalized image based on a pre-defined condition.

[0216] ​In an embodiment of the disclosure, the electronic device 2000 can re-generate the personalized image based on a similarity between the personalized image received from the server 3000 and the reference image. The electronic device 2000 can compare the similarity between the personalized image and the reference image. The electronic device 2000 can determine the similarity between the images using any one of various algorithms. For example, the electronic device 2000 can calculate the similarity between the personalized image and the reference image through pixel unit comparison or histogram comparison. Alternatively, the electronic device 2000 can extract features of the images using a neural network (e.g., CNN) for processing images, and can calculate the similarity between the personalized image and the reference image. When the similarity between the images is lower than a preset threshold value, the electronic device 2000 can transmit the prompt embedding (or prompt) and the reference embedding to the server 3000 to perform operation S1045 again. The electronic device 2000 can receive the re-generated image from the server 3000, and can display the re-generated image on the screen.

[0217] In an embodiment of the disclosure, the electronic device 2000 can re-generate the personalized image based on a user input. The electronic device 2000 can receive a user input requesting image re-generation in operation S1052. When the user input requesting image re-generation is received, the electronic device 2000 can transmit the prompt embedding (or prompt) and the reference embedding to the server 3000 to perform operation S1045 again. The electronic device 2000 can receive the re-generated image from the server 3000, and can display the re-generated image on the screen.

[0218] Figure 11 FIG. 10 is a diagram for describing an operation of an electronic device according to an embodiment of the disclosure to provide image feature information to a user.

[0219] In an embodiment of the disclosure, the electronic device 2000 can obtain image feature information. For example, the electronic device 2000 can generate a reference embedding by converting reference data including a text-image pair into a vector representation for image generation. In this case, the reference embedding is transmitted to the server 3000, and the server 3000 generates an image by using the reference embedding.

[0220] In an embodiment of the disclosure, the electronic device 2000 can store image feature information. The electronic device 2000 can store the image feature information, and when there is another request to generate a personalized image after the image feature information is stored, the image feature information can be visualized and displayed. For example, the electronic device 2000 can store the generated reference embedding. When there is another request to generate a personalized image after the reference embedding is stored, the electronic device 2000 can visualize and display the stored reference embedding. In this case, the electronic device 2000 can display the reference embedding in a form of a graph, a table, or a list. Figure 11In the middle, a graphical interface through which the electronic device 2000 can receive a prompt as input from a user is shown. Various elements can be included in the graphical interface. For example, a prompt input field 1100 and recently used reference data 1110 can be included, but some of the above elements can be omitted from the graphical interface or other elements can be added to the graphical interface through which a prompt can be received.

[0221] In an embodiment of the disclosure, the electronic device 2000 can convert the reference embedding into the reference data 1110 using a decoder that performs an inverse operation of an encoder. For example, the reference data can be a text-image pair including a keyword and a reference image, which can be expressed as {keyword, reference image}. For example, referring to the reference data 1110, the first reference data = {keyword: Chris, reference image: an image including Chris}. Similarly, the second reference data 1112 = {Joy, an image including Joy}, the third reference data = {Tom, an image including Tom}, and the fourth reference data 1114 = {my pet, an image including my pet}. Figure 11

[0222] The electronic device 2000 can receive a user input selecting the reference data. When the user of the electronic device 2000 inputs a prompt, the user can use the reference data 1110. For example, the user can use the reference data 1110 instead of inputting the sentence "draw a picture of Joy playing with my pet in the park" as a text or voice input. For example, the electronic device 2000 can receive a user input selecting the second reference data 1112 and the fourth reference data 114 from the reference data 1110.

[0223] The electronic device 2000 can transmit the reference embedding corresponding to the reference data 1110 to the server 3000 based on the user input selecting the reference data 1110. The server 3000 can generate a personalized image using a generation model based on the reference embedding, and can transmit the personalized image to the electronic device 2000.

[0224] The electronic device 2000 can omit one or more of the operations for generating the reference embedding by storing the reference embedding and reusing the stored reference. For example, all or some of the operations performed by the electronic device 2000 to generate the reference embedding, such as prompt acquisition, keyword extraction, image search, image segmentation, and vector conversion, can be omitted.

[0225] Figure 12 is a diagram for describing an operation of the electronic device generating or processing a prompt according to an embodiment of the disclosure.

[0226] The prompt generation module 1200 can include a natural language processing model for performing at least one of generating, processing, and converting a prompt.​

[0227] In an embodiment of the disclosure, the electronic device 2000 can process the obtained prompt 1210 using the prompt generation module 1200. The electronic device 2000 can obtain the prompt 1210 and can extract a keyword 1220 from the prompt 1210. For example, the electronic device 2000 can extract a puppy as the keyword 1220 from the prompt 1210 saying "draw our puppy."

[0228] The electronic device 2000 can search for one or more images 1230 corresponding to the keyword and can select one or more reference images from the searched and obtained one or more images 1230. For example, the electronic device 2000 can find one or more images 1230 corresponding to a puppy from images stored in a media storage of the electronic device 2000. In this case, a reference image can be determined from the one or more images 1230.

[0229] The electronic device 2000 can extract a text description 1240 of the reference image from the reference image. The electronic device 2000 can obtain the text description 1240 of the image using any one of various methods for extracting text from an image (e.g., captioning an image). For example, the electronic device 2000 can extract a text description 1240 saying "a Samoyed with white fur" from the reference image. Alternatively, the electronic device 2000 can extract the text description 1240 from the searched and obtained one or more images 1230.

[0230] The electronic device 2000 can change the prompt 1210 based on the text description 1240 of the reference image. The content of the text description 1240 can be included in a new prompt 1212.

[0231] In an embodiment of the disclosure, the electronic device 2000 can generate the new prompt 1212 using the prompt generation module 1200. For example, the electronic device 2000 can extract the text description 1240 from the searched and obtained one or more images 1230 and / or the reference image and can generate the new prompt 1212 based on the text description. The new prompt 1212 can include the content of the text description 1240 and can include text having a different content from the original prompt 1210.

[0232] In an embodiment of the disclosure, the electronic device 2000 can process the text description 1240 as input data for generating a personalized image. For example, the electronic device 2000 can include the text description in reference data. In this case, the reference data = {keyword, reference image, text description 1240}.

[0233] The electronic device 2000 can obtain image feature information using reference data including a keyword, a reference image, and a text description 1240. For example, the electronic device 2000 can generate a reference embedding by converting the reference data into a vector representation. The generated reference embedding can be transmitted to the server 3000.

[0234] In an embodiment of the disclosure, when the server 3000 performs an inference operation by using a generation model, the server 3000 can receive a CFG scale s t for the text description 1240 as an input parameter. That is, the server 3000 can receive a CFG scale s p for the reference, r and a CFG scale s t for the text description 1240 as input parameters. In this case, as s p increases, a personalized image similar to the prompt can be generated; as s r increases, a personalized image similar to the reference (e.g., the keyword and the reference image) can be generated; and as s t increases, a personalized image similar to the text description can be generated.

[0235] Figure 13 is a diagram for describing an operation in which an electronic device according to an embodiment of the disclosure provides a personalized image.

[0236] In an embodiment of the disclosure, the electronic device 2000 can provide a retouched image 1330 as a personalized image to the user using retouching. The retouched image 1330 can be generated based on a reference image 1300, an original image 1310, and a mask 1320.

[0237] In an embodiment of the disclosure, the electronic device 2000 can obtain the original image 1310 and the mask 1320. For example, the electronic device 2000 can receive a user input selecting the original image 1310 from the user. The original image 1310 can be stored in a local media storage of the electronic device 2000, or can be obtained (e.g., a search result) by the user from an external source. Furthermore, the electronic device 2000 can receive an input specifying the mask 1320 corresponding to the original image 1310 from the user.

[0238] The electronic device 2000 can determine the reference image 1300. Furthermore, the electronic device 2000 can generate a reference embedding using the reference image 1300. For example, the electronic device 2000 can perform all or some of prompt acquisition, keyword extraction, image search, image segmentation, and vector conversion. The operation of generating a reference embedding has been described in detail with reference to the preceding drawings, and thus a repetitive description will be omitted.

[0239] The electronic device 2000 can transmit the reference image 1300 (or reference embedding), the original image 1310, and the mask 1320 to the server 3000. The server 3000 can generate the inpainted image 1330 as a personalized image by applying the data (e.g., the prompt or prompt embedding, the reference image 1300 or reference embedding, the original image 1310, and the mask 1320) received from the electronic device 2000 as input data to the generation model. In this case, the generation model can generate the inpainted image 1330 using an image-to-image method. In detail, the generation model can identify an area to be inpainted within the original image 1310 based on the original image 1310 and the mask 1320, and can fill pixels so that the feature of the reference image 1300 is included in the identified area.

[0240] The server 3000 can transmit the inpainted image 1330 to the electronic device 2000.

[0241] The term "original image" can correspond to various expressions that can represent the same / similar concept. For example, the term "original image" can correspond to expressions such as "initial image," "base image," and "basic image."

[0242] Figure 14a is a diagram for describing an operation of an electronic device providing a personalized image according to an embodiment of the disclosure.

[0243] In an embodiment of the disclosure, the electronic device 2000 can provide a personalized image based on an input image. For example, the electronic device 2000 can provide the user with a personalized image (generated image 1420) in which the feature of the reference image 1410 is applied to the original image 1400 while the composition of the original image 1400 is maintained. The generated image 1420 can be generated based on the original image 1400 and the reference image 1410.

[0244] In an embodiment of the disclosure, the electronic device 2000 can obtain the original image 1400. For example, the electronic device 2000 can receive a user input selecting the original image 1400 from the user. The original image 1400 can be an image (e.g., search) stored in a local media storage of the electronic device 2000 or obtained by the user from an external source.

[0245] The electronic device 2000 can determine the reference image 1410. In addition, the electronic device 2000 can generate a reference embedding using the reference image 1410. For example, the electronic device 2000 can perform all or some of prompt acquisition, keyword extraction, image search, image segmentation, and vector conversion. The operation of generating a reference embedding has been described with reference to the preceding figures, and thus a repetitive description will be omitted.

[0246] The electronic device 2000 can transmit the original image 1400 and the reference image 1410 (or the reference embedding) to the server 3000. The server 3000 can generate a generated image 1420 that is a personalized image by applying data (the hint or the hint embedding, the original image 1400, and the reference image 1410 or the reference embedding) received from the electronic device 2000 as input data to the generation model. In this case, the generation model can generate the generated image 1420 using an image-to-image method. For example, the generation model can change the original image 1400 while maintaining the composition of the original image 1400 to include the features of the reference image 1410. In detail, when the hint is "change the face of the puppy," the generated image 1420 that maintains the appearance of the puppy in the original image 1400 and changes the face of the puppy to the face of the puppy in the reference image 1410 can be generated.

[0247] Figure 14b FIG. 1400 is a diagram for describing an operation of an electronic device providing a personalized image according to an embodiment of the disclosure.

[0248] In an embodiment of the disclosure, the electronic device 2000 can provide an image generated based on an image input. In Figure 14b FIG. 1400, an operation of generating a personalized image using an image-to-image method described with reference to Figure 14a FIG. 1400 will be further described.

[0249] In an embodiment of the disclosure, the electronic device 2000 can perform a data processing operation for generating a personalized image using an image-to-image method. The electronic device 2000 can provide a new image (e.g., a personalized image) generated based on an input image (an original image) and a reference image. In this case, the personalized image can be an image obtained by reflecting features of the reference image in the original image.

[0250] In operation S1410, the electronic device 2000 can obtain an original image. The original image of operation S1410 can correspond to the original image 1400 described with reference to Figure 14a FIG. 1400. The original image 1400 can be stored in a local media storage of the electronic device 2000, or can be obtained (e.g., a search result) by a user from an external source.

[0251] In operation S1420, the electronic device 2000 can determine a reference image. For example, the electronic device 2000 can determine the reference image based on a user input selecting the reference image. The reference image of operation S1420 can correspond to the reference image 1410 described with reference to Figure 14a FIG. 1400.

[0252] The electronic device 2000 can generate a reference embedding in operation S1430. The electronic device 2000 can perform all or some of prompt acquisition, keyword extraction, image search, image segmentation, and vector conversion to generate the reference embedding.

[0253] In an embodiment of the disclosure, the electronic device 2000 can generate the reference embedding by converting the reference image into a vector representation. The electronic device 2000 can convert the reference image into a vector representation, or can perform image segmentation on the reference image, and can convert the obtained segment (e.g., object) into a vector representation.

[0254] In an embodiment of the disclosure, the electronic device 2000 can generate the reference embedding by converting a text-image pair including a keyword and a reference image into a vector representation. The keyword can be obtained using any one of various methods. For example, the electronic device 2000 can analyze the reference image. The electronic device 2000 can obtain text related to the reference image using image captioning, image classification, or object recognition, and can extract a keyword from the text. For example, the electronic device 2000 can obtain user input indicating a keyword corresponding to the reference image. For example, the electronic device 2000 can obtain a prompt. The electronic device 2000 can extract a keyword from the prompt. When the electronic device 2000 obtains the keyword, image segmentation can be performed on the reference image to extract a region related to the keyword.

[0255] In an embodiment of the disclosure, the electronic device 2000 can transmit the prompt, the original image, and the reference embedding to the server 3000. The prompt transmitted to the server 3000 can be converted into a vector representation of the prompt embedding in the server 3000, and the original image transmitted to the server 3000 can be converted into a vector representation of the image embedding in the server 3000.

[0256] In an embodiment of the disclosure, the electronic device 2000 can generate a prompt embedding by converting the prompt into a vector representation, and can transmit the prompt embedding to the server 3000.

[0257] In an embodiment of the disclosure, the electronic device 2000 can generate an image embedding by converting the original image into a vector representation, and can transmit the image embedding to the server 3000.

[0258] The server 3000 can generate an image in operation S1431. The server 3000 can generate one or more images using a generation model. The generation model can use the prompt (or the prompt embedding), the original image (or the image embedding), and the reference embedding as input data. The features of the original image and the reference image can be included in the generated image. The generated image can be referred to as a personalized image. The server 3000 can transmit the personalized image to the electronic device 2000.

[0259] In operation S1440, the electronic device 2000 can display one or more images received from the server 3000.

[0260] Figure 14c is a diagram for describing an operation of an electronic device providing a personalized image according to an embodiment of the disclosure.

[0261] In an embodiment of the disclosure, the electronic device 2000 can perform an operation for generating a personalized image based on an image generated by using a generative model.

[0262] In operation S1402, the server 3000 can generate a first image using a generative model. The first image can be an image generated using only a hint (a first hint). The server 3000 can obtain the first hint. For example, the electronic device 2000 can obtain the first hint and can transmit the first hint to the server 3000. The server 3000 can receive the first hint from the electronic device 2000. In another example, the server 3000 can receive the first hint from another electronic device (e.g., a desktop personal computer as another client device) in communication with the server 3000.

[0263] The server 3000 can apply the first hint as input data to the generative model and can obtain the first image as output data. Certain data processing can be applied to the first hint. For example, the first hint can be converted into a hint embedding represented as a vector.

[0264] The electronic device 2000 can receive the first image generated based on the first hint from the server 3000. There can be one or more first images.

[0265] In operation S1406, the electronic device 2000 can display one or more first images received from the server 3000.

[0266] In operation S1422, the electronic device 2000 can determine a reference image from among the one or more first images. For example, the electronic device 2000 can determine a first image selected based on input selecting one of the one or more first images as the reference image.

[0267] In operation S1432, the electronic device 2000 can generate a reference embedding.

[0268] In an embodiment of the disclosure, the electronic device 2000 can obtain a second hint. The second hint can be another hint different from the first hint. For example, the second hint can be a hint input for generating a second image.

[0269] In an embodiment of the disclosure, the electronic device 2000 can generate a reference embedding by converting a text-image pair including a keyword and a reference image into a vector representation. The keyword can be extracted from the second prompt or can be input by a user.

[0270] The electronic device 2000 can transmit the second prompt and the reference embedding to the server 3000. The second prompt transmitted to the server 3000 can be converted into a vector representation in the server 3000. In an embodiment of the disclosure, the electronic device 2000 can generate a second prompt embedding by converting the second prompt into a vector representation, and can transmit the second prompt embedding to the server 3000.

[0271] In operation S1434, the server 3000 can generate a second image. The server 3000 can generate one or more second images using a generation model. The generation model can use the second prompt (or the second prompt embedding) and the reference embedding as input data. The features of the reference image that are the input image can be included in the second image. The reference image is determined from the one or more first images generated by the generation model, and thus, the second image can include the features of the first image. The second image can be referred to as a personalized image. The server 3000 can transmit the personalized image to the electronic device 2000.

[0272] In operation S1442, the electronic device 2000 can display one or more second images received from the server 3000.

[0273] Figure 14d is a diagram for describing an operation of an electronic device providing a personalized image according to an embodiment of the disclosure.

[0274] In an embodiment of the disclosure, the electronic device 2000 can perform an operation for generating a personalized image using an image-to-image method based on an image generated by using a generation model.

[0275] In operation S1404, the server 3000 can generate a first image by using a generation model. The first image can be generated by the generation model based on a first prompt.

[0276] In operation S1408, the electronic device 2000 can display one or more first images received from the server 3000.

[0277] In operation S1412, the electronic device 2000 can determine an original image from the one or more first images. The original image of operation S1412 can correspond to the original image 1400 described with reference to Figure 14a

[0278] ​In operation S1424, the electronic device 2000 can determine a reference image. For example, the electronic device 2000 can determine the reference image based on a user input selecting the reference image. The reference image of operation S1424 can correspond to the reference image 1410 described with reference to Figure 14a

[0279] In operation S1434, the electronic device 2000 can generate a reference embedding. To generate the reference embedding, the electronic device 2000 can perform all or some of a second prompt acquisition, keyword extraction, image search, image segmentation, and vector conversion. The second prompt can be a prompt input to generate a second image. Operation S1434 can correspond to operation S1430 of Figure 14b

[0280] In operation S1436, the server 3000 can generate a second image. The server 3000 can generate one or more second images using a generation model. The generation model can use a prompt (or a prompt embedding), an original image (or an image embedding), and a reference embedding as input data. The second image can include features of the original image and the reference image. Because the original image is determined from one or more first images generated by the generation model, the second image can include features of the first image and the reference image. The second image can be referred to as a personalized image. The server 3000 can transmit the personalized image to the electronic device 2000.

[0281] In operation S1444, the electronic device 2000 can display one or more second images received from the server 3000.

[0282] Figure 15 is a diagram for describing an operation of an electronic device providing a personalized image according to an embodiment of the disclosure.

[0283] In an embodiment of the disclosure, the electronic device 2000 can generate an image showing various appearances of a user according to a request of the user. For example, the electronic device 2000 can generate an image in which the user wears a product. In Figure 15 In, it will be assumed that the following description is made assuming that an image in which the user wears a virtual piece of clothing is generated.

[0284] The electronic device 2000 can obtain an image 1500. For example, the electronic device 2000 can receive a user input selecting the image 1500. The image 1500 can be stored in a local media storage 1522 of the electronic device 2000, or can be obtained (e.g., a search result) by the user from an external source.

[0285] ​​The electronic device 2000 can generate a hint corresponding to the image 1500 by using the hint generation module 1510. For example, the electronic device 2000 can extract a textual description of the image 1500 from the image 1500, and can generate a hint based on the extracted text. The hint generated by the hint generation module 1510 can be determined or changed based on a user input. Alternatively, the hint can be directly received from the user. In detail, for example, when the image 1500 is an image including a product "black sleeveless shirt", the hint can be "I wear a black sleeveless shirt".

[0286] The electronic device 2000 can determine one or more personalized keywords from the hint. For example, the electronic device 2000 can extract a first personalized keyword "I" and a second personalized keyword "black sleeveless shirt" from the hint "I wear a black sleeveless shirt".

[0287] The electronic device 2000 can search for an image corresponding to the personalized keyword from one or more images stored in the media storage 1522 of the electronic device 2000 by using the image search module 1520. A reference image can be determined from the searched and retrieved one or more images. For example, an image including a user as an image corresponding to the keyword "I" can be determined as a first reference image 1502.

[0288] In an embodiment of the disclosure, the electronic device 2000 can segment a region in the image 1500 by using the image segmentation module 1530. Various methods can be applied as image segmentation, and the operation has been described above, and thus a repetitive description will be omitted. As a result of image segmentation, a segment corresponding to a keyword can be determined as a reference image. For example, a segment including a black sleeveless shirt as an image corresponding to the keyword "black sleeveless shirt" can be determined as a second reference image 1504.

[0289] The electronic device 2000 can convert reference data including a personalized keyword and a reference image, {keyword, reference image}, into a vector representation. For example, the electronic device 2000 can generate a first reference embedding by converting the first reference data {I, image including a user}, and can generate a second reference embedding by converting the second reference data {black sleeveless shirt, image including a black sleeveless shirt}.

[0290] The electronic device 2000 can transmit the hint (or hint embedding) and the reference embedding to the server 3000. For example, the electronic device 2000 can transmit the hint saying "I wear a black sleeveless shirt", the first reference embedding, and the second reference embedding to the server 3000.

[0291] The server 3000 can generate the personalized image 1506 using the image generation module 1540. For example, the personalized image 1506 can include an image of the user wearing a black sleeveless shirt. The server 3000 can transmit the personalized image to the electronic device 2000.

[0292] The electronic device 2000 can display the personalized image received from the server 3000 on the screen.

[0293] Figure 16 FIG. 1 is a diagram for describing an operation of an electronic device providing information related to a personalized image according to an embodiment of the disclosure.

[0294] In Figure 16 hereinafter, it will be assumed that the personalized image 1600 has been generated. It will be assumed that the personalized image 1600 is an image of a user wearing a product as described with reference to Figure 15 hereinafter.

[0295] In an embodiment of the disclosure, the electronic device 2000 can provide information related to the personalized image 1600. For example, the electronic device 2000 can provide the user with information related to the object and the product included in the personalized image 1600.

[0296] The electronic device 2000 can detect one or more objects (e.g., products) included in the personalized image 1600.

[0297] The electronic device 2000 can extract a segment 1612 corresponding to the object in the personalized image 1600 using the image segmentation module 1610. For example, the electronic device 2000 can detect the product "black sleeveless shirt" in the personalized image 1600. Furthermore, the electronic device 2000 can detect the object in the personalized image 1600 using any one of various detection algorithms.

[0298] The electronic device 2000 can search for the object using the image search module 1620. For example, the electronic device 2000 can search for the detected product "black sleeveless shirt" in the clothing DB 1622. In an embodiment, the clothing database 1622 can be or can include a vector database.

[0299] The electronic device 2000 can provide the image search result 1630 obtained by searching for the object included in the personalized image 1600 as information related to the personalized image 1600 to the user. For example, the electronic device 2000 can provide the user with the result of searching for the product "black sleeveless shirt."

[0300] In an embodiment of the disclosure, the electronic device 2000 can provide additional information (e.g., a product name, a brand, a price, and a purchase link) of an object included in the image search result 1630 as information related to the personalized image 1600.

[0301] In an embodiment of the disclosure, the electronic device 2000 can generate another personalized image based on the information related to the personalized image 1600. For example, the electronic device 2000 can receive a user input of selecting an image including a detected object, and can generate another personalized image using the selected image as a reference image. In detail, the electronic device 2000 can receive a user input of selecting an image of other clothes from the image search result 1630, and can generate and provide an image of the user wearing the other clothes. It has been described with reference to FIGS. 16A and 16B, and thus a repeated description will be omitted. Figure 15 The operation is described, and thus a repeated description will be omitted.

[0302] Figure 17 is a diagram for describing an operation of an electronic device providing a personalized image according to an embodiment of the disclosure.

[0303] In an embodiment of the disclosure, the electronic device 2000 can provide a recommendation to a user in response to a request of the user. The recommendation provided to the user can be provided through a personalized image 1760 including recommended content.

[0304] The electronic device 2000 can obtain a user input. The user input can include, but is not limited to, a text input and a voice input. For example, the electronic device 2000 can receive a user input of a request indicating "recommend a cotton-padded jacket" from the user.

[0305] The electronic device 2000 can generate a prompt corresponding to the user input using the prompt generation module 1710. The prompt generation module 1710 can include a natural language processing model for generating / processing / converting a prompt. For example, the electronic device 2000 can generate a prompt saying "I wear a cotton-padded jacket."

[0306] The electronic device 2000 can extract one or more keywords from the prompt using the keyword extraction module 1720. For example, the electronic device 2000 can extract a first keyword "I" and a second keyword "cotton-padded jacket" from the prompt saying "I wear a cotton-padded jacket."

[0307] The electronic device 2000 can search for an image corresponding to the keyword from among images stored in a media storage 1732 of the electronic device 2000 using the first image search module 1730. For example, the electronic device 2000 can select a first image including a user corresponding to the first keyword "I" from among images stored in the media storage 1732. The electronic device 2000 can select a second image including the second keyword "cotton-padded jacket" from among images stored in the media storage 1732.

[0308] The electronic device 2000 can extract a clip from an image using the image segmentation module 1735. For example, the electronic device 2000 can segment a region corresponding to the user from the first image, and can segment a region corresponding to the cotton-padded jacket from the second image.

[0309] The electronic device 2000 can search for an image similar to a clip of an image using the second image search module 1740. For example, the electronic device 2000 can search for an image similar to the cotton-padded jacket from data stored in the clothing database 1742. In an embodiment, the clothing database 1622 can be or can include a vector database. The first image search module 1730 and the second image search module 1740 can be implemented as one image search module.

[0310] The electronic device 2000 can determine a reference image. For example, the electronic device 2000 can determine the first image corresponding to the first keyword "I" as the first reference image. Also, for example, the electronic device 2000 can determine an image searched by using the second image search module 1740 as the second reference image for recommendation. As a result, the first reference data = {I, first reference image} and the second reference data = {cotton-padded jacket, second reference image} are determined as text-image pairs.

[0311] The electronic device 2000 can generate a reference embedding by converting the first reference data and the second reference data into a vector representation. The electronic device 2000 can transmit the reference embedding to the server. The server 3000 can generate a personalized image 1760 by using a generation model. In this case, the generation model can use the reference embedding as input data. The personalized image 1760 can include a feature of the reference image. In detail, the personalized image 1760 can be an image in which the user wears the cotton-padded jacket searched in the clothing database 1742.

[0312] The electronic device 2000 can provide information related to the personalized image to the user based on the generated personalized image 1760. It has been described with reference to Figure 16 The operation is described, and thus, a repeated description will be omitted.

[0313] Figure 18a FIG. 1 is a block diagram illustrating a configuration of an electronic device according to an embodiment of the disclosure.

[0314] In an embodiment of the disclosure, the electronic device 2000 can include a communication interface 2100, a memory 2200, a processor 2300, and a display 2400.

[0315] The communication interface 2100 can perform data communication with other electronic devices under the control of the processor 2300.

[0316] The communication interface 2100 can include a communication circuit capable of performing data communication between the electronic device 2000 and another electronic device (e.g., the server 3000) by using at least one of data communication methods including, for example, wired LAN, wireless LAN, Wi-Fi, Bluetooth, ZigBee, Wi-Fi Direct (WFD), infrared communication (e.g., Infrared Data Association (IrDA)), Bluetooth Low Energy (BLE), Near Field Communication (NFC), Wireless Broadband Internet (Wibro), Worldwide Interoperability for Microwave Access (WiMAX), Shared Wireless Access Protocol (SWAP), Wireless Gigabit Alliance (WiGig), and Radio Frequency (RF) communication.

[0317] The communication interface 2100 can transmit data for providing a personalized image to the server 3000 and receive data for providing a personalized image from the server 3000. For example, the communication interface 2100 can transmit a reference embedding to the server 3000 and can receive a personalized image from the server 3000.

[0318] The memory 2200 can store instructions, data structures, and program codes readable by the processor 2300. Operations performed by the processor 2300 can be implemented by executing instructions or codes of a program stored in the memory 2200.

[0319] The memory 2200 can include a non-volatile memory and a volatile memory (RAM or SRAM), the non-volatile memory including at least one of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., a secure digital (SD) or extreme digital (XD) memory), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, or an optical disk.

[0320] The memory 2200 can store one or more instructions and / or programs that allow the electronic device 2000 to operate to provide a personalized image. For example, instructions and / or programs for performing the functions of the keyword extraction module 2210, the image search module 2220, the image segmentation module 2230, and the encoder 2240 can be stored in the memory 2200. Instructions and / or programs for performing the functions of the prompt generation module (not shown) can be further stored in the memory 2200.

[0321] The processor 2300 can control the overall operation of the electronic device 2000. For example, the processor 2300 can control the overall operation of the electronic device 2000 to provide a personalized image by executing one or more instructions of a program stored in the memory 2200. There can be one or more processors 2300.

[0322] The processor 2300 can include at least one of, for example, and without limitation, a central processing unit, a microprocessor, a graphics processing unit, an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), an application processor (AP), a neural processing unit, or a dedicated artificial intelligence (AI) processor designed with a hardware structure dedicated to processing an AI model.

[0323] The processor 2300 can extract one or more keywords from the prompt by executing the keyword extraction module 2210. The operation of the keyword extraction module 2210 has been described with reference to the preceding drawings, and thus a repetitive description will be omitted.

[0324] The processor 2300 searches for one or more images corresponding to the keyword by executing the image search module 2220. A reference image can be selected from the searched and retrieved one or more images. The keyword-reference image pair can be stored as reference data. The operation of the image search module 2220 has been described with reference to the preceding drawings, and thus a repetitive description will be omitted.

[0325] The processor 2300 can segment a region corresponding to the keyword in the reference image by executing the image segmentation module 2230. In this case, the keyword-reference image segment pair can be stored as reference data. The operation of the image segmentation module 2230 has been described with reference to the preceding drawings, and thus a repetitive description will be omitted.

[0326] The processor 2240 can generate a reference embedding by converting the reference data into a vector representation using the encoder 2240. The operation of the encoder 2240 has been described with reference to the preceding drawings, and thus a repetitive description will be omitted.

[0327] The modules stored in the memory 2200 are for convenience of explanation, and the present disclosure is not limited thereto. Other modules (e.g., a prompt generation module) can be added to implement the above-described embodiments, or certain modules (e.g., the image segmentation module) can be omitted. Furthermore, one module can be divided into a plurality of modules according to detailed functions, and some modules can be integrated into one module.

[0328] The display 2400 can output an image signal to a screen of the electronic device 2000 under the control of the processor 2300. For example, the display 2400 can output an image signal processed in the process in which the electronic device 2000 provides a personalized image to the screen as a prompt input result, an image search result, or a personalized image generation result. The display 2400 can include a touch panel. The touch panel can include one or more touch sensors for detecting a touch input. In an embodiment of the present disclosure, the prompt can be text input through the touch panel.

[0329] Although not shown in FIG. 2, Figure 18a the electronic device 2000 can further include additional elements for performing the operations described in the above embodiments. For example, the electronic device 2000 can further include a camera, a microphone, etc. The electronic device 2000 can store an image obtained using the camera in the memory 2200. The electronic device 2000 can receive a voice input indicating a prompt using the microphone.

[0330] Figure 18b is a block diagram illustrating a configuration of an electronic device according to an embodiment of the disclosure.

[0331] In an embodiment of the disclosure, the electronic device 2000 can include a communication interface 2100, a memory 2200, a processor 2300, and an input / output interface 2500. Figure 18b The communication interface 2100, the memory 2200, and the processor 2300 of the electronic device 2000 can correspond to the communication interface 2100, the memory 2200, and the processor 2300 of the electronic device 1000, respectively. Accordingly, a repetitive description will be omitted for brevity. Figure 18a The communication interface 2100, the memory 2200, and the processor 2300 of the electronic device 2000 can correspond to the communication interface 2100, the memory 2200, and the processor 2300 of the electronic device 1000, respectively. Accordingly, a repetitive description will be omitted for brevity.

[0332] The input / output interface 2500 can process input to and output from the electronic device 2000. The input / output interface 2500 can provide a connection between the electronic device 2000 and an external device. Examples of the input / output interface 2500 can include, but are not limited to, a universal serial bus (USB) port, a high-definition multimedia interface (HDMI) port, a display port (DP) port, a video graphics array (VGA) port, an RGB port, a digital video interface (DVI), and an audio jack. The electronic device 2000 can be connected to an external device such as a display, a camera, a microphone, a speaker, a keyboard, a mouse, and a touchpad through the input / output interface 2500.

[0333] The electronic device 2000 can output an image signal by using a display connected to the electronic device 2000 through the input / output interface 2500. For example, a prompt input result, an image search result, or a personalized image generation result can be displayed on a screen of the display connected to the electronic device 2000.

[0334] The electronic device 2000 can obtain input data using an input device connected to the electronic device 2000 through the input / output interface 2500. For example, the electronic device 2000 can receive a text input indicating a prompt using a keyboard, a mouse, and a touchpad connected to the electronic device 2000. For example, the electronic device 2000 can receive a voice input indicating a prompt using a microphone connected to the electronic device 2000.

[0335] When the method according to the embodiment of the disclosure includes a plurality of operations, the plurality of operations can be executed by one processor or can be executed by a plurality of processors. For example, when a first operation, a second operation, and a third operation are performed by the method according to the embodiment of the disclosure, all of the first operation, the second operation, and the third operation can be executed by a first processor, or the first operation and the second operation can be executed by a first processor (e.g., a general-purpose processor), and the third operation can be executed by a second processor (e.g., an artificial intelligence processor). The artificial intelligence processor as an example of the second processor can execute an operation for training / inferencing an artificial intelligence model. However, embodiments of the disclosure are not limited thereto.

[0336] The one or more processors according to the disclosure can be implemented as a single core processor or a multi-core processor.

[0337] When the method according to the embodiment of the disclosure includes a plurality of operations, the plurality of operations can be executed by one core or can be executed by a plurality of cores included in one or more processors.

[0338] Figure 19a is a block diagram illustrating a configuration of a server according to an embodiment of the disclosure.

[0339] In an embodiment of the disclosure, the server 3000 can include a communication interface 3100, a memory 3200, and a processor 3300. The server 3000 can be a computing device having higher performance than the electronic device 2000 and capable of performing complex calculations and tasks such as training, inference, management, and deployment of a generation model using large-scale data.

[0340] The communication interface 3100 can perform data communication with other electronic devices under the control of the processor 3300.

[0341] The communication interface 3100 can include a communication circuit capable of performing data communication between the server 3000 and another electronic device (e.g., the electronic device 2000) using at least one of data communication methods including, for example, a wired local area network (LAN), a wireless LAN, Wi-Fi, Bluetooth, ZigBee, WFD, IrDA, BLE, NFC, Wibro, WiMAX, SWAP, WiGig, and RF communication.

[0342] The communication interface 3100 can transmit and receive data for providing a personalized image to and from the electronic device 2000 under the control of the processor 3300. For example, the server 3000 can receive a prompt (or prompt embedding) and a reference embedding from the electronic device 2000 through the communication interface 3100, and can transmit a personalized image to the electronic device 2000.

[0343] The memory 3200 can store instructions, data structures, and program codes readable by the processor 3300. The operations performed by the processor 3300 can be performed by executing the instructions or codes of the program stored in the memory 3200.

[0344] The memory 3200 can include a non-volatile memory including at least one of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., an SD or xD memory), a ROM, an EEPROM, a PROM, a magnetic memory, a magnetic disk, or an optical disk, and a volatile memory such as a RAM or an SRAM.

[0345] The memory 3200 can store one or more instructions and / or programs that allow the server 3000 to operate to generate a personalized image. For example, the memory 3200 can store instructions and / or programs for performing the functions of the image generation module 3210. The image generation module 3210 can include a generation model.

[0346] The processor 3300 can control the overall operation of the server 3000. For example, the processor 3300 can control the overall operation of the server 3000 to generate a personalized image by executing one or more instructions of a program stored in the memory 3200. There can be one or more processors 3300.

[0347] The processor 3300 can include, for example, and without limitation, at least one of a central processing unit, a microprocessor, a graphic processing unit, an ASIC, a DSP, a DSPD, a PLD, an FPGA, an AP, a neural processing unit, or a dedicated AI processor designed with a hardware structure dedicated to processing an AI model.

[0348] The processor 3300 can generate a personalized image by executing the image generation module 3210. The processor 3300 can generate a personalized image by using a generation model that uses a reference embedding as input data. The operation of the image generation module 3210 has been described with reference to the preceding figures, and thus, a repeated description will be omitted.

[0349] Figure 19b is a block diagram illustrating a configuration of a server according to an embodiment of the disclosure.

[0350] In an embodiment of the disclosure, the operations performed by the electronic device 2000 and the server 3000 can be performed by the server 3000 alone. The server 3000 can obtain a prompt from a user, can obtain an image stored in the electronic device 2000 of the user as a reference image, and can generate a personalized image based on the reference image. In this case, the server 3000 can obtain access to a media storage of the electronic device 2000 of the user.

[0351] In Figure 19bIn the following description, for the sake of brevity, it will be assumed that the same description as that performed in the foregoing description will be omitted. Figure 19a

[0352] In an embodiment of the disclosure, instructions and / or programs for performing the functions of the keyword extraction module 3220, the image search module 3230, and the image segmentation module 3240 can be stored in the memory 3200. The image generation module 3210 can include a generation model. Furthermore, a prompt generation module (not shown) can be further stored in the memory 3200.

[0353] The keyword extraction module 3220, the image search module 3230, the image segmentation module 3240, and the prompt generation module can be executed by the processor 3300. The operation of the above modules has been described with reference to the foregoing drawings, and thus repetitive description will be omitted.

[0354] The disclosure relates to a method, an electronic device, and a server for generating and providing a personalized image. Furthermore, the disclosure relates to a method of generating a personalized image in which a deployed generation model reflects features of only images owned by a user without retraining the generation model. However, the technical objects to be achieved by the disclosure are not limited thereto, and other technical objects not mentioned will be apparent to those skilled in the art from the description of the disclosure.

[0355] According to an aspect of the disclosure, a method of providing a personalized image by an electronic device can be provided.

[0356] The method can include obtaining a prompt for image generation.

[0357] The method can include extracting a keyword from the prompt.

[0358] The method can include searching for one or more images corresponding to the keyword from among a plurality of images stored in the electronic device.

[0359] The method can include selecting a reference image from among the searched and retrieved one or more images.

[0360] The method can include segmenting a region related to the keyword within the reference image.

[0361] The method can include generating a reference embedding by converting a patch of the reference image into a vector representation.

[0362] The method can include transmitting the prompt and the reference embedding to a server.

[0363] The method can include receiving, from the server, a personalized image generated using a generation model that receives the prompt and the reference embedding as input.

[0364] The generation model can be deployed after completion of training.​

[0365] The generation model can be configured to use the reference embedding to generate the personalized image only during an inference operation using the generation model.

[0366] The personalized image can include a feature of the reference image that was not used to train the generation model.

[0367] Extracting the keywords can include displaying the one or more keywords extracted from the prompt.

[0368] Extracting the keywords can include determining, based on user input, a keyword for personalization from the one or more keywords.

[0369] Selecting the reference image can include displaying the one or more images searched and retrieved.

[0370] Selecting the reference image can include determining, based on user input, the one or more reference images.

[0371] Selecting the reference image can include storing the keyword and the reference image corresponding to the keyword as a text-image pair.

[0372] The method can include applying, based on user input, a weight to the reference image.

[0373] The method can include storing the reference embedding.

[0374] The method can include in response to another request to generate a personalized image after the reference embedding is stored, visualizing and displaying the stored reference embedding.

[0375] The method can include extracting, from the reference image, a textual description of the reference image.

[0376] The method can include changing the prompt based on the textual description of the reference image.

[0377] Generating the reference embedding can include converting the textual description of the reference image to a vector representation.

[0378] The method can include detecting one or more products included in the personalized image.

[0379] The method can include displaying information related to the one or more products.

[0380] The method can include displaying the personalized image.

[0381] According to an aspect of the disclosure, an electronic device for providing a personalized image can be provided.

[0382] The electronic device can include a communication interface, a memory in which one or more instructions are stored, and one or more processors configured to execute the one or more instructions stored in the memory.

[0383] The one or more processors can be configured to obtain a prompt for image generation by executing one or more instructions.

[0384] The one or more processors can be configured to extract a keyword from the prompt by executing one or more instructions.

[0385] The one or more processors can search for one or more images corresponding to the keyword from among a plurality of images stored in the electronic device by executing one or more instructions.

[0386] The one or more processors can be configured to select a reference image from among the searched and retrieved one or more images by executing one or more instructions.

[0387] The one or more processors can be configured to segment a region related to the keyword within the reference image by executing one or more instructions.

[0388] The one or more processors can be configured to generate a reference embedding by converting a patch of the reference image into a vector representation by executing one or more instructions.

[0389] The one or more processors can be configured to transmit the prompt and the reference embedding to a server through a communication interface by executing one or more instructions.

[0390] The one or more processors can be configured to receive, through the communication interface, a personalized image generated by a generation model using the received prompt and the reference embedding as input by executing one or more instructions.

[0391] The generation model can be deployed after completion of training.

[0392] The generation model can be configured to use the reference embedding to generate the personalized image only during an inference operation using the generation model.

[0393] The personalized image can include a feature of the reference image that was not used to train the generation model.

[0394] The one or more processors can be configured to display the one or more keywords extracted from the prompt by executing one or more instructions.

[0395] The one or more processors can be configured to determine a keyword for personalization from among the one or more keywords based on a user input by executing one or more instructions.

[0396] The one or more processors can be configured to display the searched and retrieved one or more images by executing one or more instructions.

[0397] The one or more processors can be configured to determine one or more reference images based on the user input by executing one or more instructions.

[0398] The one or more processors can be configured to store the keyword and the reference image corresponding to the keyword as a text-image pair by executing one or more instructions.

[0399] The one or more processors can be configured to apply a weight to the reference image based on the user input by executing one or more instructions.

[0400] The one or more processors can be configured to store the reference embedding by executing one or more instructions.

[0401] The one or more processors can be configured to visualize and display the stored reference embedding in response to another request for generating the personalized image after the reference embedding is stored by executing one or more instructions.

[0402] The one or more processors can be configured to extract a text description of the reference image from the reference image by executing one or more instructions.

[0403] The one or more processors can be configured to change the prompt based on the text description of the reference image by executing one or more instructions.

[0404] The one or more processors can be configured to convert the text description of the reference image into a vector representation by executing one or more instructions.

[0405] The one or more processors can be configured to detect one or more products included in the personalized image by executing one or more instructions.

[0406] The one or more processors can be configured to display information related to the one or more products by executing one or more instructions.

[0407] The one or more processors can be configured to display the personalized image by executing one or more instructions.

[0408] According to an aspect of the disclosure, a method of providing a personalized image by a server can be provided.

[0409] The method can include obtaining a prompt for image generation.

[0410] The method can include extracting a keyword from the prompt.

[0411] The method can include searching for one or more reference images corresponding to the keyword from a plurality of images stored in the electronic device.

[0412] The method can include segmenting an area related to the keyword in the reference image.

[0413] The method can include obtaining image feature information based on a region of the reference image related to the keyword.

[0414] The method can include generating the personalized image by using a generative model.

[0415] The generative model can be an artificial intelligence model configured to receive the prompt and the image feature information as input and output the personalized image.

[0416] According to an aspect of the disclosure, a server for providing a personalized image can be provided.

[0417] The server can include a communication interface, a memory in which one or more instructions are stored, and one or more processors configured to execute the one or more instructions.

[0418] The one or more processors can be configured to obtain a prompt for image generation by executing the one or more instructions.

[0419] The one or more processors can be configured to extract a keyword from the prompt by executing the one or more instructions.

[0420] The one or more processors can be configured to search for one or more reference images corresponding to the keyword from among a plurality of images stored in the electronic device by executing the one or more instructions.

[0421] The one or more processors can be configured to segment a region related to the keyword in the reference image by executing the one or more instructions.

[0422] The one or more processors can be configured to obtain image feature information based on the region related to the keyword in the reference image by executing the one or more instructions.

[0423] The one or more processors can be configured to generate the personalized image by using a generative model by executing the one or more instructions.

[0424] The generative model can be an artificial intelligence model configured to receive the prompt and the image feature information as input and output the personalized image.

[0425] Embodiments of the present disclosure can be implemented as a recording medium having computer-executable instructions such as program modules, which are executed by a computer. The computer-readable medium can be any available medium accessible by a computer, and examples thereof include all volatile and non-volatile media, and detachable and non-detachable media. Furthermore, examples of the computer-readable medium can include computer storage media and communication media. Examples of the computer storage media include all volatile and non-volatile media that have been implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, and other data, and detachable and non-detachable media. The communication media typically include computer-readable instructions, data structures, or other data modulated data signals such as program modules.

[0426] Furthermore, the computer-readable storage medium can be provided as a non-transitory storage medium. Here, "non-transitory" means that the storage medium does not include a signal (e.g., an electromagnetic wave), and is tangible, but does not distinguish whether data is semi-permanently or temporarily stored in the storage medium. For example, the "non-transitory storage medium" can include a buffer that temporarily stores data.

[0427] According to embodiments of the present disclosure, the method according to various embodiments of the present disclosure can be provided in a computer program product. The computer program product is a product that can be purchased between a seller and a buyer. The computer program product can be distributed in the form of a machine-readable storage medium (e.g., a compact disc read only memory (CD-ROM)) or online via an application store (e.g., downloaded or uploaded) or directly between two user devices (e.g., smart phones) (e.g., downloaded or uploaded). When distributed online, at least a portion of the computer program product (e.g., a downloadable application) can be temporarily generated or at least temporarily stored in a machine-readable storage medium such as a memory of a manufacturer's server, a server of an application store, or a relay server.

[0428] The above description of the present disclosure is provided only for explanation, and those of ordinary skill in the art will understand that various changes in form and details can be made therein without departing from the spirit and scope of the present disclosure defined by the following claims. Therefore, the above-described embodiments are merely examples in all aspects and are not limited thereto. For example, each component described as a single type can be implemented in a distributed manner, and similarly, components described as distributed can be implemented in a combined form.

[0429] The scope of the present disclosure is defined by the appended claims rather than the detailed description, and all changes or modifications within the scope of the appended claims and their equivalents will be construed as being included in the scope of the present disclosure.

Claims

1. A method for providing personalized images using an electronic device, the method comprising: Get hints for image generation; Extract keywords from the prompts; Search for one or more reference images corresponding to the keywords from a plurality of images stored in the electronic device; Segment the keyword-related region within each of the one or more reference images; Image feature information is obtained based on regions related to keywords; Send the prompt and image feature information to the server; as well as The system receives personalized images from the server, which are generated by providing prompts and image feature information as input to the generative model.

2. The method as described in claim 1, wherein, Obtain image feature information, including: Obtain text-image pairs containing the keyword and a reference image, representing the region corresponding to the keyword; and Reference embeddings are generated by converting text-image pairs into vector representations.

3. The method as described in any one of claims 1 and 2, wherein, The generative model is deployed after training and configured to generate personalized images using image feature information only during inference operations using the generative model. Personalized images include features from reference images that were not used to train the generative model.

4. The method according to any one of claims 1 to 3, wherein, Extract keywords, including: Display one or more keywords extracted from the prompt; and Based on user input, keywords for personalization are determined from one or more keywords.

5. The method according to any one of claims 1 to 4, wherein, Searching the one or more reference images includes: Search for one or more images that correspond to the keywords, and retrieve the one or more images; Display the one or more images; and Based on user input, the one or more reference images are determined.

6. The method according to any one of claims 1 to 5, further comprising: Store image feature information; as well as In response to another request to generate a personalized image after the image feature information has been stored, the stored image feature information is visualized and displayed.

7. The method according to any one of claims 1 to 6, further comprising: Extract text descriptions about each reference image; as well as The prompt is modified based on the text description of the reference image.

8. An electronic device for generating personalized images, the electronic device comprising: Communication interface; The memory is configured to store instructions; as well as One or more processors are configured to execute the instructions. When executed by the one or more processors, the instructions cause the electronic device to perform the following operations: Get hints for image generation. Extract keywords from the prompts. Search for one or more reference images corresponding to the keyword from a plurality of images stored in the electronic device. Within each of the one or more reference images, segment the region associated with the keyword. Image feature information is obtained based on regions related to keywords. The prompts and image feature information are sent to the server via the communication interface; Personalized images are received from the server via a communication interface. These personalized images are generated by providing prompts and image feature information as input to the generative model.

9. The electronic device as claimed in claim 8, wherein, When executed by the one or more processors, the instructions also cause the electronic device to perform the following operations: Obtain text-image pairs containing the keyword and a reference image, representing the region corresponding to the keyword; and Reference embeddings are generated by converting text-image pairs into vector representations.

10. The electronic device as claimed in any one of claims 8 and 9, wherein, The generative model is deployed after training and configured to generate personalized images using image feature information only during inference operations using the generative model. Personalized images include features from reference images that were not used to train the generative model.

11. The electronic device as claimed in any one of claims 8 to 10, wherein, When executed by the one or more processors, the instructions also cause the electronic device to perform the following operations: Display one or more keywords extracted from the prompt, and Based on user input, keywords for personalization are determined from one or more keywords.

12. The electronic device as claimed in any one of claims 8 to 11, wherein, When executed by the one or more processors, the instructions also cause the electronic device to perform the following operations: Search for one or more images corresponding to the keywords, and retrieve the one or more images. Display the one or more images, and Based on user input, the one or more reference images are determined.

13. The electronic device as claimed in any one of claims 8 to 12, wherein, When executed by the one or more processors, the instructions also cause the electronic device to perform the following operations: Storing image feature information, and In response to another request to generate a personalized image after the image feature information has been stored, the stored image feature information is visualized and displayed.

14. The electronic device as claimed in any one of claims 8 to 13, wherein, When executed by the one or more processors, the instructions also cause the electronic device to perform the following operations: Extract text descriptions about each reference image; and The prompt is modified based on the text description of the reference image.

15. A computer-readable recording medium having a program recorded thereon for performing the method as described in any one of claims 1 to 7 on a computer.