Electronic device and method for generating image using object information
By analyzing and selecting keywords from an original image to generate AI images, the device addresses the issue of misinterpretation in existing AI models, ensuring the generated images align with user intentions and improve scene representation.
Patent Information
- Application Number
- PCT/KR2024/017225
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-21
- Filing Date
- 2024-11-05
- Publication Date
- 2025-07-17
AI Technical Summary
Existing image generation using artificial intelligence models often fails to accurately interpret user intentions due to the lack of consideration for surrounding objects, resulting in images that do not match the intended scene or atmosphere.
An electronic device and method that analyzes an original image to extract keywords, displays relevant keywords for user selection, and generates an artificial intelligence image based on selected keywords and the original image, using a combination of image analysis and AI models to ensure the generated image aligns with the intended scene and atmosphere.
Enables the generation of images that complement or modify the original image by incorporating user-selected keywords, ensuring the final image accurately reflects the intended scene and atmosphere, thereby improving the quality of AI-generated images.
Smart Images

Figure KR2024017225_17072025_PF_FP_ABST
Abstract
Description
Electronic device and method for generating an image using object information
[0001] The following embodiments relate to a technique for generating an image using object information.
[0002] With the advancement of artificial intelligence technology, revolutionary changes are also taking place in the field of image processing. AI-generated images, in particular, are a field attracting attention in the field of computer vision. AI-generated images are created through the computer's ability to interpret and understand images using machine learning and deep learning algorithms.
[0003] By configuring a prompt using text and entering it, the AI model can generate an image that matches the prompt. Furthermore, by specifying the area the user wishes to create and then entering text as a prompt, the AI model can insert the generated image into the designated area.
[0004] To edit a photo using a generative AI model, a user first writes a text prompt expressing their desired intent, which is then used to generate an image. However, if the AI model fails to interpret the prompt as intended, the resulting image may differ from the intended intent. If the AI model generates the image based solely on the input prompt, without considering surrounding objects, it may produce an awkward image that doesn't blend in with the surrounding objects.
[0005] An electronic device and method for generating an image using object information according to an embodiment are proposed.
[0006] According to one embodiment, an electronic device includes a display unit; a memory; and a processor, wherein the processor displays a first image corresponding to an original image on the display unit in response to a user input, analyzes the first image to extract keywords, displays a plurality of keywords in relation to an object included in the first image, selects at least one keyword from among the plurality of keywords, and generates a second image corresponding to the artificial intelligence image using an artificial intelligence model based on at least a portion of the selected keyword and at least a portion of the first image.
[0007] According to one embodiment, a method for generating an artificial intelligence image may include: displaying a first image corresponding to an original image in response to a user input; analyzing the first image to extract keywords; displaying a plurality of keywords in relation to an object included in the first image; selecting at least one keyword from among the plurality of keywords; and generating a second image corresponding to the artificial intelligence image using an artificial intelligence model based on at least a portion of the selected keywords and at least a portion of the first image.
[0008] An artificial intelligence image generation method according to one embodiment can easily generate a new image that complements a lacking part or emphasizes or changes the feeling of an image through keywords appropriate for the scene, atmosphere, or object information of the image.
[0009] FIG. 1 is a schematic diagram illustrating a configuration of an electronic device according to one embodiment.
[0010] Figure 2 is a flowchart schematically illustrating a flow for generating an artificial intelligence image according to one embodiment.
[0011] Figure 3 is a flowchart illustrating a flow for generating an artificial intelligence image according to one embodiment.
[0012] FIG. 4 is a diagram illustrating an example of generating an artificial intelligence image by selecting a keyword according to one embodiment.
[0013] FIG. 5 is a diagram illustrating an example of generating an artificial intelligence image while changing the selection of keywords according to one embodiment.
[0014] FIG. 6 is a diagram illustrating an example of modifying an artificial intelligence image by loading a stored artificial intelligence image and changing the selection of keywords according to one embodiment.
[0015] FIG. 7 is a diagram illustrating an example of re-extracting keywords for generating an artificial intelligence image according to one embodiment.
[0016] FIG. 8 is a diagram illustrating an example of generating an artificial intelligence image using preferred keywords according to one embodiment.
[0017] FIG. 9 is a diagram illustrating an example of registering preferred keywords and generating an artificial intelligence image using the preferred keywords according to one embodiment.
[0018] FIG. 10 is a diagram illustrating an example of generating an artificial intelligence image using keywords included in a past category according to one embodiment.
[0019] FIG. 11 is a diagram illustrating an example of generating an artificial intelligence image using keywords included in a future category according to one embodiment.
[0020] FIG. 12 is a schematic diagram illustrating a configuration of an electronic device that generates an artificial intelligence image based on keywords according to one embodiment.
[0021] FIG. 13 is a block diagram of an electronic device within a network environment according to one embodiment.
[0022] Hereinafter, embodiments are described in detail with reference to the attached drawings. However, the embodiments may be modified in various ways, and the scope of the patent application is not limited or restricted by these embodiments. It should be understood that all modifications, equivalents, or alternatives to the embodiments are included within the scope of the patent application.
[0023] The terms used in the examples are for illustrative purposes only and should not be construed as limiting. Singular expressions include plural expressions unless the context clearly dictates otherwise. In this specification, terms such as "comprise" or "have" are intended to indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but should be understood to not preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0024] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art to which the embodiments pertain. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined herein.
[0025] In addition, when describing with reference to the attached drawings, identical components will be assigned the same reference numerals regardless of the drawing numbers, and redundant descriptions thereof will be omitted. When describing embodiments, if a detailed description of a related known technology is judged to unnecessarily obscure the gist of the embodiment, the detailed description will be omitted.
[0026] Additionally, terms such as first, second, A, B, (a), (b), etc. may be used to describe components of the embodiments. These terms are only intended to distinguish the components from other components, and the nature, order, or sequence of the components are not limited by the terms. When a component is described as being "connected," "coupled," or "connected" to another component, it should be understood that the component may be directly connected or connected to the other component, but another component may also be "connected," "coupled," or "connected" between each component.
[0027] Components included in one embodiment and components with common functions will be described using the same names in other embodiments. Unless otherwise stated, the descriptions given in one embodiment may also apply to other embodiments, and detailed descriptions will be omitted to the extent of overlap.
[0028] Hereinafter, an electronic device and method for generating an image using object information according to an embodiment of the present disclosure will be described in detail with reference to the attached FIGS. 1 to 13.
[0029] FIG. 1 is a schematic diagram illustrating a configuration of an electronic device according to one embodiment.
[0030] Referring to FIG. 1, an electronic device (100) may be configured to include a processor (110), a communication unit (120), a display unit (130), and a memory (140).
[0031] The communication unit (120) is a communication interface device including a receiver and a transmitter, which transmits and receives data wired or wirelessly. The communication unit (120) can communicate with an artificial intelligence server (150) that generates artificial intelligence images. In this case, the communication unit (120) may have a configuration corresponding to the communication module (1390) of FIG. 13, and the artificial intelligence server (150) may have a configuration corresponding to the server (1308) of FIG. 13.
[0032] The display unit (130) displays status information (or indicators), limited numbers and characters, moving pictures, and still pictures generated during the operation of the electronic device (100). In addition, the present disclosure may display a preview image, at least one artificial intelligence graphic object, and an artificial intelligence image.
[0033] The display unit (130) of the present disclosure may be a touch screen capable of touch input.
[0034] A touch screen includes a display unit that performs a screen output function and a touch sensor that performs a touch input function. Such a touch screen may have a structure in which a touch sensor is arranged on the front of the display unit. The display unit may be formed of a liquid crystal display (LCD), an organic light emitting diode (OLED), etc. The touch sensor performs a function of receiving a touch input, i.e., a touch event, a double touch event, a touch movement event, and a touch release event. That is, the touch sensor can generate a touch event when an object, for example, a user's finger, touches the touch sensor, and transmit the generated touch event to the processor (110). In addition, when a finger moves in a certain direction on the touch sensor while maintaining a touched state, the touch sensor can generate a touch movement event and transmit the touch movement event to the processor (110). The touch movement event can be divided into a flick event having a movement speed greater than a preset threshold and a drag event having a movement speed less than the threshold. In addition, after a touch event or touch movement event is generated, if the user's finger moves away from the touch sensor, the touch sensor can generate a touch release event and transmit it to the processor (110). Such a touch sensor can be formed by a pressure-sensitive method, an infrared method, a capacitive method, etc. Hereinafter, for the convenience of explanation, the display unit (130) will be described collectively as a touch screen. At this time, the display unit (130) may have a configuration corresponding to the display module (1360) of FIG. 13.
[0035] The memory (140) stores an operating system, an application program, and data for storage (compressed image files, videos, etc.) for controlling the overall operation of the electronic device (100). In addition, the memory (140) may store an original image (hereinafter referred to as a 'first image'), metadata of the image, information about preferred keywords, accumulated information about selected keywords, an artificial intelligence image (hereinafter referred to as a 'second image'), metadata of the artificial intelligence image, a final artificial intelligence image received from an artificial intelligence server (150) (hereinafter referred to as a 'third image'), and metadata of the final artificial intelligence image, according to various embodiments of the present disclosure.
[0036] The processor (110) may be configured to include an image analysis unit (111), a keyword analysis unit (112), a prompt generation unit (113), an image generation unit (114), an area management unit (116), an image synthesis unit (117), and an image storage unit (118).
[0037] At this time, the image analysis unit (111), keyword analysis unit (112), prompt generation unit (113), image generation unit (114), area management unit (116), image synthesis unit (117), and image storage unit (118) can be stored in the memory (140) in the form of instructions.
[0038] The image analysis unit (111) can analyze the first image to extract keywords by category, analyze the first image to extract keywords for objects included in the first image, and analyze the main scenes and atmosphere of the second image to extract keywords. At this time, keyword extraction can be performed by utilizing a Transformer-based artificial intelligence deep learning model (e.g., models such as Clip, blip, and Vit) that utilizes the first image as input to extract keywords in text form from the first image.
[0039] The keyword analysis unit (112) can analyze keywords obtained through the image analysis unit (111) using an artificial intelligence model and determine keywords to be displayed that correspond to scenes, atmospheres, and objects with high relevance (priority) among the extracted keywords. The keyword analysis unit (112) can classify keywords related to the recognized object around the object recognized in the first image and display them on the display unit (130).
[0040] The keyword analysis unit (112) may also display registered preferred keywords together with multiple keywords on the display unit (130). At this time, the preferred keywords may be keywords that are registered in advance and preferred by the user, or may be keywords that are selected a preset number of times in descending order of frequency among previously selected keywords, or may be keywords that are selected a preset number of times or more among previously selected keywords.
[0041] When the keyword analysis unit (112) receives a request to change a plurality of displayed keywords, it can re-select a plurality of keywords to be displayed from the extracted keywords, or request the image analysis unit (111) to extract keywords again and re-select a plurality of keywords to be displayed from the newly extracted keywords.
[0042] At this time, the keyword analysis unit (112) can utilize a deep learning model such as Transformer or CNN.
[0043] The prompt generation unit (113) can generate a prompt to be used in generative artificial intelligence by using a keyword selected by the user from among a plurality of keywords obtained through the keyword analysis unit (112) and image data as needed. The prompt generation unit (113) can, depending on the category of the selected keyword, obtain a first image, text-type data from the artificial intelligence server (150), or data from the memory (140) of the electronic device (100) and utilize the same to generate a prompt. At this time, the prompt generation unit (113) can utilize a deep learning model trained for the corresponding purpose based on a Transformer, etc.
[0044] The image generation unit (114) can generate an image that matches the selected keyword using the prompt generated by the prompt generation unit (113) as input using the generative artificial intelligence.
[0045] The area management unit (116) can estimate the depth data of each object based on the information of the image, and can estimate the area data by analyzing the area by keyword in the image.
[0046] The image synthesis unit (117) can generate a second image, which is an artificial intelligence image, by synthesizing the image generated by the image generation unit (114) with the image generated by the image generation unit (114) using the depth data and area data estimated through the area management unit (116) and the keyword to the first image.
[0047] The image synthesis unit (117) may request the artificial intelligence server (150) to generate a final second image through the communication unit (120), and may also receive a third image corresponding to the final artificial intelligence image from the artificial intelligence server (150). At this time, the received third image may have a higher resolution than the second image generated through the image synthesis unit (117).
[0048] The image storage unit (118) can be controlled to store the second image or the third image in the memory (140).
[0049] At this time, the image storage unit (118) can be controlled to store the metadata of the second image together with the second image. At this time, the metadata of the second image can include at least one of information about the first image, which is the source data used to generate the second image (for example, the metadata information of the first image can include information about the time the first image was taken, information about the location where the first image was taken, information about the owner of the first image, and information about the size of the first image), information about an object included in the first image, extracted keywords, a plurality of displayed keywords, selected keywords, and category information corresponding to each keyword.
[0050] The processor (110) can control the overall operation of the electronic device (100). In addition, the processor (110) can perform the functions of an image analysis unit (111), a keyword analysis unit (112), a prompt generation unit (113), an image generation unit (114), an area management unit (116), an image synthesis unit (117), and an image storage unit (118). The image analysis unit (111), the keyword analysis unit (112), the prompt generation unit (113), the image generation unit (114), the area management unit (116), the image synthesis unit (117), and the image storage unit (118) are illustrated separately to explain each function separately. Accordingly, the processor (110) may include at least one processor configured to perform each function of the image analysis unit (111), the keyword analysis unit (112), the prompt generation unit (113), the image generation unit (114), the area management unit (116), the image synthesis unit (117), and the image storage unit (118). In addition, the processor (110) may include at least one processor configured to perform some of the functions of the image analysis unit (111), the keyword analysis unit (112), the prompt generation unit (113), the image generation unit (114), the area management unit (116), the image synthesis unit (117), and the image storage unit (118). At this time, the processor (110) may have a configuration corresponding to the processor (1380) of FIG. 13.
[0051] Hereinafter, the method according to the present disclosure configured as above will be described with reference to the drawings below.
[0052] Figure 2 is a flowchart schematically illustrating a flow for generating an artificial intelligence image according to one embodiment.
[0053] Referring to FIG. 2, in operation 210, the electronic device (100) can display a first image, which is an original image for generating an artificial intelligence image in response to a user input.
[0054] And, in operation 220, the electronic device (100) can analyze the first image and metadata of the first image to extract keywords.
[0055] And, in operation 230, the electronic device (100) may display a plurality of keywords related to an object included in the first image by overlaying them on an object corresponding to the plurality of keywords of the first image or displaying them around the object.
[0056] And, in operation 240, when the electronic device (100) receives a selection of a keyword to be used for generating a second image, which is an artificial intelligence image, from among a plurality of keywords, the electronic device (100) can generate a second image, which is an artificial intelligence image, based on at least a portion of the keywords selected in operation 250 and at least a portion of the image. At this time, in operation 250, the selection of the keyword to be used for the second image can be performed by detecting a user input through a touch screen to determine which keyword has been selected from among the plurality of keywords.
[0057] Figure 3 is a flowchart illustrating a flow for generating an artificial intelligence image according to one embodiment.
[0058] Referring to FIG. 3, in operation 310, the electronic device (100) may display a first image, which is an original image for generating an artificial intelligence image in response to a user input.
[0059] And, in operation 312, the electronic device (100) can analyze the first image to extract keywords.
[0060] In operation 312, the electronic device (100) may further analyze metadata of the first image in addition to the first image to extract keywords. At this time, the metadata of the first image may include at least one of information on the time the first image was taken, information on the location where the first image was taken, information on the owner of the first image, and information on the size of the first image.
[0061] In operation 312 according to one embodiment, the electronic device (100) may identify and recognize at least one object included in a first image, analyze the recognized object, and extract at least one keyword corresponding to the recognized object and a category corresponding to each keyword. At this time, the category may include at least one of a category classifying a scene, a category classifying an atmosphere, a category classifying a type of object, a category indicating that at least a part of the object can be changed to the past, and a category indicating that at least a part of the object can be changed to the future.
[0062] In operation 312, the electronic device (100) can extract keywords using an artificial intelligence model. At this time, the artificial intelligence model used can extract keywords in text form from the first image by utilizing a Transformer-based deep learning model (e.g., models such as Clip, blip, and Vit).
[0063] In addition, in operation 314, the electronic device (100) may display a plurality of keywords related to an object included in the first image. In operation 314, when displaying a plurality of keywords, the electronic device (100) may classify and display keywords related to the recognized object around the recognized object.
[0064] In operation 314, the electronic device (100) may select multiple keywords to be displayed on the first image from among the extracted keywords. At this time, the selection of multiple keywords may take into account at least one of information regarding the user's preferences, information regarding the frequency of keyword usage, and information regarding the priority of preset keywords. The artificial intelligence model used to select multiple keywords may utilize a deep learning model such as a Transformer or CNN.
[0065] In addition, in operation 314, the electronic device (100) may display registered preferred keywords together with multiple keywords. At this time, the preferred keywords may be keywords that are registered in advance and preferred by the user, or a preset number of keywords selected in descending order of frequency among previously selected keywords, or keywords that are selected more frequently than a preset number of times among previously selected keywords.
[0066] And, in operation 316, the electronic device (100) can check whether a request to change the displayed plurality of keywords is received. At this time, the request to change the displayed plurality of keywords can be received when a keyword reset button displayed on a touch screen is selected by user input, or when a keyword reset menu is selected through a selection menu.
[0067] When a request to change multiple keywords displayed in the confirmation result of operation 316 is received, the electronic device (100) may return to operation 312 and perform a series of operations. That is, when the electronic device (100) receives a request to change multiple keywords displayed in operation 316, the electronic device may return to operation 312, analyze the image to extract keywords again, and reselect multiple keywords to be displayed from among the keywords extracted again in operation 314.
[0068] If a request to change the displayed plurality of keywords is not received as a result of the confirmation of operation 316, in operation 318, the electronic device (100) can confirm whether a selection of a keyword reflected in the generation of the second image, which is an artificial intelligence image, is received from among the displayed plurality of keywords. In operation 318, it can be confirmed whether the selection of the keyword is selected in response to a user's touch input for the plurality of keywords displayed on the display unit.
[0069] If the selection of a keyword to be broadcast in the creation of the second image is not received among the multiple keywords displayed as a result of the confirmation of operation 318, the process returns to operation 316 and a series of operations can be performed.
[0070] When a selection of a keyword to be broadcast for generating a second image is received from among the plurality of keywords displayed as a result of the confirmation of operation 318, in operation 320, the electronic device (100) can generate a second image, which is an artificial intelligence image, based on the selected keyword.
[0071] In operation 320, the electronic device (100) may generate a prompt based on a selected keyword to generate a second image, and input the generated prompt into an artificial intelligence model that generates the second image to generate the second image. At this time, the artificial intelligence model used to generate the prompt may utilize a deep learning model trained for the corresponding purpose based on a Transformer, and the artificial intelligence model that generates the second image may utilize a model such as Diffusion or GAN.
[0072] For example, when the keywords "full moon", "mountain", "aurora", and "cloud removal" extracted from a photo of clouds floating in a night sky with a crescent moon are selected, the electronic device (100) can generate a prompt in the form of a sentence, such as "there is a mountain, a full moon is floating over the mountain, and a cloudless sky with an aurora visible together with the full moon."
[0073] And, in operation 322, the electronic device (100) can check whether a request to change the selected keyword is received.
[0074] When a request to change a selected keyword is received as a result of confirmation of operation 322, the electronic device (100) can display a plurality of keywords by overlaying them on a second image in operation 324 and return to operation 316 to perform a series of operations.
[0075] If a request to change the selected keyword is not received as a result of the confirmation in operation 322, the electronic device (100) can confirm in operation 326 whether a request to generate a third image, which is the final artificial intelligence image, is received. At this time, the second image is a type of preview image generated to quickly confirm the user's request before generating the third image. Accordingly, the second image may be a low-resolution image, and the third image may be a high-resolution image compared to the second image.
[0076] If a request for creating a third image is not received as a result of the confirmation of operation 326, the electronic device (100) can proceed to operation 330 and perform subsequent operations.
[0077] If a request for generation of a third image is received as a result of confirmation of operation 326, in operation 328, the electronic device (100) may transmit the second image, metadata of the second image, and selected keywords to an external artificial intelligence server (150) to request generation of a third image, and receive the third image from the external artificial intelligence server (150). At this time, the third image generated by the external artificial intelligence server (150) may have a higher resolution than the second image generated by the electronic device (100).
[0078] And, when the electronic device (100) receives a storage request for the second image or the third image in operation 330, it can store the second image or the third image together with the corresponding metadata in operation 332. At this time, the metadata of the second image or the third image may include at least one of information about the first image, which is the source data used to generate the second or third image (for example, the metadata information of the first image may include information about the time the first image was taken, information about the location where the first image was taken, information about the owner of the first image, and information about the size of the first image), object information included in the second or third image, extracted keywords, a plurality of displayed keywords, selected keywords, and category information corresponding to each keyword.
[0079] Meanwhile, the electronic device (100) recognizes an object included in the first image in operation 312, and if a person is included in the recognized object, it determines whether the recognized object is a person stored in the memory, and if the recognized object is a person stored in the memory, it can extract keywords related to the past or future of the person. In addition, if the selected keywords in operation 320 include keywords related to the past or future of the person, it can generate an artificial intelligence image by considering information about the person or an image of the person stored in the memory.
[0080] FIG. 4 is a diagram illustrating an example of generating an artificial intelligence image by selecting a keyword according to one embodiment.
[0081] Referring to FIG. 4, when a user selects a first image of a moon and clouds floating in the night sky as in the first drawing (410) and issues a prompt generation command (e.g., input of a shooting button, a voice command, or selection of a prompt generation menu) to create a second image, the electronic device (100) analyzes at least one object in the first image to extract keywords, and selects multiple keywords to be displayed in the first image from among the extracted keywords to display on the screen as in the middle drawing (420). At this time, the extracted keywords may be generated through analysis of metadata or objects included in the first image.
[0082] The electronic device (100) can generate a prompt using a selected keyword (full moon, erase, mountain, aurora) selected by the user from among a plurality of keywords displayed as in the second drawing (420).
[0083] For example, the electronic device (100) may generate and display multiple keywords such as 'stars' and 'full moon' that are related to the moon near the moon, as shown in the second drawing (420). In addition, the electronic device (100) may include multiple keywords such as 'erase' to confirm to the user whether to erase the object. In addition, the electronic device (100) may analyze the atmosphere of the night sky and include additional words such as 'galaxy' and 'aurora' as multiple keywords to display.
[0084] In addition, the electronic device (100) can generate a prompt based on a keyword selected by the user from among the displayed multiple keywords. For example, if the user selects "Full moon" from among the multiple keywords near the moon, "erase" from near clouds, and "Mountain" and "Aurora" from the margin, a prompt related to these (e.g., "Full moon over the Mountain with Aurora") can be generated, and this prompt can be displayed on the display.
[0085] The electronic device (100) may generate a second image of the third picture (430) corresponding to the artificial intelligence image by referring to the generated prompt and the first image of the first picture (410). Alternatively, the electronic device (100) may transmit the generated prompt and the first image of the first picture (410) to an externally located generating artificial intelligence server (150), and receive and display the second image of the third picture (430), which is the artificial intelligence image, from the artificial intelligence server (150).
[0086] FIG. 5 is a diagram illustrating an example of generating an artificial intelligence image while changing the selection of keywords according to one embodiment.
[0087] Referring to FIG. 5, the first drawing (510) shows an example of displaying keywords related to an object around an object or displaying them as an overlay on a first image, and displaying a keyword selected by a user from among the plurality of keywords.
[0088] The second image (520) displays multiple keywords and the selected keyword as an overlay on the second image, which is an artificial intelligence image generated by the keyword selected in the first image (510), thereby enabling the user to easily change the selected keyword. In this case, the selected keywords may be full moon, cloud erase, mountain, and aurora.
[0089] The third picture (530) shows the second image changed when the keyword selected in the second picture (520) is changed, and multiple keywords and the changed selected keyword are overlaid on the changed artificial intelligence image, the second-first image. In this case, the changed selected keywords may be stars, cloud erase, and galaxy.
[0090] The fourth picture (540) is the third image, which is the final artificial intelligence image, created by using the second-first image generated in the third picture (530). At this time, the third image may also be created through the artificial intelligence server (150).
[0091] At this time, the second and second-1 images of the second picture (520) and the third picture (530) may be low-resolution images, and the third image generated in the fourth picture (540) may be a high-resolution image.
[0092] FIG. 6 is a diagram illustrating an example of modifying an artificial intelligence image by loading a stored artificial intelligence image and changing the selection of keywords according to one embodiment.
[0093] Referring to FIG. 6, the first drawing (610) shows what is displayed to modify an artificial intelligence image (e.g., the third image of FIG. 5) stored in an electronic device (100).
[0094] The second figure (620) shows an example of displaying multiple keywords and selected keywords using the metadata of the third image in the first figure (610). In this case, the displayed multiple keywords may be stars, erase of the moon, full moon, snow, thunder, rain, erase of clouds, galaxy, mountain, and aurora. In addition, the selected keywords among the multiple keywords may be stars, erase of clouds, galaxy, and the fact that they have been selected may be displayed to distinguish them from unselected keywords.
[0095] The third figure (630) shows an example of restoring the original image (e.g., the first image of Fig. 5) without displaying the third image in the second figure (620), and displaying multiple keywords and a selected keyword over the first image.
[0096] The fourth figure (640) shows an example of creating a new artificial intelligence image, the 2-2 image, by selecting a keyword that was not selected from among multiple keywords.
[0097] That is, the electronic device (100) according to various embodiments of the present disclosure can easily perform modification using an artificial intelligence image (e.g., the third image of FIG. 5) stored in the electronic device (100).
[0098] FIG. 7 is a diagram illustrating an example of re-extracting keywords for generating an artificial intelligence image according to one embodiment.
[0099] Referring to FIG. 7, the first picture (710) displays a plurality of keywords and selected keywords on an artificial intelligence image (e.g., the second image (520) of FIG. 5), and a keyword change request key (711) can be displayed to change the displayed plurality of keywords.
[0100] The second figure (720) shows a case where a keyword change request key (711) is entered in the first figure (710), multiple keywords are changed, and newly selected.
[0101] According to the second figure (720), it can be confirmed that stars, snow, thunder, rain, and galaxy are deleted from the existing displayed multiple keywords by inputting the keyword change request key (711), and comet, waterfall, lake, fantasy, and pastel are added as new multiple keywords.
[0102] The third picture (730) shows the 2-3 image, which is an artificial intelligence image generated when keywords selected from among the multiple keywords changed in the second picture (720), namely comet, full moon, cloud erase, mountain, and aurora, are selected.
[0103] That is, the electronic device (100) can change the displayed multiple keywords through a reset button such as a keyword change request key (711) when the user wants to change the displayed multiple keywords.
[0104] FIG. 8 is a diagram illustrating an example of generating an artificial intelligence image using preferred keywords according to one embodiment.
[0105] Referring to FIG. 8, the first image (810) is the first image, which is the original image that serves as the source for generating an artificial intelligence image, and is an image selected by the user. At this time, the electronic device (100) can receive input from the user for generating an artificial intelligence image by displaying a command for generating a prompt for generating an artificial intelligence image on the first image.
[0106] The second figure (820) shows an example of displaying multiple keywords, preferred keywords, and selected keywords in the first figure (810). At this time, when displaying multiple keywords, the electronic device (100) may display preferred keywords (821) (e.g., pastel and cloud) in addition to the multiple keywords, and may determine a selected keyword based on the user's selection among the multiple keywords and preferred keywords.
[0107] The third picture (830) is the second image, which is an artificial intelligence image generated by selecting stars, full moon, pastel, and cloud as keywords in the second picture (820).
[0108] That is, when the electronic device (100) provides multiple keywords to be displayed in order to generate an artificial intelligence image, it can reduce the number of times the user resets multiple keywords by providing the preferred keywords preferred by the user together with the preferred keywords, thereby allowing the user to select the preferred keywords. At this time, the preferred keywords may be keywords that are registered in advance and preferred by the user, or may be keywords that are selected a preset number of times in descending order of frequency among previously selected keywords, or may be keywords that are selected a preset number of times or more among previously selected keywords.
[0109] FIG. 9 is a diagram illustrating an example of registering preferred keywords and generating an artificial intelligence image using the preferred keywords according to one embodiment.
[0110] Referring to FIG. 9, the first drawing (910) shows an example in which a plurality of keywords are displayed on the first image, which is the original image that serves as a source for generating an artificial intelligence image, and a preferred keyword is selected and registered among the plurality of keywords.
[0111] The second figure (920) shows an example showing multiple keywords, preferred keywords and selected keywords.
[0112] The third picture (930) is a second image, which is an artificial intelligence image generated by selecting the keywords erase, stars, cherry blossom, and galaxy of the moon from among the multiple keywords and preferred keywords of the second picture (920).
[0113] That is, the electronic device (100) can allow the user to add a keyword preferred by the user among a plurality of keywords to a list of preferred keywords.
[0114] FIG. 10 is a diagram illustrating an example of generating an artificial intelligence image using keywords included in a past category according to one embodiment.
[0115] Referring to FIG. 10, the first drawing (1010) shows an example in which multiple keywords are overlaid on the first image, which is the original image, and a keyword is selected from among the multiple keywords by the user.
[0116] At this time, the first picture (1010) can display the past of salad and the past of pizza included in the past category with multiple keywords.
[0117] The second image (1020) is the second image, which is an artificial intelligence image created by selecting the past of salad and the past of pizza as keywords.
[0118] Keywords can be categorized into various types, as follows. For example, keywords can include at least one of a category that classifies a scene, a category that classifies an atmosphere, a category that classifies a type of object, a category that indicates a possible change to the past, and a category that indicates a possible change to the future.
[0119] In Fig. 10, the past of salad and the past of pizza are categories indicating that changes to the past are possible. When the past of salad and the past of pizza are determined as selected keywords, the electronic device (100) can generate a second image that returns the image of the already eaten salad and pizza to the past state of the salad and pizza before eating them.
[0120] FIG. 11 is a diagram illustrating an example of generating an artificial intelligence image using keywords included in a future category according to one embodiment.
[0121] Referring to FIG. 11, the first drawing (1110) shows an example in which multiple keywords are displayed as overlays on the first image, which is the original image, and a keyword is selected from among the multiple keywords by the user.
[0122] At this time, the first picture (1110) can display the future of place A, the future of people A, the future of people B, and the future of people C as keywords included in the future category with multiple keywords.
[0123] The second picture (1120) is the second image, which is an artificial intelligence image generated by selecting keywords such as the future of place A, the future of people A, the future of people B, and the future of people C.
[0124] In Fig. 11, the future of place A, the future of people A, the future of people B, and the future of people C are categories indicating that changes to the future are possible. When the future of place A, the future of people A, the future of people B, and the future of people C are determined as selected keywords, the electronic device (100) can generate images of people A, people B, and people C as children taken at place A in the past into second images of people A, people B, and people C as adults taken at place A after time has passed.
[0125] In FIG. 11, when the electronic device (100) detects that personal information about people B (e.g., recent images of people B, or 3D facial scan information) is stored in the memory (140), it displays privacy as a keyword and, when generating a second image, can generate a future appearance of people B using the stored personal information of people B.
[0126] That is, the electronic device (100) can create a future image using stored personal information.
[0127] FIG. 12 is a schematic diagram illustrating a configuration of an electronic device that generates an artificial intelligence image based on keywords according to one embodiment.
[0128] Referring to FIG. 12, an electronic device (1200) may be configured to include a processor (1210) and a memory (1220).
[0129] The memory (1220) stores an operating system, an application program, and storage data for controlling the overall operation of the electronic device (1200). In addition, the memory (1220) may store, according to the present disclosure, a first image as an original image, metadata of the first image, information about preferred keywords, accumulated information about selected keywords, a second image as an artificial intelligence image, metadata of the second image, a third image as a final artificial intelligence image, and metadata of the third image.
[0130] The processor (1210) may correspond to the processor (110) of the electronic device (100) of FIG. 1. That is, the processor (1210) may include the configuration of the processor (110) of FIG. 1.
[0131] The processor (1210) can display a first image corresponding to an original image on a display unit in response to a user input, display a plurality of keywords related to an object included in the first image, select at least one keyword from among the plurality of keywords, and generate a second image corresponding to the artificial intelligence image using an artificial intelligence model based on at least a portion of the selected keyword and at least a portion of the first image.
[0132] When the processor (1210) is requested to store a second image, the processor (1210) may generate metadata of the second image including at least a portion of the first image, keywords extracted from object information included in the first image, a plurality of keywords, selected keywords, or category information corresponding to each keyword, and store the second image and the metadata of the second image together in the memory (1210).
[0133] The processor (1210) is optional.
[0134] When displaying multiple keywords, the processor (1210) may display keywords related to the object around the object or display them by overlaying them.
[0135] When analyzing a first image to extract keywords, the processor (1210) can identify and recognize an object included in the first image, and if a person is included in the recognized object, it can check whether the recognized object is a person stored in memory, and if the recognized object is a person stored in memory, it can extract keywords related to the past or future of the person.
[0136] When generating a second image based on at least a portion of the selected keywords and at least a portion of the first image, the processor (1210) may generate the second image by considering information about the person or an image about the person stored in the memory, if the selected keywords include keywords related to the past or future of the person.
[0137] The processor (1210) can perform the operations of FIGS. 2 and 3. Therefore, a detailed description of the processor (1210) is omitted.
[0138] FIG. 13 is a block diagram of an electronic device within a network environment according to one embodiment.
[0139] Referring to FIG. 13, in a network environment (1300), an electronic device (1301) may communicate with an electronic device (1302) via a first network (1398) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (1304) or a server (1308) via a second network (1399) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (1301) may communicate with the electronic device (1304) via the server (1308). According to one embodiment, the electronic device (1301) may include a processor (1320), a memory (1330), an input module (1350), an audio output module (1355), a display module (1360), an audio module (1370), a sensor module (1376), an interface (1377), a connection terminal (1378), a haptic module (1379), a camera module (1380), a power management module (1388), a battery (1389), a communication module (1390), a subscriber identification module (1396), or an antenna module (1397). In one embodiment, the electronic device (1301) may omit at least one of these components (e.g., the connection terminal (1378)), or may have one or more other components added. In one embodiment, some of these components (e.g., sensor module (1376), camera module (1380), or antenna module (1397)) may be integrated into one component (e.g., display module (1360)).
[0140] The processor (1320) may control at least one other component (e.g., hardware or software component) of the electronic device (1301) connected to the processor (1320) by executing, for example, software (e.g., program (1340)), and may perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (1320) may store commands or data received from other components (e.g., sensor module (1376) or communication module (1390)) in a volatile memory (1332), process the commands or data stored in the volatile memory (1332), and store result data in a non-volatile memory (1334).
[0141] Meanwhile, the processor (1320) can perform the operation of the processor (110) of FIG. 1 or the operation of the processor (1210) of FIG. 12.
[0142] According to one embodiment, the processor (1320) may include a main processor (1321) (e.g., a central processing unit or an application processor) or an auxiliary processor (1323) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (1321). For example, when the electronic device (1301) includes the main processor (1321) and the auxiliary processor (1323), the auxiliary processor (1323) may be configured to use less power than the main processor (1321) or to be specialized for a given function. The auxiliary processor (1323) may be implemented separately from the main processor (1321) or as a part thereof.
[0143] The auxiliary processor (1323) may control at least a portion of functions or states associated with at least one component (e.g., the display module (1360), the sensor module (1376), or the communication module (1390)) of the electronic device (1301), for example, on behalf of the main processor (1321) while the main processor (1321) is in an inactive (e.g., sleep) state, or together with the main processor (1321) while the main processor (1321) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (1323) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (1380) or a communication module (1390)). In one embodiment, the auxiliary processor (1323) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (1301) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (1308)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0144] The memory (1330) can store various data used by at least one component (e.g., the processor (1320) or the sensor module (1376)) of the electronic device (1301). The data can include, for example, software (e.g., the program (1340)) and input data or output data for commands related thereto. The memory (1330) can include volatile memory (1332) or non-volatile memory (1334).
[0145] Meanwhile, the memory (1330) can perform the role of the memory (140) of FIG. 1 or the memory (1220) of FIG. 12.
[0146] The program (1340) may be stored as software in memory (1330) and may include, for example, an operating system (1342), middleware (1344), or an application (1346).
[0147] The input module (1350) can receive commands or data to be used in a component of the electronic device (1301) (e.g., a processor (1320)) from an external source (e.g., a user) of the electronic device (1301). The input module (1350) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0148] The audio output module (1355) can output audio signals to the outside of the electronic device (1301). The audio output module (1355) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. According to one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0149] The display module (1360) can visually provide information to an external party (e.g., a user) of the electronic device (1301). The display module (1360) may include, for example, a display, a holographic device, or a projector, and a control circuit for controlling the device. According to one embodiment, the display module (1360) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch. In this case, the display module (1360) may perform the role of the display unit (130) of FIG. 1.
[0150] The audio module (1370) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (1370) can acquire sound through the input module (1350), output sound through the sound output module (1355), or an external electronic device (e.g., electronic device (1302)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (1301).
[0151] The sensor module (1376) can detect the operating status (e.g., power or temperature) of the electronic device (1301) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (1376) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0152] The interface (1377) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (1301) with an external electronic device (e.g., the electronic device (1302)). In one embodiment, the interface (1377) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0153] The connection terminal (1378) may include a connector through which the electronic device (1301) may be physically connected to an external electronic device (e.g., the electronic device (1302)). According to one embodiment, the connection terminal (1378) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0154] The haptic module (1379) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (1379) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0155] The camera module (1380) can capture still images and videos. According to one embodiment, the camera module (1380) may include one or more lenses, image sensors, image signal processors, or flashes.
[0156] The power management module (1388) can manage power supplied to the electronic device (1301). According to one embodiment, the power management module (1388) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).
[0157] A battery (1389) may power at least one component of the electronic device (1301). In one embodiment, the battery (1389) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0158] The communication module (1390) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (1301) and an external electronic device (e.g., electronic device (1302), electronic device (1304), or server (1308)), and the performance of communication through the established communication channel. The communication module (1390) may operate independently from the processor (1320) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1390) may include a wireless communication module (1392) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (1394) (e.g., a local area network (LAN) communication module, or a power line communication module). Any of these communication modules may communicate with an external electronic device (1304) via a first network (1398) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (1399) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a local area network or a wide area network)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1392) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (1396) to verify or authenticate the electronic device (1301) within a communication network such as the first network (1398) or the second network (1399). The communication module (1390) can perform the role of the communication unit (120) of FIG. 1.
[0159] The wireless communication module (1392) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency communications (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (1392) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (1392) can support various technologies for securing performance in high-frequency bands, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (1392) can support various requirements specified in the electronic device (1301), an external electronic device (e.g., the electronic device (1304)), or a network system (e.g., the second network (1399)). According to one embodiment, the wireless communication module (1392) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0160] The antenna module (1397) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). According to one embodiment, the antenna module (1397) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (1397) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (1398) or the second network (1399), may be selected from the plurality of antennas by, for example, the communication module (1390). A signal or power may be transmitted or received between the communication module (1390) and an external electronic device via the selected at least one antenna. According to one embodiment, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (1397).
[0161] In one embodiment, the antenna module (1397) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.
[0162] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0163] According to one embodiment, commands or data may be transmitted or received between the electronic device (1301) and an external electronic device (1304) via a server (1308) connected to a second network (1399). Each of the external electronic devices (1302 or 1304) may be the same or a different type of device as the electronic device (1301). According to one embodiment, all or part of the operations executed in the electronic device (1301) may be executed in one or more of the external electronic devices (1302, 1304, or 1308). For example, when the electronic device (1301) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (1301) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (1301). The electronic device (1301) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (1301) may provide an ultra-low latency service using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (1304) may include an Internet of Things (IoT) device. The server (1308) may be an intelligent server utilizing machine learning and / or a neural network.According to one embodiment, an external electronic device (1304) or server (1308) may be included in the second network (1399). The electronic device (1301) may be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0164] Electronic devices according to embodiments of the present disclosure may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to embodiments of the present disclosure are not limited to the aforementioned devices.
[0165] The embodiments of the present disclosure and the terminology used herein are not intended to limit the technical features described in the present disclosure to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In the present disclosure, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among the phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0166] The term "module" used in one embodiment of the present disclosure may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0167] An embodiment of the present disclosure may be implemented as software (e.g., a program (1340)) including one or more instructions stored in a storage medium (e.g., an internal memory (1336) or an external memory (1338)) readable by a machine (e.g., an electronic device (1301)). For example, a processor (e.g., a processor (1320)) of the machine (e.g., an electronic device (1301)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0168] According to one embodiment, a method according to one embodiment of the present disclosure may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0169] According to one embodiment, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to one embodiment, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to one embodiment, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
[0170] According to one embodiment, an electronic device (100; 1200; 1300) includes a display unit (130; 1360); a memory (140, 1330); and a processor (110; 1210; 1320), wherein the processor displays a first image corresponding to an original image on the display unit (130; 1360) in response to a user input, analyzes the first image to extract a keyword, displays a plurality of keywords in relation to an object included in the first image, selects at least one keyword from among the plurality of keywords, and generates a second image corresponding to the artificial intelligence image using an artificial intelligence model based on at least a part of the selected keyword and at least a part of the first image.
[0171] According to one embodiment, when the processor (110; 1210; 1320) is requested to store the second image, the processor may generate metadata of the second image including at least a portion of the first image, object information included in the first image, the extracted keyword, the plurality of keywords, the selected keyword or category information corresponding to each keyword, and store the second image and the metadata of the second image together in the memory (140, 1330).
[0172] According to one embodiment, when displaying with the plurality of keywords, the processor (110; 1210; 1320) may select the plurality of keywords to be displayed from among the extracted keywords by considering at least one of information on user preference, frequency of use information of keywords, and priority information of preset keywords in the memory (140, 1330).
[0173] According to one embodiment, when displaying the plurality of keywords, the processor (110; 1210; 1320) may display keywords related to the object around the object or display them by overlaying them.
[0174] According to one embodiment, when analyzing the first image to extract keywords, the processor (110; 1210; 1320) may identify and recognize an object included in the first image, and if a person is included in the recognized object, determine whether the recognized object is a person stored in the memory, and if the recognized object is a person stored in the memory (140, 1330), extract keywords related to the past or future of the person.
[0175] According to one embodiment, when generating the second image based on at least a portion of the selected keyword and at least a portion of the first image, if the selected keyword includes a keyword related to the past or future of the person, the processor (110; 1210; 1320) may generate the second image by taking into consideration information about the person or an image about the person stored in the memory (140, 1330).
[0176] According to one embodiment, a method for generating an artificial intelligence image may include: displaying a first image corresponding to an original image in response to a user input; analyzing the first image to extract keywords; displaying a plurality of keywords in relation to an object included in the first image; selecting at least one keyword from among the plurality of keywords; and generating a second image corresponding to the artificial intelligence image using an artificial intelligence model based on at least a portion of the selected keywords and at least a portion of the first image.
[0177] According to one embodiment, the operation of analyzing the first image to extract a keyword includes an operation of analyzing metadata corresponding to the first image, and the metadata may include at least one of information on the time at which the first image was taken, information on the location at which the first image was taken, information on the owner of the first image, information on the size of the first image, information on an object included in the first image, the extracted keyword, the plurality of keywords, the selected keyword, and category information corresponding to each of the keywords.
[0178] According to one embodiment, the method for generating an artificial intelligence image may further include, when a request is made to store the second image, generating metadata of the second image including at least a portion of the first image, object information included in the first image, the extracted keyword, the plurality of keywords, the selected keyword or category information corresponding to each keyword; and storing the second image and the metadata of the second image together.
[0179] According to one embodiment, the operation of displaying the plurality of keywords in relation to the object included in the first image may include an operation of selecting the plurality of keywords to be displayed by considering at least one of information on user preference, information on frequency of use of keywords, and priority information of preset keywords among the extracted keywords.
[0180] According to one embodiment, the operation of displaying the plurality of keywords in relation to the object included in the first image may include an operation of displaying or overlaying keywords related to the object around the object.
[0181] According to one embodiment, the artificial intelligence image generation method may further include an operation of generating a new second image according to the changed keyword when a change in the selected keyword is requested.
[0182] According to one embodiment, the method for generating an artificial intelligence image further includes, when receiving a generation input of a third image corresponding to a final artificial intelligence image, generating the third image using an external artificial intelligence server, wherein the third image may have a higher resolution than the second image.
[0183] According to one embodiment, the artificial intelligence image generation method may further include, when receiving a request to change the plurality of keywords, analyzing at least a portion of the first image and metadata of the first image to re-extract the keywords, and re-selecting the plurality of keywords from among the re-extracted keywords.
[0184] According to one embodiment, the operation of displaying the plurality of keywords in relation to the object included in the first image may include an operation of displaying a preferred keyword registered together with the plurality of keywords.
[0185] According to one embodiment, the preferred keyword may be a keyword that is preferred by a user registered in advance, a preset number of keywords selected in descending order of frequency among previously selected keywords, or a keyword that is selected a preset number of times or more among previously selected keywords.
[0186] According to one embodiment, the operation of analyzing the first image to extract keywords includes: an operation of identifying and recognizing an object included in the first image; and an operation of analyzing the recognized object to extract keywords corresponding to the recognized object and categories corresponding to each keyword, wherein the categories may include at least one of a category classifying a scene, a category classifying an atmosphere, a category classifying a type of object, a category indicating that a change to the past is possible, and a category indicating that a change to the future is possible.
[0187] According to one embodiment, the operation of analyzing the first image to extract keywords may include: an operation of identifying and recognizing an object included in the first image; an operation of identifying whether the recognized object is a person stored in memory if the recognized object includes a person stored in memory; and an operation of extracting keywords related to the past or future of the person if the recognized object is a person stored in memory.
[0188] According to one embodiment, the operation of generating the second image based on at least a portion of the selected keywords and at least a portion of the first image may include an operation of generating the second image by taking into consideration information about the person or an image about the person stored in the memory, if the selected keywords include keywords related to the past or future of the person.
[0189] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may store program commands, data files, data structures, etc., singly or in combination. The program commands recorded on the medium may be those specially designed and configured for the embodiment or may be those known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands, such as ROMs, RAMs, and flash memories. Examples of the program commands include not only machine language codes generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operations of the embodiment, and vice versa.
[0190] Software may include a computer program, code, instructions, or a combination of one or more of these, and may configure a processing device to perform a desired operation or, independently or collectively, command the processing device. The software and / or data may be stored on any type of machine, component, physical device, virtual equipment, computer storage medium, or device, for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems, and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.
[0191] Although the embodiments described above have been described with limited drawings, those skilled in the art will appreciate that various technical modifications and variations can be applied based on the above. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.
[0192] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.
Claims
1. In electronic devices, Display section; memory; and Contains a processor, The above processor, In response to a user input, displaying a first image corresponding to an original image on the display unit, analyzing the first image to extract keywords, displaying a plurality of keywords related to an object included in the first image, selecting at least one keyword from the plurality of keywords, and generating a second image corresponding to the artificial intelligence image using an artificial intelligence model based on at least a part of the selected keyword and at least a part of the first image. Electronic devices.
2. In paragraph 1, The above processor, When a request is made to store the second image, metadata of the second image including at least a portion of the first image, object information included in the first image, the extracted keyword, the plurality of keywords, the selected keyword or category information corresponding to each keyword is generated, and the second image and the metadata of the second image are stored together in the memory. Electronic devices.
3. An action of displaying a first image corresponding to the original image in response to user input; An action of analyzing the first image above to extract keywords; An action of displaying multiple keywords in relation to an object included in the first image; An operation of selecting at least one keyword from among the above multiple keywords; and An operation of generating a second image corresponding to the artificial intelligence image using an artificial intelligence model based on at least a portion of the selected keywords and at least a portion of the first image. An artificial intelligence image generation method comprising:
4. In paragraph 3, The operation of analyzing the first image above and extracting keywords is as follows: An operation of analyzing metadata corresponding to the first image above. Including, The above metadata is, Information on the time at which the first image was taken, information on the location at which the first image was taken, information on the owner of the first image, information on the size of the first image, information on the object included in the first image, the extracted keyword, the plurality of keywords, the selected keyword, and category information corresponding to each keyword. An artificial intelligence image generation method comprising at least one of:
5. In any one of paragraphs 3 to 4, When a request is made to store the second image, an operation of generating metadata of the second image including at least a portion of the first image, object information included in the first image, the extracted keyword, the plurality of keywords, the selected keyword or category information corresponding to each keyword; and An operation of storing the second image and metadata of the second image together An artificial intelligence image generation method further comprising:
6. In any one of paragraphs 3 to 5, The operation of displaying the plurality of keywords in relation to the object included in the first image is: An operation for selecting multiple keywords to be displayed by considering at least one of the information on the user's preference, information on the frequency of keyword usage, and information on the priority of preset keywords among the extracted keywords. An artificial intelligence image generation method comprising:
7. In any one of paragraphs 3 to 6, The operation of displaying the plurality of keywords in relation to the object included in the first image is: An action to display keywords related to the object around the object or overlay them An artificial intelligence image generation method comprising:
8. In any one of paragraphs 3 to 7, When a change to the above selected keyword is requested, an operation to generate a new second image according to the changed keyword An artificial intelligence image generation method further comprising:
9. In any one of paragraphs 3 to 8, When receiving the input for generating a third image corresponding to the final AI image, an operation for generating the third image using an external AI server Including more, The third image above is, Has a higher resolution than the second image above How to generate artificial intelligence images.
10. In any one of paragraphs 3 to 9, When receiving a request to change the plurality of keywords, an operation of analyzing at least a part of the first image and metadata of the first image to re-extract the keywords, and re-selecting the plurality of keywords from among the re-extracted keywords An artificial intelligence image generation method further comprising:
11. In any one of paragraphs 3 to 10, The operation of displaying the plurality of keywords in relation to the object included in the first image is: Action to display registered preferred keywords together with the above multiple keywords An artificial intelligence image generation method comprising:
12. In any one of paragraphs 3 to 11, The above preferred keywords are, Keywords preferred by pre-registered users, or A preset number of keywords selected in order of frequency from among the previously selected keywords, or Among the keywords previously selected, a keyword whose frequency is greater than the preset number of times How to generate artificial intelligence images.
13. In any one of paragraphs 3 to 12, The operation of analyzing the first image above and extracting keywords is as follows: An action of recognizing and identifying an object contained in the first image; and An operation of analyzing the above recognized object and extracting keywords corresponding to the above recognized object and categories corresponding to each keyword. Including, The above categories are, Categories that classify scenes, Categories that classify the mood, Categories that classify the type of object, A category that indicates that changes to the past are possible. A category that indicates that changes to the future are possible. Contains at least one of How to generate artificial intelligence images.
14. In any one of paragraphs 3 to 13, The operation of analyzing the first image above and extracting keywords is as follows: An action of recognizing and confirming an object contained within the first image; If a person exists in the above recognized object, an operation of checking whether the recognized object is a person stored in memory; and If the above recognized object is a person stored in memory, an operation to extract keywords related to the past or future of the person. An artificial intelligence image generation method comprising:
15. In any one of paragraphs 3 to 14, An operation of generating the second image based on at least a portion of the selected keywords and at least a portion of the first image, If the selected keyword includes a keyword related to the past or future of the person, an operation of generating the second image by considering information about the person or an image of the person stored in the memory. An artificial intelligence image generation method comprising:
Citation Information
Patent Citations
Electronic device and method for providing filter n electronic device
KR1020160051390A
method and device for adjusting an image
KR1020180051367A
KR20200017263A