Electronic device and image management method of electronic device

The electronic device uses AI models to recognize user-defined points of interest and objects, generating detailed tags to enhance image metadata, addressing inefficiencies in existing image management systems by improving search and classification.

WO2026063646A1PCT designated stage Publication Date: 2026-03-26SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing image management systems lack the ability to efficiently and automatically update image metadata based on user input, leading to inefficiencies in image search and classification.

Method used

An electronic device equipped with artificial intelligence models, such as large language models (LLM) and large vision models (LVM), automatically modifies image metadata by recognizing user-defined points of interest and objects, generating detailed tags, and updating existing tags based on user input.

Benefits of technology

Enhances image management efficiency by providing more detailed and personalized metadata, improving search and classification capabilities through automated tag modification based on user interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025012740_26032026_PF_FP_ABST
    Figure KR2025012740_26032026_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device according to an embodiment of the present disclosure may: display an original image; receive a user input for the original image; recognize, on the basis of the user input, at least one of a point of interest (POI) of the original image and at least one object included in the original image; generate a second tag related to the at least one of the POI and the at least one object; and modify a first tag related to the original image on the basis of the second tag. Various other embodiments identified through the specification are possible.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device and image management method of the electronic device

[0001] The embodiments disclosed in this document relate to a technology for managing images of electronic devices and information related to images.

[0002] An electronic device may acquire images using a camera or provide editing functions for acquired images. For example, the electronic device may generate a new image based on user input, extract at least a portion of an image, or add new objects to an image. Images acquired and stored by the electronic device may include metadata related to the images. For example, metadata related to the images may include the date (capture) of the image, the time of acquisition, information related to camera settings, and / or information representing the image (e.g., tags). For example, metadata related to the images may be used to recognize, search, and / or classify the images.

[0003] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.

[0004] An electronic device according to one embodiment disclosed herein may include a display, a memory for storing instructions, and at least one processor. When the instructions are executed individually or collectively by the at least one processor, the electronic device may display an original image through the display, receive user input regarding the original image, recognize at least one of a point of interest (POI) of the original image or at least one object included in the original image based on the user input, generate a second tag associated with at least one of the point of interest or at least one object, and modify a first tag associated with the original image based on the second tag.

[0005] Additionally, a method according to one embodiment disclosed in this document may include: displaying an original image; receiving user input regarding the original image; recognizing at least one of a point of interest (POI) of the original image or at least one object included in the original image based on the user input; generating a second tag associated with at least one of the point of interest or at least one object; and modifying a first tag associated with the original image based on the second tag.

[0006] Additionally, a storage medium according to one embodiment disclosed in this document may store a program and / or instructions that, when executed by at least one processor of an electronic device, cause the electronic device to display an original image, receive user input regarding the original image, recognize at least one of a point of interest (POI) of the original image or at least one object included in the original image based on the user input, generate a second tag associated with at least one of the point of interest or at least one object, and modify a first tag associated with the original image based on the second tag.

[0007] FIG. 1 is a block diagram of an electronic device according to one embodiment.

[0008] FIG. 2 is a block diagram of an electronic device according to one embodiment.

[0009] FIG. 3 is a diagram illustrating an operation to modify tag information related to an image of an electronic device according to one embodiment.

[0010] FIG. 4 is a diagram illustrating an operation to modify tag information related to an image of an electronic device according to one embodiment.

[0011] FIG. 5 is a diagram illustrating an operation to modify tag information related to an image of an electronic device according to one embodiment.

[0012] FIG. 6 is a diagram illustrating an operation to modify tag information related to an image of an electronic device according to one embodiment.

[0013] FIG. 7 is a diagram illustrating an operation to modify tag information related to an image of an electronic device according to one embodiment.

[0014] FIG. 8 is a diagram illustrating an operation to modify tag information related to an image of an electronic device according to one embodiment.

[0015] FIG. 9 is a diagram illustrating an operation to modify tag information related to an image of an electronic device according to one embodiment.

[0016] FIG. 10 is a flowchart of an image management method for an electronic device according to one embodiment.

[0017] FIG. 11 is a flowchart of an image management method for an electronic device according to one embodiment.

[0018] FIG. 12 is a flowchart of an image management method for an electronic device according to one embodiment.

[0019] FIG. 13 is a flowchart of an image management method for an electronic device according to one embodiment.

[0020] FIG. 14 is a flowchart of an image management method for an electronic device according to one embodiment.

[0021] FIG. 15 is a flowchart of an image management method for an electronic device according to one embodiment.

[0022] FIG. 16 is a block diagram of an exemplary electronic device capable of performing the operations described in this document.

[0023] FIG. 17 illustrates a generative artificial intelligence system according to one embodiment.

[0024] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.

[0025] FIG. 1 is a block diagram of an electronic device according to one embodiment.

[0026] According to one embodiment, an electronic device (100) (e.g., the electronic device (200) of FIG. 2, the electronic device (1600) of FIG. 16, or the electronic device (1700) of FIG. 17) may include a display (110) (e.g., the display (1640) of FIG. 16), a memory (120) (e.g., the memory (220) of FIG. 2, the memory (1620) of FIG. 16, or the database (1730) of FIG. 17), and at least one processor (130) (e.g., the processor (1610) of FIG. 16 or the artificial intelligence framework (1720) of FIG. 17).

[0027] According to one embodiment, the display (110) can visually display information. For example, the display (110) can display at least one image (e.g., original image and / or edited image). The display (110) can display information related to at least one image (e.g., tags related to the image). According to one embodiment, the display (110) can be integrally formed with an input device (e.g., a touch panel). For example, the display (110) may include a touchscreen display (110).

[0028] According to one embodiment, the memory (120) may store instructions that control the operation of the electronic device (100) when executed individually or collectively by at least one processor (130). For example, the instructions may be stored in one memory (120) or multiple memories (120). The memory (120) may store at least temporarily information and / or data related to the operations of the electronic device (100). The memory (120) may store at least one image (e.g., original image and / or edited image). The memory (120) may store information related to the image (e.g., tags related to the image). For example, the memory (120) may store at least one artificial intelligence model. The artificial intelligence model may be learned based on user input, the original image, the image edited from the original image, the reference image input by the user, a first tag related to the original image, a region of interest of the original image, or a second tag related to at least one of at least one object related to the original image. For example, artificial intelligence models may include, but are not limited to, LLM, LVM, and / or generative AI models.

[0029] According to one embodiment, at least one processor (130) can control the operations of an electronic device (100) by executing instructions stored in memory (120) individually or collectively. For example, the operations described in this disclosure as being performed by the 'processor (130)' may be understood as being performed individually or collectively by at least one processor (130). For example, at least one processor (130) can control each of the operations of the electronic device (100) described below independently or collectively. According to one embodiment, at least one processor (130) (140) may include circuits such as a central processing unit (CPU), a micro processor unit (MPU), an application processor (AP), a communication processor (CP), a System On Chip (SoC), and / or an Integrated Circuit (IC).

[0030] According to one embodiment, the processor (130) may display an original image through the display (110). For example, the original image may include an image stored in the electronic device (100). The original image may include an image captured through the camera of the electronic device (100) and / or an image received from an external source (e.g., an external electronic device (100) or an external network). For example, the electronic device (100) (e.g., memory (120)) may include a first tag associated with the original image. For example, the first tag may be acquired together with the acquisition of the original image and may be generated by the processor (130) when the original image is acquired. For example, the processor (130) can generate the first tag using a learned artificial intelligence model (e.g., an under model (e.g., large language model (LLM)), a vision model (e.g., large vision model (LVM)), and / or a multimodal model (e.g., large multi-modal model (LMM)). For example, the learned artificial intelligence model is not limited to those mentioned above and may include at least one generative AI model learned based on an artificial neural network algorithm, a machine learning algorithm, and / or a big data-based algorithm.

[0031] According to one embodiment, the processor (130) may receive user input for an original image. For example, user input for an original image may include user input for modifying (editing) the original image. For example, user input for an original image may include input for extracting at least a portion of the original image, input for adjusting the size of the original image (e.g., enlarge or reduce), input for removing (deleting) at least one object included in the original image, input for adding at least one object to the original image, input for editing at least a portion of the original image (e.g., objects included in the original image and / or the background of the original image), input for selecting (or specifying) at least a portion of the original image, and / or input for applying at least one visual effect to the original image. For example, user input may include user touch input, gesture input, voice input, motion input, and / or input corresponding to the user's gaze, and the method of user input is not limited to those listed above.

[0032] According to one embodiment, the processor (130) may recognize at least one of a point of interest (e.g., POI or ROI) of the original image or at least one object associated with the original image based on user input. For example, at least one object associated with the original image may include an object included in the original image and / or a new object to be added to the original image. For example, the point of interest may include at least one object included in the original image. For example, the processor (130) may recognize a portion of the original image displayed enlarged (or reduced) on the display (110) as the point of interest if the user input is an input for enlarging (or reducing) the original image. The processor (130) may create and / or save at least one of the recognized point of interest or at least one object as a separate image. For example, the processor (130) may recognize the new object as an object associated with the original image if the user input received in the 1020 operation is an input for adding a new object to the original image.

[0033] According to one embodiment, the processor (130) may generate a second tag associated with at least one of a region of interest or at least one object. For example, the processor (130) may generate a second tag associated with at least one of a region of interest or at least one object using a trained artificial intelligence model (e.g., LLM, LVM, and / or LMM). For example, the second tag may contain a more detailed description associated with the image than the first tag. For example, the second tag may contain more information about the same image (at least a portion of the image) compared to the first tag. For example, the processor (130) may generate the second tag at a higher level (a level for generating more detailed information) than when generating the first tag using a trained artificial intelligence model. The processor (130) may generate a first tag containing information at a first level of description using an artificial intelligence model, and generate a second tag containing information at a second level of description (which may contain more detailed information than the information at the first level of description, and / or a greater amount of information). The processor (130) may store the second tag in memory (120). For example, the processor (130) may store the second tag in memory (120) by associating it with at least one of a region of interest or at least one object. For example, if the received user input is an input for adding a new object to the original image, the processor (130) may generate a second tag associated with the new object.

[0034] According to one embodiment, the processor (130) can modify a first tag associated with the original image based on a second tag. For example, at least one of the modified first tag or second tag may include information associated with at least one of the type, quantity, name, shape, or material of at least one of the object included in the region of interest or at least one object associated with the original image.

[0035] For example, if the received user input is an input for removing at least one object included in the original image, the processor (130) may remove information corresponding to a second tag from a first tag associated with the original image. For example, if the user input is an input for removing a first object included in the original image, the processor (130) may determine whether the original image contains a second object of the same type as the first object after removing the first object from the original image. If the original image contains a second object of the same type as the first object, the processor (130) may retain the tag of the original image. For example, if the user input is an input for removing a first object included in the original image and the first tag contains information related to the quantity of the first object, the processor (130) may modify the information related to the quantity of the first object included in the first tag after removing the first object from the original image.

[0036] For example, if the user input is an input for adding a new object to the original image, the processor (130) may modify the first tag to include information corresponding to a second tag associated with the new object.

[0037] According to one embodiment, the operation of the processor (130) modifying the first tag may include the operation of regenerating a tag related to the original image based on information corresponding to the first tag (e.g., the first tag before modification) and / or information corresponding to the second tag using an artificial intelligence model. For example, the processor (130) may modify at least a portion of the first tag related to the original image, or generate a new first tag related to the original image.

[0038] According to various embodiments, the configuration of the electronic device (100) is not limited to that shown in FIG. 1, and at least some configurations may be omitted, or at least one configuration (e.g., at least one of the components of the electronic device (200) of FIG. 2 and / or the components of the electronic device (1600) of FIG. 16) may be added.

[0039] An electronic device according to embodiments of the present disclosure can expand tags and reflect user interests in tags by automatically modifying or updating tags related to an image based on user input regarding the image, without the user having to manually modify tags related to the image (e.g., original image). An electronic device according to one embodiment can increase the convenience and efficiency of managing the user's images (e.g., image search, utilization, and / or classification) by modifying tags based on user input.

[0040]

[0041] FIG. 2 is a block diagram of an electronic device according to one embodiment.

[0042] According to one embodiment, an electronic device (200) (e.g., electronic device (100) of FIG. 1, electronic device (200) of FIG. 2, electronic device (1600) of FIG. 16, or electronic device (1700) of FIG. 17) may include a content viewer application (210), an input module (230), an LMM (240) (e.g., generative AI model (1750) of FIG. 17), and a memory (220) (e.g., memory (1) of FIG. 1, memory (1620) of FIG. 16, or database (1730) of FIG. 17). Hereinafter, the operation of the content viewer application (210) may be referred to as an operation performed by a processor of the electronic device (200) (e.g., processor (130) of FIG. 1 or at least one processor (1610) of FIG. 16) rather than the operation of a specific application.

[0043] According to one embodiment, the content viewer application (210) may include a region of interest determination module (211) and a metadata generation module (213). The content viewer application (210) may modify the original content (201) (e.g., original image) based on user input received through the input module (230), or generate and / or modify metadata related to the original content (201). For example, the content may include, but is not limited to, still images, videos, photographs, drawings, images, text content, and / or sound (sound) content.

[0044] According to one embodiment, the region of interest determination module (211) can determine a region of interest of the original content (201) based on information related to user input provided by the input module (230). For example, the region of interest determination module (211) can determine at least a portion of the original content (201) (e.g., at least one object included in the original content (201)) as a region of interest. The region of interest determination module (211) can provide information on the determined region of interest to the metadata generation module (213).

[0045] According to one embodiment, the metadata generation module (213) can generate metadata related to content (201) using an artificial intelligence model (e.g., LMM (240)). For example, the metadata generation module (213) can generate metadata (e.g., tags) related to content (201) when an electronic device acquires content (201) (e.g., when creating, capturing, receiving, and / or storing content). For example, when an electronic device acquires content (201), it can generate metadata (e.g., tags) of a first level (e.g., a first description level) related to content using an artificial intelligence model.

[0046] According to one embodiment, the metadata generation module (213) may generate metadata related to a region of interest using an artificial intelligence model (e.g., LMM (240)). For example, the metadata generation module (213) may generate tags related to a region of interest (e.g., at least one object of the original content (201)). In the present disclosure, metadata may be referred to as a concept including tags, and generating or modifying metadata may be referred to as generating or modifying tags corresponding to the metadata. For example, metadata of an image may include tags related to the image. The metadata generation module (213) may store the generated metadata in memory (220). For example, the electronic device may generate metadata (e.g., tags) of a second level (e.g., a second description level) related to a region of interest using an artificial intelligence model. For example, metadata at the second level (e.g., second description level) (e.g., second tags) may contain more information and / or more detailed information related to the content (e.g., areas of interest in the content) than metadata at the first level (e.g., first description level) (e.g., first tags). For example, metadata at the second level may contain information that describes the content in more detail than metadata at the first level.

[0047] For example, the metadata generation module (213) can modify metadata related to the original content (201) based on metadata related to the region of interest. For example, the metadata generation module (213) can modify (e.g., add, delete, and / or change) at least some of the tags related to the original content (201) based on tags related to the region of interest. The metadata generation module (213) can store the original content (201) and metadata related to the original content (201) in association with each other in memory (220). The metadata generation module (213) can store the content corresponding to the region of interest and metadata related to the region of interest in association with each other in memory (220).

[0048] According to one embodiment, the input module (230) may receive user input. The input module (230) may include, but is not limited to, a physical key, a touch panel, a microphone, a sensor, and / or a touchscreen display. For example, the input module (230) may receive user input to identify a region of interest of the original content (201) or at least one object related to the original content (201). The input module (230) may provide information related to the input user input to a content viewer application (210) (e.g., a region of interest determination module (211)).

[0049] According to one embodiment, the LMM (240) can generate metadata (e.g., tags) related to a region of interest based on information related to original content (201) and / or a region of interest provided by a content viewer application (210). According to various embodiments, the model cited by the electronic device (200) is not limited to the LMM (240) and may include at least one trained AI model (e.g., LMM (240), LVM, LMM (240), or a combination thereof).

[0050] According to one embodiment, the memory (220) may store original content (201), content corresponding to a region of interest, metadata related to the original content (201), and / or metadata related to the region of interest.

[0051] According to various embodiments, the configuration of the electronic device (200) is not limited to that shown in FIG. 2, and at least some configurations may be omitted, or at least one configuration (e.g., at least one of the components of the electronic device (100) of FIG. 1 and / or the components of the electronic device (1600) of FIG. 16) may be added.

[0052]

[0053] FIG. 3 is a diagram illustrating an operation to modify tag information related to an image of an electronic device according to one embodiment.

[0054] According to one embodiment, 301 represents the original image (310), and 303 represents the case where the region of interest (320) of the original image (310) is determined based on user input. For example, an electronic device (e.g., electronic device (100; 200; 1600; 1700)) may receive user input to separate at least one object from the original image (310). The electronic device may store the separated at least one object as separate content (e.g., an image) based on user input. The electronic device may determine the region of interest (320) as an area containing a girl, a woman, and a dog (e.g., a Golden Retriever) included in the original image (310) based on user input. For example, the electronic device may store the region of interest (320) as an image separate from the original image (310).

[0055] According to one embodiment, tags (e.g., a first tag and / or a second tag) may include keyword-based tags and / or natural language-based tags. For example, keyword-based tags may include keyword words or sentences representing objects contained in an image (e.g., a source image (310) and / or a region of interest (320)), and / or keyword words or sentences describing the image. For example, natural language-based tags may include natural language sentences representing images (e.g., a source image (310), a region of interest (320), and / or objects).

[0056] For example, assume a case where the tags associated with the original image (310) are 'dog, male, female, boy, girl, soccer ball, “Family is playing soccer.”'. For example, 'dog, male, female, boy, girl, soccer ball' may be keyword-based tags generated based on each object included in the original image (310), and 'Family is playing soccer.' may be a natural language-based tag describing the original image (310). For example, the tags associated with the original image (310) may be tags that were included in the original image (310) when the original image (310) was acquired, or they may be tags generated by an electronic device when the original image (310) was acquired. For example, when the original image (310) is acquired, the electronic device may use an artificial intelligence model to generate tags of a first level (e.g., first description level) associated with the original image (310) (e.g., 'dog, male, female, boy, girl, “People are playing soccer.”).

[0057] The electronic device may generate tags related to the region of interest (320) and / or objects included in the region of interest (320). The electronic device may use an artificial intelligence model to generate tags of a second level (e.g., a second description level) related to objects included in the region of interest (320). For example, tags of the second level may contain relatively more detailed information representing at least a part of the image (e.g., the region of interest) compared to tags of the first level. For example, when the electronic device generates tags using an artificial intelligence model, it may pass a query requesting a description of the image to the artificial intelligence model along with the level describing the image (e.g., a description level). The electronic device may obtain a response (information describing the image) corresponding to the requested level from the artificial intelligence model. For example, the electronic device may generate tags related to the region of interest (320) such as 'Golden Retriever (and / or puppy name), Wife (and / or wife name), Son (and / or son name), Daughter (and / or daughter name), “The wife is happily watching the children happily playing soccer with the puppy.” For example, 'Golden Retriever (and / or puppy name), Wife (and / or wife name), Son (and / or son name), Daughter (and / or daughter name)' may be keyword-based tags, and '“The wife is happily watching the children happily playing soccer with the puppy.”' may be a natural language-based tag. For example, an electronic device may generate second tags (e.g., 'Golden Retriever (and / or puppy name)', 'Wife (and / or wife name)', 'Son (and / or son name)', 'Daughter (and / or daughter name)') that are more specific and / or personalized at a second level (referred to as the 'second description level' or 'second analysis level') for keyword-based tags (e.g., 'dog', 'female', 'boy', 'girl') included in the first tag.For example, 'Golden Retriever' may be a tag that further specifies 'dog' of the first tag, and 'puppy name', 'wife (and / or wife name)', 'son (and / or son name)', 'daughter (and / or daughter name)' may be tags that personalize (i.e., generate based on the user's personal information) 'dog', 'female', 'boy', and 'girl' of the first tag.

[0058] The electronic device may modify tags related to the original content based on tags related to the region of interest (320) or and / or objects included in the region of interest (320). For example, the electronic device may add tags related to the region of interest (320) or and / or objects included in the region of interest (320) to tags related to the original content. For example, the electronic device may modify tags related to the original content to 'dog, male, female, boy, girl, soccer ball, “people are playing soccer.”, Golden Retriever (and / or puppy name), wife (and / or wife’s name), son (and / or son’s name), daughter (and / or daughter’s name), “wife is happily watching the children happily playing soccer with the puppy.”'. For example, the electronic device may modify (update) the tags related to the original content based on the tags related to the area of ​​interest (320) or the objects included in the area of ​​interest (320), rather than adding the tags related to the area of ​​interest (320) or the objects included in the area of ​​interest (320) as they are to the tags related to the original content. For example, the electronic device may modify the tags related to the original content to “My family is playing soccer.”, soccer ball, Golden Retriever (and / or puppy name), wife (and / or wife's name), son (and / or son's name), daughter (and / or daughter's name), “My wife is happily watching me (or username) and the children happily playing soccer with the puppy.”

[0059]

[0060] FIG. 4 is a diagram illustrating an operation to modify tag information related to an image of an electronic device according to one embodiment.

[0061] 401 represents the original image (410), 403 represents the case where the first object (420) is designated (e.g., separated) from the original image (410), and 405 represents the case where the second object (430) is designated (e.g., separated) from the original image (410).

[0062] For example, assume that the tags associated with the original image (410) are “can, candy container, tube, toothpaste, STARBU, hand cream”.

[0063] An electronic device (e.g., electronic device (100; 200; 1600; 1700)) can recognize a first object (420) (e.g., a first region of interest including the first object (420)) based on user input and generate a tag associated with the first object (420). For example, among the tags associated with the original image (410), the tag associated with the first object (420) may be “can”. The electronic device can generate a tag for the first object (420) as “can, XYLITOL Crystal lemon” based on the tag associated with the original image (410) and the analysis result of the object (e.g., the first object (420)) using a learned AI model (e.g., LVM and / or LMM). For example, the electronic device can use an AI model to generate a tag for the object (420) that includes more detailed information related to the image and / or object than the tag for the original image (410). The electronic device can use an AI model to generate a tag for an object (420) that has a higher level (e.g., a description level) than the level (e.g., a description level) of the original image (410). For example, the tag associated with the original image (410) may contain only “can” as information associated with the first object (420), but the information associated with the first object (420) may include “can, XYLITOL Crystal lemon”, which may contain relatively more information and / or more detailed information associated with the first object (420). Based on the tag of the first object (420), the electronic device can modify the tag associated with the original image (410) from “can, candy container, tube, toothpaste, STARBU, hand cream” to “can, candy container, tube, toothpaste, STARBU, hand cream, XYLITOL Crystal lemon”.

[0064] The electronic device can recognize a second object (430) (e.g., a second region of interest including the second object (430)) based on user input and generate tags associated with the second object (430). For example, among the tags associated with the original image (410), the tag associated with the second object (430) may be “candy container, STARBU”. Based on the tags associated with the original image (410) and the analysis results of the object (e.g., the second object (430)) using a learned AI model (e.g., LVM and / or LMM), the electronic device can generate tags for the second object (430) as “candy container, STARBU, STARBU Korea 2024 Sea Friends Candy, round metal box containing candy with a STARBU logo”. The electronic device can modify the tag associated with the original image (410) from “can, candy container, tube, toothpaste, STARBU, hand cream” to “can, candy container, tube, toothpaste, STARBU, hand cream, STARBU Korea 2024 Sea Friends Candy, round metal box containing candy with STARBU logo” based on the tag of the second object (430).

[0065]

[0066] FIG. 5 is a diagram illustrating an operation to modify tag information related to an image of an electronic device according to one embodiment.

[0067] For example, 501 indicates a case where the original image (510) before enlargement is displayed on the display of an electronic device (e.g., electronic device (100; 200; 1600; 1700)), and 503 indicates a case where a portion (520) of the enlarged original image is displayed on the display of the electronic device based on user input.

[0068] For example, assume a case where the tags associated with the original image (510) are “dog, male, female, boy, girl, “people are playing soccer.” The electronic device may determine an area (520) of the enlarged original image displayed on the display as the region of interest (520) based on user input (e.g., a pinch zoom-in gesture). For example, if the electronic device enlarges and displays at least a portion of the image based on user input, the electronic device may determine at least one object included in at least a portion of the enlarged image and / or at least a portion of the enlarged image as the region of interest (520). For example, the electronic device may determine at least a portion of the enlarged area as the region of interest if at least a portion of the image is enlarged beyond a specified threshold or / or remains enlarged for a specified amount of time. For example, the electronic device may generate tags associated with the region of interest (520) as “Golden Retriever, wife, son, daughter, “wife is happily watching the children happily playing soccer with the puppy.” The electronic device and the region of interest (520) Based on the related tags, the tags related to the original image (510) can be modified to “Our family is playing soccer.”, soccer ball, Golden Retriever (and / or puppy name), wife (and / or wife's name), son (and / or son's name), daughter (and / or daughter's name), “My wife is happily watching the children playing soccer with the puppy.”

[0069]

[0070] FIG. 6 is a diagram illustrating an operation to modify tag information related to an image of an electronic device according to one embodiment. For example, FIG. 6 illustrates an operation in which an electronic device modifies a tag recursively.

[0071] For example, assume that the first tag associated with the original image (610) is 'dog, male, female, boy, girl, soccer ball, “people are playing soccer”’. For example, the initial first tag associated with the original image (610) may be a tag generated at a first level (e.g., a first description level) by an electronic device using an artificial intelligence model. For example, the first tag before modification may contain information at the first level associated with the original image (610).

[0072] For example, in operation 601, an electronic device (e.g., electronic device (100; 200; 1600; 1700)) can extract at least one object (620) contained in an original image (610) based on a first user input. The electronic device can generate a first image corresponding to at least one object (620) extracted from the original image (610). The electronic device can generate a second tag associated with at least one object (620) extracted from the original image (610). For example, the electronic device can generate the second tag as 'Golden Retriever (or, puppy name), Wife (or, wife's name), Daughter (or, daughter's name), Son (or, son's name), Me (or, user name)'. For example, the electronic device can use an artificial intelligence model to generate a second tag associated with at least one object (620) at a second level (e.g., a second description level). The second level may be a level that requests more detailed information related to the image than the first level. For example, the second tag of the second level may contain more information and / or more detailed information related to the image (e.g., at least one object (620)) than the first tag of the initial first level.

[0073] For example, in operation 603, the electronic device can modify the first tag based on the second tag. For example, the electronic device can modify the first tag to 'Golden Retriever, wife, daughter, son, me, “Our family is playing soccer with the puppy.”'

[0074] For example, in operation 605, the electronic device may extract at least one object (630) from an object (620) extracted from an original image (610) based on a second user input. For example, the electronic device may extract at least one object (630) included in the first image. The electronic device may generate a third tag associated with at least one object (630) extracted from the object (620) extracted from the original image (610) (e.g., at least one object (630) extracted from the first image). For example, the electronic device may generate a third tag as “me wearing brown hair, a shirt, and shorts.” For example, the electronic device may use an artificial intelligence model to generate a third tag associated with at least one object (630) at a second level (e.g., a second description level). The second level may be a level that requests more detailed information related to the image than the first level. For example, the third tag of the second level may contain more information and / or more detailed information related to the image (e.g., at least one object (630)) than the first tag of the first level.

[0075] For example, in operation 607, the electronic device may modify the second tag based on the third tag. For example, the electronic device may modify the second tag to 'Golden Retriever (or, puppy name), wife (or, wife's name), daughter (or, daughter's name), son (or, son's name), me with brown hair wearing a shirt and shorts (or, user name)'.

[0076] For example, in operation 609, the electronic device may modify the first tag based on the third tag and / or the second tag modified based on the third tag. For example, the electronic device may modify the first tag to ‘Golden Retriever, wife, daughter, son, me with brown hair wearing a shirt and shorts, “Our family is playing soccer with a puppy (puppy name).”’

[0077]

[0078] FIG. 7 is a diagram illustrating an operation to modify tag information related to an image of an electronic device according to one embodiment.

[0079] For example, an electronic device (e.g., electronic device (100; 200; 1600; 1700)) may receive user input to insert (or add) a first object (701) into an original image (710). Based on the user input, the electronic device may generate an edited image (720) with the first object (701) inserted.

[0080] For example, assume that the original image (710) includes a second object (e.g., candy container) and a third object (e.g., hand cream), and that the tags associated with the original image (710) are “candy container, hand cream, STARBU”.

[0081] The electronic device can generate tags associated with the first object (701) using a trained AI model. For example, the electronic device can generate tags associated with the first object (701) such as “can, Xylitol candy crystal lemon”.

[0082] The electronic device can modify tags related to the original image (710) based on tags related to the first object (701). The electronic device can add information related to tags related to the first object (701) to tags related to the original image (710). For example, the electronic device can modify tags related to the original image (710) to “candy container, hand cream, STARBU, can, Xylitol candy crystal lemon”. The electronic device can modify tags related to the original image (710) to create tags for the edited image (720) (i.e., the edited original image).

[0083] According to one embodiment, after the editing of the original image (710) is completed, tag analysis can be performed again on the edited image (720) (i.e., the edited original image). For example, assume that the initial tag associated with the original image (710) is “left is a STARBU candy container, right is hand cream.” When the electronic device creates the edited image (720) by adding a first object (701) to the original image (710), it can create a tag associated with the original image (720) (e.g., “xylitol candy can”) by adding a tag associated with the first object (701) (e.g., “xylitol candy can”) to the tag associated with the original image (720) (e.g., “left is a STARBU candy container, right is hand cream, xylitol candy can”). After the editing is completed, the electronic device can perform tag analysis again on the edited image (720) to create a tag associated with the edited image (720) such as “left is a xylitol candy can, middle is a STARBU candy container, right is hand cream.” For example, a tag associated with an edited image (720) may simply include a tag associated with the original image (710) with a tag associated with the first object (701) added thereto, and / or may include a tag associated with the original image (710), a tag associated with the first object (701), and / or a newly generated tag based on the edited image (720).

[0084]

[0085] FIG. 8 is a diagram illustrating an operation to modify tag information related to an image of an electronic device according to one embodiment.

[0086] For example, an electronic device (e.g., electronic device (100; 200; 1600; 1700)) may receive user input to remove a first object (801) from an original image (810). Based on the user input, the electronic device may generate an edited image (820) from which the first object (801) has been removed.

[0087] For example, the original image (810) includes a first object (801) (e.g., a can), a second object (e.g., a candy container), and a third object (e.g., hand cream), and assumes that the tags associated with the original image (810) are “candy container, hand cream, STARBU, can”.

[0088] The electronic device can generate tags associated with the first object (801) using a trained AI model. For example, the electronic device can generate tags associated with the first object (801) such as “can, Xylitol candy crystal lemon”.

[0089] The electronic device can modify tags related to the original image (810) based on tags related to the first object (801). The electronic device can delete information related to tags related to the first object (801) from tags related to the original image (810). For example, the electronic device can modify tags related to the original image (810) to “candy container, hand cream, STARBU”. The electronic device can modify tags related to the original image (810) to create tags for the edited image (820) (i.e., the edited original image).

[0090] According to one embodiment, after the editing of the original image (810) is completed, tag analysis can be performed again on the edited image (820) (i.e., the edited original image). For example, assume that the initial tag associated with the original image (810) is “the left is a candy can, the middle is a STARBU candy container, and the right is hand cream.” When the electronic device creates the edited image (820) by deleting the first object (801) from the original image (810), it can create a tag associated with the edited image (820) (e.g., “the middle is a STARBU candy container, and the right is hand cream”) by removing information associated with the tag associated with the first object (701) (e.g., “Xylitol candy can”) from the tag associated with the original image (720). After the editing is completed, the electronic device can perform tag analysis again on the edited image (820) to create a tag associated with the edited image (820) such as “the left is a STARBU candy container, and the right is hand cream.” For example, a tag associated with an edited image (820) may simply include a tag associated with the original image (810) with the information associated with the tag associated with the first object (801) removed, and / or may include a tag associated with the original image (810), a tag associated with the first object (801), and / or a newly generated tag based on the edited image (820).

[0091]

[0092] FIG. 9 is a diagram illustrating an operation to modify tag information related to an image of an electronic device according to one embodiment.

[0093] For example, the original image (910) includes a first object (901) (e.g., a first can), a second object (903) (e.g., a second can), a third object (e.g., a candy container), and a fourth object (e.g., hand cream), and assumes that the tags associated with the original image (910) are “candy container, hand cream, STARBU, can”. For example, the first object (901) and the second object (903) may be the same object or objects of the same kind.

[0094] For example, an electronic device (e.g., electronic device (100; 200; 1500; 1600)) may receive user input to remove a second object (903) from an original image (910). Based on the user input, the electronic device may generate an edited image (920) from which the second object (903) has been removed.

[0095] For example, based on the fact that the tag associated with the original image (910) does not contain information regarding the quantity of the first object (901) and the second object (903), the electronic device may maintain the tag for the original image (910) without modifying it. For example, even though the second object (903) has been removed from the original image (910), the first object (901) identical to the second object (903) still remains in the original image (910) (i.e., because the edited image (920) contains the first object (901) identical to the second object (903)), the electronic device may not modify the tag for the original image (910).

[0096] As another example, assume a case where the tag associated with the original image (910) is “candy container, hand cream, STARBU, 2 cans”. The tag associated with the original image (910) may include information related to the quantity of the first object (901) and the second object (903), such as “2 cans”. In this case, the electronic device may generate a tag associated with the second object (903) using a trained AI model. For example, the electronic device may generate a tag associated with the second object (903) such as “can, Xylitol candy crystal lemon”. The electronic device may modify the tag associated with the original image (810) based on the tag associated with the second object (903). The electronic device may modify information related to the receipt of the first object (901) and the second object (903) in the tag associated with the original image (810). For example, the electronic device can modify the tag associated with the original image (910) to “candy container, hand cream, STARBU, 1 can” or “candy container, hand cream, STARBU, can”.

[0097]

[0098] An electronic device according to one embodiment may include at least one processor comprising a display, a memory for storing instructions, and processing circuitry.

[0099] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may display the original image through the display.

[0100] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may receive user input for the original image.

[0101] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may recognize at least one of a point of interest (POI) of the original image or at least one object associated with the original image based on the user input.

[0102] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may generate a second tag associated with at least one of the region of interest or the at least one object.

[0103] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may modify the first tag associated with the original image based on the second tag.

[0104] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may remove information corresponding to the second tag associated with the at least one object removed from the original image based on the user input from the first tag, if the user input is an input for removing at least one object included in the original image.

[0105] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may recognize a portion of the original image that is enlarged and displayed on the display as the region of interest when the user input is an input for enlarging the original image.

[0106] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may generate the second tag associated with the new object when the user input is an input for adding a new object to the original image.

[0107] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may modify the first tag to include information corresponding to the second tag associated with the new object.

[0108] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may use a learned artificial intelligence model to regenerate a tag associated with the original image based on information corresponding to the first tag and information corresponding to the second tag.

[0109] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may determine whether the original image contains a second object of the same type as the first object after removing the first object from the original image, if the user input is an input for removing a first object included in the original image.

[0110] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may retain the first tag when the original image contains the second object.

[0111] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may generate a first image corresponding to the region of interest.

[0112] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may receive a second user input for the first image.

[0113] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may recognize at least one of a region of interest of the first image or at least one object associated with the first image based on the second user input.

[0114] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may generate a third tag associated with at least one of a region of interest of the first image or at least one object associated with the first image.

[0115] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may modify the second tag based on the third tag.

[0116] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may modify the first tag based on the modified second tag.

[0117] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may store the original image and the first tag in association.

[0118] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may store an image corresponding to at least one of the region of interest or at least one of the at least one object and the second tag in association.

[0119] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may generate the first tag using a learned AI model when acquiring the original image.

[0120] At least one of the modified first tag or the second tag may include information related to at least one of the type, quantity, name, shape, or material of an object included in the region of interest or at least one of the at least one object.

[0121] An electronic device according to one embodiment can improve performance when searching, classifying, and / or utilizing images by expanding the metadata (e.g., tags) of an image based on user input.

[0122]

[0123] FIG. 10 is a flowchart of an image management method for an electronic device according to one embodiment.

[0124] According to one embodiment, in operation 1010, an electronic device (e.g., electronic device (100; 200; 1600; 1700)) may display an original image. For example, the original image may include an image previously stored in the electronic device. The original image may include an image captured through a camera of the electronic device and / or an image received from an external source (e.g., an external electronic device or an external network). For example, the electronic device may include a first tag associated with the original image. For example, the first tag may be acquired together with the acquisition of the original image and may be generated by the electronic device at the acquisition of the original image. For example, the first tag may be generated using a learned artificial intelligence model (e.g., LLM, LVM, and / or LMM).

[0125] According to one embodiment, in operation 1020, the electronic device may receive user input regarding the original image. For example, user input regarding the original image may include user input for modifying (editing) the original image. For example, user input regarding the original image may include input for extracting at least a portion of the original image, input for adjusting the size of the original image (e.g., enlarge or reduce), input for removing (deleting) at least one object included in the original image, input for adding at least one object to the original image, input for editing at least a portion of the original image (e.g., objects included in the original image and / or the background of the original image), input for selecting (or specifying) at least a portion of the original image, and / or input for applying at least one visual effect to the original image. For example, user input may include user touch input, gesture input, voice input, motion input, and / or input corresponding to the user's gaze, and the method of user input is not limited to those listed above.

[0126] According to one embodiment, in operation 1030, the electronic device may recognize at least one of a region of interest (e.g., a point of interest (POI) or a region of interest (ROI)) of the original image or at least one object associated with the original image based on user input. For example, at least one object associated with the original image may include an object contained in the original image and / or a new object to be added to the original image. For example, the region of interest may include at least one object contained in the original image. For example, if the user input is an input for zooming in (or zooming out) on the original image, the electronic device may recognize a portion of the original image displayed on the display as the region of interest. The electronic device may create and / or save at least one of the recognized region of interest or at least one object as a separate image. For example, if the user input received in operation 1020 is an input for adding a new object to the original image, the electronic device may recognize the new object as an object associated with the original image.

[0127] According to one embodiment, in operation 1040, the electronic device may generate a second tag associated with at least one of a region of interest or at least one object. For example, the electronic device may generate a second tag associated with at least one of a region of interest or at least one object using a learned artificial intelligence model (e.g., LLM, LVM, and / or LMM). The electronic device may store the second tag. For example, the electronic device may store the second tag associated with at least one of a region of interest or at least one object. For example, if the user input received in operation 1020 is an input for adding a new object to the original image, the electronic device may generate a second tag associated with the new object.

[0128] According to one embodiment, in operation 1050, the electronic device may modify a first tag associated with the original image based on a second tag. For example, at least one of the modified first tag or second tag may include information associated with at least one of the type, quantity, name, shape, or material of at least one of the object included in the region of interest or at least one object associated with the original image.

[0129] For example, if the user input received in operation 1020 is an input for removing at least one object included in the original image, the electronic device may remove information corresponding to a second tag from a first tag associated with the original image. For example, if the user input is an input for removing a first object included in the original image, the electronic device may determine whether the original image contains a second object of the same type as the first object after removing the first object from the original image. If the original image contains a second object of the same type as the first object, the electronic device may retain the tag of the original image. For example, if the user input is an input for removing a first object included in the original image and the first tag contains information related to the quantity of the first object, the electronic device may modify the information related to the quantity of the first object included in the first tag after removing the first object from the original image.

[0130] For example, if the user input is an input for adding a new object to an original image, the electronic device may modify the first tag to include information corresponding to a second tag associated with the new object.

[0131] According to one embodiment, the operation of an electronic device modifying a first tag may include an operation of regenerating a tag associated with an original image based on information corresponding to the first tag (e.g., the first tag before modification) and / or information corresponding to a second tag using an artificial intelligence model. For example, the electronic device may modify at least a portion of the first tag associated with the original image, or generate a new first tag associated with the original image.

[0132] According to one embodiment, the electronic device may perform at least some of the operations 1010 to 1050 recursively. For example, the electronic device may generate a first image corresponding to a region of interest of an original image and generate a third tag associated with at least one of the region of interest of the first image or an object associated with the first image. The electronic device may modify the first tag and / or the second tag based on the third tag. For example, the electronic device may modify the second tag associated with the first image based on the third tag and modify the first tag associated with the original image based on the modified second tag.

[0133] According to various embodiments, the order of the operations of FIG. 10 may be changed, or at least some of the operations may be performed simultaneously. At least some of the operations of FIG. 10 may be omitted, or at least one operation (e.g., at least one of the operations of FIG. 11 to 15) may be added. For example, the operations of FIG. 10 may be performed independently of the operations of FIG. 11 to 15, performed in conjunction with each other, or at least some of the operations may be integrated.

[0134] An image management method for an electronic device according to one embodiment can expand tags and reflect the user's interests in tags by automatically modifying or updating tags related to an image based on user input regarding the image, without the user having to manually modify tags related to the image (e.g., original image). By modifying tags based on user input, the image management method for an electronic device according to one embodiment can increase the convenience and efficiency of managing the user's images (e.g., image search, utilization, and / or classification).

[0135]

[0136] FIG. 11 is a flowchart of an image management method for an electronic device according to one embodiment. In the following, descriptions that overlap with the description of FIG. 10 are omitted or briefly described.

[0137] According to one embodiment, in operation 1110, an electronic device (e.g., electronic device (100; 200; 1600; 1700)) can generate a first tag of an original image. For example, the electronic device can generate the first tag using a trained AI model. For example, the electronic device can acquire the first tag of the original image together with the acquisition of the original image. The electronic device can modify the first tag of the original image when acquiring the original image using an AI model. For example, the electronic device can generate the first tag of the original image based on the time of acquisition of the original image, the location of acquisition, the acquisition path (e.g., source of the original image), and / or the state of the electronic device at the time of acquisition of the original image, or modify the first tag acquired together with the original image.

[0138] According to one embodiment, in operation 1120, the electronic device may receive user input regarding the original image. For example, the electronic device may receive user input for editing (modifying) at least a portion of the original image and / or user input for specifying at least a portion of the original image.

[0139] According to one embodiment, in operation 1130, the electronic device may determine a target object or a region of interest based on user input. For example, the region of interest may include at least a portion of the original image. The target object may include at least one object included in the original image and / or at least one object to be added to the original image based on user input.

[0140] According to one embodiment, in operation 1140, the electronic device may determine whether the first tag contains a tag related to a target object or a region of interest. For example, the electronic device may determine whether the first tag contains content related to the target object and / or the region of interest. If the first tag contains content related to the target object and / or the region of interest, the electronic device may determine that modification of the first tag is unnecessary.

[0141] If the first tag contains a tag related to the target object or region of interest, the electronic device may not perform additional actions, or may perform actions 1120 or lower again (or repeatedly). If the first tag does not contain a tag related to the target object or region of interest, the electronic device may perform action 1150.

[0142] According to one embodiment, in operation 1150, the electronic device may generate a second tag for a target object or a region of interest. The electronic device may generate the second tag using a trained AI model. For example, the second tag may be different from the first tag.

[0143] According to one embodiment, in operation 1160, the electronic device may modify the first tag based on the second tag. For example, the electronic device may delete information corresponding to the second tag from the first tag, add information corresponding to the second tag to the first tag, and / or regenerate the first tag based on at least some of the information corresponding to the second tag. For example, the electronic device may modify or regenerate the first tag using a learned AI model.

[0144] According to one embodiment, in operation 1170, the electronic device may store a modified first tag and a second tag. The electronic device may store the modified first tag by including it in the original image or by associating it with the original image. The electronic device may store the second tag by including it in the image corresponding to the region of interest and / or the target object, or by associating it with the image corresponding to the region of interest and / or the target object.

[0145] According to various embodiments, the order of the operations of FIG. 11 may be changed, or at least some of the operations may be performed simultaneously. At least some of the operations of FIG. 11 may be omitted, or at least one operation (e.g., at least one of the operations of FIG. 10 and FIG. 12 to 15) may be added. For example, the operations of FIG. 11 may be performed independently of the operations of FIG. 10 and FIG. 12 to 15, performed in conjunction with each other, or at least some of the operations may be integrated.

[0146]

[0147] FIG. 12 is a flowchart of an image management method of an electronic device according to one embodiment. In the following, descriptions that overlap with the description of FIG. 10 or FIG. 11 are omitted or briefly described. For example, FIG. 12 illustrates operations in which an electronic device manages image tags recursively.

[0148] According to one embodiment, in operation 1205, an electronic device (e.g., electronic device (100; 200; 1600; 1700)) can display the original image.

[0149] According to one embodiment, in operation 1210, the electronic device can extract a first object from an original image. For example, the electronic device can extract a region (e.g., a region of interest) containing the first object from the original image.

[0150] According to one embodiment, in operation 1215, the electronic device may determine whether the first tag for the original image contains a tag for the first object. For example, the electronic device may determine whether the first tag contains content related to the tag for the first object (e.g., the second tag). If the first tag contains a tag for the first object, the electronic device may perform operation 1220. If the first tag does not contain a tag for the first object, the electronic device may perform operation 1225.

[0151] According to one embodiment, in operation 1220, the electronic device may maintain the first tag. For example, the electronic device may not modify the first tag.

[0152] According to one embodiment, in operation 1225, the electronic device can generate a second tag for a first object. For example, the electronic device can generate a second tag for a first object using a trained AI model.

[0153] According to one embodiment, in operation 1230, the electronic device may modify the first tag based on the second tag. For example, the electronic device may modify at least a portion of the first tag based on the second tag (e.g., add and / or delete content corresponding to the second tag) or regenerate the first tag for the original image.

[0154] According to one embodiment, in operation 1235, the electronic device may determine whether a second object is extracted from a first object. For example, the electronic device may receive user input to extract a second object (e.g., a region of interest of the first image containing the second object) from a first image containing the first object. The electronic device may perform operation 1240 if the second object is extracted from the first object, and may not perform additional operations if the second object is not extracted from the first object.

[0155] According to one embodiment, in operation 1240, the electronic device may determine whether the second tag contains a tag for the second object. For example, the electronic device may determine whether the second tag contains information corresponding to a tag for the second object. If the second tag contains a tag for the second object, the electronic device may not perform any additional operation. For example, the electronic device may retain the second tag (and / or the first tag modified based on the second tag). If the second tag does not contain a tag for the second object, the electronic device may perform operation 1245.

[0156] According to one embodiment, in operation 1245, the electronic device can generate a third tag for the second object. For example, the electronic device can generate the third tag using a trained AI model.

[0157] According to one embodiment, in operation 1250, the electronic device may modify the second tag based on the third tag. For example, the electronic device may use a learned AI model to modify at least a portion of the second tag based on the third tag (e.g., adding and / or deleting content corresponding to the third tag) or regenerate the second tag.

[0158] For example, if the electronic device modifies the second tag, it may perform operation 1230 again. For example, the electronic device may modify the first tag based on the modified second tag. According to one embodiment, the electronic device may perform operations 1230 to 1250 recursively or repeatedly. For example, if a new object (or region of interest) is extracted from an extracted object (or region of interest), the electronic device may recursively modify and / or update existing tags based on the newly created tag.

[0159] According to various embodiments, the order of the operations of FIG. 12 may be changed, or at least some of the operations may be performed simultaneously. At least some of the operations of FIG. 12 may be omitted, or at least one operation (e.g., at least one of the operations of FIG. 10, 11, and 13 to 15) may be added. For example, the operations of FIG. 12 may be performed independently of the operations of FIG. 10, 11, and 13 to 15, performed in conjunction with each other, or at least some of the operations may be integrated.

[0160]

[0161] FIG. 13 is a flowchart of an image management method for an electronic device according to one embodiment. In the following, descriptions that overlap with the descriptions in FIG. 10 to 12 are omitted or briefly described.

[0162] According to one embodiment, in operation 1310, an electronic device (e.g., electronic device (100; 200; 1600; 1700)) can generate (or modify) a tag for an original image. For example, the electronic device can use an AI model to generate a tag of a first level (e.g., a first description level) for the original image. The electronic device can send a query requesting first-level description information to the AI ​​model and obtain a tag of the first level as a response from the AI ​​model.

[0163] According to one embodiment, in operation 1320, an electronic device (e.g., electronic device (100; 200; 1500; 1600)) may insert an object into an original image. For example, the inserted object may include an object corresponding to an object contained in the original image, an object extracted from the original image, an image different from the original image, an object extracted from an image different from the original image, and an object for decorating the original image (e.g., a sticker, text, a picture, a shape, and / or a visual effect).

[0164] According to one embodiment, in operation 1330, the electronic device may determine whether a tag exists for the inserted object. If no tag exists for the inserted object, the electronic device may perform operation 1330. If a tag exists for the inserted object, the electronic device may perform operation 1340.

[0165] According to one embodiment, in operation 1340, the electronic device may generate a tag for an inserted object. For example, the electronic device may analyze the inserted object based on a learned AI model and generate a tag for the inserted object. For example, the electronic device may use the AI ​​model to generate a tag of a second level (e.g., a second description level) for the object. For example, the tag of the second level may contain more information and / or more detailed information about the image (e.g., the object) than the tag of the first level.

[0166] According to one embodiment, in operation 1350, the electronic device may modify the tag for the original image based on the tag for the inserted object. For example, the electronic device may add information corresponding to the tag for the inserted object to the tag for the original image. For example, the electronic device may regenerate the tag for the original image based on the tag for the inserted object. For example, assume the case where the original image is an image containing a moon and the tag for the original image is “moon,” the inserted object is a sticker in the shape of a headband with rabbit ears, and the tag for the inserted object is “rabbit ear shaped headband.” When a sticker in the shape of a headband with rabbit ears is inserted on the moon of the original image, the electronic device may modify the tag for the original image “moon” to “moon with rabbit ears” based on the tag for the inserted object “rabbit ear shaped headband.”

[0167] For example, the electronic device can add tags for objects to the tags of the original image. The electronic device can use an AI model to modify the tags of the original image to which the tags for objects have been added. The electronic device can use an AI model to perform analysis and / or modification of the tags for the original image after adding the tags for objects. After editing, the electronic device can modify the context of the tags for the original image into natural-sounding sentences based on the original image. For example, based on the location of the inserted object, the electronic device can edit the content related to the location of each object in the tags of the original image and modify them to correspond to the image.

[0168] According to various embodiments, the order of the operations of FIG. 13 may be changed, or at least some of the operations may be performed simultaneously. At least some of the operations of FIG. 13 may be omitted, or at least one operation (e.g., at least one of the operations of FIG. 10 to 12, 14 and 15) may be added. For example, the operations of FIG. 13 may be performed independently of the operations of FIG. 10 to 12, 14 and 15, performed in conjunction with each other, or at least some of the operations may be integrated.

[0169]

[0170] FIG. 14 is a flowchart of an image management method for an electronic device according to one embodiment. In the following, descriptions that overlap with the descriptions of FIG. 10 to 13 are omitted or briefly described.

[0171] According to one embodiment, in operation 1410, an electronic device (e.g., electronic device (100; 200; 1600; 1700)) can display the original image.

[0172] According to one embodiment, in operation 1415, the electronic device can delete a first object from the original image based on user input regarding the original image.

[0173] According to one embodiment, in operation 1420, the electronic device can determine whether there is a second object corresponding to a first object in the original image. For example, the second object may include an object identical to the first object or an object of the same type.

[0174] The electronic device may perform operation 1430 if there is no second object corresponding to the first object in the original image. The electronic device may perform operation 1460 if there is a second object corresponding to the first object in the original image.

[0175] According to one embodiment, in operation 1430, the electronic device may determine whether there is a second tag for the first object. If there is a second tag for the first object, the electronic device may perform operation 1440. If there is a second tag for the first object, the electronic device may perform operation 1450.

[0176] According to one embodiment, in operation 1440, the electronic device can generate a second tag for the first object. For example, the electronic device can generate a second tag for the first object using an AI model.

[0177] According to one embodiment, in operation 1450, the electronic device may modify a first tag for an original image based on a second tag. For example, the electronic device may delete a second tag from the first tag. The electronic device may delete information corresponding to the second tag from the first tag.

[0178] According to one embodiment, in operation 1460, the electronic device may determine whether the first tag for the original image contains information regarding the quantity of the first object and the second object. For example, since the first object is the same or of the same type as the second object, the first tag may contain information regarding the quantity of the first object and the second object. The electronic device may perform operation 1470 if the first tag contains information regarding the quantity of the first object and the second object. If the first tag does not contain information regarding the quantity of the first object and the second object, the electronic device may perform operation 1480.

[0179] According to one embodiment, in operation 1470, the electronic device may modify information related to the quantity of the first object and the second object included in the first tag. For example, assume a case where the original image includes the first object and the second object, and the first object and the second object correspond to the same cup. In this case, the first tag for the original image may include information related to the quantity of the first object and the second object (e.g., “2 cups”). Since the first object (1 cup) has been deleted from the original image, the electronic device may modify the information of “2 cups” included in the first tag to “1 cup” or “cup”.

[0180] According to one embodiment, in operation 1480, the electronic device may maintain a first tag. For example, assume a case where an original image includes a first object and a second object, and the first object and the second object correspond to the same hand cream. The first tag for the original image may not include information related to the quantity of the first object and the second object, but may include information representing the first object and the second object (e.g., “hand cream”). In this case, even if the first object is deleted from the original image, the second object identical to the first object remains in the original image, so the electronic device may maintain the first tag without modifying it.

[0181] According to various embodiments, the order of the operations of FIG. 14 may be changed, or at least some of the operations may be performed simultaneously. At least some of the operations of FIG. 14 may be omitted, or at least one operation (e.g., at least one of the operations of FIG. 10 to 13 and 15) may be added. For example, the operations of FIG. 14 may be performed independently of the operations of FIG. 10 to 13 and 15, performed in conjunction with each other, or at least some of the operations may be integrated.

[0182]

[0183] FIG. 15 is a flowchart of an image management method for an electronic device according to one embodiment. In the following, descriptions that overlap with the descriptions of FIG. 10 to 14 are omitted or briefly described.

[0184] According to one embodiment, in operation 1510, an electronic device (e.g., electronic device (100; 200; 1600; 1700)) may generate a first tag associated with an image at a first analysis level (which may be referred to as 'first level' or 'first description level' in this disclosure). For example, the electronic device may generate a first tag at the first analysis level using an AI model. The electronic device may pass a query to a trained AI model requesting a tag at the first analysis level for the image. For example, the analysis level is a level indicating the degree to which information representing the image is analyzed, and the higher the analysis level, the more information and / or more detailed information associated with the image (or, region of interest and / or object) the AI ​​model may provide a response. For example, the electronic device may generate a first tag associated with the image upon the acquisition of the image (e.g., creation, reception, capture, editing, and / or storage). When an electronic device acquires an image, if the image has an initial tag, it can use an AI model to analyze the initial tag of the image at a first analysis level and modify at least some of the initial tags of the image to generate a first tag at the first analysis level.

[0185] According to one embodiment, in operation 1520, the electronic device may recognize at least one of a region of interest or at least one object in an image based on user input. For example, user input may include an input for designating at least a portion of an image as a region of interest, an input for editing the image, an input for copying, extracting, or deleting at least a portion of an image, an input for enlarging at least a portion of an image, an input for copying, removing, extracting, or selecting (designating) at least one object in the image, and / or an input for adding at least one object to the image. User input is not limited to those listed above and may include various inputs capable of identifying a region associated with the image (e.g., a region of interest) and / or at least one object associated with the image.

[0186] According to one embodiment, in operation 1530, an electronic device may generate a second tag associated with at least one of a region of interest or at least one object at a second analysis level (which may be referred to as 'second level' or 'second description level' in this disclosure). For example, the electronic device may generate the second tag at the second analysis level using an AI model. The electronic device may pass a query to a trained AI model requesting a tag at the second analysis level for at least one of the region of interest or at least one object. For example, the tag at the first analysis level may contain more information and / or more detailed information about the subject of analysis (e.g., image, region of interest, and / or object) than the tag at the first analysis level. For example, even for the same object, the information about the object included in the second tag at the second analysis level may contain more information and / or more detailed information (e.g., a more detailed description) than the information about the object included in the first tag at the first analysis level.

[0187] According to one embodiment, in operation 1540, the electronic device may modify the first tag based on the second tag. The electronic device may modify the first tag based on information corresponding to the second tag. For example, the electronic device may add and / or delete information corresponding to the second tag from the first tag, and may create a new tag for the image based on information corresponding to the second tag. For example, the electronic device may create or modify a tag associated with the image based on the image (e.g., original image and / or edited image), the first tag, and / or the second tag.

[0188] For example, if an image is edited based on user input, the electronic device may modify the first tag based on the second tag, and then perform an analysis of the modified first tag at a first analysis level using an AI model. The electronic device may perform an analysis of the modified first tag at a first analysis level based on the edited image, and may modify at least a part of the modified first tag (e.g., the context of the modified first tag) again to correspond to the edited image.

[0189] According to various embodiments, the order of the operations of FIG. 15 may be changed, or at least some of the operations may be performed simultaneously. At least some of the operations of FIG. 15 may be omitted, or at least one operation (e.g., at least one of the operations of FIG. 10 to 14) may be added. For example, the operations of FIG. 15 may be performed independently of the operations of FIG. 10 to 14, performed in conjunction with each other, or at least some of the operations may be integrated.

[0190]

[0191] An image management method for an electronic device according to one embodiment may include an operation of displaying an original image.

[0192] According to one embodiment, the method may include the operation of receiving user input for the original image.

[0193] According to one embodiment, the method may include an operation of recognizing at least one of a point of interest (POI) of the original image or at least one object associated with the original image based on the user input.

[0194] According to one embodiment, the method may include the operation of generating a second tag associated with at least one of the region of interest or at least one of the at least one object.

[0195] According to one embodiment, the method may include an operation of modifying a first tag associated with the original image based on the second tag.

[0196] According to one embodiment, the operation of modifying the first tag may include, when the user input is an input for removing at least one object included in the original image, an operation of removing information corresponding to the second tag related to the at least one object removed from the original image based on the user input from the first tag.

[0197] According to one embodiment, the operation of recognizing at least one of the region of interest or at least one of the at least one object may include, when the user input is an input for enlarging the original image, recognizing a portion of the original image that is enlarged and displayed on the display as the region of interest.

[0198] According to one embodiment, the operation of generating the second tag may include the operation of generating the second tag associated with the new object when the user input is an input for adding a new object to the original image.

[0199] According to one embodiment, the operation of modifying the first tag may include the operation of modifying the first tag to include information corresponding to the second tag associated with the new object.

[0200] According to one embodiment, the operation of modifying the first tag may include the operation of regenerating a tag related to the original image based on information corresponding to the first tag and information corresponding to the second tag using a learned artificial intelligence model.

[0201] According to one embodiment, the method may include an operation of determining whether a second object of the same type as the first object is included in the original image after removing the first object from the original image, when the user input is an input for removing a first object included in the original image.

[0202] According to one embodiment, the method may include an operation to maintain the first tag when the original image contains the second object.

[0203] According to one embodiment, the method may include the operation of generating a first image corresponding to the region of interest.

[0204] According to one embodiment, the method may include the operation of receiving a second user input for the first image.

[0205] According to one embodiment, the method may include an operation of recognizing at least one of a region of interest of a first image or at least one object associated with the first image based on the second user input.

[0206] According to one embodiment, the method may include the operation of generating a third tag associated with a region of interest of the first image or at least one of at least one object included in the first image.

[0207] According to one embodiment, the method may include an operation of modifying the second tag based on the third tag.

[0208] According to one embodiment, the method may include an operation of modifying the first tag based on the modified second tag.

[0209] According to one embodiment, the method may include the operation of associating and storing the original image and the first tag.

[0210] According to one embodiment, the method may include the operation of associating and storing an image corresponding to at least one of the region of interest or at least one of the at least one object with the second tag.

[0211] According to one embodiment, the method may include an operation of generating the first tag using a learned AI model when acquiring the original image.

[0212] A storage medium according to one embodiment may store a program and / or instructions that, when executed by at least one processor of an electronic device, cause the electronic device to display an original image, receive user input regarding the original image, recognize at least one of a point of interest (POI) of the original image or at least one object associated with the original image based on the user input, generate a second tag associated with at least one of the point of interest or at least one object, and modify a first tag associated with the original image based on the second tag.

[0213] A method and a recording medium according to one embodiment can improve performance when searching, classifying, and / or utilizing images by expanding the metadata (e.g., tags) of an image based on user input.

[0214]

[0215] FIG. 16 is a block diagram of an exemplary electronic device (1600) capable of performing the operations described in this document.

[0216] Referring to FIG. 16, the electronic device (1600) may be one of various forms of electronic devices, such as a notebook (1690), smartphones (1691) having various form factors (e.g., a bar-type smartphone (1691-1), a foldable-type smartphone (1691-2), or a sliderable (or rollable)-type smartphone (1691-3)), a tablet (1692), a cellular phone (not shown), and other similar computing devices (not shown). The components, their relationships, and their functions illustrated in FIG. 15 are illustrative only and are not intended to limit the implementations described or claimed herein. The electronic device (1600) may be referred to as a mobile device, a user device, a multifunction device, a portable device, or a server.

[0217] The electronic device (1600) may include components comprising at least one processor (1610) (hereinafter referred to as processor (1610)), at least one memory (1620) (hereinafter referred to as memory (1620)), at least one display (1640) (hereinafter referred to as display (1640)), at least one image sensor (1650) (hereinafter referred to as image sensor (1650)), at least one communication circuit (1660) (hereinafter referred to as communication circuit (1660)), and / or at least one sensor (1670) (hereinafter referred to as sensor (1670)). The components are merely exemplary. For example, the electronic device (1600) may include other components (e.g., power management integrated circuitry (PMIC), audio processing circuit, antenna, rechargeable battery, or input / output interface). For example, some components may be omitted from the electronic device (1600). For example, some components can be integrated into a single component.

[0218] The processor (1610) may be implemented as one or more IC (integrated circuit (or circuitry)) chips and may perform various data processing operations. The processor (1610) may include at least one electrical circuit and may process instructions (or programs, data, etc.) stored in memory (1620) individually or collectively in a distributed manner. The processor (1610) may include a processor assembly comprising one or more processing circuits. The processor (1610) may include any processing circuit that is operative to control the performance and operations of one or more components of the electronic device (1600) (e.g., memory (1620), display (1640), image sensor (1650), communication circuit (1660), and / or sensor (1670)). For example, a processor (1610) (e.g., an application processor (AP)) may be implemented as a system on chip (SoC) (e.g., a single chip or a chipset). For example, the processor (1610) may be implemented as a plurality of cores (or at least one core circuit), a plurality of chips, or a plurality of chipsets. For example, the processor (1610) may include one or more processing circuits. For example, the processor (1610) may include one or more processing circuits configured to perform the various functions of the present disclosure individually and / or collectively. As an example without limitation, at least a portion of the processor (1610) may be included in a first chip of the electronic device (1600), and at least another portion of the processor (1610) may be included in a second chip of the electronic device (1600) different from the first chip of the electronic device (1600).

[0219] For example, the processor (1610) may include a central processing unit (CPU) (1611), a graphics processing unit (GPU) (1612), a neural processing unit (NPU) (1613), an image signal processor (ISP) (1614), a display controller (1615), a memory controller (1616), a storage controller (1617), a communication processor (CP) (1618), and / or a sensor interface (1619). These components of the processor (1610) are merely exemplary. For example, the processor (1610) may include other components. For example, some components of the processor (1610) may be omitted from the processor (1610). For example, some components of the processor (1610) may be included as separate components of the electronic device (1600) outside of the processor (1610). For example, some components of the processor (1610) (e.g., memory controller (1616)) may be included in other components (e.g., at least part of memory (1620), an interface (e.g. available for connection to at least one component of the electronic device (100)), a display (1640) and / or an image sensor (1650)).

[0220] The processor (1610) may cause other components of the electronic device (1600) to perform various operations by executing instructions stored in memory (1620). The CPU (1611) (or central processing circuit) may be configured to control the components of the processor (1610) based on the execution of instructions stored in memory (1620) (e.g., volatile memory (1621) and / or non-volatile memory (1622)). The GPU (1612) (or graphics processing circuit) may be configured to execute parallel operations (e.g., rendering). The NPU (1613) (or neural processing circuit, or AI (artificial intelligence) chip) may be configured to execute operations for an artificial intelligence model (e.g., convolution computation). An ISP (1614) (or image signal processing circuit) may be configured to process a raw image acquired through an image sensor (1650) into a format suitable for a component within an electronic device (1600) or a component of a processor (1610). A display controller (1615) (or display control circuit, or DPU (display processing unit)) may be configured to process an image acquired from a CPU (1611), GPU (1612), ISP (1614), or memory (1620) (e.g., volatile memory (1621)) into a format suitable for a display (1640). A memory controller (1616) (or memory control circuit) may be configured to control reading data from volatile memory (1621) and writing data to volatile memory (1621). The storage controller (1617) (or storage control circuit) may be configured to control reading data from non-volatile memory (1622) and writing data to non-volatile memory (1622).The CP (1618) (communication processing circuit) may be configured to process data obtained from a component of the processor (1610) into a format suitable for transmitting to another electronic device via the communication circuit (1660), or to process data obtained from another electronic device via the communication circuit (1660) into a format suitable for processing by the component of the processor (1610). For example, the communication circuit (1660) may include one or more communication circuits. The sensor interface (1619) (or sensing data processing circuit, sensor hub) may be configured to process data regarding the state of the electronic device (1600) and / or the state around the electronic device (1600), obtained through the sensor (1670), into a format suitable for the component of the processor (1610).

[0221] Memory (1620) may include one or more storage media (or one or more storage devices). For example, memory (1620) may include a memory assembly comprising one or more storage media. For example, the one or more storage media may include a hard drive, a flash memory, a permanent memory such as ROM (read-only memory) (e.g., non-volatile memory (1622)), a semi-permanent memory such as RAM (random access memory) (e.g., volatile memory (1621)), any other suitable type of storage (or storage assembly), or any combination thereof. Memory (1620) may include a cache memory, which is one or more different types of memory used to temporarily store data for a function or feature of the electronic device (1600). As an example not limited to, the cache memory may be included within the processor (1610). The memory (1620) may be fixedly embedded within the electronic device (1600) or incorporated into one or more suitable types of components (e.g., a SIM (subscriber identity module) card and / or an SD (secure digital) card) that can be repeatedly inserted into and removed from the electronic device (1600).

[0222] For example, memory (1620) may store one or more software applications, such as operating system (or system) software applications, firmware software applications, driver software applications, plugin (e.g., add-in, add-on, and / or applet) software applications, and / or any other suitable software applications. For example, the one or more software applications may include instructions executable by the processor (1610). For example, memory (1620) may store instructions that can be called by an application programming interface (API). For example, memory (1620) may store instructions within a library.

[0223]

[0224] Referring to FIG. 17, in a generative artificial intelligence system (1700), a user question / response interface (1710) can receive user input. The user input may be in the form of natural language, images, and / or videos. Additionally, context information may be transmitted along with the user input. The context information may include various additional information at the time of user input. For example, information about the application currently being used by the user or the user's location information. Furthermore, the user input may be in a mixed form of the aforementioned natural language, images, sounds, and context information. Additionally, the user input may be in a non-natural language form, such as selecting a menu.

[0225] According to one embodiment, the user question / response interface (1710) can output results of a generative artificial intelligence system to the user. The output may be in the form of natural language or specific content, and may also be provided in the form of actions requested by the user. The user question / response interface (1710) can output results of a generative artificial intelligence system (1700) to the user. The output may be in the form of natural language or specific content, and may also be provided in the form of actions requested by the user.

[0226] According to one embodiment, the artificial intelligence framework (1720) can receive input from a user and coordinate and control each component or module necessary to perform the user's intent based on the user's query.

[0227] According to one embodiment, user input received from a user question / response interface (1710) may be transmitted to a prompt design module (1721). The prompt design module (1721) may be used to generate prompts suitable for inputting user input into a large language model (LMM) or a large multi-modal model (LMM). The prompt design module (1721) may be an artificial intelligence component that uses machine learning algorithms or neural networks to develop better prompts over time. The prompt design module (1721) may generate prompts by accessing a database (1730) containing user preference data, a prompt library, and prompt examples based on user input, and transmit the generated prompts to the LLM or LMM.

[0228] According to one embodiment, the API / plugin management module (1722) can perform the role of communicating with external information when there is a request for additional information when transmitting user input as input to the generative artificial intelligence model (1750). The API / plugin management module (1722) establishes a channel to communicate with the outside of the artificial intelligence interface through an application programming interface (API), and can enable access to various data sources through the established channel. Additionally, the API / plugin management module (1722) can request the action through the API if the application / service module (1740) needs to perform an action that executes the user input as a final step, rather than an intermediate result. The information obtained from the outside may be used to generate a prompt in the prompt design module (1721) along with the user input, or it may be transmitted as input to the generative artificial intelligence model (1750).

[0229] According to one embodiment, the transformation module (1723) can fine-tune the output from the generative artificial intelligence model (1750). For example, the transformation module (1723) can verify whether the content generated through the LLM and / or LMM is irrelevant, contains biased content, or contains harmful content. Additionally, the transformation module (1723) can determine the extent to which the output matches what the user wants and, if additional processing is required, proceed with that process. Furthermore, the transformation module (1723) can configure and provide hints to the user to avoid unwanted output.

[0230] According to one embodiment, a generative artificial intelligence model (1750) may generally refer to an artificial intelligence neural network that generates new forms of data based on user input information. The generative artificial intelligence model (1750) may include a model that generates images and / or a model that generates language. Models that generate images include, but are not limited to, GANs (generative adversarial networks) and VAEs (variational autoencoders), and examples include diffusion-based generative models that use VAEs and Transformer structures. Models that generate language are models trained to output the most statistically appropriate output value based on input values, and examples include models such as CHAT-GPT 3 and CHAT-GPT 4. There are also LMMs that can recognize various forms of data input, such as text, images, and voice, and generate new data corresponding to them.

[0231]

[0232] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as "coupled" or "connected" to another (e.g., 2nd) component, with or without the terms "functionally" or "communicationly," it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.

[0233] As used in the various embodiments of this document, the term “module” may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions.

[0234] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In an electronic device, display; Memory for storing instructions; and It includes at least one processor comprising processing circuitry, and When the above instructions are executed individually or collectively by the at least one processor, the electronic device, The original image is displayed through the above display, and Receive user input for the above original image, and Based on the above user input, at least one of a point of interest (POI) of the original image or at least one object associated with the original image is recognized, and Generate a second tag associated with at least one of the above-mentioned region of interest or at least one of the above-mentioned at least one object, and An electronic device that modifies a first tag associated with the original image based on the second tag above.

2. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that, when the above user input is an input for removing at least one object included in the above original image, removes information corresponding to the second tag associated with the at least one object removed from the above original image based on the above user input from the first tag.

3. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that recognizes a portion of the original image, which is enlarged and displayed on the display, as the region of interest when the above user input is an input for enlarging the above original image.

4. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, If the above user input is an input for adding a new object to the above original image, the above second tag related to the above new object is generated, and An electronic device that modifies the first tag to include information corresponding to the second tag associated with the new object.

5. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that uses a learned artificial intelligence model to regenerate tags associated with the original image based on information corresponding to the first tag and information corresponding to the second tag.

6. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, If the above user input is an input for removing a first object included in the original image, after removing the first object from the original image, determine whether the original image contains a second object of the same type as the first object, and An electronic device that maintains the first tag when the original image contains the second object.

7. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, A first image corresponding to the above region of interest is generated, and Receiving a second user input for the first image above, and Based on the second user input above, at least one of the region of interest of the first image or at least one object associated with the first image is recognized, and Generate a third tag associated with at least one of the region of interest of the first image or at least one object associated with the first image, and Based on the above third tag, modify the above second tag, and An electronic device that modifies the first tag based on the modified second tag.

8. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, The above original image and the above first tag are associated and stored, An electronic device that stores an image corresponding to at least one of the above-mentioned region of interest or at least one of the above-mentioned at least one object and the above-mentioned second tag in association.

9. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that generates the first tag using a trained AI model when acquiring the original image.

10. In a method for managing images of an electronic device, Action of displaying the original image; The operation of receiving user input regarding the above original image; An operation to recognize at least one of a point of interest (POI) of the original image or at least one object associated with the original image based on the above user input; The operation of generating a second tag associated with at least one of the above-mentioned region of interest or at least one of the above-mentioned at least one object; and A method comprising the operation of modifying a first tag associated with the original image based on the second tag.

11. In Claim 10, The operation of modifying the above-mentioned first tag is, A method comprising removing information corresponding to the second tag associated with the at least one object removed from the original image based on the user input, when the user input is an input for removing at least one object included in the original image.

12. In Claim 10, The operation of generating the above second tag is, If the above user input is an input for adding a new object to the above original image, the method includes the operation of generating the above second tag associated with the above new object, and The operation of modifying the above-mentioned first tag is, A method comprising the operation of modifying the first tag to include information corresponding to the second tag associated with the new object.

13. In Claim 10, If the above user input is an input for removing a first object included in the above original image, an operation to determine whether the above original image contains a second object of the same type as the first object after removing the first object from the above original image; and A method including the operation of maintaining the first tag when the second object is included in the original image.

14. In Claim 10, The operation of generating a first image corresponding to the above-mentioned region of interest; The operation of receiving a second user input for the first image above; An operation of recognizing at least one of a region of interest of the first image or at least one object associated with the first image based on the second user input; The operation of generating a third tag associated with a region of interest of the first image or at least one of at least one object included in the first image; An operation to modify the second tag based on the third tag; and A method comprising the operation of modifying the first tag based on the modified second tag.

15. In a storage medium for storing computer-readable instructions, When the above instructions are executed by at least one processor of the electronic device, the electronic device, Display the original image, and Receive user input for the above original image, and Based on the above user input, at least one of a point of interest (POI) of the original image or at least one object associated with the original image is recognized, and Generate a second tag associated with at least one of the above-mentioned region of interest or at least one of the above-mentioned at least one object, and A storage medium that modifies a first tag associated with the original image based on the second tag above.

Citation Information

Patent Citations

  • Method and apparatus for tagging of portable terminal

    KR1020100113334A

  • Method and apparatus for updating tag information of content

    KR1020140114246A

  • Mobile terminal and method for controlling the same

    KR1020150082841A

  • Water injection instrument for fire extinguishing

    KR1020230168304A

  • KR20230068780A