Electronic device for modifying image, operation method thereof, and storage medium
The electronic device uses AI models to automate image editing, enabling users to modify images by changing or adding objects and text, addressing the limitations of existing technologies in image editing automation and user interaction.
Patent Information
- Application Number
- PCT/KR2025/010514
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-31
- Filing Date
- 2025-07-17
- Publication Date
- 2026-02-05
AI Technical Summary
Existing image editing technologies lack efficient methods for automated and user-driven modification of images using artificial intelligence, particularly in changing, adding, or deleting objects, while maintaining natural image quality.
An electronic device equipped with processors and memory, utilizing AI models like CNN and GAN, allows users to modify images by changing or adding text and objects based on user input, with features like object recognition, text editing, and image enhancement.
Enables efficient and natural image modification by allowing users to change or add objects and text, improving the automation and user interaction in image editing processes.
Smart Images

Figure KR2025010514_05022026_PF_FP_ABST
Abstract
Description
Electronic device for modifying images, method of operation thereof and storage medium
[0001] The present disclosure relates to an electronic device for modifying an image, a method of operating the same, and a storage medium.
[0002] Image processing technology plays a crucial role in the creation of advertising, entertainment, and social media content. While traditional image editing, for example, was handled manually, recent advances in artificial intelligence have made automated image editing possible.
[0003] For example, image modification, such as changing, adding, and / or deleting objects in an image, can be performed using an artificial intelligence model. During image modification, for example, an artificial intelligence model (e.g., a convolution neural network (CNN) model) can be used to recognize and / or classify objects. The CNN model can be used to learn the visual features of an image to recognize and classify each object. Accordingly, various objects within the image can be recognized. During image modification, for example, a generative adversarial network (GAN) model can be used to change, add, and / or delete objects. The GAN model can be composed of two networks, namely, a generator and a discriminator. The generator attempts to generate natural results during the process of adding or modifying objects in an image, and the discriminator can determine how close the generated image is to reality, thereby enabling natural image modification. Meanwhile, in addition to the aforementioned artificial intelligence models, various artificial intelligence models, such as autoencoders and transformers, can be used for image modification.
[0004] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.
[0005] In one or more embodiments of the present disclosure, an electronic device (101) may include a display; one or more processors (120); and a memory (130) storing instructions. The instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to: provide an image and texts describing the image; change first text included in the texts to second text based on a user input; and provide a modified image in which an object is created, removed, or modified based on the second text.
[0006] In one or more embodiments of the present disclosure, a non-transitory computer-readable storage medium is provided storing one or more instructions comprising computer-executable instructions, which when individually or collectively executed by one or more processors (120) of an electronic device (101), cause the electronic device (101) to perform the following operations: providing an image and texts describing the image; changing a first text included in the texts to a second text based on a user input; and providing a modified image in which an object is created, removed, or modified based on the second text.
[0007] In one or more embodiments of the present disclosure, a method of operating an electronic device (101) may include: providing an image and texts describing the image; changing first text included in the texts to second text based on a user input; and providing a modified image in which an object is created, removed, or modified based on the second text.
[0008] Figure 1 is a block diagram of an electronic device according to one embodiment.
[0009] Figure 2 is a flowchart for explaining an operating method of an electronic device according to one embodiment.
[0010] FIG. 3a is an example of a screen provided according to various embodiments.
[0011] Figure 3b is an example of a screen provided according to various embodiments.
[0012] FIG. 3c is an example of a screen provided according to various embodiments.
[0013] FIG. 4A is a flowchart illustrating an operating method of an electronic device according to one embodiment.
[0014] FIG. 4b is a diagram illustrating text provision according to various embodiments.
[0015] FIG. 4c is a diagram illustrating text provision according to various embodiments.
[0016] FIG. 5A is a flowchart illustrating an operating method of an electronic device according to one embodiment.
[0017] FIG. 5b is a diagram for explaining candidate selection according to one embodiment.
[0018] FIG. 6 is a flowchart illustrating an operating method of an electronic device according to one embodiment.
[0019] FIG. 7A is a flowchart illustrating an operating method of an electronic device according to one embodiment.
[0020] Figure 7b is a drawing for explaining object addition according to one embodiment.
[0021] FIG. 8 is a drawing for explaining image modification by an electronic device according to one embodiment.
[0022] FIG. 9 is a drawing for explaining image modification by an electronic device according to one embodiment.
[0023] FIG. 10 is a drawing for explaining image modification by an electronic device according to one embodiment.
[0024] FIG. 11 is a drawing for explaining image modification by an electronic device according to one embodiment.
[0025] FIG. 12 is a drawing for explaining image modification by an electronic device according to one embodiment.
[0026] FIG. 13a is a diagram for explaining image modification by an electronic device according to one embodiment.
[0027] FIG. 13b is a diagram for explaining image modification by an electronic device according to one embodiment.
[0028] FIG. 14 is a drawing for explaining image modification by an electronic device according to one embodiment.
[0029] FIG. 15A is a diagram for explaining image modification by an electronic device according to various embodiments.
[0030] FIG. 15b is a diagram for explaining image modification by an electronic device according to various embodiments.
[0031] Figure 1 is a block diagram of an electronic device according to one embodiment.
[0032] According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), and / or a display (160). In some embodiments, at least one of these components may be omitted, or one or more other components may be added to the electronic device (101). In some embodiments, some of these components may be integrated into a single component, and there is no limitation on the implementation thereof. For example, the electronic device (101) may perform at least some of the operations described in the present disclosure in conjunction with the server (108). This operational configuration may be referred to as a non-stand alone (NSA) mode, in which computational tasks, data processing, or model inference may be partially or fully offloaded to the server (108) via a network (e.g., Wi-Fi, 5G, or other communication interfaces). For example, the electronic device (101) may perform at least some of the operations described in the present disclosure independently, without being associated with the server (108). This configuration may also be referred to as stand-alone (SA) mode, or on-device mode. Those skilled in the art will appreciate that each of the operations described in this disclosure may be performed by the electronic device (101), the server (108), or both entities.
[0033] The processor (120) may execute, for example, at least one instruction stored in the memory (130). The memory (130) may store at least one instruction, and the at least one instruction may be executed by the processor (120). For example, the memory (130) may include non-volatile memory and / or volatile memory, without limitation. The memory (130) may include a hard disk, a read-only memory (ROM), a random access memory (RAM), a cache memory, and / or a register, without limitation in its implementation. Some of the above-described entities (for example, but not limited to, a register) may be implemented as a part of the processor (120), without limitation in their implementation form. At least one instruction, when executed by the processor (120), may cause the electronic device (101) to perform at least one operation. For example, as at least one instruction is executed, at least one other component may be controlled, and / or various data processing or operations may be performed. As at least a part of the data processing or operations, the processor (120) may store instructions or data received from other components in at least a part of the memory (130), process the instructions or data stored in the memory (130), and store result data in the memory (130). The processor (120) may include a main processor (e.g., a central processing unit) including circuitry, or an auxiliary processor (e.g., a graphics processing unit, a neural network processing unit, an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together therewith. For example, performing a specific operation may mean that the specific operation is performed by (or under the control of) one entity (e.g., the main processor).For example, performing a specific operation may mean that the specific operation is performed by (or under the control of) a plurality of entities (e.g., but not limited to, a main processor and one or more auxiliary processors). For example, performing a plurality of operations may mean that all of the plurality of operations are performed by (or under the control of) a single entity (e.g., a main processor). For example, performing a plurality of operations may mean that some of the plurality of operations are performed by at least one entity, and some of the remaining operations are performed by at least one other entity. Meanwhile, at least one instruction for performing a specific operation may be stored in a single memory, or may be stored in a distributed manner in each of a plurality of memories.
[0034] The display (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). For example, when the electronic device (101) is implemented as a smart phone, a tablet PC, a video see through (VST) device, or a head-mounted display (HMD), the display (160) may be implemented to include, for example, a liquid crystal display (LCD) and a control means (for example, a display driving integrated circuit (DDI)). The display (160) may further include a touch screen panel (TSP) for touch detection and / or a control means (for example, a TSP integrated circuit (TSP IC)). For example, when the electronic device (101) is implemented as an augmented reality (AR) glasses device, the display (160) may be implemented to include, for example, a light irradiation device, an optical waveguide, and / or a control means. The processor (120) may control the display (160) to express an object. The expression of the object may be visual presentation, rendering, or digital on the display. Alternatively, it may represent the representation of a graphic object, which will be described in more detail later. For example, the processor (120) may generate control data to initiate or manage the representation of an object and transmit it to the display (160), which may be expressed as the processor (120) controlling the display (160).
[0035] Meanwhile, although not separately illustrated in FIG. 1, the electronic device (101) may further include a communication interface (170). The communication interface (170) may be implemented by any one or any combination of a digital modem, a radio frequency (RF) modem, a communication circuit, an antenna circuit, a WiFi chip, and related software and / or firmware. The electronic device (101) may transmit / receive data to / from an external electronic device, for example, a server (108), through the communication interface (170). The electronic device (101) may, for example, request the server to perform some or all of the operations performed by the electronic device (101) in various embodiments of the present disclosure, and may receive the performance results in response to the request. As will be described in more detail below, the electronic device (101) may provide text describing an object included in an image, modify a portion of the text, and / or modify an image based on the modification. The electronic device (101) may request the server (108) to provide text describing an object included in an image, modify a portion of the text, and / or modify part or all of the image based on the modification, and may also receive a result of the performance as a response to the request. The electronic device (101) may provide a result of performing a function based on the result of the performance received from the server (108).
[0036] FIG. 2 is a flowchart illustrating an operating method of an electronic device according to one embodiment.
[0037] The embodiment of Fig. 2 will be described with reference to Figs. 3a, 3b, and 3c.
[0038] FIG. 3a is an example of a screen provided according to various embodiments.
[0039] FIG. 3b is an example of a screen provided according to various embodiments.
[0040] FIG. 3c is an example of a screen provided according to various embodiments.
[0041] According to one embodiment, the electronic device (101) may, in operation 201, provide texts (321, 322, 323, 324, 325, 326) to describe an image (310) and one or more objects (311, 312, 313, 314, 315) included in the image (310), as illustrated in FIG. 3A. The term "text" may represent a character, a word, a phrase, or a sentence. For example, the electronic device (101) may provide texts (321, 322, 323, 324, 325, 326) based on an editing request for the image (310), but those skilled in the art will understand that this is exemplary and that there is no limitation to the provision event of the texts (321, 322, 323, 324, 325, 326). The electronic device (101) may display sentences or phrases (e.g., "Me walking with Minah on the beach with the sunset at my back") that provide an overall description of the image (310) along with visual segmentations corresponding to individual texts (321, 322, 323, 324, 325, 326). For example, the electronic device (101) may display bounding boxes around the texts (321, 322, 323, 324, 325, 326), and the user may select one or more for editing.
[0042] In the embodiment of FIG. 3A, the image (310) and the texts (321, 322, 323, 324, 325, 326) are depicted as being provided together on the same display screen, but this is merely exemplary. The image (310) and the texts (321, 322, 323, 324, 325, 326) may each be provided on different display screens, and there is no limitation on the order in which they are provided. In the embodiment of FIG. 3A, the texts (321, 322, 323, 324, 325, 326) are depicted as being provided so as not to overlap the image (310), but this is merely exemplary. For example, those skilled in the art will appreciate that the texts (321, 322, 323, 324, 325, 326) may be expressed to be positioned on at least a portion of the image (310). For example, based on extracting an object from an image, recognizing the extracted object, and / or providing a recognition result, text describing the object included in the image may be provided. The provision of text describing the object included in the image may be provided, for example, as an inference result of artificial intelligence, and the artificial intelligence model may include, but is not limited to, ResNet (residual network), VGGNet (visual geometry group network), Inception (inception network), YOLO (You Only Look Once), Faster R-CNN (region-based convolutional neural network), Mask R-CNN, or a Scene Recognition model. The provision of text will be described with reference to FIGS. 4a, 4b, and 4c.
[0043] For example, referring to FIG. 3A, text corresponding to an image (310) (e.g., "Me walking with Minah on the beach with my back to the sunset sky") may include a first part (321) corresponding to an object (311) in the image (310). For example, the first part (321) may be determined based on the recognition result of the object (311) being "sunset sky." The text corresponding to the image (310) may include a third part (323) corresponding to multiple objects (312, 313). For example, the text "beach" may be identified based on the recognition result of the objects (312, 313) being the sea and the surface, and the text "at" may be identified based on the attribute of the text being a place. For example, the text corresponding to the image (310) may include a fourth part (324) corresponding to an object (314). The text of "Min-ah" can be confirmed based on the recognition result of the object (314) being confirmed as a person stored as "Min-ah", and the text of "wa" can be confirmed based on the existence of multiple persons. The text corresponding to the image (310) can include a sixth part (326) corresponding to the object (315). The text of "I" can be confirmed based on the recognition result of the object (315) being confirmed as a person stored as "user". The text corresponding to the image (310) can include a second part (322) and a fifth part (325) that are confirmed based on the relationship between the objects (312, 313, 314, 315). For example, the second part (322) can be confirmed based on the depth value of the objects (314, 315) being smaller than the depth value of the object (311) and a part of the objects (314, 315) being recognized as a face (or, a front view of a human). For example, based on the relationship between the third part (323) of "on the beach" and the texts corresponding to the characters (324, 326), the fifth part (325) of the verb form "to walk" can be identified.Object recognition processes for identifying objects (311, 312, 313, 314, 315) and identification of texts (321, 322, 323, 324, 325, 326) corresponding to the objects can be performed by the electronic device (101) in a standalone mode or in cooperation with the server (108).
[0044] In one embodiment, if some texts are editable and others are not, the electronic device (101) may, but is not limited to, display editable texts among the texts (321, 322, 323, 324, 325, 326) as distinct from other texts. Those skilled in the art will appreciate that the electronic device (101) may be implemented to support editing corresponding to all texts.
[0045] The electronic device (101) may, in operation 203, identify a user input that causes a change of a first text included in the texts to a second text. The electronic device (101) may, in operation 205, change the first text to the second text based on the user input. In one example, the electronic device (101) may, based on the user input (327) for the first text (321), provide a plurality of replaceable candidates (331, 332, 333, 334), as illustrated in FIG. 3B . There is no limitation on the provision location and / or presentation method of the list (330) including the plurality of candidates (331, 332, 333, 334). Based on the confirmation of the selection (335) of the second text (332) among the multiple candidates (331, 332, 333, 334), the electronic device (101) can change the first text to the second text. In one example, the electronic device (101) can change the first text to the second text based on a user input of the second text to replace the first text (for example, but not limited to, input via a soft input panel (SIP)). The soft input panel (SIP) can represent an on-screen keyboard or a virtual keyboard that enables text input without physical keys. In some embodiments, the electronic device (101) can change the first text to the second text based on the analysis result of the user's voice from the user.
[0046] The electronic device (101) may, in operation 207, provide a modified image (350) including a modified object (351) based on the second text (371), as illustrated in FIG. 3C . The modified image (350) may include "blue sky" instead of "sunset sky" as the modified object (351) based on the second text (371). In this example, attributes of the object (e.g., color, brightness, and / or saturation, but not limited thereto) based on the selected text may be changed. For example, although it has been described that the shape of the object is maintained while the attributes are changed, this is exemplary and not limiting. The modified object (351) may be provided based on, for example, an inference result of an artificial intelligence model, but is not limited thereto. For example, the electronic device (101) may input a prompt for changing the first text to the second text and an image (310) prior to modification into a generative artificial intelligence model. Although FIG. 3C illustrates an example of modifying an existing object within the original image, embodiments of the present disclosure are not limited thereto. The modified image may also be generated by adding a new object to the original image, for example, by replacing the first text (321) with the second text (371).
[0047] According to one embodiment, a generative artificial intelligence model may provide an image (350) including a modified object (351). The generative AI model may include, but is not limited to, a generative adversarial network (GAN), a conditional GAN, Pix2pix, CycleGAN, or Deep image matting, for example. The remaining objects (352, 353, 354, 355) of the modified image (350) may be identical to each of the objects (312, 313, 314, 315) of the image (310) before modification, and / or may be generated by modifying some of the objects (312, 313, 314, 315) of the image (310) before modification. For example, the brightness levels of figures under a blue sky may be different from the brightness levels of figures under a sunset sky, and those skilled in the art will understand that each of the objects (352, 353, 354, 355) may be modified or maintained based on the modification of the object (351).
[0048] FIG. 4A is a flowchart illustrating a method of operating an electronic device according to one embodiment.
[0049] The embodiment of Fig. 4a will be described with reference to Figs. 4b and 4c.
[0050] FIG. 4b is a diagram illustrating text provision according to various embodiments.
[0051] FIG. 4c is a diagram illustrating text provision according to various embodiments.
[0052] According to one embodiment, the electronic device (101) may, in operation 401, extract one or more objects included in an image (431) as illustrated in FIG. 4B. The electronic device (101) may recognize one or more objects. Segmentation and / or recognition of segmented objects can be performed based on, for example, Otsu's method, edge detection method, region based method, k-means clustering method, random forest method, FCN (fully convolutional networks), U-Net method, SegNet method, Mask R-CNN method, DeepLab method, PSPNet method, HRNet (high-resolution network) method, dellLabv3 method, Semantic segmentation networks method (deeplab, PSPNet), instance segmentation network (PANet, YOKACT), transformer-based model (DETR, Swin transformer method), graph-based method (GCN), attention method (self-attention, non-local networks), multi-task learning (multi-task networks), etc., but there is no limitation on the method.
[0053] The electronic device (101) can identify a first part of the texts by recognizing one or more extracted objects in operation 403. The electronic device (101) can identify a second part of the texts by identifying at least one adjective corresponding to the one or more extracted objects in operation 405. For example, referring to FIG. 4B, the electronic device (101) can recognize an object (441) corresponding to a person, an object (442) corresponding to a pet, and another object (443) from an image (431). The electronic device (101) can recognize an object (451) corresponding to an unknown person (a person without a known recognition result), an object (452) corresponding to terrain, an object (453) corresponding to the sky, an object (454) corresponding to the sea, an object (455) corresponding to a major building, or an object (456) corresponding to a background from the image (431). Each object can be recognized based on an artificial intelligence model trained specifically for that object and / or a general image recognition model, and there is no limitation on the type and / or number of the artificial intelligence models.
[0054] The electronic device (101) can identify features (461, 462, 463, 464, 465, 471, 472, 473, 474) associated with objects (441, 442, 443, 451, 452, 453, 454, 455, 456). The features (461, 462, 463, 464, 465, 471, 472, 473, 474) can be, for example, text of adjectives to describe the objects (441, 442, 443, 451, 452, 453, 454, 455, 456), but are not limited thereto. In one example, features (461, 462, 463, 464, 465, 471, 472, 473, 474) may be identified as part of the object recognition result. Alternatively, or additionally, those skilled in the art will understand that features (461, 462, 463, 464, 465, 471, 472, 473, 474) may be identified based on the inference results of an additional artificial intelligence model for the object recognition result. For example, an artificial intelligence model for recognizing a posture / pose, which is a feature (461), may be implemented independently from an artificial intelligence model for identifying the type of object, or may be implemented as one artificial intelligence model. In the case where the artificial intelligence model for recognizing the feature is independent from the artificial intelligence model for identifying the type of object, those skilled in the art will understand that the electronic device (101) may select an artificial intelligence model to additionally utilize depending on the type of the identified object.
[0055] The electronic device (101) can provide images and texts in operation 407. For example, as illustrated in FIG. 4C, the recognition results of objects (485, 486, 487, 488, 489, 490, 491) can be confirmed based on the segmentation results (481, 482, 483). The electronic device (101) can also confirm the characteristics of objects (485, 486, 487, 488, 489, 490, 491) as described above (e.g., the color of the sky is blue, etc.). The electronic device (101) can confirm the text (495) corresponding to the image (431) based on the recognition results of objects, the characteristics of objects, and / or the relationships between objects.
[0056] FIG. 5A is a drawing for explaining an operation method of an electronic device according to one embodiment.
[0057] The embodiment of Fig. 5a will be described with reference to Fig. 5b.
[0058] FIG. 5b is a diagram for explaining candidate selection according to one embodiment.
[0059] According to one embodiment, the electronic device (101) may, in operation 501, provide an image and associated texts. Since the method of providing texts for describing one or more objects included in the image has been described above, the description thereof will not be repeated here. The electronic device (101) may, in operation 503, confirm a selection of a first portion among the texts. There is no limitation on the method for selecting the first portion. Based on the selection of the first portion, the electronic device (101) may, in operation 505, provide one or more candidate texts for the first portion. For example, the electronic device (101) may assign a priority (511) to one or more objects included in the image and / or addable objects, as illustrated in FIG. 5B . For example, the electronic device (101) may confirm multiple candidates for a createable group (520). For example, the electronic device (101) can check candidate texts (521, 522, 523) for the text of "me frowning." For example, the candidate text (521) of "me smiling" can be given the highest priority, the candidate text (522) of "me waving" can be given a second priority, and the candidate text (523) of "me standing" can be given a third priority. For example, when "me frowning" is selected among the texts for the image, the electronic device (101) can provide at least some of the candidate texts (521, 522, 523) to enable the user to select.
[0060] For example, the electronic device (101) can identify applicable modification operations for a person based on whether "Frowning Me" corresponds to the person. The applicable modification operations can be, for example, preset. For example, an artificial intelligence model executed by (or accessible to) the electronic device (101) can support the available modification operations, and the electronic device (101) can identify the supported modification operations. The electronic device (101) can provide text for at least some of the pre-specified supported modification operations as candidate text.
[0061] For example, the electronic device (101) can inquire about possible editable tasks and identify candidate texts as responses to the inquiry. For example, the candidate texts can be provided based on a question-and-answer interaction based on a chat-based method. For example, the electronic device (101) can inquire about possible editable tasks for at least some of the texts describing the image to an artificial intelligence model (e.g., a generative artificial intelligence model based on a chat-based method, but is not limited thereto). The artificial intelligence model can provide texts for possible editable tasks as responses, and the electronic device (101) can provide texts for possible editable tasks identified based on the responses as candidate texts. Meanwhile, those skilled in the art will understand that the above-described method of identifying candidate texts is exemplary and is not limited thereto.
[0062] For example, candidate texts (521, 522, 523) may be arranged in order of priority, but there is no limitation. For example, if only some of the candidate texts (521, 522, 523) are provided, candidate texts of interest may be provided in order of priority. For example, the priority may be set to be customized for the user based on usage history. For example, based on the fact that the history of changes to "smiling me" is relatively numerous, a relatively high priority may be given to the candidate text (521) of "smiling me." For example, the priority may be set based on the evaluation of the modification (or the performance of the artificial intelligence model). For example, in the case of changing "frowning me" to "smiling me," the expected evaluation score for the modification may be relatively high because the object within the face is not associated with other objects as it is modified. For example, if an object corresponding to "me" in an image is sitting, changing it to "me standing" requires not only a change in the appearance of the object corresponding to "me" but also boundary processing with other surrounding objects and / or modification of other objects, so the expected evaluation score may be relatively low. The electronic device (101) may also assign a relatively high priority if the expected evaluation score is relatively high. Meanwhile, those skilled in the art will understand that the above-described priority determination method is exemplary and there is no limitation. Meanwhile, those skilled in the art will understand that the provision of candidate texts based on priority is merely exemplary, and that candidate texts may be provided without being based on priority.
[0063] Referring to FIG. 5B, the electronic device (101) can verify and / or provide candidate texts (524, 525, 526) for “cloudy sky,” which is another object of a createable group (520). The priority setting of the candidate texts (524, 525, 526) has been described above, and thus will not be repeated here. The electronic device (101) can verify and / or provide candidate texts (531, 532) for a deleteable group (530). The electronic device (101) can verify and / or provide candidate texts (541, 542, 543) for an addable group (540). The candidate texts related to creation, deletion, and / or addition may be provided based on pre-specified information as described above, and / or may be verified based on a question-and-answer session with artificial intelligence, but there is no limitation on the verification method.
[0064] Referring back to FIG. 5A, the electronic device (101) may, in operation 507, confirm a selection of one or more candidate texts. In operation 509, the electronic device (101) may provide texts and images including the selected candidate text. The images may be, for example, modified images based on the selected candidate text. For example, based on the selection of candidate text (521) among the provided candidate texts (521, 522, 523) for the text "Frowning Me," an image in which the frowning face of the object corresponding to "Me" is modified to a smiling face may be provided. For example, based on the selection of candidate text (531), the electronic device (101) may provide a modified image in which people that failed to be recognized are deleted. For example, based on the selection of candidate text (541), the electronic device (101) may provide an image in which a puppy is added.
[0065] FIG. 6 is a flowchart illustrating an operating method of an electronic device according to one embodiment.
[0066] According to one embodiment, the electronic device (101) may provide an image in operation 601. The electronic device (101) may confirm a selection of a first object in the image in operation 603. For example, the electronic device (101) may confirm a selection of the first object based on a user's touch (or, without limitation, another type of gesture) on the first object in the image, but there is no limitation on the method of selection.
[0067] The electronic device (101) may, in operation 605, provide a first text for the selected first object. For example, based on confirmation of a user touch on the first object corresponding to "Frowning Me" in the image, the electronic device (101) may provide "Frowning Me" as the first text for describing the first object.
[0068] The electronic device (101) may, in operation 607, determine a user input that causes a change of the first text to a second text. In one example, the electronic device (101) may determine the change to the second text based on an input corresponding to the second text (e.g., input via SIP or input based on user voice, but is not limited thereto). In one example, the electronic device (101) may provide a plurality of candidate texts and determine whether any one of the plurality of candidate texts is selected as the second text.
[0069] The electronic device (101) can, in operation 609, change the first text to the second text based on the user input.
[0070] The electronic device (101) may, in operation 611, provide a modified image including a modified object based on the second text. In this case, those skilled in the art will understand that the representation of the second text may be omitted.
[0071] As described above, the electronic device (101) may be configured to modify an image based on a text provision and a corresponding text change command for a specific object selected by the user. For example, if an object that cannot be modified is selected, the electronic device (101) may refrain from providing text or provide text indicating that modification is not possible.
[0072] FIG. 7A is a flowchart illustrating an operating method of an electronic device according to one embodiment. The embodiment of FIG. 7A will be described with reference to FIG. 7B.
[0073] Figure 7b is a drawing for explaining object addition according to one embodiment.
[0074] According to one embodiment, the electronic device (101), in operation 701, may provide texts (713) describing an image (711) and one or more objects included in the image (711), as illustrated in FIG. 7B.
[0075] The electronic device (101), in operation 703, may identify a user input that causes the addition of a second text. For example, as illustrated in FIG. 7B , the electronic device (101) may provide an object (715) for the addition of the second text. The object (715) may include text (717) for the object that can be added. The electronic device (101) may identify an object (719) that causes the modification of the image. The object (715) may be an input field that allows the user to enter text (e.g., text (717) that reads “with pepper”), or may be a suggestion box that displays one or more suggested text options for the user to select. The object (719) may represent a button or icon that, when selected or clicked by the user, triggers an image generation function.
[0076] The electronic device (101) may, for example, confirm selection of an object (719) in which text (717) is expressed within an object (715) as a user input, and add a second text in operation 705. Accordingly, the electronic device (101) may provide texts (725) including existing texts (713) and the added second text. Those skilled in the art will understand that, depending on the implementation, the expression of the texts (725) may be omitted.
[0077] The electronic device (101) may, in operation 707, provide a modified image (721) including an object (723) added based on the second text.
[0078] For example, an object that can be added can be set based on a recognition target confirmed based on the analysis results of a plurality of images stored in association with the electronic device (101) or a user account. For example, an object that can be added can be a pet dog recognized as a result of the analysis of a plurality of images. The electronic device (101) can express that the pet dog can be added, for example, along with the recognition result (e.g., the name Pepper), and can add the object based on confirmation of a user command for addition. Meanwhile, the determination of an object to be added based on the analysis results of previously acquired images is merely exemplary. Those skilled in the art will understand that the electronic device (101) can also express an object other than an object confirmed based on previously acquired images as an object that can be added. For example, a person skilled in the art will understand that the electronic device (101) may express an object associated with a scene as an object that can be added based on the results of scene analysis of the image, and there is no limitation on the type, number, and / or confirmation method of the objects that can be added.
[0079] FIG. 8 is a drawing for explaining image modification by an electronic device according to one embodiment.
[0080] According to one embodiment, the electronic device (101) may provide an image (801). For example, the electronic device (101) may provide an image (801) based on executing a gallery application, or may provide an image (801) captured through a camera application, but there is no limitation on the providing event. The electronic device (101) may provide an object (802) that causes the provision of text to describe the image (801). The electronic device (101) may provide an object (803) that causes the generation of a modified image. The object (803) may represent a button or icon that triggers an image generation function when selected or clicked by a user.
[0081] Based on the confirmation of selection of the object (802), the electronic device (101) may provide texts (811) to describe the image (801). The texts (811) may include, but are not limited to, texts associated with objects included in the image (801), such as "sunset sky," "back to the beach," "with Min-ah," "walking," and "me." The method of confirming and / or providing the text has been described above, and thus the description thereof will not be repeated herein. For example, the electronic device (101) may confirm the selection of the text "sunset sky."
[0082] Based on the selection of the text "sunset sky", the electronic device (101) may provide candidate texts (813) for "sunset sky" as described above. The candidate texts (813) may include, but are not limited to, "blue sky", "redder sky", "night sky", and "aurora sky" to describe modifications that can replace "sunset sky". The method for confirming the candidate texts has been described above and will not be repeated here. For example, the electronic device (101) may confirm the selection of the candidate text "blue sky". Thereafter, the electronic device (101) may confirm the selection of the object (803) that causes the generation of the modified image.
[0083] The electronic device (101) may provide texts (813) containing candidate texts of "blue sky." The electronic device (101) may, for example, modify (or update) at least some of the existing texts (811) based on the selected candidate text. The electronic device (101) may provide a modified image (831) based on the selected candidate text. For example, the modified image (831) may be generated by changing an object corresponding to the selected text to an object corresponding to the candidate text. For example, the modified image (831) may be generated by changing at least some of the shape and / or properties of surrounding objects based on the effect of the object corresponding to the selected text on the surrounding objects. As described above, the change from "sunset sky" to "blue sky" results in an observable increase in the amount of light in the environment, which may be quantized as an enhancement in the luminance of the surrounding scene. This change may affect various objects, such as people, the ground, or other environmental elements within the scene. The electronic device (101) may generate a modified image (831) by applying an effect (e.g., an amount of light to surrounding objects such as people, the ground, etc.) corresponding to the increase in ambient light levels using an image processing algorithm. The term "effect" may refer to specific visual adjustments or image processing algorithms that the electronic device (101) applies to render the modified image (831) as if it were captured under new lighting conditions (e.g., a transition from a sunset sky to a blue sky). Such "effects" may include brightness adjustments, tone mapping, saturation or color enhancement, exposure compensation, shading or shadow effects, and similar modifications. For example, the brightness of surrounding objects may increase with increased light, but this is exemplary and there is no limitation on the type and / or application of the effect.The electronic device (101) may provide an object (804) that triggers the regeneration of the modified image. The object (804) may represent a button or icon that triggers an image regeneration function when selected or clicked by a user. If the electronic device (101) confirms the selection of another text and / or another candidate text and then confirms that the object (804) has been selected, the electronic device (101) may provide a modified image based on the newly selected text and / or candidate text. The electronic device (101) may provide an object (805) that triggers the completion of the modification. Based on the selection of the object (805) and / or the selection of an object that triggers additional storage, the modified image may be stored within the electronic device (101) or in a data storage (e.g., cloud storage, but not limited to) accessible based on a user account. Alternatively, an object (833) may be expressed to indicate the object to which the modification has been applied, but is not limited to this.
[0084] According to one embodiment, the electronic device (101) may visually distinguish between non-editable text and editable text. For example, among the texts (811), the electronic device (101) may further display circular objects around the texts "Sunset Sky," "At the Beach," "With Min-ah," and "Walking," as editable text. For example, among the texts (811), the electronic device (101) may not display circular objects around the texts "Backing Down" and "I," as non-editable text.
[0085] Meanwhile, those skilled in the art will understand that the representation of a circular object surrounding text is merely exemplary, and that there is no limitation to the method of distinguishing between editable and uneditable text. For example, the electronic device (101) may determine that "backward" is an uneditable text if there is no AI model for changing the object corresponding to "backward" and / or the AI model does not support changing the object. For example, the electronic device (101) may determine that the text is an uneditable text if the text is a designated text and / or a designated part of speech. For example, "I" may be designated as an uneditable text, and the electronic device (101) may accordingly determine that "I" is an uneditable text. Meanwhile, the method of determining whether text is an uneditable text as described above is exemplary, and there is no limitation to the method.
[0086] FIG. 9 is a drawing for explaining image modification by an electronic device according to one embodiment.
[0087] According to one embodiment, the electronic device (101) may provide a modified image (831) generated based on the text “sunset sky” being changed to “blue sky”, for example, as described with reference to FIG. 8. The electronic device (101) may provide texts (815) corresponding to the modified image (831). The texts (815) may include texts prior to modification (e.g., “back to the ground,” “at the beach,” “with Min-ah,” “walking,” and “me”) and modified text (e.g., “blue sky”). The electronic device (101) may confirm a selection of the modified text (e.g., “blue sky”), for example.
[0088] The electronic device (101) may provide candidate texts (821) that can replace "blue sky" based on the confirmation of the selection of the text "blue sky." For example, the candidate texts (821) may include "cloudless blue sky" and "sunny blue sky," but there is no limitation on the type and / or number thereof. For example, the electronic device (101) may provide "cloudless blue sky" and "sunny blue sky" as candidate texts (821) associated with "blue sky" based on the confirmation that a change from "sunset sky" to "blue sky" has been performed. For example, the "cloudless blue sky" of the candidate texts (821) may include the modified text "blue sky," but this is exemplary and not limiting. For example, the electronic device (101) may provide “cloudless blue sky” and “sunny blue sky” as candidate texts (821) associated with “blue sky” by giving a relatively high priority to candidate texts containing “blue sky” among the replaceable candidate texts, but there is no limitation on the method of providing them. Meanwhile, this is exemplary, and the electronic device (101) may also be configured to provide candidate texts unrelated to “blue sky” depending on the selection of “blue sky.”
[0089] For example, the electronic device (101) can confirm the selection of the candidate text of "cloudless blue sky". Based on the selection of the candidate text of "cloudless blue sky", the electronic device (101) can change the existing text of "blue sky" to the candidate text of "cloudless blue sky". Accordingly, texts (817) for describing an image including "cloudless blue sky" can be provided. The electronic device (101) can confirm the selection of the text of "cloudless blue sky". The electronic device (101) can confirm the selection of the object (803) after the selection of the text of "cloudless blue sky". Based on the confirmation of the selection of the object (803), the electronic device (101) can provide a modified image (833) reflecting the object corresponding to "cloudless blue sky". The electronic device (101) may, but is not limited to, represent an object (e.g., an object corresponding to a sky from which objects such as clouds have been removed) to represent a changed object on the modified image (833). The electronic device (101) may provide texts (817) to describe the image (833) including the changed text (e.g., “cloudless blue sky”). For example, based on a confirmation of selection for an object (805) that causes storage, the electronic device (101) may provide an object (806) that causes storage. Based on a confirmation of selection for an object (806), the modified image (833) may be stored in a storage accessible to the electronic device (101) and / or a user account.
[0090] FIG. 10 is a drawing for explaining image modification by an electronic device according to one embodiment.
[0091] According to one embodiment, the electronic device (101) may provide an image (1001). The electronic device (101) may provide texts (1005) to describe the image (1001). The texts (1005) may include, for example, “blue sky,” “below,” “crowded people,” “between,” “husband holding ice cream,” and “husband.” As described above, the texts (1005) may be provided based on the recognition results for the image (1001). If a plurality of objects are included in the image (1001), as described above, for example, only texts for some of the objects may be provided, and texts for the remaining objects may not be provided. However, the user may also intend to modify objects other than the provided texts (1005).
[0092] The electronic device (101) may identify a touch (1007) (or other gestures, such as a long press, flick, double click, etc.) on an object (1006) within the image (1001) as a user input that initiates or triggers a text-changing function. The user input may trigger activation of an editable mode for the texts (1005). Based on the identification of the user input causing the text change, the electronic device (101) may provide text for the object corresponding to the user input (e.g., “Next to the parasol”). For example, the image (1001) may initially be displayed without the texts (1005). When the user input (e.g., the touch (1007)) activates the text-editable mode, the texts (1005) are presented together with the image (1001). The electronic device (101) may display the texts (1005) such that editable texts (e.g., "ice cream," "crowded people," and "blue sky") are visually distinct from non-editable texts (e.g., "husband holding," "between," and "under") within the texts (1005). In the embodiment of FIG. 10, the electronic device (101) is illustrated as having text for an object corresponding to a user input (e.g., "beside a parasol") replace the existing text "under a blue sky," but this is exemplary.
[0093] Those skilled in the art will appreciate that the electronic device (101) may be implemented to add text (e.g., "next to the parasol") for an object corresponding to a user input while maintaining the existing text. For example, the electronic device (101) may replace "under the blue sky", which is identified as meaning "place," with "next to the parasol", which is specified by the user, but this is exemplary. The electronic device (101) may provide texts (1015) to describe an image (1001) containing "next to the parasol", as described above. The electronic device (101) may confirm a selection of the text "parasol," for example. Based on the confirmation of the selection of the text "parasol," the electronic device (101) may provide candidate texts (1017). For example, the electronic device (101) may confirm a selection of the candidate text "tree" among the candidate texts (1017). After confirming the selection of the candidate text of "tree", the electronic device (101) can confirm the selection of the object (1003) that causes the generation of the modified image. Based on the confirmation of the selection of the object (1003), the electronic device (101) can provide a modified image (1021) by changing the object (1006) corresponding to the text of "parasol" to the object (1018) corresponding to the candidate text of "tree". The electronic device (101) can provide texts (1023) to describe the modified image (1021). The texts (1023) can include the selected candidate text of "tree". The electronic device (101) can also provide, but is not limited to, an object (1025) that causes regeneration and / or an object (1027) that causes the completion of modification.
[0094] FIG. 11 is a drawing for explaining image modification by an electronic device according to one embodiment.
[0095] According to one embodiment, the electronic device (101) may provide an image (1101). The electronic device (101) may provide texts (1105) to describe the image (1101). The texts (1105) may include, for example, “sunset sky,” “with my back to the camera,” “in the grass,” and “me sitting” to describe objects included in the image (1101). The electronic device (101) may provide an object (1103) that causes the creation of a modified image. The electronic device (101) may provide an object (1106) that causes the addition of an object and / or text, for example. The object (1106) may represent a button or icon (e.g., a plus symbol) that enables the addition of new text within the texts (1105). Based on the confirmation of the selection for the object (1106), the electronic device (101) can provide candidate texts (1107) for objects that can be added to the image (1101). For example, the electronic device (101) can provide objects that can be added to the image (1101) based on recognition results from images stored in a storage accessible to the electronic device (101) and / or a user account, but is not limited thereto. The electronic device (101) can provide candidate texts based on the recognition results.
[0096] The electronic device (101) may provide candidate texts based on at least some of the recognition results, for example, based on scene analysis of the image (1101). For example, the electronic device (101) may, as at least some of the scene analysis, determine that the recognition result of the person in the image (1101) is "me." The electronic device (101) may, based on the analysis results of the stored images, provide priorities for each of the recognition results based on the number of times they are recognized together with "me." For example, "pepper," "me," "myunghwa," "myunghwa," and "husband" may be used to construct candidate texts based on the fact that the number of times a pet dog recognized as "pepper," a person recognized as "so-un," a person recognized as "kyunghwa," and a person recognized as "husband" are recognized together with "me" is greater than the number of times corresponding to other recognition results. Meanwhile, prioritizing based on the number of times images are recognized together is merely exemplary, and there are no limitations on how priorities are determined. For example, the electronic device (101) may also prioritize based on the date the images were captured.
[0097] The electronic device (101) can identify the modifiers "sitting," "standing," "smiling," and "sitting" for "pepper," "small," "hardening," and "husband" based on a new analysis of pre-stored images. For example, the pre-stored image may include "sitting pepper." The electronic device (101) can recognize "sitting pepper" from the pre-stored image. The pre-stored image may be, for example, an image in which "I" is recognized, but is not limited thereto. For example, the electronic device (101) can identify the modifier "standing" corresponding to "small," identify the modifier "smiling" corresponding to "hardening," and / or identify the modifier "sitting" corresponding to "husband" based on a new analysis of the pre-stored images. In this case, the modifiers associated with the candidate text and / or the objects generated based on the candidate text may depend on pre-stored images.
[0098] The electronic device (101) can identify the modifiers "sitting," "standing," "smiling," and "sitting" for "pepper," "snow," "hardening," and "husband" based on the analysis of the image (1101) to be modified. For example, the electronic device (101) can identify the position, size, and / or posture of a person in the image (1101). For example, the electronic device (101) can identify that the person in the image (1101) is in a sitting posture. Based on the fact that the posture of the person in the image (1101) is sitting, the electronic device (101) can set the modifier of "pepper" to "sitting." In this case, the modifier associated with the candidate text and / or the object generated based on the candidate text may not depend on a pre-stored image, but on the image (1101) to be modified.
[0099] The electronic device (101) can confirm the selection of "sitting pepper" among the candidate texts (1017). Based on the confirmation of the selection of "sitting pepper," the electronic device (101) can provide a modified image (1121) to which an object (1141) corresponding to "sitting pepper" has been added. The object (1141) corresponding to "sitting pepper" can be acquired and / or generated, for example, based on a pre-stored image, but is not limited thereto.
[0100] FIG. 12 is a drawing for explaining image modification by an electronic device according to one embodiment.
[0101] According to one embodiment, the electronic device (101) may provide an image (1201). The electronic device (101) may provide texts (1205) to describe the image (1201). The texts (1205) may include, for example, “blue sky,” “below,” “in front of people,” “desert,” “in,” “standing,” and “me” to describe objects included in the image (1201). For example, “in front of people” may be text corresponding to multiple people included in the image (1201), but is not limited thereto. The electronic device (101) may provide an object (1203) that causes modification of the image. The electronic device (101) may confirm a selection of the text “in front of people.”
[0102] The electronic device (101) may provide a plurality of candidate texts (1207) in response to a selection of the text "in front of people." The plurality of candidate texts (1207) may include "remove people" and "blur people's faces." In the present embodiment, the candidate texts may include text for deleting and / or processing an object for the selected text, rather than text that replaces the text "in front of people." For example, the electronic device (101) may provide processing related to the background, such as deletion and / or blurring, based on whether the object (1241) corresponding to "in front of people" is included in the background or is smaller than a specified size, but there is no limitation thereto. For example, the electronic device (101) may determine that the object (1241) corresponding to "in front of people" is included in the background based on whether the object (1241) corresponding to "in front of people" is identified in the background recognition process, but there is no limitation on the method of such identification.
[0103] In the present embodiment, the electronic device (101) has been described as providing candidate text for deletion and / or processing of an object corresponding to the selected text based on confirmation of selection of the text, but this is exemplary. The electronic device (101) may also be configured to provide candidate text for deletion and / or processing of an object based on confirmation of additional user input for deletion and / or processing after a specific text has been selected.
[0104] The electronic device (101) can confirm the selection of the candidate text "Remove People" among the candidate texts (1207). The electronic device (101) can confirm the selection of the object (1203) after the selection of the candidate text "Remove People". The electronic device (101) can perform object deletion as a modification corresponding to the candidate text based on the confirmation of the selection of the candidate text and / or the object (1203). The electronic device (101) can provide a modified image (1221) generated by deleting the object (1241) included in the existing image (1201). The electronic device (101) can perform inpainting, which deletes the object (1241) included in the image (1201) and draws the deleted portion to correspond to the surrounding background. For example, GAN, CNN, DeepFill, EdgeConnect, etc. can be used, but there is no limitation.
[0105] The electronic device (101) may provide texts (1222) to describe the image (1221). For example, text (1227) corresponding to a deleted object may be indicated using a strikethrough (also referred to as a "delete line" or "delete indicator"). However, it should be understood that this is only one possible representation, and alternative methods of indicating a deleted object may be used, or no specific indication may be provided at all. The electronic device (101) may also provide an object (1223) for regeneration and / or an object (1225) for causing completion.
[0106] FIG. 13A is a diagram for explaining image modification by an electronic device according to one embodiment.
[0107] According to one embodiment, the electronic device (101) may perform modifications to the image (1301). As described above, the electronic device (101) may provide texts to describe the image (1301). For example, the electronic device (101) may provide texts such as “under the blue sky,” “located by the riverside,” and “buildings” as texts to describe the image (1301). The electronic device (101) may identify a user input that causes the text “under the blue sky” to be changed to “under the sunset sky,” for example. Based on the user input, the electronic device (101) may provide a modified image (1311). The modified image (1311) may include, for example, an object (1311) corresponding to the changed text “under the sunset sky,” i.e., a sunset sky.
[0108] Meanwhile, as described above, the electronic device (101) may apply an effect indicating an influence based on a change to the object (1311) along with a change to the object (1311). For example, the color of the object (1303) corresponding to "riverside" in the image (1301) before modification may be different from the color of the object (1312) corresponding to "riverside" in the modified image (1311). For example, an effect corresponding to an influence due to a change to a specific object may be applied based on CycleGAN, Pix2Pix, NST (neural style transfer), etc., but there is no limitation.
[0109] FIG. 13b is a diagram for explaining image modification by an electronic device according to one embodiment.
[0110] According to one embodiment, the electronic device (101) can perform modification on the image (1341). As described above, the electronic device (101) can provide texts to describe the image (1341). For example, the electronic device (101) can provide texts such as “under the blue sky,” “walking down the street,” “with people,” and “me” as texts to describe the image (1341). The electronic device (101) can confirm a user input to remove an object (1342) corresponding to the text “with people,” for example. Based on the user input, the electronic device (101) can provide a modified image (1351). The modified image (1351) can be generated by, for example, deleting an object (1342) corresponding to “with people” selected as a deletion target.
[0111] Meanwhile, the electronic device (101) can delete the object (1343), i.e., the shadow, associated with the object (1342) corresponding to "people". For example, the electronic device (101) can be configured to delete the object (1343) together with the deletion of the object (1342) based on the relationship between the object (1343) and the object (1342). For this purpose, inpainting can be performed, and examples thereof, such as GAN, CNN, DeepFill, EdgeConnect, etc., can be used, but are not limited thereto. As described with reference to FIGS. 13a and 13b, a modified image can be generated by applying not only a change to an object designated by text, but also an effect indicating the influence of the change.
[0112] FIG. 14 is a drawing for explaining image modification by an electronic device according to one embodiment.
[0113] According to one embodiment, the electronic device (101) may provide an image (1401). The electronic device (101) may provide texts (1411) to describe the image (1401). The texts (1411) may include, for example, “each other,” “looking at,” “smiling,” “yellow hood,” “wearing,” “woman with,” “brown hat,” “wearing,” and “man” to describe objects included in the image (1401). The electronic device (101) may confirm a selection of, for example, “brown hat.” In response to a selection of “brown hat,” the electronic device (101) may provide candidate texts (1412). The candidate texts (1412) may include, for example, “black hat,” “Santa Claus hat,” “brown beanie,” and “blue swim cap.” The electronic device (101) may, for example, confirm a selection for the candidate text of "Santa Claus hat." Based on the confirmation of the selection for the candidate text of "Santa Claus hat," the electronic device (101) may provide a modified image (1403) including an object corresponding to the candidate text of "Santa Claus hat." The electronic device (101) may provide texts (1413) for describing the modified image (1403). The texts (1413) for describing the modified image (1403) may include "each other," "looking at," "smiling," "wearing a yellow hood," "wearing," "woman," "wearing a Santa Claus hat," "wearing," and "man."
[0114] The electronic device (101) may then additionally confirm the selection of the text "yellow hood." Based on the confirmation of the selection of the text "yellow hood," the electronic device (101) may provide candidate texts (1415). Based on the image modification associated with the "Santa Claus hat," the electronic device (101) may also provide candidate texts (1415) based on the previously modified information. For example, the electronic device (101) may provide "Santa suit," "red suit," "Christmas dress," and "Rudolph suit," which are semantically associated with the "Santa Claus hat," as candidate texts (1415). For example, candidate texts (1415) such as “Santa suit”, “red suit”, “Christmas dress”, and “Rudolph suit” can be confirmed based on both the modified text “Santa Claus hat” and the selected text “yellow hood”, but there is no limitation. If no modification related to “Santa Claus hat” is performed, the electronic device (101) may provide candidate texts unrelated to “Santa suit”, “red suit”, “Christmas dress”, and “Rudolph suit”. For example, the candidate text “red suit” among the candidate texts (1415) can be selected. Based on the confirmation of the selection of the candidate text “red suit”, the electronic device (101) can provide a modified image (1405) including an object corresponding to the candidate text “red suit”. The electronic device (101) may provide texts (1417) to describe the modified image (1405). The texts (1417) may include "each other", "looking at", "smiling", "in red clothes", "wearing", "woman with", "wearing a Santa Claus hat", "wearing", and "man".
[0115] As described above, the electronic device (101) may also set candidate texts for other texts based on previously modified information related to a specific text.
[0116] FIG. 15A is a diagram for explaining image modification by an electronic device according to various embodiments.
[0117] FIG. 15b is a diagram for explaining image modification by an electronic device according to various embodiments.
[0118] According to one embodiment, referring to FIG. 15A, an electronic device (101) may provide an image (1501). The electronic device (101) may provide texts (1503) to describe the image (1501). The texts (1503) may include, for example, “brown hair,” “Jay’s,” and “selfie.” The electronic device (101) may confirm a selection for, for example, “selfie.” Based on confirming the selection for “selfie,” the electronic device (101) may provide candidate texts (1505).
[0119] The texts (1503) and / or candidate texts (1505) may be set based on, for example, a user's editing history. For example, based on a user's history of editing related to "face," the texts (1503) may include "brown hair," "Jay's," "selfie," and / or the candidate texts (1505) may include "big-eyed selfie," "sad selfie," and "tearful selfie." For example, "big-eyed selfie," "sad selfie," and "tearful selfie" may be set based on editing operations previously performed by the user, but are not limited thereto. For example, the candidate texts (1505) may include "edited selfie." For example, if "Edited Selfie" is selected, the electronic device (101) can perform a modification saved by the user (e.g., a modification that relatively brightens the skin color of the facial area, but is not limited thereto). For example, the user can manually perform a modification on the image and save it as a modification set by the user.
[0120] The electronic device (101) can store information related to the modification (e.g., the degree of brightness value adjustment for a face portion) as a modification set by the user. The electronic device (101) can perform image modification (e.g., modification based on the degree of brightness value adjustment for a face portion) based on the stored information, based on a selection of candidate text corresponding to the modification set by the user, such as "edited selfie."
[0121] According to one embodiment, referring to FIG. 15B, the electronic device (101) may provide an image (1501). The electronic device (101) may provide texts (1513) to describe the image (1501). The texts (1513) may include, for example, “sweater,” “wearing,” “Jay’s,” and “selfie.”
[0122] As described above, the electronic device (101) can set texts (1513) and / or candidate texts (1515) based on the modification history of the user. For example, the electronic device (101) can provide texts (1513) different from the texts (1503) of FIG. 15A based on the history of modifications the user has made related to "clothes." The electronic device (101) can confirm a selection of, for example, "selfie." Based on the confirmation of the selection of "selfie," the electronic device (101) can provide candidate texts (1515). The candidate texts (1515) can be set based on, for example, the modification history of the user. For example, based on the history of modifications the user has made related to "clothes," the candidate texts (1515) can include "sleeveless," "hoodie," and "Santa suit." For example, "sleeveless," "hoodie," and "Santa outfit" can be set based on modifications already performed by the user, but there is no limitation. The electronic device (101) can modify the image (1501) based on a selected candidate text among the candidate texts (1515).
[0123] In one or more embodiments of the present disclosure, an electronic device (101) may include: a display; one or more processors (120); and a memory storing instructions. The instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to: provide an image and texts describing the image; change first text included in the texts to second text based on a user input; and provide a modified image in which an object is created, removed, or modified based on the second text.
[0124] The above instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to: provide a user interface for receiving the second text, and verify the second text based on the user input entered through the user interface.
[0125] The above instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to: verify the user input specifying the second text input through a virtual input panel for inputting a plurality of characters.
[0126] The instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to: provide at least one candidate text corresponding to the first text, and to confirm a user input indicating a selection of the second text from among the at least one candidate text.
[0127] The at least one candidate text may be set based on the priority of each of the plurality of candidate texts corresponding to the first text.
[0128] The above at least one candidate text may be set based on the user's image modification history.
[0129] The instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to: provide the modified image including a second object as the modified object corresponding to the second text, replacing the first object corresponding to the first text.
[0130] The instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to: provide the modified image including the modified object by changing a first attribute of the object corresponding to the first text to a second attribute corresponding to the second text.
[0131] The instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to: provide the modified image including the modified object by applying a visual effect to a surrounding object of the first object corresponding to the first text based on the second text.
[0132] The instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to: identify the user input causing the addition of third text to the texts; display, based on the user input, a modified version of the texts including the third text; and provide the modified image including the generated object corresponding to the third text.
[0133] The instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to: identify a user input causing deletion of a fourth text included in the texts; based on the user input, exclude the fourth text or display a modified version of the texts including a deletion indicator applied to the fourth text; and provide the modified image in which an object corresponding to the fourth text has been deleted.
[0134] The instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to: provide at least one candidate text associated with the second text based on determining a selection of the modified object corresponding to the second text among the objects included in the modified image; and provide an additional modified image including an additional modified object corresponding to the selected candidate text based on determining a selection of one of the at least one candidate text associated with the second text. At least a portion of the at least one candidate text associated with the second text may include at least a portion of the second text.
[0135] The instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to: provide at least one candidate text associated with the second text and the sixth text based on a determination of a selection of an object corresponding to a sixth text different from the second text; and provide an additional modified image including an additional modified object corresponding to the selected candidate text based on a determination of a selection of one of the at least one candidate text associated with the second text and the sixth text.
[0136] The above instructions, when individually or collectively executed by the one or more processors (120), may cause the electronic device (101) to: display editable portions of the text so as to be visually distinct from non-editable portions of the text.
[0137] In one or more embodiments of the present disclosure, a non-transitory computer-readable storage medium is provided storing one or more instructions comprising computer-executable instructions, which when individually or collectively executed by one or more processors (120) of an electronic device (101) cause the electronic device (101) to: provide an image and texts describing the image; change a first text included in the texts to a second text based on a user input; and provide a modified image in which an object is created, removed, or modified based on the second text.
[0138] The one or more instructions may cause the electronic device (101) to perform the following actions: providing a user interface for receiving the second text; and confirming the second text based on the user input entered through the user interface.
[0139] The one or more instructions may cause the electronic device (101) to perform the following actions: providing at least one candidate text corresponding to the first text; and confirming the user input indicating a selection of the second text from among the at least one candidate text.
[0140] The action of providing the modified image may include: providing the modified image including a second object as the modified object corresponding to the second text, replacing the first object corresponding to the first text.
[0141] The operation of providing the modified image may include: providing the modified image including the modified object by changing a first property of the object corresponding to the first text to a second property corresponding to the second text.
[0142] In one or more embodiments of the present disclosure, a method of operating an electronic device (101) may include: providing an image and texts describing the image; changing first text included in the texts to second text based on a user input; and providing a modified image in which an object is created, removed, or modified based on the second text.
[0143] The operation of providing the texts describing the image and one or more objects included in the image according to the recognition result of the image may include an operation of expressing a changeable first part and an unchangeable second part among the texts so as to be distinguished.
[0144] Electronic devices according to the embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments disclosed in this document are not limited to the aforementioned devices.
[0145] The embodiments of this document and the terminology used herein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0146] The term "module" used in the embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0147] One embodiment of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0148] According to one embodiment, the method according to one embodiment disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0149] According to one embodiment, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and arranged in other components. According to one embodiment, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to one embodiment, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In an electronic device (101), display; One or more processors (120); Includes a memory (130) that stores instructions, The above instructions, when individually or collectively executed by one or more processors (120), cause the electronic device (101) to: Provide images and text describing said images, Based on user input, change the first text contained in the above texts to the second text, and An electronic device (101) that causes an object to provide a modified image, wherein the image is generated, removed, or modified based on the second text.
2. In paragraph 1, The above instructions, when individually or collectively executed by one or more processors (120), cause the electronic device (101) to: Providing a user interface for inputting the second text, and An electronic device (101) that causes the second text to be verified based on the user input entered through the user interface.
3. In paragraph 1 or 2, The above instructions, when individually or collectively executed by one or more processors (120), cause the electronic device (101) to: An electronic device (101) that causes the user to confirm the second text input through a virtual input panel for receiving multiple characters.
4. In any one of paragraphs 1 to 3, The above instructions, when individually or collectively executed by one or more processors (120), cause the electronic device (101) to: providing at least one candidate text corresponding to the first text, and An electronic device (101) that causes the user to confirm a selection of the second text among the at least one candidate text.
5. In any one of paragraphs 1 to 4, An electronic device (101) wherein the at least one candidate text is set based on the priority of each of a plurality of candidate texts corresponding to the first text.
6. In any one of paragraphs 1 to 5, An electronic device (101) wherein at least one candidate text is set based on the user's image modification history.
7. In any one of paragraphs 1 to 6, The above instructions, when individually or collectively executed by one or more processors (120), cause the electronic device (101) to: An electronic device (101) that causes a modified image to be provided, the modified image including a second object as the modified object corresponding to the second text, by replacing the first object corresponding to the first text.
8. In any one of paragraphs 1 to 7, The above instructions, when individually or collectively executed by one or more processors (120), cause the electronic device (101) to: An electronic device (101) that causes a modified image including the modified object to be provided by changing a first attribute of an object corresponding to the first text to a second attribute corresponding to the second text.
9. In any one of paragraphs 1 to 8, The above instructions, when individually or collectively executed by one or more processors (120), cause the electronic device (101) to: An electronic device (101) that causes a modified image including the modified object to be provided by applying a visual influence to surrounding objects of the first object corresponding to the first text based on the second text.
10. In any one of paragraphs 1 to 9, The above instructions, when individually or collectively executed by the one or more processors (120), cause the electronic device (101) to: identify the user input causing the addition of a third text to the texts; Based on the user input, display a modified version of the texts including the third text, and An electronic device (101) causing the modified image to be provided, the modified image including the generated object corresponding to the third text.
11. In any one of paragraphs 1 to 10, The above instructions, when individually or collectively executed by one or more processors (120), cause the electronic device (101) to: Identify the user input that causes the deletion of the fourth text contained in the above texts, Based on said user input, excluding said fourth text or displaying a modified version of said text that includes a deletion indicator applied to said fourth text, and An electronic device (101) causing the object corresponding to the fourth text to be provided with the modified image deleted.
12. In any one of paragraphs 1 to 11, The above instructions, when individually or collectively executed by one or more processors (120), cause the electronic device (101) to: Providing at least one candidate text associated with the second text based on confirming a selection of the modified object corresponding to the second text among the objects included in the modified image, and Based on the confirmation of a selection of one of the at least one candidate text associated with the second text, causing an additional modified image to be provided that includes an additional modified object corresponding to the selected candidate text; An electronic device (101) wherein at least a portion of said at least one candidate text associated with said second text comprises at least a portion of said second text.
13. In any one of paragraphs 1 to 12, The above instructions, when individually or collectively executed by one or more processors (120), cause the electronic device (101) to: Based on the selection of an object corresponding to a sixth text different from the second text, providing at least one candidate text associated with the second text and the sixth text, and An electronic device (101) that causes a user to provide an additional modified image including an additional modified object corresponding to the selected candidate text based on a selection of at least one candidate text associated with the second text and the sixth text.
14. In any one of paragraphs 1 to 13, The above instructions, when individually or collectively executed by one or more processors (120), cause the electronic device (101) to: An electronic device (101) that causes the editable portions of the above texts to be displayed so as to be visually distinguished from the non-editable portions of the above texts.
15. A non-transitory computer-readable storage medium storing one or more instructions comprising instructions executable by a computer, The above instructions, when individually or collectively executed by one or more processors (120) of the electronic device (101), cause the electronic device (101) to: An action that provides an image and text describing said image; An action of changing a first text included in the above texts into a second text based on user input; and A storage medium that causes an object to perform an action that provides a modified image, wherein the object is created, removed, or modified based on the second text.
Citation Information
Patent Citations
Noise determination apparatus and method for actuator apparatus
KR1020210087908A
Operation part for surgical instrument and surgical instrument for electrocautery equipped with the operation part
KR1020250157048A
Dual powernet system for vehicle and operating method thereof
KR102275136B1
Apparatus and method for correcting sentence using test and image
KR102388599B1
Image generating device for generating 3D images corresponding to user input sentences and operation method thereof
KR102597074B1