Electronic device, method, and non-transitory computer-readable storage medium for synthesizing visual object with image

By using user input strokes to specify visual object locations within images, the electronic device improves the accuracy and convenience of integrating visual elements into images through an image generation model, addressing the limitations of natural language-based prompts.

WO2026059068A1PCT designated stage Publication Date: 2026-03-19SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing systems face difficulties in accurately specifying the location of visual objects within images using natural language-based prompts, making it cumbersome to composite or outpaint visual elements effectively.

Method used

An electronic device employs user input in the form of strokes to specify the location of visual objects, utilizing an image generation model to composite or outpaint these objects onto expanded images, allowing for more precise placement and integration of visual elements.

Benefits of technology

This approach enables more accurate and convenient integration of visual objects by allowing users to specify locations through direct input, enhancing the precision and efficiency of image editing processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025010227_19032026_PF_FP_ABST
    Figure KR2025010227_19032026_PF_FP_ABST
Patent Text Reader

Abstract

This electronic device may comprise: a memory storing instructions; a display; and at least one processor. The instructions may cause the electronic device to: display an original image via the display; receive a user input for adding one or more strokes on a peripheral area with respect to an area of the display on which the original image is displayed; and acquire, on the basis of at least receiving the user input, an edited image including a first portion at least partially corresponding to the original image and a second portion corresponding to the peripheral area and synthesized with a visual object represented by one or more strokes.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device, method, and non-transient computer-readable storage medium for synthesizing a visual object with an image

[0001] The present disclosure relates to an electronic device, a method, and a non-transient computer-readable storage medium for synthesizing a visual object with an image.

[0002] The electronic device may include a trained model. The electronic device may receive user input representing a natural language-based prompt. By providing the received prompt to the trained model, the electronic device may acquire an image of the prompt. The electronic device may display the acquired image through a display.

[0003] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure.

[0004] No claim or determination is made as to whether any of the foregoing can be applied as prior art related to the present disclosure.

[0005] An electronic device is described. The electronic device may include a memory comprising one or more storage media for storing instructions. The electronic device may include one or more sensors. The electronic device may include at least one processor comprising a processing circuit. The electronic device may include a display. The instructions may cause the electronic device to display an original image through the display when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to receive user input to add one or more strokes to a surrounding area of ​​the display where the original image is displayed when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to obtain an edited image, comprising, at least based on receiving the user input, a first part corresponding at least partially to the original image, and a second part corresponding to the surrounding area and composited with a visual object represented by the one or more strokes.

[0006] A method is provided. The method may be executed within an electronic device having a display. The method may include an operation of displaying an original image through the display. The method may include an operation of receiving user input for adding one or more strokes to a surrounding area of ​​the display where the original image is displayed. The method may include, at least based on receiving the user input, an operation of acquiring an edited image comprising a first portion that corresponds at least partially to the original image, and a second portion that corresponds to the surrounding area and is composited with a visual object represented by the one or more strokes.

[0007] A non-transient computer-readable storage medium is provided. The non-transient computer-readable storage medium may store one or more programs. The one or more programs may include instructions that cause the electronic device to display an original image through the display when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to receive user input to add one or more strokes to a surrounding area of ​​the display where the original image is displayed when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain an edited image, comprising, at least based on receiving the user input, a first portion corresponding at least partially to the original image, and a second portion corresponding to the surrounding area and composited with a visual object represented by the one or more strokes.

[0008] An electronic device is described. The electronic device may include a memory comprising one or more storage media for storing instructions. The electronic device may include one or more sensors. The electronic device may include at least one processor comprising a processing circuit. The electronic device may include a display. The instructions may cause the electronic device to display an original image through a first area of ​​the display when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to obtain a first edited image representing the original image, based on a first user input for adding one or more first strokes within the first area when executed individually or collectively by the at least one processor, by synthesizing a first visual object represented by the one or more first strokes. When the above instructions are executed individually or collectively by the at least one processor, based on a second user input for adding one or more second strokes within a second region adjacent to the first region, the second visual object represented by the one or more second strokes is synthesized, and the electronic device may obtain a second edited image including the original image.

[0009] A method is provided. The method may be executed within an electronic device having a display. The method may include an operation of displaying an original image through a first area of ​​the display. The method may include an operation of synthesizing a first visual object represented by one or more first strokes based on a first user input for adding one or more first strokes within the first area, and acquiring a first edited image representing the original image. The method may include an operation of synthesizing a second visual object represented by one or more second strokes based on a second user input for adding one or more second strokes within a second area adjacent to the first area, and acquiring a second edited image including the original image.

[0010] A non-transient computer-readable storage medium is provided. The non-transient computer-readable storage medium may store one or more programs. The one or more programs may include instructions that cause the electronic device to display an original image through a first area of ​​the display when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain a first edited image representing the original image, based on a first user input to add one or more first strokes within the first area when executed by the electronic device, and a first visual object represented by the one or more first strokes is composited. The one or more programs may include instructions that cause the electronic device to obtain a second edited image including the original image, based on a second user input to add one or more second strokes within a second area adjacent to the first area when executed by the electronic device, and a second visual object represented by the one or more second strokes is composited.

[0011] Figure 1 illustrates an example of an electronic device that acquires an image using a prompt.

[0012] Figure 2 is a simplified block diagram of an exemplary electronic device.

[0013] Figure 3 is a flowchart illustrating the operation of an electronic device for acquiring an edited image.

[0014] FIG. 4 illustrates an exemplary operation of an electronic device that acquires an edited image based on user input adding one or more strokes.

[0015] FIG. 5 illustrates an exemplary operation of an electronic device receiving user input to reduce the size of an original image.

[0016] FIG. 6a illustrates an exemplary operation of an electronic device receiving user input in which the size of an original image is changed.

[0017] FIGS. 6b through 6d illustrate exemplary operation of an electronic device that acquires another edited image based on an edited image.

[0018] FIG. 7 illustrates an exemplary operation of an electronic device receiving data from an AI (artificial intelligence) server.

[0019] FIG. 8 illustrates an exemplary operation of an electronic device that acquires an edited image based on time identified from a software application for time.

[0020] FIGS. 9 and 10 illustrate exemplary operation of an electronic device that displays a screen when an edited image is acquired.

[0021] FIG. 11 illustrates an exemplary operation of an electronic device receiving other user input to avoid synthesis.

[0022] FIG. 12 illustrates an exemplary operation of an electronic device receiving user input for adding one or more strokes expressed as text.

[0023] FIG. 13 illustrates an exemplary operation of an electronic device receiving other user input to determine at least a portion of an original image.

[0024] FIG. 14 is a flowchart illustrating the operation of an electronic device for acquiring a first edited image or a second edited image.

[0025] FIG. 15 is a block diagram of an electronic device in a network environment according to various embodiments.

[0026] Figure 16 is a schematic diagram of an exemplary AI system.

[0027] Figure 1 illustrates an example of an electronic device that acquires an image using a prompt.

[0028] Referring to FIG. 1, an electronic device (100) may be used to generate or acquire an image. For example, the electronic device (100) may acquire an image using an image generation model (e.g., the image generation model (714) of FIG. 7). For example, a state (110) may be described as a state in which an image is acquired using an image (112) and a prompt (114). For example, a state (120) may be described as a state in which an image (122) generated based on the image (112) and the prompt (114) is displayed. For example, when the electronic device (100) displays the image (112) through a display (e.g., the display (208) of FIG. 2), it may receive user input indicating a prompt (114) instructing to generate the image (122). For example, the electronic device (100) can obtain an image (122) from an image (112) by providing a prompt (114) and an image (112) to the image generation model. For example, the electronic device (100) can display the obtained image (122) through the display.

[0029] The electronic device (100) may receive user input indicating a prompt (114) for changing the composition of the image (112). For example, the user of the electronic device (100) may provide user input indicating a prompt (114) that causes a visual object to be included. For example, the electronic device (100) may receive or obtain a prompt (114) that causes a visual object to be included when performing inpainting and / or outpainting using the image (112). For example, the inpainting may be described as a technique for restoring damaged parts or defects of the image using a trained model. For example, the outpainting may be described as a technique for obtaining an expanded image by creating a new image outside the boundaries of the image using a trained model.

[0030] The electronic device (100) can use the image generation model to composite a visual object indicated by a prompt (114) onto an expanded image on which outpainting has been performed on the image (112). For example, the electronic device (100) can use the image generation model to display an image (122) containing a visual object indicated by a prompt (114) through the display. For example, the electronic device (100) can receive user input for a prompt (114) indicating a visual object to be added to the image (122). For example, the electronic device (100) can identify or determine a visual object to be added to the image (122) from the prompt (114) based on receiving the user input. For example, the electronic device (100) can obtain the visual object from the prompt (114) using the image generation model or a multimodal model (e.g., the multimodal model (712) of FIG. 7).

[0031] The prompt (114) may be based on natural language. For example, the prompt (114) may include text based on natural language. For example, an electronic device (100) may receive a natural language-based prompt (114) indicating the location of the visual object. For example, the electronic device (100) may add or compose the visual object at the location identified by the prompt (114). For example, the electronic device (100) may display an image (122) containing the visual object located in the area specified by the prompt (114) through the display.

[0032] For example, the electronic device (100) may determine the location of the image (122) where the visual object is to be displayed based on the prompt (114). For example, the prompt (114) may include text indicating to add the visual object at specific coordinates. For example, the prompt (114) may include text indicating to add the visual object at a location of a first pixel from the left and a second pixel from the bottom. For example, the prompt (114) may include text indicating to add the visual object at a specific direction or a specific distance from a visual element (e.g., a house or a tree) included in the image (112).

[0033] It may be difficult to specify the location of a visual object to be included in an image (122) using a natural language-based prompt (114). The electronic device (100) may be required to receive user input to specify the location of the visual object. For example, the user input may include user input to add one or more strokes. For example, based on receiving user input to add one or more strokes, the electronic device (100) may acquire or generate an image (122) containing a visual object represented by the one or more strokes in a second part of the image (122) corresponding to a first part where the user input was received.

[0034] The electronic device (100) can determine the portion of the image (122) where the visual object is to be displayed by receiving user input for adding the one or more strokes. For example, specifying the location where the visual object is to be displayed by the prompt (114) may be more difficult than specifying the location where the visual object is to be displayed by user input for adding the one or more strokes. For example, specifying the location of the visual object by the user input may be more convenient than specifying the location of the visual object using the prompt (114).

[0035] For example, the electronic device (100) may include hardware components used to perform or execute the above operations. The hardware components are described and illustrated with reference to FIG. 2.

[0036] Figure 2 is a simplified block diagram of an exemplary electronic device.

[0037] Referring to FIG. 2, the electronic device (100) may include at least one processor (207), memory (206), and display (208). The electronic device (100) may include a communication circuit (205). However, it is not limited thereto. The communication circuit (205) may be an optional component.

[0038] At least one processor (207) may include a hardware component for processing data using instructions stored in memory (206). The hardware component for processing data may include a CPU (central processing unit) (e.g., including processing circuits). The hardware component for processing data may include a GPU (graphic processing unit) (e.g., including processing circuits). The hardware component for processing data may include a DPU (display processing unit) (e.g., including processing circuits). The hardware component for processing data may include a NPU (neural processing unit) (e.g., including processing circuits).

[0039] At least one processor (207) may include one or more cores. For example, at least one processor (207) may have the structure of a multi-core processor such as a dual core, a quad core, or a hexa core.

[0040] Memory (206) may include a hardware component for storing data and / or instructions that are input to and / or output from at least one processor (207). Memory (206) may include, for example, volatile memory such as random-access memory (RAM) and / or non-volatile memory such as read-only memory (ROM). Volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). Non-volatile memory may include, for example, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, hard disk, compact disk, and embedded multimedia card (EMMC).

[0041] The display (208) can output visualized information. For example, the display (208) can output visualized information to a user under the control of at least one processor (207). The display (208) may include hardware components of an electronic device (100) used to display a screen. For example, the display (208) may include light-emitting elements and circuits (e.g., transistors) that control the light-emitting elements to emit light. For example, each of the light-emitting elements may include an organic light-emitting diode (OLED) or a micro LED. However, it is not limited thereto. For example, the display (208) may include a liquid crystal display (LCD).

[0042] The communication circuit (205) may include hardware components to support the transmission and / or reception of signals between the electronic device (100) and an external electronic device. The communication circuit (205) may include, for example, at least one of a modem, an antenna, and an O / E (optic / electronic) converter. The communication circuit (205) may support the transmission and / or reception of signals based on various types of protocols such as Ethernet, LAN (local area network), WAN (wide area network), WiFi (wireless fidelity), Bluetooth, BLE (Bluetooth low energy), Zigbee, LTE (long term evolution), and 5G NR (new radio).

[0043] At least one processor (207) can display an original image (e.g., original image (412) of FIG. 4) through a display (208). At least one processor (207) can receive user input for adding one or more strokes to a surrounding area of ​​the display (208) where the original image is displayed. At least one processor (207) can obtain an edited image (e.g., edited image (422) of FIG. 4) which, at least based on receiving the user input, includes a first part corresponding at least partially to the original image and a second part corresponding to the surrounding area, includes content of the original image extended to the second part based on outpainting, and a visual object represented by the one or more strokes is composited to the second part. For example, at least one processor (207) can transmit an original image (e.g., the original image (412) of FIG. 4) to an artificial intelligence (AI) server (e.g., the AI ​​server (710) of FIG. 7) through a communication circuit (205). For example, at least one processor (207) can receive an edited image (e.g., the edited image (422) of FIG. 4) from the AI ​​server. For example, at least one processor (207) can store the edited image in memory (206).

[0044] FIG. 3 is a flowchart illustrating the operation of an electronic device for acquiring an edited image. This method may be executed by the electronic device (100) illustrated in FIG. 2 or by at least one processor (207) of the electronic device (100).

[0045] Referring to FIG. 3, in operation 310, the electronic device (100) can display an original image (e.g., the original image (412) of FIG. 4) through a display (208). For example, when the electronic device (100) displays the original image, it may receive user input to change the size of the original image. For example, while displaying the original image, the electronic device (100) may receive user input to reduce the size of the original image. For example, based on receiving the user input, the electronic device (100) may display the original image and a UI (user interface) object (e.g., the UI object (526) of FIG. 5) through the display (208). For example, the UI object may be referenced as a visual element.

[0046] In operation 320, the electronic device (100) may receive user input to add one or more strokes to a surrounding area (e.g., surrounding area (522) of FIG. 5) for an area (e.g., area (524) of FIG. 5) of the display (208) on which the original image is displayed. For example, when the electronic device (100) displays the original image, it may receive user input to add one or more strokes to a UI object (e.g., UI object (526) of FIG. 5) displayed in the surrounding area. For example, the electronic device (100) may display the original image in the area (e.g., area (524) of FIG. 5) and display the UI object in the surrounding area (e.g., surrounding area (522) of FIG. 5). For example, the electronic device (100) may display one or more strokes (e.g., one or more strokes (414) of FIG. 4) on a peripheral area of ​​the display (208) (e.g., a peripheral area (522) of FIG. 5) based on receiving the user input.

[0047] In operation 330, the electronic device (100) may, at least based on receiving the user input, obtain an edited image (e.g., an edited image (422) of FIG. 4) comprising a first part corresponding at least partially to the original image and a second part corresponding to the surrounding area, and including content of the original image extended to the second part based on outpainting, and a visual object (e.g., a visual object (426) of FIG. 6a) represented by one or more strokes is composited to the second part. For example, the electronic device (100) may, at least based on receiving the user input, generate or obtain the edited image by providing information about the original image and the user input to an image generation model (e.g., an image generation model (714) of FIG. 7).

[0048] For example, the electronic device (100) may obtain an edited image (e.g., an edited image of FIG. 4 (422)) containing a visual object (e.g., a visual object of FIG. 4 (426)) at a designated location based on receiving user input to add one or more strokes (e.g., one or more strokes (414) of FIG. 4). For example, adding a visual object based on a prompt may be more inconvenient than adding a visual object based on a stroke. For example, specifying a location based on touch input may be more convenient or accurate than specifying a location based on a prompt. For example, the electronic device (100) may cause the visual object to be displayed at an accurate location by receiving user input based on a stroke.

[0049] For example, the electronic device (100) can display the acquired edited image through a display (208). The acquisition of the edited image is described and illustrated in more detail with reference to FIG. 4.

[0050] FIG. 4 illustrates an exemplary operation of an electronic device that acquires an edited image based on user input adding one or more strokes.

[0051] Referring to FIG. 4, the state (410) can be described as a state in which the original image (412) is displayed. For example, the electronic device (100) may display the original image (412) through a display (208). For example, while displaying the original image (412), the electronic device (100) may receive user input for adding one or more strokes. For example, the user input may be received through the display (208). For example, the display (208) may be a touch-sensitive display. For example, the user input may be based on a part of the user of the electronic device (100) (e.g., a finger). For example, the user input may be performed by a stylus pen. However, it is not limited thereto. For example, the user input may include user input by a mouse.

[0052] The electronic device (100) may display one or more strokes (414) through the display (208) based on receiving user input for adding one or more strokes. For example, the electronic device (100) may receive the user input on a surrounding area (e.g., surrounding area (522) of FIG. 5) of an area (e.g., area (524) of FIG. 5) where the original image (412) is displayed. For example, after completing the reception of the user input, the electronic device (100) may receive another user input regarding the executable object (416). For example, the other user input may be described as user input for acquiring or creating an edited image (422). An operation in which an edited image (422) is acquired based on receiving another user input regarding the executable object (416) is described, but embodiments are not limited thereto. For example, the electronic device (100) can acquire an edited image (422) based on the completion of receiving user input for adding one or more strokes.

[0053] The state (420) can be described as a state in which an edited image (422) is obtained from an original image (412). For example, an electronic device (100) may obtain or generate an edited image (422) by providing the original image (412) and one or more strokes (414) to an image generation model (e.g., the image generation model (714) of FIG. 7). For example, the electronic device (100) may display the edited image (422) through a display (208) based on obtaining the edited image (422).

[0054] The edited image (422) may include at least a portion of the original image (412) and a visual object (426) represented by one or more strokes. For example, the edited image (422) may be described as an image in which a visual object (426) is added to an extended image in which outpainting has been performed using the original image (412). For example, the visual object (426) may correspond to one or more strokes (414) represented by one or more strokes. For example, the electronic device (100) may determine or acquire the visual object (426) based on providing one or more strokes (414) to a multimodal model (e.g., the multimodal model (712) of FIG. 7). For example, when determining the visual object (426), the electronic device (100) may use other visual objects (e.g., a house or a tree) included in the original image (412). For example, the electronic device (100) can determine the visual object (426) by identifying the other visual object included in the original image (412) by providing the original image (412) to the multimodal model. For example, the electronic device (100) can obtain a prompt indicating one or more strokes (414) by providing one or more strokes (414) to the multimodal model (e.g., the multimodal model (712) of FIG. 7). For example, the electronic device (100) can obtain the visual object (426) represented by one or more strokes by providing a prompt indicating one or more strokes (414) to the image generation model (e.g., the image generation model (714) of FIG. 7). For example, the electronic device (100) can obtain an edited image (422) by compositing the visual object (426) onto the expanded image. For example, the electronic device (100) can obtain an edited image (422) by performing inpainting on the expanded image and visual object (426).For example, the electronic device (100) can obtain one or more edited images, each including a visual object (426) corresponding to one or more strokes (414), by using the original image (412). For example, the electronic device (100) can obtain a plurality of edited images, each including an edited image (422), by using the original image (412). For example, the electronic device (100) can obtain or generate a plurality of edited images based on the expanded image and the visual object (426).

[0055] According to one embodiment, the size of the original image (412) of the state (410) may be smaller than the size of the edited image (422) of the state (420). For example, since the edited image (422) includes the original image (412), the size of the edited image (422) may be larger than the size of the original image (412). However, it is not limited thereto. For example, the size of the edited image (422) may be equal to or smaller than the size of the original image (412). For example, the size of the edited image (422) that includes at least a portion of the original image (412) may be displayed as being smaller than the original image (412). For example, the electronic device (100) may control the size of the edited image (422) based on user input to change the size of the edited image (422).

[0056] According to one embodiment, the electronic device (100) may obtain or generate another edited image different from the edited image (422) based on receiving user input regarding the executable object (424). For example, the electronic device (100) may receive user input regarding the executable object (424) when displaying the edited image (422). For example, the electronic device (100) may generate another edited image using the original image (412) and one or more strokes (414) based on receiving user input regarding the executable object (424). For example, the electronic device (100) may generate another edited image different from the edited image (422) using an image generation model (e.g., the image generation model (714) of FIG. 7). For example, the electronic device (100) may obtain another edited image different from the edited image (422) because it uses the image generation model.

[0057] For example, the electronic device (100) may obtain an edited image (422) comprising a visual object (426) displayed in a second part of the edited image (422) corresponding to a first part where user input for adding one or more strokes is received. For example, the visual object (426) may be located at another location in the edited image (422) corresponding to a location where one or more strokes (414) are displayed. For example, the location of the visual object (426) within the edited image (422) is described and illustrated in more detail with reference to FIG. 5.

[0058] FIG. 5 illustrates an exemplary operation of an electronic device receiving user input to reduce the size of an original image.

[0059] Referring to FIG. 5, state (510) can be described as a state in which an original image (412) of a first size is displayed through a display (208). State (520) can be described as a state in which an original image (412) of a second size smaller than the first size is displayed through a display (208). For example, while displaying the original image (412) of the first size through the display (208), the electronic device (100) may receive user input to change the size of the original image (412). For example, based on receiving the user input, the electronic device (100) may display the original image (412) of the first size in a second size smaller than the first size. For example, the user input may include a touch input representing a pinch gesture, but is not limited thereto. For example, the user input may include user input to a UI object to change the size of the original image (412).

[0060] According to one embodiment, a visual object (416) for obtaining an edited image (422) may not be activated in a state (510) where only the original image (412) is displayed. For example, the visual object (416) may be activated in a state (520) or state (530) where user input related to the original image (412) is received. However, it is not limited thereto. For example, the visual object (416) may be displayed in state (520) and state (530) among state (510), state (520), and state (530). For example, the visual object (416) may not be displayed in state (510).

[0061] For example, the electronic device (100) may receive other user input indicating the addition of a surrounding area (522) while displaying the original image (412). For example, based on receiving the other user input, the electronic device (100) may display a surrounding area (522) that at least partially surrounds the original image (412) together with the original image (412).

[0062] For example, while the electronic device (100) displays the original image (412) of the second size in an area (524) of the display (208), it may display a UI object (526) in a surrounding area (522) of the display (208). For example, the electronic device (100) may display a UI object (526) on the surrounding area (522) capable of receiving user input for adding one or more strokes based on a reduction in the size of the original image (412). For example, the UI object (526) may include a white canvas. However, it is not limited thereto. For example, the UI object (526) may be referred to as an empty canvas. Referring to FIG. 5, the UI object (526) is depicted as a white or empty canvas, but the embodiment is not limited thereto. For example, the UI object (526) may not be white. For example, the UI object (526) may have a grid pattern.

[0063] The peripheral area (522) may be described as an area surrounding the area (524) where the original image (412) of the second size is displayed. For example, the peripheral area (522) may be an area adjacent to the area (524). For example, the electronic device (100) may receive user input to add one or more strokes to the peripheral area (522). For example, the electronic device (100) may receive user input to add one or more strokes to a UI object (526) displayed in the peripheral area (522). For example, the electronic device (100) may display one or more strokes (414) through the display (208) based on receiving the user input.

[0064] State (520) is illustrated as a state in which a touch input is received to reduce the size of the original image (412), but embodiments are not limited thereto. For example, the electronic device (100) may receive user input regarding an executable object to reduce the size of the original image (412). For example, the electronic device (100) may change the size of the original image (412) based on receiving user input regarding the executable object.

[0065] The state (530) can be described as a state in which the original image (412) of the second size and one or more strokes (414) are displayed. For example, when the electronic device (100) displays the original image (412) on the area (524), it may receive other user input to change the position of the original image (412). For example, based on receiving other user input to change the position of the original image (412), the electronic device (100) may change the position where the original image (412) of the second size is displayed from the area (524) to the area (532). For example, based on receiving other user input, the electronic device (100) may display the original image (412) of the second size on the area (532).

[0066] According to one embodiment, an electronic device (100) can obtain an edited image (422) by using information about the surrounding area of ​​the area (532) where the original image (412) is displayed. For example, when performing outpainting, the electronic device (100) can provide information about the surrounding area of ​​the area (532) where the original image (412) is displayed to an image generation model (e.g., image generation model (714) of FIG. 7). For example, the information may include at least one of the size, orientation, and / or positional relationship with the area where the original image (412) is displayed. For example, the electronic device (100) can obtain an edited image (422) by performing outpainting based on the location where the original image (412) is displayed within a blank canvas. For example, the edited image (422) may be dependent on the surrounding area of ​​the area where the original image (412) is displayed. For example, the electronic device (100) can obtain an expanded image as a result of outpainting using the original image (412) based on the surrounding area. For example, the edited image obtained when the area where the original image (412) is displayed is the top-left (e.g., area (532)) may differ from the edited image obtained when the area where the original image (412) is displayed is the bottom-right. For example, the electronic device (100) can obtain an edited image (422) based on the size of the UI object (526) displayed on the surrounding area of ​​the area where the original image (412) is displayed and / or the positional relationship between the original image (412) and the UI object (526).

[0067] According to one embodiment, the electronic device (100) may display a different image different from the original image (412) within a surrounding area (522). For example, the electronic device (100) may receive other user input for displaying the different image within the surrounding area (522). For example, the electronic device (100) may display the different image on the surrounding area (522) based on receiving the other user input regarding the surrounding area (522). For example, the electronic device (100) may obtain an edited image (422) using one or more strokes (414) and the different image. However, it is not limited thereto. For example, the electronic device (100) may obtain an edited image (422) without one or more strokes (414) using the different image and the original image (412). For example, the electronic device (100) can obtain an edited image (422) in which another visual object represented by the other image is composited onto a second part of the edited image (422) corresponding to a surrounding area (522) where the other user input for displaying the other image is received. For example, the electronic device (100) can obtain an edited image (422) in which another visual object corresponding to the other image is composited by attaching data for the other image to the surrounding area (522).

[0068] According to one embodiment, when the electronic device (100) generates an edited image (422), it may generate different edited images (422) depending on the area where the original image (412) is displayed. For example, the electronic device (100) may generate an expanded image by performing outpainting using information about the surrounding area (532) where the original image (412) is displayed. For example, since the electronic device (100) generates the expanded image based on the area (532) where the original image (412) is displayed, the edited image (422) may be dependent on the area (532) where the original image (412) is displayed.

[0069] When displaying the original image (412), the electronic device (100) may display one or more strokes (414) on the surrounding area (522) based on receiving user input for adding one or more strokes. For example, the electronic device (100) may display an edited image (422) through a display (208) that includes a visual object (426) located at another location of the edited image (422) corresponding to the location of the surrounding area (522) where the user input is received. For example, the edited image (422) may include at least a portion of the original image (412). For example, the edited image (422) may include a first portion corresponding at least partially to the original image (412) and a second portion corresponding to the surrounding area (522). For example, the electronic device (100) may include or place content of the original image (412) on which outpainting has been performed on the original image (412) in the second portion. For example, the first part may be described as a portion of an edited image (422) containing the original image (412). For example, the second part may be described as the remaining portion of the edited image (422) excluding the first part. For example, the electronic device (100) may display content based on outpainting performed using the original image (412) on the second part of the edited image (422).

[0070] An example is described in which a UI object (526) is displayed on a surrounding area (522) based on receiving user input to change the size of the original image (412), but the embodiment is not limited thereto. For example, the electronic device (100) may display the original image (412) of the first size without displaying the UI object (526), ​​and the original image (412) of the second size, which is smaller than the first size, through the display (208) based on receiving user input to change the form of the electronic device (100). For example, the display of the original image (412) of the second size is described and illustrated in more detail with reference to FIG. 6a.

[0071] FIG. 6a illustrates an exemplary operation of an electronic device receiving user input in which the size of an original image is changed.

[0072] Referring to FIG. 6a, the state (610) can be described as a state in which the original image (412) is displayed through a first display of an electronic device (100) that includes a plurality of displays. For example, the electronic device (100) can display the original image (412) through the first display. For example, the state (620) can be described as a state in which the original image (412) is displayed through a second display of an electronic device (100) that includes a plurality of displays. For example, the size of the first display may be smaller than the size of the second display. For example, when the electronic device (100) displays the original image (412) through the first display, it may display a UI object (526) on a surrounding area (622) based on receiving user input to change the display in which the original image (412) is displayed from the first display to the second display. For example, the electronic device (100) may display the original image (412) on an area (624) and display a UI object (526) on a surrounding area (622) of the area (624), based on a decision that the display for displaying the original image (412) is changed from the first display to the second display. For example, the electronic device (100) may receive user input to add one or more strokes to the surrounding area (622). For example, the electronic device (100) may receive user input to add one or more strokes to the UI object (526) on the surrounding area (622).

[0073] For example, the state (610) may be referred to as the folding state. For example, the state (620) may be referred to as the unfolding state. For example, when the electronic device (100) is in the folding state, the original image (412) may be displayed without displaying the UI object. For example, when the electronic device (100) is in the unfolding state, the original image (412) may be displayed together with the UI object.

[0074] According to one embodiment, the electronic device (100) can obtain an edited image (422) of a size corresponding to the size of the UI object (526) in the unfolded state.

[0075] According to one embodiment, the electronic device (100) can obtain an edited image (422) having the same aspect ratio (or size) as the first display included in the electronic device (100). For example, the electronic device (100) can obtain another edited image having the same aspect ratio (or size) as the other aspect ratio (or size) of a second display included in the electronic device (100) that is different from the first display. For example, the electronic device (100) can obtain the edited image and the other edited image together.

[0076] According to one embodiment, the electronic device (100) can acquire an edited image (422) of the same size (or aspect ratio) as an external electronic device (e.g., a smart watch, smart glasses, or a VST (Video See-Through) device) connected to the electronic device (100). However, it is not limited thereto. For example, the electronic device (100) can acquire an edited image (422) that represents a dimension (e.g., 2D (two-dimensional space) or 3D (three-dimensional space)) supported by the external electronic device.

[0077] For example, the electronic device (100) may acquire another edited image based on the edited image. For example, the electronic device (100) may receive user input to acquire the edited image based on a folding state or an unfolding state. For example, the acquisition of the other edited image is described and illustrated in more detail with reference to FIGS. 6b through 6d.

[0078] FIGS. 6b through 6d illustrate exemplary operation of an electronic device that acquires another edited image based on an edited image.

[0079] Referring to FIG. 6b, the state (630) can be described as a state in which the original image (632) is displayed through the display (208). For example, the electronic device (100) may include a multi-foldable electronic device. For example, the state (630) can be described as a state in which the electronic device (100) is fully folded. For example, the fully folded state can be described as a third state.

[0080] State (640) may be described as a state that has been unfolded once from a third state. For example, the state that has been unfolded once may be referred to as a fourth state or a half-folded state. For example, the electronic device (100) may display a visual element (642) through a display (208) based on the transition of the state of the electronic device (100) from state (630) to state (640). For example, the visual element (642) may be described as an element for receiving one or more strokes. For example, the electronic device (100) may receive user input indicating one or more strokes (644) for the visual element (642). For example, the electronic device (100) may display one or more strokes (644) on the visual element (642) based on receiving the user input.

[0081] For example, the position and / or size of the visual element (642) may be based on the size of the display (208). For example, the position and / or size of the visual element (642) may be based on the order or method (or direction) of unfolding the folded electronic device (100). For example, referring to FIG. 6b, the visual element (642) is shown to be located to the left of the original image (632), but the embodiment is not limited thereto. For example, depending on the direction in which the electronic device (100) is unfolded, the visual element (642) may be located to the right of the original image (632).

[0082] Referring to FIG. 6c, state (650) can be described as a state in which another user input for creating an edited image is received in state (640). For example, the electronic device (100) may acquire a first edited image (652) based on receiving the other user input in state (640). For example, the electronic device (100) may display the first edited image (652), which includes a visual object (654) corresponding to one or more strokes (644), through a display (208) based on receiving the other user input. For example, the visual object (654) may be positioned at a location in the first edited image (652) corresponding to a location where user input for adding one or more strokes (644) is received.

[0083] Referring to FIG. 6d, the state (660) can be described as a state in which the electronic device (100) is fully unfolded. For example, the fully unfolded state can be referred to as a fifth state. For example, the electronic device (100) can display a visual element (662) through a display (208) based on the transition of the state of the electronic device (100) from a half-folded state to a fully unfolded state. For example, the visual element (662) can be described as an element for receiving user input to add one or more strokes.

[0084] For example, the position and / or size of the visual element (662) may be based on the size of the display (208). For example, the position and / or size of the visual element (662) may be based on the order or method (or direction) of unfolding the folded electronic device (100). For example, referring to FIG. 6c, the visual element (662) is shown to be located to the left of the first edited image (652), but the embodiment is not limited thereto. For example, depending on the direction in which the electronic device (100) is unfolded, the visual element (662) may be located to the right of the first edited image (652).

[0085] The electronic device (100) may receive user input to add one or more strokes (664) in a state (660). For example, the electronic device (100) may display one or more strokes (664) through a display (208) based on receiving the user input. For example, the electronic device (100) may receive other user input to acquire or create a second edited image when displaying one or more strokes (664) and a first edited image (652).

[0086] A state (670) may be described as a state in which a second edited image (672) is displayed based on the other user input. For example, the electronic device (100) may acquire a second edited image (672) using one or more strokes (664) and a first edited image (652) based on receiving the other user input, and display the second edited image (672) through a display (208). For example, the second edited image (672) may include a visual object (674) corresponding to one or more strokes (664). For example, the visual object (674) may be located at a location in the second edited image (672) corresponding to the location where user input for adding one or more strokes (664) is received. For example, the visual object (674) may be represented by one or more strokes (664).

[0087] The electronic device (100) can generate or acquire an edited image (422) using a multimodal model (e.g., the multimodal model (712) of FIG. 7) and / or an image generation model (e.g., the image generation model (714) of FIG. 7). For example, the electronic device (100) may be connected to an AI server (e.g., the AI ​​server (710) of FIG. 7) via a communication circuit (205). The generation of the edited image (422) based on the AI ​​server is described and illustrated in more detail with reference to FIG. 7.

[0088] FIG. 7 illustrates an exemplary operation of an electronic device receiving data from an AI (artificial intelligence) server.

[0089] Referring to FIG. 7, the electronic device (100) can transmit data to the AI ​​server (710) through the communication circuit (205). For example, the electronic device (100) can receive data from the AI ​​server (710) through the communication circuit (205). For example, the electronic device (100) can obtain an edited image (422) from an original image (412) by communicating with the AI ​​server (710). For example, the electronic device (100) can transmit information representing the original image (412) and the user input through the communication circuit (205) to the AI ​​server (710), based on receiving user input to add one or more strokes when displaying the original image (412).

[0090] The multimodal model (712) can be described as a model trained by machine learning (or deep learning) techniques. The multimodal model (712) can be described as a model trained to output text-based prompts by receiving data of two or more types (e.g., images and text). For example, the AI ​​server (710) can obtain a prompt describing the original image (412) by providing the original image (412) to the multimodal model (712). For example, the AI ​​server (710) can obtain a prompt indicating an object related to the original image (412) (e.g., an object related to a boat—water, an object related to a car—road) by providing the original image (412) to the multimodal model (712). For example, the AI ​​server (710) can obtain a prompt indicating the style of the original image (412) (e.g., cartoon and realistic) by providing the original image (412) to the multimodal model (712). For example, the electronic device (100) can identify the style of the original image (412) using the original image (412). For example, the electronic device (100) can obtain an edited image (422) containing a visual object (426) that is expressed according to the style.

[0091] The image generation model (714) can be described as a model trained through machine learning (or deep learning) techniques. For example, the image generation model (714) can obtain an edited image (422) by receiving a prompt. For example, the AI ​​server (710) can obtain an edited image (422) by providing the image generation model (714) with a prompt obtained from the multimodal model (712). For example, the AI ​​server (710) can obtain or generate an edited image (422) by providing the image generation model (714) with a prompt obtained from the original image (412) and / or the multimodal model (712). For example, the AI ​​server (710) can transmit the edited image (422) to the electronic device (100). For example, the electronic device (100) can receive the edited image (422) from the AI ​​server (710) through a communication circuit (205).

[0092] Referring to FIG. 7, an AI server (710) including a multimodal model (712) and an image generation model (714) is shown as being located outside the electronic device (100), but is not limited thereto. For example, the electronic device (100) may include a multimodal model (712) and / or an image generation model (714). For example, the electronic device (100) may obtain an edited image (422) from an original image (412) by using the multimodal model (712) and / or an image generation model (714) included in the electronic device (100).

[0093] According to one embodiment, the multimodal model (712) and the image generation model (714) are illustrated as being included in the AI ​​server (710), but the embodiment is not limited thereto. For example, the multimodal model (712) may be included in a first electronic device. For example, the image generation model (714) may be included in a second electronic device. For example, the first electronic device or the second electronic device may be an electronic device (100). For example, one of the multimodal model (712) or the image generation model (714) may be included in the electronic device (100), and the other may be included in the AI ​​server (710).

[0094] For example, the electronic device (100) may obtain an edited image (422) by providing one or more strokes (414) and the original image (412) to a trained model (e.g., a multimodal model (712) or an image generation model (714)) based on receiving user input to add one or more strokes (414). For example, the electronic device (100) may obtain a prompt to generate an edited image (422) by providing the original image (412) and one or more strokes (414) to the multimodal model (712). For example, the electronic device (100) may obtain an edited image (422) using the prompt. For example, the prompt may instruct to composite another visual object (e.g., road, river) related to a visual object (426) (e.g., car, boat) represented by one or more strokes (414) into a second part of the edited image (422) corresponding to the surrounding area (522).

[0095] Referring to FIG. 7, the multimodal model (712) and the image generation model (714) are shown as distinct, but the embodiments are not limited thereto. For example, the multimodal model (712) and the image generation model (714) may be the same. For example, the electronic device (100) can obtain an edited image (422) from an original image (412) using the image generation model (714) without the multimodal model (712). For example, the electronic device (100) can acquire an edited image (422) by using a trained model (e.g., a multimodal model (712) or an image generation model (714)) to perform outpainting of the original image (412) to extend the content of the original image (412) along the direction from the first part of the edited image (422) to the second part of the edited image (422), and to perform inpainting of the original image (412) so that a visual object (426) is composited into the second part.

[0096] According to one embodiment, the image generation model (714) may be composed of a model for performing outpainting and a model for performing inpainting. For example, the model for performing outpainting may be described as a model trained to output an expanded image by performing outpainting on a provided image. For example, the model for performing inpainting may be described as a model trained to output an edited image by performing inpainting on a provided image.

[0097] According to one embodiment, the electronic device (100) can obtain information about objects included in the original image (412) by performing object recognition on the original image (412). For example, the electronic device (100) can obtain information (e.g., country, product) related to objects (e.g., national flag, trademark) included in the original image (412) by performing the object recognition. For example, the electronic device (100) can perform object recognition by providing the original image (412) to a multimodal model (712). For example, the electronic device (100) can obtain a prompt indicating the information based on performing the object recognition. For example, the electronic device (100) can perform outpainting by providing the prompt indicating the information obtained by performing the object recognition to an image generation model (714). For example, the electronic device (100) can generate content included in a second part of an edited image (422) corresponding to a surrounding area (522) by providing a prompt indicating the information to an image generation model (714). For example, the electronic device (100) can obtain a prompt indicating features of the original image (412) as it performs object recognition by providing the original image (412) to a multimodal model (712). For example, the electronic device (100) can obtain an edited image (422) containing other visual objects indicating the features by using the prompt. For example, the electronic device (100) can obtain a prompt indicating the automotive industry by identifying the original image (412) indicating a country where the automotive industry is developed as it performs object recognition. For example, the electronic device (100) can obtain an edited image containing other visual objects indicating the automotive industry (e.g., trademarks or designs of automotive companies).

[0098] The electronic device (100) can obtain a prompt representing one or more strokes (414) by providing one or more strokes (414) to a multimodal model (712). For example, the electronic device (100) can determine, create, or obtain a visual object (426) represented by one or more strokes by providing the prompt to an image generation model (714). For example, the electronic device (100) can determine a visual object (426) corresponding to one or more strokes (414) by providing the prompt to an image generation model (714). However, it is not limited thereto. For example, the electronic device (100) can directly obtain or determine a visual object (426) corresponding to one or more strokes (414) by providing one or more strokes (414) to an image generation model (714). For example, the electronic device (100) can obtain an edited image (422) containing a visual object (426) corresponding to one or more strokes (414) by providing one or more strokes (414) and an original image (412) to an image generation model (714).

[0099] According to one embodiment, when the electronic device (100) receives user input for adding one or more strokes, it may receive further user input for specifying a second part of an edited image (422) where a visual object (426) represented by the one or more strokes is to be displayed. For example, the electronic device (100) may determine, based on receiving the other user input, the location on the second part of the edited image (422) where the visual object (426) is to be displayed in a precise (or detailed) manner. For example, the electronic device (100) may obtain a prompt for specifying the part (or location) where the visual object (426) is to be displayed by providing the other user input to a multimodal model (712). For example, the electronic device (100) may obtain an edited image (422) in which the visual object (426) is located in the specified part by providing the prompt to an image generation model (714).

[0100] For example, the electronic device (100) may receive other user input on a surrounding area (522) for selecting or determining the location of a visual object (426). For example, the electronic device (100) may obtain an edited image (422) in which the visual object (426) is composited into the second part based on the location, based on receiving the other user input. For example, the electronic device (100) may receive the other user input before or after receiving the user input for adding one or more strokes (414).

[0101] According to one embodiment, the electronic device (100) may determine or generate a visual object based on other user input for selecting or determining the location of the visual object (426). For example, the electronic device (100) may determine a visual object corresponding to one or more strokes (414) based on the location identified by the other user input. For example, the electronic device (100) may determine a visual object in the shape of an airplane based on the determination that the location is an area corresponding to the sky on an extended image. For example, the electronic device (100) may determine a visual object in the shape of a car based on the determination that the location is an area corresponding to the ground on an extended image.

[0102] When the electronic device (100) acquires an edited image (422), it may use information identified from a software application. For example, the electronic device (100) may acquire an edited image (422) by using a time identified by a software application for time. The creation of the edited image (422) based on the time is described and illustrated in more detail with reference to FIG. 8.

[0103] FIG. 8 illustrates an exemplary operation of an electronic device that acquires an edited image based on time identified from a software application for time.

[0104] Referring to FIG. 8, the state (810) can be described as a state of receiving user input to add one or more strokes. For example, the electronic device (100) may receive user input to add one or more strokes when displaying the original image (412). For example, the electronic device (100) may display one or more strokes (812) and one or more strokes (814) through the display (208). For example, one or more strokes (812) may represent a street light.

[0105] The state (820) can be described as a state in which an edited image (826) is obtained using a time identified by a software application for time. For example, when an electronic device (100) generates an edited image (826), it may obtain information about time from a software application for time. For example, when an electronic device (100) generates an edited image (826), it may use said information to reflect the time. For example, if the time identified by said information is night, the electronic device (100) may obtain an edited image (826) containing a visual object (822) corresponding to said time. For example, if the visual object (822) is a street light, the visual object (822) may represent a street light emitting light. For example, if the time is night, the electronic device (100) may obtain an edited image (826) that represents a dark background. For example, the above edited image (826) may include a visual object (824) corresponding to one or more strokes (814).

[0106] For example, the electronic device (100) can identify the current time based on one or more strokes (812) indicating a change in time represented by the original image (412). For example, the electronic device (100) can identify the current time using a software application for time. For example, the electronic device (100) can obtain an edited image (826) containing the content of the original image (412) changed according to the current time.

[0107] According to one embodiment, an electronic device (100) can obtain weather information from a software application regarding weather. For example, the electronic device (100) can obtain an edited image reflecting the weather using the information. For example, when the electronic device (100) identifies the weather using the information, it can obtain an edited image dependent on the identified weather. For example, the electronic device (100) can obtain an edited image expressing snowy weather based on identifying snowy weather. For example, the electronic device (100) can obtain an edited image expressing rainy weather based on identifying rainy weather.

[0108] When the electronic device (100) acquires an edited image (422) from an original image (412), it may display a user interface through a display (208) indicating that the edited image (422) is being created or acquired. For example, the user interface is described and illustrated in more detail with reference to FIGS. 9 and 10.

[0109] FIGS. 9 and 10 illustrate exemplary operation of an electronic device that displays a screen when an edited image is acquired.

[0110] Referring to FIG. 9, a state (910) can be described as a state in which an original image (412) and one or more strokes (414) are displayed. For example, when an electronic device (100) displays an original image (412) and one or more strokes (414), it may receive user input to create an edited image (422). For example, states (920) and (930) can be described as states in which a screen is displayed before the edited image (422) created based on the user input is displayed. For example, the electronic device (100) may display an extended image created using the original image (412) through a display (208) to guide the creation of the edited image (422). For example, state (920) can be described as a state in which the extended image (925) is displayed with a first transparency. For example, the state (930) may be described as a state in which the expanded image (935) is displayed with a second transparency lower than the first transparency. For example, the electronic device (100) may obtain the expanded image by performing outpainting on the original image (412). For example, the electronic device (100) may perform the outpainting by providing the original image (412) to an image generation model (714).

[0111] For example, the electronic device (100) may display text indicating that an edited image (422) is being generated in states (920) and (930). However, it is not limited thereto. For example, the electronic device (100) may display a visual affordance, visual effect, visual element, or icon to indicate that an edited image (422) is being generated in states (920) and (930). For example, a user of the electronic device (100) may recognize that an edited image (422) is being generated based on the text and / or the expanded image of the first transparency and the expanded image of the second transparency. For example, the transparency of the expanded image output before the edited image (422) is acquired may gradually change over time.

[0112] The state (940) can be described as a state that displays an edited image (422). For example, the electronic device (100) can switch the screen displayed through the display (208) from the expanded image of the second transparency to the edited image (422) based on the completion of the creation of the edited image (422).

[0113] For example, the electronic device (100) may display an executable object (424) for creating another edited image based on the completion of the creation of the edited image (422) through the display (208). For example, the electronic device (100) may activate the executable object (424) based on the completion of the creation of the edited image (422).

[0114] Referring to FIG. 10, a state (1010) may be described as a state in which an original image (412) and one or more strokes (414) are displayed. For example, when the original image (412) and one or more strokes (414) are displayed, the electronic device (100) may receive user input to create an edited image (422). A state (1020) and a state (1030) may be described as a state in which a screen is displayed to guide that an edited image (422) is being created. For example, a state (1020) may be described as a state in which a first part (1025) of the edited image (422) is displayed. For example, the first part (1025) of the edited image (422) may include at least a portion of the edited image (422). For example, the electronic device (100) may transition from a state (1020) to a state (1030) based on the passage of time. For example, the state (1030) may be described as a state in which a second part (1035) of the edited image (422) is displayed. For example, the second part (1035) of the edited image (422) may include a first part (1025) of the edited image (422). For example, the size of the second part (1035) of the edited image (422) may be larger than the other size of the first part (1025) of the edited image (422). For example, the electronic device (100) may display a part of the edited image (422) that gradually expands over time through the display (208).

[0115] For example, the electronic device (100) may display text indicating that an edited image (422) is being created in state (1020) and state (1030). For example, a user of the electronic device (100) may recognize that an edited image (422) is being created based on the text and / or a first part (1025) of the edited image (422) and a second part (1035) of the edited image (422).

[0116] The state (1040) can be described as a state in which an edited image (422) is displayed. For example, the electronic device (100) may display the entire edited image (422) through the display (208) based on the completion of the creation of the edited image (422) over time. For example, the electronic device (100) may display an executable object (424) for creating another edited image through the display (208) based on the completion of the creation of the edited image (422). For example, the electronic device (100) may activate the executable object (424) based on the completion of the creation of the edited image (422).

[0117] For example, the electronic device (100) may refrain from, bypass, or block the synthesis of a visual object corresponding to one or more strokes (e.g., one or more strokes (1112) of FIG. 11) based on receiving another user input for adding one or more strokes. The operation of refraining from synthesizing the visual object is described and illustrated in more detail with reference to FIG. 11.

[0118] FIG. 11 illustrates an exemplary operation of an electronic device receiving other user input to avoid synthesis.

[0119] Referring to FIG. 11, the state (1110) may be described as a state in which another user input is received to add one or more other strokes (1114) to one or more strokes (1112). For example, the electronic device (100) may receive user input to add one or more strokes (414) when displaying the original image (412). For example, the electronic device (100) may receive user input to add one or more strokes (1112) when displaying the original image (412). For example, the electronic device (100) may display one or more strokes (414) and one or more strokes (1112) through the display (208) based on receiving the user input. For example, the electronic device (100) may receive another user input to add one or more strokes (1114) to one or more strokes (1112). For example, one or more strokes (1114) may be in the shape of an X. However, they are not limited thereto.

[0120] A state (1120) can be described as a state in which an edited image (422) is displayed upon receiving user input for an executable object (416) of a state (1110). For example, the electronic device (100) may receive user input that causes the acquisition of an edited image (422) and the display of an edited image (422) upon displaying an original image (412), one or more strokes (414), one or more strokes (1112), and one or more strokes (1114) for an executable object (416). For example, the electronic device (100) may display the edited image (422) through a display (208) based on receiving user input that causes the acquisition of an edited image (422) and the display of an edited image (422).

[0121] The edited image (422) displayed in state (1120) may include a visual object (426) corresponding to one or more strokes (414). The edited image (422) displayed in state (1120) may not include a visual object (e.g., a boat) corresponding to one or more strokes (1112). For example, the electronic device (100) may obtain a prompt to refrain from synthesizing a visual object represented by one or more strokes (1112) by providing the multimodal model (712) with one or more strokes (1114) displayed overlapping with one or more strokes (1112) (or one or more strokes (1114) displayed adjacent to one or more strokes (1112) within a specified distance). For example, the electronic device (100) can obtain an edited image (422) that does not contain a visual object represented by one or more strokes (1112) by providing the prompt to an image generation model (714). For example, the action of refraining from synthesizing a visual object corresponding to one or more strokes (1112) based on one or more strokes (1114) can be referred to as a negative prompt or negative drawing input.

[0122] According to one embodiment, the electronic device (100) may receive other user input for adding one or more strokes (1114) to the original image (412). For example, the electronic device (100) may obtain an edited image (422) that does not contain said visual object based on receiving other user input for adding one or more strokes (1114) to said visual object contained in the original image (412). For example, the electronic device (100) may receive other user input for adding one or more strokes (1114) to the first visual object among the first visual object contained in the original image (412) and the second visual object contained in the original image (412). For example, the electronic device (100) may obtain an edited image (422) containing the second visual object among the first visual object and the second visual object, based on receiving other user input for adding one or more strokes (1114) to the first visual object.

[0123] According to one embodiment, the electronic device (100) can obtain an edited image (422) without a visual object (e.g., a boat) corresponding to one or more strokes (1112) being composited into a second part of the edited image (422) based on providing one or more strokes (1114) displayed on one or more strokes (1112) to an image generation model (714). For example, the electronic device (100) can display one or more strokes (e.g., one or more strokes (414) and one stroke (1112)) on a surrounding area (522) based on receiving user input to add one or more strokes (414). For example, the electronic device (100) can receive other user input to add one or more other strokes (1114) for at least a portion of the one or more strokes. For example, the electronic device (100) may acquire an edited image (422) based on receiving the other user input without including at least a portion (e.g., a boat) of a visual object (e.g., a car and a boat) corresponding to at least a portion of the one or more strokes in the second portion. For example, the electronic device (100) may acquire an edited image (422) in which a visual object (426) corresponding to the remaining portion (e.g., a car) among the at least portion of the one or more strokes and the remaining portion different from the at least portion is composited in the second portion.

[0124] According to one embodiment, the electronic device (100) may receive negative drawing input based on a phase. For example, in a first phase, the electronic device (100) may receive user input for adding one or more strokes (e.g., one or more strokes (414) and one or more strokes (1112)) corresponding to a visual object (426) to be composited onto an edited image (422). For example, the electronic device (100) may switch the phase of the electronic device (100) from the first phase to a second phase based on the completion of receiving the user input. For example, in the second phase, the electronic device (100) may receive another user input for adding one or more strokes (1114) to at least some of the one or more strokes added in the first phase. For example, the electronic device (100) may receive other user input for adding one or more strokes (1114) to at least some of one or more strokes (e.g., one or more strokes (1112)). For example, based on receiving the other user input, the electronic device (100) may display one or more strokes (1114) superimposed on one or more strokes (1112). For example, based on receiving the other user input, the electronic device (100) may obtain an edited image (422) that includes a visual object (426) corresponding to one or more strokes (414) and does not include a visual object (e.g., a boat) corresponding to one or more strokes (1112).

[0125] The electronic device (100) may receive user input for adding one or more strokes. For example, the one or more strokes may be in the form of text. For example, the electronic device (100) may obtain an edited image (422) using the one or more strokes expressed in the form of text. For example, the one or more strokes expressed in the form of text are described and illustrated in more detail with reference to FIG. 12.

[0126] FIG. 12 illustrates an exemplary operation of an electronic device receiving user input for adding one or more strokes expressed as text.

[0127] Referring to FIG. 12, the state (1210) can be described as a state in which one or more strokes (1214) expressed as text are displayed. For example, the electronic device (100) may receive user input to add one or more strokes when the original image (412) is displayed. For example, the electronic device (100) may display one or more strokes (1214) through the display (208) based on receiving the user input. For example, one or more strokes (1214) may be expressed in the form of text. For example, the electronic device (100) may obtain an edited image (422) by providing the original image (412) and one or more strokes (1214) to a trained model (e.g., a multimodal model (712) and / or an image generation model (714)).

[0128] A state (1220) can be described as a state in which an edited image (422) generated using one or more strokes (1214) is displayed. For example, an electronic device (100) can obtain a prompt indicating one or more strokes (1214) by providing one or more strokes (1214) to a multimodal model (712). For example, an electronic device (100) can obtain an edited image (422) by providing said prompt to an image generation model (714). For example, an electronic device (100) can obtain an edited image (422) containing a visual object (426) represented by one or more strokes (1214) based on providing the original image (412) and the prompt obtained by providing one or more strokes (1214) to the multimodal model (712) to the image generation model (714). For example, a visual object (426) may be positioned at another location on the edited image (422) corresponding to the location where user input for adding one or more strokes (1214) is received.

[0129] According to one embodiment, when the electronic device (100) acquires an edited image (422), it may use information stored in memory (206). For example, the electronic device (100) may determine information available to acquire the edited image (422) by using one or more strokes. For example, the electronic device (100) may identify information about the user of the electronic device (100) through memory (206) based on the determination that one or more strokes represent a user. For example, the electronic device (100) may acquire the edited image (422) by using the information identified through memory (206). For example, the electronic device (100) may acquire the edited image (422) by using the user's image stored in memory (206).

[0130] The electronic device (100) may receive user input to determine at least a portion of the original image (412) to be used to generate an edited image (422). For example, the electronic device (100) may determine or specify at least a portion (1315) of the original image (412) based on receiving the user input. The determination of the at least portion (1315) is described and illustrated in more detail with reference to FIG. 13.

[0131] FIG. 13 illustrates an exemplary operation of an electronic device receiving other user input to determine at least a portion of an original image.

[0132] Referring to FIG. 13, the state (1310) can be described as a state in which at least a portion (1315) of the original image (412) is determined. For example, when displaying the original image (412), the electronic device (100) may receive user input to determine at least a portion (1315) of the original image (412). For example, based on receiving the user input, the electronic device (100) may use at least a portion (1315) of the original image (412) as a source for performing outpainting.

[0133] A state (1320) can be described as a state in which, when at least a portion (1315) of the original image (412) is displayed, user input is received to add one or more strokes (414). For example, an electronic device (100) may receive user input to add one or more strokes (414) when at least a portion (1315) of the original image (412) is displayed. For example, based on receiving the user input, the electronic device (100) may display at least a portion (1315) of the original image (412) and one or more strokes (414) through a display (208). For example, based on receiving the user input, the electronic device (100) may obtain an edited image (422) using at least a portion (1315) of the original image (412). For example, an electronic device (100) may obtain an edited image (422) containing a visual object (426) represented by one or more strokes (414) by using at least a portion (1315) of the original image (412) and the remaining portion of the original image (412). For example, when the electronic device (100) obtains the edited image (422) by using at least a portion (1315) of the original image (412), it may refrain from, bypass, or block the remaining portion that is different from at least a portion (1315) of the original image (412) in order to obtain the edited image (422). For example, the edited image (422) obtained by using at least a portion (1315) may not include the remaining portion of the original image (412) that is different from at least a portion (1315) of the original image (412). For example, the electronic device (100) may receive other user input for selecting at least a portion (1315) of the original image (412).For example, the electronic device (100) can obtain an edited image (422) including at least a portion (1315) of the original image (412) and a remaining portion different from at least a portion (1315) of the original image (412), based on receiving user input for adding one or more strokes (414) and the other user input.

[0134] When displaying the original image (412), the electronic device (100) may acquire an edited image based on the location where user input for adding one or more strokes is received. For example, the electronic device (100) may acquire a first edited image when receiving the user input for a first area where the original image (412) is displayed. For example, the electronic device (100) may acquire a second edited image (e.g., edited image (422)) when receiving the user input for a second area different from the first area. For example, the acquisition of the first edited image and the second edited image is described and illustrated in more detail with reference to FIG. 14.

[0135] FIG. 14 is a flowchart illustrating the operation of an electronic device for acquiring a first edited image or a second edited image. This method may be executed by the electronic device (100) illustrated in FIG. 2 or at least one processor (207) of the electronic device (100).

[0136] Referring to FIG. 14, in operation 1410, the electronic device (100) can display the original image (412) through a first area of ​​the display (208). For example, the operation 1410 may correspond to operation 310 of FIG. 3.

[0137] In operation 1420, the electronic device (100) may obtain a first edited image representing the original image (412) in which a first visual object represented by one or more first strokes is composited based on a first user input for adding one or more first strokes within the first area. For example, the electronic device (100) may receive a first user input for adding one or more first strokes with respect to the first area where the original image (412) is displayed. For example, the electronic device (100) may receive the first user input for the first area without receiving the first user input for the second area. For example, the electronic device (100) may obtain a first edited image in which a first visual object represented by one or more first strokes is composited in an area of ​​the first edited image corresponding to the first area. For example, the acquisition of the first edited image may be described by a sketch-to-image technique.

[0138] In operation 1430, the electronic device (100) can obtain a second edited image (e.g., edited image (422)) including the original image (412), by synthesizing a second visual object represented by one or more second strokes based on a second user input for adding one or more second strokes within a second area adjacent to the first area. For example, operation 1430 may correspond to operation 330 of FIG. 3.

[0139] For example, the second area may be an area surrounding the first area where the original image (412) is displayed. For example, the second edited image may be composited with a second visual object (e.g., visual object (426)) represented by one or more second strokes (e.g., one or more strokes (414)). For example, the second edited image may include at least a portion of the original image (412).

[0140] For example, at least a portion of the original image (412) may be located in a first portion of the second edited image corresponding to the first area. For example, the second visual object may be located in a second portion of the second edited image corresponding to the second area. For example, the second edited image may include a first portion corresponding to a first area where the original image (412) is displayed, and a second portion corresponding to a second area adjacent to the first area.

[0141] For example, the electronic device (100) may obtain a prompt instructing the creation of a second edited image by providing the original image (412) and one or more second strokes to the multimodal model (712) based on receiving the second user input. For example, the electronic device (100) may obtain or create the second edited image by providing the prompt to the image creation model (714).

[0142] FIG. 15 is a block diagram of an electronic device in a network environment according to various embodiments.

[0143] FIG. 15 is a block diagram of an electronic device (1501) in a network environment (1500) according to various embodiments. Referring to FIG. 15, in the network environment (1500), the electronic device (1501) may communicate with an electronic device (1502) through a first network (1598) (e.g., a short-range wireless communication network) or may communicate with at least one of an electronic device (1504) or a server (1508) through a second network (1599) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (1501) may communicate with the electronic device (1504) through a server (1508). According to one embodiment, the electronic device (1501) may include a processor (1520), memory (1530), input module (1550), sound output module (1555), display module (1560), audio module (1570), sensor module (1576), interface (1577), connection terminal (1578), haptic module (1579), camera module (1580), power management module (1588), battery (1589), communication module (1590), subscriber identification module (1596), or antenna module (1597). In some embodiments, at least one of these components (e.g., connection terminal (1578)) may be omitted from the electronic device (1501), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (1576), camera module (1580), or antenna module (1597)) may be integrated into a single component (e.g., display module (1560)).

[0144] The processor (1520) can, for example, execute software (e.g., program (1540)) to control at least one other component (e.g., hardware or software component) of the electronic device (1501) connected to the processor (1520) and perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (1520) can store commands or data received from other components (e.g., sensor module (1576) or communication module (1590)) in volatile memory (1532), process the commands or data stored in volatile memory (1532), and store the resulting data in non-volatile memory (1534). According to one embodiment, the processor (1520) may include a main processor (1521) (e.g., a central processing unit or an application processor) or an auxiliary processor (1523) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (1501) includes a main processor (1521) and an auxiliary processor (1523), the auxiliary processor (1523) may be configured to use lower power than the main processor (1521) or to be specialized for a specified function. The auxiliary processor (1523) may be implemented separately from the main processor (1521) or as part thereof.

[0145] The auxiliary processor (1523) may control at least some of the functions or states associated with at least one component of the electronic device (1501) (e.g., display module (1560), sensor module (1576), or communication module (1590)) on behalf of the main processor (1521) while the main processor (1521) is in an inactive (e.g., sleep) state, or together with the main processor (1521) while the main processor (1521) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (1523) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (1580) or communication module (1590)). According to one embodiment, the auxiliary processor (1523) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (1501) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (1508)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.

[0146] The memory (1530) can store various data used by at least one component of the electronic device (1501) (e.g., processor (1520) or sensor module (1576)). The data may include, for example, input data or output data for software (e.g., program (1540)) and related commands. The memory (1530) may include volatile memory (1532) or non-volatile memory (1534).

[0147] The program (1540) may be stored as software in memory (1530) and may include, for example, an operating system (1542), middleware (1544), or an application (1546).

[0148] The input module (1550) can receive commands or data to be used for a component of the electronic device (1501) (e.g., processor (1520)) from outside the electronic device (1501) (e.g., user). The input module (1550) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0149] The sound output module (1555) can output a sound signal to the outside of the electronic device (1501). The sound output module (1555) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.

[0150] The display module (1560) can visually provide information to an external (e.g., user) of the electronic device (1501). The display module (1560) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (1560) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.

[0151] The audio module (1570) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (1570) can acquire sound through an input module (1550) or output sound through an audio output module (1555) or an external electronic device (e.g., electronic device (1502)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (1501).

[0152] The sensor module (1576) can detect the operating state of the electronic device (1501) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (1576) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0153] The interface (1577) may support one or more specified protocols that can be used for the electronic device (1501) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (1502)). According to one embodiment, the interface (1577) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0154] The connection terminal (1578) may include a connector through which the electronic device (1501) can be physically connected to an external electronic device (e.g., electronic device (1502)). According to one embodiment, the connection terminal (1578) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0155] The haptic module (1579) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (1579) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.

[0156] The camera module (1580) can capture still images and video. According to one embodiment, the camera module (1580) may include one or more lenses, image sensors, image signal processors, or flashes.

[0157] The power management module (1588) can manage the power supplied to the electronic device (1501). According to one embodiment, the power management module (1588) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).

[0158] The battery (1589) can supply power to at least one component of the electronic device (1501). According to one embodiment, the battery (1589) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0159] The communication module (1590) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (1501) and an external electronic device (e.g., electronic device (1502), electronic device (1504), or server (1508)), and the performance of communication through the established communication channel. The communication module (1590) may include one or more communication processors that operate independently of the processor (1520) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1590) may include a wireless communication module (1592) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (1594) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (1504) through a first network (1598) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (1599) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1592) can identify or authenticate the electronic device (1501) within a communication network such as the first network (1598) or the second network (1599) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (1596).

[0160] The wireless communication module (1592) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (1592) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (1592) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (1592) can support various requirements specified in the electronic device (1501), external electronic device (e.g., electronic device (1504)), or network system (e.g., second network (1599)). According to one embodiment, the wireless communication module (1592) can support a Peak data rate (e.g., 20 Gbps or more) for realizing eMBB, loss coverage (e.g., 164 dB or less) for realizing mMTC, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for realizing URLLC.

[0161] An antenna module (1597) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (1597) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (1597) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (1598) or a second network (1599), may be selected from the plurality of antennas, for example, by a communication module (1590). A signal or power may be transmitted or received between the communication module (1590) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (1597).

[0162] According to various embodiments, the antenna module (1597) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.

[0163] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.

[0164] According to one embodiment, commands or data may be transmitted or received between the electronic device (1501) and an external electronic device (1504) through a server (1508) connected to a second network (1599). Each of the external electronic devices (1502, or 1504) may be the same or a different type of device as the electronic device (1501). According to one embodiment, all or part of the operations performed on the electronic device (1501) may be performed on one or more of the external electronic devices (1502, 1504, or 1508). For example, if the electronic device (1501) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (1501) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (1501). The electronic device (1501) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (1501) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (1504) may include an Internet of Things (IoT) device. The server (1508) may be an intelligent server using machine learning and / or neural networks.According to one embodiment, an external electronic device (1504) or server (1508) may be included within the second network (1599). The electronic device (1501) may be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0165] Some of the operations described above may be executed (or performed) through an artificial intelligence (AI) system described with reference to FIG. 16.

[0166] Figure 16 is a schematic diagram of an exemplary AI system.

[0167] Referring to FIG. 16, the AI ​​system (1600) may include an input / output interface (1610), an AI framework (1620), a generative AI model (1630), and / or a knowledge repository (1690).

[0168] The input / output interface (1610) may receive input. The input may include user input and / or data obtained or generated by an electronic device (e.g., the electronic device (100) or electronic device (1501) described above). The data may include images, videos, and / or sensor data generated by at least one processor of the electronic device (e.g., at least one processor (207) or processor (1520)), such as illuminance data around the electronic device obtained from a sensor or sensor hub (e.g., auxiliary processor (1523), attitude data (or orientation data) of the electronic device, temperature inside the electronic device (e.g., display (208)), or temperature of at least one processor (207), size information of the display area of ​​the display (208), and / or images obtained through an image sensor of the electronic device (e.g., included in a camera module (1580)). The user input may include natural language, touch data obtained through a touch circuit included within the display panel (e.g., used to identify input from a finger and / or stylus), an image displayed (and / or to be displayed) on the display panel, and / or video. By example, without limitation, the user input may be received by an input / output interface (1610) along with context information. The context information may be described as additional information obtained in relation to the user input. The context information may be related to the state at the time the user input is received (e.g., the state of the electronic device and / or the state of the surroundings of the electronic device (e.g., user state)). For example, the context information may include information about one or more software applications executed within the electronic device at the time the user input is received.For example, the above situation information may include information about the location of the electronic device (or the location of the user of the electronic device) at the time the user input is received. For example, the user input may be integrated with the situation information. For example, the user input with the situation information integrated as input may be received by the input / output interface (1610).

[0169] The input / output interface (1610) may transmit (or provide) an output. The output may include a result (or result information) generated or obtained by the AI ​​system (1600) based on at least part of the input. The format of the output may vary. For example, the output may include natural language. For example, the output may include content (e.g., media content and / or multimedia content). For example, the output may include actions related to the user of the electronic device. For example, the output may have a format according to the user settings of the electronic device.

[0170] The input / output interface (1610) can be described as a user question / response interface (1610).

[0171] The AI ​​framework (1620) can be used to obtain information (or data) about the input from the input / output interface (1610) and to control one or more components related to the AI ​​system (1600) using the obtained information.

[0172] For example, a prompt design component (1621) within an AI framework (1620) can generate or obtain prompts for a generative AI model (1630) (e.g., including a large language model (LLM) or a large multimodal model (LMM)) using the acquired information. For example, the prompt design component (1621) may be described as an AI component that uses a learning algorithm and / or a neural network to provide prompts that are enhanced over time. For example, the prompt design component (1621) can generate or obtain prompts by accessing a knowledge component (e.g., a knowledge repository (1690)) containing user preference data, a prompt library, and / or prompt examples using the acquired information. The generated prompts may be provided to the generative AI model (1630) (e.g., including an LLM or LMM).

[0173] For example, an API / plugin management component (1622) within the AI ​​framework (1620) may be used to support communication for additional information requested (or induced) in relation to the prompt provided (or to be provided) to the generative AI model (1630). For example, the API / plugin management component (1622) may be used to create or establish a channel for communication with various data sources (e.g., knowledge repository (1690)). For example, the API / plugin management component (1622) may support access to at least some of the data sources. For example, the API / plugin management component (1622) may be used to request another component (e.g., application / service component (1680)) that performs feedback (or response) according to the prompt. As a non-limiting example, information obtained (or generated) through the API / plugin management component (1622) may be provided to the prompt design component (1621) for generating a prompt. As a non-limiting example, information obtained (or generated) through the API / plugin management component (1622) may be provided to the generative AI model (1630).

[0174] For example, an improvement component (1623) within the AI ​​framework (1620) can at least partially tune (or adjust) (or change) the result (e.g., content) obtained (or output) from the generative AI model (1630). For example, the improvement component (1623) can determine or verify whether the content obtained from the generative AI model (1630) is related to the input. For example, the improvement component (1623) can determine or verify whether the content obtained from the generative AI model (1630) contains biased content. For example, the improvement component (1623) can determine or verify whether the content obtained from the generative AI model (1630) contains harmful content. For example, the improvement component (1623) can support or assist in performing additional processing to improve the content obtained from the generative AI model (1630). For example, the improvement component (1623) may support providing a hint to the user to improve the content.

[0175] A generative AI model (1630) can be described as an artificial intelligence neural network that generates feedback in response to a prompt. For example, the feedback may include additional data and / or information relative to the prompt, but relative to the prompt. For example, the feedback may include new content relative to the prompt. For example, the generative AI model (1630) may include a model that generates images and / or a model that generates language. For example, the model that generates images may include a generative adversarial network (GAN) and / or a variational autoencoder (VAE). For example, the model that generates images may include a diffusion-based generative model (e.g., a transformer VAE). For example, the model that generates language may include CHAT-GPT 3 and / or CHAT-GPT 4. For example, the generative AI model (1630) may include an LMM that generates the feedback by recognizing text, images, and / or speech.

[0176] As an example without limitation, the AI ​​framework (1620) and / or generative AI model (1630) may be included within an AI module (e.g., including a processing circuit) within the electronic device. For example, the AI ​​module may be operatively coupled with at least one processor of the electronic device (e.g., at least one processor (207) or processor (1520)). For example, the AI ​​module may be operatively coupled with a display driving circuit of the electronic device (e.g., a display driving circuit or DDI). For example, the AI ​​module may be operatively coupled with a sensor hub of the electronic device for one or more sensors within the electronic device.

[0177] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure pertains.

[0178] An electronic device as described above (e.g., electronic device (100)) may include a memory (e.g., memory (206)) for storing instructions. The electronic device may include a display (e.g., display (208)). The electronic device may include at least one processor (e.g., at least one processor (207)). The instructions may cause the electronic device to display an original image (e.g., original image (412)) through the display when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to receive user input to add one or more strokes (e.g., one or more strokes (414)) on a surrounding area of ​​the display where the original image is displayed when executed individually or collectively by the at least one processor. The above instructions may cause the electronic device to obtain an edited image (e.g., edited image (422)), comprising, at least based on receiving the user input, a first part corresponding at least partially to the original image, and a second part corresponding to the surrounding area and in which a visual object (e.g., visual object (426)) represented by the one or more strokes is composited.

[0179] According to one embodiment, the instructions may cause the electronic device to acquire the edited image by providing the one or more strokes and the original image to a trained model (e.g., an image generation model (714)) based on receiving the user input, when executed individually or collectively by the at least one processor.

[0180] According to one embodiment, the instructions may cause the electronic device to obtain a prompt for generating the edited image by providing the original image and the one or more strokes to a multimodal model (e.g., multimodal model (712)) when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to obtain the edited image by using the prompt when executed individually or collectively by the at least one processor.

[0181] According to one embodiment, the prompt may instruct to composite another visual object related to the visual object represented by the one or more strokes into the second part.

[0182] According to one embodiment, the instructions may cause the electronic device to obtain a prompt indicating features of the original image as it performs object recognition by providing the original image to a multimodal model when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to obtain an edited image including another visual object indicating the features by using the prompt when executed individually or collectively by the at least one processor.

[0183] According to one embodiment, the instructions may cause the electronic device to identify the style of the original image using the original image when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to acquire the edited image including the visual object represented according to the style when executed individually or collectively by the at least one processor.

[0184] According to one embodiment, the instructions may cause the electronic device to identify the current time based on the one or more strokes indicating to change the time represented by the original image when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to acquire the edited image including the content of the original image changed according to the current time when executed individually or collectively by the at least one processor.

[0185] According to one embodiment, the instructions may cause the electronic device to receive other user input on the surrounding area for determining the location of the visual object when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to obtain the edited image in which the visual object is composited into the second part based on the location, based on receiving the other user input when executed individually or collectively by the at least one processor.

[0186] According to one embodiment, the instructions may cause the electronic device to display the one or more strokes on the surrounding area based on receiving the user input when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to receive another user input for adding one or more other strokes for at least a portion of the one or more strokes when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to acquire the edited image without including at least a portion of the visual object corresponding to the at least a portion of the one or more strokes in the second portion, based on receiving the other user input when executed individually or collectively by the at least one processor.

[0187] According to one embodiment, the instructions may cause the electronic device to receive another user input indicating the addition of the surrounding area while displaying the original image, when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to display the surrounding area that at least partially surrounds the original image, together with the original image, based on receiving the other user input, when executed individually or collectively by the at least one processor.

[0188] According to one embodiment, the instructions may cause the electronic device to receive another user input for selecting at least a portion of the original image when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to obtain the edited image comprising the at least portion of the original image among the at least portion of the original image and the remaining portion different from the at least portion of the original image, based on receiving the user input and the other user input when executed individually or collectively by the at least one processor.

[0189] According to one embodiment, the instructions may cause the electronic device to acquire the edited image by using a trained model to perform outpainting of the original image to extend the content of the original image along the direction from the first part to the second part, and to perform inpainting of the original image so that the visual object is composited into the second part, when executed individually or collectively by the at least one processor.

[0190] According to one embodiment, the instructions may cause the electronic device to display the edited image through the display, based on acquiring the edited image, when executed individually or collectively by the at least one processor.

[0191] A method performed by an electronic device (e.g., electronic device (100)) having a display (e.g., display (208)) as described above may include an operation of displaying an original image (e.g., original image (412)) through the display. The method may include an operation of receiving user input for adding one or more strokes on a surrounding area (e.g., surrounding area (522)) for an area of ​​the display where the original image is displayed. The method may include an operation of obtaining an edited image (e.g., edited image (422)) comprising, at least based on receiving the user input, a first part corresponding at least partially to the original image, and a second part corresponding to the surrounding area and composited with a visual object (e.g., visual object (426)) represented by the one or more strokes.

[0192] According to one embodiment, the method may include the operation of obtaining the edited image by providing the one or more strokes and the original image to a trained model (e.g., an image generation model (714)) based on receiving the user input.

[0193] According to one embodiment, the method may include the operation of obtaining a prompt for generating the edited image by providing the original image and the one or more strokes to a multimodal model (e.g., a multimodal model (712)). The method may include the operation of obtaining the edited image using the prompt.

[0194] According to one embodiment, the prompt may instruct to composite another visual object related to the visual object represented by the one or more strokes into the second part.

[0195] According to one embodiment, the method may include an operation of obtaining a prompt representing a feature of the original image as object recognition is performed by providing the original image to a multimodal model. The method may include an operation of obtaining the edited image including another visual object representing the feature using the prompt.

[0196] According to one embodiment, the method may include an operation of identifying the style of the original image using the original image. The method may include an operation of obtaining the edited image including the visual object expressed according to the style.

[0197] According to one embodiment, the method may include an operation of identifying a current time based on one or more strokes indicating to change the time represented by the original image. The method may include an operation of obtaining an edited image including the content of the original image changed according to the current time.

[0198] According to one embodiment, the method may include receiving another user input for determining the location of the visual object on the surrounding area. The method may include, based on receiving the other user input, obtaining the edited image in which the visual object is composited on the second part based on the location.

[0199] According to one embodiment, the method may include an operation of displaying the one or more strokes on the surrounding area based on receiving the user input. The method may include an operation of receiving another user input for adding one or more other strokes for at least a portion of the one or more strokes. The method may include an operation of acquiring the edited image based on receiving the other user input without including at least a portion of the visual object corresponding to the at least a portion of the one or more strokes in the second portion.

[0200] According to one embodiment, the method may include receiving another user input indicating the addition of the surrounding area while displaying the original image. Based on receiving the other user input, the method may include displaying the surrounding area that at least partially encloses the original image together with the original image.

[0201] According to one embodiment, the method may include an operation of receiving another user input for selecting at least a portion of the original image. Based on receiving the user input and the other user input, the method may include an operation of obtaining the edited image comprising the at least portion of the original image among the at least portion of the original image and a remaining portion different from the at least portion of the original image.

[0202] According to one embodiment, the method may include the operation of obtaining the edited image using a trained model to perform outpainting of the original image to extend the content of the original image along the direction from the first part to the second part, and to perform inpainting of the original image so that the visual object is composited into the second part.

[0203] According to one embodiment, the method may include an operation of displaying the edited image through the display based on acquiring the edited image.

[0204] In a computer-readable storage medium in which one or more programs are stored as described above, the one or more programs may include instructions that cause the electronic device (e.g., electronic device (100)) having a display (e.g., display (208)) to display an original image (e.g., original image (412)) through the display when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to receive user input to add one or more strokes (e.g., strokes) on a peripheral area (e.g., peripheral area (522)) for an area of ​​the display where the original image is displayed when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to obtain an edited image (e.g., edited image (422)), comprising, at least based on receiving the user input when executed by the electronic device, a first part corresponding at least partially to the original image, and a second part corresponding to the surrounding area and in which a visual object (e.g., visual object (426)) represented by the one or more strokes is composited.

[0205] According to one embodiment, the one or more programs may include instructions that cause the electronic device to obtain the edited image by providing the one or more strokes and the original image to a trained model (e.g., an image generation model (714)) based on receiving the user input when executed by the electronic device.

[0206] According to one embodiment, the one or more programs may include instructions that cause the electronic device to obtain a prompt for generating the edited image by providing the original image and the one or more strokes to a multimodal model (e.g., multimodal model (712)) when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain the edited image by using the prompt when executed by the electronic device.

[0207] According to one embodiment, the prompt may instruct to composite another visual object related to the visual object represented by the one or more strokes into the second part.

[0208] According to one embodiment, the one or more programs may include instructions that cause the electronic device to obtain a prompt indicating a feature of the original image as it performs object recognition by providing the original image to a multimodal model when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain the edited image including another visual object indicating the feature by using the prompt when executed by the electronic device.

[0209] According to one embodiment, the one or more programs may include instructions that cause the electronic device to identify the style of the original image using the original image when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain the edited image including the visual object represented according to the style when executed by the electronic device.

[0210] According to one embodiment, the one or more programs may include instructions that cause the electronic device to identify the current time based on the one or more strokes indicating to change the time represented by the original image when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain the edited image containing the content of the original image changed according to the current time when executed by the electronic device.

[0211] According to one embodiment, the one or more programs may include instructions that cause the electronic device to receive other user input on the surrounding area for determining the location of the visual object when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain the edited image in which the visual object is composited into the second part based on the location, based on receiving the other user input when executed by the electronic device.

[0212] According to one embodiment, the one or more programs may include instructions that cause the electronic device to display the one or more strokes on the surrounding area based on receiving the user input when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to receive other user input for adding one or more other strokes for at least a portion of the one or more strokes when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to acquire the edited image without including at least a portion of the visual object corresponding to the at least a portion of the one or more strokes in the second portion, based on receiving the other user input when executed by the electronic device.

[0213] According to one embodiment, the one or more programs may include instructions that cause the electronic device to receive other user input indicating the addition of the surrounding area while displaying the original image when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to display the surrounding area that at least partially encloses the original image together with the original image, based on receiving the other user input when executed by the electronic device.

[0214] According to one embodiment, the one or more programs may include instructions that cause the electronic device to receive other user input for selecting at least a portion of the original image when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain the edited image comprising the at least portion of the original image among the at least portion of the original image and the remaining portion different from the at least portion of the original image, based on receiving the user input and the other user input when executed by the electronic device.

[0215] According to one embodiment, the one or more programs may include instructions that cause the electronic device to obtain the edited image by using a trained model to perform outpainting of the original image to extend the content of the original image along the direction from the first part to the second part when executed by the electronic device, and to perform inpainting of the original image so that the visual object is composited into the second part.

[0216] According to one embodiment, the one or more programs may include instructions that cause the electronic device to display the edited image through the display, based on acquiring the edited image when executed by the electronic device.

[0217] An electronic device as described above (e.g., electronic device (100)) may include a memory (e.g., memory (206)) for storing instructions. The electronic device may include a display (e.g., display (208)). The electronic device may include at least one processor (e.g., at least one processor (207)). The instructions may cause the electronic device to display an original image (e.g., original image (412)) through a first area of ​​the display when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to obtain a first edited image representing the original image, based on a first user input for adding one or more first strokes within the first area, and a first visual object represented by the one or more first strokes, based on the first user input for adding one or more first strokes within the first area. When the above instructions are executed individually or collectively by the at least one processor, based on a second user input for adding one or more second strokes (e.g., one or more strokes (414)) within a second area (e.g., a surrounding area (522)) adjacent to the first area, the electronic device may cause a second visual object (e.g., a visual object (426)) represented by the one or more second strokes to be synthesized and a second edited image including the original image to be obtained.

[0218] According to one embodiment, at least a portion of the original image may be located in a first portion of the second edited image corresponding to the first area. The second visual object may be located in a second portion of the second edited image corresponding to the second area.

[0219] According to one embodiment, the instructions may cause the electronic device to obtain a prompt for generating the second edited image by providing the original image and the one or more second strokes to a multimodal model based on receiving the second user input when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to obtain the second edited image by providing the prompt to a trained model when executed individually or collectively by the at least one processor.

[0220] A method performed by an electronic device (e.g., electronic device (100)) having a display (e.g., display (208)) as described above may include an operation of displaying an original image (e.g., original image (412)) through a first area of ​​the display. The method may include an operation of synthesizing a first visual object represented by one or more first strokes based on a first user input for adding one or more first strokes within the first area, and obtaining a first edited image representing the original image. The method may include an operation of synthesizing a second visual object (e.g., visual object (426)) represented by one or more second strokes based on a second user input for adding one or more second strokes (e.g., one or more strokes (414)) within a second area (e.g., surrounding area (522)) adjacent to the first area, and obtaining a second edited image (e.g., edited image (422)) including the original image.

[0221] According to one embodiment, at least a portion of the original image may be located in a first portion of the second edited image corresponding to the first area. The second visual object may be located in a second portion of the second edited image corresponding to the second area.

[0222] According to one embodiment, the method may include an operation of obtaining a prompt for generating the second edited image by providing the original image and the one or more second strokes to a multimodal model based on receiving the second user input. The method may include an operation of obtaining the second edited image by providing the prompt to a trained model.

[0223] In a computer-readable storage medium in which one or more programs are stored as described above, the one or more programs may include instructions that cause the electronic device (e.g., electronic device (100)) having a display (e.g., display (208)) to display an original image (e.g., original image (412)) through a first area of ​​the display. The one or more programs may include instructions that cause the electronic device to obtain a first edited image representing the original image, based on a first user input to add one or more first strokes within the first area, and to synthesize a first visual object represented by the one or more first strokes. The above one or more programs may include instructions that cause the electronic device to acquire a second edited image (e.g., edited image (422)) including the original image, based on a second user input for adding one or more second strokes (e.g., one or more strokes (414)) within a second area (e.g., surrounding area (522)) adjacent to the first area when executed by the electronic device.

[0224] According to one embodiment, at least a portion of the original image may be located in a first portion of the second edited image corresponding to the first area. The second visual object may be located in a second portion of the second edited image corresponding to the second area.

[0225] According to one embodiment, the one or more programs may include instructions that cause the electronic device to obtain a prompt for generating the second edited image by providing the original image and the one or more second strokes to a multimodal model based on receiving the second user input when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain the second edited image by providing the prompt to a trained model when executed by the electronic device.

[0226] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs.

[0227] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.

[0228] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or instruct the processing unit independently or collectively. Software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.

[0229] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may continuously store a computer-executable program, or temporarily store it for execution or download. Additionally, the medium may be various recording or storage means in the form of a single or several combined hardware, and may not be limited to a medium directly connected to a computer system but may exist distributed over a network. Examples of media may include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and media configured to store program instructions, including ROM, RAM, and flash memory. Additionally, other examples of media may include recording or storage media managed by app stores that distribute applications or sites and servers that supply or distribute various other software.

[0230] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results may be achieved even if the described techniques are performed in a different order than described, and / or the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.

[0231] Therefore, other implementations, other embodiments, and equivalents to the claims set forth below are also within the scope of the claims. According to one embodiment, the method according to the various embodiments disclosed herein may be provided as a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created in a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0232] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In an electronic device, display; Memory comprising one or more storage media for storing instructions; and It includes at least one processor comprising processing circuitry, and When the above instructions are executed individually or collectively by the at least one processor, Display the original image through the above display, and Receiving user input for adding one or more strokes on a surrounding area of ​​the area of ​​the display where the original image is displayed, and Based at least on receiving the above user input, to obtain an edited image comprising a first portion corresponding at least partially to the original image, and a second portion corresponding to the surrounding area and having a visual object represented by the one or more strokes composited therein. causing the above electronic device, Electronic device.

2. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, Based on receiving the above user input, the one or more strokes and the original image are provided to a trained model to obtain the edited image, causing the above electronic device, Electronic device.

3. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, By providing the original image and the one or more strokes to a multimodal model, a prompt for generating the edited image is obtained, and To obtain the above edited image using the above prompt, causing the above electronic device, Electronic device.

4. In claim 3, the prompt is, Instructing to composite another visual object related to the visual object expressed by the one or more strokes above into the second part, Electronic device.

5. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, By providing the above original image to a multimodal model, a prompt representing the features of the above original image is obtained as object recognition is performed, and Using the above prompt, to obtain the above edited image including other visual objects representing the above features, causing the above electronic device, Electronic device.

6. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, Using the above original image, identify the style of the above original image, and To obtain the edited image including the visual object expressed according to the above style, causing the above electronic device, Electronic device.

7. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, Identifying the current time based on the one or more strokes indicating to change the time represented by the original image above, and To obtain the edited image including the content of the original image changed according to the current time, causing the above electronic device, Electronic device.

8. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, Receiving other user input for determining the location of the visual object on the surrounding area, and Based on receiving the above other user input, the visual object obtains the edited image composited in the second part based on the above location, causing the above electronic device, Electronic device.

9. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, Based on receiving the above user input, display the one or more strokes on the surrounding area, and Receiving other user input for adding one or more other strokes, for at least some of the one or more strokes, and Based on receiving the other user input above, to acquire the edited image without including at least a portion of the visual object corresponding to at least a portion of the one or more strokes in the second portion, causing the above electronic device, Electronic device.

10. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, While displaying the original image above, receiving other user input indicating the addition of the surrounding area, and Based on receiving the other user input above, to display the surrounding area that at least partially encloses the original image together with the original image, causing the above electronic device, Electronic device.

11. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, Receiving other user input for selecting at least a portion of the original image, and Based on receiving the above user input and the above other user input, to obtain the edited image including the above at least part of the original image among the above at least part of the original image and the remaining part different from the above at least part of the original image, causing the above electronic device, Electronic device.

12. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, Using a trained model to perform outpainting of the original image to extend the content of the original image along the direction from the first part to the second part, and to perform inpainting of the original image so that the visual object is composited into the second part, to obtain the edited image, causing the above electronic device, Electronic device.

13. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, Based on obtaining the above edited image, to display the above edited image through the display, causing the above electronic device, Electronic device.

14. In an electronic device, display; Memory comprising one or more storage media for storing instructions; and It includes at least one processor comprising processing circuitry, and When the above instructions are executed individually or collectively by the at least one processor, The original image is displayed through the first area of ​​the display, and Based on a first user input for adding one or more first strokes within the first area, a first visual object represented by the one or more first strokes is synthesized, and a first edited image representing the original image is obtained, and Based on a second user input for adding one or more second strokes within a second area adjacent to the first area, a second visual object represented by the one or more second strokes is synthesized, and a second edited image including the original image is obtained. causing the above electronic device, Electronic device.

15. In claim 14, at least a portion of the original image is, Located in the first part of the second edited image corresponding to the first area, and The above second visual object is, Located in the second part of the second edited image corresponding to the second area, Electronic device.

Citation Information

Patent Citations

  • Image inpainting apparatus and method for thereof

    KR102486300B1

  • KR20230023437A