Electronic device, non-transitory computer-readable storage medium, and method for executing automation based on object recognition for image

WO2026192200A1PCT designated stage Publication Date: 2026-09-17SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/000834
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-31
Filing Date
2026-01-14
Publication Date
2026-09-17

Smart Images

  • Figure KR2026000834_17092026_PF_FP_ABST
    Figure KR2026000834_17092026_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device may comprise an image sensor, a display, a processor, and a memory storing instructions. The instructions, when individually or collectively executed by the processor, may cause the electronic device to: acquire an image including an object through the image sensor; display the image through the display; determine description information on the basis of characteristics of the object; determine an executable application in association with the description information; and generate and store a routine for executing the application corresponding to the description information.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device for executing automation based on object recognition of an image, non-transient computer-readable storage medium, and method

[0001] The following descriptions relate to an electronic device for executing automation based on object recognition of an image, a non-transient computer-readable storage medium, and a method.

[0002] An electronic device may provide an automation program for performing an operation corresponding to said condition based on the satisfaction of said condition. For example, a user of said electronic device may set conditions and operations for said automation using said automation program executed within said electronic device.

[0003] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.

[0004] An electronic device is provided. The electronic device may include at least one processor comprising an image sensor, a display, and a processing circuit, and a memory comprising one or more storage media for storing instructions. The instructions may cause the electronic device to acquire a first image comprising a first object through the image sensor when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to display the first image through the display when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to determine first description information based on at least some of the characteristics of the first object when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to determine a first application executable in association with the first description information when executed individually or collectively by the at least one processor. The above instructions may cause the electronic device to generate and store a first routine for executing the first application corresponding to the first description information when executed individually or collectively by the at least one processor.

[0005] A non-transient computer-readable storage medium is provided. The non-transient computer-readable storage medium may store one or more programs. The one or more programs may include instructions that cause the electronic device to acquire a first image including a first object through the image sensor when executed by the electronic device having an image sensor and a display. The one or more programs may include instructions that cause the electronic device to display the first image through the display when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to determine first description information based on at least some of the characteristics of the first object when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to determine a first application executable in association with the first description information when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to generate and store a first routine for executing the first application corresponding to the first description information when executed by the electronic device.

[0006] A method is provided. The method may be performed by an electronic device having an image sensor and a display. The method may include an operation of acquiring a first image including a first object through the image sensor. The method may include an operation of displaying the first image through the display. The method may include an operation of determining first description information based on at least some of the characteristics of the first object. The method may include an operation of determining a first application executable in association with the first description information. The method may include an operation of generating and storing a first routine for executing the first application corresponding to the first description information.

[0007] An electronic device is provided. The electronic device may include at least one processor comprising a display and a processing circuit, and a memory comprising one or more storage media for storing instructions. The instructions may cause the electronic device to display a first image through the display when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to receive user input through the electronic device for capturing the first image while the first image is displayed when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to determine tag information for at least one object of the first image in response to the user input when executed individually or collectively by the at least one processor. The above instructions may cause the electronic device to identify that the tag information for the first image corresponds to the tag information for at least one object of the second image, based on at least a portion of the determination of the tag information for the first image, when executed individually or collectively by the at least one processor. The second image may be captured to execute a function of a software application. The above instructions may cause the electronic device to display, through the display, a user interface (UI) object that executes the function of the software application using the first image, based on at least a portion of the identification of the tag information for the first image corresponding to the tag information for the second image, when executed individually or collectively by the at least one processor.

[0008] A non-transient computer-readable storage medium is provided. The non-transient computer-readable storage medium may store one or more programs. The one or more programs may include instructions that cause the electronic device to display a first image through the display when executed by the electronic device having a display. The one or more programs may include instructions that cause the electronic device to receive user input for capturing the first image through the electronic device while the first image is displayed when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to determine tag information for at least one object of the first image in response to the user input when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to identify, when executed by the electronic device, that the tag information for the first image corresponds to tag information for at least one object of the second image, based on at least a portion of the determination of the tag information for the first image. The second image may be captured to execute a function of a software application. The above one or more programs may include instructions that cause the electronic device to display, through the display, a user interface (UI) object that executes the function of the software application using the first image, based on at least a portion of the identification of the tag information for the first image corresponding to the tag information for the second image, when executed by the electronic device.

[0009] A method is provided. The method may be performed by an electronic device having a display. The method may include an operation of displaying a first image through the display. The method may include an operation of receiving user input through the electronic device for capturing the first image while the first image is displayed. The method may include an operation of determining tag information for at least one object among the objects of the first image in response to the user input. The method may include an operation of identifying that the tag information for the first image corresponds to tag information for at least one object of the second image based on at least a part of the determination of the tag information for the first image. The second image may be captured to execute a function of a software application. The method may include an operation of displaying a user interface (UI) object that executes the function of the software application using the first image through the display, based on at least a part of the identification of the tag information for the first image corresponding to the tag information for the second image.

[0010] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.

[0011] Figure 1 is a schematic view of an exemplary electronic device.

[0012] Figure 2 is a block diagram of an environment for executing the function of a routine for automation based on object recognition of an image.

[0013] Figure 3 is a flowchart illustrating a method for executing the function of a routine for automation based on object recognition of an image.

[0014] FIGS. 4a, FIGS. 4b, FIGS. 4c, and FIGS. 4d are examples of user interfaces provided to execute the function of a routine for automation based on user input for capturing an image acquired through a camera.

[0015] FIGS. 5A, FIGS. 5B, FIGS. 5C, and FIGS. 5D are drawings for explaining a method of generating a routine for automation based on object recognition of an image using an automation program executed within an electronic device.

[0016] Figures 6a and 6b are examples of user interfaces provided to guide the creation of routines for automation based on object recognition of images.

[0017] Figure 7 is a flowchart illustrating a method for executing the function of a routine identified by object recognition of an image among routines for automation.

[0018] FIGS. 8A, FIGS. 8B, and FIGS. 8C are examples of user interfaces provided to guide the execution of routines for automation based on object recognition of images.

[0019] FIG. 9 is a flowchart illustrating a method for selectively guiding the creation of a routine for automation or the execution of a routine for automation based on object recognition of an image and whether data for the routine for automation is stored.

[0020] FIGS. 10a, FIGS. 10b, FIGS. 10c, and FIGS. 10d are examples of user interfaces provided to guide routines for automation based on object recognition of images.

[0021] FIGS. 11a and FIGS. 11b are examples of user interfaces provided to execute the function of a routine for automation based on user input for capturing a screen displayed through the display of an electronic device.

[0022] FIGS. 12a and FIGS. 12b are examples of environments in which the functions of routines for the automation of external electronic devices are executed based on object recognition of images performed within the electronic device.

[0023] FIG. 13 is a block diagram of an electronic device in a network environment according to various embodiments.

[0024] FIG. 14 is a schematic diagram of an exemplary artificial intelligence (AI) system according to one embodiment.

[0025] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily practice them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and brevity.

[0026] Figure 1 is a schematic view of an exemplary electronic device.

[0027] Referring to FIG. 1, the electronic device (101) may include at least one processor (110), memory (120), camera (130), display (140), and at least one sensor (150). The electronic device (101) may include at least a part of the electronic device (1301) of FIG. 13 or correspond to at least a part of the electronic device (1301) of FIG. 13.

[0028] At least one processor (110) may include a processing circuit. At least one processor (110) may include a single processor or multiple processors. At least one processor (110) may control the memory (120) and / or one or more components (e.g., camera (130), display (140), and at least one sensor (150)) of the electronic device (101). For example, at least one processor (110) may include at least a part of the processor (1320) of FIG. 13 or correspond to at least a part of the processor (1320) of FIG. 13.

[0029] Memory (120) may store one or more programs configured to be executed individually and / or collectively by at least one processor (110). The one or more programs may include instructions. The instructions may cause an electronic device (101) to perform operations described with reference to FIGS. 2 through 12b. Memory (120) may include one or more storage media. At least some of the one or more programs may be available to manage, control, and / or execute an automation program for creating and executing routines for automation, as described below. For example, memory (120) may include at least a portion of the memory (1330) of FIG. 13 or correspond to at least a portion of the memory (1330) of FIG. 13.

[0030] The camera (130) can capture (or take) images (e.g., still images) and video. The camera (130) may include an image sensor (131), one or more lenses (not shown), and / or flashes (not shown). For example, the image sensor (131) can acquire an image corresponding to a subject by converting light emitted or reflected from the subject and transmitted through the lens into an electrical signal. According to one embodiment, the image sensor (131) may include one image sensor selected from image sensors with different properties, such as an RGB (red green blue) sensor, a BW (black and white) sensor, an IR (infrared) sensor, or a UV (ultraviolet) sensor, a plurality of image sensors having the same properties, or a plurality of image sensors having different properties. For example, each image sensor included in the image sensor (131) may be implemented using a CCD (charged coupled device) sensor or a CMOS (complementary metal oxide semiconductor) sensor. For example, the camera (130) may include at least a part of the camera module (1380) of FIG. 13 or correspond to at least a part of the camera module (1380) of FIG. 13.

[0031] A display (140) can visually provide information to an external (e.g., user) outside of an electronic device (101). For example, the display (140) may include a display panel and / or a touch sensor. For example, the display panel may be used to display visual information (e.g., images, screens, objects, UI (user interface), GUI (graphic user interface) and / or visual objects). For example, the display panel may have a display area capable of receiving touch input. For example, the touch sensor may be used to obtain data about an external object located on the display panel. For example, the touch sensor may be located within or on the display panel to provide an area of ​​the display panel capable of receiving the touch input. For example, the touch sensor may be configured to obtain data about contact points on at least a portion of the area. For example, the display (140) may include at least a part of the display module (1360) of FIG. 13 or correspond to at least a part of the display module (1360) of FIG. 13.

[0032] At least one sensor (150) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state or illuminance level) and generate an electrical signal or data value corresponding to the detected state. For example, at least one sensor (150) may include an illuminance sensor for identifying an illuminance level indicating the intensity of light reaching the electronic device (101). For example, at least one sensor (150) may include at least a part of the sensor module (1376) of FIG. 13 or correspond to at least a part of the sensor module (1376) of FIG. 13.

[0033] The electronic device (101) may provide an automation program that performs an operation of the routine corresponding to the condition based on the condition of the routine for automation being satisfied. For example, the routine may include a condition set in association with the routine and an operation of the electronic device (101) that is executed based on the condition being satisfied (e.g., a function of a software application).

[0034] For example, the operation of generating the routine may be described as an operation of generating routine data for the routine and storing the generated routine data. For example, the operation of executing the routine may be described as an operation of executing the function of the routine. For example, the routine data may include condition data representing the conditions of the routine for automation and operation data representing the operation of the routine for automation. For example, the routine data may represent the conditions of the routine and the operation of the routine (e.g., the function of a software application). For example, an electronic device (101) may generate routine data for the routine for automation using the automation program and store the routine data in memory (120). For example, the electronic device (101) may execute the function of the routine for automation by using the routine data stored in memory (120) and performing the operation of the routine corresponding to the condition based on the satisfaction of the condition of the routine.

[0035] For example, an electronic device (101) can determine tag information (or description information) for an object in an image based on object recognition of an image displayed through a display (140). For example, the electronic device (101) can set an image tag condition corresponding to the tag information as a condition for a routine for automation, and set the function of a software application to be executed based on the satisfaction of the condition as the operation of the routine. For example, the electronic device (101) can further maximize the convenience of the user utilizing the routine for automation by executing the function of the routine for automation based on object recognition of the image. A method of operation for creating and / or executing a routine for automation based on object recognition of the image is described below with reference to FIGS. 2 to FIGS. 12b.

[0036] Figure 2 is a block diagram of an environment for executing the function of a routine for automation based on object recognition of an image.

[0037] Referring to FIG. 2, an environment (200) for executing the function of a routine for automation based on object recognition of an image is illustrated. The environment (200) may include a routine generation control unit (210), an image analysis unit (220), a routine execution control unit (230), a routine estimation unit (240), and a database (DB) (250). For example, the database (250) may include at least a portion of the memory (120) of FIG. 1 or correspond to at least a portion of the memory (120) of FIG. 1. For example, the database (250) may store and manage routine data representing a routine for automation. For example, the routine data may include name information of the routine for automation, condition data representing the condition of the routine for automation, and operation data representing the operation of the routine for automation.

[0038] The electronic device (101) can manage and / or control the creation of routines for automation using the routine creation control unit (210). For example, the electronic device (101) can generate routine data based on user input (e.g., image input) entered through the user interface of the automation program using the routine creation control unit (210). For example, the routine data may include name information of the routine for automation, condition data indicating the conditions of the routine for automation, and operation data indicating the operation of the routine for automation. For example, the electronic device (101) can store the routine data in a database (250) using the routine creation control unit (210).

[0039] The electronic device (101) can determine tag information (or description information) for an image based on user input (e.g., image input) by using an image analysis unit (220). For example, the tag information may include tag information for at least one object among the objects of the image (e.g., keyword of the object). For example, the tag information for the object may be associated with the type, attribute, and / or category of the object. For example, the electronic device (101) can determine tag information for an object in an image based on computer vision recognition of the content of the image (e.g., object recognition, scene analysis, pattern classification, and / or optical character recognition (OCR)) by using an image analysis unit (220). In one embodiment, the tag information may include tag information indicating the location of the electronic device (101) when the image is captured. In one embodiment, the tag information may include tag information indicating the connection status (e.g., electrical connection status or communication connection status) between the electronic device (101) and an external electronic device when an image is captured. In one embodiment, the tag information may include tag information indicating the illuminance level identified through at least one sensor (150) when an image is captured.

[0040] The electronic device (101) can manage and / or control the execution of a routine for automation using a routine execution control unit (230). For example, the electronic device (101) can use the routine execution control unit (230) to identify user input for capturing the image while the image is displayed through the display (140). For example, the electronic device (101) can use the routine execution control unit (230) to identify a routine among the routines for automation represented by routine data stored in the database (250) that satisfies a condition corresponding to the tag information for the image, based on the user input for capturing the image. For example, the electronic device (101) can use the routine execution control unit (230) to display a UI (user interface) object for performing an action represented by the routine data for the identified routine through the display (140).

[0041] The electronic device (101) can determine (or estimate) a routine for automation based on tag information for the image while the image is displayed through the display (140) by using the routine estimation unit (240). For example, the electronic device (101) can display a UI (user interface) object for executing the routine determined (or estimated) based on the tag information for the image displayed through the display (140) by using the routine estimation unit (240) through the display (140).

[0042] Figure 3 is a flowchart illustrating a method for executing the function of a routine for automation based on object recognition of an image.

[0043] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.

[0044] Referring to FIG. 3, according to one embodiment, in operation 301, at least one processor (110) may display a first image including at least one object through a display (140). For example, the first image may be described as a preview image obtained through an image sensor (131) of a camera (130). For example, the first image may be described as an image corresponding to at least a portion of a screen displayed through the display (140). By example, without limitation, the first image may correspond to at least a portion of an execution screen of a software application running on an electronic device (101).

[0045] According to one embodiment, in operation 302, at least one processor (110) may receive user input for capturing the first image through an electronic device (101) while the first image is displayed. For example, the user input for capturing may be described as an input for capturing the first image using a camera software application while the first image is displayed. For example, the user input for capturing may be described as an input for capturing a screen containing the first image while the first image is displayed (e.g., an input for a screenshot).

[0046] According to one embodiment, in operation 303, at least one processor (110) may determine tag information (or description information) for at least one object among the objects of the first image based on the user input for capturing the first image. For example, at least one processor (110) may determine tag information for the first image based on the characteristics of the object of the first image. For example, at least one processor (110) may determine tag information for at least one object among the objects of the image by performing computer vision recognition (e.g., object recognition, scene analysis, pattern classification, and / or optical character recognition (OCR)) on the content of the image.

[0047] In one embodiment, the tag information may be associated with user information (or personal information) stored in memory (120). For example, at least one processor (110) may determine tag information for at least one object associated with the user information among the objects of the first image. By example, without limitation, the object associated with the user information may be described as an object associated with user pattern information, schedule information, device information, personal identification information, location information, internet usage information, social media information, financial information, and / or health information.

[0048] According to one embodiment, in operation 304, at least one processor (110) can identify, based on the determination of the tag information for the first image, that the tag information for the first image corresponds to tag information for at least one object of the second image captured prior to the first image.

[0049] For example, the second image may be described as an image captured prior to the first image to execute a function (e.g., image transmission) of a software application (e.g., a message software application) executed according to user input. For example, at least one processor (110) may display the second image through a display (140) before the first image is captured. For example, the second image may be described as a preview image obtained through an image sensor (131) of a camera (130). For example, the second image may be described as an image corresponding to at least a portion of a screen displayed through the display (140). By example, without limitation, the second image may correspond to at least a portion of an execution screen of a software application executed on the electronic device (101). For example, at least one processor (110) may receive user input to capture the second image through the electronic device (101) while the second image is being displayed. For example, at least one processor (110) can determine tag information (or description information) for at least one object among the objects of the second image based on the user input for capturing the second image. For example, at least one processor (110) can determine tag information for the object of the second image based on the characteristics of the object of the second image.

[0050] According to one embodiment, in operation 305, at least one processor (110) may display, through a display (140), a UI (user interface) object for executing the function of the software application executed using the second image using the first image, based on the identification of the tag information for the first image corresponding to the tag information for the second image. For example, at least one processor (110) may execute the function of a routine for automation corresponding to the tag information for the first image by executing the function of the software application using the first image based on user input to the UI object. For example, the UI object may be displayed based on user input to allow the electronic device (101) to execute the function of the routine.

[0051] In one embodiment, at least one processor (110) can identify that a condition represented by routine data for a routine stored in memory (120) is satisfied based on the determination of tag information for the first image. For example, the routine data may represent a condition that tag information for at least one object of the image to be captured corresponds to the tag information for the second image, and a function of a software application executed based on the satisfaction of the condition. For example, at least one processor (110) may display, through a display (140), the UI object executing the function of the software application represented by the routine data based on the satisfaction of the condition.

[0052] In one embodiment, at least one processor (110) may receive user input for capturing the second image through an electronic device (101) while the second image is displayed on a display (140). For example, at least one processor (110) may determine tag information for at least one object of the second image based on the user input for capturing the second image. For example, at least one processor (110) may display through the display (140) a UI object that generates routine data to be stored in memory (120) based on the determination of the tag information for the second image. For example, the routine data may set a condition corresponding to the tag information for the second image as an image tag condition of the routine. For example, the image tag condition may be such that the tag information for at least one object of the image corresponds to the tag information for at least one object of the second image.

[0053] In one embodiment, at least one processor (110) may display at least one UI object through the display (140) which is used to select at least one object of the second image while the second image is displayed on the display (140). For example, the at least one UI object may correspond to the at least one object of the second image. For example, while the at least one UI object is displayed, at least one processor (110) may receive user input for one or more of the at least one UI objects through the electronic device (101). For example, based on the user input for the one or more UI objects, at least one processor (110) may determine tag information for one or more objects corresponding to the one or more of the at least one objects of the second image as tag information for the second image.

[0054] In one embodiment, at least one processor (110) can execute the function of the software application using the first image based on user input for the UI object executing the function of the software application using the routine data. For example, the routine data may represent one or more conditions for displaying the UI object executing the function of the software application using the image to be captured. For example, the one or more conditions may include an image tag condition in which tag information for at least one object of the image corresponds to tag information for at least one object of the second image.

[0055] In one embodiment, the routine data may be generated based on the execution of an automation program using the second image. For example, the automation program may be described as a program for generating routine data representing an operation performed through an electronic device (101) based on a condition set in association with the routine and the satisfaction of said condition. A description of the automation program will be given later with reference to FIGS. 5a, 5b, and 5c.

[0056] In one embodiment, the one or more conditions may include a position condition in which the electronic device (101) is located within a defined area associated with the tag information for the second image when a user input for capturing the first image is received. For example, the defined area may be described as an area including the location of the electronic device (101) identified when a user input for capturing the second image is received while the second image is displayed through the display (140).

[0057] In one embodiment, the one or more conditions may include an illuminance condition in which an illuminance level identified by at least one sensor (150) when a user input for capturing the first image is received corresponds to a defined range (e.g., a range below a defined threshold value) associated with the tag information for the second image. For example, the defined range may be described as a range including the illuminance level identified by at least one sensor (150) when a user input for capturing the second image is received while the second image is displayed through the display (140).

[0058] In one embodiment, the one or more conditions may include a condition that a time (a time of day) identified by the electronic device (101) when a user input for capturing the first image is received is included in a time interval defined in association with the tag information for the second image. For example, the defined time interval may be described as a time interval including a time identified by the electronic device (101) when a user input for capturing the second image is received while the second image is displayed through the display (140).

[0059] In one embodiment, the condition of the routine for automation may be set in natural language generated using tag information for an object in an image (e.g., "if [#car number] is included and [#building column] contains a number or an alphabet"). For example, at least one processor (110) may use a large language model (LLM) to determine whether the condition is satisfied. For example, at least one processor (110) may generate a prompt for the LLM using tag information, condition data indicating the condition of the routine, and an image based on user input, and determine whether the condition is satisfied by applying the generated prompt to the LLM.

[0060] LLM refers to an artificial neural network-based language model that has been trained on a large amount of text data through prior learning. LLMs can include significantly more parameters (e.g., over 10 billion) than general language models. As a type of machine learning model used in the field of natural language processing, LLMs can be utilized to perform predictions on new text data by learning from massive amounts of text data. LLMs can be applied to tasks such as natural language understanding, sentence generation, translation, grammar correction, and summarization.

[0061] In one embodiment, at least one processor (110) can set and execute a routine for automation linked to a software application (e.g., a calendar software application). For example, at least one processor (110) can create a routine to transmit an image to an external electronic device owned by friend B when a user captures an image at a location corresponding to region C during period A (from January 2 to January 7) if a schedule such as 'travel to region C with friend B during period A' is registered in the calendar software application. For example, the routine can be executed during period A.

[0062] FIGS. 4a, FIGS. 4b, FIGS. 4c, and FIGS. 4d are examples of user interfaces provided to execute the function of a routine for automation based on user input for capturing an image acquired through a camera.

[0063] Referring to FIG. 4a, a user interface (400a) is shown that is provided to generate the routine based on user input for capturing an image (411a) acquired through a camera (130) before the routine for automation using the image (411b) shown in FIG. 4b is executed.

[0064] In the display state (410a), the electronic device (101) can display, through the display (140), an execution screen of a camera software application including an image (411a) acquired through the camera (130). The image (411a) can be described as a preview image displayed through the display (140) while being acquired through the camera (130). For example, the electronic device (101) can display, through the display (140), a UI object (412a) for capturing the image (411a) displayed through the display (140). For example, the electronic device (101) can store the image (411a) in memory (120) in association with a gallery software application based on user input regarding the UI object (412a). For example, the electronic device (101) can display a UI object (413a) for displaying an execution screen of the gallery software application including an image (411a) through a display (140). For example, the electronic device (101) can switch the display state (410a) to a display state (420a) based on user input regarding the UI object (413a).

[0065] In a display state (420a), the electronic device (101) can display, through the display (140), the execution screen of the gallery software application, including an image (411a) captured through the camera (130), based on the user input to the UI object (413a). For example, the image (411a) may include an object (421a), an object (422a), and an object (423a). For example, the object (421a) may correspond to text indicating the parking location of a vehicle (e.g., A-18). For example, the object (422a) may correspond to text indicating a vehicle registration number (VRN) (e.g., 12A 1234). For example, the object (423a) may correspond to a vehicle. The electronic device (101) can determine tag information for an object (421a) (e.g., A Tower), tag information for an object (422a) (e.g., 12A 1234), and tag information for an object (423a) (e.g., parking) based on user input for capturing an image (411a). For example, the electronic device (101) can display a UI object (424a) through a display (140) to execute a function of another software application distinct from the gallery software application using an image (411a) captured through a camera (130). For example, the electronic device (101) can switch a display state (420a) to a display state (430a) based on user input for the UI object (424a).

[0066] In the display state (430a), the electronic device (101) can display an image (411a) captured through the camera (130) via the display (140) based on the user input for the UI object (424a). The electronic device (101) can display, through a display (140), a UI object (431a) for executing a function (e.g., image transmission) of a first software application (e.g., data sharing software application) using an image (411a), a UI object (432a) for executing a function (e.g., image transmission) of a second software application (e.g., message software application) using an image (411a), a UI object (433a) for executing a function (e.g., image registration) of a third software application (e.g., wallet software application) using an image (411a), and a UI object (434a) for executing a function (e.g., notification setting) of a fourth software application (e.g., reminder software application) using an image (411a).

[0067] For example, the electronic device (101) can generate routine data for a routine to execute the function of the first software application using an image (411a) to be captured (e.g., image (411b) of FIG. 4b) by executing the function of the first software application using an image (411a) based on user input for a UI object (431a), and can store the routine data in memory (120) (e.g., database (250)). For example, the electronic device (101) can generate routine data for a routine to execute the function of the second software application using an image (411a) to be captured (e.g., image (411b) of FIG. 4b) by executing the function of the second software application using an image (411a) based on user input for a UI object (432a), and can store the routine data in memory (120) (e.g., database (250)). For example, the electronic device (101) can generate routine data for a routine to execute the function of the third software application using an image (411a) to be captured (e.g., image (411b) of FIG. 4b) by executing the function of the third software application using an image (411a) based on user input for a UI object (433a), and can store the routine data in memory (120) (e.g., database (250)). For example, the electronic device (101) can generate routine data for a routine to execute the function of the fourth software application using an image (411a) to be captured (e.g., image (411b) of FIG. 4b) by executing the function of the fourth software application using an image (411a) based on user input for a UI object (434a), and can store the routine data in memory (120) (e.g., database (250)).

[0068] Referring to FIG. 4b, a user interface (400b) is shown that is provided to execute the function of a routine for automation based on user input for capturing an image (411b) acquired through a camera (130) after the function of a software application using an image (411a) is executed in the user interface (400a) shown in FIG. 4a.

[0069] In the display state (410b), the electronic device (101) can display, through the display (140), an execution screen of a camera software application including an image (411b) acquired through the camera (130). The image (411b) can be described as a preview image displayed through the display (140) while being acquired through the camera (130). For example, the electronic device (101) can display, through the display (140), a UI object (412b) for capturing the image (411b) displayed through the display (140). For example, the electronic device (101) can store the image (411b) in memory (120) in association with a gallery software application based on user input regarding the UI object (412b). For example, the electronic device (101) can display a UI object (413b) for displaying an execution screen of the gallery software application including an image (411b) through a display (140). For example, the electronic device (101) can switch the display state (410b) to a display state (420b) based on user input regarding the UI object (413b).

[0070] In the display state (420b), the electronic device (101) can display, through the display (140), the execution screen of the gallery software application, including an image (411b) captured through the camera (130), based on the user input for the UI object (413b). For example, the image (411b) may include an object (421b), an object (422b), and an object (423b). For example, the object (421b) may correspond to text indicating the parking location of a vehicle (e.g., A-19). For example, the object (422b) may correspond to text indicating a vehicle registration number (VRN) (e.g., 12A 1234). For example, the object (423b) may correspond to a vehicle. The electronic device (101) can determine tag information for an object (421b) (e.g., A Tower), tag information for an object (422b) (e.g., 12A 1234), and tag information for an object (423b) (e.g., parking) based on user input for capturing an image (411b). For example, the electronic device (101) can display a container (or snack bar) (424b) for executing the function of a routine for automation through the display (140) based on identifying that the tag information for the objects (421b, 422b, 423b) of the image (411b) corresponds to the tag information for the objects (421a, 422a, 423a) of the image (411a). The container (424b) may include a UI object (425b) and text (426b). The UI object (425b) may be used to execute the function of a software application executed using the image (411a) using the image (411b).Text (426b) may be described as text representing the name of a routine for executing the function of the software application (e.g., parking location sending routine). For example, text (426b) may be described as text obtained from a large language model (LLM). For example, the electronic device (101) may execute a function of the software application using an image (411b) (e.g., image transmission) based on user input to a UI object (425b). For example, the electronic device (101) may refrain from displaying the container (424b) after a defined time has elapsed since displaying the container (424b) through the display (140).

[0071] According to one embodiment, the function of the routine may include one or more functions. For example, the one or more functions may be executed through a single software application or different software applications. For example, the one or more functions may include a series of functions of software applications, such as sending a message and setting an alarm sound (or vibration pattern). For example, the one or more functions may be performed sequentially or in parallel.

[0072] According to one embodiment, routine data stored in memory (120) may be stored as metadata for the image. For example, when the electronic device (101) transmits the image to an external electronic device, the routine data may be transmitted to the external electronic device along with the image. For example, when the electronic device (101) transmits the image to an external electronic device through a program such as a message software application, an email software application, or QuickShare, the external electronic device may display a notification message such as “Would you like to save the routine stored with the image?” upon receiving the image. For example, the external electronic device may save the routine based on user input for saving the routine.

[0073] According to one embodiment, with reference to user interface (400a) and user interface (400b), the electronic device (101) can generate a first routine and a second routine using a single image.

[0074] For example, the electronic device (101) may set the first routine to a function of a software application (e.g., sending a message, sending a voice message, or sending a social network service (SNS) message) for sharing information about a parking location with another user (e.g., family) based on user input. For example, the electronic device (101) may set each of the functions of the software applications in different formats according to user input. For example, the electronic device (101) may perform the operation of transmitting an image two or more times using multiple software applications. For example, the electronic device (101) may identify map information corresponding to parking lot information where a vehicle is parked using a map software application based on GPS (global positioning system) information stored in the image, and transmit at least a portion of the map information and the image to an external electronic device. For example, the electronic device (101) can display summarized content generated through the artificial intelligence (AI) briefing function of the electronic device (101) based on text information related to the first routine, the map information, or at least part of the image through the display (140) and transmit it to external electronic devices of other people (e.g., family) registered in the user's contacts. For example, each of the external electronic devices can display the location where the vehicle is parked in the form of a pin using a map software application based on user input regarding the received map information. For example, each of the external electronic devices can store the first routine based on user input regarding the first routine.

[0075] For example, the electronic device (101) may set the second routine to store GPS information included in the metadata of an image captured through the camera (130), and to display information about the image and the parking location through the display (140) when the user returns to a location corresponding to the GPS information.

[0076] According to one embodiment, the electronic device (101) can disable a function associated with at least one other routine based on the execution of one of the plurality of routines when a plurality of routines are set.

[0077] According to one embodiment, when the electronic device (101) sets a routine for sharing an image with an external electronic device or a routine for reminding an image through the electronic device (101), the execution of the routine can be set as a condition of another routine.

[0078] According to one embodiment, the electronic device (101) can execute the function of a routine linked to the artificial intelligence (AI) briefing function of the electronic device (101) by identifying the user's schedule using an artificial intelligence (AI) model. For example, when the user captures an image related to parking, the electronic device (101) can display information obtained through the AI ​​briefing function using the user's schedule, parking location, and the image, based on the user's schedule registered in the electronic device (101), through a display (140). For example, the obtained information may include map information corresponding to the location for the schedule, information regarding the schedule, map information corresponding to the parking location, at least a portion of the image, expected route information for moving from the parking location to the location for the schedule, expected travel time information, and / or expected arrival time information.

[0079] Referring to FIG. 4c, a user interface (400c) is shown that is provided to execute the function of a routine for automation based on user input to allow the execution of the routine for automation after the function of a software application using an image (411a) in the user interface (400a) shown in FIG. 4a has been executed.

[0080] In the display state (410c), the electronic device (101) can display the execution screen of a gallery software application, including an image (411b) stored in memory (120), through the display (140). The electronic device (101) can identify that tag information for objects (421b, 422b, 423b) of the image (411b) corresponds to tag information for objects (421a, 422a, 423a) of the image (411a). While the image (411b) is being displayed, the electronic device (101) can receive a swipe input (412c) according to a specified direction (e.g., downward).

[0081] In a display state (420c), the electronic device (101) can display a notification panel screen through the display (140) based on a swipe input (412c). For example, the electronic device (101) can display a notification message (421c) to guide the execution of a routine for automation within the notification panel screen based on identifying that tag information for objects (421b, 422b, 423b) of image (411b) corresponds to tag information for objects (421a, 422a, 423a) of image (411a). The electronic device (101) can switch the display state (420c) to a display state (420b) based on user input (e.g., touch input) for the notification message (421c). For example, the user input for the notification message (421c) can be described as a user input that allows the execution of a function of a software application executed using an image (411a) based on routine data for a routine for automation using an image (411b).

[0082] In the display state (420b), the electronic device (101) can display, through the display (140), a container (or snack bar) (424b) for executing the function of a routine for automation based on the user input for the notification message (421c).

[0083] Referring to FIG. 4d, a user interface (400d) is shown to guide the execution of a routine for automation after the function of a software application using an image (411a) is executed in the user interface (400a) shown in FIG. 4a.

[0084] In the display state (420b), the electronic device (101) can display, through the display (140), an execution screen of a gallery software application including an image (411b) captured through the camera (130). For example, the image (411b) may include an object (421b), an object (422b), and an object (423b). The electronic device (101) can display, through the display (140), a container (or snack bar) (424b) for executing the function of a routine for automation based on identifying that tag information for the objects (421b, 422b, 423b) of the image (411b) corresponds to tag information for the objects (421a, 422a, 423a) of the image (411a). The container (424b) may include a UI object (425b) and text (426b). The electronic device (101) may switch the display state (420b) to the display state (420d) based on user input for an area corresponding to the text (426b).

[0085] In a display state (420d), the electronic device (101) may display a pop-up window (421d) through the display (140) to guide the execution of a routine for automation based on the user input (e.g., touch input) for the area corresponding to the text (426b). The pop-up window (421d) may include an image (422d) and text (423d). For example, the image (422d) may include at least a part of the object (421b), at least a part of the object (422b), and at least a part of the object (423b). For example, the text (423d) may include tag information (or description information) for at least one of the objects (421b, 422b, 423b) for the image (411b), tag information indicating the location of the electronic device (101) when the image (411b) is captured, and guide information for guiding the execution of a routine for automation.

[0086] FIGS. 5A, FIGS. 5B, FIGS. 5C, and FIGS. 5D are drawings for explaining a method of generating a routine for automation based on object recognition of an image using an automation program executed within an electronic device.

[0087] Referring to FIG. 5a, a flowchart is illustrated to explain a method for generating a routine for automation based on object recognition of an image using an automation program executed within an electronic device (101).

[0088] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.

[0089] According to one embodiment, in operation 501, at least one processor (110) may display a user interface for generating routine data for a routine for automation through a display (140). For example, the routine data may include condition data indicating the conditions of the routine for automation and operation data indicating the operation of the routine for automation. For example, the user interface may be described as a user interface of an automation program for setting the conditions of the routine for automation and the operation of the routine for automation. For example, at least one processor (110) may execute the function of the routine for automation by using the routine data stored in memory (120) to perform the operation of the routine corresponding to the condition based on the satisfaction of the condition of the routine. For example, the automation program may be described as a program for an operation performed through an electronic device (101) based on the satisfaction of the condition indicated by the routine data. For example, the routine data may be generated using the automation program. For example, the routine data may represent one or more conditions for the function of a software application to be executed using an image to be captured through the camera (130). For example, the one or more conditions may include an image tag condition, a location condition, a time condition, and an illumination condition.

[0090] According to one embodiment, in operation 502, at least one processor (110) may receive user input through an electronic device (101) to set a condition of a routine as an image tag condition while the user interface is displayed. For example, the image tag condition may include a condition that tag information for an image corresponds to tag information represented by routine data stored in memory (120).

[0091] According to one embodiment, in operation 503, at least one processor (110) may set a condition corresponding to tag information (or description information) for at least one object among the objects of an image as the image tag condition based on the user input for setting the condition of the routine as the image tag condition.

[0092] According to one embodiment, in operation 504, at least one processor (110) may receive user input through an electronic device (101) for setting at least one other condition of the routine that is distinct from the image tag condition. In one embodiment, the at least one other condition may include a position condition that the electronic device (101) is located within an area defined in association with the tag information for the image. In one embodiment, the at least one other condition may include an illuminance condition that an illuminance level identified through at least one sensor (150) corresponds to a range defined in association with the tag information for the image. In one embodiment, the at least one other condition may include a time condition that a time (a time of day) identified by the electronic device (101) falls within a time interval defined in association with the tag information for the image. For example, at least one processor (110) may generate condition data for the image tag condition and the at least one other condition.

[0093] According to one embodiment, in operation 505, at least one processor (110) may receive user input for setting the operation of the routine through an electronic device (101). For example, at least one processor (110) may generate operation data based on the user input for setting the operation.

[0094] According to one embodiment, in operation 506, at least one processor (110) may generate routine data for a routine to be stored in memory (120) in association with the automation program. For example, the routine data may include the condition data and the operation data. For example, at least one processor (110) may store the routine data in memory (120) (e.g., database (250)).

[0095] Referring to FIG. 5b, a user interface (500b) is shown that generates condition data for a routine for automation based on object recognition of an image using an automation program executed within an electronic device (101).

[0096] In the display state (510), the electronic device (101) can display, through the display (140), an execution screen of an automation program for generating routine data for a routine for automation. For example, the automation program may be described as a program for performing an action through the electronic device (101) based on the satisfaction of a condition. The electronic device (101) can display a UI object (511) for setting conditions for a routine for automation and a UI object (512) for setting actions for a routine for automation through the display (140). For example, the electronic device (101) can switch the display state (510) to the display state (520) based on user input for the UI object (511).

[0097] In the display state (520), the electronic device (101) may display a user interface including a list (521), a list (522), and a list (523) that can be set for conditions of a routine for automation based on the user input to the UI object (511). For example, the list (521) may be used to set a time condition as a condition for a routine for automation. For example, the time condition may be set as a condition for the time when the user of the electronic device (101) wakes up, the time before the user goes to sleep, the time when the user is driving, or a time set according to user input. For example, the list (522) may be used to set a communication condition for the electronic device (101) as a condition for a routine for automation. For example, the above communication condition may be set as a condition regarding the connection status of WiFi (wireless fidelity) of the electronic device (101), the connection status of Bluetooth of the electronic device (101), whether the airplane mode of the electronic device (101) is enabled, or whether the hot spot of the electronic device (101) is enabled. For example, the list (523) may be used to set the image tag condition as a condition for a routine for automation. For example, the list (523) may include a UI object (524), a UI object (525), and a UI object (526). For example, the UI object (524) may be used to set tag information for one or more objects determined by user input among at least one object of an image acquired through the camera (130) as tag information for the image tag condition.For example, a UI object (525) may be used to set tag information for one or more objects determined by user input among at least one object of an image stored in memory (120) in association with a gallery software application as tag information for the image tag condition. For example, a UI object (526) may be used to set the image tag condition according to user input for determining tag information for one or more objects. For example, an electronic device (101) may switch the display state (520) to the display state (530) based on user input regarding the UI object (526).

[0098] In the display state (530), the electronic device (101) may display a user interface for setting image tag conditions according to user input for determining tag information for one or more objects based on the user input for the UI object (511). For example, in the user interface of the display state (530), image tag conditions including tag information (531) corresponding to a vehicle, tag information (532) corresponding to a vehicle registration number (VRN), and tag information (533) corresponding to a parking lot may be determined according to user input. For example, in the user interface of the display state (530), the electronic device (101) may display a UI object (534) for adding additional tag information to the image tag conditions.

[0099] Referring to FIG. 5c, a user interface (500c) is shown that generates operation data for a routine for automation based on object recognition of an image using an automation program executed within an electronic device (101).

[0100] For example, the electronic device (101) can switch the display state (510) to the display state (540) based on user input to the UI object (512).

[0101] In the display state (540), the electronic device (101) may display UI objects (541), UI objects (542), UI objects (543), UI objects (544), UI objects (545), UI objects (546), and UI objects (547) through the display (140) to set the operation of a routine for automation based on the user input for the UI object (512). The UI object (541) may be used to set the operation of a routine for automation to execute a software application set according to user input or to execute a function of said software application. The UI object (542) may be used to set the operation of a routine for automation to terminate the execution of a software application set according to user input. The UI object (543) may be used to set the operation of a routine for automation to execute a software application set according to user input. A UI object (544) can be used to set the action of turning on an alarm at a time set according to user input as an action of a routine for automation. A UI object (545) can be used to set the action of turning off an alarm set according to user input as an action of a routine for automation. A UI object (546) can be used to set the action of turning on a timer (e.g., a down-count timer) corresponding to a time set according to user input as an action of a routine for automation. A UI object (547) can be used to set the action of turning on a stopwatch (e.g., an up-town timer) as an action of a routine for automation.

[0102] Referring to FIG. 5d, a user interface (500d) is shown that displays a routine for automation generated using an automation program executed within an electronic device (101).

[0103] In a display state (550), the electronic device (101) may set conditions for a routine for automation in the user interface (500b) shown in FIG. 5b, set the operation of said routine based on user input for a UI object (544) in the user interface (500c) shown in FIG. 5c, and then display a UI object (551) for said routine through the display (140). For example, the UI object (551) may be used to display a user interface indicating the conditions of said routine and the settings of said operation of said routine. For example, said routine may include the function of an alarm software application that turns on an alarm at a time set according to user input based on an image tag condition set according to user input and the satisfaction of said image tag condition.

[0104] Figures 6a and 6b are examples of user interfaces provided to guide the creation of routines for automation based on object recognition of images.

[0105] Referring to FIG. 6a, a user interface (600a) is shown to guide the creation of a routine for automation of at least one object included in an image displayed through a display (140).

[0106] In the display state (610a), the electronic device (101) can display an execution screen of a gallery software application including an image (611a) through the display (140). The image (611a) may include an object (612a) corresponding to text (e.g., A-18) indicating the parking location of a vehicle. The electronic device (101) can display a UI object (613a) used to select the object (612a) through the display (140).

[0107] In a display state (620a), the electronic device (101) may display a visual effect (621a) for an object (612a) through the display (140) based on user input (e.g., touch input) for a UI object (613a). The electronic device (101) may display a container (or snack bar) (622a) through the display (140) for generating routine data for a routine for automation corresponding to the object (612a) based on the user input for the UI object (613a). The container (622a) may include a UI object (623a) and text (624a). The UI object (623a) may be used to generate routine data for a routine for automation corresponding to the object (612a). For example, the routine data may include condition data indicating a condition in which the image tag condition of the routine for automation corresponds to tag information for the object (612a). Text (624a) may be described as text for guiding the generation of routine data for the routine including the condition data. For example, the electronic device (101) may refrain from displaying the container (622a) after a defined time has elapsed since displaying the container (622a) through the display (140).

[0108] In a display state (630a), the electronic device (101) may display a pop-up window (631a) through the display (140) to guide the creation of a routine for automation based on user input for a UI object (623a). The pop-up window (631a) may include an image (632a), text (633a), and a UI object (634a). For example, the image (632a) may include at least a portion of an object (612a). For example, the text (633a) may be described as text representing condition data for a routine for automation and operation data for a routine for automation, including image tag conditions and location conditions. For example, the UI object (634a) may be used to create routine data for the routine, including the condition data and the operation data.

[0109] According to one embodiment, with reference to the user interface (600a), the electronic device (101) may display a user interface (e.g., a user interface of the display state (510) shown in FIG. 5b, a user interface of the display state (520), and / or a user interface of the display state (530) for setting an image tag condition as a condition of a routine based on user input for a UI object (623a). For example, a candidate for a routine displayed through the user interface based on the user input for the UI object (623a) may be determined from a plurality of routines for routine data stored in memory (120) that include a condition corresponding to tag information for an object (612a). For example, the electronic device (101) may set the routine determined by the user input among the one or more routines as a routine generated based on the user input for the UI object (623a).

[0110] According to one embodiment, with reference to the user interface (600a), the electronic device (101) may display an image acquired through the camera (130) (e.g., a captured image) or an image stored in association with a gallery software application through the display (140). For example, the electronic device (101) may display a user interface that generates a routine in relation to at least one object included in the image. For example, the user interface that generates the routine may be displayed in an area corresponding to the quick view screen displayed when the image is captured. For example, the user interface that generates the routine may be displayed based on the thumbnail or another button distinct from the thumbnail being pressed.

[0111] According to one embodiment, with reference to the user interface (600a), the electronic device (101) may display another user interface that includes at least a portion of at least one object of an image based on user input for generating a routine. For example, the at least portion of the at least one object displayed within the other user interface may be described as at least a portion of an actual image, a pre-stored image, or an image generated using generative artificial intelligence (AI). For example, the electronic device (101) may display tag information related to the image and / or information used for the routine in text form below the image included in the other user interface. For example, the tag information included in the other user interface may be described as tag information generated based on recognition of the object included in the image or tag information based on user input. For example, the electronic device (101) may display guide information associated with the routine using the tag information in the other user interface. For example, the electronic device (101) may display the guide information set according to user input within the other user interface. For example, the electronic device (101) may display the guide information in the form of text of a routine generated from generative artificial intelligence (AI) within the other user interface.

[0112] Referring to FIG. 6b, a user interface (600b) is shown to guide the creation of a routine for automating multiple objects included in an image displayed through a display (140).

[0113] In the display state (610b), the electronic device (101) can display an execution screen of a gallery software application including an image (611b) through the display (140). The image (611b) may include an object (612b), an object (613b), and an object (614b). For example, the object (612b) may correspond to text indicating the parking location of a vehicle (e.g., A-18). For example, the object (613b) may correspond to text indicating a vehicle registration number (VRN) (e.g., 12A 1234). For example, the object (614b) may correspond to a vehicle. The electronic device (101) can display a UI object (615b) used to select an object (612b), a UI object (616b) used to select an object (613b), and a UI object (617b) used to select an object (614b) through a display (140).

[0114] In the display state (620b), the electronic device (101) can display a visual effect (621b) for an object (612b) through the display (140) based on user input (e.g., touch input) for a UI object (615b). The electronic device (101) can display a visual effect (622b) for an object (613b) through the display (140) based on user input (e.g., touch input) for a UI object (616b). The electronic device (101) can display a container (or snack bar) (623b) for generating routine data for a routine for automation corresponding to the object (612b) and the object (613b) through the display (140) based on the user input for the UI object (615b) and the user input for the UI object (616b). The container (623b) may include a UI object (624b) and text (625b). The UI object (624b) may be used to generate routine data for a routine for automation corresponding to the object (612b) and the object (613b). For example, the routine data may include condition data indicating a condition in which the image tag condition of the routine for automation corresponds to tag information for the object (612b) and the object (613b). The text (625b) may be described as text for guiding the generation of routine data for the routine including the condition data. For example, the electronic device (101) may refrain from displaying the container (623b) after a defined time has elapsed since displaying the container (623b) through the display (140).

[0115] In a display state (630b), the electronic device (101) may display a pop-up window (631b) through the display (140) to guide the creation of a routine for automation based on user input for a UI object (624b). The pop-up window (631b) may include an image (632b), text (633b), and a UI object (634b). For example, the image (632b) may include at least a part of an object (612b), at least a part of an object (613b), and at least a part of an object (614b). For example, the text (633b) may be described as text representing condition data for a routine for automation and operation data for a routine for automation, including image tag conditions and location conditions. For example, the UI object (634b) may be used to create routine data for the routine, including the condition data and the operation data.

[0116] According to one embodiment, with reference to the user interface (600b), the electronic device (101) may display an image acquired through the camera (130) (e.g., a captured image) or an image stored in association with a gallery software application through the display (140). The image may include a plurality of objects, including a first object and a second object. For example, when a plurality of objects are displayed in the image, the electronic device (101) may display a recommendation routine for each of the plurality of objects or a recommendation routine for all of the plurality of objects. For example, the electronic device (101) may select each of the plurality of objects based on user input for generating a routine and display a recommendation routine for the selected object among the plurality of objects. For example, the electronic device (101) may display a recommendation routine by displaying another user interface related to the generation of the routine based on user input for the selected object. For example, the other user interface may include a first object, a second object, an object of interest among one or more objects selected according to user input, and / or an AI object obtained by applying the first object and the second object to generative artificial intelligence (AI).

[0117] Although not illustrated in the user interface (600b), according to one embodiment, the electronic device (101) may recognize each of a plurality of objects in a preview image in real time based on user input for generating a routine or execution of the camera (130), and may display a recommendation routine through a user selection for each of the plurality of objects, or may automatically display a recommendation routine without a user selection. For example, if another object is included in the preview image as the angle of view of the camera (130) changes, the electronic device (101) may display a recommendation routine for the object included in the preview image before the angle of view changes and the other object included in the preview image after the angle of view changes. For example, if a user photographs a vehicle and the license plate of the vehicle through the camera (130) and photographs the number of a pillar next to the vehicle as the angle of view of the camera (130) changes, the electronic device (101) may automatically display a recommendation routine for the vehicle, the license plate, and the pillar number, or may generate the recommendation routine according to user input. For example, the electronic device (101) may execute the function of the recommendation routine when the user sequentially photographs the vehicle, the license plate, and the pillar number through the camera (130) after the recommendation routine is generated, or when the user photographs the vehicle, the license plate, and the pillar number simultaneously within the field of view of the camera (130).

[0118] Although not shown in the user interface (600b), according to one embodiment, the routine for automation may be recommended by artificial intelligence (AI) using data (e.g., data on captured images and data on previously stored routines) based on the user's personal preferences, or may be displayed as a recommendation routine obtained from a server connected to the electronic device (101).

[0119] Although not illustrated in the user interface (600b), according to one embodiment, the setting of the routine may be performed on a wearable device (e.g., a video see-through (VST) device). For example, the electronic device (101) may be implemented as a device that identifies the user's gaze, such as a VST device or an augmented reality (AR) glasses device. For example, the electronic device (101) may use a camera that tracks the user's gaze to identify whether the user gazes at a specific object in a preview image displayed through a display or a specific object in the real world visible through the display for more than a specified amount of time. For example, the electronic device (101) may display a user interface to determine whether to create a routine based on the identification, or provide information about the routine to the user through an acoustic output module (e.g., a speaker) and / or a haptic module (e.g., vibration).

[0120] In one embodiment, the electronic device (101) may display the user interface (600b) through a display (e.g., a display of a VST device) or output information about the user interface (600b) through a speaker based on user input for generating a routine through an input module (e.g., a microphone) or a display.

[0121] Figure 7 is a flowchart illustrating a method for executing the function of a routine identified by object recognition of an image among routines for automation.

[0122] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.

[0123] Referring to FIG. 7, according to one embodiment, in operation 701, at least one processor (110) may display an image through a display (140). For example, the image may include at least one object. For example, the image may be described as a preview image obtained through an image sensor (131) of a camera (130). For example, the image may be described as an image corresponding to at least a portion of a screen displayed through the display (140). By example, without limitation, the image may correspond to at least a portion of an execution screen of a software application running within an electronic device (101).

[0124] According to one embodiment, in operation 702, at least one processor (110) may receive user input for capturing the image through an electronic device (101) while the image is being displayed. For example, the user input for capturing may be described as an input for capturing a still image and / or video through a camera (130). For example, the user input for capturing may be described as an input for capturing a screen displayed through a display (140) (e.g., an input for a screenshot). At least one processor (110) may determine tag information (or description information) for at least one object among the objects of the image based on the user input for capturing the image. For example, at least one processor (110) may determine tag information for the object of the image based on the characteristics of the object of the image.

[0125] According to one embodiment, in operation 703, at least one processor (110) can identify a routine corresponding to the tag information for the image among routines for automation represented by routine data stored in memory (120) (e.g., database (250)), based on the determination of the tag information for the at least one object of the image.

[0126] According to one embodiment, in operation 704, at least one processor (110) can display, through a display (140), a UI (user interface) object that executes the function of a software application represented by the routine data, using routine data for the identified routine.

[0127] According to one embodiment, in operation 705, at least one processor (110) can execute the function of the software application using the image based on user input for the UI object.

[0128] FIGS. 8A, FIGS. 8B, and FIGS. 8C are examples of user interfaces provided to guide the execution of routines for automation based on object recognition of images.

[0129] Referring to FIG. 8a, a user interface (800a) is shown to guide the execution of a routine for automation of one object included in an image displayed through a display (140).

[0130] In the display state (810a), the electronic device (101) can display, through the display (140), an execution screen of a camera software application including an image (811a) acquired through the camera (130). The image (811a) can be described as a preview image displayed through the display (140) while being acquired through the camera (130). The image (811a) may include an object (812a) corresponding to text (e.g., A-19) indicating the parking location of a vehicle. For example, the electronic device (101) can display a UI object (813a) through the display (140) to capture the image (811a) displayed through the display (140). For example, the electronic device (101) can store the image (811a) in memory (120) in association with a gallery software application based on user input for the UI object (813a). For example, the electronic device (101) can display a UI object (814a) for displaying an execution screen of the gallery software application including an image (811a) through a display (140). For example, the electronic device (101) can switch a display state (810a) to a display state (820a) based on user input regarding the UI object (814a).

[0131] In the display state (820a), the electronic device (101) can display, through the display (140), the execution screen of the gallery software application including an image (811a) captured through the camera (130) based on the user input for the UI object (814a). The electronic device (101) can determine tag information (e.g., A Tower) for the object (812a) based on the user input for capturing the image (811a). The electronic device (101) can display, through the display (140), a container (or snack bar) (821a) for executing the function of the routine using routine data for the automation routine corresponding to the object (812a). The container (821a) may include a UI object (822a) and text (823a). A UI object (822a) can be used to execute a function of a software application represented by routine data corresponding to an object (812a) using an image (811a). For example, the routine data may include condition data representing a condition in which an image tag condition of a routine for automation corresponds to tag information for the object (812a). Text (823a) may be described as text representing the name of a routine for executing the function of the software application (e.g., parking alarm routine). For example, text (823a) may be described as text obtained from a large language model (LLM). For example, an electronic device (101) may execute a function of a software application using an image (811a) (e.g., alarm activation) based on user input to the UI object (822a). For example, the electronic device (101) may refrain from displaying the container (821a) after a defined time has elapsed since displaying the container (821a) through the display (140).

[0132] According to one embodiment, with reference to the user interface (800a), the electronic device (101) may display an image acquired through the camera (130) (e.g., a captured image) or an image stored in association with a gallery software application through the display (140). For example, the electronic device (101) may display a user interface that executes a routine function in relation to at least one object included in the image. For example, the user interface that executes the routine function may be displayed in an area corresponding to the quick view screen displayed when the image is captured. For example, the user interface that executes the routine function may be displayed based on the thumbnail or another button distinct from the thumbnail being pressed.

[0133] According to one embodiment, with reference to the user interface (800a), the electronic device (101) may display another user interface including at least a portion of at least one object of an image based on user input for executing a function of a routine. For example, the at least portion of the at least one object displayed within the other user interface may be described as at least a portion of an actual image, a pre-stored image, or an image generated using generative artificial intelligence (AI). For example, the electronic device (101) may display tag information related to the image and / or information used for the routine in text form below the image included in the other user interface. For example, the tag information included in the other user interface may be described as tag information generated based on recognition of the object included in the image or tag information based on user input. For example, the electronic device (101) may display guide information associated with the routine using the tag information in the other user interface. For example, the electronic device (101) may display the guide information set according to user input within the other user interface. For example, the electronic device (101) may display the guide information in the form of text of a routine generated from generative artificial intelligence (AI) within the other user interface.

[0134] Referring to FIG. 8b, a user interface (800b) is shown to guide the execution and creation of a routine for automation of multiple objects included in an image displayed through a display (140).

[0135] In the display state (810b), the electronic device (101) can display, through the display (140), an execution screen of a camera software application including an image (811b) acquired through the camera (130). The image (811b) can be described as a preview image displayed through the display (140) while being acquired through the camera (130). The image (811b) may include an object (812b), an object (813b), and an object (814b). The object (812b) may correspond to text indicating the parking location of a vehicle (e.g., A-19). The object (813b) may correspond to text indicating a vehicle registration number (VRN) (e.g., 12A 1234). The object (814b) may correspond to a vehicle. For example, the electronic device (101) may display a UI object (815b) for capturing an image (811b) displayed through the display (140) via the display (140). For example, the electronic device (101) may store the image (811b) in memory (120) in association with a gallery software application based on user input regarding the UI object (815b). For example, the electronic device (101) may display a UI object (816b) for displaying an execution screen of the gallery software application containing the image (811b) via the display (140). For example, the electronic device (101) may switch the display state (810b) to a display state (820b) based on user input regarding the UI object (816b).

[0136] In the display state (820b), the electronic device (101) can display the execution screen of the gallery software application, including an image (811b) captured through the camera (130) based on the user input for the UI object (816b), through the display (140). The electronic device (101) can determine tag information for an object (812b) (e.g., A Tower), tag information for an object (813b) (e.g., 12A 1234), and tag information for an object (814b) (e.g., parking) based on the user input for capturing the image (811b). For example, the electronic device (101) can display, through the display (140), a container (or snack bar) (821b) for executing the function of said routine using routine data for an automation routine corresponding to an object (812b), an object (813b), and an object (814b).

[0137] For example, the container (821b) may include a UI object (822b), text (823b), a UI object (824b), and text (825b). The UI object (822b) may be used to execute a function of a software application represented by routine data for a first routine corresponding to an object (812b), an object (813b), and an object (814b). For example, the routine data for the first routine may include condition data representing a condition in which an image tag condition of the first routine corresponds to tag information for the object (812b), the object (813b), and the object (814b). The text (823b) may be described as text representing the name of the first routine (e.g., parking alarm routine) for executing the function of the software application. For example, the electronic device (101) can execute a function of a software application using an image (811b) (e.g., alarm activation) based on user input to a UI object (822b). The UI object (824b) can be used to generate routine data for a second routine corresponding to objects (812b), objects (813b), and objects (814b). For example, the routine data for the second routine may include condition data indicating conditions where the image tag conditions of the second routine correspond to tag information for objects (812b), objects (813b), and objects (814b). Text (825b) may be described as text indicating the name of the second routine (e.g., parking location sending routine). For example, each of text (823b) and text (825b) may be described as text obtained from a large language model (LLM).For example, the electronic device (101) may refrain from displaying the container (821b) after a defined time has elapsed since displaying the container (821b) through the display (140).

[0138] Referring to FIG. 8c, a user interface (800c) is shown to guide the execution of a routine for automating multiple objects included in an image displayed through a display (140).

[0139] In the display state (810b), the electronic device (101) can display, through the display (140), an execution screen of a camera software application including an image (811b) obtained through the camera (130).

[0140] In the display state (820c), the electronic device (101) can display, through the display (140), an execution screen of a gallery software application including an image (811b) captured through the camera (130) based on the user input for the UI object (816b). The electronic device (101) can display a visual effect (823c) in an area (821c) corresponding to the object (812b) and a visual effect (824c) in an area (822c) corresponding to the object (813b) based on the identification of routine data corresponding to the object (812b) and the object (813b). For example, the electronic device (101) can display, through the display (140), a container (or snack bar) (821b) for executing the function of said routine using routine data for an automation routine corresponding to the object (812b) and the object (813b). The container (821b) may include a UI object (822b) and text (823b). The UI object (822b) may be used to execute the operation of a software application represented by routine data corresponding to the object (812b) and the object (813b). For example, the routine data may include condition data representing a condition in which an image tag condition of a routine for automation corresponds to tag information for the object (812b) and the object (813b).

[0141] According to one embodiment, with reference to the user interface (800c), the electronic device (101) may display an image acquired through the camera (130) (e.g., a captured image) or an image stored in association with a gallery software application through the display (140). The image may include a plurality of objects, including a first object and a second object. For example, when a plurality of objects are displayed in the image, the electronic device (101) may display a recommendation routine for each of the plurality of objects or a recommendation routine for all of the plurality of objects. For example, the electronic device (101) may select each of the plurality of objects based on user input for executing the function of the routine and display a recommendation routine for the selected object among the plurality of objects. For example, the electronic device (101) may display the recommendation routine by displaying another user interface related to the execution of the routine based on user input for the selected object. For example, the other user interface may include a first object, a second object, an object of interest among one or more objects selected according to user input, and / or an AI object obtained by applying the first object and the second object to generative artificial intelligence (AI).

[0142] Although not illustrated in the user interface (800c), according to one embodiment, the electronic device (101) may recognize each of a plurality of objects in real time in a preview image based on user input or execution of the camera (130) for executing the function of a routine, and may display a recommendation routine through a user selection for each of the plurality of objects, or may display a recommendation routine automatically without a user selection. For example, if another object is included in the preview image as the angle of view of the camera (130) changes, the electronic device (101) may display a recommendation routine for the object included in the preview image before the angle of view changed and the other object included in the preview image after the angle of view changed. For example, the electronic device (101) may automatically display a recommendation routine for the vehicle, the license plate, and the pillar number when the user photographs the vehicle and the license plate of the vehicle through the camera (130) and photographs the pillar number of the pillar next to the vehicle as the field of view of the camera (130) changes, or may generate the recommendation routine according to user input. For example, after the recommendation routine is generated, the electronic device (101) may execute the function of the recommendation routine when the user photographs the vehicle, the license plate, and the pillar number sequentially through the camera (130), or may execute the function of the recommendation routine when the vehicle, the license plate, and the pillar number are photographed simultaneously within the field of view of the camera (130).

[0143] Although not shown in the user interface (800c), according to one embodiment, the routine for automation may be recommended by artificial intelligence (AI) using data (e.g., data on captured images and data on previously stored routines) based on the user's personal preferences, or may be displayed as a recommendation routine obtained from a server connected to the electronic device (101).

[0144] Although not illustrated in the user interface (800c), according to one embodiment, the setting of the routine may be performed on a wearable device (e.g., a video see-through (VST) device). For example, the electronic device (101) may be implemented as a device that identifies the user's gaze, such as a VST device or an augmented reality (AR) glasses device. For example, the electronic device (101) may use a camera that tracks the user's gaze to identify whether the user gazes at a specific object in a preview image displayed through a display or a specific object in the real world visible through the display for more than a specified amount of time. For example, the electronic device (101) may display a user interface to determine whether to execute the function of the routine based on the identification, or provide information about the routine to the user through an acoustic output module (e.g., a speaker) and / or a haptic module (e.g., vibration).

[0145] In one embodiment, the electronic device (101) may display a user interface (800c) through a display (e.g., a display of a VST device) or output information about the user interface (800c) through a speaker based on user input for executing a routine function through an input module (e.g., a microphone) or a display.

[0146] FIG. 9 is a flowchart illustrating a method for selectively guiding the creation of a routine for automation or the execution of a routine for automation based on object recognition of an image and whether data for the routine for automation is stored.

[0147] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.

[0148] Referring to FIG. 9, according to one embodiment, in operation 901, at least one processor (110) can acquire an image to be displayed through a display (140). For example, at least one processor (110) can acquire the image through an image sensor (131) of a camera (130). For example, at least one processor (110) can acquire the image stored in memory (120). For example, at least one processor (110) can acquire the image from an external electronic device (e.g., a server (1308)).

[0149] According to one embodiment, in operation 902, at least one processor (110) may determine tag information (or description information) of the object in the image based on the characteristics of the object in the image. For example, tag information for the object may be described as a keyword (or label) of the object. For example, a keyword for the object may be associated with the type, characteristics, and / or category of the object. For example, at least one processor (110) may determine tag information for at least one object among the objects in the image by performing computer vision recognition (e.g., object recognition, scene analysis, pattern classification, and / or optical character recognition (OCR)) on the content of the image.

[0150] According to one embodiment, in operation 903, at least one processor (110) may acquire location information of the electronic device (101), connection information of the electronic device (101), and sensing data while the image is acquired. At least one processor (110) may determine tag information (or description information) for the image corresponding to the location information, tag information (or description information) for the image corresponding to the connection information, and tag information (or description information) for the image corresponding to the sensing data. For example, the location information may indicate the location of the electronic device (101). For example, at least one processor (110) may acquire the location information based on a global positioning system (GPS) signal received through the electronic device (101). For example, the connection information may indicate an electrical connection state or a communication connection state between an external electronic device and the electronic device (101). For example, the sensing data may be acquired through at least one sensor (150). For example, the sensing data may include data representing an illuminance level identified through at least one sensor (150).

[0151] According to one embodiment, in operation 904, at least one processor (110) can identify whether routine data for a routine corresponding to tag information for the image is stored in memory (120) (e.g., database (250)).

[0152] According to one embodiment, in operation 905, at least one processor (110) can display, through a display (140), a UI object that generates routine data to be stored in memory (120) based on the fact that the routine data corresponding to tag information for the image is not stored in memory (120).

[0153] According to one embodiment, in operation 906, at least one processor (110) can display, through a display (140), a UI object that executes the function of a software application represented by the routine data, based on the routine data corresponding to tag information for the image being stored in memory (120).

[0154] In one embodiment, at least one processor (110) may acquire a first image including a first object through an image sensor (131). For example, at least one processor (110) may display the first image through a display (140). For example, at least one processor (110) may determine first tag information (or first description information) based on the characteristics of the first object. For example, at least one processor (110) may determine a first software application executable in association with the first tag information. For example, at least one processor (110) may generate and store a first routine for executing the first software application corresponding to the first tag information. For example, at least one processor (110) may display the at least one graphic effect for the first object through the display (140) while the first image is displayed. For example, at least one processor (110) may display a popup window through a display (140) that includes at least a portion of the first object, the first tag information, and a graphic object corresponding to the first application after the determination of the first software application. For example, at least one processor (110) may generate and store the first routine for executing the first software application based on user input regarding the graphic object included in the popup window. For example, at least one processor (110) may refrain from displaying the popup window after a defined time has elapsed since the popup window was displayed.

[0155] In one embodiment, at least one processor (110) may acquire a second image including a second object through an image sensor (131). For example, at least one processor (110) may display the second image through a display (140). For example, at least one processor (110) may determine second tag information (or second description information) based on the characteristics of the second object. For example, at least one processor (110) may execute the first software application using the first routine based on a judgment that the second tag information corresponds to the first tag information while the second image is displayed. For example, at least one processor (110) may determine a second software application executable in association with the second tag information based on a judgment that the second tag information does not correspond to the first tag information while the second image is displayed, and may create and store a second routine for executing the second software application corresponding to the second tag information. For example, at least one processor (110) may display a UI (user interface) object corresponding to the first routine through a display (140) based on the determination that the second tag information corresponds to the first tag information while the second image is displayed. For example, at least one processor (110) may execute the first software application using the first routine based on user input regarding the UI object.

[0156] In one embodiment, at least one processor (110) may acquire a third image including a third object and a fourth object through an image sensor (131). For example, at least one processor (110) may display the third image through a display (140). For example, at least one processor (110) may determine third tag information (or third description information) based on the characteristics of the third object and fourth tag information (or fourth description information) based on the characteristics of the fourth object. For example, at least one processor (110) may display a UI object corresponding to the first routine in a first area of ​​the display (140) based on a determination that the third tag information corresponds to the first tag information while the third image is being displayed. For example, at least one processor (110) may execute the first software application using the first routine based on user input regarding the UI object corresponding to the first routine. For example, at least one processor (110) may display a UI object corresponding to the second routine in a second area of ​​the display (140) based on the determination that the fourth tag information corresponds to the second tag information while the third image is being displayed. For example, at least one processor (110) may execute the second software application using the second routine based on user input regarding the UI object corresponding to the second routine.

[0157] In one embodiment, at least one processor (110) may receive user input regarding the third object and the fourth object within the third image while the third image is displayed. For example, at least one processor (110) may display through a display (140) a UI object that generates a third routine for executing a third software application that is executable in association with the third tag information and the fourth tag information, based on the user input regarding the third object and the fourth object, the third tag information, and the fourth tag information. At least one processor (110) may generate and store the third routine for executing the third software application corresponding to the third tag information and the fourth tag information based on the user input regarding the UI object for generating the third routine.

[0158] In one embodiment, the at least one processor (110) can perform the decision of the first software application based on history information indicating that the first software application was executed after acquiring the fourth image associated with the first tag information.

[0159] FIGS. 10a, FIGS. 10b, FIGS. 10c, and FIGS. 10d are examples of user interfaces provided to guide routines for automation based on object recognition of images.

[0160] Referring to FIG. 10a, a user interface (1000a) is shown to guide the execution of routines for automation of one object included in an image displayed through a display (140).

[0161] In the display state (1010a), the electronic device (101) can display, through the display (140), an execution screen of a camera software application including an image (1011a) obtained through the camera (130). The image (1011a) may include an object (1012a) corresponding to text (e.g., A-19) indicating the parking location of a vehicle. For example, the electronic device (101) can display, through the display (140), a UI object (1013a) for capturing the image (1011a) displayed through the display (140). For example, the electronic device (101) can display, through the display (140), a UI object (1014a) for displaying an execution screen of the gallery software application including the image (1011a). For example, the electronic device (101) can switch the display state (1010a) to the display state (1020a) based on user input to the UI object (1014a).

[0162] In a display state (1020a), the electronic device (101) can display, through the display (140), the execution screen of the gallery software application including an image (1011a) captured through the camera (130) based on the user input for the UI object (1014a). The electronic device (101) can determine tag information (e.g., A Tower) for an object (1012a) based on the user input for capturing the image (1011a). The electronic device (101) can display, through the display (140), a container (or snack bar) (1021a) for executing the function of said routine using routine data for an automation routine corresponding to the object (1012a). A container (1021a) may include a UI object (1022a), text (1023a), a UI object (1024a), and text (1025a). Each of the UI object (1022a) and the UI object (1024a) may be used to execute different routines. For example, the UI object (1022a) may be used to execute the function of a first routine, and the UI object (1024a) may be used to execute the function of a second routine. For example, the text (1023a) may be described as text representing the name of the first routine (e.g., parking alarm routine). For example, the text (1025a) may be described as text representing the name of the second routine (e.g., parking location sending routine).

[0163] Referring to FIG. 10b, a user interface (1000b) is shown to guide the execution of routines for automation of multiple objects included in an image displayed through a display (140).

[0164] In the display state (1010b), the electronic device (101) can display, through the display (140), an execution screen of a camera software application including an image (1011b) acquired through the camera (130). The image (1011b) may include an object (1012b), an object (1013b), and an object (1014b). The object (1012b) may correspond to text indicating the parking location of a vehicle (e.g., A-19). The object (1013b) may correspond to text indicating a vehicle registration number (VRN) (e.g., 12A 1234). The object (1014b) may correspond to a vehicle. For example, the electronic device (101) may display a UI object (1015b) for capturing an image (1011b) displayed through the display (140) via the display (140). For example, the electronic device (101) may display a UI object (1016b) for displaying an execution screen of the gallery software application including the image (1011b) via the display (140). For example, the electronic device (101) may switch the display state (1010b) to the display state (1020b) based on user input regarding the UI object (1016b).

[0165] In a display state (1020b), the electronic device (101) can display, through the display (140), the execution screen of the gallery software application including an image (1011b) captured through the camera (130) based on the user input for the UI object (1016b). The electronic device (101) can determine tag information for an object (1012b) (e.g., A Tower), tag information for an object (1013b) (e.g., 12A 1234), and tag information for an object (1014b) (e.g., parking) based on the user input for capturing the image (1011b). The electronic device (101) may display a visual effect (1023b) in an area (1021b) corresponding to an object (1012b) and a visual effect (1024b) in an area (1022b) corresponding to an object (1013b), based on the identification of routine data corresponding to an object (1012b) and an object (1013b). For example, the electronic device (101) may display a container (or snack bar) (1025b) for executing the function of said routine using routine data for an automation routine corresponding to an object (1012b) and an object (1013b) through a display (140). The container (1025b) may include a UI object (1026b), text (1027b), a UI object (1028b), and text (1029b). UI object (1026b) and UI object (1028b) can each be used to execute different routines. For example, UI object (1026b) can be used to execute the function of a first routine, and UI object (1028b) can be used to execute the function of a second routine. For example, text (1027b) can be described as text representing the name of the first routine (e.g., parking alarm routine).For example, text (1029b) can be described as text representing the name of the second routine (e.g., parking location sending routine).

[0166] Referring to FIG. 10c, a user interface (1000c) is shown to guide the execution of routines for automation of multiple objects included in an image displayed through a display (140).

[0167] In a display state (1020b), the electronic device (101) can display the execution screen of the gallery software application, including an image (1011b) captured through the camera (130), through the display (140). The electronic device (101) can switch the display state (1020b) to a display state (1030c) based on user input (e.g., touch input) for an object (1014b).

[0168] In a display state (1030c), the electronic device (101) may display a visual effect (1032c) in an area (1031c) corresponding to the object (1014b) based on the user input (e.g., touch input) for the object (1014b). For example, the electronic device (101) may display, through the display (140), a container (or snack bar) (1033c) for executing the function of said routine using routine data for the object (1012b), the object (1013b), and the routine for automation corresponding to the object (1014b), based on the user input (e.g., touch input) for the object (1014b). The container (1033c) may include a UI object (1028b) and text (1029b).

[0169] Referring to FIG. 10d, a user interface (1000d) is shown to guide the execution of routines for automation of multiple objects included in an image displayed through a display (140).

[0170] In the display state (1010b), the electronic device (101) can display, through the display (140), an execution screen of a camera software application including an image (1011b) obtained through the camera (130).

[0171] In a display state (1020d), the electronic device (101) can display, through the display (140), the execution screen of the gallery software application, including an image (1011b) captured through the camera (130), based on the user input for the UI object (1016b). The electronic device (101) can display a visual effect (1022d) in an area (1021d) corresponding to the object (1012b) based on the identification of routine data corresponding to the object (1012b). For example, the electronic device (101) can display, through the display (140), a container (or snack bar) (1023d) for executing the function of the routine using routine data for the automation routine corresponding to the object (1012b). The container (1023d) may include a UI object (1024d) and text (1025d). For example, a UI object (1024d) may be used to execute the function of the first routine. For example, text (1025d) may be described as text representing the name of the first routine (e.g., parking location sending routine).

[0172] In a display state (1030d), the electronic device (101) may display a visual effect (1032d) in an area (1031d) corresponding to the object (1013b) based on user input (e.g., touch input) for the object (1013b). For example, the electronic device (101) may display a container (or snack bar) (1033d) for executing the function of said routine using routine data for the object (1012b) and the routine for automation corresponding to the object (1013b) through the display (140). The container (1033d) may include a UI object (1034d) and text (1035d). For example, the UI object (1034d) may be used to execute the function of the second routine. For example, text (1035d) may be described as text representing the name of the second routine (e.g., parking alarm routine).

[0173] FIGS. 11a and FIGS. 11b are examples of user interfaces provided to execute the function of a routine for automation based on user input for capturing a screen displayed through the display of an electronic device.

[0174] Referring to FIG. 11a, a user interface (1100a) is shown that is provided to generate the routine based on user input for capturing an image (1111a) included in the execution screen of a software application before the routine for automation using the image (1111b) shown in FIG. 11b is executed.

[0175] In a display state (1110a), the electronic device (101) can display an execution screen of a software application including an image (1111a) through a display (140). For example, the electronic device (101) can switch the display state (1110a) to a display state (1120a) based on user input for capturing the execution screen (e.g., input for a screenshot).

[0176] In the display state (1120a), the electronic device (101) can determine tag information (e.g., barcode) for an object (1121a) included in the image (1111a) and tag information (e.g., coupon) for an object (1122a) included in the image (1111a) based on the user input for capturing the execution screen. The electronic device (101) can display a UI object (1123a) through the display (140) to execute a function of another software application distinct from the software application using the captured image (1111a) based on the user input for capturing the execution screen. For example, the electronic device (101) can switch the display state (1120a) to the display state (1130a) based on the user input for the UI object (1123a).

[0177] In the display state (1130a), the electronic device (101) can display the captured image (1111a) through the display (140) based on the user input for the UI object (1123a). The electronic device (101) can display, through a display (140), a UI object (1131a) for executing a function (e.g., image transmission) of a first software application (e.g., data sharing software application) using an image (1111a), a UI object (1132a) for executing a function (e.g., image transmission) of a second software application (e.g., message software application) using an image (1111a), a UI object (1133a) for executing a function (e.g., image registration) of a third software application (e.g., wallet software application) using an image (1111a), and a UI object (1134a) for executing a function (e.g., notification setting) of a fourth software application (e.g., reminder software application) using an image (1111a).

[0178] For example, the electronic device (101) can generate routine data for a routine to execute the function of the first software application using an image (1111a) to be captured (e.g., image (1111b) of FIG. 11b) by executing the function of the first software application using an image (1111a) based on user input for the UI object (1131a), and can store the routine data in memory (120) (e.g., database (250)). For example, the electronic device (101) can generate routine data for a routine to execute the function of the second software application using an image (1111a) to be captured (e.g., image (1111b) of FIG. 11b) by executing the function of the second software application using an image (1111a) based on user input for the UI object (1132a), and can store the routine data in memory (120) (e.g., database (250)). For example, the electronic device (101) can generate routine data for a routine to execute the function of the third software application using an image (1111a) to be captured (e.g., image (1111b) of FIG. 11b) by executing the function of the third software application using an image (1111a) based on user input for a UI object (1133a), and can store the routine data in memory (120) (e.g., database (250)). For example, the electronic device (101) can generate routine data for a routine to execute the function of the fourth software application using an image (1111a) to be captured (e.g., image (1111b) of FIG. 11b) by executing the function of the fourth software application using an image (1111a) based on user input for a UI object (1134a), and can store the routine data in memory (120) (e.g., database (250)).

[0179] Referring to FIG. 11b, a user interface (1100b) is shown to be provided for executing a routine for automation based on user input to capture an image (1111b) included in the execution screen of a software application after the function of a software application using an image (1111a) in the user interface (1100a) shown in FIG. 11a has been executed.

[0180] In a display state (1110b), the electronic device (101) can display an execution screen of a software application including an image (1111b) through a display (140). For example, the image (1111b) may include an object (1112b) and an object (1113b). For example, the electronic device (101) can switch the display state (1110b) to a display state (1120b) based on user input for capturing the execution screen (e.g., input for a screenshot).

[0181] In a display state (1120b), the electronic device (101) can determine tag information (e.g., barcode) for an object (1112b) included in an image (1111b) and tag information (e.g., coupon) for an object (1113b) included in an image (1111b) based on the user input for capturing the execution screen. For example, the electronic device (101) can display a container (or snack bar) (1121b) for executing the function of a routine for automation through the display (140) based on identifying that the tag information for the objects (1112b, 1113b) of the image (1111b) corresponds to the tag information for the objects (1112a, 1113a) of the image (1111a). A container (1121b) may include a UI object (1122b) and text (1123b). The UI object (1122b) may be used to execute a function of a software application executed using an image (1111a) using the image (1111b). The text (1123b) may be described as text representing the name of a routine for executing the function of the software application (e.g., coupon adding routine). For example, the text (1123b) may be described as text obtained from a large language model (LLM). For example, an electronic device (101) may execute a function (e.g., image registration) of a software application (e.g., wallet software application) using the image (1111b) based on user input to the UI object (1122b). For example, the electronic device (101) may refrain from displaying the container (1121b) after a defined time has elapsed since displaying the container (1121b) through the display (140).

[0182] According to one embodiment, routine data stored in memory (120) may be stored as metadata for the image. For example, when the electronic device (101) transmits the image to an external electronic device, the routine data may be transmitted to the external electronic device along with the image. For example, when the electronic device (101) transmits the image to an external electronic device via a message software application, an email software application, or a QuickShare program, the external electronic device may display a notification message such as “Would you like to save the routine stored with the image?” upon receiving the image. For example, the external electronic device may save the routine based on user input for saving the routine.

[0183] According to one embodiment, with reference to user interface (1100a) and user interface (1100b), the electronic device (101) can generate a plurality of different routines using a single image.

[0184] According to one embodiment, the electronic device (101) can set a routine for the function of a software application (e.g., message transmission, voice message transmission, or SNS message transmission) for sharing an image with another user (e.g., family) according to user input. For example, the electronic device (101) can set each of the functions of the software applications in different formats according to user input. For example, the electronic device (101) can perform the operation of transmitting an image two or more times using a plurality of software applications. For example, each of the external electronic devices can store the routine based on user input for the routine.

[0185] According to one embodiment, the electronic device (101) can disable a function associated with at least one other routine based on the execution of one of the plurality of routines when a plurality of routines are set.

[0186] According to one embodiment, when the electronic device (101) sets a routine for sharing an image with an external electronic device or a routine for reminding an image through the electronic device (101), the execution of the routine can be set as a condition of another routine.

[0187] FIGS. 12a and FIGS. 12b are examples of environments in which the functions of routines for the automation of external electronic devices are executed based on object recognition of images performed within the electronic device.

[0188] Referring to FIG. 12a, an environment (1200a) is illustrated for executing the function of a routine for the automation of an external electronic device (1201) based on object recognition of an image performed within an electronic device (101) implemented as a smartphone. Referring to FIG. 12b, an environment (1200b) is illustrated for executing the function of a routine for the automation of an external electronic device (1201) based on object recognition of an image performed within an electronic device (101) implemented as a smart glass (e.g., augmented glass) including a camera (130). For example, environments (1200a) and (1200b) represent a multi-device experience (MDE) environment.

[0189] Referring to FIGS. 12a and 12b, the electronic device (101) may be in a state where it can communicate with an external electronic device (1201). For example, the electronic device (101) may take an image including an object (1202) (e.g., a door) and an object (1203) (e.g., a door number), and by transmitting the taken image to the external electronic device (1201), cause the external electronic device (1201) to generate a routine.

[0190] According to one embodiment, the electronic device (101) may transmit the image to the external electronic device (1201) or transmit data related to the routine to the external electronic device (1201) in order to set the routine of the external electronic device (1201) corresponding to the object (1202) and the object (1203). For example, the electronic device (101) may set the routine of the external electronic device (1201) using metadata transmitted along with the image. The external electronic device (1201) may store a routine that executes the function of a software application based on the metadata and transmit the result of the execution of the routine to the electronic device (101). For example, when the electronic device (101) transmits the image including the object (1202) and the object (1203) to the external electronic device (1201) or when the image is captured, the external electronic device (1201) may be caused to execute the function of the routine for the function of the software application. For example, the electronic device (101) can transmit metadata (e.g., tag information) transmitted along with the image to an external electronic device (1201). For example, after the external electronic device (1201) stores the routine based on the metadata, when the electronic device (101) transmits the image including object (1202) and object (1203), the external electronic device (1201) can execute the function of the routine based on the image. For example, the external electronic device (1201) transmits the result of the execution of the routine to the electronic device (101), and the electronic device (101) can display the result of the execution through a display (140) or output it through an audio output module (e.g., a speaker). For example, when the external electronic device (1201) captures the image including object (1202) and object (1203) through a camera, the external electronic device (1201) can execute the function of the routine.

[0191] The electronic device (101) may correspond to the electronic device (1301) described with reference to FIG. 13 below.

[0192] FIG. 13 is a block diagram of an electronic device in a network environment according to various embodiments.

[0193] Referring to FIG. 13, in a network environment (1300), an electronic device (1301) may communicate with an electronic device (1302) through a first network (1398) (e.g., a short-range wireless communication network) or with at least one of an electronic device (1304) or a server (1308) through a second network (1399) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (1301) may communicate with the electronic device (1304) through a server (1308). According to one embodiment, the electronic device (1301) may include a processor (1320), memory (1330), input module (1350), sound output module (1355), display module (1360), audio module (1370), sensor module (1376), interface (1377), connection terminal (1378), haptic module (1379), camera module (1380), power management module (1388), battery (1389), communication module (1390), subscriber identification module (1396), or antenna module (1397). In some embodiments, at least one of these components (e.g., connection terminal (1378)) may be omitted from the electronic device (1301), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (1376), camera module (1380), or antenna module (1397)) may be integrated into a single component (e.g., display module (1360)).

[0194] The processor (1320) can, for example, execute software (e.g., program (1340)) to control at least one other component (e.g., hardware or software component) of the electronic device (1301) connected to the processor (1320) and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (1320) can store commands or data received from other components (e.g., sensor module (1376) or communication module (1390)) in volatile memory (1332), process the commands or data stored in volatile memory (1332), and store the resulting data in non-volatile memory (1334). According to one embodiment, the processor (1320) may include a main processor (1321) (e.g., a central processing unit or an application processor) or an auxiliary processor (1323) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (1301) includes a main processor (1321) and an auxiliary processor (1323), the auxiliary processor (1323) may be configured to use less power than the main processor (1321) or to be specialized for a specified function. The auxiliary processor (1323) may be implemented separately from the main processor (1321) or as part thereof.

[0195] The auxiliary processor (1323) may control at least some of the functions or states associated with at least one component of the electronic device (1301) (e.g., display module (1360), sensor module (1376), or communication module (1390)) on behalf of the main processor (1321) while the main processor (1321) is in an inactive (e.g., sleep) state, or together with the main processor (1321) while the main processor (1321) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (1323) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (1380) or communication module (1390)). According to one embodiment, the auxiliary processor (1323) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (1301) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (1308)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.

[0196] The memory (1330) can store various data used by at least one component of the electronic device (1301) (e.g., processor (1320) or sensor module (1376)). The data may include, for example, input data or output data for software (e.g., program (1340)) and related commands. The memory (1330) may include volatile memory (1332) or non-volatile memory (1334).

[0197] The program (1340) may be stored as software in memory (1330) and may include, for example, an operating system (1342), middleware (1344), or an application (1346).

[0198] The input module (1350) can receive commands or data to be used for a component of the electronic device (1301) (e.g., processor (1320)) from outside the electronic device (1301) (e.g., user). The input module (1350) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0199] The sound output module (1355) can output an audio signal to the outside of the electronic device (1301). The sound output module (1355) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.

[0200] The display module (1360) can visually provide information to an external (e.g., user) of the electronic device (1301). The display module (1360) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (1360) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.

[0201] The audio module (1370) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (1370) can acquire sound through an input module (1350) or output sound through an audio output module (1355) or an external electronic device (e.g., electronic device (1302)) (e.g., speaker or headphones) that is directly or wirelessly connected to the electronic device (1301).

[0202] The sensor module (1376) can detect the operating state of the electronic device (1301) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (1376) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0203] The interface (1377) may support one or more specified protocols that can be used for the electronic device (1301) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (1302)). According to one embodiment, the interface (1377) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0204] The connection terminal (1378) may include a connector through which the electronic device (1301) can be physically connected to an external electronic device (e.g., electronic device (1302)). According to one embodiment, the connection terminal (1378) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0205] The haptic module (1379) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be perceived by the user through tactile or kinesthetic senses. According to one embodiment, the haptic module (1379) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.

[0206] The camera module (1380) can capture still images and video. According to one embodiment, the camera module (1380) may include one or more lenses, image sensors, image signal processors, or flashes.

[0207] The power management module (1388) can manage power supplied to the electronic device (1301). According to one embodiment, the power management module (1388) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).

[0208] The battery (1389) can supply power to at least one component of the electronic device (1301). According to one embodiment, the battery (1389) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0209] The communication module (1390) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (1301) and an external electronic device (e.g., electronic device (1302), electronic device (1304), or server (1308)), and the performance of communication through the established communication channel. The communication module (1390) may include one or more communication processors that operate independently of the processor (1320) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1390) may include a wireless communication module (1392) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (1394) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (1304) through a first network (1398) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (1399) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1392) can identify or authenticate the electronic device (1301) within a communication network such as the first network (1398) or the second network (1399) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (1396).

[0210] The wireless communication module (1392) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (1392) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (1392) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (1392) can support various requirements specified in the electronic device (1301), external electronic device (e.g., electronic device (1304)), or network system (e.g., second network (1399)). According to one embodiment, the wireless communication module (1392) may support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.

[0211] An antenna module (1397) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (1397) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (1397) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (1398) or a second network (1399), may be selected from the plurality of antennas, for example, by a communication module (1390). A signal or power may be transmitted or received between the communication module (1390) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (1397).

[0212] According to various embodiments, the antenna module (1397) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.

[0213] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.

[0214] According to one embodiment, commands or data may be transmitted or received between an electronic device (1301) and an external electronic device (1304) through a server (1308) connected to a second network (1399). Each of the external electronic devices (1302, or 1304) may be the same or a different type of device as the electronic device (1301). According to one embodiment, all or part of the operations performed on the electronic device (1301) may be performed on one or more of the external electronic devices (1302, 1304, or 1308). For example, if the electronic device (1301) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (1301) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (1301). The electronic device (1301) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (1301) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In one embodiment, the external electronic device (1304) may include an Internet of Things (IoT) device. The server (1308) may be an intelligent server using machine learning and / or neural networks.According to one embodiment, an external electronic device (1304) or server (1308) may be included within the second network (1399). The electronic device (1301) may be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0215] FIG. 14 is a schematic diagram of an exemplary artificial intelligence (AI) system according to one embodiment.

[0216] Referring to FIG. 14, the AI ​​system (1400) may include an input / output interface (1410), an AI (artificial intelligence) framework (1420), a generative AI model (1430), an application / service component (1480), and / or a knowledge repository (1490).

[0217] The input / output interface (1410) can receive input. The input may include user input and / or data acquired or generated by an electronic device (e.g., the electronic device (101) or electronic device (1301) described above). The data may include images, videos, and / or sensor data generated by at least one processor of the electronic device (e.g., at least one processor (110) or processor (1320)), such as illuminance data around the electronic device acquired from a sensor or sensor hub (e.g., auxiliary processor (1323), attitude data (or orientation data) of the electronic device, temperature inside the electronic device (e.g., display (140)), or temperature of at least one processor (110), size information of the display area of ​​the display, and / or images acquired through an image sensor of the electronic device (e.g., included in a camera module (1380)). The user input may include natural language, touch data obtained through a touch circuit included within the display panel (e.g., used to identify input from a finger and / or stylus), an image displayed (and / or to be displayed) on the display panel, and / or video. By example, without limitation, the user input may be received by an input / output interface (1410) along with context information. The context information may be described as additional information obtained in relation to the user input. The context information may be related to the state at the time the user input is received (e.g., the state of the electronic device and / or the state of the surroundings of the electronic device (e.g., user state)). For example, the context information may include information about one or more software applications executed within the electronic device at the time the user input is received.For example, the above situation information may include information about the location of the electronic device (or the location of the user of the electronic device) at the time the user input is received. For example, the user input may be integrated with the situation information. For example, the user input with the situation information integrated as input may be received by the input / output interface (1410).

[0218] The input / output interface (1410) may transmit (or provide) an output. The output may include a result (or result information) generated or obtained by the AI ​​system (1400) based on at least part of the input. The format of the output may vary. For example, the output may include natural language. For example, the output may include content (e.g., media content and / or multimedia content). For example, the output may include actions related to the user of the electronic device. For example, the output may have a format according to the user settings of the electronic device.

[0219] The input / output interface (1410) can be described as a user question / response interface (1410).

[0220] The AI ​​framework (1420) can be used to obtain information (or data) about the input from the input / output interface (1410) and to control one or more components related to the AI ​​system (1400) using the obtained information.

[0221] For example, a prompt design component (1421) within an AI framework (1420) can generate or obtain prompts for a generative AI model (1430) (e.g., including a large language model (LLM) or a large multimodal model (LMM)) using the acquired information. For example, the prompt design component (1421) may be described as an AI component that utilizes a learning algorithm and / or a neural network to provide prompts that are enhanced over time. For example, the prompt design component (1421) can generate or obtain prompts by accessing a knowledge component (e.g., a knowledge repository (1490)) containing user preference data, a prompt library, and / or prompt examples using the acquired information. The generated prompts may be provided to the generative AI model (1430) (e.g., including an LLM or LMM).

[0222] For example, an API / plugin management component (1422) within the AI ​​framework (1420) may be used to support communication for additional information requested (or induced) in relation to the prompt provided (or to be provided) to the generative AI model (1430). For example, the API / plugin management component (1422) may be used to create or establish a channel for communication with various data sources (e.g., knowledge repository (1490)). For example, the API / plugin management component (1422) may support access to at least some of the data sources. For example, the API / plugin management component (1422) may be used to request another component (e.g., application / service component (1480)) that performs feedback (or response) according to the prompt. As a non-limiting example, information obtained (or generated) through the API / plugin management component (1422) may be provided to the prompt design component (1421) for generating a prompt. As a non-limiting example, information obtained (or generated) through the API / plugin management component (1422) may be provided to the generative AI model (1430).

[0223] For example, an improvement component (1423) within the AI ​​framework (1420) can at least partially tune (or adjust) (or change) the result (e.g., content) obtained (or output) from the generative AI model (1430). For example, the improvement component (1423) can determine or verify whether the content obtained from the generative AI model (1430) is related to the input. For example, the improvement component (1423) can determine or verify whether the content obtained from the generative AI model (1430) contains biased content. For example, the improvement component (1423) can determine or verify whether the content obtained from the generative AI model (1430) contains harmful content. For example, the improvement component (1423) can support or assist in performing additional processing to improve the content obtained from the generative AI model (1430). For example, the improvement component (1423) may support providing a hint to the user to improve the content.

[0224] A generative AI model (1430) can be described as an artificial intelligence neural network that generates feedback in response to a prompt. For example, the feedback may include additional data and / or information relative to the prompt, but relative to the prompt. For example, the feedback may include new content relative to the prompt. For example, the generative AI model (1430) may include a model that generates images and / or a model that generates language. For example, the model that generates images may include a generative adversarial network (GAN) and / or a variational autoencoder (VAE). For example, the model that generates images may include a diffusion-based generative model (e.g., a transformer VAE). For example, the model that generates language may include CHAT-GPT 3 and / or CHAT-GPT 4. For example, a generative AI model (1430) may include an LMM that generates the feedback by recognizing text, images, and / or speech.

[0225] As an example without limitation, the AI ​​framework (1420) and / or generative AI model (1430) may be included within an AI module (e.g., including a processing circuit) within the electronic device. For example, the AI ​​module may be operatively coupled with at least one processor of the electronic device (e.g., at least one processor (110) or processor (1320)). For example, the AI ​​module may be operatively coupled with a display driving circuit of the electronic device. For example, the AI ​​module may be operatively coupled with a sensor hub of the electronic device for one or more sensors within the electronic device.

[0226] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure pertains.

[0227] As described above, an electronic device (e.g., electronic device (101)) may include an image sensor (e.g., image sensor (131)), a display (e.g., display (140)); at least one processor (e.g., at least one processor (110)) including a processing circuit; and a memory (e.g., memory (120)) that stores instructions and includes one or more storage media. The instructions may cause the electronic device to, when executed individually or collectively by the at least one processor: to acquire a first image including a first object through the image sensor; to display the first image through the display; to determine first description information based on at least some of the characteristics of the first object; to determine a first application executable in association with the first description information; and to generate and store a first routine for executing the first application corresponding to the first description information.

[0228] For example, the above instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to: acquire a second image including a second object through the image sensor; display the second image through the display; determine second description information based on at least some of the characteristics of the second object; execute the first application using the first routine based on at least some of the judgment that the second description information corresponds to the first description information while the second image is displayed; and determine a second application executable in association with the second description information based on at least some of the judgment that the second description information does not correspond to the first description information while the second image is displayed, and to generate and store a second routine for executing the second application corresponding to the second description information.

[0229] For example, when the above instructions are executed individually or collectively by the at least one processor: displaying a first UI (user interface) object corresponding to the first routine through the display based on at least part of the judgment that while the second image is displayed, the second depiction information corresponds to the first depiction information; and causing the electronic device to execute the first application using the first routine in response to a first user input to the first UI object.

[0230] For example, when the above instructions are executed individually or collectively by the at least one processor: acquiring a third image including a third object and a fourth object through the image sensor; displaying the third image through the display; determining third description information based on at least some of the characteristics of the third object and fourth description information based on at least some of the characteristics of the fourth object; displaying a second UI object corresponding to the first routine in a first area of ​​the display based on at least some of the judgment that the third description information corresponds to the first description information while the third image is displayed; executing the first application using the first routine in response to a second user input regarding the second UI object corresponding to the first routine; and displaying a third UI object corresponding to the second routine in a second area of ​​the display based on at least some of the judgment that the fourth description information corresponds to the second description information while the third image is displayed; And, in response to a third user input to the third UI object corresponding to the second routine, the electronic device may be caused to execute the second application using the second routine.

[0231] For example, the above instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to: receive a fourth user input for the third object and the fourth object within the third image while the third image is displayed; display the third description information and the fourth description information through the display in response to the fourth user input, a fourth UI object that generates a third routine for executing an executable third application associated with the third description information and the fourth description information; and generate and store the third routine for executing the third application corresponding to the third description information and the fourth description information in response to a fifth user input for the fourth UI object to generate the third routine.

[0232] For example, when the above instructions are executed individually or collectively by the at least one processor: after the determination of the first application, a pop-up window including at least a portion of the first object, the first description information, and a graphic object corresponding to the first application is displayed through the display; and in response to a sixth user input regarding the graphic object included in the pop-up window, the electronic device may generate and store the first routine for executing the first application.

[0233] For example, when the above instructions are executed individually or collectively by the at least one processor: the electronic device may be caused to refrain from displaying the popup window after a defined time has elapsed since the popup window was displayed.

[0234] For example, when the above instructions are executed individually or collectively by the at least one processor: the electronic device may cause at least one graphic effect for the first object to be displayed through the display while the first image is displayed.

[0235] For example, when the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to perform the decision of the first application based on at least a portion of history information indicating that the first application was executed after acquiring the fourth image associated with the first description information.

[0236] For example, the electronic device may further include a camera (e.g., camera (130)) including the image sensor. The instructions, when executed individually or collectively by the at least one processor: receiving, through the electronic device, a seventh user input for capturing the first image through the camera while the first image is displayed; determining, in response to the seventh user input, the first description information for the first object of the first image; identifying, based on at least part of the determination of the first description information for the first image, that the first description information for the first image corresponds to second description information for a second object of the second image, and the second image is captured to execute the function of the first application; And, based on at least a portion of the identification of the first description information for the first image corresponding to the second description information for the second image, the electronic device may be able to display, through the display, a sixth UI (user interface) object that executes the function of the first application using the first image.

[0237] For example, when the above instructions are executed individually or collectively by the at least one processor: identifying that a condition represented by data for the first routine stored in the memory is satisfied based on at least a part of the determination of the first description information for the first image, and that the data represents the function of the first application executed based on at least a part of the satisfaction of the condition that description information for at least one object of the image to be captured through the camera corresponds to the second description information for the second image; and causing the electronic device to display, through the display, the fifth UI object executing the function of the first application represented by the data for the first routine based on at least a part of the satisfaction of the condition.

[0238] For example, when the above instructions are executed individually or collectively by the at least one processor: receiving, through the electronic device, a fifth user input for capturing the second image while the second image is displayed through the display; determining, in response to the fifth user input for capturing the second image, the second description information for the second object of the second image; and, based on at least part of the determination of the second description information for the second image, causing the electronic device to display, through the display, a sixth UI object that generates the data for the first routine to be stored in the memory.

[0239] For example, when the above instructions are executed individually or collectively by the at least one processor: the electronic device may be caused to execute the function of the first application using the first image in response to the ninth user input to the fifth UI object, using data for the first routine. For example, the data may represent one or more conditions for displaying the fifth UI object that executes the function of the first application using an image to be captured through the camera. The one or more conditions may include a condition that the description information for at least one object of the image corresponds to the second description information for the second image.

[0240] For example, the above data may be generated based on at least part of the execution of an automation program using the second image. The automation program may be intended to generate data for a routine representing an operation performed through the electronic device based on at least part of satisfying a condition.

[0241] For example, one or more of the above conditions may include a condition that the electronic device is located within a region defined in association with the second description information for the second image.

[0242] For example, the electronic device may further include at least one sensor (e.g., at least one sensor (150)). The one or more conditions may include a condition that the illuminance level identified through the at least one sensor corresponds to a range defined in association with the second description information for the second image.

[0243] For example, one or more of the above conditions may include a condition that a time (a time of day) identified by the electronic device falls within a time interval defined in association with the second description information for the second image.

[0244] For example, the fifth UI object may be displayed in response to a first-th user input to allow the electronic device to execute the function of a routine.

[0245] For example, the first description information for the first image may correspond to description information associated with user information.

[0246] A non-transient computer-readable storage medium as described above may store one or more programs. The one or more programs may include instructions that cause the electronic device (e.g., electronic device (101)) having an image sensor (e.g., image sensor (131)) and a display (e.g., display (140)) to acquire a first image including a first object through the image sensor; display the first image through the display; determine first description information based on at least some of the characteristics of the first object; determine a first application executable in association with the first description information; and generate and store a first routine for executing the first application corresponding to the first description information.

[0247] A method as described above may be performed by an electronic device (e.g., electronic device (101)) having an image sensor (e.g., image sensor (131)) and a display (e.g., display (140)). The method may include an operation of acquiring a first image including a first object through the image sensor. The method may include an operation of displaying the first image through the display. The method may include an operation of determining first description information based on at least some of the characteristics of the first object. The method may include an operation of determining a first application executable in association with the first description information. The method may include an operation of creating and storing a first routine for executing the first application corresponding to the first description information.

[0248] For example, the above method may include an operation of acquiring a second image containing a second object through the image sensor. The above method may include an operation of displaying the second image through the display. The above method may include an operation of determining second description information based on at least some of the characteristics of the second object. The above method may include an operation of executing the first application using the first routine based on at least some of the judgment that the second description information corresponds to the first description information while the second image is displayed. The above method may include an operation of determining a second application executable in association with the second description information based on at least some of the judgment that the second description information does not correspond to the first description information while the second image is displayed. The above method may include an operation of creating and storing a second routine for executing the second application corresponding to the second description information based on at least some of the judgment that the second description information does not correspond to the first description information while the second image is displayed.

[0249] For example, the above method may include an operation of displaying a first UI (user interface) object corresponding to the first routine through the display, based on at least part of the judgment that while the second image is displayed, the second description information corresponds to the first description information. The above method may include an operation of executing the first application using the first routine in response to a first user input regarding the first UI object.

[0250] For example, the above method may include an operation of acquiring a third image including a third object and a fourth object through the image sensor. The above method may include an operation of displaying the third image through the display. The above method may include an operation of determining third description information based on at least some of the characteristics of the third object and fourth description information based on at least some of the characteristics of the fourth object. The above method may include an operation of displaying a second UI object corresponding to the first routine in a first area of ​​the display based on at least some of the judgment that the third description information corresponds to the first description information while the third image is displayed. The above method may include an operation of executing the first application using the first routine in response to a second user input regarding the second UI object corresponding to the first routine. The above method may include an operation of displaying a third UI object corresponding to the second routine in a second area of ​​the display based on at least some of the judgment that the fourth description information corresponds to the second description information while the third image is displayed. The above method may include an operation of executing the second application using the second routine in response to a third user input for the third UI object corresponding to the second routine.

[0251] For example, the above method may include an operation of receiving a fourth user input regarding the third object and the fourth object within the third image while the third image is displayed. The above method may include an operation of displaying, through the display, a fourth UI object that generates a third routine for executing a third application executable in association with the third description information and the fourth description information in response to the fourth user input, the third description information, and the fourth description information. The above method may include an operation of generating and storing the third routine for executing the third application corresponding to the third description information and the fourth description information in response to a fifth user input regarding the fourth UI object for generating the third routine.

[0252] For example, the above method may include the operation of displaying a popup window through the display, after the determination of the first application, the popup window including at least a portion of the first object, the first description information, and a graphic object corresponding to the first application. The above method may include the operation of generating and storing the first routine for executing the first application in response to a sixth user input regarding the graphic object included in the popup window.

[0253] For example, the above method may include an action of refraining from displaying the popup window after a defined time has elapsed since the popup window was displayed.

[0254] For example, the above method may include the operation of displaying at least one graphic effect for the first object through the display while the first image is displayed.

[0255] For example, the above method may include an operation of performing the decision of the first application based on at least a portion of history information indicating that the first application was executed after acquiring the fourth image associated with the first description information.

[0256] For example, the electronic device may further include a camera including the image sensor. The method may include an operation of receiving a seventh user input through the electronic device for capturing the first image through the camera while the first image is displayed. The method may include an operation of determining the first description information for the first object of the first image in response to the seventh user input. The method may include an operation of identifying that the first description information for the first image corresponds to second description information for a second object of the second image based on at least a portion of the determination of the first description information for the first image. The second image may be captured to execute a function of the first application. The method may include an operation of displaying a fifth UI (user interface) object through the display that executes the function of the first application using the first image based on at least a portion of the identification of the first description information for the first image corresponding to the second description information for the second image.

[0257] For example, the above method may include an operation of identifying that a condition represented by data for the first routine stored in the memory is satisfied, based on at least a portion of the determination of the first description information for the first image. The data may represent the function of the first application executed based on at least a portion of satisfying the condition that the description information for at least one object of the image to be captured through the camera corresponds to the second description information for the second image. The above method may include an operation of displaying, through the display, the UI fifth object executing the function of the first application represented by the data for the first routine, based on at least a portion of satisfying the condition.

[0258] For example, the method may include receiving an eighth user input for capturing the second image through the electronic device while the second image is displayed through the display. The method may include determining second description information for the second object of the second image in response to the eighth user input for capturing the second image. The method may include displaying a sixth UI object through the display that generates data for the first routine to be stored in the memory based on at least a portion of the determination of the second description information for the second image.

[0259] For example, the above method may include an operation of executing the function of the first application using the first image in response to a ninth user input to the fifth UI object, using data for the first routine. The data may represent one or more conditions for displaying the fifth UI object that executes the function of the first application using an image to be captured through the camera. The one or more conditions may include a condition that description information for at least one object of the image corresponds to the second description information for the second image.

[0260] For example, the above data may be generated based on at least part of the execution of an automation program using the second image. The automation program may be intended to generate data for a routine representing an operation performed through the electronic device based on at least part of satisfying a condition.

[0261] For example, one or more of the above conditions may include a condition that the electronic device is located within a region defined in association with the second description information for the second image.

[0262] For example, the electronic device may further include at least one sensor. The one or more conditions may include a condition that the illuminance level identified through the at least one sensor corresponds to a range defined in association with the second description information for the second image.

[0263] For example, one or more of the above conditions may include a condition that a time (a time of day) identified by the electronic device falls within a time interval defined in association with the second description information for the second image.

[0264] For example, the fifth UI object may be displayed in response to a first-th user input to allow the electronic device to execute the function of a routine.

[0265] For example, the first description information for the first image may correspond to description information associated with user information.

[0266] As described above, an electronic device (e.g., electronic device (101)) may include a display (e.g., display (140)); at least one processor (e.g., at least one processor (110)) including a processing circuit; and a memory (e.g., memory (120)) that stores instructions and includes one or more storage media. When the instructions are executed individually or collectively by the at least one processor, a first image is displayed through the display; while the first image is displayed, a user input for capturing the first image is received through the electronic device; based on the user input, tag information for at least one object of the first image is determined; based on the determination of the tag information for the first image, it is identified that the tag information for the first image corresponds to tag information for at least one object of the second image, and the second image is captured to execute a function of a software application. Based on the identification of the tag information for the first image corresponding to the tag information for the second image, the electronic device may be able to display, through the display, a UI (user interface) object that executes the function of the software application using the first image.

[0267] A non-transient computer-readable storage medium as described above may store one or more programs. The one or more programs may include instructions such as, when executed by an electronic device (e.g., electronic device (101)) having a display (e.g., display (140)), display a first image through the display; receive user input through the electronic device to capture the first image while the first image is displayed; determine tag information for at least one object of the first image based on the user input; identify that the tag information for the first image corresponds to tag information for at least one object of the second image based on the determination of the tag information for the first image, and the second image is captured to execute a function of a software application; and cause the electronic device to display, through the display, a UI (user interface) object that executes the function of the software application using the first image based on the identification of the tag information for the first image corresponding to the tag information for the second image.

[0268] A method as described above may be performed by an electronic device (e.g., electronic device (101)) having a display (e.g., display (140)). The method may include: an operation of displaying a first image through the display; an operation of receiving user input through the electronic device to capture the first image while the first image is displayed; an operation of determining tag information for at least one object among the objects of the first image based on the user input; an operation of identifying that the tag information for the first image corresponds to tag information for at least one object of the second image based on the determination of the tag information for the first image, wherein the second image is captured to execute a function of a software application; and an operation of displaying, through the display, a UI (user interface) object that executes the function of the software application using the first image based on the identification of the tag information for the first image corresponding to the tag information for the second image.

[0269] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs.

[0270] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.

[0271] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as "coupled" or "connected" to another (e.g., 2nd) component, with or without the terms "functionally" or "communicationly," it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.

[0272] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0273] Various embodiments of the present document may be implemented as software (e.g., program (1340)) comprising one or more instructions stored in a storage medium (e.g., internal memory (1336) or external memory (1338)) readable by a machine (e.g., electronic device (1301)). For example, a processor (e.g., processor (1320)) of the machine (e.g., electronic device (1301)) may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0274] According to one embodiment, the method according to the various embodiments disclosed herein may be provided as included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0275] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In an electronic device, Image sensor; display; At least one processor including a processing circuit; and Memory that stores instructions and includes one or more storage media, When the above instructions are executed individually or collectively by the at least one processor: Acquiring a first image including a first object through the above image sensor; Displaying the first image through the above display; Based on at least some of the characteristics of the first object, first description information is determined; Determining a first application executable in association with the above first description information; and To create and store a first routine for executing the first application corresponding to the first description information, The above electronic device, causing, Electronic device.

2. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor: Acquiring a second image including a second object through the above image sensor; Displaying the second image through the above display; Based on at least some of the characteristics of the second object, determine second description information; Execute the first application using the first routine based on at least part of the determination that the second depiction information corresponds to the first depiction information while the second image is displayed; and Based on at least part of the judgment that the second depiction information does not correspond to the first depiction information while the second image is displayed: Determining a second application executable in association with the above second description information, and To create and store a second routine for executing the second application corresponding to the second description information, The above electronic device, causing, Electronic device.

3. In Claim 2, When the above instructions are executed individually or collectively by the at least one processor: Based on at least part of the judgment that while the second image is displayed, the second depiction information corresponds to the first depiction information, a first UI (user interface) object corresponding to the first routine is displayed through the display; and In response to the first user input to the first UI object, to execute the first application using the first routine, The above electronic device, causing, Electronic device.

4. In Claim 2, When the above instructions are executed individually or collectively by the at least one processor: Acquiring a third image including a third object and a fourth object through the image sensor; Displaying the third image through the above display; Determining third description information based on at least some of the characteristics of the third object and fourth description information based on at least some of the characteristics of the fourth object; Based on at least part of the determination that the third depiction information corresponds to the first depiction information while the third image is displayed, a second UI object corresponding to the first routine is displayed in the first area of ​​the display; In response to a second user input for the second UI object corresponding to the first routine, the first application is executed using the first routine; Based on at least part of the determination that while the third image is displayed, the fourth depiction information corresponds to the second depiction information, a third UI object corresponding to the second routine is displayed in the second area of ​​the display; and In response to a third user input to the third UI object corresponding to the second routine, the second application is executed using the second routine. The above electronic device, causing, Electronic device.

5. In Claim 4, When the above instructions are executed individually or collectively by the at least one processor: While the third image is displayed, a fourth user input for the third object and the fourth object is received within the third image; In response to the fourth user input, a fourth UI object that generates a third routine for executing a third application executable in association with the third description information and the fourth description information, the third description information, and the fourth description information are displayed through the display; and In response to the fifth user input to the fourth UI object for generating the third routine, the third routine is generated and stored to execute the third application corresponding to the third description information and the fourth description information. The above electronic device, causing, Electronic device.

6. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor: After the determination of the first application, a popup window including at least a portion of the first object, the first description information, and a graphic object corresponding to the first application is displayed through the display; and In response to the sixth user input regarding the graphic object included in the pop-up window, to generate and store the first routine for executing the first application, The above electronic device, causing, Electronic device.

7. In Claim 6, When the above instructions are executed individually or collectively by the at least one processor: Refrain from displaying the popup window after a defined time has elapsed since the above popup window was displayed. The above electronic device, causing, Electronic device.

8. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor: While the first image is displayed, at least one graphic effect for the first object is displayed through the display. The above electronic device, causing, Electronic device.

9. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor: To perform the decision of the first application based on at least a portion of history information indicating that the first application was executed after acquiring the fourth image associated with the first description information. The above electronic device, causing, Electronic device.

10. In Claim 1, The above electronic device is, A camera including the above image sensor is further included, When the above instructions are executed individually or collectively by the at least one processor: While the first image is displayed, a seventh user input for capturing the first image through the camera is received through the electronic device; In response to the above seventh user input, the first description information for the first object of the first image is determined; Based on at least a portion of the determination of the first description information for the first image, identifying that the first description information for the first image corresponds to second description information for a second object of the second image, and the second image is captured to execute the function of the first application; and A fifth UI (user interface) object that executes the function of the first application using the first image, based on at least a portion of the identification of the first description information for the first image corresponding to the second description information for the second image, to be displayed through the display. The above electronic device, causing, Electronic device.

11. In Claim 10, When the above instructions are executed individually or collectively by the at least one processor: Identifying that a condition represented by data for the first routine stored in the memory is satisfied based on at least a portion of the determination of the first description information for the first image, and that the data represents the function of the first application executed based on at least a portion of the satisfaction of the condition that the description information for at least one object of the image to be captured through the camera corresponds to the second description information for the second image; and Based on at least some of the above conditions being satisfied, to display, through the display, the UI 5 object that executes the function of the first application represented by the data for the first routine. The above electronic device, causing, Electronic device.

12. In Claim 11, When the above instructions are executed individually or collectively by the at least one processor: While the second image is displayed through the display, an eighth user input for capturing the second image is received through the electronic device; In response to the eighth user input for capturing the second image, the second description information for the second object of the second image is determined; and To display, through the display, a sixth UI object that generates the data for the first routine to be stored in the memory based on at least a portion of the determination of the second description information for the second image, The above electronic device, causing, Electronic device.

13. In Claim 10, When the above instructions are executed individually or collectively by the at least one processor: In response to the ninth user input to the fifth UI object, the function of the first application using the first image is to be executed using the data for the first routine. The above electronic device, causing, The above data represents one or more conditions for displaying the fifth UI object that executes the function of the first application using an image to be captured through the camera, and The above one or more conditions include a condition that the description information for at least one object of the image corresponds to the second description information for the second image. Electronic device.

14. In a non-transient computer-readable storage medium storing one or more programs, said one or more programs, when executed by an electronic device having an image sensor and a display: Acquiring a first image including a first object through the above image sensor; Displaying the first image through the above display; Based on at least some of the characteristics of the first object, first description information is determined; Determining a first application executable in association with the above first description information; and To create and store a first routine for executing the first application corresponding to the first description information, Instructions including those that cause the above electronic device Non-transient computer-readable storage media.

15. A method performed by an electronic device having an image sensor and a display, The operation of acquiring a first image including a first object through the image sensor; The operation of displaying the first image through the above display; An operation to determine first description information based on at least some of the characteristics of the first object; An operation to determine a first application executable in association with the above first description information; and A method including the operation of creating and storing a first routine for executing the first application corresponding to the first description information. method.