Electronic device and control method therefor

WO2026205746A1PCT designated stage Publication Date: 2026-10-01SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/001775
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-04-30
Filing Date
2026-01-29
Publication Date
2026-10-01

Smart Images

  • Figure KR2026001775_01102026_PF_FP_ABST
    Figure KR2026001775_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides a electronic device control method. The method may comprise the operations of: acquiring first data; receiving a first user input for the first data; on the basis of receiving the first user input, providing a user interface that includes one or more first images related to the first data, the one or more first images including a subject and the subject including a person and / or an animal; receiving a second user input for selecting one or more target images from among the one or more first images, the target image including a target subject; acquiring subject data related to the target subject; providing a user interface that includes one or more user interface subjects related to the subject data; receiving a third user input for selecting one or more target user interface subjects from among the one or more user interface subjects, the target user interface subject being related to target subject data from among the subject data; and acquiring second data on the basis of at least one of the first data, the target image or the target subject data, the second data including an image and / or a recommended activity related to the target subject.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device and control method thereof

[0001] The present disclosure relates to an electronic device and method capable of obtaining a proposal.

[0002] Artificial Intelligence (AI) systems are technologies that learn and make decisions on their own, replacing existing rule-based systems and becoming increasingly sophisticated.

[0003] Artificial intelligence technology is applied across various fields. Linguistic understanding technology encompasses functions that recognize and process human language, such as natural language processing, machine translation, and speech recognition. Visual understanding technology is used to analyze and interpret visual information, including object recognition, image search, and the recognition of people and scenes. Reasoning and prediction technologies make logical judgments based on data and perform optimization and recommendations. Knowledge representation technology is utilized to automatically build and manage knowledge, while motion control technology plays a role in coordinating the movements of autonomous vehicles or robots. Artificial intelligence technology can also be used for image acquisition.

[0004] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. None of the foregoing is to be claimed as prior art related to the present disclosure, nor is it to be used to determine prior art.

[0005] The present disclosure provides an electronic device and a method for obtaining a proposal. The proposal may include a proposal regarding an object. The proposal regarding an object may include an image and / or recommendation activity regarding the object.

[0006] According to one embodiment, the electronic device comprises a memory including at least one storage medium for storing instructions; and at least one processor including a processing circuit, wherein when the instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to perform at least one operation. The at least one operation may include an operation of acquiring first data. The at least one operation may include an operation of receiving a first user input regarding the first data. The at least one operation may include an operation of providing a user interface including one or more first images regarding the first data based on receiving the first user input. The one or more first images may include a subject, and the subject may include a person and / or an animal. The at least one operation may include an operation of receiving a second user input for selecting one or more target images among the one or more first images. The target images may include a target object. The at least one operation may include an operation of acquiring object data regarding the target object. At least one operation may include an operation of providing a user interface comprising one or more user interface objects regarding the object data. At least one operation may include an operation of receiving a third user input for selecting one or more target user interface objects among the one or more user interface objects. The target user interface object may be regarding target object data among the object data. At least one operation may include an operation of obtaining second data based on at least one of the first data, the target image, or the target object data. The second data may include an image and / or recommendation activity regarding the target object.

[0007] According to one embodiment, a control method for an electronic device may include at least one operation. At least one operation may include an operation of acquiring first data. At least one operation may include an operation of receiving a first user input regarding the first data. At least one operation may include an operation of providing a user interface including one or more first images regarding the first data based on receiving the first user input. The one or more first images include a subject, and the subject may include a person and / or an animal. At least one operation may include an operation of receiving a second user input selecting one or more target images among the one or more first images. The target images may include a target object. At least one operation may include an operation of acquiring object data regarding the target object. At least one operation may include an operation of providing a user interface including one or more user interface objects regarding the object data. At least one operation may include an operation of receiving a third user input selecting one or more target user interface objects among the one or more user interface objects. The target user interface object may relate to the target object data among the object data. At least one operation may include an operation to acquire second data based on at least one of the first data, the target image, or the target object data. The second data may include an image and / or recommendation activity regarding the target object.

[0008] According to one embodiment, a storage medium may be provided for storing at least one instruction readable by a computer. The at least one instruction may cause the electronic device to perform at least one operation when executed by at least a part of at least one processor of the electronic device. The at least one operation may include an operation of acquiring first data. The at least one operation may include an operation of receiving a first user input regarding the first data. The at least one operation may include an operation of providing a user interface including one or more first images regarding the first data based on receiving the first user input. The one or more first images include a subject, and the subject may include a person and / or an animal. The at least one operation may include an operation of receiving a second user input for selecting one or more target images among the one or more first images. The target image may include a target object. The at least one operation may include an operation of acquiring object data regarding the target object. The at least one operation may include an operation of providing a user interface including one or more user interface objects regarding the object data. At least one operation may include receiving a third user input that selects one or more target user interface objects among the one or more user interface objects. The target user interface objects may relate to target object data among the object data. At least one operation may include acquiring second data based on at least one of the first data, the target image, or the target object data. The second data may include images and / or recommendation activities related to the target object.

[0009] According to embodiments of the present disclosure, an electronic device and a method for controlling the same can automatically obtain suggestions, such as images and / or recommended activities, based on data selected by a user. The electronic device and the method for controlling the same can obtain suggestions using more efficient and specific prompts by receiving simple text data, obtaining data associated with the text data from a user's personal database, and providing a user interface that allows the user to select desired data from the obtained data.

[0010] The effects obtainable from the exemplary embodiments of the present disclosure are not limited to those mentioned above, and other unmentioned effects can be clearly derived and understood by those skilled in the art to which the exemplary embodiments of the present disclosure belong from the description below. That is, unintended effects resulting from the implementation of the exemplary embodiments of the present disclosure can also be derived by those skilled in the art from the exemplary embodiments of the present disclosure.

[0011] FIG. 1 is a block diagram of an electronic device in a network environment according to various embodiments.

[0012] FIGS. 2a, 2b, 2c, 2d, and 2e illustrate execution screens of an application according to one embodiment.

[0013] FIG. 3 is a block diagram showing the configuration of an electronic device according to one embodiment.

[0014] FIG. 4 illustrates a method for obtaining an image including a face according to one embodiment.

[0015] FIG. 5 illustrates a method for acquiring object data according to one embodiment.

[0016] FIG. 6 is a flowchart illustrating the operation of an electronic device according to one embodiment.

[0017] FIG. 7 is a flowchart illustrating the operation of an electronic device according to one embodiment.

[0018] FIGS. 8a, 8b, 8c, and 8d illustrate execution screens of an application according to one embodiment.

[0019] FIGS. 9a and 9b illustrate execution screens of an application according to one embodiment.

[0020] FIG. 10 is a block diagram of a generative artificial intelligence (AI) system according to one embodiment.

[0021] FIG. 11 is a block diagram of an AI framework according to one embodiment.

[0022] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily practice them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and brevity.

[0023] An embodiment of the present disclosure will be described below with reference to the attached drawings. In the present disclosure, the term "user" may refer to a person using an electronic device or a device using an electronic device (e.g., an artificial intelligence electronic device).

[0024] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to various embodiments.

[0025] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).

[0026] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)), for example, and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., sensor module (176) or communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., central processing unit or application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., graphics processing unit, neural processing unit (NPU), image signal processor, sensor hub processor, or communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use lower power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.

[0027] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.

[0028] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).

[0029] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0030] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0031] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.

[0032] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.

[0033] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).

[0034] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0035] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multi-media interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0036] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0037] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be perceived by the user through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.

[0038] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0039] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).

[0040] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0041] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).

[0042] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or largescale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) can support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.

[0043] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).

[0044] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.

[0045] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.

[0046] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the request may perform at least part of the requested function or service, or additional functions or services related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the above result as is or additionally and provide it as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0047] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.

[0048] FIGS. 2a, 2b, 2c, 2d, and 2e illustrate execution screens (200, 210, 220, 230, 240, 250, 260, 270, 280, 290) of an application according to one embodiment.

[0049] In FIGS. 2a, 2b, 2c, 2d, and 2e, an electronic device (e.g., the electronic device (101) of FIG. 1) may obtain second data based on first data using image and / or relationship information stored in a user's (e.g., user of the electronic device (101)) personal database. For example, the electronic device (101) may generate a prompt for image generation based on the first data and / or data stored in the user's personal database (e.g., image and / or relationship information), and obtain second data including an image (e.g., an image associated with the first data) generated through an AI model (e.g., a generative AI model) based on the prompt. In this case, the first data may serve as the basis for a primary prompt used to generate the corresponding prompt (e.g., a final prompt). In one embodiment, the electronic device (101) may obtain the first data based on user input. The first data may be any type of data that can be obtained by user input. For example, the first data may be at least one of text data, image data, gesture data, or sound data, but is not limited thereto. The second data may be a suggestion, for example, an image and / or recommendation activity, but the present disclosure is not limited thereto. In one embodiment, the electronic device (101) may execute an application (e.g., a sketch application). For example, the application may display an object based on drawing input and may generate an image based on drawing input. FIGS. 2a through 2c illustrate a method for obtaining the second data based on user input received through a user interface (e.g., a graphical user interface) of one application according to one embodiment. However, the present disclosure is not limited thereto, and the electronic device (101) may obtain the second data based on user input received through the user interfaces of a plurality of applications.

[0050] In one embodiment, the electronic device (101) may acquire first data (201). For example, as illustrated in the execution screen (200) of the application, the electronic device (101) may receive user input that inputs the first data (201). For example, the user input may include keyboard input and / or drawing input. The electronic device (101) may receive drawing input using the user's body (e.g., fingers) and / or input device (e.g., a pen and / or a ring). The electronic device (101) may acquire the first data (201) by the user input (e.g., keyboard input and / or drawing input). For example, the first data (201) may include words and / or sentences in natural language form associated with the image the user intends to create. For example, the first data (201) may include the word "Fire Work".

[0051] In one embodiment, the electronic device (101) may receive a first user input for the first data (201). The electronic device (101) may be configured to identify the first user input as a control action requesting the acquisition of the second data. The electronic device (101) may identify the first user input as a control action requesting the completion of the input of the first data (201) and the performance of an action to acquire the second data based on the first data (201). The first user input may be predefined. The first user input may include at least one of an object selection input, a gesture input, or a voice input. For example, a gesture input may include the input of a circular, square, or freeform object (211) that includes at least a portion of the area where the first data (201) is displayed. For example, as illustrated in the execution screen (210) of the application, the first user input may include the input of a circular object (211) that includes an area in which the first data (201) is displayed. The first user input may include not only drawing input of simple shapes such as circles or squares, but also object selection input using a character input tool and / or a user interface, such as arrows, special symbols, numbers, gesture input text (e.g., handwritten text), keyboard input text, and / or UI menu selection input. The electronic device (101) may predefine the first user input through the application's setting function. The electronic device (101) may set a specific first user input that limits the acquisition of a specific data type when performing an operation to acquire object data regarding a target object (as further described in this disclosure).For example, a gesture input including the input of a rectangular object can be set as a control action requesting that other types of data, such as image data and / or sound data, not be acquired in addition to text type when performing a subject data acquisition action. The electronic device (101) can perform different actions depending on the type of the first user input.

[0052] In one embodiment, the electronic device (101) may provide a user interface (e.g., a graphical user interface) comprising one or more first images (221: 221-1, 221-2, 221-3, 221-4) relating to the first data (201) based on receiving a first user input. The electronic device (101) may obtain the first images (221) from a user's personal database (e.g., an image database of a gallery application). The first images (221) may include images associated with the first data (201). The first images (221) may include one or more subjects, and the subjects may include people (e.g., human subjects) and / or animals (e.g., animal subjects). For example, in the present disclosure, a subject may mean a primary subject of an image among a plurality of objects included in and / or that may be included in the image. For example, a subject may be a type of object. For example, a person and / or an animal may be recognized as a subject in the image of the present disclosure (e.g., the first image).

[0053] In one embodiment, the first image (221) may be a thumbnail image containing the face of an object. A method for obtaining the first image (221) will be described later with reference to FIG. 4. For example, as shown in the application execution screen (220), the electronic device (101) may provide a user interface including the first image (221) containing one or more objects as an image of the first data (201) (e.g., "Fire Work") based on receiving the first user input (221). For example, the user interface may display the first image (221) around an object (e.g., a circular object (211)) according to the first user input. The user interface shown in the application execution screen (220) is merely an example, and the electronic device (101) may provide various user interfaces including the first image (221). For example, the user interface may display the first image (221) at the bottom of the object (211) according to the first user input. For example, the user interface may display the names of the subjects included in each of the first images (221) together with each of the first images (221).

[0054] In one embodiment, the electronic device (101) may receive a second user input for selecting one or more target images (221-1) among one or more first images (221). The target images may include target objects. The second user input may be an object selection input. For example, the objects may be user interface objects (222: 222-1, 222-2, 222-3). For example, as illustrated in the application execution screen (220), the electronic device (101) may display a user interface object (e.g., a "+" symbol) (222) corresponding to each of the first images (221). The electronic device (101) may display a user interface object (222) corresponding to the vicinity of each of the first images (221) (e.g., a location adjacent to where the first image (221) is displayed). The electronic device (101) can identify that a corresponding first image (221) has been selected based on receiving a user input that selects (e.g., touches) any one of the user interface objects (222). For example, as shown in the execution screen (230) of the application, the electronic device (101) can receive a second user input that selects a target image (221-1) among one or more first images (221). The electronic device (101) can identify a user input that selects (e.g., touches) a user interface object (222-1) as the second user input that selects the target image (221-1).

[0055] In one embodiment, the electronic device (101) may provide a user interface (e.g., a graphical user interface) containing information about an object included in the first image (221) based on receiving user input touching any one of the first images (221). For example, the electronic device (101) may display information about an object included in the first image (221) in a table format. For example, after checking information about an object included in each of the first images (221) (e.g., the name of the object) on the application execution screen (220), the user may select one or more target images using a user interface object (222).

[0056] In one embodiment, the electronic device (101) may provide a user interface (e.g., a graphical user interface) that indicates that user input selecting a target image (221-1) has been received. For example, as shown in the application execution screen (240), the electronic device (101) may provide a user interface that displays the target image (221-1) based on the identification that the target image (221-1) has been selected, but does not display the first images (221-2, 221-3, and 221-4) that have not been selected. For example, the electronic device (101) may display the periphery of the target image (221-1) in a specific color (e.g., yellow).

[0057] In one embodiment, the electronic device (101) may acquire object data regarding a target object and provide a user interface (e.g., a graphical user interface) comprising one or more user interface objects (251: 251-1, 251-2, 251-3) regarding the object data. The operation of acquiring object data regarding a target object will be described later with reference to FIG. 5. The electronic device (101) may acquire object data regarding a target object from a user's personal database. The user's personal database may include data acquired from at least one of the user's schedule, social network service (SNS) activity, contacts, or notes. The electronic device (101) may determine the shape of the user interface object (251) based on the content and / or source of the object data. The user may check the content and / or source of the object data (e.g., source application) through the user interface object (251) and select the desired object data. For example, as illustrated in the application execution screen (250), the electronic device (101) may provide a user interface including a user interface object (251-1) in the shape of a hotel reservation application icon, a user interface object (251-2) in the shape of a calendar, and a user interface object (251-3) in the shape of a ticket. For example, the user interface object (251-1) may be object data originating from a hotel reservation application. For example, the user interface object (251-2) may be object data originating from a calendar application. For example, the user interface object (251-3) may be object data regarding a memo containing flight reservation details.

[0058] In one embodiment, the electronic device (101) may receive a third user input selecting one or more target user interface objects (251-1 and 251-2) among one or more user interface objects (251). The target user interface objects (251-1 and 251-2) may be related to the target object data among the object data. The third user input may be an object selection input. For example, the object may be a user interface object (252: 252-1, 252-2, 252-3). For example, as illustrated in the application execution screen (250), the electronic device (101) may display a user interface object (e.g., a "+" symbol) (252) corresponding to each of the user interface objects (251). The electronic device (101) may display a user interface object (252) corresponding to the vicinity of each of the user interface objects (251). The electronic device (101) can identify that a corresponding user interface object (251) has been selected based on receiving a user input that selects (e.g., touches) any one of the user interface objects (252). For example, as illustrated in the execution screen (260) of the application, the electronic device (101) can receive a third user input that selects one or more target user interface objects (251-1 and 251-2) among the one or more user interface objects (251). The electronic device (101) can identify a user input that selects (e.g., touches) the user interface objects (252-1 and 252-2) as the third user input that selects the target user interface objects (251-1 and 251-2). In one embodiment, the third user input that selects the target user interface objects may not be received. If the third user input is not received, the electronic device (101) can obtain second data based on the first data and / or the target image.

[0059] In one embodiment, the electronic device (101) may provide a user interface (e.g., a graphical user interface) including object data represented by a user interface object (251) based on receiving a user input that touches one of the user interface objects (251). For example, after checking the object data represented by each user interface object (251) on the application execution screen (250), the user may select one or more target user interface objects using a user interface object (252). For example, the electronic device (101) may display object data represented by a calendar-shaped user interface object (251-2) (e.g., information regarding a schedule (date and / or content) with a target object) based on receiving a user input that touches a calendar-shaped user interface object (251-2). The user may check through the user interface whether the data represented by the user interface object (251) corresponds to their intention.

[0060] In one embodiment, the electronic device (101) may provide a user interface (e.g., a graphical user interface) that indicates that user input selecting target user interface objects (251-1 and 251-2) has been received. The electronic device (101) may determine at least some of the user interface elements as target user interface elements based on the user input. For example, as illustrated in the application execution screen (270), the electronic device (101) may provide a user interface that displays the target user interface objects (251-1 and 251-2) based on the identification that the target user interface objects (251-1 and 251-2) are selected, but does not display the unselected user interface object (251-3). For example, the electronic device (101) may display the periphery of the target user interface objects (251-1 and 251-2) in a specific color (e.g., yellow). The electronic device (101) may provide a user interface that displays data selected by the user. The electronic device (101) may provide a user interface that displays a target image (221-1) and target user interface objects (251-1 and 251-2) selected by the user. Although not illustrated, the remaining user interface objects (251-3) are removed, except for the target user interface objects (251-1 and 251-2) representing the data finally selected by the user, and additional data associated with the target object data represented by the target user interface objects (251-1 and 251-2) may be displayed at the location where the user interface objects (251-3) were removed. For example, the electronic device (101) may identify that the user frequently books performances based on the user's personal database and the results of checking the user's consumption patterns.When user input is received selecting a calendar-shaped user interface object (251-2) among the user interface objects (251), the electronic device (101) can crawl information regarding future performances and display the crawled information as a user interface object. Based on the data displayed by the user interface object (251), the electronic device (101) can enable the user to additionally input a prompt associated with the data.

[0061] In one embodiment, the electronic device (101) may acquire second data based on at least one of the first data (201), the target image (222-1), or the target object data. The second data may include images and / or recommendation activities regarding the target object. For example, as shown in the application execution screen (270), the electronic device (101) may identify that the selection of the target user interface objects (251-1 and 251-2) is complete. The electronic device (101) may identify that the selection of the target user interface objects (251-1 and 251-2) is complete based on the reception of a selection for a user interface object (not shown) indicating that the selection is complete. The electronic device (101) may identify that the selection of the target user interface objects (251-1 and 251-2) is complete based on the fact that no user input for further selecting the user interface object (251) is received for a certain period of time. When it is identified that the selection of target user interface objects (251-1 and 251-2) is completed, the electronic device (101) can acquire second data based on at least one of the first data (201), target image (222-1), or target object data.

[0062] In one embodiment, the electronic device (101) may display second data. For example, as shown in the application execution screen (280), the electronic device (101) may generate an image and provide a user interface (281) to check whether the user wants to view the generated image. For example, the user interface (281) may display a message such as "Image generation is complete. Would you like to view it?". The electronic device (101) may receive user input (e.g., touch input for a YES button) acknowledging the display of the generated image. Based on the reception of user input acknowledging the display of the generated image, the electronic device (101) may display the generated image (291), as shown in the application execution screen (290). The electronic device (101) may receive user input (e.g., touch input for a YES button) acknowledging the saving of the generated image (291). The electronic device (101) can store the generated image (291) based on receiving user input that approves the storage of the generated image (291).

[0063] In one embodiment, as illustrated in the application execution screen (280), the electronic device (101) may generate a recommended activity (e.g., scheduling an event) and provide a user interface (e.g., a graphical user interface) (282) to determine whether to execute the recommended activity. For example, the user interface (282) may display a message such as "Would you like to schedule an additional event with John?". The electronic device (101) may receive user input (e.g., a touch input to a YES button) acknowledging the execution of the recommended activity. For example, the user interface (282) may display user interface objects (284, 285, 286) representing data selected by the user for obtaining first data (283) and second data. The user interface objects (284, 285, 286) may include a user interface object (286) corresponding to a target image and a user interface object (284, 285) corresponding to a target user interface object. A user interface object (286) corresponding to a target image may include an image identical to the target image (221-1). A user interface object (284, 285) corresponding to a target user interface object may have the same shape as the target user interface object (251-1, 251-2). The electronic device (101) may display information regarding the selected user interface object (284, 285, 286) based on receiving a user input selecting (e.g., touching) any one of the user interface objects (284, 285, 286). The user may check the information and approve or disapprove the execution of a recommendation activity. The electronic device (101) may execute a recommendation activity based on receiving a user input approving the execution of the recommendation activity (e.g., a touch input to a YES button).For example, based on receiving user input approving to schedule an additional schedule with John, the electronic device (101) can book accommodation based on accommodation booking history through an accommodation application, schedule data from a calendar application, and / or conversation history with John through an SNS application.

[0064] In one embodiment, the execution screens of the application (200, 210, 220, 230, 240, 250, 260, 270, 280, 290) are not only displayed sequentially, but, depending on the embodiment, the display of some execution screens of the application (an action related thereto) may be omitted, and the order of display of the execution screens of the application may be changed.

[0065] In one embodiment, the electronic device (101) may provide a user interface (e.g., a graphical user interface) comprising one or more first images (2011: 2011-1, 2011-2) regarding the first data (201), as shown in the application execution screen (2010) of FIG. 2e, based on receiving a first user input for the first data (201) in the application execution screen (210) shown in FIG. 2a, for example. For example, the first image (2011) may include an image corresponding to an object with established person relationships. The establishment of person relationships will be described later with reference to FIG. 7. The electronic device (101) may search for images in the user's personal database based on the first data (201). The user's personal database may be provided in an electronic device (101), an external electronic device (e.g., the external electronic device (102 or 104) of FIG. 1), or a server (e.g., the server (108) of FIG. 1). The electronic device (101) can identify objects (e.g., people and / or animals) within a retrieved image. The electronic device (101) can identify whether the identified object is an object with established person relationships. Based on the retrieval of an image containing an object with established person relationships, the electronic device (101) may provide a user interface containing object information. For example, the object information may include an image and / or name of the object (e.g., Chloe, Chris). The image of the object may include an object image clustered based on the face recognition of the object in the user's personal database (e.g., the image database of a gallery application). For example, the object information may include object information (e.g., name, relationship information, object image) that the user has manually set (or registered) directly in the electronic device (101).Object information may be displayed at the bottom of the display as shown in the application execution screen (2010) of FIG. 2e, or around the object (211) as shown in the application execution screen (2020) of FIG. 2e.

[0066] In one embodiment, the electronic device (101) may receive a second user input selecting one or more target images (2011-1) among the first images (2011) obtained based on the first data (201). The target images (2011-1) may include target objects. For example, as illustrated in the execution screen (2010) of the application, the electronic device (101) may display a user interface object (e.g., a "+" symbol) (2012) corresponding to each of the first images (2011). The electronic device (101) may display a user interface object (2012) corresponding to each of the first images (2011) near each of the first images (2011). The electronic device (101) may identify that the corresponding first image (2011) is selected based on receiving a user input selecting (e.g., touching) any one of the user interface objects (2012). For example, as illustrated in the application execution screen (2020), the electronic device (101) may receive a second user input selecting a target image (2011-1) among one or more first images (2011). The electronic device (101) may identify a user input selecting (e.g., touching) a user interface object (2012-1) as the second user input selecting the target image (2011-1).

[0067] In one embodiment, the electronic device (101) may acquire second data based on receiving a second user input selecting a target image (2021-1). For example, the electronic device (101) may acquire object data regarding a target object (e.g., Chloe) based on receiving a second user input selecting a target image (2021-1) and may provide a user interface (e.g., a graphical user interface) comprising one or more user interface objects (2031: 2031-1, 2031-2) regarding the object data. The electronic device (101) may receive a third user input selecting one or more target user interface objects among the user interface objects (2031). The electronic device (101) may acquire second data based on at least one of the first data (201), the target image (2011-1), or the target object data (2031-1 and / or 2032-2). For example, object data may include keywords of at least one image containing a target object.

[0068] In one embodiment, the electronic device (101) may acquire keywords of at least one image containing a target object. The electronic device (101) may acquire at least one image containing a target object based on facial feature data of the target object (e.g., see FIG. 7). The electronic device (101) may analyze at least one image containing a target object and acquire keywords (e.g., beach, sunset) corresponding to the main subject and / or context included in at least one image. For example, the electronic device (101) may acquire keywords based on caption information of at least one image. For example, as shown in the execution screen (2030) of the application in FIG. 2e, the electronic device (101) may display user interface objects (2031: 2031-1, 2031-2) corresponding to the keywords. The electronic device (101) may receive user input for selecting user interface objects (2031) corresponding to the keywords. For example, the user input may be an object selection input. For example, the object may be a user interface object (2032: 2032-1, 2032-2). For example, as illustrated in the application execution screen (2030), the electronic device (101) may display a user interface object (e.g., a "+" symbol) (2032) corresponding to each keyword. The electronic device (101) may display a corresponding user interface object (2032) near a user interface object (2031) corresponding to a keyword. The electronic device (101) may identify that a corresponding keyword has been selected based on receiving user input selecting (e.g., touching) any one of the user interface objects (2032). The electronic device (101) may generate an image using the keyword based on receiving user input selecting the keyword.The electronic device (101) can acquire (or extract) features from an image corresponding to a keyword (e.g., an image in which the keyword is acquired) and acquire a prompt for image generation based on the features.

[0069] FIG. 3 is a block diagram showing the configuration of an electronic device according to one embodiment (e.g., the electronic device (101) of FIG. 1).

[0070] In FIG. 3, the electronic device (101) may include or be composed of at least one of a pre-processor (310), a first search module (320), a second search module (330), or a suggestion generator (340). The components included in the electronic device (101) may be software or firmware executed by the processor (120). The components included in the electronic device (101) may be hardware modules included in the processor (120) or existing independently.

[0071] In one embodiment, the preprocessor (310) can take input data received by the electronic device (101) as input, analyze the input data, and output the result of the analysis. The preprocessor (310) can receive data entered by a user while the image generation application is running. The preprocessor (310) can analyze the format and / or meaning of the input data. For example, the input data may include at least one of text, image, voice, or gesture (e.g., drawing), but is not limited thereto. For example, the preprocessor (310) can receive text in the form of natural language and analyze the format and meaning of the input text.

[0072] In one embodiment, the preprocessor (310) may include a large language model (LLM) and / or an optical character recognition (OCR) model. An LLM is a deep learning algorithm capable of performing natural language processing (NLP) tasks by learning from a vast amount of data. For example, natural language processing tasks may include machine translation, document summarization, question answering, and / or text generation. An LLM is based on a transformer architecture. An LLM may be pre-trained to solve problems such as machine translation, document summarization, question answering, and / or text generation, and then fine-tuned to be optimized for specific tasks. An LLM may be a multi-modal LLM capable of receiving various forms of data other than text. For example, an LLM may include ChatGPT, BERT (bidirectional encoder representations from transformers), and / or LLaMA (large language model meta AI). OCR models can recognize text in documents and / or images through OCR processing.

[0073] In one embodiment, the preprocessor (310) receives data in the form of an image through a display (e.g., the display module (160) of FIG. 1) and can identify text from the input data through software. For example, if a user inputs text on the display (160) using a gesture or an external input device (e.g., a pen and / or a ring), the electronic device (101) can identify the input text as an image type. The LLM can receive text information of the image type input by the user. The LLM can analyze the input data. For example, the LLM can analyze whether there are typos in the text input by the user and, if necessary, correct the typos and replace them with the correct words. If the user inputs text that is not horizontally aligned (e.g., text aligned diagonally or inverted text) on the display (160), the LLM can analyze the text intended by the user (e.g., horizontally aligned text) and output the analysis result. When text is horizontally aligned in an image, the OCR model can recognize the text. The text recognition rate of an OCR model can be improved through preprocessing such as horizontal alignment. An OCR model can learn how to recognize text from input data by using input data in an ideal form (e.g., horizontally aligned text) and ground truth data corresponding to that input data. If optical character recognition (OCR) is performed without preprocessing, it may be relatively difficult to recognize text. An LLM is used in the preprocessing process to convert input data into a horizontal form (e.g., horizontally aligned text), thereby increasing the accuracy of OCR. The corrected user input data output from the LLM is passed as input to the OCR model, and the OCR model can extract and output text from the user input data. The preprocessor (310) can transmit the output data to the first search module (320).In one embodiment, when a user inputs text through a keyboard, the text can be transmitted to the first search module (320) without preprocessing.

[0074] In one embodiment, the OCR model can use a text recognition algorithm to extract data from input data (e.g., scanned documents, camera images, and / or image-based PDF files) and change the use of the extracted data. For example, the OCR model can select characters from an image to form words and convert the formed words into sentences, thereby enabling the electronic device (101) to access and edit the input data (original content). The OCR model can prevent data duplication that may occur during the manual data input process. The OCR model can utilize a combination of hardware and software to convert image-type data (e.g., scanned documents) into machine-readable text. For example, the electronic device (101) can recognize text through the OCR model after receiving image-type data input via a touch screen on the display (160).

[0075] In one embodiment, the first search module (320) receives text-type data output from the preprocessor (310) and can output an image associated with the data. The image may include an object (e.g., a person and / or an animal). The first search module (320) may include a large vision model (LVM) (321) and / or a face recognition model (322). The first search module (320) receives text-type data output from the preprocessor (310) and can use the LVM (321) and / or the face recognition model (322) to search for and output data associated with the text from the user's personal database. The personal database may be stored in an electronic device (101) and / or a server (e.g., the server (108) of FIG. 1). For example, the personal database may include an image database of a gallery application installed on the electronic device (101). LVM (321) can access images stored in a personal database (e.g., gallery images (323) stored in the image database of a gallery application). The gallery images (323) or the image database in which the gallery images (323) are stored may be internal or external images of the first search module (320), or internal or external databases. LVM (321) can output images by taking text data (e.g., corrected and / or analyzed text) received from the preprocessor (310) as input. LVM (321) may be a large-scale model designed to facilitate bidirectional data movement and analysis between text-to-image and image-to-text. By using LVM (321), the first search module (320) can search for data (e.g., image) that is contextually related to the input data in the personal database based on user input data (e.g., text).For example, if a user inputs the word "firework" into an electronic device (101) via an external input device (e.g., a pen), the LVM (321) can analyze "firework" into various meanings such as fire play, campfire, or fireworks. The LVM (321) can extract images related to the text with various meanings from a personal database. The LVM (321) can pass the extracted images to a face recognition model (322). The LVM (321) can determine the meaning of natural language in the form of sentences as well as words. The user's personal database (e.g., an image database of a gallery application) can support natural language search in the form of sentences as well as words. The LVM (321) can search for related images in the personal database using text with various meanings (e.g., words and / or sentences). The LVM (321) can pass the images related to the text with various meanings to a face recognition model.

[0076] In one embodiment, the face recognition model (322) may receive images related to text from the LVM (321). The face recognition model (322) may recognize the face of an object. The face recognition model (322) may filter images containing an object (e.g., a person and / or an animal) among the images searched by the LVM (321) and output images containing the object. The face recognition model (322) may exhibit more efficient performance by recognizing faces using fewer parameters than the LVM (321). The first search module (320) may receive text-type data output from the preprocessor (310), filter images associated with the data using the LVM (321), and filter images containing a face among the filtered images using the face recognition model (322). In one embodiment, the first search module (320) receives text-type data output from the preprocessor (310) and can filter images containing faces among images associated with the data using LVM (321). However, since LVM (321) uses a large number of parameters and is implemented with a deep layer structure, detecting faces using LVM (321) may be relatively inefficient.

[0077] In one embodiment, the first search module (320) may acquire one or more first images using the coordinates of the face region of an object in an image containing an object, which is acquired through face recognition. The first image may include the face of the object. The first image may be a thumbnail-type image (hereinafter referred to as a thumbnail image). The first image may be provided as a user interface object. The first image may be displayed on a display (160). The user may select one or more target images from among the one or more first images displayed on the display. The target image may include the face of the target object. The electronic device (101) may receive user input selecting one or more target images from among the one or more first images. The first search module (320) may transmit the target images to the second search module (330).

[0078] In one embodiment, the second search module (330) receives a target image selected by the user and may recommend a prompt for obtaining second data using the target image based on the user's personal database. The second data may include images and / or recommendation activities regarding the target object, but the present disclosure is not limited thereto. The second search module (330) may include at least one of a face recognition model, a personal database (DB) manager, or a user interface object creation model. The second search module (330) may obtain information regarding the target object from the user's personal database using the face recognition model and the personal database manager. The user's personal database may be a relationship information database. The relationship information database may store and / or manage information about objects (e.g., people and / or animals) that have a relationship with the user. The information about the object may include information such as the object's name, contact information, image, and events with the user (e.g., conversations, activities). The relationship information database may be stored in an electronic device (101) and / or a server (e.g., the server (108) of FIG. 1). In one embodiment, the second search module (330) receives a target image and can obtain information about a target object included in the target image from the relationship information database.

[0079] In one embodiment, the second search module (330) can obtain information regarding a target object using a face recognition model. The face recognition model of the second search module (330) may be substantially the same as the face recognition model of the first search module (320). The face recognition model may receive a target image as input. The face recognition model may extract feature information of the face based on the face of the target object in the target image. The face recognition model may extract feature information of the face through face detection, face alignment, face normalization, and / or face representation. Face detection is a process of extracting coordinate information of the face, and may use the position coordinate information of the face extracted from the face recognition model of the first search module (310). Face alignment is a process of finding the location of the face in the image based on the position coordinates of the face and then extracting the main landmark points of the face (eyes, nose, mouth, face contour, etc.). Face normalization is a process of rotating the face region based on the extracted landmark points and changing it to a state where face matching is possible. Face representation is the process of expressing a normalized face as an n-dimensional (where n is a natural number) feature vector through an embedding process. Embedding is a process of transforming high-dimensional data into a low-dimensional space (mapping it into a vector of a fixed size) to retain only the information necessary for recognition. Since the feature vectors transformed by the embedding process have unique values ​​for each object, it is possible to identify objects by comparing vectors.

[0080] In one embodiment, information regarding an object including an object's face feature vector may be stored in a user's personal database. The second search module (330) can obtain substantially matching object information by comparing the face feature vector output by the face recognition model with the face feature vectors stored in the user's personal database. The substantially matching object information may be information regarding an object in which the matching score between the face feature vectors is greater than or equal to a threshold.

[0081] In one embodiment, the personal database manager may receive data output by a face recognition model (e.g., face feature vector). The personal database manager may use the received data (e.g., face feature vector) to search for information regarding an object matching that data in the user's personal database. When searching for information regarding an object, data transmission and reception may occur repeatedly between the user's personal database and the personal database manager. The user's personal database may store information regarding the object as a vector in the form of a dictionary. The personal database manager may search for and output data corresponding to the object's face feature vector from the user's personal database. The personal database manager may output data to a user interface object creation model. In one embodiment, data corresponding to the object's face feature vector may not be stored in the user's personal database. In this case, the personal database manager may build a database regarding the object in the user's personal database. The personal database manager may store the object's face feature vector in the user's personal database.

[0082] In one embodiment, data corresponding to an object's face may be composed of various types of sub-data. For example, the sub-data may include a schedule with the target object in the case of a calendar application, conversation content and / or post content with the target object in the case of a social media application, and / or text recording an event with the target object in the case of a diary application. The sub-data may include images as well as text. When a large amount of data is stored in a user's personal database, bottlenecks may occur when retrieving the data. To prevent this problem in advance, the user's personal database may store data by vectorizing it.

[0083] In one embodiment, the second search module (330) can generate a user interface object representing data regarding a target object using a user interface object generation model. The user interface object generation model can receive data output by a personal database manager and generate a user interface object. The user interface object can display information regarding the target object on a display (160) so that the user can verify information regarding the target object. The user can select one or more target object data among the object data regarding the target object to generate an input prompt for a generative artificial intelligence model (hereinafter further described in this disclosure). The electronic device (101) can receive user input selecting one or more target object data among the object data regarding the target object. The second search module (330) can transmit the target object data to a suggestion generator (340).

[0084] In one embodiment, the suggestion generator (340) may obtain suggestions based on an input prompt using a generative artificial intelligence model. For example, the suggestion may include, but is not limited to, images and / or recommendation activities regarding an object. The input prompt may utilize first data, target images, and / or target object data. The generative artificial intelligence model may include a computer algorithm that automatically generates and / or modifies images. The generative artificial intelligence model may generate images through pattern recognition and / or learning based on a training data set. For example, the generative artificial intelligence model may include models such as Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), and / or Diffusion, but is not limited thereto.

[0085] In one embodiment, the electronic device (101) can acquire an image using a prompt generated based on the user's selection. In one embodiment, the electronic device (101) can generate an image corresponding to the user's intention using a prompt associated with an object. The electronic device (101) can perform various additional tasks in addition to image generation by using sub-data corresponding to the user's intention based on relationship information with the object.

[0086] In one embodiment, the electronic device (101) can perform additional tasks by driving an action model based on sub-data selected by the user. According to one embodiment, the electronic device (101) can generate an image by configuring a prompt using an image and / or data regarding an object selected by the user. In this case, the electronic device (101) can generate an image that matches the user's intent. The electronic device (101) can generate an image using a prompt generated based on the target object selected by the user and data regarding the target object (e.g., person relationship information stored in the electronic device (101)). The electronic device (101) can perform various additional tasks, including generating an image that matches the user's intent, by using the sub-data selected by the user among the data regarding the target object selected by the user. For example, when user input is received selecting both sub-data obtained from an accommodation reservation application and a calendar application, the electronic device (101) can determine that there is a schedule with the target object and that accommodation reservation is required for that schedule. If a future schedule with the target object is set in the calendar application, the action model can confirm with the user whether to reserve accommodation for that schedule. When user input approving a lodging reservation is received, the electronic device (101) can reserve the lodging using an action model. In this way, the electronic device (101) can perform additional operations using an action model when an association between sub-data selected by the user is confirmed. According to one embodiment, the electronic device (101) can perform additional operations using an action model when an association between prompts used for image generation is confirmed.

[0087] FIG. 4 illustrates a method for obtaining an image including a face according to one embodiment.

[0088] In FIG. 4, the face image filtering module (400) may include at least one of an image search model (410) or a face recognizer (420). The face image filtering module (400) may be one of the modules constituting the first search module (320) of FIG. 3. The image search model (410) and the face recognizer (420) may be the LVM (large vision model) (321) and the face recognition model (322) of the first search module (320), respectively.

[0089] In one embodiment, the face image filtering module (400) receives text data (401) and / or image data (402) and outputs a face-containing image (404). The text data (401) may be first data entered by a user. The image data (402) may be data containing images stored in the user's personal database (e.g., images from a gallery application).

[0090] In one embodiment, the image search model (410) receives text data (401) and / or image data (402) as input and can output one or more images (403) corresponding to the text data (401). The image search model (410) can output one or more images (403) that are high in relation to the text data (401). The image search model (410) can determine how similar the meaning of the text data (401) is to the images in the image data (402) using a contrastive language-image pretraining (CLIP) model. The image search model (410) can calculate a CLIP score between each image in the image data (402) and the text data (401). The image search model (410) can select and output images that received relatively high CLIP scores. The image search model (410) can transmit one or more images (403) corresponding to the text data (401) to a face recognition device (420).

[0091] In one embodiment, the face recognizer (420) receives one or more images (403) corresponding to text data (401) and can output a face-containing image (404). The face recognizer (420) can filter images containing at least a portion of the object's face among one or more images (403). In one embodiment, the image search model (410) does not transmit one or more images (403) to the face recognizer (420) and can filter images containing at least a portion of the object's face among one or more images (403). The face-containing image (404) may be an image containing at least a portion of the object's face among images associated with the text data (401).

[0092] FIG. 5 illustrates a method for acquiring object data according to one embodiment.

[0093] In FIG. 5, the object data filtering module (500) may include at least one of a face recognizer (510), a personal database manager (520), or a type organizer (530). The object data filtering module (500) may be one of the modules constituting the second search module (330) of FIG. 3. The face recognizer (510) and the personal database manager (520) may be the face recognition model and the personal database manager of the second search module (330), respectively. The face recognizer (510) may be the same as the face recognizer (420) of FIG. 4.

[0094] In one embodiment, the object data filtering module (500) may receive a face-containing image (501) and output object data (504). The face-containing image (501) may be the face-containing image (404) of FIG. 4. The object data (504) may be data regarding a target object included in the face-containing image (501).

[0095] In one embodiment, the face recognizer (510) receives a face-containing image (501) and can output face feature data (502) containing information regarding the features of the face included in the face-containing image (501). For example, the face feature data (502) may include coordinate values ​​of the face region included in the face-containing image (501) and / or a face feature vector. The face included in the face-containing image (501) may be the face of a target object.

[0096] In one embodiment, a personal database manager (520) may receive face feature data (502) and output corresponding data (503). The personal database manager (520) may search for face feature data that matches the face feature data (502) in the user's personal database. The user's personal database may store information about an object containing the object's face feature data (e.g., the object's name and / or relational title (e.g., uncle)) in a table format. The personal database manager (520) may obtain data (503) corresponding to the face feature data (502) from the user's personal database. The personal database manager (520) may obtain and output the corresponding data (503) containing the face feature data that matches the face feature data (502) in a table format. The corresponding data (503) may include data about the target object included in the face feature data (502).

[0097] In one embodiment, the type organizer (530) receives corresponding data (503) and can output object data (504) by converting the type of the corresponding data (503). In the user's personal database, data may be stored at a vector level to optimize storage space. The type organizer (530) can convert the corresponding data (503) into a natural language form. The object data (504) may be data in a natural language form.

[0098] FIG. 6 is a flowchart illustrating the operation of an electronic device (101) according to one embodiment.

[0099] In the following operational embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, or at least two operations may be performed in parallel.

[0100] According to one embodiment, operations 610 to 680 may be understood to be performed by a processor of the electronic device (101) (e.g., the preprocessor (310) of FIG. 3, the first search module (320), the second search module (330), and / or the suggestion generator (340)).

[0101] In one embodiment, an electronic device (101) (e.g., a preprocessor (310)) can acquire first data in operation 610 (e.g., see screen (200) in FIG. 2a). The electronic device (101) can receive user input that inputs the first data. The electronic device (101) can acquire the first data by user input. Refer to FIG. 2a through 2d for a description of the first data.

[0102] In one embodiment, the electronic device (101) may receive a first user input for first data in operation 620 (e.g., see screen (220) in FIG. 2a). The electronic device (101) may be configured to identify the first user input as a control operation requesting the acquisition of second data. Refer to FIG. 2a through 2d for a description of the first user input.

[0103] In one embodiment, an electronic device (101) (e.g., a first search module (320)) may provide a user interface including one or more first images regarding the first data in operation 630 (e.g., see screen (220), screen (2010) and screen (2020) of FIG. 2a). The first images may include one or more subjects. The subjects may include people and / or animals. Based on receiving a first user input, the electronic device (101) may provide a user interface including one or more first images regarding the first data. The electronic device (101) may obtain text data regarding the first data. The text data regarding the first data may be text data including the content and / or meaning of the first data. The electronic device (101) can analyze first data (e.g., text in the form of natural language and / or text-based images (e.g., text input from a drawing) using an AI-based language model (e.g., a large-scale language model (LLM)). The AI-based language model (e.g., LLM) can take the first data as input, analyze the first data to correct errors, correct the first data if it is input in an unideal state, and output the correction result. For example, if the first data is text input from a drawing, the electronic device (101) can obtain text data by performing OCR processing on the first data. The electronic device (101) can obtain text data by performing OCR based on the correction result output by the AI-based language model (e.g., LLM). For example, if the first data is image data, the electronic device (101) can obtain text data that describes the image data through image captioning. For example, if the first data is sound data, the electronic device (101) can use speech recognition technology (e.g., STT (speech-to-text) Sound data can be obtained by converting it into text data using the technology.The electronic device (101) can acquire one or more images associated with text data regarding the first data. The electronic device (101) can acquire one or more images associated with the text data from the user's personal database (e.g., an image database of a gallery application). For example, the electronic device (101) can search for words and / or sentences in natural language form in the image database of a gallery application using text data acquired through OCR. The electronic device (101) can acquire images associated with the text data among the gallery images stored in the image database of the gallery application through the search. The electronic device (101) can acquire images associated with the text data among the gallery images using an AI-based language model (e.g., LLM). One or more images associated with the text data may contain the content of the text data. The electronic device (101) can acquire a second image containing one or more objects among the one or more images associated with the text data. The second image may include the face of the object. The electronic device (101) can obtain an image containing a face from among images extracted by an AI-based language model (e.g., LLM) through face recognition filtering. There may be one or more second images. The electronic device (101) can obtain one or more first images using the second images. The second images may be identical to or different from the first images. In one embodiment, the first image may be a thumbnail image containing the face of an object included in the second image.

[0104] In one embodiment, the electronic device (101) may receive a second user input in operation 640 for selecting one or more target images among one or more first images (e.g., see screen (220) in FIG. 2a). The second user input for selecting target images may be a user input for selecting, for example, an image containing the face of a specific object to be used for generating input prompts for a generative artificial intelligence model as the target image. The electronic device (101) may identify the image selected by the second user input among one or more first images as the target image. The target image may include a target object. The target object may include one or more objects.

[0105] In one embodiment, an electronic device (101) (e.g., a second search module (330)) can obtain object data regarding a target object in operation 650. The electronic device (101) can obtain object data regarding a target object from a user's personal database. The electronic device (101) can obtain object data regarding a target object by obtaining facial features of the target object and obtaining data corresponding to facial features of the target object from a user's personal database (e.g., see FIG. 5).

[0106] In one embodiment, the electronic device (101) (e.g., the second search module (330)) may provide a user interface comprising one or more user interface objects regarding object data in operation 660 (e.g., see screen (250) in FIG. 2b and screen (2030) in FIG. 2e). The user interface object regarding object data may correspond to the source application of the object data. The user interface object regarding object data may be in the form of an icon of the source application of the object data. For example, the object data may include a schedule with an object extracted from a calendar application. In this case, the user interface object regarding object data may be in the form of an icon of the calendar application.

[0107] In one embodiment, the electronic device (101) may receive a third user input in operation 670 that selects one or more target user interface objects among one or more user interface objects (e.g., see screen (260) in FIG. 2c). The target user interface object may be related to target object data among object data. The electronic device (101) may identify the user interface object selected by the third user input among one or more user interface objects as the target user interface object.

[0108] In one embodiment, an electronic device (101) (e.g., a proposed generator (340)) may acquire second data based on at least one of first data, a target image, or target object data in operation 680 (e.g., see screen (280) and screen (290) of FIG. 2d, and screen (2030) of FIG. 2e). Refer to FIG. 2a through 2e for a description of the second data.

[0109] FIG. 7 is a flowchart illustrating the operation of an electronic device (101) according to one embodiment.

[0110] In the following operational embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, or at least two operations may be performed in parallel.

[0111] According to one embodiment, operations 710 to 790 may be understood to be performed by a processor of the electronic device (101) (e.g., the processor (120) of FIG. 1).

[0112] In one embodiment, the electronic device (101) may execute an application in operation 710. The application may be a sketch application that displays an object according to drawing input. The application may generate an image in response to an image generation request.

[0113] In one embodiment, the electronic device (101) may receive text input in operation 720. The user may input text in the form of natural language using the user's body (e.g., fingers), an external input device (e.g., a pen and / or ring), and / or a keyboard. The text may include sentences as well as words.

[0114] In one embodiment, the electronic device (101) can identify text in operation 730. The electronic device (101) can identify the meaning of the text. When text is input using a keyboard, the electronic device (101) can identify the meaning of the text relatively accurately. When text is input through an external input device, the electronic device (101) may find it relatively difficult to identify the meaning of the text. The electronic device (101) can recognize text input through an external input device as an image. The electronic device (101) can use a large language model (LLM) to analyze the meaning of text contained in the image and correct errors (typos). The electronic device (101) can use an OCR model to convert text contained in the image into a text type. The output result of the LLM is transmitted as input to an OCR model that operates based on an AI model, and the OCR model can convert the text information contained in the image into a text type and output it. The electronic device (101) can prevent the generation of inappropriate images (e.g., adult content, violent images, and / or hate images) by filtering inappropriate words and / or sentences (e.g., discriminatory and / or hate messages). The electronic device (101) can filter inappropriate words and / or sentences using an nsfw (non-safe for work) filter and / or a harness filter. When inappropriate words and / or sentences are identified, the electronic device (101) can provide a notification to the user that inappropriate words and / or sentences have been identified and can reset the entered text.

[0115] In one embodiment, the electronic device (101) may, in operation 740, acquire an associated image and display it on the display (160). The electronic device (101) may acquire an image associated with the text from the user's personal database (e.g., an image database of a gallery application) by utilizing the semantics of the identified text. The electronic device (101) may search for images associated with text in the form of words and / or sentences in the personal database (e.g., an image database of a gallery application) using an AI-based language model (e.g., a large vision model (LVM)). For example, since an AI-based language model operates in the gallery application, text in the form of words and / or sentences may be searched. For example, if an image is searched using the word 'Christmas', the AI-based language model may output images associated with Christmas, such as a Christmas tree, a snowman, falling snow, Santa Claus, or a chimney. The electronic device (101) may acquire an image containing an object among the images output by the AI-based language model as an associated image by utilizing a face recognition model. The associated image may contain an object. The face recognition model can receive input images output by an AI-based language model. The face recognition model can output coordinate values ​​of the face region in each of the input images that include a face. The electronic device (101) can extract only the face region from the image in a thumbnail form based on the coordinate values. The electronic device (101) can obtain a user interface (UI) object (e.g., a graphical user interface object) (e.g., the first image (221) of FIG. 2a)) based on the face data extracted in a thumbnail form. The electronic device (101) can obtain a user interface object in a thumbnail form.The electronic device (101) can display a user interface object through the user interface of an application and receive user input for selecting the user interface object. The user can select a user interface object containing a person for whom a related image is to be generated. The associated image selected by the user can be referenced as a target image. The target image may include a target object. The target image may include at least a portion of the face region of the target object.

[0116] In one embodiment, the electronic device (101) can retrieve associated data from a personal database in operation 750. The associated data may be relationship information with a target object included in a target image. For example, the relationship information with the object may include information regarding activities between the object and the user (e.g., a travel itinerary with the object). The electronic device (101) can use a face recognition model to identify an object (e.g., a person and / or animal) included in a user interface object selected by user input. The face recognition model may receive an image of the user interface object selected by user input. The face recognition model may receive coordinate values ​​of a face region along with the image. The face recognition model may extract a face feature vector of the object based on the coordinate values. The electronic device (101) can use the face feature vector of the object to retrieve relationship information with the object from the user's personal database. The user's personal database may store an image containing at least a part of the object's face and relationship information with the object corresponding to the image as a set. The relationship information stored in the user's personal database may include sub-data configured based on information regarding interaction with an object through an application installed on the electronic device (101). For example, if the user has a conversation with an object through a social network service (SNS) application, data summarizing the conversation content may be vectorized and stored. If the user selects the object as an object for image generation, this data may be obtained as associated data.

[0117] In one embodiment, the electronic device (101) can identify whether associated data is obtained in operation 760. The electronic device (101) may not obtain data that substantially matches the face feature vector in the user's personal database. The electronic device (101) may not obtain face feature data that substantially matches the face feature vector in the user's personal database. The electronic device (101) may obtain face feature data that substantially matches the face feature vector in the user's personal database, but may not obtain relationship information with the target object. In this case, the electronic device (101) does not obtain associated data.

[0118] In one embodiment, if associated data is not obtained, the electronic device (101) can establish a person relationship in operation 770. Establishing a person relationship may mean building a database regarding an object. The electronic device (101) can store information regarding an object in the user's personal database. For example, the user can converse with a specific object through an application installed on the electronic device (101) (e.g., a social network service application). The electronic device (101) can obtain the content of the conversation between the user and the object. The electronic device (101) can vectorize the information regarding the object (e.g., the content of the conversation between the user and the object) and store it in the personal database. In one embodiment, the electronic device (101) can automatically collect information regarding the object corresponding to the image of the object (e.g., an image including the face area of ​​the object) and establish a person relationship. If the interaction associated with an image of an object stored in the electronic device (101) (e.g., an image including the face area of ​​the object) does not satisfy the minimum conditions for establishing a person relationship (e.g., insufficient data that can be collected using the image of the object), the user's personal database may contain an image including the face area of ​​the object, but may not contain relationship information with the object corresponding to that image. If the minimum conditions for establishing a person relationship are not satisfied, the electronic device (101) cannot automatically establish a person relationship. In one embodiment, the electronic device (101) may receive user input that manually inputs information about the object to establish a person relationship. The electronic device (101) can establish a person relationship based on the user input.

[0119] In one embodiment, when associated data is obtained, the electronic device (101) may display the associated data in operation 780. The electronic device (101) may display the associated data on a display (160) via a UI (e.g., see screen (250) in FIG. 2b). The electronic device (101) may display sub-data associated with a target object included in an associated image based on receiving user input selecting (e.g., touching) a specific user interface object. The user may view the data associated with the target object and select the data among the associated data that they wish to use for image generation. The user may select one or more sub-data, or may not select any sub-data. When sub-data is selected, the electronic device (101) may pass the text entered by the user, the associated image, and / or sub-data as input to a generative AI model (e.g., suggestion generator (340) in FIG. 3). If sub-data is not selected, the electronic device (101) may transmit only the text entered by the user and / or associated images as input to a generative AI model (e.g., the suggestion generator (340) of FIG. 3). The electronic device (101) may transmit the text entered by the user, associated images, and / or associated data selected by the user as input to a generative AI model (e.g., the suggestion generator (340) of FIG. 3). The associated data selected by the user may include vectorized information.

[0120] In one embodiment, the electronic device (101) can generate an image in operation 790. The electronic device (101) can generate an image using a generative AI model. The generative AI model can generate an image by receiving a prompt consisting of a target image and / or sub-data selected by the user. The generated images may be one or multiple. The electronic device (101) can set the number of images generated according to user input.

[0121] FIGS. 8a, 8b, 8c, and 8d illustrate execution screens (810, 820, 830, 840) of an application according to one embodiment.

[0122] In FIGS. 8a, 8b, 8c, and 8d, the electronic device (101) may display recommendation data that can configure a prompt for obtaining second data (e.g., images and / or recommendation activities) and receive user input selecting at least some of the recommendation data. Based on receiving user input selecting at least some of the recommendation data, the electronic device (101) may display sub-data of the selected recommendation data as recommendation data. Based on receiving user input selecting recommendation data, the electronic device (101) may repeatedly perform the operation of providing sub-data of the recommendation data as recommendation data. The electronic device (101) may perform the operation of providing sub-data of the recommendation data as recommendation data n times (n is a natural number greater than or equal to 1). The electronic device (101) may configure a prompt for obtaining second data based on a series of recommendation data selected by the user. The electronic device (101) may configure a prompt for obtaining second data based on finally selected recommendation data.

[0123] In one embodiment, the electronic device (101) may receive user input selecting first recommendation data. Based on receiving user input selecting first recommendation data, the electronic device (101) may provide second recommendation data. The second recommendation data may be sub-data of the first recommendation data. The electronic device (101) may receive user input selecting second recommendation data. Based on receiving user input selecting second recommendation data, the electronic device (101) may provide third recommendation data. The third recommendation data may be sub-data of the second recommendation data. The electronic device (101) may receive user input selecting third recommendation data. The electronic device (101) may configure a prompt for obtaining second data based on the third recommendation data. The electronic device (101) may configure a prompt for obtaining second data based on at least one of the first recommendation data, the second recommendation data, or the third recommendation data.

[0124] In one embodiment, the electronic device (101) may be a foldable device. For example, the foldable device may include one or more hinge structures, a foldable housing comprising two or more housing portions rotatably connected by one or more hinge structures, and a flexible display disposed on the front of the foldable housing. For example, as illustrated in FIGS. 8a through 8d, the electronic device (101) may include a first display area and a second display area separated by a single hinge structure.

[0125] In one embodiment, as illustrated in FIG. 8a, the electronic device (101) may include a first display area (810) and a second display area (820). The electronic device (101) may display an execution screen of an application (e.g., a sketch application) in the first display area (810). Through the first display area (810), first data (811) and first user input may be obtained. For example, the first data (811) may include the word "Duck Boat". For example, the first user input may include a drawing input of a circular object (812). Based on the acquisition of the first data (811) and the first user input, the electronic device (101) may obtain one or more first images regarding the first data from images stored in the user's personal database (e.g., images from a gallery application). The first images may include objects. The electronic device (101) may display one or more first images in the second display area (820). The electronic device (101) may receive a second user input for selecting one or more target images (821-1) among one or more first images displayed in a second display area (820). The target image (821-1) may include a target object.

[0126] In one embodiment, as illustrated in FIG. 8b, the electronic device (101) may display a thumbnail image (831) based on a target image (821-1) in a first display area (830) based on receiving a second user input selecting one or more target images (821-1). The electronic device (101) may obtain object data regarding a target object included in the target image (821-1). The electronic device (101) may identify a target object (e.g., a father) included in the target image (821-1). The electronic device (101) may obtain object data regarding a target object from a user's personal database. The electronic device (101) may display the obtained object data (841) in a second display area (840). The electronic device (101) may display the object data (841) in a table format. For example, object data (841) may include data such as the object's relationship title (e.g., Dad), data from a social media application (e.g., conversation content shared with the object through a social media application), data from a calendar application (e.g., schedule with the object), the object's credit information (e.g., credit card information), data from an over-the-top (OTT) application (e.g., the object's OTT account information), the object's birthday, and / or data from a calling application (e.g., call recording data and / or data summarizing the call content). The electronic device (101) may receive user input selecting (e.g., touch) at least some of the object data (841) displayed in the second display area (840). For example, the electronic device (101) may receive user input selecting the value (842) of a social media item. The electronic device (101) may provide a user interface (843) that displays at least some of the items of the object data (841) in a list format. The user interface (843) may include a message such as “Please check the table and select the data to use for image generation.”The electronic device (101) can display a user interface (843) in a second display area (840). The electronic device (101) can receive user input selecting at least some of the items displayed in a list format. For example, the electronic device (101) can receive user input selecting a social media item (e.g., an Instagram item (844)). The electronic device (101) can receive user input indicating that the user input selecting at least some of the items displayed in a list format is complete (e.g., a touch on the "Create" button (845)).

[0127] In one embodiment, as illustrated in FIG. 8c, the electronic device (101) may display a first user interface object (851) representing the selected object data (842 and / or 844) based on receiving a user input selecting at least some of the object data (841) (842 and / or 844). The electronic device (101) may display the first user interface object (851) in a first display area (850). The electronic device (101) may display sub-data (861) of the selected object data in a second display area (860). The electronic device (101) may display the sub-data (861) of the selected object data in a table format. For example, the sub-data (861) may include data such as SNS posts and / or images. The electronic device (101) may receive a user input selecting (e.g., touching) at least some of the sub-data (861) displayed in the second display area (860). For example, the electronic device (101) may receive user input selecting a second SNS post (862) posted on December 13, 2024. The electronic device (101) may provide a user interface (863) that displays at least some items of sub-data (861) in a list format. The user interface (843) may include a message such as "Please check the table and select the data to use for image generation." The electronic device (101) may display the user interface (863) in a second display area (860). The electronic device (101) may receive user input selecting at least some of the items displayed in a list format. For example, the electronic device (101) may receive user input selecting an SNS post (864) posted on December 13, 2024.The electronic device (101) can receive user input (e.g., a touch on the "Create" button (865)) indicating that user input for selecting at least some of the items displayed in a list form has been completed.

[0128] In one embodiment, as illustrated in FIG. 8d, the electronic device (101) may display a second user interface object (872) representing selected sub-data (862 and / or 864) based on receiving user input selecting at least some of the sub-data (861) (862 and / or 864). The second user interface object (872) may include a thumbnail image generated based on content (e.g., a post) corresponding to the sub-data (862 and / or 864) selected by the user. The electronic device (101) may display the second user interface object (872) in a first display area (870). The electronic device (101) may display the second user interface object (872) overlapping at least partially with the first user interface object (851). The second user interface object (872) may be underlaid below the first user interface object (851), and the first user interface object (851) may be overlaid on top of the second user interface object (872). The electronic device (101) may display sub-data (881) of sub-data (862 and / or 864) in the second display area (880). The electronic device (101) may display the sub-data (881) in a table format. For example, the sub-data (881) may include data such as URL addresses associated with or included in SNS posts (e.g., SNS posts (864)) (e.g., URL addresses used to purchase products associated with SNS posts (863)). The electronic device (101) may receive user input selecting (e.g., touching) at least some of the sub-data (881) displayed in the second display area (880). The electronic device (101) may provide a user interface (882) that displays at least some items of the sub-data (881) in a list format. The user interface (882) may include a message such as "Please check the table and select the data to use for image generation."The electronic device (101) may display a user interface (882) in a second display area (880). The user interface (882) may display sub-data (881) in a list format. The electronic device (101) may receive user input selecting at least some of the items displayed in the list format. The electronic device (101) may receive user input indicating that the user input selecting at least some of the items displayed in the list format has been completed (e.g., a touch on a "Create" button). In one embodiment, the electronic device (101) may generate second data (e.g., an image) based on data regarding first data (e.g., "Duck Boat") displayed in the first display area (870), a target image (e.g., an image containing a target object riding a duck boat), and / or first and second user interface objects (851 and 872).

[0129] In one embodiment, the electronic device (101) may be a bar type, a sliderable device, and / or a rollable device. The user interface illustrated in FIGS. 10a and 10b may be provided in a similar manner in a bar type, a sliderable device, and / or a rollable device.

[0130] FIGS. 9a and 9b illustrate execution screens (910, 920, 930) of an application according to one embodiment.

[0131] In FIGS. 9a and 9b, the electronic device (101) can run an application. The electronic device (101) can acquire second data (e.g., images and / or recommendation activities) using an application (e.g., a sketch application). The electronic device (101) can receive user input for selecting data (e.g., target images and / or target user interface objects) for acquiring second data. Based on the reception of user input for selecting data for acquiring second data, the electronic device (101) can acquire second data using the selected data. The electronic device (101) can execute recommendation activities using an application installed on the electronic device (101). For example, if the data for acquiring second data includes a schedule with a specific person and / or a summary of conversations communicated with a specific person through a social media application, the electronic device (101) can save the schedule with the person to a calendar application and book a restaurant and / or accommodation based on the data.

[0132] In one embodiment, relationship information with objects may be stored in the user's personal database according to data type. The data type may be determined based on the application type. The electronic device (101) may suggest a recommendation activity based on the data type selected by the user and select an application to execute the recommendation activity based on the data type. The electronic device (101) may execute the recommendation activity using the selected application.

[0133] In one embodiment, as illustrated in the screen (910) of FIG. 9a, the electronic device (101) may display user interface objects (911, 912, 913) representing data obtained by user input. For example, the user interface objects (911, 912, 913) may include first data (e.g., "Christmas") (911), an object (912) entered by the first user input, and / or a user interface object (913: 913-1, 913-2, 913-3) selected by the user. The user interface object (913) may include one or more target images (913-1 and / or 913-2) and one or more target user interface objects (913-3). The electronic device (101) may receive user input selecting one or more target images (913-1 and 913-2). One or more target images (913-1 and 913-2) may contain one or more target objects. If information regarding a target object is obtained from a specific application (e.g., a hotel booking application), the electronic device (101) may display a user interface object in the shape of the application icon. The electronic device (101) may receive user input selecting a target user interface object (913-3). Based on the reception of user input selecting a user interface object (913), the electronic device (101) may determine whether to obtain second data through the user interface (914). The user interface (914) may include a message (915) and a user interface object (916) for requesting the acquisition of second data. For example, the message (915) may include a message such as, "Would you like to create an image and schedule using the selected [Christina] and [Hotels]?"

[0134] In one embodiment, as illustrated in the screen (920) of FIG. 9a, the electronic device (101) may provide a user interface that displays a generated image (921) in response to receiving a user input requesting the acquisition of a second data. The electronic device (101) may generate the image (921) based on information corresponding to the first data (e.g., "Christmas") (911) and / or a user interface object (913) selected by the user. For example, the information corresponding to the user interface object (913) may include information corresponding to a target user interface object (913-3) (e.g., information obtained from an accommodation reservation application). For example, the user interface may include messages (922: 922-1, 922-2) and / or the generated image (921). For example, the message (922) may include a first message indicating that image creation is complete (e.g., "Image creation is complete") (922-1) and / or a second message describing the content of the created image (921) (e.g., "I am having a home party with Christina for Christmas 2023. I have a tree decoration at home and am holding a Santa hat and props") (922-2). The electronic device (101) may obtain the second message (922-2) describing the content of the created image (921) through image captioning. The electronic device (101) may provide a user interface (923) that can edit the created image (921). The user interface (923) may provide one or more tools for editing the image (921). For example, the user interface (923) may include sticker elements (924: 924-1, 924-2, 924-3) that can be added to the image (921). The sticker elements (924) may include three sticker elements (924-1, 924-2, 924-3), but the number of sticker elements (924) is not limited to a specific number.For example, the user interface (923) may include a user interface object (e.g., "Add to image") (925) for receiving user input to add a sticker element (924) to an image (921).

[0135] In one embodiment, as illustrated in the screen (930) of FIG. 9b, the electronic device (101) may acquire and execute a recommendation activity in response to receiving a user input requesting the acquisition of a second data. The electronic device (101) may acquire data regarding a target object included in the target images (913-1 and 913-2). For example, data regarding the target object may be acquired from a first application (e.g., a social network service (SNS) application) and / or a second application (e.g., an accommodation booking application). Data regarding the target object may be stored in the user's personal database by data type based on the application type. For example, data acquired from the first application (e.g., a social network service (SNS) application) (e.g., conversation content between the user and the target object) and data acquired from the second application (e.g., an accommodation booking application) (e.g., accommodation booking history) may be stored in the user's personal database by data type. The electronic device (101) may suggest a recommended activity (e.g., restaurant and / or accommodation reservation) based on data types (e.g., conversation content between a user and a target object and / or accommodation reservation history) obtained from a first application (e.g., a social network service application) and / or a second application (e.g., an accommodation reservation application), and may execute the recommended activity using the second application (e.g., an accommodation reservation application). Based on the completion of the recommended activity execution, the electronic device (101) may display a message indicating that the recommended activity execution is complete (e.g., "Schedule created") (931) and / or content regarding the recommended activity. For example, the recommended activity may include flight reservation (935), hotel reservation (936), and / or schedule registration (937).The electronic device (101) can display icons (932, 933, 934) of an application that executes a recommendation activity along with recommendation activities (935, 936, 937). The electronic device (101) can display icons (932, 933, 934) of a corresponding application corresponding to each of the recommendation activities (935, 936, 937).

[0136] FIG. 10 is a block diagram of a generative artificial intelligence (AI) system (1000) according to one embodiment.

[0137] In FIG. 10, the generative AI system (1000) may include a user interface (1010), an AI framework (1020), a generative AI model (1030), a knowledge repository (1040), and an application / service module (1050). These components may be operated on one or more of an electronic device (101), an external electronic device (102 or 104), or a server (108). For example, the user interface (1010) and the AI ​​framework (1020) may be operated on the electronic device (101), and the knowledge repository (1040) and the generative AI model (1030) may be operated on the server (108).

[0138] According to one embodiment, a user interface (1010) may receive user input (e.g., user query). User input may be received in the form of text, images, voice (e.g., natural language), video, menu selection, or a combination thereof. The user interface (1010) may include various context information (e.g., running application or user location) related to the generative artificial intelligence system (1000) at the time the user input is received, in addition to or instead of the user input. The user interface (1010) may provide the user input or the context information to the AI ​​framework (1020) and provide the result of processing therefrom to the user, for example, through the AI ​​framework (1020). According to one embodiment, in addition to user input, the electronic device may provide context information obtained using information included on the screen to the AI ​​framework (1020). The result may be provided in the form of text, images, voice, video, an action requested by the user (e.g., execution of a specified function or app), or a combination thereof.

[0139] According to one embodiment, the AI ​​framework (1020) can identify (e.g., estimate) a user intent based on at least some user input or context information received from a user interface (1010), control each of the relevant modules (e.g., 1021, 1023, or 1025) to perform a function or action corresponding to the identified user intent, and coordinate collaboration between two or more modules. The AI ​​framework (1020) may include a prompt design module (1021), an API / plugin management module (1023), and an output modification module (1025), as illustrated in FIG. 10.

[0140] According to one embodiment, the prompt design module (1021) can generate a prompt to be input to a generative AI model (1030) based at least partially on user input or context information received from a user interface (1010). For example, the prompt design module (1021) can generate a prompt using user preferences, a prompt library, or prompt examples stored in a knowledge repository (1040) based at least partially on user input or context information.

[0141] According to one embodiment, an API / plug-in management module (1023) may communicate, for example, via an API, with various resources (e.g., a knowledge repository (1040)) that provide said additional information when there is a request for said additional information in relation to user input. Additionally or alternatively, when a specified action (e.g., a function, app, or service) is performed in response to said user input, the API / plug-in management module (1023) may request an application / service module (1050) to perform said specified action via a corresponding API. The API / plug-in management module (1023) may provide information obtained from the knowledge repository (1040), the application / service module 1050, or another external resource to the prompt design module (1021). The obtained information may be used by the prompt design module (1021) to generate a prompt together with the user input or provided to a generative AI model (1030).

[0142] According to one embodiment, an output processing module (1025) can fine-tune the results obtained through a generative AI model (1030) as at least part of the response to a user input (e.g., user query). For example, the output processing module (1025) can determine whether the content of the response obtained through the generative AI model (1030) is appropriate as a response to a request made by the user input. For example, the output processing module (1025) can determine the degree of relevance of the difference between the response obtained through the generative AI model (1030) and the user input, the degree of bias (e.g., political or social bias), or the degree of harmfulness (e.g., sexual or profanity). Additionally or generally, the output processing module (1025) can request that additional AI processing be performed on the obtained response, or provide the user with a hint to avoid unwanted output. For example, the response can be obtained again through the generative AI model (1030) by generating additional prompts through a prompt design module.

[0143] According to one embodiment, the generative AI model (1030) may form at least part of an artificial intelligence neural network and may include a model that generates images or a model that generates language. The image generation model may include, for example, a generative adversarial network (GAN), a variational autoencoder (VAE), or a Diffusion-based model using a VAE and a Transformer. The language generation model may include, for example, a large language model (LLM), a large multimodal model (LMM), a large vision model (LVM), or a large action model (LAM). The LAM may automatically generate actions for an environment (e.g., a robot, a car, an electronic device (101), or a program (140)). Additionally, for at least some AI models (e.g., LLM), there may be a low-rank adaptation (LoRA) adapter fine-tuned for, for example, a specific task or a specific situation.

[0144] FIG. 11 is a block diagram of an AI framework (1020) according to one embodiment.

[0145] In FIG. 11, the AI ​​framework (1020) may have on-device AI processing capabilities. In this case, the AI ​​framework (1020) may generate and learn a response to the user input using resources within the device, instead of sending the user input received through a user interface (1010) operating on the same device (e.g., electronic device (101)) to a generative AI model (1030) operating on an external device (e.g., server (108)), or additionally. Referring to FIG. 11, the AI ​​framework (1020) may include a cross-application action module (1110), a personal data management module (1130), an on-device AI model (1150), and an orchestration module (1170).

[0146] According to one embodiment, the cross-application action module (1110) determines one or more additional applications required for the operation of an executed application (e.g., an assistant app) and may connect or suggest operations between the app and at least one additional application, or between a plurality of additional applications. For example, the cross-application action module (1110) may execute one or more additional applications to be used to respond to a user request through the assistant app sequentially or at least partially and simultaneously. Additionally, the cross-application action module (1110) may communicate with the additional applications so that the result of the execution of one additional application (e.g., content) can be shared with other additional applications.

[0147] According to one embodiment, the personal data management module (1130) may provide personal information (e.g., schedule, contact, or message information) about a user of the application (e.g., assistant app) or the additional application running on the device (e.g., electronic device 101) or other related individuals (e.g., family or friends) to another module of the AI ​​framework (1020) or a related module (e.g., generative AI model (1030)) running on another device.

[0148] According to one embodiment, the on-device AI model (1150) may include at least one model among one or more AI models (e.g., GAN, VAE, LLM, LMM, LVM, or LAM) operated on an external device (e.g., server (108)) or a corresponding lightweight AI model. Additionally, for said model or said lightweight model, there may be, for example, a LoRA adapter.

[0149] According to one embodiment, the orchestration module (1170) may select one or more AI models to be used to obtain a response to user input (e.g., user query). For example, the orchestration module (1170) may select one or more AI models from among an on-device AI model (1150), an AI model operating on an external device (e.g., server (108)) (e.g., generative AI model (1030)), or an 11th AI model (not shown) operating on another external device. When multiple AI models are selected, the orchestration module (1170) may communicate with the selected models or devices so that the operation between the selected AI models and the processing of the results thereof can be coordinated between the relevant models or devices.

[0150] According to one embodiment, two or more modules of a generative AI system (1000) (e.g., a cross-application action module (1110) and an orchestration module (1170)) may be implemented as a single module to maintain the same functionality. Various variations are possible.

[0151] In the present disclosure, the ‘artificial intelligence model’ may be identical or similar to the configuration of the generative AI model (1030) of FIG. 10 or the on-device AI model (1150) of FIG. 11, or may be an artificial intelligence model included in the generative AI model (1030) of FIG. 10 or the on-device AI model (1150) of FIG. 11. In the present disclosure, the ‘artificial intelligence system’ may be the generative artificial intelligence system (1000) of FIG. 10.

[0152] In one embodiment, the electronic device (101) comprises: a memory (130) including at least one storage medium for storing instructions; and at least one processor (120) including a processing circuit, wherein when the instructions are executed individually or collectively by at least one processor, the electronic device comprises: an operation of acquiring first data; an operation of receiving a first user input regarding the first data; an operation of providing a user interface including one or more first images regarding the first data based on receiving the first user input, wherein the one or more first images include a subject, and the subject includes a person and / or an animal; an operation of receiving a second user input for selecting one or more target images among the one or more first images, wherein the target images include a target object; an operation of acquiring object data regarding the target object; and an operation of providing a user interface including one or more user interface objects regarding the object data. It may cause to perform an action of receiving a third user input selecting one or more target user interface objects among the one or more user interface objects above—wherein the target user interface object relates to target object data among the object data—; and an action of obtaining second data based on at least one of the first data, the target image, or the target object data—wherein the second data includes an image and / or recommendation activity regarding the target object.

[0153] In one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to perform an action of executing the recommendation activity in response to receiving a user input requesting to execute the recommendation activity.

[0154] In one embodiment, the operation of providing a user interface including one or more first images related to the first data may include: an operation of obtaining text data based on the first data; an operation of obtaining one or more images associated with the text data from a user's personal database; an operation of obtaining one or more second images among one or more images related to the text data, wherein the one or more second images include an object, and the object includes a person and / or an animal; and an operation of obtaining one or more first images using the one or more second images.

[0155] In one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to perform an operation of setting a specific first user input that specifies to acquire a specific data type when performing an operation to acquire the object data regarding the target object.

[0156] In one embodiment, the user interface object regarding the object data may correspond to the source application of the object data.

[0157] In one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to perform an operation of displaying the object data through the user interface including the one or more user interface objects regarding the object data.

[0158] In one embodiment, the operation of acquiring object data regarding the target object may include the operation of acquiring facial features of the target object; and the operation of acquiring data corresponding to facial features of the target object from a user's personal database.

[0159] In one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: perform an operation of building a database regarding the target object when the object data regarding the target object is not obtained.

[0160] In one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to perform: an operation of providing a user interface including the second image; and an operation of providing a user interface capable of editing the second image.

[0161] In one embodiment, the first data includes at least one of text data, image data, voice data, or gesture data, and the first user input may include at least one of object selection input, drawing input, or voice input.

[0162] In one embodiment, a control method of an electronic device (101) comprises: an operation of acquiring first data; an operation of receiving a first user input regarding the first data; an operation of providing a user interface including one or more first images regarding the first data based on receiving the first user input, wherein the one or more first images include a subject, and the subject includes a person and / or an animal; an operation of receiving a second user input for selecting one or more target images among the one or more first images, wherein the target images include a target object; an operation of acquiring object data regarding the target object; an operation of providing a user interface including one or more user interface objects regarding the object data; an operation of receiving a third user input for selecting one or more target user interface objects among the one or more user interface objects, wherein the target user interface objects are regarding the target object data among the object data; and may include an operation of acquiring second data based on at least one of the first data, the target image, or the target object data, wherein the second data includes an image and / or recommendation activity regarding the target object.

[0163] In one embodiment, the control method may include an operation to execute the recommendation activity in response to receiving a user input requesting to execute the recommendation activity.

[0164] In one embodiment, the operation of providing a user interface including one or more first images related to the first data may include: an operation of obtaining text data based on the first data; an operation of obtaining one or more images associated with the text data from a user's personal database; an operation of obtaining one or more second images among one or more images related to the text data, wherein the one or more second images include an object, and the object includes a person and / or an animal; and an operation of obtaining one or more first images using the one or more second images.

[0165] In one embodiment, when performing an operation to acquire object data regarding the target object, the operation may include setting a specific first user input that specifies acquiring a specific data type.

[0166] In one embodiment, the user interface object regarding the object data may correspond to the source application of the object data.

[0167] In one embodiment, the operation of displaying the object data through the user interface including one or more user interface objects regarding the object data may be included.

[0168] In one embodiment, the operation of acquiring object data regarding the target object may include the operation of acquiring facial features of the target object; and the operation of acquiring data corresponding to facial features of the target object from a user's personal database.

[0169] In one embodiment, the control method may include an operation of building a database regarding the target object when the object data regarding the target object is not obtained.

[0170] In one embodiment, the control method may include: an operation of providing a user interface including the second image; and an operation of providing a user interface capable of editing the second image.

[0171] In one embodiment, the first data includes at least one of text data, image data, voice data, or gesture data, and the first user input may include at least one of object selection input, drawing input, or voice input.

[0172] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as "coupled" or "connected" to another (e.g., 2nd) component, with or without the terms "functionally" or "communicationly," it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.

[0173] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0174] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

[0175] According to one embodiment, the method according to the various embodiments disclosed herein may be provided as included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0176] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In an electronic device (101), Memory (130) including at least one storage medium for storing instructions; and It includes at least one processor (120) including a processing circuit, and When the above instructions are executed individually or collectively by at least one processor, the electronic device: Operation of acquiring first data; The operation of receiving a first user input for the above first data; An operation of providing a user interface comprising one or more first images relating to the first data based on receiving the first user input, wherein the one or more first images comprise a subject, and the subject comprises a person and / or an animal. An operation of receiving a second user input selecting one or more target images among the above one or more first images - said target images include target objects -; An operation to acquire object data regarding the above-mentioned target object; An operation of providing a user interface including one or more user interface objects regarding the above object data; An operation to receive a third user input selecting one or more target user interface objects among the above one or more user interface objects - the target user interface object relates to the target object data among the object data -; and An operation to acquire second data based on at least one of the first data, the target image, or the target object data - the second data includes an image and / or recommendation activity regarding the target object - An electronic device that causes to perform.

2. In Paragraph 1, When the above instructions are executed individually or collectively by at least one processor, the electronic device: Action of executing the recommendation activity in response to receiving user input requesting the execution of the recommendation activity. An electronic device that causes to perform.

3. In Paragraph 1 or 2, The operation of providing a user interface including one or more first images relating to the first data above is, An operation to obtain text data based on the first data above; The operation of obtaining one or more images associated with the above text data from the user's personal database; The operation of acquiring one or more second images among one or more images relating to the above text data - said one or more second images include an object, said object includes a person and / or an animal -; and The operation of obtaining one or more first images using one or more second images. An electronic device including 4. In any one of paragraphs 1 through 3, When the above instructions are executed individually or collectively by at least one processor, the electronic device: An operation to set a specific first user input that specifies to acquire a specific data type when performing an operation to acquire the object data regarding the above-mentioned target object. An electronic device that causes to perform.

5. In any one of paragraphs 1 through 4, The user interface object regarding the object data is an electronic device corresponding to the source application of the object data.

6. In any one of paragraphs 1 through 5, When the above instructions are executed individually or collectively by at least one processor, the electronic device: An operation of displaying the object data through the user interface comprising one or more user interface objects relating to the object data. An electronic device that causes to perform.

7. In any one of paragraphs 1 through 6, The operation of acquiring object data regarding the above-mentioned target object is, An operation to acquire facial features of the above-mentioned target object; and The operation of obtaining data corresponding to the facial features of the aforementioned target object from the user's personal database. An electronic device including 8. In any one of paragraphs 1 through 7, When the above instructions are executed individually or collectively by at least one processor, the electronic device: The operation of constructing a database regarding the target object based on the fact that the object data regarding the target object is not obtained. An electronic device that causes to perform.

9. In any one of paragraphs 1 through 8, When the above instructions are executed individually or collectively by at least one processor, the electronic device: An operation of providing a user interface including the second image above; and Operation of providing a user interface capable of editing the second image above An electronic device that causes to perform.

10. In any one of paragraphs 1 through 9, The above first data includes at least one of text data, image data, voice data, or gesture data, and An electronic device wherein the first user input comprises at least one of an object selection input, a drawing input, or a voice input.

11. A method for controlling an electronic device (101), Operation of acquiring first data; The operation of receiving a first user input for the above first data; An operation of providing a user interface comprising one or more first images relating to the first data based on receiving the first user input, wherein the one or more first images comprise a subject, and the subject comprises a person and / or an animal. An operation of receiving a second user input selecting one or more target images among the above one or more first images - said target images include target objects -; An operation to acquire object data regarding the above-mentioned target object; An operation of providing a user interface including one or more user interface objects regarding the above object data; An operation to receive a third user input selecting one or more target user interface objects among the above one or more user interface objects - the target user interface object relates to the target object data among the object data -; and An operation to acquire second data based on at least one of the first data, the target image, or the target object data - the second data includes an image and / or recommendation activity regarding the target object - A control method including 12. In Paragraph 11, Action of executing the recommendation activity in response to receiving user input requesting the execution of the recommendation activity. A control method including 13. In Paragraph 11 or 12, The operation of providing a user interface including one or more first images relating to the first data above is, An operation to obtain text data based on the first data above; The operation of obtaining one or more images associated with the above text data from the user's personal database; The operation of acquiring one or more second images among one or more images relating to the above text data - said one or more second images include an object, said object includes a person and / or an animal -; and The operation of obtaining one or more first images using one or more second images. A control method including 14. In any one of paragraphs 11 through 13, An operation to set a specific first user input that specifies to acquire a specific data type when performing an operation to acquire the object data regarding the above-mentioned target object. A control method including 15. In any one of paragraphs 11 through 14, A control method in which the user interface object regarding the object data corresponds to the source application of the object data.