Artificial intelligence-based answer providing method, and electronic device and storage medium therefor

WO2026205795A1PCT designated stage Publication Date: 2026-10-01SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/003120
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-05-09
Filing Date
2026-02-25
Publication Date
2026-10-01

Smart Images

  • Figure KR2026003120_01102026_PF_FP_ABST
    Figure KR2026003120_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is an electronic device comprising at least one camera, at least one microphone, a memory, and at least one processor. According to one embodiment, the electronic device may be configured to: receive a query from a user using the at least one microphone; identify an object of interest associated with the query from an image obtained using the at least one camera; and while outputting a first answer corresponding to the query, change the length of the answer on the basis of the state of the object of interest.
Need to check novelty before this filing date? Find Prior Art

Description

AI-based answer provision method, electronic device and storage medium for the same

[0001] The embodiments disclosed in this document relate to a method for providing answers based on artificial intelligence (AI), an electronic device for the same, and a storage medium.

[0002] As the use of artificial intelligence increases, various AI systems that provide responses based on input data are being developed. For example, an electronic device can receive a user's voice command and provide a response corresponding to the voice command. The electronic device can generate a response corresponding to the voice command by converting the voice command into text data and inputting the text data into a large language model (LLM). The electronic device can generate a response based not only on voice commands but also on visual information. For example, the electronic device can generate a response by processing voice commands and visual information together using a large multi-modal model (LMM). U.S. Patent Application Publication No. 2022 / 0382989 (published December 1, 2022) discloses a method for providing a response using voice and visual information.

[0003] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art in relation to the present disclosure.

[0004] An electronic device according to one embodiment disclosed in this document may include at least one camera, at least one microphone, a memory, and at least one processor. The at least one processor may be communicatively connected to the at least one camera, at least one microphone, and the memory, and may include at least one processing circuit. The memory may store instructions that, when executed individually or in combination by the at least one processor, cause the electronic device to receive a query from a user using the at least one microphone, identify an object of interest associated with the query from an image acquired using the at least one camera, and change the length of the answer based on the state of the object of interest while outputting a first answer corresponding to the query.

[0005] Additionally, a method for changing an answer of an electronic device according to one embodiment disclosed in this document may include receiving a user query, identifying an object of interest associated with the query from an image, and changing the length of the answer based on the state of the object of interest while outputting a first answer corresponding to the query.

[0006] A computer-readable storage medium according to one embodiment disclosed in this document may store instructions that cause the electronic device to perform a method for changing the answer when executed by a processor of the electronic device.

[0007] An electronic device according to one embodiment disclosed herein may include at least one camera, at least one microphone, a memory, and at least one processor. The at least one processor may include at least one processing circuit operatively connected to the at least one camera, at least one microphone, and the memory. The memory may store instructions that, when executed individually or in combination by the at least one processor, allow the electronic device to receive a query from a user while capturing the external environment of the electronic device using the at least one camera, and to identify state information associated with said object based at least in part on a determination that an object associated with said query exists in said external environment. The instructions may, when executed individually or in combination by the at least one processor, allow the electronic device to provide a first response of a first level to said query if the state information corresponds to a first condition, and to provide a second response of a second level different from the first level to said query if the state information corresponds to a second condition.

[0008] Additionally, a method for providing a response to an electronic device according to an embodiment disclosed in this document may include: receiving a query from a user while capturing an external environment of the electronic device; identifying state information associated with an object based at least partially on a determination that an object associated with the query exists in the external environment; providing a first response of a first level to the query when the state information corresponds to a first condition; and providing a second response of a second level different from the first level to the query when the state information corresponds to a second condition.

[0009] A computer-readable storage medium according to one embodiment disclosed in this document may store instructions that cause the electronic device to perform the method for providing the answer when executed by a processor of the electronic device.

[0010] Figure 1 illustrates an example of an environment for providing answers.

[0011] FIG. 2 illustrates a block diagram of an electronic device according to one embodiment.

[0012] FIG. 3 illustrates a block diagram of a question-and-answer providing system according to one embodiment.

[0013] FIG. 4 illustrates the structure of an artificial intelligence system according to one embodiment.

[0014] FIG. 5 illustrates the structure of an artificial intelligence model module according to one embodiment.

[0015] FIG. 6 illustrates a flowchart of an object of interest tracking method according to one embodiment.

[0016] FIG. 7 illustrates a flowchart of a method for providing answers according to one embodiment.

[0017] Figure 8a illustrates an example of providing an answer.

[0018] Figure 8b illustrates one example of answer reduction.

[0019] Fig. 8c illustrates an example of providing additional answers based on interest.

[0020] Fig. 8d illustrates an example of providing additional answers based on posture.

[0021] Figure 8e illustrates an example of providing additional answers based on distance.

[0022] Figure 8f illustrates an example of providing multiple additional answers.

[0023] Fig. 8g illustrates one example of a change in the answer level.

[0024] FIG. 9a illustrates a flowchart of a method for changing an answer according to one embodiment.

[0025] FIG. 9b illustrates a flowchart of a method for changing an answer according to one embodiment.

[0026] FIG. 10 illustrates an example of logging an electronic device according to one embodiment.

[0027] FIG. 11 illustrates an example of object of interest-based logging of an electronic device according to one embodiment.

[0028] FIG. 12 is a block diagram of an exemplary electronic device capable of performing the operations described in this document.

[0029] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.

[0030] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily practice them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and brevity.

[0031] The terms used in this disclosure are used merely to describe specific embodiments and are not intended to limit the scope of other embodiments. Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as generally understood by those skilled in the art described in this disclosure. Terms used in this disclosure that are defined in a general dictionary may be interpreted as having the same or similar meaning as they have in the context of the relevant technology, and are not to be interpreted in an ideal or overly formal sense unless explicitly defined in this disclosure. In some cases, even terms defined in this disclosure are not to be interpreted to exclude the embodiments of this disclosure.

[0032] Additionally, in this disclosure, expressions such as "greater than" or "less than" may be used to determine whether a specific condition is satisfied or fulfilled; however, this is merely for the purpose of expressing an example and does not exclude descriptions of "greater than" or "less than." Conditions described as "greater than" may be replaced with "greater than," conditions described as "less than" may be replaced with "less than," and conditions described as "greater than and less than" may be replaced with "greater than and less than." Furthermore, "A" to "B" below refer to at least one of the elements from A (including A) to B (including B). Below, "C" and / or "D" refers to at least one of "C" or "D," i.e., including {"C", "D", "C" and "D"}.

[0033] Figure 1 illustrates an example of an environment for providing answers.

[0034] Referring to FIG. 1, examples of various electronic devices of the present disclosure may be described. For example, a user (1) may provide input to an electronic device using a query (90) (e.g., speech). The electronic device may provide a response to the query (90) to the user (1) based on input and context information. In the present disclosure, any device that obtains the query (90) and provides a response may correspond to the electronic device (10) described below in relation to FIG. 2. For example, a mobile phone (10a), a wearable device (10b) (e.g., an AI pin), and / or a head-mounted device (HMD) (10c) (e.g., smart glasses) of FIG. 1 may be referred to as an "electronic device" of the present disclosure. In relation to FIG. 1, a query (90) through speech is exemplified, but the query (90) of the present disclosure may be transmitted to the electronic device (10) through any input. For example, the query (90) may include touch input, input via peripheral devices (e.g., keyboard and / or mouse), input via a virtual keyboard, and / or text input.

[0035] For example, the electronic device (10) of FIG. 2 may correspond to a mobile phone (10a). A user (1) may utter a query (90) while photographing an object of interest (2) using the mobile phone (10a). For example, the query (90) may be "What is this?". The mobile phone (10a) may acquire the query (90) using a microphone and identify the object of interest (2) corresponding to the query (90) using a camera. Through the identification of the object of interest (2), the mobile phone (10a) may recognize that the query (90) is related to the object of interest (2). The mobile phone (10a) may generate an answer to the query (90) using an artificial intelligence model (e.g., LMM) and provide the answer to the user (1).

[0036] In this case, the provided answer may not match the user's (1) intention. For example, the provided answer may be too brief or too detailed. For example, the length of the provided answer may be too long or too short. For example, in response to a query (90), the mobile phone (10a) may provide model information (e.g., product name) of the object of interest (2). If the user (1) entered the query (90) to know what the use of the object of interest (2) is, the model information may not match the user's intention. If the user (1) wants more specific information about the object of interest (2), the user (1) may need to make an additional query. According to embodiments of the present disclosure, the electronic device may variably adjust the answer based on the state information of the object of interest (2). The electronic device may track the object of interest (2). For example, the electronic device may track the object of interest (2) by acquiring a plurality of images on a time axis and identifying the object of interest (2) from the plurality of images. The electronic device identifies the state information of the object of interest (2) by tracking the object of interest (2) and can dynamically change the answer based on the state information of the object of interest (2) without additional querying. For example, the electronic device can change the length, detail, tone, and / or level of the answer. By dynamically adjusting the answer, the answer can be provided without interrupting the response to the user's (1) query (90). Additionally, the user's (1) convenience can be further increased by dynamically changing the answer without additional querying.

[0037] In the example of FIG. 1, the mobile phone (10a) is described as performing the acquisition of a query (90) and the provision of an answer, but embodiments of the present disclosure are not limited thereto. In acquiring the query (90) and contextual information and providing an answer, the mobile phone (10a) may use other electronic devices. For example, the mobile phone (10a) may acquire audio data corresponding to the query (90) through an audio device (20a) (e.g., earbuds and / or a headset). Similarly, the mobile phone (10a) may acquire audio data corresponding to the query (90) using a wearable device (10b), an HMD (10c), and / or a smart watch (20b). For example, a mobile phone (10a) can obtain contextual information (e.g., visual information and / or sensor data) using other electronic devices (e.g., audio device (20a), wearable device (10b), HMD (10c), and / or smart watch (20b)).

[0038] In the example of FIG. 1, the mobile phone (10a) is described as providing an answer to the query (90), but embodiments of the present disclosure are not limited thereto. The mobile phone (10a) may provide an answer through other electronic devices. For example, the mobile phone (10a) may provide an answer through an audio device (20a), a wearable device (10b), an HMD (10c), and / or a smart watch (20b).

[0039] Hereinafter, various embodiments of the present disclosure are described with reference to FIGS. 2 through 12. Unless otherwise noted, the details described above in relation to FIG. 1 may apply to the embodiments below. The details described above in relation to FIG. 1 are intended to describe various examples of electronic devices described below and are not intended to limit the types of "electronic devices" of the present disclosure. A person skilled in the art will understand that any device capable of performing the methods described below corresponds to the "electronic devices" of the present disclosure. For example, if the electronic device that mainly performs the operations described below is an HMD (10c), the HMD (10c) may correspond to the electronic device (10) of FIG. 2.

[0040] The shapes of the mobile phone (10a), wearable device (10b), HMD (10c), audio device (20a), and / or smart watch (20b) illustrated in FIG. 1 are examples, and embodiments of the present disclosure are not limited thereto. For example, the HMD (10c) may include a VR (virtual reality) device, an XR (extended reality) device, an AR (augmented reality) device, a MR (mixed reality) device, and / or a VST (video-see-through) device. The HMD (10c) may include any electronic device combined with a head-worn assembly.

[0041] FIG. 2 illustrates a block diagram of an electronic device according to one embodiment.

[0042] Referring to FIG. 2, according to one embodiment, the electronic device (10) may include a processor (120), a memory (130), a sensor circuit (140), a display (160), a camera (170), an interface (180), and / or a communication circuit (190). The electronic device (10) may correspond to the mobile phone (10a), wearable device (10b), and / or HMD (10c) of FIG. 1. The electronic device (10) may correspond to the electronic device (1200) described below in relation to FIG. 12. The electronic device (10) may include a configuration similar to the electronic device (1200) described below in relation to FIG. 12. For example, the processor (120) may correspond to at least one processor (1210) of FIG. 12. For example, the memory (130) may correspond to the memory (1220) of FIG. 12. For example, the sensor circuit (140) may correspond to the sensor interface (1219) and / or sensor (1270) of FIG. 12. For example, the display (160) may correspond to the display (1240) of FIG. 12. For example, the camera (170) may include the image sensor (1250) of FIG. 12. For example, the communication circuit (190) may correspond to the communication circuit (1260) of FIG. 12. The configuration of the electronic device (10) shown in FIG. 2 is exemplary and is not limited thereto. For example, the electronic device (10) may further include a configuration not shown in FIG. 2 (e.g., at least one of the configurations of the electronic device (1200) of FIG. 12). For example, the electronic device (10) may not include at least one of the configurations shown in FIG. 2. A person skilled in the art will understand that, where the electronic device (10) does not include the configuration necessary to perform the operations described below, the electronic device (10) may perform at least some of the operations described below by using another electronic device connected in a communicable manner.

[0043] The processor (120) may be connected communicatively, electrically, operatively, or functionally to memory (130), sensor circuit (140), display (160), camera (170), interface (180), and / or communication circuit (190). In various embodiments of the present disclosure, when one component is connected “operatively” to another component, it may mean that the component is connected to enable the other component to operate. For example, the component may enable the other component by transmitting a control signal to the other component directly or through another component. In various embodiments of the present disclosure, when one component is connected “functionally” to another component, it may mean that the component is connected to enable the function of the other component. For example, the component may enable the function of the other component by transmitting a control signal to the other component directly or through another component.

[0044] The processor (120) may include at least one processor. The at least one processor may include at least one processing circuit. For example, the processor (120) may include an application processor (AP), a central processing unit (CPU), an image signal processor (ISP), a graphics processing unit (GPU), a neural processing unit (NPU), a tensor processing unit (TPU), and / or a communication processor (CP). The processor (120) may include at least one chip or one chipset. In the present disclosure, the processor (120) may be referred to as a hardware component having an architecture by at least one processing circuit. For example, the processor (120) may be placed on a substrate (e.g., a printed circuit board) located inside the electronic device (10) and may communicate with other components of the electronic device (10) through at least one conductive path formed in the substrate, a flexible printed circuit board (FPCB), a cable, and / or any communication channel (e.g., a wireless communication channel).

[0045] The memory (130) can store instructions. When the instructions are executed by the processor (120), they can cause the electronic device (10) to perform various operations. For example, the instructions can cause the electronic device (10) to perform various operations by being executed individually or collectively by at least one processor. In various embodiments of the present disclosure, the operation of the electronic device (10) may be referred to as an operation performed by the processor (120) by executing instructions stored in the memory (130). The memory (130) may be referred to as a hardware component for storing data.

[0046] The sensor circuit (140) may be configured to detect information associated with the electronic device (10) (e.g., contextual information). Contextual information may include optical information of the surroundings of the electronic device (10), movement information of the electronic device (10), location information of the electronic device (10), and / or information of an object of interest. For example, the sensor circuit (140) may include an image sensor (141), an inertial sensor (143), a position sensor (145), and / or a proximity sensor (147). The configuration of the sensor circuit (140) shown in FIG. 2 is exemplary and the embodiments of the present disclosure are not limited thereto.

[0047] The image sensor (141) may be configured to detect optical signals. For example, the image sensor (141) may be configured to detect light signals, infrared signals, and / or ultraviolet signals. The image sensor (141) may be configured to acquire depth information. For example, the image sensor (141) may include a light detection and ranging (LiDAR). For example, the electronic device (10) may use the image sensor (141) to identify the distance between an object of interest and the electronic device (10).

[0048] The inertial sensor (143) may be configured to detect the movement of the electronic device (10). For example, the inertial sensor (143) may include an inertial measurement unit (IMU) configured to detect the orientation, magnetic force, angular rate, and / or specific force of the electronic device (10). The inertial sensor (143) may include an accelerometer, a gyroscope, and / or a magnetometer. For example, the electronic device (10) may use the inertial sensor (143) to detect the movement of a user wearing the electronic device (10).

[0049] The location sensor (145) may include any sensor configured to acquire the geographical location of the electronic device (10). For example, the location sensor (145) may acquire geographical location information based on a global navigation satellite system (GNSS). For example, the location sensor (145) may acquire geographical location information based on the reception strength and / or angle of arrival of an external signal. In one example, the electronic device (10) may acquire geographical location information based on reception information from a base station.

[0050] The proximity sensor (147) may include any sensor for identifying whether the electronic device (10) is blocked. For example, the electronic device (10) may use the proximity sensor (147) to detect blockage of the electronic device (10) by a nearby object. Based on the detection of blockage, the electronic device (10) may adjust the length of the answer. In one example, the electronic device (10) may reduce the length of the answer if the electronic device (10) is blocked during the provision of the answer (e.g., if the electronic device (10) is placed in a bag or pocket).

[0051] In one example, the sensor circuit (140) may further include a sensor not shown in FIG. 2. The sensor circuit (140) may include a sensor configured to detect the shape of the electronic device (10). For example, the electronic device (10) may be a foldable device. In this case, the sensor circuit (140) may include a sensor for detecting the folded, unfolded, and / or half-folded states of the electronic device (10). For example, the electronic device (10) may be a sliderable device. In this case, the sensor circuit (140) may include a sensor for detecting the slide-in, slide-out, and / or half-slide-out states of the electronic device (10). For example, the electronic device (10) may be a rollable device. In this case, the sensor circuit (140) may include a sensor for detecting roll-in, roll-out, and / or half-roll-out states of the electronic device (10). The electronic device (10) may adjust the length of the response based on the detection of a change in the shape of the electronic device (10).

[0052] The display (160) may include at least one pixel configured to display an image. In one example, the display (160) may include a plurality of displays. For example, the display (160) may include a left-eye display and a right-eye display. The display (160) may include a front display and / or a rear display. The display (160) may include at least one of a see-through display, a three-dimensional display, a flexible display, a rollable display, a foldable display and / or a rigid display. In one example, the display (160) may include at least one projector for projecting an image.

[0053] The camera (170) may include at least one camera. If the electronic device (10) includes a plurality of cameras, each of the plurality of cameras may differ in at least one of the facing direction, magnification, pixels, or field of view. For example, the first camera of the electronic device (10) may be configured to acquire an image in a direction corresponding to the wearer's field of view. For example, the second camera of the electronic device (10) may be configured to acquire an image corresponding to at least a part of the wearer's body (e.g., the wearer's body, hands, face, and / or eyes). For example, the camera (170) may be configured to acquire body information (e.g., face information) for authenticating the user.

[0054] The interface (180) may include any human interface device (HID). The interface (180) may include at least one device configured to receive input. For example, the interface (180) may include a touch circuit (181) (e.g., a touch screen display) configured to receive touch input. The interface (180) may include at least one microphone (e.g., a microphone (182)) configured to receive voice input. In one example, the electronic device (10) may receive voice input corresponding to at least one designated direction by performing beamforming using a plurality of microphones. The interface (180) may include a button (185) for receiving button input. For example, the interface (180) may include a button (185). For example, the interface (180) may include a biometric sensor (e.g., a fingerprint sensor). According to one embodiment, the processor (120) may be configured to receive input using the interface (180) and to process the received input. The interface (180) may include at least one device for output. For example, the interface (180) may include a haptic module for tactile output, at least one speaker for sound output (e.g., speaker (184)), and / or an indicator (e.g., light-emitting diode).

[0055] The communication circuit (190) may be configured to perform short-range wireless communication and / or long-range wireless communication. The communication circuit (190) may include a plurality of transceiver circuits (e.g., transceivers) corresponding to each communication method. The communication circuit (190) may be connected to at least one antenna. The communication circuit (190) may include a network interface card (NIC). The processor (120) may communicate with other external electronic devices based on wireless communication and / or wired communication, for example, using the communication circuit (190). The processor (120) may communicate with external devices via an internet protocol (IP) network, for example, using the communication circuit (190). As described above in relation to FIG. 1, the electronic device (10) may be connected to at least one other electronic device so as to be communicable using the communication circuit (190). When input is obtained from another electronic device using the communication circuit (190) or data is output through another electronic device using the communication circuit (190), the communication circuit (190) may be referred to as part of the “interface” of the present disclosure.

[0056] The electronic device (10) may provide a response using an artificial intelligence system (320) described later in relation to FIGS. 3 to 5. For example, the electronic device (10) may obtain a user's query using an interface (180) and / or a communication circuit (190). The electronic device (10) may obtain contextual information based on the acquisition of the query. The electronic device (10) may obtain contextual information using an interface (180), a camera (170), and / or a sensor circuit (140). Contextual information may include, for example, visual information, auditory information, and / or sensor data. Contextual information may include, for example, operation history information of the electronic device (10).

[0057] According to one embodiment, the electronic device (10) may acquire contextual information based on the acquisition of a query. For example, the electronic device (10) may identify an object of interest associated with the query based on visual information. The electronic device (10) may identify the state of the object of interest and adjust the length, detail, tone, and / or level of the response based on the state of the object of interest. In one example, the electronic device (10) may adjust the length, detail, tone, and / or level of the response in real time based on the state of the object of interest while providing the response. In generating and adjusting the response, the electronic device (10) may use an artificial intelligence system (320) described below in connection with FIGS. 3 to 5.

[0058] FIG. 3 illustrates a block diagram of a question-and-answer providing system according to one embodiment.

[0059] Referring to FIGS. 2 and FIGS. 3, according to one embodiment, the electronic device (10) may provide a response to a query based on the query and context information, and may adjust the length, detail, tone, and / or level of the response. The configurations of the electronic device (10) described in relation to FIG. 3 may be software modules (e.g., threads, functions, databases, and / or programs) implemented by the electronic device (10) executing instructions stored in memory (130) using a processor (120). In the example of FIG. 3, the electronic device (10) may include an input and output interface module (310), an artificial intelligence system (320), and / or a response control module (330). In one example, some of the configurations of the electronic device (10) may be implemented by an external device (e.g., a server). For example, at least a part of the artificial intelligence system (320) may be implemented by an external device. In this case, the electronic device (10) can transmit and receive information by communicating with an external device using a communication circuit (190).

[0060] According to one embodiment, the input and output interface module (310) may be configured to provide an interface between a user and an artificial intelligence system (320). The input and output interface module (310) may correspond to a software module configured to perform processing of input data and output data. The input and output interface module (310) may be configured to provide input data to an answer control module (330). For example, the input and output interface module (310) may be configured to provide an interface between a database or at least one application and the artificial intelligence system (320). For example, the input and output interface module (310) may include a model context protocol. For example, through the model context protocol, an image acquired through an application (e.g., a camera application) may be transmitted to the artificial intelligence system (320) according to specified permissions or specified conditions, without the process of creating a prompt.

[0061] The input and output interface module (310) can process input data and provide the processed input data to the artificial intelligence system (320). For example, the input and output interface module (310) can acquire text data through the interface (180). The input and output interface module (310) can provide the acquired text data to the artificial intelligence system (320). For example, the input and output interface module (310) can acquire voice input (e.g., audio data) using a microphone (182) or acquire voice input from an external device using a communication circuit (190). The input and output interface module (310) can convert the voice input into a form (e.g., text) that can be input to the artificial intelligence system (320) and provide the converted voice input to the artificial intelligence system (320). For example, the input and output interface module (310) can perform preprocessing on visual data (e.g., depth data, still images and / or videos) acquired using the camera (170) or received from an external electronic device. The input and output interface module (310) can provide the preprocessed visual data to the artificial intelligence system (320). For example, the input and output interface module (310) can perform preprocessing on sensor data acquired using the sensor circuit (140) or received from an external electronic device. The input and output interface module (310) can provide the preprocessed sensor data to the artificial intelligence system (320). Similarly, the input and output interface module (310) can provide input data (e.g., text data, visual data, audio data, and / or sensor data) to the answer control module (330). In one example, the input and output interface module (310) can provide input data to the artificial intelligence system (320) and / or the answer control module (330) based on the reception of a query.

[0062] The input and output interface module (310) may be configured to output output data. For example, the input and output interface module (310) may obtain an answer generated based on input data (e.g., a question and / or a response generated based on input data) from an artificial intelligence system (320). The input and output interface module (310) may output the obtained response using the components of the electronic device (10). For example, the input and output interface module (310) may provide a visual response using a display (160). For example, the input and output interface module (310) may provide an auditory response using a speaker (184). For example, the input and output interface module (310) may provide a tactile response using a haptic module. In one example, the input and output interface module (310) may provide a response including at least two of an auditory response, a visual response, or a tactile response. The input and output interface module (310) can provide a response using a communication circuit (190). The input and output interface module (310) can provide a response through at least one external electronic device (not shown) that is communicably connected to the electronic device (10).

[0063] According to one embodiment, the artificial intelligence system (320) can process input data based on artificial intelligence. The artificial intelligence system (320) may include, for example, a generative artificial intelligence model. The generative artificial intelligence model may include, for example, a large language model (LLM), a large vision model (LVM), and / or a large multi-modal model (LMM). The artificial intelligence system (320) may be configured to process various types of input data (e.g., text data, visual data, audio data, and / or sensor data). The artificial intelligence system (320) may use the generative artificial intelligence model to generate an answer based on a query (e.g., a query based on input data). According to one embodiment, the artificial intelligence system (320) may receive an answer control prompt from an answer control module (330) and generate a response based on the answer control prompt. The response generated by the artificial intelligence system (320) may be output through an input and output interface module (310). The structure of the artificial intelligence system (320) may be described later in relation to FIGS. 4 and FIGS. 5.

[0064] According to one embodiment, the answer control module (330) can control the response based on input data. For example, the answer control module (330) can generate an answer control prompt based on input data and control the length, detail, tone, and / or level of the answer corresponding to the query by inputting the answer control prompt into the artificial intelligence system (320). For example, the answer control module (330) may include an object detection and tracking module (331), an object analysis module (333), an object state identification module (335), and / or an answer control prompt generation module (337).

[0065] The object detection and tracking module (331) may be configured to identify objects of interest based on queries and visual data, and to track the identified objects of interest. For example, the object detection and tracking module (331) may identify and / or track one or more objects of interest. The object detection and tracking module (331) may identify objects of interest based on user queries and visual information. For example, the object detection and tracking module (331) may identify objects of interest based on user queries, answers to queries, and visual information. For example, the object detection and tracking module (331) may identify objects of interest based on salient object detection (SOD). The object detection and tracking module (331) may identify the most prominent (salient) object within the image (e.g., an object in the foreground) as an object of interest. For example, the object detection and tracking module (331) may identify objects of interest based on the user's gaze direction, distance from the user, and / or interaction with the user. The object detection and tracking module (331) can identify an object located at a position corresponding to the user's gaze as an object of interest. The object detection and tracking module (331) can identify an object closest to the user as an object of interest. The object detection and tracking module (331) can identify an object that interacts with the user (e.g., an object held by the user) as an object of interest.

[0066] For example, the object detection and tracking module (331) can identify an object of interest using an artificial intelligence system (320). The object detection and tracking module (331) can identify an object of interest by inputting a prompt to the artificial intelligence system (320) that inquires about the identification of an object of interest associated with the query, along with the query. In the present disclosure, an object of interest may be referred to as an object that is the subject of the query, an object that requires observation to answer the query, or an object associated with the answer to the query. For example, an object that is the subject of the query may include an object indicated by the query or an object identified as the subject of the query when the query and visual information are combined. An object that requires observation to answer the query may include an object that is not the subject of the query but contains information necessary for the answer to the query. For example, an object associated with the answer to the query may include an object of the same type as an object included in the answer to the query. Identification of the object of interest may include obtaining information about the region of the object of interest within the image (e.g., information related to location coordinates and / or size).

[0067] For example, the object detection and tracking module (331) can transmit additional prompts to the artificial intelligence system (320) along with the query and visual information, such as “If the object mentioned in the question exists in the image, let me know in the form of (x, y, width, height)”. The object detection and tracking module (331) can identify the location of the object of interest within the image based on the answer from the artificial intelligence system (320) to the additional prompts.

[0068] For example, even if a specific object is not explicitly indicated by the query, an object of a type that may be included in the answer to the query may be referenced as an object of interest. The object detection and tracking module (331) may identify the object of interest by using an additional prompt to query for an object corresponding to the information included in the answer. In one example, a user may utter a query such as "What would be a good gift for my girlfriend's anniversary?" while looking at a flower. In this case, visual information including the flower may be acquired by the camera (170) of the electronic device (10). The artificial intelligence system (320) may generate "A letter, flowers, accessories, etc. would make a good gift" as an answer to the query. The object detection and tracking module (331) may generate an additional prompt such as "If the object related to the answer is in the video, tell me the location of the object." The object detection and tracking module (331) may input the generated additional prompt into the artificial intelligence system (320). The artificial intelligence system (320) can identify the location of a "flower" associated with the answer within an image and provide information about the identified location to the object detection and tracking module (331). The answer control module (330) can input information about the identified object (e.g., flower) and additional prompts to the artificial intelligence system (320). For example, additional prompts may include "Tell me the answer about the object associated with the answer," "Use the object currently in view in the answer," and / or "Use the flower currently in view in the additional answer." Based on the additional prompts and visual information (e.g., information about the object provided to the answer control module (330)), the artificial intelligence system (320) can generate an answer such as "Since the flower language of the flower you are looking at is farewell, it may not be a good choice as an anniversary gift."

[0069] In the examples described above with respect to the object detection and tracking module (331), methods for identifying objects of interest using an artificial intelligence system (320) have been described, but the embodiments of the present disclosure are not limited thereto. For example, the object detection and tracking module (331) may identify at least one object from an image based on any object identification method. The object detection and tracking module (331) may select an object of interest from the at least one identified object. For example, the object detection and tracking module (331) may select an object of interest among the identified objects that is highly related to the query, an object located in the center of the image, and / or an object with the largest weight within the image.

[0070] In one example, the object detection and tracking module (331) can identify multiple objects (e.g., multiple candidate objects) from an image. The object detection and tracking module (331) can identify an object of interest to be tracked among the multiple objects. The object detection and tracking module (331) can identify at least one object of interest based on the content of a query. For example, the object detection and tracking module (331) can identify the number of objects for which an answer is required by the query, and identify at least one object of interest among the multiple objects based on the number of objects for which an answer is required. The object detection and tracking module (331) can identify at least one object of interest based on the user's previous question history, functions executed by the user on the electronic device (10) (e.g., the most recently executed function by the user), and / or user preferences based on the user's profile (e.g., preferences set by user input and / or statistical preferences based on user information). For example, the object detection and tracking module (331) can identify the object of interest that is the subject of the question by considering the user's previous question history and the current question together. For example, the object detection and tracking module (331) can identify an object associated with a recently executed function (e.g., an object corresponding to the purpose of the recently executed function, the name of the recently executed function, and / or the type of the recently executed function) as the object of interest. For example, the object detection and tracking module (331) can identify the object with the highest user preference among a plurality of objects as the object of interest.

[0071] The object detection and tracking module (331) and / or the artificial intelligence system (320) can analyze a user's query and identify the number of objects for which an answer is required by the query. For example, the query may include a plural object, such as "What kind of flowers are these?". In this case, the object detection and tracking module (331) and / or the artificial intelligence system (320) can determine that the number of objects for which an answer is required by the query is plural. For example, the query may include a singular object, such as "What is this flower?". In this case, the object detection and tracking module (331) and / or the artificial intelligence system (320) can determine that the number of objects for which an answer is required by the query is singular.

[0072] If the number of objects for which an answer is required is singular, the object detection and tracking module (331) may select one of the multiple objects identified from the image as a candidate object. If the number of objects for which an answer is required is multiple, the object detection and tracking module (331) and / or the artificial intelligence system (320) may determine whether the number of objects for which an answer is required is indicated in the query. For example, if the query is "What kind of flowers are these?", the number of objects for which an answer is required is not indicated by the query. In this case, the object detection and tracking module (331) may select up to N objects (e.g., N is an integer greater than or equal to 2) among the objects identified from the image as candidate objects. For example, N may be the maximum number of objects trackable by the electronic device (10) or a specified number. For example, if the query is "Tell me the types of the 3 flowers visible," the number of objects for which an answer is required is indicated as "3" by the query. In this case, the object detection and tracking module (331) can select a specified number (e.g., 3) of objects identified from the image as candidate objects.

[0073] The object detection and tracking module (331) can identify at least one object of interest among at least one candidate object. If only one candidate object is selected, the object detection and tracking module (331) can identify that candidate object as the object of interest. If multiple candidate objects are selected, the object detection and tracking module (331) can identify at least one object of interest to be tracked. In one example, the object detection and tracking module (331) can identify at least one object of interest among multiple candidate objects based on the similarity of the object's appearance and / or traceability. The object detection and tracking module (331) can identify at least one object of interest among multiple candidate objects based on the size of the multiple candidate objects within an image and / or the distance from the electronic device (10). Among the multiple candidate objects, the size within the image can be identified based on the size of the area corresponding to each candidate object (e.g., a mask or rectangular area containing the candidate object). For example, the object detection and tracking module (331) can identify a candidate object among a plurality of candidate objects whose corresponding area size is greater than or equal to a specified value as an object of interest. For example, the object detection and tracking module (331) can identify a candidate object among a plurality of candidate objects whose distance from the electronic device (10) is within a specified distance as an object of interest. If the external similarity between the plurality of candidate objects (e.g., similarity calculated based on color and shape) is greater than or equal to a specified value, the object detection and tracking module (331) can identify the largest candidate object among the plurality of candidate objects, a candidate object among the plurality of candidate objects that is greater than or equal to a specified size, or a candidate object among the plurality of candidate objects that is closest to the electronic device (10) as an object of interest.

[0074] The object detection and tracking module (331) can perform tracking on an identified object of interest. The object detection and tracking module (331) can perform tracking on an object of interest, for example, for a specified period of time or while providing a response to a query. The object detection and tracking module (331) can identify the movement of the object of interest and / or the presence of the object of interest within an image through tracking the object of interest. The tracking results of the object detection and tracking module (331) can be used by the object analysis module (333). The object detection and tracking module (331) can track an object of interest by identifying the object of interest from a plurality of images (e.g., video). In one example, the object detection and tracking module (331) can perform marking on the identified object of interest and transmit an image containing the marked object of interest to the object analysis module (333).

[0075] The object analysis module (333) can analyze an object based on information received from the object detection and tracking module (331). The object analysis module (333) can identify the current state of the object based on the received information. For example, the current state may include object appearance time information, distance information, pose information, size information, and / or appearance information.

[0076] Object appearance time information may include information indicating the acquisition time of an image when an object of interest is present in the image (e.g., camera (170)). The object analysis module (333) can determine whether an object of interest is present in the image based on information received from the object detection and tracking module (331). If an object of interest is present in the image, the object analysis module (333) can identify the image acquisition time information as object appearance time information.

[0077] Distance information may represent information regarding the distance between an object of interest and an electronic device (10) (e.g., camera (170)). For example, the electronic device (10) may include a sensor for measuring distance (e.g., infrared sensor, LIDAR, and / or laser sensor, UWB (ultra-wideband) sensor). In this case, the object analysis module (333) may identify the distance between the object of interest and the electronic device (10) based on information obtained from the sensor. In one example, the object analysis module (333) may identify the distance information based on the size of the object of interest identified within an image. The distance information may include information regarding the absolute distance and / or relative distance between the electronic device (10) and the object of interest.

[0078] Pose information may include the movement and rotation of the object of interest identified based on the image of the object of interest within the image. In one example, the object analysis module (333) may identify rotations about three orthogonal axes based on 6-degree-of-freedom (DoF) pose estimation. In one example, the object analysis module (333) may estimate the rotation of the object of interest based on changes within the image of the object of interest. In one example, the object analysis module (333) may identify the rotation of the object of interest by recognizing the object of interest using an artificial intelligence model.

[0079] Appearance information may include, for example, the color, shape, and / or feature values ​​of the object of interest. For example, the appearance information of the object of interest may include shape and color. For example, the appearance information of the object of interest may include feature values ​​or vector values ​​extracted from an image of the identified object of interest. For example, feature values ​​may include information on feature points extracted based on image differentiation, edge extraction, corner extraction, point of interest extraction, Hough line transformation, SITF (scale-invariant feature transformation), SURF (speeded-up robust features), FAST (features from accelerated segment test), and / or any image feature extraction technique. For example, vector values ​​may include a vector of feature values ​​generated by encoding the image.

[0080] Size information may include information indicating the size of an image within an object of interest. For example, size information may include the size (e.g., width) of a mask corresponding to the object of interest. For example, size information may include the size (e.g., width or the length of the diagonal between two vertices) of a rectangle (e.g., frame) indicating the object of interest. In one example, the object analysis module (333) may identify distance information based on size information.

[0081] In one example, the electronic device (10) may provide a visual frame surrounding the object of interest or an identifier indicating the object of interest. When the state of the object of interest (e.g., pose, distance, appearance, and / or size) analyzed by the object analysis module (333) changes, the electronic device (10) may provide an indication indicating the state change through the visual frame. For example, when a significant change in the state of the object of interest (e.g., a change exceeding a specified threshold level) is detected, the electronic device (10) may provide an indication by changing the color, size, and / or shape of the visual frame.

[0082] The object state identification module (335) can identify the state of an object of interest using information of the object of interest analyzed by the object analysis module (333). For example, the object analysis module (333) can identify the current state information of the object of interest for each image (e.g., frame). The object state identification module (335) can identify the object state based on the current state information of each of the multiple images. For example, the object state identification module (335) can compare the current state information of the previous frame and the subsequent frame, and identify the state of the object of interest based on the comparison.

[0083] The object state identification module (335) can identify the appearance time period information of an object of interest based on object appearance time information. The appearance time period information may include information indicating how long the object of interest has appeared in the images. For example, the object state identification module (335) can set the state of the object of interest to 'continuously observed' if the appearance time period information exceeds a specified time length (e.g., 20 seconds).

[0084] The object state identification module (335) can identify the relative movement of the object of interest to the electronic device (10) based on size information and / or distance information. For example, the object state identification module (335) can set the state of the object of interest to 'closer' if the object of interest moves within a specified first distance. For example, the object state identification module (335) can set the state of the object of interest to 'further away' if the object of interest moves outside a second distance.

[0085] The object state identification module (335) can identify the rotation of the object of interest based on pose information. For example, if the object of interest rotates by more than a specified angle (e.g., 30 degrees) relative to an arbitrary axis, the object state identification module (335) can set the state of the object of interest to 'rotated'.

[0086] The object state identification module (335) can identify changes in the appearance of an object of interest based on appearance information. For example, if the color or shape of the object of interest is changed, the object state identification module (335) can set the state of the object of interest to 'appearance changed'. In one example, changes in the shape of the object of interest may include changes in the captured image of the object of interest due to rotation of the object of interest.

[0087] The object state identification module (335) can identify the state of an object of interest based on the tracking result of the object of interest. For example, the object detection and tracking module (331) may fail to identify the object of interest from an image. If the identification of the object of interest fails for a specified amount of time or longer, the object state identification module (335) may set the state of the object of interest to 'tracking failed'.

[0088] According to one embodiment, the answer control prompt generation module (337) may determine whether to generate an answer control prompt based on the state of the object of interest. Additionally, the answer control prompt generation module (337) may generate an answer control prompt based on the determination. For example, if the state of the object of interest identified by the object state identification module (335) satisfies a specified condition, the answer control prompt generation module (337) may generate an answer control prompt. If the state of the object of interest identified by the object state identification module (335) does not satisfy a specified condition, the answer control prompt generation module (337) may not generate an answer control prompt. For example, the answer control prompt (337) may track the state of the identified object of interest while an answer to a query is being provided. Based on the tracked state of the object of interest, the answer control prompt generation module (337) may determine whether to generate an answer control prompt. When it is determined that an answer control prompt is created, the answer control prompt creation module (337) creates an answer control prompt and can input the created answer control prompt into the artificial intelligence system (320).

[0089] For example, the state of the object of interest may remain 'continuously observed' even at the time when the answer to the query is completed. In this case, the user's intention may be to request additional information about the object of interest. For example, the answer control prompt generation module (337) may determine the generation of an answer control prompt. The answer control prompt generation module (337) may generate an answer control prompt that requests additional information about the object of interest. The answer control prompt generation module (337) may provide additional answers about the object of interest by inputting the answer control prompt into the artificial intelligence system (320). For example, the additional answers may include visual, auditory, and / or tactile provision of additional information. By providing the previously provided answers and additional answers, the electronic device (10) may change the length, level, tone, and / or detail of the answers.

[0090] For example, the status of the object of interest may change to 'Tracking failed' before the answer to the query is completed. In this case, the user may have lost interest in the object of interest. For example, the answer adjustment prompt generation module (337) may generate an answer adjustment prompt to reduce the length of the answer. For example, the answer adjustment prompt generation module (337) may generate an answer adjustment prompt to provide a shortened answer. The answer adjustment prompt generation module (337) may provide a simplified answer to the object of interest by inputting the answer adjustment prompt into the artificial intelligence system (320). In this case, the electronic device (10) may change the length, level, tone, and / or detail of the answer by stopping the provision of the answer currently being provided or by providing a simplified answer instead of the answer currently being provided.

[0091] For example, while providing an answer to a query, the posture and / or appearance of the object of interest may change. In this case, the answer adjustment prompt generation module (337) may generate an answer adjustment prompt that requests additional information based on the changed posture and / or appearance.

[0092] In one example, the answer control prompt generation module (337) may generate an answer control prompt based on the state and / or context information of the object of interest. The state of the object of interest may include the shape of the object of interest, the time of appearance, and / or distance. The shape of the object of interest may include the posture, appearance, and / or direction (e.g., the direction in which it is observed) of the object of interest. The time of appearance of the object of interest may indicate the length of time the object of interest was observed by the electronic device (10). The distance of the object of interest may indicate the difference between the distance at which the object of interest was first observed and the distance at which the object of interest is currently observed. The context information may include any information stored in the electronic device (10). For example, the context information may include a history of queries similar to the query, content stored in the clipboard, content stored in the buffer, the function of the electronic device (10) used immediately before the query, the location of the electronic device (10), and a user profile (e.g., the user's age group, gender, and / or field of interest).

[0093] For example, while providing an answer to a query, the distance of the object of interest may change. If the distance of the object of interest decreases, the user may want to observe the object of interest more specifically. In this case, the answer adjustment prompt generation module (337) may generate an answer adjustment prompt requesting additional information about the object of interest. If the distance of the object of interest increases, the answer adjustment prompt generation module (337) may generate an answer adjustment prompt requesting a reduction of the answer regarding the object of interest.

[0094] In the examples described above in relation to FIG. 3, examples of response control based on the state of the object of interest have been described. Changes in the state of the object of interest may be, for example, caused by interaction between the user and the object of interest. Thus, response control based on the state of the object of interest may be referred to as response control based on interaction. For example, as the user's interaction with the object of interest increases, the electronic device (10) may increase the length, level, and / or detail of the response. If the user's interaction with the object of interest decreases or stops, the electronic device (10) may decrease the length, level, and / or detail of the response. In the present disclosure, increasing the length of the response may include providing additional responses. In the present disclosure, decreasing the length of the response may include discontinuing the provision of a previously provided response and / or providing a shortened response instead of a previously provided response. In the present disclosure, an increase in the detail of an answer may include an increase in the amount of information about the object of interest through the provision of additional answers, an increase in the modality of the answer, and / or an increase in the amount of information about the object of interest through modification of an existing answer. In the present disclosure, a decrease in the detail of an answer may include providing an abbreviated version of an existing answer, discontinuing the provision of an existing answer, and / or a decrease in the modality of the answer. In the present disclosure, a change in the level of an answer may include a change in length, a change in detail, and / or a change in modality.

[0095] FIG. 4 illustrates the structure of an artificial intelligence system according to one embodiment.

[0096] Referring to FIG. 4, in an artificial intelligence system (320), an input and output interface module (310) can receive user input. User input may include forms such as natural language, images, and / or videos. Additionally, contextual information may be transmitted along with the user input. Contextual information may include various additional information at the time of user input. For example, contextual information may include information about the application currently being used by the user or the user's location information. Additionally, user input may be in a mixed form of the aforementioned natural language, images, sounds, and contextual information. Furthermore, user input may include input in a form other than natural language, such as selecting a menu.

[0097] According to one embodiment, the input and output interface module (310) can output results of the artificial intelligence system (320) to a user. The input and output interface module (310) can output results in the form of natural language or specific content. The output of results can be provided in the form of an action requested by the user.

[0098] According to one embodiment, the artificial intelligence framework (420) receives user input through the input and output interface module (310) and can coordinate and control each component or module necessary to perform the user's intention based on the user's query. The artificial intelligence framework (420) can be configured to perform orchestration of the artificial intelligence system (320).

[0099] According to one embodiment, user input received from an input and output interface module (310) may be transmitted to a prompt design module (421). The prompt design module (421) may generate a prompt suitable for a large language model (LLM) or a large multi-modal model (LMM) based on the user input. For example, the prompt design module (421) may be an artificial intelligence component using a machine learning algorithm or a neural network. The prompt design module (421) may generate a prompt by accessing a database (430) containing user preference data, a prompt library, and prompt examples based on the user input, and transmit the generated prompt to an artificial intelligence model module (450). In one example, the database (430) may store content stored in the clipboard, question history, content stored in the buffer, function execution history, and / or user profile.

[0100] According to one embodiment, the API (application programming interface) / plugin management module (422) can obtain information from outside the artificial intelligence system (320) using the API and / or plugin. For example, the API / plugin management module (422) can obtain information from outside based on a request for additional information when transmitting user input to the artificial intelligence model module (450). For example, the API / plugin management module (422) can establish a channel to communicate with the outside of the artificial intelligence system (320) using the API and / or plugin, and can access various data sources through the established channel. Information obtained from the outside can be used to generate a prompt by the prompt design module (421) or as input to the artificial intelligence model module (450).

[0101] According to one embodiment, the API / plugin management module (422) may provide an interface between the artificial intelligence framework (420) and the application and / or application / service module (440). For example, an action that performs user input, rather than an intermediate result, may need to be performed in the application / service module (440). In this case, the API / plugin management module (422) may request the action from the application via the API.

[0102] According to one embodiment, the transformation module (423) can finely tune the output of the artificial intelligence model module (450). For example, the transformation module (423) can evaluate the relevance of the content and query generated through LLM and / or LMM, the bias of the content, or the harmfulness of the content. If the transformation module (423) evaluates that the relevance between the generated content and the query is low, it can tune the output result through an additional process. If the transformation module (423) evaluates that the content is biased or determines that the content is harmful, it can tune the output result through an additional process or refrain from outputting the content. Additionally, the transformation module (423) can configure and provide a hint to the user to avoid unwanted output.

[0103] According to one embodiment, the artificial intelligence model module (450) may generally refer to an artificial intelligence neural network that generates new forms of data based on user input information. For example, the artificial intelligence model module (450) may include a model that generates images and / or a model that generates language. Models that generate images include, typically, GANs (generative adversarial networks) and VAEs (variational autoencoders), and examples include diffusion-based generative models using VAEs and Transformer structures. Models that generate language are models trained to output the statistically most appropriate output value based on input values, and examples include models such as CHAT-GPT 3 and CHAT-GPT 4. Additionally, the artificial intelligence model module (450) may include an LMM capable of recognizing various forms of data input, such as text, images, and voice, and generating new data corresponding to them.

[0104] FIG. 5 illustrates the structure of an artificial intelligence model module according to one embodiment.

[0105] Referring to FIG. 5, the structure of an artificial intelligence model module (450) according to one embodiment may be illustrated. For example, the artificial intelligence model module (450) may include an input conversion module (510) and an artificial intelligence model (530). The structure of the artificial intelligence model module (450) illustrated in FIG. 5 is an example, and the embodiments of the present disclosure are not limited thereto. In the example of FIG. 5, an artificial intelligence model module (450) configured to process various types of inputs is illustrated, but in one example, the artificial intelligence model module (450) may be configured to process only a specified type of data (e.g., text data). For example, at least a part of the input conversion module (510) may be included in the conversion module (423) of FIG. 4.

[0106] The input conversion module (510) can convert various types of data into a form that can be processed by the artificial intelligence model (530). The input conversion module (510) can convert, for example, image data (501), video data (502), audio data (503), and / or sensor data (504) into a form that can be processed by the artificial intelligence model (530). The input conversion module (510) can, for example, convert the data into a text format to generate input data (520) together with text data (505). The input conversion module (510) can, for example, input at least a portion of the data into an internal layer of the artificial intelligence model (530).

[0107] For example, the artificial intelligence model (530) may include an embedding layer (531) that receives input data (520) as input. In the example of FIG. 5, the artificial intelligence model (530) is shown to include a first inner layer (532), a second inner layer (533), and a third inner layer (534), but the number of inner layers is not limited to that shown in FIG. 5. The artificial intelligence model (530) may be configured to generate output data (540) using the input data (520) and / or the converted input of the input conversion module (510).

[0108] In one example, the input data (520) may correspond to text data (505). The text data (505) may correspond to text generated by the prompt design module (421) of FIG. 4. For example, the input conversion module (510) may include an encoder and a resampler configured to process various types of data. Image data (501), video data (502), audio data (503), and / or sensor data (504) may be encoded through an encoder configured to process the corresponding type of data. The encoded data may be converted into tokens of a preset number of modalities through a resampler. The tokens converted by the input conversion module (510) may be input into a cross-attention layer among the internal layers of the artificial intelligence model (530). Through the cross-attention layer, the tokens may be fused with the input data (520).

[0109] In one example, the input transformation module (510) can encode data using an encoder and align inputs of different modalities. The aligned inputs can be processed by a separate cross-attention layer and then input into the internal layers of the artificial intelligence model (530). Fusion between the modalities of the data can be performed using the separate cross-attention layer.

[0110] In one example, the input conversion module (510) can convert the input data into a text format. The converted text can form the input data (520) together with the text data (505). In this case, the fusion of multiple modality inputs can occur at the input stage of the artificial intelligence model (530).

[0111] In one example, the input conversion module (510) can tokenize the input data. The tokenized data can form the input data (520) together with the text data (505).

[0112] Output data (540) may include, for example, data in text format. Output data (540) may be provided to the user in one or more formats. For example, output data (540) may be in text format. In this case, the electronic device may provide output data (540) audibly and / or visually. In one example, output data (540) may include information (e.g., code, command, and / or link) to which data can be output. The electronic device may provide a response corresponding to output data (540) by performing an action based on output data (540). In this case, the response may include a visual, auditory, and / or tactile response. In some examples, output data (540) may include image data, video data, audio data, and / or text data. Similar to the examples described above, the electronic device may provide a visual, auditory, and / or tactile response based on output data.

[0113] FIG. 6 illustrates a flowchart of an object of interest tracking method according to one embodiment.

[0114] Referring to FIGS. 2 and FIGS. 6, according to one embodiment, an electronic device (10) can perform identification, tracking, and state identification of an object of interest. An operation described below in relation to FIGS. 6 may be referred to as an operation performed in the processor (120) of the electronic device (10) of FIGS. 2. The order of the operations described below in relation to FIGS. 6 is an example, and embodiments of the present disclosure are not limited thereto. For example, at least some of the operations may be performed differently from the order of FIGS. 6 or may be performed substantially simultaneously with other operations of FIGS. 6.

[0115] According to one embodiment, in operation 605, the electronic device (10) can obtain a query. For example, the electronic device (10) may have a voice assistant activated. The voice assistant may be activated by user input (e.g., touch input and / or wake-up utterance). In one example, the electronic device (10) can obtain a query from a user's utterance. The electronic device (10) can obtain audio data corresponding to the user's utterance using a microphone (182) and obtain text data from the audio data based on speech recognition of the audio data. The electronic device (10) can determine whether the text data contains a query based on natural language understanding of the text data. If the text data contains a query, the electronic device (10) can obtain the query from the user's utterance. In one example, the electronic device (10) can obtain a query from user input (e.g., user input via a touchscreen and / or button). The electronic device (10) can obtain a query from user input by performing natural language understanding on text data entered by the user.

[0116] According to one embodiment, in operation 610, the electronic device (10) can identify an object. For example, the electronic device (10) can identify an object based on the acquisition of a query. The electronic device (10) can acquire at least one image using a camera (170). The electronic device (10) can activate the camera (170) in response to the acquisition of a query or in response to the reception of user input, and can acquire at least one image using the activated camera (170). The electronic device (10) can identify an object from at least one image. For example, the electronic device (10) can identify at least one object using any object identification technique. For example, the electronic device (10) can identify at least one object from at least one image by inputting at least one image into the artificial intelligence system (320) of FIG. 3.

[0117] In one example, the electronic device (10) may fail to identify an object from at least one image. In this case, the electronic device (10) may provide a response based on an acquired query without considering the image.

[0118] According to one embodiment, in operation 615, the electronic device (10) can identify an object of interest. The electronic device (10) can identify at least one object of interest among the identified objects. The object of interest may be an object associated with a query and / or answer. For example, the electronic device (10) can identify the object of interest using an artificial intelligence system (320). The electronic device (10) can identify the object of interest using the object detection and tracking module (331) of FIG. 3. For example, the electronic device (10) can identify the foreground and background of an image. The electronic device (10) can identify candidate objects included in the foreground. The electronic device (10) can identify the names and / or types of the candidate objects. The electronic device (10) can identify at least one of the identified candidate objects as at least one object of interest based on query history, execution function history, and / or preference.

[0119] According to one embodiment, in operation 620, the electronic device (10) can perform tracking of an object of interest and state identification. The electronic device (10) can perform tracking of an object of interest and state identification based at least partially on the provision of an answer to a query. For example, the electronic device (10) can perform tracking of an object of interest and state identification from the time of receiving the query until the time when the provision of the answer is completed (e.g., within a specified time after the completion of the provision of the answer). For example, tracking of an object of interest may include tracking of the location, distance, pose, size, and / or appearance of the object of interest. The state of the object of interest may be based on object appearance time information, distance information, pose information, size information, and / or appearance information. As described above in relation to FIG. 3, the electronic device (10) can identify the state of an object of interest and detect changes in the state of the object of interest by tracking the object of interest. For example, the electronic device (10) can detect a significant change in the state of the object of interest (e.g., a change exceeding a threshold level). For example, the electronic device (10) can perform tracking and status identification of an object of interest according to the method described above in relation to the object detection and tracking module (331), object analysis module (333), and object status identification module (335) of FIG. 3.

[0120] Hereinafter, with reference to FIG. 7, a method of providing a response by an electronic device (10) may be described. For example, while the methods of FIG. 7 are being performed, it may be assumed that the electronic device (10) performs operation 620.

[0121] FIG. 7 illustrates a flowchart of a method for providing answers according to one embodiment.

[0122] Referring to FIGS. 2 and FIGS. 7, according to one embodiment, an electronic device (10) may provide an answer to a query. For example, the electronic device (10) may be assumed to have obtained a query in accordance with operation 605 of FIG. 6. An operation described below in relation to FIG. 7 may be referred to as an operation performed in the processor (120) of the electronic device (10) of FIG. 2. The order of the operations described below in relation to FIG. 7 is an example, and embodiments of the present disclosure are not limited thereto. For example, at least some of the operations may be executed differently from the order of FIG. 7 or may be executed substantially simultaneously with other operations of FIG. 7.

[0123] According to one embodiment, in operation 705, the electronic device (10) may begin providing an answer to a query. For example, the electronic device (10) may obtain a query through operation 605 of FIG. 6. The electronic device (10) may generate an answer to the query by inputting the obtained query into the artificial intelligence system (320) of FIG. 3. In one example, the electronic device (10) may generate an answer to the query by inputting context information (e.g., visual information, auditory information, and / or sensor data) along with the obtained query into the artificial intelligence system (320). In one example, the electronic device (10) may generate an answer to the query by inputting information of an identified object of interest along with the obtained query into the artificial intelligence system (320). The electronic device (10) may provide the generated answer auditorily, visually, or tactilely.

[0124] According to one embodiment, in operation 710, the electronic device (10) can identify the state of an object of interest. For example, the electronic device (10) can identify the state of an object of interest using the object state identification module (335) of FIG. 3. The state of the object of interest can be identified based on the tracking results of the identified object of interest. The state of the object of interest may include, for example, a change in distance of the object of interest, a change in size of the object of interest, a change in attitude of the object of interest, a change in appearance of the object of interest, and / or the tracking state of the object of interest.

[0125] According to one embodiment, in operation 715, the electronic device (10) may determine whether to adjust the length of the state-based answer. For example, the electronic device (10) may adjust the length of the answer if the state of the object of interest satisfies a specified condition. For example, in operation 705, the electronic device (10) may provide a linguistic form of answer (e.g., a response via voice and / or text). In this case, the adjustment of the length of the answer may include at least one of changing the number of sentences included in the answer (e.g., a decrease and / or increase) or changing the length of a single sentence (e.g., abbreviation and / or addition of explanation). For example, the electronic device (10) may adjust the length of the answer by inputting a prompt to the artificial intelligence model to adjust the length of the answer. As described below, the electronic device (10) may adjust the length of the answer by using an answer having a hierarchical structure. The electronic device (10) may adjust the length of the answer by inputting a prompt to the artificial intelligence model to change the hierarchy of the answer having a hierarchical structure. For example, a prompt instructing a change in hierarchy may include "provide an answer for the upper hierarchy (or upper concept)" or "provide an answer for the lower hierarchy (or lower concept)." For example, the electronic device (10) may convey examples of the upper hierarchy (or upper concept) and the lower hierarchy (or lower concept) to the artificial intelligence model via the prompt. For example, the electronic device (10) may convey examples of answer adjustments based on changes in the state of an object to the artificial intelligence model via the prompt. The electronic device (10) may determine whether the state of the object of interest satisfies a specified condition based on the interaction of the user of the electronic device (10) with the object of interest.

[0126] In one example, the specified conditions may include a condition for shortening the answer and a condition for adding the answer. If the condition for shortening the answer is satisfied, the electronic device (10) may stop providing the answer currently being provided or provide a shortened answer instead of the current answer currently being provided. If the condition for shortening the answer is satisfied, the electronic device (10) may reduce the modality of the answer currently being provided. If the condition for adding the answer is satisfied, the electronic device (10) may provide an additional answer (e.g., additional sentences) in addition to the answer currently being provided or increase the level of detail of the answer currently being provided (e.g., addition of explanations). If the condition for increasing the answer is satisfied, the electronic device (10) may increase the modality of the answer currently being provided.

[0127] In one example, the answers may have a hierarchical structure. The answers in the upper layer may have a relatively shorter length (e.g., relatively less information) compared to the answers in the lower layer. The answers in the lower layer may have a relatively longer length (e.g., relatively more information) compared to the answers in the upper layer. According to one embodiment, the electronic device (10) may reduce the answers by providing the answers in the upper layer or increase the answers by providing the answers in the lower layer.

[0128] For example, the answer may include multiple segments. A segment may correspond to a sentence, a phrase, and / or a clause. For example, the answer may consist of segments ABC. In one example, the electronic device (10) may provide segments ABC sequentially. For example, segment C may consist of multiple layers. For example, segment C may include a top layer C1, a middle layer C2, and a lower layer C3. In this example, it may be assumed that the electronic device (10) is providing segment B. If a specified condition for answer adjustment is not satisfied during the provision of segment B, for example, the electronic device (10) may provide an answer of a preset layer (e.g., middle layer C2). If a condition for answer reduction is satisfied during the provision of segment B, the electronic device (10) may provide an answer of a higher layer (e.g., upper layer C1). If the condition for increasing the answer is satisfied during the provision of segment B, the electronic device (10) may provide a lower-level answer (e.g., lower-level C3). As described above, to increase the answer, the electronic device (10) may add a new segment (e.g., segment D). Additionally, the electronic device (10) may not output the remaining segment (e.g., segment C) to reduce the answer.

[0129] The electronic device (10) may determine that the answer shortening condition is satisfied if it identifies that the user’s interaction with the object of interest has stopped or decreased. For example, the electronic device (10) may fail to track the object of interest during the provision of the answer. In this case, the user’s interaction with the object of interest may have ended. For example, the electronic device (10) may identify that the user is holding the object of interest and then puts it down. In this case, the interaction with the object of interest may have ended. If the distance of the object of interest increases or the size of the object of interest decreases (e.g., the size within the image of the object of interest captured by the camera (170)), the electronic device (10) may determine that the interaction with the object of interest has decreased. If the object of interest moves from the center to the periphery of the image captured by the camera (170), the electronic device (10) may determine that the interaction with the object of interest has decreased.

[0130] The electronic device (10) may determine that the condition for adding an answer is satisfied when it is identified that the user's interaction with the object of interest is maintained or increased. For example, the object of interest may be identified from an image acquired by the camera (170) at the time when the output for the query is completed. In this case, it may be determined that the user's interaction with the object of interest is maintained. For example, the electronic device (10) may determine that the user is still holding the object of interest in their hand at the time when the output for the query is completed. If the user is still holding the object of interest in their hand, the electronic device (10) may determine that the interaction with the object of interest is maintained. For example, if the size of the object of interest increases or the distance of the object of interest decreases, the electronic device (10) may determine that the interaction with the object of interest has increased. For example, if the state of the object of interest changes (e.g., color change by user operation and / or rotation by user), the electronic device (10) may determine that the interaction with the object of interest has increased.

[0131] In one embodiment, when adjusting the answer length (e.g., operation 715-YES) in operation 720, the electronic device (10) can adjust the length of the answer based on the answer adjustment prompt. For example, the electronic device (10) can generate an answer adjustment prompt based on the satisfaction of a specified condition. For example, the electronic device (10) can generate an answer adjustment prompt using the answer adjustment prompt generation module (337) of FIG. 3.

[0132] If the condition for shortening the answer is satisfied, the electronic device (10) may generate an answer control prompt to reduce the length of the answer, or stop the output of the current answer currently being output. For example, reducing the length of the answer may include stopping the provision of the answer, providing a shortened answer, and / or reducing the modality of the answer. For example, the answer control prompt may request the shortening of the current answer. The electronic device (10) may generate a shortened answer by inputting the answer control prompt into the artificial intelligence system (320). For example, the answer control prompt may request the shortening of a part of the current answer that has not yet been output. The electronic device (10) may change the length of the answer by providing a shortened answer. The electronic device (10) may stop the output of the current answer currently being output and output a shortened answer.

[0133] If the condition for adding an answer is satisfied, the electronic device (10) may generate an answer adjustment prompt that causes the length of the answer to increase. Increasing the length of the answer may include, for example, providing an additional answer and / or increasing the modality of the answer. For example, the answer adjustment prompt may request more detailed information about the object of interest and / or information different from that in the current answer about the object of interest. The electronic device (10) may generate an additional answer by inputting the answer adjustment prompt into the artificial intelligence system (320). The electronic device (10) may change the length of the answer by providing an additional answer. The electronic device (10) may provide an additional answer after the output of the current answer currently being output is completed, output an additional answer together with the current answer, or output an additional answer instead of the current answer.

[0134] In one embodiment, when the answer length is not adjusted (e.g., operation 715-NO), in operation 725, the electronic device (10) can determine whether the answer provision is complete. For example, the electronic device (10) can determine that the answer provision is complete when the output of the answer to the query is completed, or when a specified time interval has elapsed after the output of the answer is completed. If an additional answer is provided or the length of the answer is changed, the electronic device (10) can determine that the answer provision is complete when the additional answer provision is completed or when the provision of the answer with changed length is completed.

[0135] If the provision of the answer is not completed (e.g., operation 725-NO), the electronic device (10) may continue to monitor the status of the object of interest. If the provision of the answer is completed (e.g., operation 725-YES), the electronic device (10) may terminate the identification of the status of the object of interest based on the completion of the provision of the answer in FIG. 7. In this case, the electronic device (10) may terminate the tracking of the object of interest. For example, the electronic device (10) may disable the camera (170) for tracking the object of interest.

[0136] Figure 8a illustrates an example of providing an answer.

[0137] Referring to FIGS. 2 and FIGS. 8a, for example, an electronic device (10) may acquire a first utterance (91) of a user (1) at time t0. For example, the first utterance (91) may include a query. As described above in relation to FIGS. 1 through 7, the electronic device (10) may identify the query from the first utterance (91) and identify an object of interest associated with the query from an image. For example, the first image (851) may represent an object of interest identified from an image acquired using a camera (170). The electronic device (10) may provide a first answer (801) by inputting information of the object of interest (e.g., object recognition result and / or image) and the query into an artificial intelligence system (e.g., the artificial intelligence system (320) of FIG. 3). For example, the first answer (801) may have a length from time t1 to time t3.

[0138] For example, the first utterance (91) may be "What kind of phone is this smartphone?". Based on the first utterance (91), the electronic device (10) may identify an image of a mobile phone from the first image (851). The electronic device (10) may identify a mobile phone corresponding to the "smartphone" indicated by the query from the first image (851). The electronic device (10) may generate input data based on the query and the first image (851), and generate a first response (801) by inputting the input data into an artificial intelligence system. The electronic device (10) may output the first response (801) as visual and / or auditory information. For example, the first response (801) may be "This phone is the BBB model from company AAA, released in January 2024."

[0139] In the following, examples of changing the length of various answers may be described with reference to FIGS. 8b through 8g. Unless otherwise noted, the examples described in relation to FIG. 8a may also apply to FIGS. 8b through 8g. Examples of changing the length of the answer without explicit instruction from the user (1) are described, but the electronic device (10) of the present disclosure is not unable to change the length of the answer based on explicit instruction. For example, if there is explicit input from the user (1), the electronic device (10) may change the length of the first answer (801).

[0140] Figure 8b illustrates one example of answer reduction.

[0141] Referring to FIGS. 8a and 8b, according to one embodiment, the electronic device (10) can reduce the length of the response. For example, at time t21, the electronic device (10) can acquire a second image (852). For example, the electronic device (10) can acquire the second image (852) while providing a response to the first utterance (91). The electronic device (10) can detect the end of the interaction with the object of interest from the second image (852). For example, the electronic device (10) can identify from the second image (852) that the object of interest has failed to be identified or that the object of interest has been removed from the user's (1) hand. In this case, the electronic device (10) can provide a second response (802) that is relatively shorter in length compared to the first response (801). The electronic device (10) may generate a response adjustment prompt and provide a second response (802) based on the response adjustment prompt. For example, the second response (802) may be provided between time t1 and time t22. For example, at least part of the second response (802) may be provided between time t21 and time t22. In one example, the second response (802) may be a response in which part of the first response (801) is omitted. In one example, the second response (802) may be a response in which part of the first response (801) is replaced by a shortened response.

[0142] For example, the first response (801) may be "This phone is a BBB model from company AAA, released in January 2024." For example, the electronic device (10) may provide a second response (802) by omitting part of the first response (801). In this case, the second response (802) may be "This phone is a BBB model from company AAA." For example, the electronic device (10) may provide a second response (802) by abbreviating part of the first response (801). In this case, the second response (802) may be "This phone is a BBB released in January from company AAA."

[0143] In one example, the first response (801) may be output visually and audibly. The electronic device (10) may reduce the modality of the first response (801). For example, the electronic device (10) may provide the response audibly and visually, and then provide the second response (802) audibly only by stopping the visual output.

[0144] Fig. 8c illustrates an example of providing additional answers based on interest.

[0145] Referring to FIGS. 8a and 8c, according to one embodiment, the electronic device (10) can increase the length of the response. For example, the electronic device (10) can increase the length of the response by providing additional responses. In the example of FIG. 8c, the electronic device (10) can acquire a third image (853) at time t3 when the provision of the first response (801) is completed. In one example, the electronic device (10) can acquire a third image (853) at time when a part of the first response (801) is provided (e.g., any time between time t1 and time t3). For example, the electronic device (10) can identify that the user (1)'s interaction with the electronic device has been maintained by acquiring multiple images between time t1 and time t3. For example, the electronic device (10) can identify from multiple images that an object of interest (e.g., the electronic device) has been grasped by the user (1). For example, grasping of an object of interest by the user (1) may imply interaction with the object of interest. By detecting the maintenance of grasping of the object of interest, the electronic device (10) may detect the maintenance of interaction with the object of interest. The electronic device (10) may generate a response control prompt at time t3 and provide a third response (803) generated based on the response control prompt. For example, at least part of the third response (803) may be provided between time t3 and time t4. In the example of FIG. 8c, the electronic device (10) is described as generating a response control prompt based on time t3, but embodiments of the present disclosure are not limited thereto. For example, the electronic device (10) may generate a response control prompt at any time during the provision of the first response (801) (e.g., between time t1 and time t3) and generate a third response (803) based on the response control prompt.

[0146] For example, the third response (803) may include information about the object of interest that is not provided by the first response (801). In one example, the third response (803) may include information such as the price and / or performance specifications of the object of interest.

[0147] Fig. 8d illustrates an example of providing additional answers based on posture.

[0148] Referring to FIGS. 8a and 8d, according to one embodiment, the electronic device (10) may increase the length of the answer based on the state of the object of interest. For example, at time t3, the electronic device (10) may acquire a fourth image (854). In the example of FIG. 8d, the electronic device (10) is illustrated as acquiring the fourth image (854) at time t3, but embodiments of the present disclosure are not limited thereto. For example, the electronic device (10) may acquire the fourth image (854) at any time between time t1 and time t3.

[0149] Based on the fourth image (854), the electronic device (10) can identify that the orientation of the object of interest has been rotated. As illustrated in the fourth image (854), for example, the object of interest may be rotated laterally relative to the first image (851). In this case, the electronic device (10) may use the changed state information of the object of interest in generating additional answers. For example, the electronic device (10) may include information indicating that the object of interest has been rotated laterally in the answer adjustment prompt. For example, the electronic device (10) may include the fourth image (854) in the answer adjustment prompt. For example, at least part of the fourth response (804) may be provided between time t3 and time t4. For example, the fourth response (804) may include information based on the state of the object of interest. The fourth response (804) may include information about a component observed from the side of the object of interest. For example, the fourth response could be, "This button is the volume button, and the power button is on the bottom."

[0150] In the example of FIG. 8d, it is described that the electronic device (10) generates an additional response based on time t3, but embodiments of the present disclosure are not limited thereto. For example, the electronic device (10) may generate a response adjustment prompt at any time during the provision of the first response (801) (e.g., between time t1 and time t3) and generate a fourth response (804) based on the response adjustment prompt.

[0151] Figure 8e illustrates an example of providing additional answers based on distance.

[0152] Referring to FIGS. 8a and 8e, according to one embodiment, the electronic device (10) may increase the length of the answer based on the distance of the object of interest. For example, at time t3, the electronic device (10) may acquire a fifth image (855). In the example of FIG. 8e, the electronic device (10) is illustrated as acquiring the fifth image (855) at time t3, but embodiments of the present disclosure are not limited thereto. For example, the electronic device (10) may acquire the fifth image (855) at any time between time t1 and time t3.

[0153] Based on the fifth image (855), the electronic device (10) can identify that the distance of the object of interest to the electronic device (10) has decreased. In this case, the electronic device (10) can generate a response adjustment prompt requesting additional information. The electronic device (10) can provide a fifth response (805) based on the response adjustment prompt. For example, the electronic device (10) can provide at least part of the fifth response (805) between time t3 and time t4. The fifth response (805) may include information about the object of interest that was not provided by the first response (801).

[0154] In the example of FIG. 8e, it is described that the electronic device (10) generates an additional response based on time t3, but embodiments of the present disclosure are not limited thereto. For example, the electronic device (10) may generate a response adjustment prompt at any time during the provision of the first response (801) (e.g., between time t1 and time t3) and generate a fifth response (805) based on the response adjustment prompt.

[0155] Figure 8f illustrates an example of providing multiple additional answers.

[0156] Referring to FIGS. 8a and 8f, according to one embodiment, the electronic device (10) may provide a plurality of additional responses. Although the provision of one additional response has been described in relation to FIGS. 8c through 8e, embodiments of the present disclosure are not limited thereto. As described above in relation to FIG. 7, the electronic device (10) may dynamically adjust the length of the response before the provision of the response is completed. For example, while an additional response is being provided, the electronic device (10) may decide to provide another additional response. For example, the electronic device (10) may decide to reduce the additional response.

[0157] In the example of FIG. 8f, the electronic device (10) may acquire a third image (853) at time t3 (or any time between time t1 and t3) and provide a third response (803) based on the third image (853). For example, the electronic device (10) may provide at least a portion of the third response (803) between time t3 and time t4. For example, the third response (803) may be referred to as an additional response to the first response (801). At time t4 (or any time between time t3 and t4), the electronic device (10) may acquire a fourth image (854). In this case, the electronic device (10) may provide a fourth response (804) based on the fourth image (854). For example, the electronic device (10) may provide at least a portion of the fourth response (803) between time t4 and time t5. In this case, the fourth response (804) may be referenced as an additional answer to the third response (803).

[0158] Fig. 8g illustrates one example of a change in the answer level.

[0159] Referring to FIGS. 8a and 8f, according to one embodiment, the electronic device (10) can change the level of the response. The electronic device (10) can change the level by increasing the modality of the response. For example, the electronic device (10) can acquire a fifth image (855) at time t2 during the provision of the first response (801). Based on the fifth image (855), the electronic device (10) can identify that the distance between the object of interest and the electronic device (10) has decreased. Based on the fifth image (855), the electronic device (10) can additionally provide a sixth response (806). For example, the electronic device (10) can output the first response (801) audibly. The electronic device (10) can output the sixth response (806) visually. By outputting the sixth response (806) visually, the electronic device (10) can increase the modality of the response. For example, the electronic device (10) may provide at least a portion of the sixth response (806) between time t2 and time t3. As another example, the electronic device (10) may provide the sixth response (806) after time t3.

[0160] FIG. 9a illustrates a flowchart of a method for changing an answer according to one embodiment.

[0161] Referring to FIG. 2 and FIG. 9a, according to one embodiment, the electronic device (10) may change the length of the answer to a query. An operation described below in relation to FIG. 9a may be referred to as an operation performed in the processor (120) of the electronic device (10) of FIG. 2. The order of the operations described below in relation to FIG. 9a is an example, and embodiments of the present disclosure are not limited thereto. For example, at least some of the operations may be executed differently from the order of FIG. 9a or may be executed substantially simultaneously with other operations of FIG. 9a.

[0162] For example, the electronic device (10) may include at least one camera (e.g., camera (170) of FIG. 2), at least one microphone (e.g., microphone (182) of FIG. 2), memory (130), and at least one processor (e.g., processor (120) of FIG. 2). The at least one processor may be communicably connected to at least one camera, at least one microphone, and memory, and may include at least one processing circuit. The memory (130) may enable the electronic device (10) to perform the operations described below when executed individually or collectively by at least one processor.

[0163] According to one embodiment, in operation 905, the electronic device (10) can receive a query. For example, the electronic device (10) can acquire a user's speech using at least one microphone. The electronic device (10) can acquire the query through speech recognition and natural language understanding of the acquired speech. For example, the electronic device (10) can acquire the query from an external electronic device through a communication circuit (190). For example, the electronic device (10) can acquire the query through any interface (e.g., a touch screen display and / or buttons). The electronic device (10) can receive the query according to operation 605 of FIG. 6.

[0164] According to one embodiment, in operation 910, the electronic device (10) can identify an object of interest associated with a query. The electronic device (10) can acquire at least one image using at least one camera. The electronic device (10) can identify at least one object of interest from at least one image. For example, the object of interest may be an object that is the subject of the query. For example, the object of interest may be an object that may be included in the answer to the query. For example, the electronic device (10) can identify the object of interest according to operation 610 and / or operation 615 of FIG. 6.

[0165] According to one embodiment, in operation 915, the electronic device (10) may change the length of the answer based on the state of the object of interest while outputting a first answer corresponding to the query. For example, the electronic device (10) may decrease or increase the length of the answer. The electronic device (10) may generate an answer adjustment prompt based on a change in the state of the object of interest and change the length of the answer based on the answer adjustment prompt. For example, the state of the object of interest may include information about at least one of the posture, appearance, or size of the object of interest. For example, the electronic device (10) may change the length of the answer when the condition of operation 715 of FIG. 7 is satisfied. The electronic device (10) may adjust the length of the answer according to operation 720 of FIG. 7.

[0166] The electronic device (10) may reduce the length of the answer by terminating the output of the first answer early or by outputting the content of the first answer in a shortened form. In one example, the electronic device (10) may reduce the length of the answer by generating an answer adjustment prompt, generating a shortened first answer based on the answer adjustment prompt, and outputting the shortened first answer. For example, the electronic device (10) may acquire an image using at least one camera during the output of the first answer. The electronic device (10) may reduce the length of the answer if the object of interest is not identified from the image.

[0167] The electronic device (10) can increase the length of the answer based on the state of the object of interest. For example, the electronic device (10) can increase the length of the answer by outputting an additional second answer after the output of the first answer. For example, the electronic device (10) can obtain the second answer by generating an answer adjustment prompt and inputting the answer adjustment prompt into an artificial intelligence model. The electronic device (10) can generate an answer adjustment prompt based on a change in the state of the object of interest. For example, the electronic device (10) can acquire an image using at least one camera at the time when the output of the first answer is completed. If the object of interest is identified from the acquired image, the electronic device (10) can output the second answer. For example, the electronic device (10) can output the second answer if an interaction between the user and the object of interest is detected during the output of the first answer, or if the distance between the object of interest and the electronic device (10) decreases. In one example, the second answer may contain a relatively larger amount of information compared to the first answer.

[0168] FIG. 9b illustrates a flowchart of a method for changing an answer according to one embodiment.

[0169] Referring to FIG. 2 and FIG. 9b, according to one embodiment, the electronic device (10) may change the answer to a query. In the example of FIG. 9a, examples in which the initial answer is changed after the electronic device (10) has provided an initial answer are described, but the embodiments of the present disclosure are not limited thereto. For example, the electronic device (10) may determine the initial answer based on the state of an object (e.g., an object of interest).

[0170] The operations described below in relation to FIG. 9b may be referred to as operations performed in the processor (120) of the electronic device (10) of FIG. 2. The order of the operations described below in relation to FIG. 9b is an example, and embodiments of the present disclosure are not limited thereto. For example, at least some of the operations may be performed differently from the order of FIG. 9b, or may be performed substantially simultaneously with other operations of FIG. 9b.

[0171] According to one embodiment, in operation 955, the electronic device (10) may receive a query while capturing the external environment. For example, the electronic device (10) may receive the query through a microphone or any interface. For example, the electronic device (10) may be a wearable device (e.g., a glasses-type device). The electronic device (10) may capture the external environment of the electronic device (10) using at least one camera while, for example, being worn by a user. For example, the electronic device (10) may capture the external environment prior to receiving the query.

[0172] According to one embodiment, in operation 960, the electronic device (10) can identify the state of an object of interest associated with a query. For example, the electronic device (10) can identify state information of the object of interest based at least in part on a determination that an object associated with a query exists within an external environment. The state information of the object of interest may include information related to the pose, appearance, and / or orientation of the object of interest. The state information of the object of interest may include the time (e.g., appearance time or appearance time interval) during which the object of interest is observed in the field of view of a camera. The state information of the object of interest may include information related to the distance between the object of interest and the electronic device (10). The state information of the object of interest may include information indicating a state changed by user interaction from the initial observed state of the object of interest.

[0173] According to one embodiment, in operation 965, the electronic device (10) can determine whether the state of the object of interest corresponds to a first condition or a second condition. For example, the first condition and the second condition may be different from each other. In one example, the first condition and the second condition may be mutually exclusive. If the first condition is satisfied, the second condition may not be satisfied. Conversely, if the second condition is satisfied, the first condition may not be satisfied. In the present disclosure, state information corresponding to the first condition may mean that the state information satisfies the first condition. In the present disclosure, state information corresponding to the second condition may mean that the state information satisfies the second condition. The first condition and the second condition may be predefined conditions. The first condition and the second condition may be dynamically defined according to the state of the object of interest.

[0174] According to one embodiment, in operation 970, when the state of the object of interest corresponds to the first condition, the electronic device (10) may provide a first response of the first level. According to one embodiment, in operation 975, when the state of the object of interest corresponds to the second condition, the electronic device (10) may provide a second response of the second level. For example, the second level may be different from the first level.

[0175] For example, the first condition may correspond to a first pose of the object of interest that initially appears in the camera's field of view. For example, the second condition may correspond to a second pose that is different from the first pose. During the shooting of the object of interest, if the state information of the object of interest indicates a change in the pose of the object of interest (e.g., a change from a first pose to a second pose), the electronic device (10) may provide a second response. During the shooting of the object of interest, if the state information of the object of interest indicates maintaining the pose of the object of interest (e.g., maintaining the first pose), the electronic device (10) may provide a first response.

[0176] For example, the first condition may correspond to a first appearance time in which the object of interest exists within the field of view of the camera (170). The second condition may correspond to a second appearance time in which the object of interest exists within the field of view of the camera (170). The second appearance time may correspond to a relatively longer time compared to the first appearance time. The appearance time of the object of interest may represent the length of the appearance time of the object of interest observed between the initial observation time of the camera (170) and the time of reception of the query. For example, the electronic device (10) may determine that if the appearance time is less than a specified time, the state information corresponds to the first condition, and if the appearance time is greater than or equal to the specified time, the state information corresponds to the second condition.

[0177] For example, the first condition may correspond to the case where the distance between the electronic device (10) and the object of interest is the first distance. The second condition may correspond to the case where the distance between the electronic device (10) and the object of interest is the second distance. The second distance may be different from the first distance. For example, the second distance may be shorter than the first distance. For example, the electronic device (10) may determine that the state information corresponds to the first condition if the distance of the object of interest is greater than or equal to the specified distance, and that the state information corresponds to the second condition if the distance of the object of interest is less than the specified distance.

[0178] The first response may be a lower-level response compared to the second response. For example, responses of different levels may differ in modality, output path, amount of information, length, examples of responses, and / or tone. For example, the second-level response may contain a greater amount of content compared to the first-level response. For example, the second-level response may contain lower-level items compared to the first-level response.

[0179] For example, higher-level responses may differ in modality from lower-level responses. Level 1 responses may be output audibly. Level 2 responses, which are higher than Level 1, may be output visually. Level 3 responses, which are higher than Level 2, may be output visually and audibly. Level 4 responses, which are higher than Level 3, may be output visually, audibly, and tactilely.

[0180] For example, higher-level responses may have many modalities compared to lower-level responses. Level 1 responses may be output, for example, auditorily. Level 2 responses, which are higher than Level 1, may be output, for example, visually and auditorily. Level 3 responses, which are higher than Level 2, may be output, for example, visually, auditorily, and tactilely.

[0181] For example, the output paths for the first level response and the second level response may be different. For example, the electronic device (10) may output the first level response from the electronic device (10). For example, the electronic device (10) may output the second level response from an external electronic device that is communicationally connected to the electronic device (10). For example, the electronic device (10) may output the second level response using a plurality of electronic devices (e.g., the electronic device (10) and an external electronic device).

[0182] For example, a higher-level response may have a relatively larger amount of information and / or a relatively longer length compared to a lower-level response. A first-level response may have a lower level of detail compared to, for example, a second-level response.

[0183] For example, responses at different levels can have different tones. For instance, the second response at level 2 may have a formal tone. The first response at level 1 may have a relatively non-formal tone compared to the second response.

[0184] With reference to FIG. 9b, a method for determining an initial response has been described, but embodiments of the present disclosure are not limited thereto. According to one embodiment, as described above with reference to FIG. 9a, the electronic device (10) may change the length of the initial response during the provision of the initial response.

[0185] For example, the electronic device (10) may determine that state information has been changed to a state corresponding to a second condition while providing a first response. For example, the state information of an object of interest may be changed by a user to a state corresponding to a second condition. In this case, the electronic device (10) may provide a second response that has been switched from the first response. For example, the electronic device (10) may change the response according to the method described above in relation to FIG. 9a.

[0186] For example, the electronic device (10) may provide additional responses related to a query based on context information after providing a first response or a second response. The electronic device (10) may provide additional responses based on context information identified after providing a first response or a second response. For example, context information may include a history of similar questions, clipboard stored content, buffer stored content, functions used immediately before the question, the location of the electronic device (10), observation time, and / or user profile (e.g., age, gender, and / or interests). The history of similar questions, clipboard stored content, buffer stored content, functions used immediately before the question, the location of the electronic device (10), and / or user profile may be used to generate an answer control prompt for generating additional responses. By including context information in the answer control prompt, the electronic device (10) may provide additional responses that correspond to the user's context. The observation time may correspond, for example, to the time when the object of interest was observed after receiving the query. If the observation time indicates that the object of interest has been observed for more than a specified time, the electronic device (10) may provide additional responses related to the query.

[0187] According to one embodiment, the electronic device (10) may generate an additional response based on situational information and / or environmental information when providing an additional response. For example, environmental information may include ambient sounds (e.g., content, type, and / or volume of ambient sounds), surrounding objects, usage time of the electronic device (10), user reactions (e.g., movement and / or voice), and / or tasks performed by the user after the initial response.

[0188] For example, the electronic device (10) may determine the order of levels at which a response is to be provided based on environmental information. Environmental information may indicate, for example, that the user of the electronic device (10) is in a noisy environment. For example, the electronic device (10) may identify a noisy environment based on ambient sounds and / or user reactions. In this case, the electronic device (10) may output a first response of a first level visually, and a second response of a second level visually and audibly. For example, environmental information may indicate a situation of high mobility. For example, the electronic device (10) may identify mobility based on the user's movement and / or the location of the electronic device (10). In this case, the electronic device (10) may output a first response of a first level audibly, and a second response of a second level visually and audibly.

[0189] For example, the electronic device (10) can generate additional responses based on environmental information. For example, the electronic device (10) can include information about surrounding objects in the response adjustment prompt. By generating additional responses based on information about surrounding objects, the electronic device (10) can provide additional responses related to the query, objects of interest, and surrounding objects. For example, the electronic device (10) can generate additional responses based on tasks performed by the user after the initial response. The electronic device (10) can include information about tasks performed after the initial response in the response adjustment prompt. Based on information about the performed tasks, information about the performed tasks may be reflected in the additional response, or tasks that can be performed subsequently to the performed task or content that can be provided subsequently may be provided through the additional response. For example, the electronic device (10) may output the initial response (e.g., the first response or the second response) audibly, and output the additional response audibly and / or visually.

[0190] FIG. 10 illustrates an example of logging an electronic device according to one embodiment.

[0191] Referring to FIG. 2 and FIG. 10, according to one embodiment, an electronic device (10) can generate a log (e.g., a life log) based on visual information. For example, a user (1) may request the generation of a life log based on an image. For example, a second utterance (92) may be “Analyze the image and make a life log.” The electronic device (10) can generate a life log based on the reception of the second utterance (92). The electronic device (10) can analyze the image and generate a life log in text form based on the analyzed image. By generating a life log in text instead of an image, the life log capacity may be reduced.

[0192] For example, the electronic device (10) may acquire images at specified intervals or based on events (e.g., a change in location). In the example of FIG. 10, the electronic device (10) may acquire a first image (1001) at 7:00 AM and a second image (1002) at 9:00 AM. A third image (1003) may be acquired at 12:30 PM. The electronic device (10) may store a life log by mapping information indicating the acquisition time of each image with text data describing each image. For example, the electronic device (10) may acquire text data by inputting the acquired images into an artificial intelligence system. For example, the life log associated with the first image (1001) may be “AM 07:00: A quiet living room with flowers.” For example, the life log associated with the second image (1002) may be “AM 09:00: A quiet city street.” For example, the life log associated with the third image (1003) may be “PM 1230: delicious Italian pizza”.

[0193] FIG. 11 illustrates an example of object of interest-based logging of an electronic device according to one embodiment.

[0194] Referring to FIG. 2 and FIG. 11, according to one embodiment, an electronic device (10) can generate a log (e.g., a life log) based on visual information. For example, a user (1) may request the generation of a life log based on an image. For example, a user (1) may request the generation of a life log for a designated object of interest. The electronic device (10) may acquire an image using a camera and generate a life log for the object of interest when the object of interest is recognized from the image.

[0195] For example, when an object designated as an object of interest is recognized from an image, the electronic device (10) inputs the image into an artificial intelligence system to generate information about the object of interest and can store the generated information as a life log. In one example, if specific information about the object of interest is specified by the user (1), the electronic device (10) can input a prompt to the artificial intelligence system to analyze specific information from the image. In one example, the electronic device (10) can analyze the difference between the object of interest at the time of creating the previous life log and the current object of interest by inputting the previous life log together with the current image.

[0196] For example, the third utterance (93) may be “make a life log for this tree.” The electronic device (10) may create a life log based on the reception of the third utterance (93). The electronic device (10) may acquire a first image (1101) of an object of interest (e.g., a tree). The electronic device (10) may identify the state of the object of interest for the first image (1101). In the example of FIG. 11, the object of interest is described as being designated by the third utterance (93) of the user (1), but embodiments of the present disclosure are not limited thereto. For example, the electronic device (10) may identify an object of interest based on a query history and / or gaze. The electronic device (10) may identify a specific object as an object of interest if the user (1) has a history of querying while watching a specific object. For example, if an image and / or video containing an object of interest is stored in memory (130), the electronic device (10) may use the image and / or video to generate a life log of the identified object of interest. In generating the life log, the electronic device (10) may adjust the level of detail (e.g., length) of the life log based on the user's (1) interactions (e.g., query history and / or time spent looking). For example, the electronic device (10) may be configured to generate a more detailed life log as the interaction with the object of interest becomes longer.

[0197] For example, the state of the object of interest identified from the first image (1101) may be “Color: Brown, Length: 30 cm”. The state of the object of interest may be input into an artificial intelligence system along with the first image (1101). The electronic device (10) may use the artificial intelligence system to generate a first life log for the first image (1101). The first life log may be, for example, “The tree in front of the house in April is a 30 cm sapling with no leaves”.

[0198] For example, the state of the object of interest identified from the second image (1102) may be “Color: brown and green, Length: 100 cm”. The state of the object of interest may be input into an artificial intelligence system along with the second image (1102). The electronic device (10) may use the artificial intelligence system to generate a second life log for the second image (1102). The second life log may be, for example, “The tree in front of the house in July has green leaves and has become 100 cm long”.

[0199] For example, the state of the object of interest identified from the third image (1102) may be “Color: brown and yellow, Length: 150 cm”. The state of the object of interest may be input into an artificial intelligence system along with the third image (1103). The electronic device (10) may generate a third life log for the third image (1103) using the artificial intelligence system. The third life log may be, for example, “In November, the green leaves of the tree in front of the house turned yellow, and its length became 150 cm”. For example, the electronic device (10) may input a second life log along with the third image (1103) into the artificial intelligence system. The artificial intelligence system may generate a third life log containing information about changes in the object of interest by comparing the contents of the second life log with the current state of the object of interest.

[0200] FIG. 12 is a block diagram of an exemplary electronic device (1200) capable of performing the operations described in this document.

[0201] Referring to FIG. 12, the electronic device (1200) may be one of various forms of electronic devices, such as a notebook (1290), smartphones (1291) having various form factors (e.g., a bar-type smartphone (1291-1), a foldable-type smartphone (1291-2), or a sliderable (or rollable)-type smartphone (1291-3)), a tablet (1292), a cellular phone (not shown), and other similar computing devices (not shown). The components, their relationships, and their functions illustrated in FIG. 12 are illustrative only and are not intended to limit the implementations described or claimed herein. The electronic device (1200) may be referred to as a mobile device, a user device, a multifunction device, a portable device, or a server.

[0202] The electronic device (1200) may include components comprising at least one processor (1210) (hereinafter referred to as processor (1210)), at least one memory (1220) (hereinafter referred to as memory (1220)), at least one display (1240) (hereinafter referred to as display (1240)), at least one image sensor (1250) (hereinafter referred to as image sensor (1250)), at least one communication circuit (1260) (hereinafter referred to as communication circuit (1260)), and / or at least one sensor (1270) (hereinafter referred to as sensor (1270)). The components are merely exemplary. For example, the electronic device (1200) may include other components (e.g., power management integrated circuitry (PMIC), audio processing circuit, antenna, rechargeable battery, or input / output interface). For example, some components may be omitted from the electronic device (1200). For example, some components can be integrated into a single component.

[0203] The processor (1210) may be implemented as one or more integrated circuit (or circuitry) chips and may perform various data processing operations. The processor (1210) may include at least one electrical circuit and may process instructions (or programs, data) stored in memory (1220) individually or collectively in a distributed manner. The processor (1210) may include a processor assembly comprising one or more processing circuits. The processor (1210) may include any processing circuit that is operative to control the performance and operations of one or more components of the electronic device (1200) (e.g., memory (1220), display (1240), image sensor (1250), communication circuit (1260), and / or sensor (1270)). For example, the processor (1210) (e.g., application processor (AP)) may be implemented as a system on chip (SoC) (e.g., a single chip or chipset). For example, the processor (1210) may be implemented with a plurality of cores (or at least one core circuit), a plurality of chips, or a plurality of chipsets. For example, the processor (1210) may include one or more processing circuits. For example, the processor (1210) may include one or more processing circuits configured to perform the various functions of the present disclosure individually and / or collectively. As an example without limitation, at least a portion of the processor (1210) may be included in a first chip of the electronic device (1200), and at least another portion of the processor (1210) may be included in a second chip of the electronic device (1200) different from the first chip of the electronic device (1200).

[0204] For example, the processor (1210) may include a central processing unit (1211), a graphics processing unit (1212), a neural processing unit (1213), an image signal processor (1214), a display controller (1215), a memory controller (1216), a storage controller (1217), a communication processor (1218), and / or a sensor interface (1219). These components of the processor (1210) are merely exemplary. For example, the processor (1210) may include other components. For example, some components of the processor (1210) may be omitted from the processor (1210). For example, some components of the processor (1210) may be included as separate components of the electronic device (1200) outside of the processor (1210). For example, some components of the processor (1210) (e.g., memory controller (1216)) may be included in other components (e.g., at least part of memory (1220), an interface (e.g. available for connection to at least one component of the electronic device (100)), a display (1240) and / or an image sensor (1250)).

[0205] The processor (1210) may cause other components of the electronic device (1200) to perform various operations by executing instructions stored in memory (1220). The CPU (1211) (or central processing circuit) may be configured to control the components of the processor (1210) based on the execution of instructions stored in memory (1220) (e.g., volatile memory (1221) and / or non-volatile memory (1222)). The GPU (1212) (or graphics processing circuit) may be configured to execute parallel operations (e.g., rendering). The NPU (1213) (or neural processing circuit, or AI (artificial intelligence) chip) may be configured to execute operations for an artificial intelligence model (e.g., convolution computation). An ISP (1214) (or image signal processing circuit) may be configured to process a raw image acquired through an image sensor (1250) into a format suitable for a component within an electronic device (1200) or a component of a processor (1210). A display controller (1215) (or display control circuit, or DPU (display processing unit)) may be configured to process an image acquired from a CPU (1211), GPU (1212), ISP (1214), or memory (1220) (e.g., volatile memory (1221)) into a format suitable for a display (1240). A memory controller (1216) (or memory control circuit) may be configured to control reading data from volatile memory (1221) and writing data to volatile memory (1221). A storage controller (1217) (or storage control circuit) may be configured to control reading data from non-volatile memory (1222) and writing data to non-volatile memory (1222).The CP (1218) (communication processing circuit) may be configured to process data obtained from a component of the processor (1210) into a format suitable for transmitting to another electronic device via the communication circuit (1260), or to process data obtained from another electronic device via the communication circuit (1260) into a format suitable for processing by the component of the processor (1210). For example, the communication circuit (1260) may include one or more communication circuits. The sensor interface (1219) (or sensing data processing circuit, sensor hub) may be configured to process data regarding the state of the electronic device (1200) and / or the state around the electronic device (1200), obtained through the sensor (1270), into a format suitable for the component of the processor (1210).

[0206] Memory (1220) may include one or more storage media (or one or more storage devices). For example, memory (1220) may include a memory assembly comprising one or more storage media. For example, the one or more storage media may include a hard drive, a flash memory, a permanent memory such as ROM (read-only memory) (e.g., non-volatile memory (1222)), a semi-permanent memory such as RAM (random access memory) (e.g., volatile memory (1221)), any other suitable type of storage (or storage assembly), or any combination thereof. Memory (1220) may include a cache memory, which is one or more different types of memory used to temporarily store data for a function or feature of the electronic device (1200). As an example not limited to, the cache memory may be included within the processor (1210). The memory (1220) may be fixedly embedded within the electronic device (1200) or incorporated into one or more suitable types of components (e.g., a SIM (subscriber identity module) card and / or an SD (secure digital) card) that can be repeatedly inserted into and removed from the electronic device (1200).

[0207] For example, memory (1220) may store one or more software applications, such as operating system (or system) software applications, firmware software applications, driver software applications, plugin (e.g., add-in, add-on, and / or applet) software applications, and / or any other suitable software applications. For example, the one or more software applications may include instructions executable by the processor (1210). For example, memory (1220) may store instructions that can be called by an application programming interface (API). For example, memory (1220) may store instructions within a library.

[0208] Referring again to FIGS. 1 through 12, an electronic device (10) according to one embodiment may include at least one camera (170), at least one microphone (182), a memory (130), and at least one processor (120). The at least one processor (120) may be communicatively connected to the at least one camera, the at least one microphone, and the memory, and may include at least one processing circuit. The memory may store instructions that cause the electronic device to perform the operations described below when executed individually or collectively by the at least one processor.

[0209] According to one embodiment, the electronic device receives a query from a user using the at least one microphone, identifies an object of interest associated with the query from an image acquired using the at least one camera, and while outputting a first answer corresponding to the query, may change the length of the answer based on the state of the object of interest. For example, the object of interest may be the subject of the query.

[0210] According to one embodiment, the electronic device may reduce the length of the answer by terminating the output of the first answer early or by summarizing and outputting the content of the first answer based on the state of the object of interest. For example, the electronic device may reduce the length of the answer if, during the output of the first answer, the object of interest is not identified from an image acquired using at least one camera.

[0211] According to one embodiment, the electronic device may increase the length of the answer by outputting an additional second answer after the output of the first answer, based on the state of the object of interest. For example, the electronic device may output the second answer if the object of interest is identified from an image acquired using at least one camera after the completion of the output of the first answer. For example, the electronic device may output the second answer if, during the output of the first answer, an interaction between the user and the object of interest is detected, or if the distance between the object of interest and the electronic device decreases. For example, the electronic device may obtain the second answer by generating an answer adjustment prompt and inputting the answer adjustment prompt into an artificial intelligence model. For example, the state of the object of interest may include information regarding at least one of the posture, appearance, or size of the object of interest. For example, the electronic device may generate the answer adjustment prompt based on a change in the state of the object of interest. For example, the second answer may contain a relatively larger amount of information compared to the first answer.

[0212] An artificial intelligence-based answer providing method for an electronic device according to one embodiment may include: receiving a query from a user using at least one microphone; identifying an object of interest associated with the query from an image acquired using at least one camera; and changing the length of the answer based on the state of the object of interest while outputting a first answer corresponding to the query. For example, the object of interest may be the subject of the query.

[0213] According to one embodiment, the operation of changing the length of the answer may include an operation of reducing the length of the answer by terminating the output of the first answer early or by summarizing and outputting the content of the first answer based on the state of the object of interest. For example, the operation of reducing the length of the answer may include an operation of reducing the length of the answer when, during the output of the first answer, the object of interest is not identified from an image acquired using at least one camera.

[0214] According to one embodiment, the operation of changing the length of the answer may include an operation of increasing the length of the answer by outputting an additional second answer after the output of the first answer, based on the state of the object of interest. For example, the operation of increasing the length of the answer may include an operation of outputting the second answer when the object of interest is identified from an image acquired using at least one camera after the completion of the output of the first answer. For example, the operation of increasing the length of the answer may include an operation of outputting the second answer when an interaction between the user and the object of interest is detected during the output of the first answer, or when the distance between the object of interest and the electronic device is reduced. For example, the operation of increasing the length of the answer may include an operation of generating an answer adjustment prompt, and an operation of obtaining the second answer by inputting the answer adjustment prompt into an artificial intelligence model. For example, the state of the object of interest may include information regarding at least one of the posture, appearance, or size of the object of interest. For example, the operation of generating an answer control prompt may include the operation of generating the answer control prompt based on a change in the state of the object of interest. For example, the second answer may contain a relatively larger amount of information compared to the first answer.

[0215] An electronic device (10) according to one embodiment may include at least one camera (170), at least one microphone (182), a memory (130), and at least one processor (120). The at least one processor (120) may include at least one processing circuit and is operatively connected to the at least one camera, the at least one microphone, and the memory. The memory may store instructions that cause the electronic device to perform the operations described below when executed individually or collectively by the at least one processor.

[0216] According to one embodiment, the electronic device receives a query from a user while capturing an external environment of the electronic device using at least one camera, and identifies state information associated with the object based at least partially on a determination that an object associated with the query exists in the external environment, and if the state information corresponds to a first condition, provides a first response of a first level to the query, and if the state information corresponds to a second condition, provides a second response of a second level different from the first level to the query. For example, the first response may be provided audibly, and the second response may be provided audibly and visually.

[0217] For example, the state information may include the pose of the object. The first condition may correspond to a first pose of the object that initially appears in the field of view of at least one camera. For example, the second condition may correspond to a second pose that is different from the first pose.

[0218] For example, the state information may include an appearance time interval of the object. The first condition may correspond to a first appearance time interval during which the object exists within the field of view of the at least one camera. The second condition may correspond to a second appearance time interval during which the object exists within the field of view of the at least one camera. The second appearance time interval may be longer than the first appearance time interval.

[0219] For example, the state information may include the location of the object. The first condition may correspond to a first distance between the electronic device and the object. The second condition may correspond to a second distance between the electronic device and the object. The second distance may be different from the first distance.

[0220] For example, the state information may be changed by the user from a state corresponding to the first condition to a state corresponding to the second condition. According to one embodiment, the electronic device may provide the second response switched from the first response based on a determination that the state information was changed from a state corresponding to the first condition to a state corresponding to the second condition while providing the first response.

[0221] According to one embodiment, the electronic device may provide an additional response related to the query based on context information identified by the electronic device after providing the first response or the second response. For example, the context information may include the object being observed in the external environment for a specified period of time or longer.

[0222] According to one embodiment, the electronic device may further include a communication circuit. For example, the first response is output from the electronic device, and at least a portion of the second response may be output from an external electronic device wirelessly connected through the communication circuit.

[0223] A method for providing an AI-based response for an electronic device according to one embodiment may include: receiving a query from a user while capturing an external environment of the electronic device using at least one camera; identifying state information associated with an object based at least partially on a determination that an object associated with the query exists in the external environment; providing a first response of a first level to the query when the state information corresponds to a first condition; and providing a second response of a second level different from the first level to the query when the state information corresponds to a second condition. For example, the first response may be provided audibly, and the second response may be provided audibly and visually. For example, the first response may be output from the electronic device, and at least a portion of the second response may be output from an external electronic device wirelessly connected through a communication circuit of the electronic device.

[0224] For example, the state information may include the pose of the object. The first condition may correspond to a first pose of the object that initially appears in the field of view of at least one camera. For example, the second condition may correspond to a second pose that is different from the first pose.

[0225] For example, the state information may include an appearance time interval of the object. The first condition may correspond to a first appearance time interval during which the object exists within the field of view of the at least one camera. The second condition may correspond to a second appearance time interval during which the object exists within the field of view of the at least one camera. The second appearance time interval may be longer than the first appearance time interval.

[0226] For example, the state information may include the location of the object. The first condition may correspond to a first distance between the electronic device and the object. The second condition may correspond to a second distance between the electronic device and the object. The second distance may be different from the first distance.

[0227] For example, the state information may be changed by the user from a state corresponding to the first condition to a state corresponding to the second condition. According to one embodiment, the operation of providing the second response may include the operation of providing the second response switched from the first response based on a determination that the state information was changed from a state corresponding to the first condition to a state corresponding to the second condition while providing the first response.

[0228] According to one embodiment, the method may further include an operation of providing an additional response related to the query based on context information identified by the electronic device after providing the first response or the second response. For example, the context information may include the object being observed in the external environment for a specified period of time or longer.

Claims

1. In an electronic device (10), At least one camera (170); At least one microphone (182); Memory (130); and The apparatus includes at least one camera, at least one microphone, and at least one processor (120) that is communicatively connected to the memory and includes at least one processing circuit, and When the above memory is executed individually or collectively by the above at least one processor, the electronic device: Receive a query from a user using at least one microphone as described above, and Identifying an object of interest associated with the query from an image acquired using at least one camera, and An electronic device that stores instructions for changing the length of an answer based on the state of the object of interest while outputting a first answer corresponding to the above query.

2. In Paragraph 1, The object of interest above is an electronic device that is the subject of the above query.

3. In Paragraph 2, When the above instructions are executed individually or in combination by the at least one processor, the electronic device: An electronic device that reduces the length of the answer by terminating the output of the first answer early or outputting the content of the first answer in a shortened form, based on the state of the object of interest.

4. In Paragraph 3, When the above instructions are executed individually or in combination by the at least one processor, the electronic device: An electronic device that reduces the length of the answer when, during the output of the first answer, the object of interest is not identified from an image obtained using at least one camera.

5. In Paragraph 1, When the above instructions are executed individually or in combination by the at least one processor, the electronic device: An electronic device that increases the length of the answer by outputting an additional second answer after outputting the first answer, based on the state of the object of interest.

6. In Paragraph 5, When the above instructions are executed individually or in combination by the at least one processor, the electronic device: An electronic device that outputs the second answer when the object of interest is identified from an image obtained using at least one camera after the completion of outputting the first answer.

7. In Paragraph 5, When the above instructions are executed individually or in combination by the at least one processor, the electronic device, An electronic device that outputs the second answer when, during the output of the first answer, an interaction between the user and the object of interest is detected, or when the distance between the object of interest and the electronic device is reduced.

8. In Paragraph 5, When the above instructions are executed individually or in combination by the at least one processor, the electronic device, Create an answer control prompt, An electronic device that obtains the second answer by inputting the above answer adjustment prompt into an artificial intelligence model.

9. In Paragraph 8, The state of the object of interest includes information about at least one of the posture, shape, or size of the object of interest, and When the above instructions are executed individually or in combination by the at least one processor, the electronic device: An electronic device that generates the answer control prompt based on a change in the state of the object of interest.

10. In Paragraph 8, The above second answer is an electronic device containing a relatively large amount of information compared to the above first answer.

11. In an electronic device (10), At least one camera (170); At least one microphone (182); Memory (130); and It includes at least one camera, at least one microphone, and at least one processor (120) operatively connected to the memory and comprising at least one processing circuit, and When the above memory is executed individually or collectively by the above at least one processor, the electronic device: While capturing the external environment of the electronic device using at least one camera, a query by a user is received, and Based at least in part on the determination that an object associated with the above query exists in the above external environment, state information associated with said object is identified, and If the above state information corresponds to the first condition, a first response of the first level is provided for the above query, and An electronic device that stores instructions to provide a second response of a second level different from the first level to the query when the above state information corresponds to a second condition.

12. In Paragraph 11, The above state information includes the posture of the object, and The above first condition corresponds to a first pose of the object that initially appears in the field of view of at least one camera, and The above second condition is an electronic device corresponding to a second pose different from the above first pose.

13. In Paragraph 11, The above state information includes the appearance time interval of the object, and The first condition above corresponds to a first appearance time interval in which the object exists within the field of view of at least one camera, and The above second condition corresponds to a second appearance time interval in which the object exists within the field of view of at least one camera, and An electronic device in which the second appearance time interval is longer than the first appearance time interval.

14. In Paragraph 11, The above state information includes the location of the object, and The above first condition corresponds to a first distance between the electronic device and the object, and The second condition above corresponds to a second distance between the electronic device and the object, and The above second distance is different from the above first distance, an electronic device.

15. In Paragraph 11, The above state information is an electronic device that changes, by the user, from a state corresponding to the first condition to a state corresponding to the second condition.