Method for providing response based on modality determination and electronic device therefor

By determining input modality and selectively activating sensors, the electronic device optimizes processing of single-modal or multi-modal inputs, improving response correlation and reducing power consumption, addressing inefficiencies in existing systems.

WO2026005387A1PCT designated stage Publication Date: 2026-01-02SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/008529
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-23
Filing Date
2025-06-19
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing electronic devices face challenges in efficiently processing multi-modal inputs, leading to reduced correlation between responses and user intentions, increased power consumption, and degraded user experience due to unnecessary activation of cameras and inconsistent modality determination.

Method used

The electronic device determines the input modality based on context information, activating only necessary sensors and cameras to process either single-modal or multi-modal inputs, thereby enhancing the correlation between responses and user inputs while reducing power consumption.

Benefits of technology

This approach improves the correlation between device responses and user intentions, reduces power consumption, and enhances user experience by optimizing sensor usage based on determined input modality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025008529_02012026_PF_FP_ABST
    Figure KR2025008529_02012026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is an electronic device comprising at least one sensor, a camera, a microphone, a memory, and a processor. The electronic device may acquire a voice input by using a microphone and determine whether a field of view of a camera corresponds to a field of view of a user. When a field of view of the camera corresponds to a field of view of the user, the electronic device may generate at least one first prompt based on an image and a voice input, and provide a response based on first result data generated by inputting the at least one first prompt into an artificial intelligence model.
Need to check novelty before this filing date? Find Prior Art

Description

Method for providing response based on modality determination and electronic device therefor

[0001] Embodiments disclosed in this document relate to a method for providing a response based on determination of modality and an electronic device therefor.

[0002] A variety of services utilizing generative AI are being developed. Early services supported generating text-based responses or generating related images based on images. With the advancement of generative AI, services supporting various input formats are being introduced. For example, a service is being developed that performs image editing using a specified image and text data specifying the desired modification. By utilizing generative AI services that support various input formats, users can generate results that align with their intended intent.

[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art in connection with the present disclosure.

[0004] An electronic device according to an embodiment disclosed in the present document may include at least one sensor, a camera, a microphone, a memory, and at least one processor communicatively connected to the at least one sensor, the camera, the microphone, and the memory. The memory may store instructions that, when individually or in combination, are executed by the at least one processor, cause the electronic device to obtain a voice input from a user using the microphone and determine whether a field of view of the camera and a field of view of the user correspond to each other. The instructions, when individually or in combination, are executed by the at least one processor, cause the electronic device to generate at least one first prompt based on an image obtained by the camera and the voice input, input the at least one first prompt into an artificial intelligence model to obtain first result data, and provide a response associated with the voice input based on the first result data.

[0005] In addition, a method for providing a response based on modality determination according to an embodiment disclosed in the present document may include an operation of acquiring a voice input and an operation of determining whether a field of view of a camera and a field of view of a user correspond to each other. The method for providing a response based on modality determination may include an operation of generating at least one first prompt based on an image acquired by the camera and the voice input when the field of view of the camera and the field of view of the user correspond to each other, an operation of acquiring first result data by inputting the at least one first prompt into an artificial intelligence model, and an operation of providing a response associated with the voice input based on the first result data.

[0006] A computer-readable storage medium according to an embodiment disclosed in this document can store instructions that, when executed by a processor of an electronic device, cause the electronic device to perform a method for providing a response based on the modality determination.

[0007] Figure 1a illustrates examples of electronic devices.

[0008] FIG. 1b illustrates a block diagram of an electronic device according to one embodiment.

[0009] Figure 2 illustrates a block diagram of an artificial intelligence system according to one embodiment.

[0010] Figure 3 illustrates the structure of an artificial intelligence model according to one embodiment.

[0011] Figure 4 is a flowchart of a task performing method according to one embodiment.

[0012] Figure 5 is a flowchart of a method for determining an operation mode according to one embodiment.

[0013] Figure 6a is a flowchart of a method for determining an operation mode according to an embodiment.

[0014] Figure 6b is a flowchart of a method for determining an operation mode according to one embodiment.

[0015] FIG. 7A illustrates a first mode operating environment of an electronic device according to one embodiment.

[0016] FIG. 7b illustrates a second mode operating environment of an electronic device according to one embodiment.

[0017] FIG. 7c illustrates a second mode operating environment of an electronic device according to one embodiment.

[0018] FIG. 7d illustrates a third mode operating environment of an electronic device according to one embodiment.

[0019] FIG. 8A illustrates an electronic device according to one embodiment.

[0020] FIG. 8b illustrates an electronic device according to one embodiment.

[0021] FIG. 9 is a flowchart of a method for providing a response based on modality determination according to one embodiment.

[0022] FIG. 10 is a flowchart of a method for providing a response based on modality determination according to one embodiment.

[0023] FIG. 11 is a block diagram of an exemplary electronic device capable of performing the operations described in this document.

[0024] In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components.

[0025] Hereinafter, various embodiments of the present invention will be described with reference to the attached drawings. However, this is not intended to limit the present invention to specific embodiments, and it should be understood that the present invention encompasses various modifications, equivalents, and / or alternatives of the embodiments.

[0026] Figure 1a illustrates examples of electronic devices.

[0027] Referring to FIG. 1A, according to one embodiment, an electronic device may support multiple modalities. For example, the electronic device may be configured to receive and process various types of input. For example, a "modality" may refer to a channel for interaction between the electronic device and a user. In this case, voice input and text input via an interface (e.g., a virtual keyboard) may be considered different modalities. For example, a "modality" may refer to a format of data input to an artificial intelligence model (e.g., a generative artificial intelligence model). For example, a user's voice input may be converted into text data via STT (speech-to-text) and NLU (natural language understanding) and input to the artificial intelligence model. In this case, from the perspective of the artificial intelligence model, voice input and text input via an interface (e.g., a virtual keyboard) may be considered the same modality. Hereinafter, a "modality" may refer to a channel between a user and the electronic device and / or a format of input data to the artificial intelligence model. In this disclosure, 'modality' may be referred to as 'input type'.

[0028] In one example, the first electronic device (10a) may be smart glasses. The first electronic device (10a) may obtain a user's voice input using a microphone. The first electronic device (10a) may obtain an image input using a camera. The first electronic device (10a) may obtain a user input using an interface (e.g., a button and / or a touchpad). The first electronic device (10a) may be configured to obtain a voice input, an image input, and / or a user input, and to process the obtained input.

[0029] In one example, the second electronic device (10b) may be a wearable device configured to be attached to the user's body or clothing. For example, the second electronic device (10b) may include an attachment structure that can be attached to the user's body or clothing. The second electronic device (10b) may obtain input via, for example, a microphone, a camera, and / or an interface.

[0030] In one example, the third electronic device (10c) may be a mobile device. The third electronic device (10c) may be a handheld device. The third electronic device (10c) may, for example, obtain input through a microphone, a camera, and / or an interface.

[0031] The first electronic device (10a), the second electronic device (10b), and the third electronic device (10c) are examples of the electronic devices (10) described with reference to FIG. 2. The electronic devices (10) of the present disclosure are not limited to the first, second, and third electronic devices (10a, 10b, 10c) of FIG. 1. For example, the electronic devices (10) may be any user devices that support multiple modalities. The electronic devices (10) may include a video see-through (VST) device, a watch-type electronic device, a head mounted device (HMD), and / or a vehicle infotainment system.

[0032] FIG. 1b illustrates a block diagram of an electronic device according to one embodiment.

[0033] Referring to FIG. 1B, according to one embodiment, the electronic device (10) may include a processor (120), a memory (130), a sensor circuit (140), a display (160), a camera (170), an interface (180), and / or a communication circuit (190). The electronic device (10) may correspond to the first electronic device (10a), the second electronic device (10b), the third electronic device (10c) of FIG. 1, and / or the electronic device (1100) of FIG. 11. The electronic device (10) may include a configuration similar to the electronic device (1100) described below with reference to FIG. 11. For example, the processor (120) may correspond to at least one processor (1110) of FIG. 11. For example, the memory (130) may correspond to the memory (1120) of FIG. 11. For example, the sensor circuit (140) may correspond to the sensor interface (1119) and / or the sensor (1170) of FIG. 11. For example, the display (160) may correspond to the display (1140) of FIG. 11. For example, the camera (170) may include the image sensor (1150) of FIG. 11. For example, the communication circuit (190) may correspond to the communication circuit (1160) of FIG. 11. The configuration of the electronic device (10) illustrated in FIG. 1B is exemplary, and the configuration of the electronic device (10) is not limited thereto. For example, the electronic device (10) may further include a configuration not illustrated in FIG. 1B (e.g., at least one of the configurations of the electronic device (1100) of FIG. 11). For example, the electronic device (10) may not include at least one of the components illustrated in FIG. 1B (e.g., the second microphone (183), the speaker (184), and / or the sensor circuit (140)).

[0034] The processor (120) may be communicatively, electrically, operatively, or functionally connected to the memory (130), the sensor circuit (140), the display (160), the camera (170), the interface (180), and / or the communication circuit (190). In various embodiments of the present disclosure, when a component is “operatively” connected to another component, it may mean that the component is connected so as to be able to operate the other component. For example, the component may operate the other component by transmitting a control signal to the other component, either directly or via another component. In various embodiments of the present disclosure, when a component is “functionally” connected to another component, it may mean that the component is connected so as to be able to execute a function of the other component. For example, the component may execute a function of the other component by transmitting a control signal to the other component, either directly or via another component.

[0035] The processor (120) may include at least one processor. For example, the processor (120) may include an application processor (AP), a central processing unit (CPU), an image signal processor (ISP), a graphics processing unit (GPU), a neural processing unit (NPU), a tensor processing unit (TPU), and / or a communication processor (CP). The processor (120) may include at least one chip or one chipset. In the present disclosure, the processor (120) may be referred to as a hardware component having an architecture by at least one processing circuit. For example, the processor (120) may be disposed on a substrate (e.g., a printed circuit board) located within the electronic device (10) and may communicate with other components of the electronic device (10) through at least one conductive path formed on the substrate.

[0036] The memory (130) can store instructions. When executed by the processor (120), the instructions can cause the electronic device (10) to perform various operations. For example, the instructions can be individually or collectively executed by at least one processor to cause the electronic device (10) to perform various operations. In various embodiments of the present disclosure, the operation of the electronic device (10) can be referred to as an operation performed by the processor (120) by executing instructions stored in the memory (130). The memory (130) can be referred to as a hardware component for data storage.

[0037] The sensor circuit (140) may be configured to detect information (e.g., context information) associated with the electronic device (10). The context information may include optical information surrounding the electronic device (10), movement information of the electronic device (10), attachment status information of the electronic device (10), and / or information of adjacent objects. For example, the sensor circuit (140) may include an image sensor (141), a motion sensor (143), an attachment sensor (145), and / or a proximity sensor (147). The configuration of the sensor circuit (140) illustrated in FIG. 1B is exemplary, and embodiments of the present disclosure are not limited thereto.

[0038] The image sensor (141) may be configured to detect optical signals. For example, the image sensor (141) may be configured to detect light signals, infrared signals, and / or ultraviolet signals. The image sensor (141) may be configured to acquire depth information. For example, the image sensor (141) may include a light detection and ranging (LiDAR).

[0039] The motion sensor (143) may be configured to detect movement of the electronic device (10). For example, the motion sensor (143) may include an inertial measurement unit (IMU) configured to detect orientation, magnetism, angular rate, and / or specific force of the electronic device (10). The motion sensor (143) may include an accelerometer, a gyroscope, and / or a magnetometer.

[0040] The attachment sensor (145) may be configured to detect the attachment status of the electronic device (10). For example, the attachment sensor (145) may detect the attachment status of the electronic device (10) by detecting an attachment of the electronic device. The attachment sensor (145) may include a Hall sensor configured to detect the magnetic force of the attachment. The operation of the attachment sensor (145) may be described below with reference to FIGS. 8A and 8B .

[0041] The proximity sensor (147) may be configured to detect an external object in proximity to the electronic device (10). The electronic device (10) may detect a blockage state of the electronic device (10) using the proximity sensor (147). For example, the proximity sensor (147) may include a capacitive sensor, a photoelectric sensor, a touch switch, an illuminance sensor, and / or an inductive sensor.

[0042] The display (160) may include at least one pixel configured to display an image. In one example, the display (160) may include multiple displays. For example, the display (160) may include a left-eye display and a right-eye display. The display (160) may include a front display and / or a rear display. The display (160) may include at least one of a see-through display, a flexible display, a rollable display, a foldable display, and / or a rigid display. In one example, the display (160) may include at least one projector for projecting an image.

[0043] The camera (170) may include at least one camera. When the electronic device (10) includes multiple cameras, each of the multiple cameras may have at least one different facing direction, magnification, or field of view. In one example, the electronic device (10) may select a camera for acquiring an image from the multiple cameras. The electronic device (10) may select a camera for acquiring an image based on user context information (e.g., utterance and / or movement direction). For example, the electronic device (10) may identify an image including an object corresponding to a user utterance from among the multiple cameras through image recognition. The electronic device (10) may use the camera that acquired the identified image to acquire an image corresponding to a voice input. For example, the electronic device (10) may select a camera facing the movement direction of the electronic device (10) from among the multiple cameras. In one example, the electronic device (10) may select one camera from among the multiple cameras based on a user input.

[0044] The interface (180) may include at least one device configured to receive an input. For example, the interface (180) may include a touch circuit (181) configured to receive a touch input (e.g., a touch screen display). The interface (180) may include at least one microphone (e.g., a first microphone (182) and / or a second microphone (183)) configured to receive a voice input. In one example, the electronic device (10) may perform beamforming using the first microphone (182) and the second microphone (183) to identify the location of a speaker of the voice input (e.g., a relative direction of the speaker with respect to the electronic device (10)). The interface (180) may include at least one device for output. For example, the interface (180) may include a haptic module for tactile output, at least one speaker (e.g., speaker (184)) for sound output, and / or an indicator. The interface (180) may include any human interface device (HID). For example, the interface (180) may include a button (185). In one embodiment, the processor (120) may be configured to receive input using the interface (180) and process the received input.

[0045] The communication circuit (190) may be configured to perform short-range wireless communication and / or long-range wireless communication. The communication circuit (190) may include a network interface card (NIC). The processor (120) may communicate with other external electronic devices based on wireless communication and / or wired communication, for example, using the communication circuit (190). The processor (120) may communicate with an external device via an IP (Internet Protocol) network, for example, using the communication circuit (190).

[0046] According to one embodiment, the electronic device (10) may obtain an input (e.g., a request to perform a task) and process the input to generate a result. For example, the electronic device (10) may generate a result by processing the input using a generative artificial intelligence model. The electronic device (10) may generate a result using a single-modal input and / or a multi-modal input. The electronic device (10) may generate a result using, for example, an artificial intelligence system (200) described below with reference to FIG. 2. The electronic device (10) may provide a response to the input using the generated result. The electronic device (10) may provide the response using the interface (180) and / or the display (160). The electronic device (10) may transmit information corresponding to the response to an external electronic device using a communication circuit (190), thereby causing the external electronic device (not shown) to provide a response.

[0047] As an example unrelated to the embodiments of the present disclosure, when a user input (e.g., voice input) is received, an electronic device (10) supporting multi-modal input may generate a result by inputting the user input and an image acquired using the camera (170) into an artificial intelligence model. If the user input is unrelated to the image acquired using the camera (170), the correlation between the response and the user input may be reduced. For example, a response unrelated to the user's intention may be provided. Furthermore, if the user specifies the modality through a separate input, the user experience may be degraded. Furthermore, to process multi-modal input, the electronic device (10) may constantly activate the camera (170). In this case, the power consumption of the electronic device (10) may increase.

[0048] According to one embodiment of the present disclosure, the electronic device (10) may determine an input modality based on the reception of a user input. For example, the electronic device (10) may determine the input modality using context information of the electronic device (10) (e.g., information related to the operating environment). The electronic device (10) may obtain the context information using the sensor circuit (140), the camera (170), and / or the interface (180). If processing of a user input by a single modal input is determined, the electronic device (10) may generate a response by processing the user input using an artificial intelligence model. If processing of a user input by a multimodal input is determined, the electronic device (10) may generate a response by processing the user input and context information using an artificial intelligence model. By determining the input modality, the electronic device (10) may increase the correlation between the response and the user input and reduce the power consumption of the electronic device (10).

[0049] Figure 2 illustrates a block diagram of an artificial intelligence system according to one embodiment.

[0050] Referring to FIGS. 1B and 2 , according to one embodiment, the artificial intelligence system (200) may be configured to process received input using an artificial intelligence model and provide a result generated using the artificial intelligence model. In one example, the artificial intelligence system (200) may be implemented by an electronic device (10). Components of the artificial intelligence system (200) may be software modules (e.g., threads, functions, databases, and / or programs) implemented by the electronic device (10) executing instructions stored in the memory (130) using the processor (120). In one example, a part of the artificial intelligence system (200) may be implemented by an external device (e.g., a server). An instruction database (260) and / or an artificial intelligence model database (280) may be implemented by the external device. In this case, the electronic device (10) may transmit and receive information by communicating with the external device using a communication circuit (190).

[0051] In the example of FIG. 2, the artificial intelligence system (200) may include an I / O (input / output) interface (210), an artificial intelligence framework (220), a knowledge DB (260), an application (270), an artificial intelligence model DB (280), and / or an operation model determination module (290). The configurations of the artificial intelligence system (200) illustrated in FIG. 2 are examples, and at least some of the configurations may be implemented as a single software module.

[0052] According to one embodiment, the I / O (input / output) interface (210) may provide user input and / or context information to the artificial intelligence framework (220) and / or the action model determination module (290).

[0053] The user input may include a voice input (e.g., natural language input) obtained using the first microphone (182) and / or the second microphone (183). The user input may include a verbal input and / or a non-verbal input (e.g., an input for selecting a menu). The user input may include a text input obtained by performing STT for the voice input, an input entered through a button (185), a keyboard, and / or a text input obtained through an interface of a virtual keyboard. The user input may include an input obtained from any human interface device (HID) or external device communicatively connected to the electronic device (10). The user input may include visual information obtained using a camera (170) and / or an image sensor (141). The user input may include each of the above-described pieces of information or any combination of the above-described pieces of information.

[0054] Context information may include information related to the environment of the electronic device (10). For example, the context information may include a battery SoC (state of charge) of the electronic device (10), background sound acquired using the first microphone (182) and / or the second microphone (183), movement information acquired using the motion sensor (143), attachment information acquired using the attachment sensor (145), and / or shielding status information of the electronic device (10). The shielding status information may be acquired by detecting an external object using at least one of the camera (170), the communication circuit (190), the image sensor (141), and / or the proximity sensor (147). The context information may include a surrounding image of the electronic device (10) acquired using the camera (170) and / or the image sensor (141). The context information may include information on an application currently running on the electronic device (10) and / or location information of the electronic device (10). The context information may include each of the above-described pieces of information or any combination of the above-described pieces of information.

[0055] The I / O (input / output) interface (210) may be configured to output results generated by an artificial intelligence model. For example, the I / O interface (210) may receive results from the AI ​​framework (220). The I / O interface (210) may output the results in the form of natural language or content in any form. For example, the I / O interface (210) may visually output the results using the display (160). For example, the I / O interface (210) may output the results as audio content using the speaker (184). For example, the I / O interface (210) may transmit the results to an external device using the communication circuit (190), thereby causing the external device to output the results.

[0056] The artificial intelligence framework (220) may be configured to receive user input and / or context information from the I / O interface (210), and control components to perform an action corresponding to an intention (e.g., performing a task) corresponding to a user query based on the user query included in the user input. For example, the artificial intelligence framework (220) may include a prompt manager (230), an application manager (240), and / or an output manager (250).

[0057] The prompt manager (230) can generate prompts for input into an artificial intelligence model from user input and / or context information. For example, the prompt manager (230) can process the user input and / or context information into a form of prompts for input into a large language model (LLM), a large multi-modal model (LMM), and / or a large vision model (LVM). For example, if the input is a voice input, text conversion and natural language understanding can be performed on the voice input. For example, if the input includes an image, feature extraction or object identification can be performed on the image. The prompt manager (230) can generate prompts using text information, feature information, and / or object information extracted from the voice input. In one example, the prompt manager (230) can utilize a trained machine learning algorithm or an artificial intelligence neural network to generate prompts.

[0058] The knowledge database (260) can store history information related to the electronic device (10). For example, the knowledge database (260) can store user preference data, a prompt library, and / or prompt example data generated based on user input. The prompt manager (230) can generate prompts from user input and / or context information using data stored in the knowledge database (260). The prompt manager (230) can transmit the generated prompts to an artificial intelligence model (e.g., LLM, LVM, or LMM).

[0059] The application manager (240) may provide an interface between external components (e.g., knowledge DB (260), application (270), artificial intelligence model DB (280), and / or operation mode determination module (290)) and the artificial intelligence framework (220). For example, the application manager (240) may establish a channel between the external components and the artificial intelligence framework (220) using an application programming interface (API) or a plug-in.

[0060] For example, when user input is transmitted to an artificial intelligence model or when a prompt is generated, additional information may be requested. The application manager (240) can obtain the additional information by communicating with the application (270). The application manager (240) can provide a notification requesting additional information using the application (270) and can receive the additional information from the application (270). The application manager (240) can obtain the additional information by accessing an external database (e.g., a knowledge database (260). The application manager (240) can transmit the additional information to the prompt manager (230) and / or the artificial intelligence model.

[0061] In one example, user input may include an intent to perform a specified task. In this case, the output generated by the AI ​​model may include an action or a sequence of actions for performing the task corresponding to the intent. The application manager (240) may execute the action or sequence of actions included in the output using at least one application or service associated with the task.

[0062] The output manager (250) can perform fine-tuning on the output from the AI ​​model. For example, fine-tuning may include tuning the output based on content policies (e.g., user terms of use, harmfulness policies, or prohibited content policies) and / or relevance (e.g., relevance between user input and the output).

[0063] For example, the output manager (250) may determine whether the output contains prohibited content (e.g., racially or politically biased content). For example, the output manager (250) may determine whether the output contains harmful content (e.g., content requiring age verification or content that violates public order and morals). For example, the output manager (250) may determine whether content that violates the Terms of Service exists. If prohibited content, harmful content, and / or content that violates the Terms of Service exists, the output manager (250) may exclude the content from the output or replace the content with other content. In one example, if prohibited content, harmful content, and / or content that violates the Terms of Service exists, the output manager (250) may provide a notification informing the user that the content cannot be provided. In one example, the output manager (250) may provide guidance information to the user to prevent prohibited content, harmful content, and / or content that violates the Terms of Service from being output.

[0064] For example, the output manager (250) can verify the correlation between the output and the user input (e.g., the intent of the user input). For example, the output manager (250) can utilize an artificial intelligence model to extract information about the output. The output manager (250) can identify the degree of correlation between the extracted information and the user input by comparing the extracted information and the user input. If the degree of correlation is below a threshold, the output manager (250) can perform additional actions. For example, the output manager (250) can modify the user input and provide the modified user input to the prompt manager (230). The artificial intelligence system (200) can process the prompt generated based on the modified user input using the artificial intelligence model, thereby generating an output that matches the intent of the user input.

[0065] The application (270) may include any application installed on the electronic device (10). The application (270) may function as a user interface (e.g., front-end) and / or a client-end for the artificial intelligence framework (220). The application (270) may provide a user interface for obtaining input. In one example, the application (270) may be configured to process actions directed by the application manager (240).

[0066] The artificial intelligence model DB (280) can store at least one artificial intelligence model. The at least one artificial intelligence model can include an LMM, an LVM, and / or an LMM. For example, the artificial intelligence model can include at least one generative artificial intelligence model. The artificial intelligence model can include an artificial intelligence model trained to generate images and / or language. For example, the artificial intelligence model can include a generative adversarial network (GAN), a variational autoencoder (VAE), and / or a diffusion-based generative model for generating images. The diffusion-based generative model can utilize a VAE and a transformer structure. For example, the artificial intelligence model can include a model (e.g., CHAT-GPT 3, CHAT-GPT 4) trained to generate statistical output values ​​based on input values ​​for generating text data.

[0067] The operation mode determination module (290) may be configured to determine the operation mode of the electronic device (10) based on user input and / or context information. For example, the operation mode determination module (290) may determine the operation mode of the electronic device (10) based on a field of view (FoV) of the camera (170), an image acquired using the camera (170), information acquired using the sensor circuit (140), and / or information acquired through the interface (180).

[0068] For example, the operation mode determination module (290) can determine the operation mode based on the direction of the user's gaze and the direction of the camera's gaze (170). For example, the operation mode determination module (290) can determine the operation mode based on whether the electronic device (10) is shielded. The operation mode determination module (290) can determine the operation mode based on the attachment state of the attachment of the electronic device (10). The operation mode determination method of the operation mode determination module (290) is described in more detail in the examples described below with respect to FIGS. 4 to 10.

[0069] The operation mode determination module (290) can control the input to be used according to the determined operation mode. For example, the first mode and the second mode may be modes that support multi-modal input. For example, the third mode may be a mode that supports single-modal input. In the present disclosure, "determining the operation mode" may be referred to as determining the input modality or input type.

[0070] According to one embodiment, the operation mode determination module (290) may control the I / O interface (210) to obtain input according to the determined operation mode. For example, in the case of an operation mode that only uses user input, the operation mode determination module (290) may control the I / O interface (210) to obtain only user input. In this case, a module for obtaining context information (e.g., a sensor circuit (140) and / or a camera (170)) may not be activated.

[0071] According to one embodiment, the operation mode determination module (290) can control the prompt manager (230) to use only inputs according to the determined operation mode. For example, in the case of an operation mode that uses only user input, the operation mode determination module (290) can control the prompt manager (230) to use only the user input among the acquired user input and context information.

[0072] Figure 3 illustrates the structure of an artificial intelligence model according to one embodiment.

[0073] Referring to FIG. 3, according to one embodiment, the artificial intelligence model DB (280) of FIG. 2 may include at least one artificial neural network model (300). The artificial neural network model (300) may include a plurality of hidden layers (320). The hidden layers (320) are positioned between the input layer (310) and the output layer (330), and may include at least one layer learned while transmitting data (x1, x2, x3, ..., xn) (n is an integer greater than or equal to 4) transmitted from the input layer (310) to the output layer (330). For example, the hidden layers (320) may include a first hidden layer (320-1), a second hidden layer (320-2), and third and Mth hidden layers (320-M) (M is an integer greater than or equal to 3). The number of hidden layers illustrated in FIG. 3 is an example, and embodiments of the present disclosure are not limited thereto. For example, the artificial neural network model (300) may include more hidden layers than the number of hidden layers illustrated in FIG. 3, or may include fewer hidden layers than the number of hidden layers illustrated in FIG. 3. In one example, the artificial neural network model (300) may correspond to a multi-layer perception (MLP) model. Each of the hidden layers (320) may include a plurality of nodes (N1, N2, ..., Nk). The weight values ​​of each node may be learned using input data.

[0074] For example, the electronic device (10) of FIG. 1B can input the generated prompt as input data to an artificial neural network model (300). The input data is calculated by the artificial neural network model (300), and the artificial neural network model (300) can output output data (Y).

[0075] The artificial neural network model (300) is an example of an artificial intelligence model, and the embodiments of the present disclosure are not limited thereto. In the present disclosure, the term "artificial intelligence model" may include a generative artificial intelligence model. For example, the artificial intelligence model may include a large language model (LLM), a large multi-modal model (LMM), and / or a large vision model (LVM).

[0076] Figure 4 is a flowchart of a task performing method according to one embodiment.

[0077] Referring to FIGS. 1B and 4 , according to one embodiment, the electronic device (10) may perform a task based on user input. For example, the electronic device (10) may obtain a user input requesting the performance of a task for a specified project.

[0078] The operations described below with reference to FIG. 4 may be referred to as the operations of the electronic device (10) of FIG. 1B. The order of the operations described below with reference to FIG. 4 is merely an example, and embodiments of the present disclosure are not limited thereto. For example, at least some of the operations may be executed in a different order than that of FIG. 4, or may be executed substantially simultaneously with other operations of FIG. 4.

[0079] In operation 405, the electronic device (10) may detect a trigger event. A trigger event may be referred to as any event that causes the electronic device (10) to initiate performance of a task corresponding to a user's request. A trigger event may include a user's request to perform the task (e.g., an input instructing the task), an input subsequent to the task performance request (e.g., a task execution input subsequent to the input instructing the task), or an input preceding the task performance request. For example, a trigger event may include a voice input, a touch input, and / or a button input. The electronic device (10) may detect a trigger event when a voice input is received from the user using a microphone (e.g., a first microphone (182) and / or a second microphone (183)). The electronic device (10) may detect a trigger event when data corresponding to a voice input is received from an external device using a communication circuit (190). The electronic device (10) may detect a trigger event when a touch input or a button input is received through the interface (180).

[0080] In the example of FIG. 4, the user input may be referred to as an input that instructs a task. The input that instructs a task may include an intention for a designated task or an input that describes a designated task. If the trigger event is an input that instructs a task, the electronic device (10) may obtain the user input through operation 405. If the trigger event is an input that follows the input that instructs a task, the electronic device (10) may obtain the user input prior to operation 405. If the trigger event is an input that precedes the input that instructs a task, the electronic device (10) may obtain the user input after operation 410.

[0081] In operation 410, the electronic device (10) may identify an operating mode. For example, the electronic device (10) may identify an operating mode based on detection of a trigger event. The electronic device (10) may identify an operating mode when a trigger event is detected. In one example, the electronic device (10) may identify an operating mode using information about the operating mode stored in the memory (130). In this case, the operating mode may be predetermined prior to operation 410. In one example, the electronic device (10) may identify an operating mode by determining an operating mode when a trigger event is detected. A method for determining an operating mode may be described below with reference to FIG. 5.

[0082] According to an example, the operating mode may include a first mode, a second mode, and a third mode. The first mode and the second mode may be referred to as operating modes that support multiple modalities. The third mode may be referred to as an operating mode that supports a single modality (e.g., an operating mode that does not support multiple modalities). In the present disclosure, an “operating mode” may be referred to as an input modality or an input type. For example, identifying an operating mode of an electronic device (10) may be referred to as identifying an input modality or an input type of the electronic device (10).

[0083] When the operation mode is predetermined, the electronic device (10) can detect user input according to the predetermined operation mode. When the predetermined operation mode is a first mode or a second mode, the electronic device (10) can activate multiple components to detect multi-modal user input. For example, in the first mode or the second mode, the electronic device (10) can keep the microphone (e.g., the first microphone (182) and / or the second microphone (183)) and the camera (170) in an activated state. In this case, the electronic device (10) can obtain multi-modal user input using the microphone and the camera (170). When the predetermined operation mode is a third mode, the electronic device (10) can keep the microphone in an activated state to detect a single-modal user input. In this case, the electronic device (10) can keep the camera (170) in an inactive state or an idle state (e.g., a state in which image processing is not performed).

[0084] At step 415, the electronic device (10) may determine whether the identified operating mode supports multi-modality. If the identified operating mode is the first mode or the second mode, the electronic device (10) may determine that the operating mode supports multi-modality. If the identified operating mode is the third mode, the electronic device (10) may determine that the operating mode does not support multi-modality.

[0085] If the operation mode supports multi-modality (e.g., operation 415-YES), at operation 420, the electronic device (10) may generate a prompt based on user input and context information. For example, the electronic device (10) may generate a prompt using the prompt manager (230) described above with respect to FIG. 2. For example, the electronic device (10) may obtain a user input from a voice input obtained using a microphone. The electronic device (10) may obtain at least one image obtained using a camera (170) as context information. In this case, the electronic device (10) may generate a prompt (e.g., at least one first prompt) using the voice input and at least one image.

[0086] If the operation mode does not support multi-modality (e.g., operation 415-NO), in operation 425, the electronic device (10) may generate a user input-based prompt. In operation 425, the electronic device (10) may generate the prompt without using context information. The electronic device (10) may obtain a user input from a voice input obtained using a microphone, and may generate a prompt (e.g., a second prompt) from the obtained user input.

[0087] In operation 430, the electronic device (10) can generate a prompt-based result and output the generated result. The electronic device (10) can generate the result using the prompt of operation 420 (e.g., at least one first prompt) or the prompt of operation 425 (e.g., a second prompt). The electronic device (10) can generate the result by inputting the prompt into at least one artificial intelligence model (e.g., the artificial intelligence model DB (280) of FIG. 2). As described above with respect to FIG. 2, the electronic device (10) can generate the result by adjusting (e.g., fine tuning) the output of the artificial intelligence model. The electronic device (10) can output the generated result using the display (160) and / or the speaker (184). The electronic device (10) can transmit information about the generated result to an external device using the communication circuit (190) and cause the external device to output the result.

[0088] Figure 5 is a flowchart of a method for determining an operation mode according to one embodiment.

[0089] Referring to FIGS. 1B and 5, according to one embodiment, the electronic device (10) can determine an operating mode of the electronic device (10).

[0090] The operations described below with reference to FIG. 5 may be referred to as operations of the electronic device (10) of FIG. 1B. The order of the operations described below with reference to FIG. 5 is merely an example, and embodiments of the present disclosure are not limited thereto. For example, at least some of the operations may be executed in a different order from that of FIG. 5, or may be executed substantially simultaneously with other operations of FIG. 5.

[0091] In operation 505, the electronic device (10) may detect a trigger event. The trigger event of operation 505 may be referred to as an event that causes the electronic device (10) to determine an operation mode. For example, the trigger event may include a turn-on, activation, boot-up, reset, attachment state change, a specified period, a movement, or reception of a user input of the electronic device (10). The electronic device (10) may execute operation 510, which will be described later, upon turn-on, activation, reset, or boot-up. The electronic device (10) may execute operation 510 when movement information detected using the motion sensor (143) exceeds a threshold value. The electronic device (10) may execute operation 510, which will be described later, when a change in attachment state is detected using the attachment sensor (145). The electronic device (10) may execute operation 510 at every specified period. The electronic device (10) can execute operation 510 when user input is received.

[0092] In operation 510, the electronic device (10) may determine an operating mode using a camera (170) and / or a sensor circuit (140). The electronic device (10) may store information on the determined operating mode in a memory (130). If operation 510 is performed before receiving a user input instructing a task, the electronic device (10) may receive a user input (e.g., a multi-modal input or a single-modal input) according to the determined operating mode.

[0093] The electronic device (10) may determine the operating mode as a mode that supports multi-modality (e.g., a first mode or a second mode) when, for example, the field of view of the camera (170) corresponds to the field of view of the user. The electronic device (10) may determine the operating mode as a mode that does not support multi-modality (e.g., a third mode) when the field of view of the camera (170) does not correspond to the field of view of the user. The case where the field of view of the camera (170) corresponds to the field of view of the user may include the case where the direction of the field of view of the camera (170) and the direction of the field of view of the user are substantially the same or the case where the field of view of the camera (170) and the field of view of the user face each other. As described above with respect to FIG. 1B, the electronic device (10) may determine whether a camera (170) corresponding to the field of view of the user exists among a plurality of cameras based on the user's context information. For example, if a camera corresponding to the field of view of the user exists, the electronic device (10) may determine that the field of view of the camera (170) and the field of view of the user correspond.

[0094] The first mode may correspond to an operating mode of the electronic device (10) when the user's field of view and the camera's (170) field of view are oriented in the same direction. In the first mode, at least a portion of the user's field of view and at least a portion of the camera's (170) field of view may overlap each other. In the first mode, the camera (170) may capture an image of at least a portion of the field of view in the direction the user is looking or an image larger than the user's field of view. In this case, the electronic device (10) may obtain visual context information corresponding to the user's field of view using the camera (170). By utilizing the user input and the visual context, the electronic device (10) may provide a response that is more consistent with the intent indicated by the user input.

[0095] In one example, the first mode may correspond to the operating mode of the electronic device (10) when the content of the user's speech corresponds to the field of view of the camera (170). For example, the electronic device (10) may perform voice recognition on the user's voice input and identify at least one entity (e.g., a keyword referring to an object) from the voice input. The electronic device (10) may acquire an image using the camera (170) based on reception of the voice input and perform object recognition on the image. If the object recognized from the image matches the entity identified from the voice input, the electronic device (10) may determine the operating mode of the electronic device (10) as the first mode.

[0096] The second mode may correspond to the operating mode of the electronic device (10) when the user is positioned within the field of view of the camera (170). In the second mode, the field of view of the camera (170) and the field of view of the user may face each other, and a portion of the field of view of the camera (170) and the field of view of the user may overlap. For example, the electronic device (10) may be mounted at a certain location. The electronic device (10) may be placed at a designated location or may be held by the user. With the electronic device (10) mounted at the designated location, the user may look at the camera (170) of the electronic device (10) and speak a voice corresponding to the user input. In this case, the electronic device (10) may identify an image corresponding to the user from images acquired using the camera (170) and acquire visual context information using the identified image. For example, in the second mode, the user may speak, “How do I look today?” The electronic device (10) may provide a response corresponding to a user input using a user image acquired using a camera (170) and a user input corresponding to a voice input. For example, the electronic device (10) may provide a response to a user input based on the user's appearance, outfit, and / or clothing identified from the user's image.

[0097] The third mode may correspond to an operation mode in which context information is not required for processing user input. In the third mode, the user's field of view and the field of view of the camera (170) may not correspond to each other. As described above with respect to FIG. 1B, the electronic device (10) may determine whether a camera (170) corresponding to the user's field of view exists among a plurality of cameras based on the user's context information. For example, if there is no camera corresponding to the user's field of view, the electronic device (10) may determine that the field of view of the camera (170) and the user's field of view do not correspond.

[0098] For example, the electronic device (10) in the third mode may process user input without acquiring an image using the camera (170). In the third mode, the electronic device (10) may not receive input using the camera (170) or may not process the received input (e.g., by transmitting it to an artificial intelligence model). In the third mode, the electronic device (10) may be in a shielded state (e.g., in a state placed in a user's pocket or bag) or attached in a state where the user cannot be identified. In the third mode, there may be cases where camera input must be added to the processing of user input. For example, the user input may include a task for an image acquired using the camera (170). In this case, the electronic device (10) may output guide information suggesting a change in the position and / or mounting state of the electronic device (10) for changing the operation mode.

[0099] In the first mode and the second mode, the electronic device (10) can activate the microphone (182, 183) and the camera (170) to receive multi-modal input. For example, the electronic device (10) can continuously activate the microphone (182, 183) and the camera (170). In this case, the electronic device (10) can provide a response using context information acquired prior to the user input. For example, the user input can be “Isn’t that David who passed by earlier?” The electronic device (10) can provide a response to the user input by analyzing images acquired using the camera (170) prior to the user input.

[0100] In the first mode and the second mode, the electronic device (10) may constantly activate only the microphones (182, 183) to reduce power consumption. For example, when a user input is detected, the electronic device (10) may activate the camera (170) based on the detection of the user input. In one example, the electronic device (10) may activate the camera (170) when a speech is detected. By activating the camera (170) before the recognition of the voice input is completed, the delay between the user's speech and context information (e.g., an image acquired using the camera (170)) can be reduced.

[0101] In one example, the electronic device (10) may activate the microphone (182, 183) and the camera (170) when a specified input (e.g., input via a button (185), a gesture input, or a touch input) is received. The electronic device (10) may determine an operating mode after receiving the specified input.

[0102] In one embodiment, the electronic device (10) may determine an operation mode using a plurality of images acquired using the camera (170) and information from the motion sensor (143). For example, if the motion information based on the images acquired using the camera (170) corresponds to the motion information of the electronic device (10) acquired using the motion sensor (143), the electronic device (10) may determine that the field of view of the camera (170) and the field of view of the user correspond to each other. In this case, the electronic device (10) may determine the operation mode of the electronic device (10) as the first mode. If the motion information based on the images and the motion information of the electronic device (10) acquired using the motion sensor (143) do not correspond to each other, the electronic device (10) may determine the operation mode as the second mode if a user image exists in the images acquired using the camera (170). If the motion information based on the images does not correspond to the motion information of the electronic device (10) obtained using the motion sensor (143), and if there is no user image in the image obtained using the camera (170), the electronic device (10) may determine the operation mode as the third mode. The determination of the operation mode using the image and motion sensor (143) may be described later with reference to FIG. 6A.

[0103] In one embodiment, the electronic device (10) may determine an operation mode based on attachment information acquired using the attachment sensor (145). For example, the electronic device (10) may determine whether the electronic device (10) is worn by the user. If the electronic device (10) is worn by the user, the electronic device (10) may determine the operation mode as the first mode. If the electronic device (10) is not worn by the user, the electronic device (10) may determine the operation mode as the second mode if a user image exists in an image acquired using the camera (170). If the electronic device (10) is not worn by the user, the electronic device (10) may determine the operation mode as the third mode if a user image does not exist in an image acquired using the camera (170). Determination of an operation mode based on an attachment state may be described later with reference to FIGS. 8A and 8B .

[0104] In one embodiment, the electronic device (10) can determine an operation mode based on information acquired using microphones (182, 183). The electronic device (10) can perform beamforming using the first microphone (182) and the second microphone (183). Using the beamforming technology, the electronic device (10) can identify a relative position (e.g., relative direction) of a speaker with respect to the electronic device (10). If the position of the speaker is located at the rear of the electronic device (10) (e.g., facing in the opposite direction from the direction in which the camera (170) of the electronic device (10) faces), the electronic device (10) can determine the operation mode as the first mode. Determining the operation mode using a microphone can be described later with reference to FIG. 8A.

[0105] In one embodiment, the electronic device (10) may determine an operating mode based on whether it is shielded. The electronic device (10) may use a communication circuit (190), an image sensor (141), and / or a proximity sensor (147) to determine whether the electronic device (10) is in a shielded state (e.g., the electronic device (10) is positioned in a user's pocket or bag). If the electronic device (10) is in a shielded state, the electronic device (10) may determine the operating mode to be a third mode.

[0106] In one example, information for determining the operating mode may be insufficient. In one example, the electronic device (10) may provide feedback to the user to obtain the insufficient information. For example, the electronic device (10) may facilitate the acquisition of information for determining the operating mode of the electronic device (10) by providing guidance to the user, such as “Please show me again” or “Please position the camera of the electronic device in the same direction as the line of sight.” In one example, when information for determining the operating mode is insufficient, the electronic device (10) may determine the operating mode of the electronic device (10) in a designated mode (e.g., multi-modal input support mode).

[0107] Those skilled in the art will appreciate that various methods for setting an operating mode may be used in addition to the methods described above. For example, when a designated input (e.g., a button press, a designated gesture, or a designated number of button presses) is received, the electronic device (10) may set an operating mode mapped to the designated input.

[0108] Figure 6a is a flowchart of a method for determining an operation mode according to one embodiment.

[0109] Referring to FIGS. 1B and 6A, according to one embodiment, the electronic device (10) may determine an operating mode based on movement information. For example, the electronic device (10) may determine an operating mode based on information acquired using a camera (170) and information acquired using a sensor circuit (140).

[0110] The operations described below with respect to FIG. 6A may be referred to as operations of the electronic device (10) of FIG. 1B. The order of the operations described below with respect to FIG. 6A is merely an example, and embodiments of the present disclosure are not limited thereto. For example, at least some of the operations may be executed in a different order from that of FIG. 6A, or may be executed substantially simultaneously with other operations of FIG. 6A. For example, the operations described below with respect to FIG. 6A may correspond to operation 510 of FIG. 5.

[0111] In operation 605, the electronic device (10) may acquire a plurality of images using the camera (170). For example, the electronic device (10) may acquire a plurality of images when a trigger event for determining an operation mode is detected (e.g., operation 505 of FIG. 5). For example, the electronic device (10) may be configured to continuously store images for a certain period of time using the camera (170), and when a trigger event is detected, the electronic device (10) may use images within a certain time range from the time at which the trigger event is detected as a plurality of images. As described above with respect to FIG. 1B, the electronic device (10) may include a plurality of cameras. In this case, the electronic device (10) may acquire a plurality of images using a camera (170) determined based on context information of a user of the plurality of cameras.

[0112] In operation 610, the electronic device (10) can obtain movement information of the electronic device (10) using the motion sensor (143). The electronic device (10) can obtain movement information of the electronic device (10) that is time-synchronized with a plurality of images. For example, the electronic device (10) can obtain movement information of the electronic device (10) that is time-synchronized with a plurality of images by substantially simultaneously activating the camera (170) and the motion sensor (143).

[0113] In operation 615, the electronic device (10) may determine whether the image-based motion identified from the plurality of images corresponds to motion information. The electronic device (10) may identify a vector representing the movement of pixels between the images using consecutive images. The electronic device (10) may identify the image-based motion information using the vector values ​​of the images. If the image-based motion information corresponds to (e.g., is substantially the same as) the motion information of the electronic device (10) (e.g., motion information acquired using the motion sensor (143), the electronic device (10) may determine that the field of view of the camera (170) corresponds to the field of view of the user.

[0114] If the image-based movement corresponds to the movement information (e.g., operation 615-YES), in operation 620, the electronic device (10) may determine the operation mode as the first mode. In one example, if the user moves while holding the electronic device (10) in his / her hand (e.g., carrying the electronic device (10), the image-based movement and the movement information of the electronic device (10) may correspond. In this case, in order to distinguish between the first mode and the second mode, the electronic device (10) may determine whether a user image is identified from a plurality of images (e.g., operation 625). If the user image is identified, the electronic device (10) may determine the operation mode as the second mode. If the user image is not identified, the electronic device (10) may determine the operation mode as the first mode.

[0115] If the image-based motion does not correspond to the motion information (e.g., operation 615-NO), in operation 625, the electronic device (10) may determine whether a user image is identified from the plurality of images. For example, the electronic device (10) may perform image recognition on the plurality of images. As a result of the image recognition, if an image corresponding to the user (e.g., a stored face image of the user) is identified from at least some of the plurality of images, the electronic device (10) may determine that the user image is identified from the plurality of images.

[0116] If the user image is identified (e.g., operation 625-YES), in operation 630, the electronic device (10) may determine the operation mode as the second mode. If the user image is not identified (e.g., operation 625-NO), in operation 635, the electronic device (10) may determine the operation mode as the third mode.

[0117] In one example, if the electronic device (10) determines that the operating mode is not the first mode or the second mode, the electronic device (10) may determine the operating mode to be the third mode. In one example, if the electronic device (10) determines that the electronic device (10) is in a shielded state, the electronic device (10) may determine the operating mode to be the third mode. In this case, the electronic device (10) may not perform the operations illustrated in FIG. 6A.

[0118] In FIG. 6A, operation 625 is illustrated as being performed subsequent to operation 615, but embodiments of the present disclosure are not limited thereto. For example, operation 625 may be performed prior to operation 615. If a user image is identified, the electronic device (10) may determine the operation mode as the second mode. If the user image is not identified, the electronic device (10) may determine the operation mode as the first mode or the third mode according to operation 615. If the image-based movement corresponds to the movement information, the electronic device (10) may determine the operation mode as the first mode. If the image-based movement does not correspond to the movement information, the electronic device (10) may determine the operation mode as the third mode.

[0119] Figure 6b is a flowchart of a method for determining an operation mode according to one embodiment.

[0120] Referring to FIGS. 1B and 6B , according to one embodiment, the electronic device (10) may determine an operating mode based on whether voice and image inputs correspond. For example, the electronic device (10) may determine an operating mode based on information acquired using a camera (170) and voice input.

[0121] The operations described below with respect to FIG. 6B may be referred to as operations of the electronic device (10) of FIG. 1B. The order of the operations described below with respect to FIG. 6B is merely an example, and embodiments of the present disclosure are not limited thereto. For example, at least some operations may be executed differently from the order of FIG. 6B, or may be executed substantially simultaneously with other operations of FIG. 6B. For example, the operations described below with respect to FIG. 6B may correspond to operation 510 of FIG. 5. For example, it may be assumed that the electronic device (10) has acquired a voice input through operation 505.

[0122] In operation 655, the electronic device (10) may acquire at least one image using the camera (170). For example, the electronic device (10) may acquire at least one image when a trigger event for determining an operation mode is detected (e.g., operation 505 of FIG. 5). For example, the electronic device (10) may be configured to continuously store images for a certain period of time using the camera (170), and when a trigger event is detected, the electronic device (10) may use images within a certain time range from the time at which the trigger event is detected. As described above with respect to FIG. 1B, the electronic device (10) may include a plurality of cameras. In this case, the electronic device (10) may acquire at least one image using the camera (170) determined based on context information of a user of the plurality of cameras.

[0123] In operation 660, the electronic device (10) may perform image recognition on at least one image. For example, the electronic device (10) may identify at least one object from the image using any algorithm for image recognition.

[0124] At operation 665, the electronic device (10) may determine whether a voice-responsive object is identified. For example, the electronic device (10) may acquire a voice input through a trigger event (e.g., operation 505 of FIG. 5 ). The electronic device (10) may identify at least one keyword included in the voice input through voice recognition of the voice input. The electronic device (10) may determine that the voice-responsive object is identified if an object corresponding to the identified keyword is identified from at least one image.

[0125] If a voice-responsive object is identified (e.g., operation 665-YES), the operating mode may be determined as the first mode in operation 670. In operation 670, the electronic device (10) may also determine the operating mode as the second mode. For example, if a voice-responsive object is identified and an image of the user is also identified from the image, the electronic device (10) may determine the operating mode as the second mode.

[0126] If the voice response object is not identified (e.g., operation 665-NO), in operation 675, the electronic device (10) may determine the operation mode based on whether or not there is eye contact. For example, the electronic device (10) may determine the operation mode according to the method described above with respect to FIG. 6A. Without being limited to operation 675, the electronic device (10) may determine the operation mode according to any operation mode determination method described in the present disclosure. Hereinafter, methods for providing a response according to the operation mode of the electronic device (10) may be described with reference to FIGS. 7A to 7D . In the examples of FIGS. 7A to 7D , examples are described using the second electronic device (10b) of FIG. 1A , but one of ordinary skill in the art will understand that the embodiments of the present disclosure may also be applied to other types of electronic devices (e.g., the electronic device (10), the first electronic device (10a), or the third electronic device (10c)).

[0127] FIG. 7A illustrates a first mode operating environment of an electronic device according to one embodiment.

[0128] Referring to FIG. 7A, the second electronic device (10b) may be attached to the clothing of the user (701). In this case, the field of view (705) of the user (701) and the field of view (710) of the second electronic device (10b) (e.g., the field of view of the camera of the second electronic device (10b)) may face the same direction, and at least a portion of the field of view (705) of the user (701) and at least a portion of the field of view (710) of the second electronic device (10b) may overlap. The second electronic device (10b) may acquire an image of at least a portion of an area within the field of view (705). In this case, when the second electronic device (10b) receives a voice input from the user (701), the second electronic device (10b) may generate a prompt using information about the surrounding environment (e.g., context information) that the user (701) may be looking at. As described above with respect to FIG. 1B, the electronic device (10) may include multiple cameras. In this case, the electronic device (10) can use a camera among multiple cameras determined based on the user's context information.

[0129] For example, a user (701) may utter a voice input such as “It’s really cool.” In this case, the second electronic device (10b) may analyze an image captured using a camera. By performing image analysis on the captured image, the second electronic device (10b) may identify an image of the sea in the evening. In this case, the second electronic device (10b) may generate a prompt such as “I’m looking at the calm sea with the sunset” as a context-based prompt. In addition, the second electronic device (10b) may generate a prompt such as “A response to my exclamation of “It’s really cool” based on the voice input of the user (701). The second electronic device (10b) may generate an output by inputting the context-based prompt and the voice input-based prompt into an artificial intelligence model. In one example, the second electronic device (10b) may use information on an image captured using a camera as a context-based prompt. For an artificial intelligence model that supports image input and text input, the second electronic device (10b) can generate a result by processing an image using the artificial intelligence model together with a voice input-based prompt.

[0130] FIG. 7b illustrates a second mode operating environment of an electronic device according to one embodiment.

[0131] Referring to FIG. 7b, a user (702) may be moving in a second direction (795). At the same time, the user (702) may move a second electronic device (10b) in a first direction (790). In this case, the motion information of the second electronic device (10b) detected by the motion sensor of the second electronic device (10b) may be a value in which the movement in the second direction (795) is offset from the movement in the first direction (790). On the other hand, the motion information based on a plurality of images acquired using a camera of the second electronic device (10b) (e.g., a camera selected based on user context information) may correspond to the movement in the first direction (790). In this case, the motion information of the second electronic device (10b) and the image-based motion information may not correspond to each other. The second electronic device (10b) may determine the operation mode as the second mode.

[0132] In one example, the second electronic device (10b) may determine the operating mode of the second electronic device (10b) as the second mode based on the identification of an image corresponding to the user (702) within an image using the camera.

[0133] For example, a user (702) can hold a second electronic device (10b) in his / her hand and alternately capture images of his / her own face and another user's face (not shown) using the second electronic device (10b). While capturing, the user (702) can utter a voice input such as "Who looks older?" In this case, when capturing the user (702), the electronic device (10) can set the operation mode to the second mode and tag the image captured in the second mode with information indicating that the image was captured in the second mode. For example, the information indicating that the image was captured in the second mode can include information indicating that the person in the image is the user (702). When capturing another user, the electronic device (10) can set the operation mode to the first mode and tag the image captured in the first mode with information indicating that the image was captured in the first mode. For example, the information indicating that the image was captured in the first mode can include information indicating that the person in the image is not the user (702).

[0134] For example, a user (702) may utter, “My back hurts.” In this case, a second electronic device (10b) in a second mode may analyze an image of the user (702). The second electronic device (10b) may generate a context-based prompt, “User sitting in a chair and reading a book,” from the images of the user (702). Based on the utterance of the user (702), the second electronic device (10b) may generate a user input-based prompt, “User having a back hurt.” The second electronic device (10b) may process the context-based prompt and the user input-based prompt using an artificial intelligence model to generate a result. The second electronic device (10b) may provide a response based on the result. When providing a response, the second electronic device (10b) may also provide an image of the user associated with the context-based prompt (e.g., an image of the user sitting in a chair).

[0135] FIG. 7c illustrates a second mode operating environment of an electronic device according to one embodiment.

[0136] Referring to FIG. 7C, according to an embodiment, the second electronic device (10b) in the second mode may provide a proactive suggestion. In the second mode, the second electronic device (10b) may provide a proactive suggestion based on changes in the user (703), facial expressions of the user (703), poses of the user (703), actions of the user (703), changes in the surrounding environment of the user (703), or gestures of the user (703) without a voice input from the user (703). As described above with respect to FIG. 1B, the electronic device (10) may include multiple cameras. In this case, the electronic device (10) may use a camera determined based on context information of the user among the multiple cameras. For example, the electronic device (10) may use a camera among the multiple cameras that is used to acquire an image in which the user (703) is identified.

[0137] In the second mode, the second electronic device (10b) can analyze an image of the user (703) acquired using a camera. The second electronic device (10b) can provide suggestions to the user (703) based on the analyzed information. For example, the second electronic device (10b) can provide a suggestion such as, “User, you look tired. How about going to the bedroom?” The second electronic device (10b) can provide a notification based on contextual information (e.g., an image acquired using a camera and / or information acquired using a sensor circuit). The second electronic device (10b) can be configured to provide suggestions for a specified situation based on user preferences or user settings.

[0138] FIG. 7d illustrates a third mode operating environment of an electronic device according to one embodiment.

[0139] Referring to FIG. 7d, the second electronic device (10b) may be stored in the bag of the user (704). In this case, the second electronic device (10b) may detect that the second electronic device (10b) is in a shielded state. Based on the detection of the shielded state, the second electronic device (10b) may determine the operating mode of the second electronic device (10b) as a third mode. In the third mode, the second electronic device (10b) may operate based on the voice input of the user (704). For example, the second electronic device (10b) may keep the camera in an inactive or idle state.

[0140] FIG. 8A illustrates an electronic device according to one embodiment.

[0141] Referring to FIG. 8A, the second electronic device (10b) may include a first microphone (882) (e.g., the first microphone (182) of FIG. 1B), a second microphone (883) (e.g., the second microphone (183) of FIG. 1B), a camera (870) (e.g., the camera (170) of FIG. 1B), and a display (860) (e.g., the display (160) of FIG. 1B). For example, the second electronic device (10b) may include a main body (801) and an attachment part (802). For example, the main body (801) and the attachment part (802) may be coupled based on magnetic force. In FIG. 8A, examples are described focusing on the second electronic device (10b) for convenience of explanation, but those skilled in the art will understand that the same examples may be applied to any electronic device having a similar structure.

[0142] For example, the attachment portion (802) may include at least one magnet for coupling with the main body portion (801). The attachment portion (802) may include a battery and / or a processing circuit. The attachment portion (802) may supply power to the main body portion (801) by electromagnetically communicating with the main body portion (801). The attachment portion (802) may include at least some of the components of the second electronic device (10b). The attachment portion (802) may implement the functions of the second electronic device (10b) together with the main body portion (801) through electromagnetic coupling with the main body portion (801). For example, when the attachment portion (802) is coupled with the main body portion (801), the attachment portion (802) may supply power to the main body portion (801) using a battery. Based on the coupling of the attachment portion (802), the second electronic device (10b) may be booted. The second electronic device (10b) can perform an operation for determining the operation mode (e.g., operation 510 of FIG. 5) at boot time. In one example, after the main body (801) and the attachment part (802) are combined, the second electronic device (10b) can perform an operation for determining the operation mode when a trigger event for determining the operation mode is detected (e.g., operation 505 of FIG. 5).

[0143] According to one embodiment, the second electronic device (10b) may be mounted on the user's clothing, body, or object in various forms. For example, the second electronic device (10b) may detect the attachment state of the attachment portion (802). The second electronic device (10b) may detect the attachment state using an attachment sensor (e.g., the attachment sensor (145) of FIG. 1B). When the second electronic device (10b) is not attached to another object, the main body (801) and the attachment portion (802) may be in close contact (e.g., not worn). When the second electronic device (10b) is attached to another object (e.g., worn), a gap may be generated between the main body (801) and the attachment portion (802) due to another object (e.g., clothes, a bag). As the distance (d) between the main body (801) and the attachment part (802) changes, the magnitude of the magnetic force detectable by the main body (801) can change. Accordingly, the second electronic device (10b) can detect the attachment state by detecting the difference in magnetic force generated in the wearing state and the non-wearing state. For example, the second electronic device (10b) can determine the attachment state as the non-wearing state if the detected magnetic force exceeds a threshold value. The second electronic device (10b) can determine the attachment state as the wearing state if the detected magnetic force is a value within a specified range below the threshold value.

[0144] With reference to FIG. 8A, operations for determining the attachment state between the main body (801) and the attachment portion (802) based on magnetic force by the second electronic device (10b) have been described; however, embodiments of the present disclosure are not limited thereto. For example, the second electronic device (10b) may include an attachment sensor that may cause mechanical deformation depending on the attachment state. The second electronic device (10b) may identify the attachment state by detecting the mechanical deformation.

[0145] According to one embodiment, the second electronic device (10b) may determine the operation mode based on the attachment state. For example, if the attachment state of the second electronic device (10b) is a worn state, the second electronic device (10b) may determine the operation mode as the first mode. Even if the attachment state is a worn state, if the second electronic device (10b) is shielded, the second electronic device (10b) may determine the operation mode as the third mode. If the attachment state is a non-wearing state, the second electronic device (10b) may determine the operation mode as the second mode or the third mode. For example, the second electronic device (10b) may determine the operation mode according to the method described above with respect to FIGS. 4 to 6b.

[0146] In the first mode of the worn state, the second electronic device (10b) can perform a life-log operation. According to the life-log operation, the second electronic device (10b) can continuously record audio and video (e.g., for a certain period of time) even without user input. In one example, the second electronic device (10b) can perform the life-log operation according to user settings. During the performance of the life-log operation, the second electronic device (10b) can generate a context-based prompt using the acquired image and generate a user input-based prompt based on the user's voice.

[0147] In a non-wearing state, the second electronic device (10b) can use a motion sensor to determine whether the second electronic device (10b) is in a carry state or a stand state. The carry state may be referred to as a state in which the second electronic device (10b) is carried by the user. If the second electronic device (10b) detects a movement exceeding a threshold value using the motion sensor in the non-wearing state, the second electronic device (10b) may determine that the second electronic device (10b) is in a carry state. The stand state may be referred to as a state in which the second electronic device (10b) is placed in a fixed position. If the second electronic device (10b) does not detect a movement exceeding a threshold value in the non-wearing state, the second electronic device (10b) may be determined to be in a stand state.

[0148] When in the carry state, the second electronic device (10b) can determine the operation mode based on whether the user image is recognized. If the user image is recognized from an image acquired using the camera (870), the second electronic device (10b) can determine the operation mode as the first mode. If the user image is not recognized from an image acquired using the camera (870), the second electronic device (10b) can determine the operation mode as the second mode.

[0149] According to one embodiment, the second electronic device (10b) can determine the operating mode based on the attachment state and the attachment location. For example, the second electronic device (10b) can detect the attached location of the second electronic device (10b) in a worn state. The second electronic device (10b) can detect the attached location of the second electronic device (10b) using a sensor capable of detecting magnetic force, a proximity sensor, and / or a sensor configured to detect a physical connection state. The second electronic device (10b) can determine the operating mode based on the operating mode set for the attached location. In one example, the second electronic device (10b) can determine that the attached location is the user's body. In this case, the second electronic device (10b) can control the settings of the camera (870) to acquire an image in a state of high movement. For example, the second electronic device (10b) can adjust image stabilization (e.g., optical image stabilization, video digital image stabilization), auto focus, white balance, exposure value, and / or shutter speed.

[0150] According to one embodiment, when in a stand state, the second electronic device (10b) may determine an operation mode (e.g., operation 510 of FIG. 5) based on detection of movement, passage of a specified period of time, or user input. If movement greater than a specified value is detected in the stand state, the second electronic device (10b) may determine an operation mode. If a specified period of time passes without detecting movement greater than a specified value or user input in the stand state, the second electronic device (10b) may determine an operation mode. If a user input is detected in the stand state, the second electronic device (10b) may determine an operation mode. In the stand state, the second electronic device (10b) may wait with the microphones (882, 883) activated without activating the camera (870). If movement is detected, a specified period of time passes, or a user input is received during the standby, the second electronic device (10b) may activate the camera (870) to determine an operation mode.

[0151] According to one embodiment, the second electronic device (10b) may determine the operation mode based on the position of the speaker. For example, the second electronic device (10b) may receive the user's speech using the first microphone (882) and the second microphone (883). Based on beamforming technology, the second electronic device (10b) may identify the relative position of the speaker (e.g., the user) with respect to the second electronic device (10b). If the speaker is located at the rear of the second electronic device (10b) (e.g., in the opposite direction from the direction in which the camera (870) faces), the second electronic device (10b) may determine the operation mode as the first mode. For example, if the speaker is located at the rear when the second electronic device (10b) is worn, the second electronic device (10b) may determine that the field of view of the camera (870) and the field of view of the user correspond to each other when the speaker is located at the rear when the second electronic device (10b) is worn. When the speaker is not located at the rear of the second electronic device (10b) in the attached state, the second electronic device (10b) can determine the operation mode as the third mode.

[0152] In one example, the second electronic device (10b) may determine an operating mode based on the speaker's location based on the attachment location of the second electronic device (10b). For example, the second electronic device (10b) may determine an operating mode based on the speaker's location when mounted at a designated location, mounted on a designated accessory, or mounted on a designated body part.

[0153] FIG. 8b illustrates an electronic device according to one embodiment.

[0154] Referring to FIG. 8B, according to one embodiment, the second electronic device (10b) may have a cylindrical shape. Since the main body (801) and the attachment portion (802) are formed in a cylindrical shape, a user can rotate the attachment portion (802) relative to the main body (801) in a wearing or non-wearing state. The second electronic device (10b) may receive user input based on the rotation of the attachment portion (802). For example, the second electronic device (10b) may set an operation mode based on the rotation direction.

[0155] FIG. 9 is a flowchart of a method for providing a response based on modality determination according to one embodiment.

[0156] Referring to FIGS. 1B and 9 , according to one embodiment, the electronic device (10) may provide a response based on a user input. For example, the electronic device (10) may provide a response using an input based on a determined modality.

[0157] The operations described below with reference to FIG. 9 may be referred to as operations of the electronic device (10) of FIG. 1B. The order of the operations described below with reference to FIG. 9 is merely an example, and embodiments of the present disclosure are not limited thereto. For example, at least some of the operations may be executed in a different order than that of FIG. 9, or may be executed substantially simultaneously with other operations of FIG. 9.

[0158] For example, the electronic device (10) may include at least one sensor (e.g., a sensor circuit (140), a camera (170), a microphone (e.g., a first microphone (182) and / or a second microphone (183)), a memory (130), and at least one processor (e.g., a processor (120). The at least one processor may be communicatively connected to at least one sensor, a camera (170), a microphone, and a memory (130). The memory (130) may store instructions that, when individually or collectively executed by the at least one processor, cause the electronic device (10) to perform the operations described below.

[0159] In operation 905, the electronic device (10) can obtain a voice input. For example, the electronic device (10) can obtain a voice input from a user using a microphone.

[0160] In operation 910, the electronic device (10) may determine whether the field of view of the camera (170) corresponds to the field of view of the user. For example, the electronic device (10) may identify the relative position of the user with respect to the electronic device (10) from a voice input acquired using a plurality of microphones (e.g., a first microphone (182) and a second microphone (183)). If the relative position is located in the opposite direction to the direction corresponding to the field of view of the camera (170) (e.g., toward the rear of the electronic device (10)), the electronic device (10) may determine that the field of view of the camera (170) and the field of view of the user correspond to each other. For example, the electronic device (10) may acquire movement information of the electronic device (10) using the motion sensor (143) and acquire a plurality of images using the camera (170). When the motion information of the electronic device (10) and the image-based motion information acquired from a plurality of images correspond to each other, the electronic device (10) may determine that the field of view of the camera (170) and the field of view of the user correspond to each other. In one example, the electronic device (10) may acquire at least one image using the camera (170) based on a voice input. When an image corresponding to the user is identified from the at least one image, the electronic device (10) may determine that the field of view of the camera (170) and the field of view of the user correspond to each other. For example, when the blocking of the camera (170) is detected, the electronic device (10) may determine that the field of view of the camera (170) and the field of view of the user do not correspond to each other. As described above with respect to FIG. 1B, the camera (170) may include a camera selected from a plurality of cameras.

[0161] Operation 910 may correspond to an operation mode determination operation. For example, the electronic device (10) may determine the operation mode according to the operation mode determination operation described above with reference to FIGS. 4 to 8B. For example, if the operation mode is determined to be the first mode or the second mode, the electronic device (10) may determine that the field of view of the camera (170) corresponds to the user's field of view. For example, if the operation mode is determined to be the third mode, the electronic device (10) may determine that the field of view of the camera (170) does not correspond to the user's field of view.

[0162] As described above, the electronic device (10) can perform operation 910 based on a trigger event. For example, if the attachment state of the electronic device (10) detected using the attachment sensor (145) is a wearing state, the electronic device (10) can determine whether the field of view of the camera (170) corresponds to the field of view of the user. For example, if the attachment state of the electronic device (10) detected using the attachment sensor (145) is a non-wearing state, the electronic device (10) can acquire at least one image using the camera (170). If an image corresponding to the user is identified from the at least one image, the electronic device (10) can determine that the field of view of the camera (170) and the field of view of the user correspond.

[0163] For example, the electronic device (10) may include a main body (e.g., main body (801) of FIGS. 8A and 8B) and an attachment portion (e.g., attachment portion (802) of FIGS. 8A and 8B). The attachment portion may be configured to be coupled to the main body based on magnetic force. The attachment sensor (145) may be configured to detect the attachment state of the electronic device (10) based on magnetic force.

[0164] As described above, the operating mode of the electronic device (10) may be determined prior to acquisition of the voice input (e.g., operation 905). In this case, operation 910 may be omitted from the method for providing a response based on modality determination. The electronic device (10) may generate a prompt using a multi-modal input or a single-modal input depending on the determined operating mode.

[0165] If the field of view of the camera (170) does not correspond to the field of view of the user (e.g., operation 910-NO), the electronic device (10) may perform operations according to reference point A. Operations according to reference point A may be described later with reference to FIG. 10.

[0166] If the field of view of the camera (170) corresponds to the field of view of the user (e.g., operation 910-YES), in operation 915, the electronic device (10) may generate at least one prompt (e.g., at least one first prompt) based on an image (e.g., an image acquired using the camera (170)) and a voice input. The at least one prompt may include a user input-based prompt and a context information-based prompt. For example, the electronic device (10) may generate a user input-based prompt based on a voice input. The electronic device (10) may generate a context information-based prompt (e.g., text data generated by recognizing the image or the image data itself) using an image. For example, the electronic device (10) may perform image recognition on the acquired image and generate at least one prompt based on a combination of a result of the image recognition (e.g., text data) and a voice input (e.g., text data recognized from the voice input). The electronic device (10) may generate at least one prompt according to operation 420 of FIG. 4.

[0167] In one example, the electronic device (10) may determine the operating mode as a multi-modal operating mode (e.g., the first mode or the second mode) even when the field of view of the camera (170) does not correspond to the field of view of the user (e.g., operation 910-NO). For example, the electronic device (10) may determine the operating mode according to the method described above with respect to FIG. 6B. The operating mode when the field of view of the camera (170) and the field of view of the user do not correspond may be referred to as a separate fourth mode. In this case, the fourth mode may be referred to as a multi-modal mode that utilizes both voice input and camera input. In the fourth mode, the electronic device (10) may generate a prompt that includes a description of an image related to the user's utterance. The electronic device (10) may generate a response by inputting the prompt into a generative artificial intelligence model and provide the generated response.

[0168] In operation 920, the electronic device (10) can obtain result data (e.g., first result data). The electronic device (10) can obtain the result data by inputting at least one generated prompt into an artificial intelligence model (e.g., an artificial intelligence model of the artificial intelligence model DB (280) of FIG. 2).

[0169] At operation 925, the electronic device (10) may provide a response associated with the voice input based on the result data. For example, the electronic device (10) may provide the response according to the methods described above with respect to the I / O interface (210) of FIG. 2. As described above with respect to the output manager (250) of FIG. 2, the electronic device (10) may perform fine tuning on the result data.

[0170] FIG. 10 is a flowchart of a method for providing a response based on modality determination according to one embodiment.

[0171] Referring to FIGS. 1B and 10 , according to one embodiment, the electronic device (10) may provide a response according to the second mode or the third mode. The operations described below with respect to FIG. 10 may be referred to as operations of the electronic device (10) of FIG. 1B . The order of the operations described below with respect to FIG. 10 is merely an example, and embodiments of the present disclosure are not limited thereto. For example, at least some of the operations may be executed differently from the order of FIG. 10 , or may be executed substantially simultaneously with other operations of FIG. 10 . When the field of view of the camera (170) does not correspond to the field of view of the user (e.g., operation 910-NO of FIG. 9 ), the electronic device (10) may perform the operations of FIG. 10 .

[0172] In operation 1010, the electronic device (10) may generate at least one prompt (e.g., a second prompt) based on a voice input (e.g., the voice input of operation 905 of FIG. 9 ). In this case, the electronic device (10) may generate a user input-based prompt based solely on the voice input, without using contextual information (e.g., an image). For example, the electronic device (10) may generate a prompt according to operation 425.

[0173] In operation 1015, the electronic device (10) can obtain result data (e.g., second result data). The electronic device (10) can obtain the result data by inputting at least one generated prompt into an artificial intelligence model (e.g., an artificial intelligence model of the artificial intelligence model DB (280) of FIG. 2).

[0174] In operation 1020, the electronic device (10) may provide a response associated with the voice input based on the result data. For example, the electronic device (10) may provide the response according to the methods described above with respect to the I / O interface (210) of FIG. 2. As described above with respect to the output manager (250) of FIG. 2, the electronic device (10) may perform fine tuning on the result data.

[0175] The operations of the electronic device (10) described above with reference to FIGS. 9 and 10 are part of the embodiments described above with reference to FIGS. 1A to 8B , and a person skilled in the art will understand that the electronic device (10) can perform the embodiments described above with reference to FIGS. 1A to 8B . For example, in the embodiments of FIGS. 9 and 10 , the electronic device (10) does not distinguish between the first mode and the second mode, but the electronic device (10) can perform an operation to determine the first mode or the second mode when the field of view of the camera (170) corresponds to the field of view of the user. For example, unlike the prompt generated in the first mode, the prompt generated in the second mode may include information indicating that the image corresponds to the user's image. Accordingly, the result data in the first mode may be different from the result data in the second mode.

[0176] FIG. 11 is a block diagram of an exemplary electronic device (1100) capable of performing the operations described in this document.

[0177] Referring to FIG. 11, the electronic device (1100) may be one of various forms of electronic devices, such as a notebook (1190), smartphones (1191) having various form factors (e.g., a bar-type smartphone (1191-1), a foldable-type smartphone (1191-2), or a sliderable (or rollable) type smartphone (1191-3)), a tablet (1192), a cellular phone (not shown), and other similar computing devices (not shown). The components, their relationships, and their functions illustrated in FIG. 11 are exemplary only and do not limit the implementations described or claimed in this document. The electronic device (1100) may be referred to as a mobile device, a user device, a multi-function device, a portable device, or a server.

[0178] The electronic device (1100) may include components including at least one processor (1110) (hereinafter referred to as processor (1110)), at least one memory (1120) (hereinafter referred to as memory (1120)), at least one display (1140) (hereinafter referred to as display (1140)), at least one image sensor (1150) (hereinafter referred to as image sensor (1150)), at least one communication circuit (1160) (hereinafter referred to as communication circuit (1160)), and / or at least one sensor (1170) (hereinafter referred to as sensor (1170)). The above components are merely exemplary. For example, the electronic device (1100) may include other components (e.g., power management integrated circuitry (PMIC), audio processing circuitry, an antenna, a rechargeable battery, or an input / output interface). For example, some components may be omitted from the electronic device (1100). For example, several components can be combined into one component.

[0179] The processor (1110) may be implemented as one or more IC (integrated circuit (or circuitry)) chips and may perform various data processing. The processor (1110) may include at least one electrical circuit and may individually or collectively perform distributed processing of instructions (or programs, data, etc.) stored in the memory (1120). The processor (1110) may include a processor assembly including one or more processing circuits. The processor (1110) may include any processing circuit operative to control the performance and operations of one or more components of the electronic device (1100) (e.g., the memory (1120), the display (1140), the image sensor (1150), the communication circuit (1160), and / or the sensor (1170)). For example, the processor (1110) (e.g., an application processor (AP)) may be implemented as a system on chip (SoC) (e.g., a single chip or chipset). For example, the processor (1110) may be implemented as multiple cores (or at least one core circuit), multiple chips, or multiple chipsets. For example, the processor (1110) may include one or more processing circuits. For example, the processor (1110) may include one or more processing circuits configured to individually and / or collectively perform various functions of the present disclosure. As a non-limiting example, at least a portion of the processor (1110) may be included in a first chip of the electronic device (1100), and at least another portion of the processor (1110) may be included in a second chip of the electronic device (1100) that is different from the first chip of the electronic device (1100).

[0180] For example, the processor (1110) may include a central processing unit (CPU) (1111), a graphics processing unit (GPU) (1112), a neural processing unit (NPU) (1113), an image signal processor (ISP) (1114), a display controller (1115), a memory controller (1116), a storage controller (1117), a communication processor (CP) (1118), and / or a sensor interface (1119). These components of the processor (1110) are merely exemplary. For example, the processor (1110) may further include other components. For example, some components of the processor (1110) may be omitted from the processor (1110). For example, some components of the processor (1110) may be included as separate components of the electronic device (1100) outside the processor (1110). For example, some components of the processor (1110) (e.g., memory controller (1116)) may be included within other components (e.g., at least a portion of memory (1120), an interface (e.g., available for connection to at least one component of the electronic device (100)), a display (1140) and / or an image sensor (1150)).

[0181] The processor (1110) may cause other components of the electronic device (1100) to perform various operations by executing instructions stored in the memory (1120). The CPU (1111) (or central processing circuit) may be configured to control components of the processor (1110) based on the execution of instructions stored in the memory (1120) (e.g., volatile memory (1121) and / or non-volatile memory (1122)). The GPU (1112) (or graphics processing circuit) may be configured to execute parallel operations (e.g., rendering). The NPU (1113) (or neural processing circuit, or artificial intelligence (AI) chip) may be configured to execute operations for an artificial intelligence model (e.g., convolution computation). The ISP (1114) (or image signal processing circuit) may be configured to process a raw image acquired through the image sensor (1150) into a format suitable for a component within the electronic device (1100) or a component of the processor (1110). The display controller (1115) (or display control circuit, or display processing unit (DPU)) may be configured to process an image acquired from the CPU (1111), the GPU (1112), the ISP (1114), or the memory (1120) (e.g., the volatile memory (1121)) into a format suitable for the display (1140). The memory controller (1116) (or memory control circuit) may be configured to control reading data from the volatile memory (1121) and writing data to the volatile memory (1121). The storage controller (1117) (or storage control circuit) may be configured to control reading data from and writing data to the nonvolatile memory (1122).The CP (1118) (communication processing circuit) may be configured to process data obtained from a component of the processor (1110) into a format suitable for transmission to another electronic device via the communication circuit (1160), or to process data obtained from another electronic device via the communication circuit (1160) into a format suitable for processing by the component of the processor (1110). For example, the communication circuit (1160) may include one or more communication circuits. The sensor interface (1119) (or sensing data processing circuit, sensor hub) may be configured to process data on the state of the electronic device (1100) and / or the state of the surroundings of the electronic device (1100), obtained via the sensor (1170), into a format suitable for the component of the processor (1110).

[0182] The memory (1120) may include one or more storage media (or one or more storage devices). For example, the memory (1120) may include a memory assembly including one or more storage media. For example, the one or more storage media may include permanent memory (e.g., non-volatile memory (1122)) such as a hard drive, flash memory, read-only memory (ROM), semi-permanent memory (e.g., volatile memory (1121)) such as random access memory (RAM), any other suitable type of storage (or storage assembly), or any combination thereof. The memory (1120) may include cache memory, which is one or more different types of memory used to temporarily store data for a function or feature of the electronic device (1100). As a non-limiting example, the cache memory may be included within the processor (1110). The memory (1120) may be fixedly embedded within the electronic device (1100) or incorporated into one or more suitable types of components (e.g., a subscriber identity module (SIM) card and / or a secure digital (SD) card) that may be repeatedly inserted into and removed from the electronic device (1100).

[0183] For example, the memory (1120) may store one or more software applications, such as an operating system (or system) software application, a firmware software application, a driver software application, a plug-in (e.g., add-in, add-on, and / or applet) software application, and / or any other suitable software applications. For example, the one or more software applications may include instructions executable by the processor (1110). For example, the memory (1120) may store instructions callable by an application programming interface (API). For example, the memory (1120) may store instructions within a library.

Claims

1. In an electronic device (10), At least one sensor (140); Camera (170); Mike (182, 183); memory (130); and At least one processor (120) communicatively connected to the at least one sensor, the camera, the microphone, and the memory, The above memory, when executed individually or in combination by the at least one processor, causes the electronic device to: Obtain voice input from the user using the above microphone, Determine whether the field of view of the above camera and the field of view of the above user correspond to each other, When the field of view of the camera and the field of view of the user correspond to each other, at least one first prompt is generated based on the image acquired by the camera and the voice input, Obtaining first result data by inputting at least one first prompt into an artificial intelligence model, An electronic device storing instructions for providing a response associated with the voice input based on the first result data.

2. In paragraph 1, Includes multiple microphones, The above instructions, when individually or in combination executed by the at least one processor, cause the electronic device to: By obtaining the voice input using the plurality of microphones, the relative position of the user with respect to the electronic device is identified, An electronic device that determines that the field of view of the camera and the field of view of the user correspond to each other when the relative position is located in the opposite direction to the direction corresponding to the field of view of the camera.

3. In paragraph 1, Including more motion sensors, The above instructions, when individually or in combination executed by the at least one processor, cause the electronic device to: Obtaining movement information of the electronic device using the above motion sensor, By acquiring multiple images using the above camera, image-based motion information is acquired from the multiple images, An electronic device that determines that the field of view of the camera and the field of view of the user correspond to each other when the movement information of the electronic device and the image-based movement information correspond to each other.

4. In paragraph 1, The above instructions, when individually or in combination executed by the at least one processor, cause the electronic device to: Based on the above voice input, at least one image is acquired using the camera, An electronic device that determines that the field of view of the camera and the field of view of the user correspond to each other when an image corresponding to the user is identified from the at least one image.

5. In paragraph 1, The above instructions, when individually or in combination executed by the at least one processor, cause the electronic device to: If the field of view of the camera and the field of view of the user do not correspond to each other, a second prompt based on the voice input is generated, By inputting the above second prompt into the artificial intelligence model, the second result data is obtained, An electronic device that provides a response related to the voice input based on the second result data.

6. In paragraph 1, The above instructions, when individually or in combination executed by the at least one processor, cause the electronic device to: An electronic device that determines that the field of view of the camera and the field of view of the user do not correspond to each other when the blocking of the camera is detected.

7. In paragraph 1, The above instructions, when individually or in combination executed by the at least one processor, cause the electronic device to: Perform image recognition on the acquired image, An electronic device configured to generate at least one first prompt based on a combination of the results of the image recognition and the voice input.

8. In paragraph 1, Further comprising an attachment sensor (145) set to detect the attachment status of the electronic device; The above instructions, when individually or in combination executed by the at least one processor, cause the electronic device to: An electronic device that determines that the field of view of the camera and the field of view of the user correspond to each other when the attachment state detected using the above attachment sensor is a wearing state.

9. In paragraph 8, The above instructions, when individually or in combination executed by the at least one processor, cause the electronic device to: If the attachment state detected using the attachment sensor is a non-wearing state, at least one image is acquired using the camera, An electronic device that determines that the field of view of the camera and the field of view of the user correspond when an image corresponding to the user is identified from at least one of the images.

10. A method for providing a response based on determining the modality of an electronic device, An action to obtain voice input; and Including an operation of determining whether the view of the camera of the electronic device and the view of the user correspond to each other; The above method, when the field of view of the camera and the field of view of the user correspond to each other: An operation of generating at least one first prompt based on an image acquired by the camera and the voice input; An operation of obtaining first result data by inputting at least one first prompt into an artificial intelligence model; and A method further comprising an action of providing a response associated with the voice input based on the first result data.

11. In paragraph 10, The operation of determining whether the field of view of the camera of the above electronic device corresponds to the field of view of the user is: An operation of identifying a relative position of the user with respect to the electronic device by obtaining the voice input using a plurality of microphones of the electronic device; and A method comprising an action of determining that the field of view of the camera and the field of view of the user correspond to each other when the relative position is located in a direction opposite to the direction corresponding to the field of view of the camera.

12. In paragraph 10, The operation of determining whether the field of view of the camera of the above electronic device corresponds to the field of view of the user is: An action of obtaining movement information of the electronic device using a motion sensor of the electronic device; An operation of acquiring image-based motion information from a plurality of images by acquiring a plurality of images using the camera; and A method comprising an action of determining that the field of view of the camera and the field of view of the user correspond to each other when the movement information of the electronic device and the image-based movement information correspond to each other.

13. In paragraph 10, The operation of determining whether the field of view of the camera of the above electronic device corresponds to the field of view of the user is: An operation of acquiring at least one image using the camera based on the voice input; and A method comprising an action of determining that a field of view of the camera and a field of view of the user correspond to each other when an image corresponding to the user is identified from at least one image.

14. In paragraph 10, An action of generating a second prompt based on the voice input when the field of view of the camera and the field of view of the user do not correspond to each other; An operation of obtaining second result data by inputting the second prompt into an artificial intelligence model; and A method further comprising an action of providing a response associated with the voice input based on the second result data.

15. In paragraph 10, The operation of determining whether the field of view of the camera of the above electronic device corresponds to the field of view of the user is: A method comprising an action of determining that the field of view of the camera and the field of view of the user do not correspond to each other when the blocking of the camera is detected.

Citation Information

Patent Citations

  • Method of configuring sensor network for monitoring a patient with sensor nodes

    KR1020230106314A

  • Method and apparatus for predicting forest fires combining heterogeneous data using ai-based object identification model

    KR1020240007883A

  • Juicer with dehydration function

    KR102423451B1

  • Method and apparatus for interaction with an intelligent personal assistant

    US11323665B2

  • KR20210020219A