Dialog assistance method and electronic device therefor
The electronic device addresses the lack of personalized reactions in extended reality by using AI to analyze verbal and non-verbal cues, enhancing conversational interactions with contextually relevant responses.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-10-29
- Publication Date
- 2026-06-04
AI Technical Summary
Existing electronic devices for extended reality lack the ability to provide personalized and context-aware recommended reactions during conversations, failing to consider non-verbal cues and emotional states of users.
An electronic device equipped with cameras, microphones, and AI agents that analyze verbal and non-verbal expressions, biometric data, and contextual information to generate personalized recommended reactions for users in real-time.
Enhances conversational interactions by providing contextually relevant and emotionally sensitive responses, improving user engagement and interaction quality in augmented and mixed reality environments.
Smart Images

Figure KR2025017457_04062026_PF_FP_ABST
Abstract
Description
Conversation assistance method and electronic device for the same
[0001] The embodiments disclosed in this document relate to a method for assisting dialogue and an electronic device for the same.
[0002] Electronic devices for providing extended reality (XR) are being researched. XR may include augmented reality (AR), virtual reality (VR), and / or mixed reality (MR). Electronic devices can provide extended reality by synthesizing virtual objects and / or information onto a real-world environment. For example, electronic devices can provide extended reality by displaying graphic objects on a transparent display. For example, electronic devices can provide extended reality by displaying graphic objects along with images of the real environment on an opaque display (e.g., via video-see-through, VST methods).
[0003] In augmented reality, a user of an electronic device may be conversing with another person. For example, the other person may be a person located in real space or a graphic image corresponding to the other person placed within the augmented space. The content of the conversation can be used to learn human behavior. For example, a large language model (LM) can be trained based on the content of the conversation. Reinforcement learning from human feedback (RLHF) can be used to learn human behavior. For example, human feedback can be converted into binary values or scores and used for reinforcement learning. For example, a reward based on human feedback can be set, and reinforcement learning based on the reward can be performed.
[0004] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art in relation to the present disclosure.
[0005] An electronic device according to one embodiment disclosed herein may include at least one camera, at least one microphone, a memory, and at least one processor. The at least one processor may be communicatively connected to the at least one camera, at least one microphone, and the memory, and may include at least one processing circuit. The memory may store instructions that, when executed individually or in combination by the at least one processor, cause the electronic device to obtain a first utterance by a user of the electronic device using the at least one microphone. The memory may store instructions that, when executed individually or in combination by the at least one processor, cause the electronic device to obtain a response to the first utterance from a conversation partner using at least one of the at least one camera or the at least one microphone. The memory may store instructions that, when executed individually or in combination by the at least one processor, cause the electronic device to generate a prompt including a relationship between the user and the conversation partner based on the response. The memory may store instructions that, when executed individually or in combination by the at least one processor, cause the electronic device to obtain a recommended reaction corresponding to the response based on the prompt and provide the recommended reaction to the user.
[0006] Additionally, a method for providing a recommended reaction of an electronic device according to an embodiment disclosed in this document may include: acquiring a first utterance of a user of the electronic device; acquiring a response from a conversation partner to the first utterance; generating a prompt including a relationship between the user and the conversation partner based on the response; acquiring a recommended reaction corresponding to the response based on the prompt; and providing the recommended reaction to the user.
[0007] A computer-readable storage medium according to one embodiment disclosed in this document may store instructions that cause the electronic device to perform the method for providing the recommended reaction when executed by the processor of the electronic device.
[0008] Figure 1 illustrates an example of augmented reality.
[0009] FIG. 2 illustrates a block diagram of an electronic device according to one embodiment.
[0010] FIG. 3a shows a perspective view of an electronic device according to one embodiment.
[0011] FIG. 3b illustrates an example of internal hardware of an electronic device according to one embodiment.
[0012] FIG. 4a shows a rear perspective view of an electronic device according to one embodiment.
[0013] FIG. 4b shows a front perspective view of an electronic device according to one embodiment.
[0014] FIG. 5a shows a front perspective view of a detachable electronic device according to one embodiment.
[0015] FIG. 5b shows a rear perspective view of a detachable electronic device according to one embodiment.
[0016] FIG. 6 illustrates examples of electronic devices according to an embodiment.
[0017] FIG. 7 illustrates a block diagram of software configurations of an electronic device according to one embodiment.
[0018] FIG. 8a is a flowchart of a conversational assistance triggering method of an electronic device according to one embodiment.
[0019] FIG. 8b is a flowchart of a method for providing a recommended reaction of an electronic device according to one embodiment.
[0020] FIG. 8c is a flowchart of a learning method of an electronic device according to one embodiment.
[0021] FIG. 9 illustrates an example of a recommended reaction of an electronic device according to one embodiment.
[0022] FIG. 10 illustrates a recommended reaction providing environment of an electronic device according to one embodiment.
[0023] FIG. 11 is a flowchart of a method for providing a recommended reaction of an electronic device according to one embodiment.
[0024] FIG. 12 is a block diagram of an exemplary electronic device capable of performing the operations described in this document.
[0025] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0026] Hereinafter, various embodiments of the present invention are described with reference to the accompanying drawings. However, this is not intended to limit the present invention to specific embodiments and should be understood to include various modifications, equivalents, and / or alternatives of the embodiments of the present invention.
[0027] Figure 1 illustrates an example of augmented reality.
[0028] Referring to FIG. 1, according to one embodiment, an electronic device (10) may be configured to provide extended reality (e.g., augmented reality, mixed reality, and / or virtual reality). For example, the electronic device (10) may provide augmented reality and / or mixed reality by displaying a virtual object mapped to the real space where the user (1) is located. For example, the electronic device (10) may provide augmented reality by displaying a virtual object within any virtual space. For example, the electronic device (10) may provide extended reality by configuring at least some of the surrounding spaces of the user (1) as virtual space and the remaining parts as real space.
[0029] In the present disclosure, “actual space” may include physically existing space and physical objects within the physical space. In the present disclosure, “virtual space” may be referred to as a virtual space provided by an electronic device (10). The virtual space and the actual space may be distinguished based on a visual representation provided by the electronic device (10). For example, if the electronic device (10) captures and provides the actual space, the provided space may be referred to as the actual space. For example, if the electronic device (10) visually reconstructs the actual space, the provided space may be referred to as the virtual space.
[0030] In the present disclosure, a “virtual object” may be referred to as a graphic object that does not exist in real space but is displayed by an electronic device (10). The electronic device (10) may display the virtual object in real space and / or virtual space.
[0031] In the example of FIG. 1, the electronic device (10) can display a first virtual object (8a) and a second virtual object (8b) in real space. For example, the electronic device (10) can display the first virtual object (8a) and the second virtual object (8b) mapped to a real object (e.g., a counterpart (2)). In the present disclosure, “mapping display” may include displaying virtual objects by mapping them to a real object or a real location. For example, the electronic device (10) can display the first virtual object (8a) and the second virtual object (8b) by mapping them to the counterpart (2). The electronic device (10) can display the first virtual object (8a) and the second virtual object (8b) at a location adjacent to the counterpart (2) when viewed by the user (1). The electronic device (10) can display the first virtual object (8a) and the second virtual object (8b) such that the first virtual object (8a) and the second virtual object (8b) have a depth similar to that of the opponent (2). For example, the electronic device (10) can set the depth of the first virtual object (8a) and the second virtual object (8b) such that the distance from the user (1) to the first virtual object (8a) and the second virtual object (8b) is similar to the distance from the user (1) to the opponent (2). In one example, the electronic device (10) can display the first virtual object (8a) and the second virtual object (8b) at any location that does not obstruct the view of the opponent (2).
[0032] According to one embodiment, the electronic device (10) may support multiple modalities. For example, the electronic device may be configured to receive and process various types of inputs. For example, 'modality' may refer to a channel for interaction between the electronic device and the user. In this case, voice input and text input through an interface (e.g., a virtual keyboard) may be considered different modalities. For example, 'modality' may refer to the format of data input to an artificial intelligence model (e.g., a generative artificial intelligence model). For example, the user's voice input may be converted into text data through STT (speech to text) and NLU (natural language understanding) and input to the artificial intelligence model. In this case, from the perspective of the artificial intelligence model, voice input and text input through an interface (e.g., a virtual keyboard) may be considered substantially the same modality. Hereinafter, 'modality' may refer to a channel between the user and the electronic device and / or the format of input data to the artificial intelligence model. In the present disclosure, 'modality' may be referred to as 'input type'.
[0033] In one example, the electronic device (10) may be a wearable device. The electronic device (10) may acquire voice input from a user using a microphone. The electronic device (10) may acquire image input using a camera. The electronic device (10) may acquire user input using an interface (e.g., a button and / or a touchpad). The electronic device (10) may acquire user input from an external device (not shown) that is communicably connected. The electronic device (10) may be configured to acquire voice input, image input and / or user input, and to process the acquired input.
[0034] In the example of FIG. 1, the user (1) may be conversing with the counterpart (2). For example, the conversation may include a conversation of a specified format that includes turn-taking. For example, a query (11) from the user (1) may occur. The counterpart (2) may perform a responsive action in response to the query (11). The responsive action may include, for example, a response utterance (21) as well as a response gesture (22). The user (1) may provide a reaction (13) in response to the responsive action. The example of FIG. 1 illustrates an example of a conversation in which turn-taking occurs, and the form of the dialogue of the present disclosure is not limited to the above form.
[0035] According to one embodiment, the electronic device (10) may provide a recommended reaction. For example, the electronic device (10) may provide a recommended reaction based on a query (11), a response utterance (21), and / or a response gesture (22). The electronic device (10) may provide information of the recommended reaction through a virtual image (e.g., a first virtual object (8a) and a second virtual object (8b)) or through an audio signal.
[0036] According to one embodiment, the electronic device (10) can generate a recommended reaction using an artificial intelligence agent associated with a conversation partner (e.g., partner (2)). The electronic device (10) can generate a recommended reaction by considering, for example, the relationship between the user (1) and the partner (2). Through this, the electronic device (10) can generate a personalized recommended reaction for the user (1). Through the personalized recommended reaction, the electronic device (10) can provide a recommended reaction that corresponds to the current situation (e.g., time, place, and relationship with the conversation partner (2)).
[0037] According to one embodiment, the electronic device (10) can analyze conversation using not only verbal expressions but also physical expressions (e.g., gestures and / or facial expressions) and / or biometric data (e.g., heart rate, oxygen saturation and / or pupil response). By using physical expressions and / or biometric data, the electronic device (10) can provide improved recommended reactions.
[0038] According to one embodiment, the electronic device (10) can train a personalized AI agent using verbal expressions, physical expressions, and / or biometric data. Using the personalized AI agent, the electronic device (10) can provide a recommended reaction that corresponds to the current situation (e.g., time, place, and relationship of conversation).
[0039] In the present disclosure, the term “personalized” may include being dedicated to a specific group as well as to a specific individual. The term “personalized” may represent the attributes of a result specialized for a specific individual or a specific group as a result generated using a relationship-based AI model. The relationships in the present disclosure may include relationships between individuals and / or relationships between an individual and a group. The relationships in the present disclosure may include relationships between real people as well as relationships between real people and virtual people. In the following, a relationship-based AI agent or a relationship-based AI model may include an AI agent or an AI model learned based on a specific relationship. The example described above in relation to FIG. 1 is for illustrative purposes only and the embodiments of the present disclosure are not limited thereto. For example, unlike the example in FIG. 1, the counterpart (2) may include a person in a real space, people connected via communication, a group in a real space, and / or a virtual avatar. However, unless otherwise stated, the details described above in relation to FIG. 1 may apply to the embodiments described below. In the following, embodiments of the present disclosure may be described with reference to FIGS. 2 to FIGS. 12.
[0040] FIG. 2 illustrates a block diagram of an electronic device according to one embodiment.
[0041] Referring to FIG. 2, according to one embodiment, the electronic device (10) may include a processor (120), memory (130), sensor circuit (140), display (160), camera (170), interface (180), and / or communication circuit (190). The electronic device (10) may correspond to the electronic device (300) of FIG. 3a and 3b, the electronic device (400) of FIG. 4a and 4b, the electronic device (500) of FIG. 5a and 5b, the smart glasses (600) of FIG. 6, the wearable device (610), and / or the electronic device (1200) of FIG. 12. The electronic device (10) may include a configuration similar to the electronic device (1200) described below in relation to FIG. 12. For example, the processor (120) may correspond to at least one processor (1210) of FIG. 12. For example, the memory (120) may correspond to the memory (1220) of FIG. 12. For example, the sensor circuit (140) may correspond to the sensor interface (1219) and / or sensor (1270) of FIG. 12. For example, the display (160) may correspond to the display (1240) of FIG. 12. For example, the camera (170) may include the image sensor (1250) of FIG. 12. For example, the communication circuit (190) may correspond to the communication circuit (1260) of FIG. 12. The configuration of the electronic device (10) shown in FIG. 2 is exemplary and the configuration of the electronic device (10) is not limited thereto. For example, the electronic device (10) may further include a configuration not shown in FIG. 2 (e.g., at least one of the configurations of the electronic device (1200) of FIG. 12). For example, the electronic device (10) may not include at least one of the configurations shown in FIG. 2.
[0042] The processor (120) may be connected communicatively, electrically, operatively, or functionally to memory (130), sensor circuit (140), display (160), camera (170), interface (180), and / or communication circuit (190). In various embodiments of the present disclosure, when one component is connected “operatively” to another component, it may mean that the component is connected to enable the other component to operate. For example, the component may enable the other component by transmitting a control signal to the other component directly or through another component. In various embodiments of the present disclosure, when one component is connected “functionally” to another component, it may mean that the component is connected to enable the function of the other component. For example, the component may enable the function of the other component by transmitting a control signal to the other component directly or through another component.
[0043] The processor (120) may include at least one processor. The at least one processor may include at least one processing circuit. For example, the processor (120) may include an application processor (AP), a central processing unit (CPU), an image signal processor (ISP), a graphics processing unit (GPU), a neural processing unit (NPU), a tensor processing unit (TPU), and / or a communication processor (CP). The processor (120) may include at least one chip or one chipset. In the present disclosure, the processor (120) may be referred to as a hardware component having an architecture by at least one processing circuit. For example, the processor (120) may be placed on a substrate (e.g., a printed circuit board) located inside the electronic device (10) and may communicate with other components of the electronic device (10) through at least one conductive path formed in the substrate, a flexible printed circuit board (FPCB), a cable, and / or any communication channel (e.g., a wireless communication channel).
[0044] The memory (130) can store instructions. When the instructions are executed by the processor (120), they can cause the electronic device (10) to perform various operations. For example, the instructions can cause the electronic device (10) to perform various operations by being executed individually or collectively by at least one processor. In various embodiments of the present disclosure, the operation of the electronic device (10) may be referred to as an operation performed by the processor (120) by executing instructions stored in the memory (130). The memory (130) may be referred to as a hardware component for storing data.
[0045] The sensor circuit (140) may be configured to detect information associated with the electronic device (10), e.g., context information. The context information may include optical information of the surroundings of the electronic device (10), movement information of the electronic device (10), attachment status information of the electronic device (10), location information of the electronic device (10), and / or information of adjacent objects. For example, the sensor circuit (140) may include an image sensor (141), an inertial sensor (143), a position sensor (145), and / or a biosensor (147). The configuration of the sensor circuit (140) shown in FIG. 2 is exemplary and the embodiments of the present disclosure are not limited thereto.
[0046] The image sensor (141) may be configured to detect optical signals. For example, the image sensor (141) may be configured to detect light signals, infrared signals, and / or ultraviolet signals. The image sensor (141) may be configured to acquire depth information. For example, the image sensor (141) may include a light detection and ranging (LiDAR). For example, the electronic device (10) may use the image sensor (141) to acquire information about a conversation partner.
[0047] The inertial sensor (143) may be configured to detect the movement of the electronic device (10). For example, the inertial sensor (143) may include an inertial measurement unit (IMU) configured to detect the orientation, magnetic force, angular rate, and / or specific force of the electronic device (10). The inertial sensor (143) may include an accelerometer, a gyroscope, and / or a magnetometer. For example, the electronic device (10) may use the inertial sensor (143) to detect the movement of a user wearing the electronic device (10).
[0048] The location sensor (145) may include any sensor configured to acquire the geographical location of the electronic device (10). For example, the location sensor (145) may acquire geographical location information based on a global navigation satellite system (GNSS). For example, the location sensor (145) may acquire geographical location information based on the reception strength and / or angle of arrival of an external signal. In one example, the electronic device (10) may acquire geographical location information based on reception information from a base station.
[0049] The biosensor (147) may include any sensor configured to acquire biometric information of the user of the electronic device (10). For example, the biosensor (147) may be configured to measure the wearer's stress, oxygen saturation, heart rate, and / or blood flow. For example, the biosensor (147) may include an infrared sensor, a photoplethysmography (PPG) sensor, and / or an electrocardiography (ECG) sensor. For example, the electronic device (10) may acquire biometric information using the biosensor (147) and determine the wearer's emotional state based on the acquired biometric information.
[0050] The display (160) may include at least one pixel configured to display an image. In one example, the display (160) may include a plurality of displays. For example, the display (160) may include a left-eye display and a right-eye display. The display (160) may include a front display and / or a rear display. The display (160) may include at least one of a see-through display, a flexible display, a rollable display, a foldable display and / or a rigid display. In one example, the display (160) may include at least one projector for projecting an image.
[0051] The camera (170) may include at least one camera. For example, the camera (170) may include a first camera (171) and a second camera (172). If the electronic device (10) includes a plurality of cameras, each of the plurality of cameras may differ in at least one of the facing direction, magnification, or field of view. For example, the first camera (171) may be configured to acquire an image in a direction corresponding to the wearer's field of view. For example, the second camera (172) may be configured to acquire an image corresponding to at least a part of the wearer's body (e.g., the wearer's body, hands, face, and / or eyes).
[0052] The interface (180) may include at least one device configured to receive input. For example, the interface (180) may include a touch circuit (181) (e.g., a touch screen display) configured to receive touch input. The interface (180) may include at least one microphone (e.g., a first microphone (182) and / or a second microphone (183)) configured to receive voice input. In one example, the electronic device (10) can identify the location of the speaker of the voice input (e.g., the relative direction of the speaker to the electronic device (10)) by performing beamforming using the first microphone (182) and the second microphone (183). The interface (180) may include at least one device for output. For example, the interface (180) may include a haptic module for tactile output, at least one speaker for sound output (e.g., a speaker (184)), and / or an indicator. The interface (180) may include any human interface device (HID). For example, the interface (180) may include a button (185). According to one embodiment, the processor (120) may be configured to receive input using the interface (180) and to process the received input.
[0053] The communication circuit (190) may be configured to perform short-range wireless communication and / or long-range wireless communication. The communication circuit (190) may include a network interface card (NIC). The processor (120) may communicate with other external electronic devices based on wireless communication and / or wired communication, for example, using the communication circuit (190). The processor (120) may communicate with external devices via an internet protocol (IP) network, for example, using the communication circuit (190).
[0054] According to one embodiment, the electronic device (10) can monitor a user's actual conversation in real time and non-invasively using a sensor circuit (140) and / or an interface. For example, the electronic device (10) can identify turn-taking in the conversation by monitoring the conversation. By recognizing turn-taking within the conversation, the electronic device (10) can determine, without explicit input from the user, whether a response was provided as feedback for a specific context. The electronic device (10) can generate learning data based on the recognition of turn-taking.
[0055] According to one embodiment, the electronic device (10) can detect changes in the user's involuntary biosignals (e.g., changes in heart rate, pupil constriction, and / or pupil dilation) using a sensor circuit (140). Based on changes in biosignals, the electronic device (10) can detect immediate emotional changes of the user during conversation. The electronic device (10) can acquire verbal expressions and physical expressions of conversation participants using an interface (180) and a camera (170). In generating learning data, the electronic device (10) can store verbal expressions, physical expressions, and emotions by mapping them to each conversation. The electronic device (10) can generate learning data based on the real-time responses of conversation participants and train an AI agent (e.g., an interface including an AI model) using the generated learning data. Since learning data is generated for each conversation partner, the AI agent can be aligned to a community or an individual.
[0056] According to one embodiment, the electronic device (10) can provide recommended reactions using a trained AI agent. Since the AI agent is trained using not only verbal expressions but also non-verbal expressions (e.g., emotions and / or physical expressions), the electronic device (10) can provide recommended reactions that include verbal expressions and / or non-verbal expressions. Since the AI agent is trained based on feedback from the user of the electronic device (10), the electronic device (10) can provide personalized recommended reactions to the user.
[0057] FIG. 3a shows a perspective view of an electronic device according to one embodiment.
[0058] Referring to FIG. 3a, the electronic device (300) may correspond to an example of the electronic device (10) of FIG. 1. For example, the electronic device (300) may include at least one display (350) (e.g., the display (160) of FIG. 2) and a frame supporting at least one display (350). For example, the electronic device (300) may be worn on a part of a user's body. For example, the electronic device (300) may include a glasses-type device configured to provide augmented reality.
[0059] The frame may be formed as a physical structure that allows the electronic device (300) to be worn on the user's body. The frame may be configured so that when the user wears the electronic device (300), the first display (350-1) and the second display (350-2) can be positioned corresponding to the user's left and right eyes. The frame may support at least one display (350).
[0060] The frame may include a first rim (301) covering at least a portion of a first display (350-1), a second rim (302) covering at least a portion of a second display (350-2), a bridge (303) positioned between the first rim (301) and the second rim (302), a first pad (311) positioned along a portion of the edge of the first rim (301) from one end of the bridge (303), a second pad (312) positioned along a portion of the edge of the second rim (302) from the other end of the bridge (303), a first temple (304) extending from the first rim (301) and fixed to a portion of the wearer's ear, and a second temple (305) extending from the second rim (302) and fixed to a portion of the ear opposite to the ear. The first pad (311) and the second pad (312) may come into contact with a part of the user's nose, and the first temple (304) and the second temple (305) may come into contact with a part of the user's face and a part of the ear. The temples (304, 305) may be rotatably connected to the rim through the hinge units (306, 307) of FIG. 3B. The first temple (304) may be rotatably connected to the first rim (301) through a first hinge unit (306) positioned between the first rim (301) and the first temple (304). The second temple (305) may be rotatably connected to the second rim (302) through a second hinge unit (307) positioned between the second rim (302) and the second temple (305).
[0061] The frame may include a contact area (320) in which at least a portion comes into contact with a part of the user's body when the user wears the electronic device (300). For example, the contact area (320) may include a nose pad (310), a first temple (304), and a second temple (305).
[0062] The electronic device (300) can provide visual information to a user through at least one display (350). For example, at least one display (350) may include a transparent or translucent lens. At least one display (350) may include a first display (350-1) and / or a second display (350-2) spaced apart from the first display (350-1). For example, the first display (350-1) and the second display (350-2) may be positioned at locations corresponding to the user's left and right eyes, respectively.
[0063] FIG. 3b illustrates an example of internal hardware of an electronic device according to one embodiment.
[0064] Referring to FIGS. 3a and 3b, according to one embodiment, an electronic device (300) may include hardware that performs various functions (e.g., hardware described above in relation to the block diagram of FIG. 2). For example, the hardware may include a battery module (370), an antenna module (375), optical devices (382, 384), speakers (392-1, 392-2), microphones (394-1, 394-2, 394-3), a light-emitting module (not shown), and / or a printed circuit board (390). The various hardware may be placed within a frame.
[0065] At least one display (350) (e.g., the display (160) of FIG. 2) may form a display area on a lens to provide a user wearing an electronic device (300) with visual information that is distinct from the visual information, along with the visual information contained in the external light passing through the lens. The lens may be formed based on at least one of a Fresnel lens, a pancake lens, or a multi-channel lens. The display area formed by the at least one display (350) may be formed on the second surface (332) of the first surface (331) and the second surface (332) of the lens. When the user wears the electronic device (300), the external light may be transmitted to the user by being incident on the first surface (331) and transmitted through the second surface (332). As another example, the at least one display (350) may display a virtual reality image to be combined with a real-world image transmitted through the external light. The virtual reality image output from at least one display (350) can be transmitted to the user's eye through one or more hardware included in the electronic device (300) (e.g., optical devices (382, 384), and / or at least one waveguide (333, 334)).
[0066] According to one embodiment, an electronic device (300) may include waveguides (333, 334) that diffract light transmitted from at least one display (350) and relayed by optical devices (382, 384) and transmit it to a user. The waveguides (333, 334) may be formed based on at least one of glass, plastic, or polymer. A nano pattern may be formed on the exterior or at least a portion of the interior of the waveguides (333, 334). The nano pattern may be formed based on a polygonal and / or curved grating structure. Light incident on one end of the waveguides (333, 334) may be propagated to the other end of the waveguides (333, 334) by the nano pattern. Waveguides (333, 334) may include at least one diffractive element (e.g., DOE (diffractive optical element), HOE (holographic optical element)) and at least one reflective element (e.g., a reflective mirror). For example, waveguides (333, 334) may be placed within an electronic device (300) to guide a screen displayed by at least one display (350) to the user's eye. For example, the screen may be transmitted to the user's eye based on total internal reflection (TIR) occurring within the waveguides (333, 334).
[0067] According to one embodiment, microphones (394-1, 394-2, 394-3) of an electronic device (300) (e.g., the first microphone (182) and / or the second microphone (183) of FIG. 2) are positioned on at least a portion of a frame to acquire a sound signal. Although the first microphone (394-1) positioned on the nose pad (310), the second microphone (394-2) positioned on the second rim (302), and the third microphone (394-3) positioned on the first rim (301) are shown in FIG. 3b, the number and position of the microphones (394) are not limited to the embodiment of FIG. 3b. If there are two or more microphones (394) included in the electronic device (300), the electronic device (300) can identify the direction of the sound signal using a plurality of microphones positioned on different portions of the frame.
[0068] In one embodiment, the camera (340) (e.g., the camera (170) of FIG. 2) may include an eye tracking camera (ET CAM) (340-1a, 340-1b), a motion recognition camera (340-2), and / or a shooting camera (340-3). The shooting camera (340-3), the eye tracking camera (340-1a, 340-1b), and the motion recognition camera (340-2) may be positioned at different locations on the frame and may perform different functions. The eye tracking camera (340-1a, 340-1b) may output data representing the gaze of a user wearing the electronic device (300). For example, the electronic device (300) may detect the gaze from an image containing the user's pupils obtained through the eye tracking camera (340-1a, 340-1b). An example in which the eye-tracking cameras (340-1a, 340-1b) are positioned toward both eyes of the user is illustrated in FIG. 3b, but the embodiment is not limited thereto, and the eye-tracking cameras (340-1a, 340-1b) may be positioned alone toward the user's left or right eye.
[0069] A shooting camera (340-3) can capture a real image or background to be matched with a virtual image in order to implement augmented reality or mixed reality content. The shooting camera can capture an image of a specific object located at the position where the user is looking and provide the image to at least one display (350). The at least one display (350) can display a single image in which information regarding a real image or background including the image of the specific object obtained using the shooting camera is superimposed with a virtual image provided through optical devices (382, 384). In one embodiment, the shooting camera may be placed on a bridge (303) positioned between the first rim (301) and the second rim (302).
[0070] The eye tracking cameras (340-1a, 340-1b) can achieve more realistic augmented reality by tracking the gaze of a user wearing the electronic device (300), thereby matching the user's gaze with visual information provided to at least one display (350). For example, when the user looks straight ahead, the electronic device (300) can naturally display environmental information related to the user's front on at least one display (350) at the location where the user is situated. The eye tracking cameras (340-1a, 340-1b) may be configured to capture an image of the user's pupil to determine the user's gaze. For example, the eye tracking cameras (340-1a, 340-1b) may receive a gaze detection light reflected from the user's pupil and track the user's gaze based on the position and movement of the received gaze detection light. In one embodiment, the eye tracking cameras (340-1a, 340-1b) may be positioned at locations corresponding to the user's left and right eyes. For example, the eye-tracking cameras (340-1a, 340-1b) may be positioned within the first rim (301) and / or the second rim (302) to face the direction in which the user wearing the electronic device (300) is located.
[0071] A motion recognition camera (340-2) can provide a specific event on a screen provided on at least one display (350) by recognizing the movement of the user's entire body or part thereof, such as the user's torso, hands, or face. The motion recognition camera (340-2) can recognize the user's motion (e.g., gesture recognition), acquire a signal corresponding to said motion, and provide a display corresponding to said signal on at least one display (350). A processor can perform a designated function based on identifying the signal corresponding to said motion. In one embodiment, the motion recognition camera (340-2) may be placed on the first rim (301) and / or the second rim (302).
[0072] According to one embodiment, an electronic device (300) can analyze an object included in a real-world image collected through a shooting camera (340-1a, 340-1b), combine a virtual object corresponding to an object among the analyzed objects that is the target of augmented reality provision, and display it on at least one display (350). The virtual object may include at least one of text and an image regarding various information related to the object included in the real-world image. The electronic device (300) can analyze the object based on a multi-camera such as a stereo camera. For the object analysis, the electronic device (300) can perform time-of-flight (ToF) and / or simultaneous localization and mapping (SLAM) supported by the multi-camera.
[0073] The camera (340) included in the electronic device (300) is not limited to the eye-tracking camera (340-1a, 340-1b) and / or motion recognition camera (340-2) described above. For example, the electronic device (300) can identify external objects included within the field of view (FoV) by using a shooting camera (340-3) positioned toward the user's field of view (FoV). The identification of external objects by the electronic device (300) can be performed based on a sensor for identifying the distance between the electronic device (300) and the external object, such as a depth sensor and / or a time of flight (ToF) sensor. The camera (340) positioned toward the FoV can support an autofocus function and / or an optical image stabilization (OIS) function. For example, the electronic device (300) may include a camera (340) (e.g., a face tracking camera) positioned toward the face to acquire an image including the face of a user wearing the electronic device (300).
[0074] The battery module (370) can supply power to the electronic components of the electronic device (300). In one embodiment, the battery module (370) may be placed within the first temple (304) and / or the second temple (305). For example, the battery module (370) may be a plurality of battery modules (370). The plurality of battery modules (370) may each be placed in the first temple (304) and the second temple (305). In one embodiment, the battery module (370) may be placed at the end of the first temple (304) and / or the second temple (305).
[0075] The antenna module (375) can transmit a signal or power to the outside of the electronic device (300) or receive a signal or power from the outside. The antenna module (375) can be electrically and / or operatively connected to a communication circuit within the electronic device (300) (e.g., communication circuit (190) of FIG. 2). In one embodiment, the antenna module (375) may be placed within the first temple (304) and / or the second temple (305). For example, the antenna module (375) may be placed close to one side of the first temple (304) and / or the second temple (305).
[0076] Speakers (392-1, 392-2) (e.g., speaker (184) of FIG. 2) can output an acoustic signal to the outside of the electronic device (300). In one embodiment, the speakers (392-1, 392-2) may be placed within a first temple (304) and / or a second temple (305) to be positioned adjacent to the ears of a user wearing the electronic device (300). For example, the electronic device (300) may include a second speaker (392-2) positioned adjacent to the user's left ear within the first temple (304), and a first speaker (392-1) positioned adjacent to the user's right ear within the second temple (305).
[0077] The electronic device (300) may include a printed circuit board (390) on a PCB. The PCB (390) may be included in at least one of a first temple (304) or a second temple (305). The PCB (390) may include an interposer disposed between at least two sub-PCBs. One or more hardware components included in the electronic device (300) may be disposed on the PCB (390). The electronic device (300) may include a flexible PCB (FPCB) for interconnecting the hardware components.
[0078] The electronic device (300) may further include configurations not illustrated in relation to FIGS. 3a and 3b. For example, the electronic device (300) may further include a light source (e.g., an LED (light emitting diode)) that emits light toward a subject (e.g., a user's eye, face, and / or an object outside the FoV) being photographed using the camera (340). The light source may include an LED of infrared wavelength. The light source may be placed in at least one of the frame and hinge units (306, 307).
[0079] For example, the electronic device (300) can identify an external object touching the frame (e.g., a user's fingertip) and / or a gesture performed by said external object by using a touch sensor, a grip sensor, and / or a proximity sensor formed on at least a portion of the surface of the frame.
[0080] For example, the electronic device (300) may include at least one of a gyroscope sensor, a gravity sensor, and / or an accelerometer sensor (e.g., sensor circuit (140) of FIG. 2) for detecting the posture of the electronic device (300) and / or the posture of a body part (e.g., head) of a user wearing the electronic device (300).
[0081] FIG. 4a shows a rear perspective view of an electronic device according to one embodiment. FIG. 4b shows a front perspective view of an electronic device according to one embodiment.
[0082] Referring to FIGS. 4a and 4b, the electronic device (400) may correspond to an example of the electronic device (10) of FIG. 1.
[0083] Referring to FIG. 4a, according to one embodiment, a first surface (410) of an electronic device (400) may have a wearable form on a part of a user's body (e.g., the user's face). Although not illustrated, the electronic device (400) may further include a strap for securing to a part of a user's body and / or one or more temples (e.g., a first temple (304) and / or a second temple (305) of FIG. 3a and FIG. 3b). A first display (450-1) for outputting an image to the user's left eye and a second display (450-2) for outputting an image to the user's right eye may be disposed on the first surface (410). The electronic device (400) may further include a rubber or silicone packing formed on the first surface (410) to prevent interference by light different from light emitted from the first display (450-1) and the second display (450-2) (e.g., ambient light).
[0084] According to one embodiment, the electronic device (400) may include cameras (440-1, 440-2) (e.g., camera (170) of FIG. 2) for photographing and / or tracking both eyes of an adjacent user. The cameras (440-1, 440-2) may be referred to as ET (eye tracking) cameras. According to one embodiment, the electronic device (400) may include cameras (440-3, 440-4) (e.g., camera (170) of FIG. 2) for photographing and / or recognizing the face of a user. The cameras (440-3, 440-4) may be referred to as FT cameras.
[0085] Referring to FIG. 4b, on a second surface (420) opposite to the first surface (410) of FIG. 4a, a camera (e.g., cameras (440-5, 440-6, 440-7, 440-8, 440-9, 440-10)) (e.g., camera (170) of FIG. 2)) and / or a sensor (e.g., depth sensor (430)) may be placed to acquire information related to the external environment of the electronic device (400). For example, cameras (440-5, 440-6, 440-7, 440-8, 440-9, 440-10) may be placed on the second surface (420) to recognize external objects different from the electronic device (400). For example, by using cameras (440-9, 440-10), the electronic device (400) can acquire images and / or media to be transmitted to each of the user's two eyes. Camera (440-9) may be placed on the second surface (420) of the electronic device (400) to acquire an image to be displayed through a second display (450-2) corresponding to the right eye among the two eyes. Camera (440-10) may be placed on the second surface (420) of the electronic device (400) to acquire an image to be displayed through a first display (450-1) corresponding to the left eye among the two eyes.
[0086] According to one embodiment, the electronic device (400) may include a depth sensor (430) (e.g., sensor circuit (140)) disposed on a second surface (420) to identify the distance between the electronic device (400) and an external object. Using the depth sensor (430), the electronic device (400) can obtain spatial information (e.g., depth map) for at least a portion of the FoV of a user wearing the electronic device (400).
[0087] Although not shown, a microphone (e.g., the first microphone (182) and / or the second microphone (183) of FIG. 2) for acquiring sound output from an external object may be placed on the second surface (420) of the electronic device (400). The number of microphones may be one or more depending on the embodiment.
[0088] As described above, according to one embodiment, the electronic device (400) may have a form factor for being worn on a user's head. The electronic device (400) may provide a user experience based on augmented reality, virtual reality, and / or mixed reality while being worn on the head. Using a first display (450-1) and a second display (450-2) positioned toward each of the user's two eyes, the electronic device (400) may display a screen containing an external object. Within the screen, the electronic device (400) may register the external object or display the result of executing a function assigned to the external object.
[0089] FIG. 5a shows a front perspective view of a detachable electronic device according to one embodiment. FIG. 5b shows a rear perspective view of a detachable electronic device according to one embodiment.
[0090] Referring to FIGS. 5a and 5b, according to one embodiment, the electronic device (500) may be configured to be detachably attached to the housing (501). For example, the electronic device (500) may be inserted into a slot (510) of the housing (501) or may be physically coupled to the housing (501) by any structure of the housing (501).
[0091] When combined, the display (550) of the electronic device (500) (e.g., the display (160) of FIG. 2) can be viewed through a left eye lens (551-1) and a right eye lens (551-2). The electronic device (500) can display a left eye image in an area of the display (550) corresponding to the left eye lens (551-1) and display a right eye image in an area of the display (550) corresponding to the right eye lens (551-2).
[0092] Although not illustrated, the housing (501) may further include a strap for securing to a part of the user's body and / or one or more temples (e.g., a first temple (304) and / or a second temple (305) in FIG. 3a and 3b). The housing (501) may further include rubber or silicone packing to prevent interference by ambient light.
[0093] According to one embodiment, the electronic device (500) may include a camera (540) (e.g., camera (170) of FIG. 2) and / or a sensor (e.g., a depth sensor (not shown)) for acquiring information related to the external environment of the electronic device (500). The housing (501) may include a camera hole (541) for the camera (540) to be exposed to the outside when combined with the electronic device (500). For example, the electronic device (500) may acquire spatial information (e.g., a depth map) for at least a portion of the FoV of a user wearing the electronic device (500) using a plurality of cameras.
[0094] Although not shown, the electronic device (500) may include a microphone (e.g., the first microphone (182) and / or the second microphone (183) of FIG. 2) for acquiring sound output from an external object.
[0095] The housing (501) may provide a structure for physical coupling with the electronic device (500) and for user wearing. In one example, the housing (501) may be electrically coupled with the electronic device (500). In this case, the housing (501) may include at least some of the components of the electronic device (10) described above in relation to FIG. 2. The electronic device (500) may receive information obtained by the components of the housing (501) through communication with the housing (501) or output information using the components of the housing (501).
[0096] FIG. 6 illustrates examples of electronic devices according to an embodiment.
[0097] With reference to FIG. 6, the functions of the electronic device (10) described above in relation to FIG. 2 may be implemented through a plurality of electronic devices. In the examples of FIG. 3a through FIG. 5b, the electronic device is described as a head-mounted device (HMD), but embodiments of the present disclosure are not limited thereto. In the example of FIG. 6, a user (601) may use a plurality of electronic devices connected to communicate with each other.
[0098] For example, visual information may be acquired by a camera of a wearable device (610) (e.g., artificial intelligence (AI)), a mobile device (620), and / or smart glasses (600). For example, auditory information may be acquired by a microphone of smart glasses (600), a wearable device (610), a mobile device (620), and / or an audio device (630). For example, biometric information may be acquired by a smart ring (not shown) of smart glasses (600) and / or a smart watch (not shown). For example, visual information may be provided through a display of a device having a display (e.g., smart glasses (600) and / or a mobile device (620)). For example, auditory information may be provided through a device having a speaker (e.g., smart glasses (600), a wearable device (610), a mobile device (620), and / or an audio device (630)).
[0099] The “electronic device” of the present disclosure may be referred to as a device that generates a recommended reaction. For example, a mobile device (620) may acquire information using other devices and generate a recommended reaction based on the acquired information. The mobile device (620) may provide the recommended reaction through smart glasses (600) and / or an audio device (630). Even if one electronic device does not support some of the functions described above in relation to the electronic device (10) of FIG. 2, the device that generates the recommended reaction may be referred to as the electronic device of the present disclosure.
[0100] The examples described above in connection with FIGS. 3a through 6 are examples of forms of electronic devices and are for illustrative purposes only and are not intended to limit the embodiments of the present disclosure. A person skilled in the art will understand that, in addition to the examples described above in connection with FIGS. 3a through 5b, various forms of electronic devices may perform the embodiments of the present disclosure. Furthermore, a person skilled in the art will understand that, as described above in connection with FIG. 6, a plurality of electronic devices may perform the embodiments of the present disclosure through communication.
[0101] FIG. 7 illustrates a block diagram of software configurations of an electronic device according to one embodiment.
[0102] Referring to FIGS. 2 and FIGS. 7, according to one embodiment, an electronic device (10) can train an AI agent based on conversation and provide a recommended reaction using the trained AI agent. The components of the electronic device (10) described in relation to FIG. 7 may be software modules (e.g., threads, functions, databases, and / or programs) implemented by the electronic device (10) executing instructions stored in memory (130) using a processor (120). In one example, some of the components of the electronic device (10) may be implemented by an external device (e.g., a server). For example, some of the AI module (270) may be implemented by an external device. In this case, the electronic device (10) may transmit and receive information by communicating with the external device using a communication circuit (190).
[0103] In the example of FIG. 7, the electronic device (10) may include a sensor data collection module (210), a recognition module (220), an analysis module (240), a reinforcement learning module (250), a conversational assistance module (260), and / or an AI module (270).
[0104] According to one embodiment, the sensor data collection module (210) may be configured to provide data acquired through the sensor circuit (140) and / or interface (180) to another module. For example, data collected by the sensor data collection module (210) may be processed by the recognition module (220). The analysis module (240) may perform an analysis of the conversation using the data processed by the recognition module (220) and / or the sensor data.
[0105] According to one embodiment, the recognition module (220) may include a voice recognition module (721), a speech recognition module (722), a tone analysis module (723), a biosignal analysis module (724), a user input processing module (725), a face recognition module (726), a face movement recognition module (727), a body gesture recognition module (728), an eye tracking module (729), and / or a pupil analysis module (730). The configurations of the illustrated recognition module (220) are exemplary, and any configuration for the recognition of linguistic expressions, physical expressions, and / or biometric data may be included in the recognition module (220). For example, the recognition module (220) may not include at least some of the illustrated configurations.
[0106] The voice recognition module (721) can identify the voice of a conversation partner based on an audio signal obtained using the first microphone (182) and / or the second microphone (183). For example, the voice recognition module (721) can identify the conversation partner by comparing the obtained voice of the partner with a stored voice. For example, if the voice recognition result indicates that the partner corresponds to a stored voice, the electronic device (10) can select an AI agent corresponding to the conversation partner (e.g., a first agent (771), a second agent (772), and / or a third agent (773)). For example, if the voice recognition result does not correspond to a stored voice, the electronic device (10) can create an AI agent corresponding to a new partner or select a stored AI agent that has a similar relationship. Here, if the voice signal of the conversation partner is similar to the stored voice by more than a threshold value, the electronic device (10) can determine that the partner corresponds to the stored voice.
[0107] The speech recognition module (722) can generate linguistic expressions of a conversation as text data based on audio signals obtained using the first microphone (182) and / or the second microphone (183). The speech recognition module (722) can be configured to perform speech to text (STT). For example, the electronic device (10) can analyze relationships, conversations, emotions, and / or agreement using the speech recognition results.
[0108] The tone analysis module (723) can analyze the tone of a conversation based on an audio signal obtained using the first microphone (182) and / or the second microphone (183). The tone analysis module (723) can identify the pitch of the speech of the conversation participants. For example, the electronic device (10) can identify the emotions of the conversation participants based on the tone analysis results.
[0109] The biosignal analysis module (724) can be configured to process data obtained by the biosensor (147). The biosignal analysis module (724) can convert the data received by the biosensor (147) into a form that can be processed by the analysis module (240). For example, the electronic device (10) can identify the emotions of a conversation participant using the biosignal data processed by the biosignal analysis module (724).
[0110] The user input processing module (725) may be configured to process user input received through the interface (180). For example, the electronic device (10) may provide a recommended reaction or identify turn-taking based on the reception of user input. In one example, the user input processing module (725) may process a gesture recognized by the body gesture recognition module (728) as user input. In one example, the user input processing module (725) may be configured to receive user input for an AI agent and perform an action corresponding to the received user input.
[0111] The face recognition module (726) can identify the face of a conversation partner from at least one image captured using the camera (170). In one example, the electronic device (10) can identify the conversation partner based on the face recognition result. The electronic device (10) can identify the conversation partner by comparing the identified face with previously stored face data. In one example, the face recognition module (726) can identify a region corresponding to a face in the image.
[0112] The face movement recognition module (727) can identify the facial expression of a conversation participant from at least one image captured using the camera (170). For example, the face movement recognition module (727) can identify the movement of each body part (e.g., lips, eyes, eyebrows, etc.) within a face region (e.g., a region recognized by the face recognition module (726)) from a plurality of images acquired using the first camera (171).
[0113] The face motion recognition module (727) can identify the facial expression of a conversation partner based on identified movements. For example, the face motion recognition module (727) can identify the movements of body parts within the face region from multiple images acquired using the second camera (172), and identify the user's facial expression based on the identified movements. The face motion recognition module (727) can store the facial expression in a specified format (e.g., a set of blendshape coefficients).
[0114] For example, the electronic device (10) can identify the emotions of a conversation participant based on the identified facial expression. In one example, the electronic device (10) can identify the mouth movements of a conversation participant using a face movement recognition module (727) and identify the speaker of the current speech based on the identified movements.
[0115] The body gesture recognition module (728) can identify physical expressions of conversation participants using at least one image acquired using the camera (170). For example, the body gesture recognition module (728) can identify gestures based on the posture, hand movements, and / or head movements of the conversation participants from the images. The body gesture recognition module (728) can store gestures based on a set of joints constituting the body of the conversation participants. For example, the body gesture recognition module (728) can identify and store gestures based on a set of joints having 6 degrees of freedom, the 3D position of the joints, and the 3D rotation of the joints. For example, the electronic device (10) can identify the emotions of the conversation participants based on physical expressions identified by the body gesture recognition module (728).
[0116] The eye tracking module (729) can identify the eye movements of a conversation participant using at least one image obtained using a camera (170). For example, the eye tracking module (729) can identify the gaze of the conversation partner from images obtained using a first camera (171). For example, the eye tracking module (729) can identify the user's gaze from images obtained using a second camera (172). The eye tracking module (729) can store the identified gazes in the form of a gaze vector. For example, the electronic device (10) can identify the emotion of a conversation participant based on the gaze identified by the eye tracking module (727).
[0117] The pupil analysis module (730) may be configured to identify the pupil response (e.g., dilation and / or constriction of the pupil) of a conversation participant using at least one image acquired using the camera (170). For example, the electronic device (100) may identify the emotion of the conversation participant based on the pupil response identified by the pupil analysis module (730).
[0118] According to one embodiment, the analysis module (240) may be configured to analyze information related to a conversation using data processed by the recognition module (220). For example, the analysis module (240) may include a relationship analysis module (741), a conversation analysis module (742), an emotion analysis module (743), an agreement analysis module (744), a gesture synthesis module (745), a focus analysis module (746), and / or a confidence analysis module (747). The configurations of the illustrated analysis module (240) are exemplary, and any configuration for the analysis of verbal expressions, physical expressions, and / or biometric data may be included in the analysis module (240). For example, the analysis module (240) may not include at least some of the illustrated configurations.
[0119] The relationship analysis module (741) can identify the relationship between the user of the electronic device (10) and the conversation partner. For example, the relationship analysis module (741) can identify the conversation partner based on the recognition results of the voice recognition module (721) and / or the face recognition module (726). The relationship analysis module (741) can identify the relationship between the user and the conversation partner by using the recognized conversation partner relationship information stored in the memory (130). For example, the relationship analysis module (741) can identify the relationship between the user and the conversation partner based on user input. If the user input explicitly indicates the relationship between the user and the conversation partner, the relationship analysis module (741) can identify the relationship between the user and the conversation partner as indicated by the user input. In one example, the user input (e.g., a voice signal obtained using the first microphone (182) and / or the second microphone (183)) may include the user's utterance included in the conversation. In one example, the user's utterance may include an explicit designation of the conversation partner (e.g., mother, father, and / or friend, etc.). The relationship analysis module (741) can identify the relationship between the user and the conversation partner using the explicit designation of the partner by the user.
[0120] In one example, the relationship analysis module (741) can identify the relationship between the user and the conversation partner based on the content of the conversation. The relationship analysis module (741) can identify the relationship based on the content of the conversation recognized by the voice recognition module (722). The relationship analysis module (741) can determine the intimacy between the user and the conversation partner, for example, by analyzing the recognized conversation content or based on the analysis results of the conversation analysis module (742). If the identified intimacy exceeds a threshold, the relationship analysis module (741) can determine that the conversation is a private conversation. If it is a private conversation, the relationship analysis module (741) can identify that the relationship between the user and the conversation partner is a private relationship. If the identified intimacy is below the threshold, the relationship analysis module (741) can determine that the conversation is a public conversation. If it is a public conversation, the relationship analysis module (741) can identify that the relationship between the user and the conversation partner is a public relationship.
[0121] The conversation analysis module (742) can analyze the content of the conversation recognized by the speech recognition module (722). For example, the conversation analysis module (742) can analyze the semantic relationships between words and sentences and identify the type of utterance (e.g., question, response utterance, and / or reaction) based on the analysis results. The conversation analysis module (742) can store text data corresponding to the utterance and map the type of the utterance. The conversation analysis module (742) can perform labeling on the input by mapping and storing the type for the input corresponding to the utterance.
[0122] In one example, the conversation analysis module (742) can generate utterance pair data based on turn-taking. For example, the electronic device (10) can identify turn-taking based on a voice signal. The electronic device (10) can identify interruptions in the voice signal (e.g., the occurrence of a silent interval or the failure to detect voice for more than a specified time). When interruptions in the voice signal are identified, the electronic device (10) can determine whether turn-taking has occurred. For example, the electronic device (10) can determine whether the speaker has changed based on the recognition results of the voiceprint recognition module (721) and / or the recognition results of the face movement recognition module (727). When a change in the speaker is identified, the electronic device (10) can determine that turn-taking has occurred. When turn-taking has occurred, the conversation analysis module (742) can generate pairs between the utterance before the turn-taking and the utterance after the turn-taking. For example, the conversation analysis module (742) can generate question-response behavior pairs and / or response behavior-reaction pairs between labeled (e.g., utterance type mapped) inputs. By forming pairs between data along with labeling, training data for the AI agent can be generated.
[0123] The emotion analysis module (743) can analyze the emotions of a conversation participant using the tone analysis module (723), the biosignal analysis module (724), the facial movement recognition module (727), the body gesture recognition module (728), and / or the pupil analysis module (730). For example, the emotion analysis module (743) can quantify valence and / or arousal. For example, the emotion analysis module (743) can identify the speaker's emotions based on the speaker's tone. For example, the emotion analysis module (743) can identify the user's emotions based on the user's biosignals. For example, the emotion analysis module (743) can identify the speaker's emotions based on body gestures. For example, the emotion analysis module (743) can identify the speaker's emotions based on pupil response. In one example, the emotion analysis module (743) can identify emotions based on at least two of the data described above.
[0124] The agreement analysis module (744) may be configured to quantify the degree of agreement with the other party's statement. For example, the agreement analysis module (744) may analyze the degree of agreement of a conversation participant by using the conversation analysis module (742) and / or the gesture synthesis module (745). For example, a nod may be recognized as affirmation or acceptance of the statement. For example, an utterance containing agreement with the statement may be recognized as affirmation or acceptance of the statement. For example, a gesture of crossing arms or shaking the head may be recognized as negation of the statement. For example, an utterance containing negative emotion regarding the statement may be recognized as negation of the statement.
[0125] The gesture synthesis module (745) can identify a gesture (e.g., physical expression) based on the analysis results of the face movement recognition module (727), the body gesture recognition module (728), and / or the eye tracking module (729). For example, the gesture synthesis module (745) can identify a gesture by integrating a facial expression (e.g., facial expression based on blendshape coefficients), a body gesture (e.g., body gesture based on joints), and / or eye information (e.g., gaze based on gaze vectors). For example, the opponent's facial expression may be a furrowed brow expression, the opponent's body gesture may be a movement of supporting the chin with a hand, and the opponent's gaze may be directed in a direction other than the user. In this case, the gesture synthesis module (745) can determine that the opponent's gesture is a "gesture of thinking about something."
[0126] The concentration analysis module (746) can quantify the concentration level of a conversation participant. For example, the concentration analysis module (746) can analyze the concentration level of a conversation participant using a biosignal analysis module (724), a body gesture recognition module (728), and / or a pupil analysis module (730). For example, if the user's pupils dilate, the concentration analysis module (746) can determine that the user's concentration level has increased. For example, if the user's yawn is detected, the concentration analysis module (746) can determine that the user's concentration level has decreased. For example, the concentration analysis module (746) can identify the user's concentration level based on changes in the user's heart rate.
[0127] The confidence analysis module (747) can quantify the confidence of a conversation participant based on the tone analysis module (723), the biosignal analysis module (724), the facial movement recognition module (727), and / or the eye tracking module (729). For example, the confidence analysis module (747) can identify confidence based on the high and / or low pitch of the voice. For example, the confidence analysis module (747) can identify confidence based on changes in heart rate. For example, the confidence analysis module (747) can identify confidence based on facial expressions. For example, the confidence analysis module (747) can identify confidence based on physical expressions. For example, the confidence analysis module (747) can identify confidence based on eye contact. In one example, the confidence analysis module (747) can identify confidence based on at least two of the information described above.
[0128] According to one embodiment, the reinforcement learning module (250) may be configured to train an AI agent based on a conversation. For example, the reinforcement learning module (250) may train the AI agent based on the analysis results of the analysis module (240). The reinforcement learning module (250) may train the AI agent when the conversation ends, for example, by synthesizing multiple turn-takings. For example, the reinforcement learning module (250) may include a reward engine (751) and / or a relational module (752).
[0129] The reward engine (751) can identify reward scores for training an AI agent based on the analysis results of the analysis module (240). For example, a question-answer behavior pair of a conversation can be mapped to a score of the response behavior to the question. For example, a response behavior-reaction pair of a conversation can be mapped to a score of the reaction to the response behavior. The scores can be set based, for example, on emotion, agreement, focus, and / or confidence. The reward engine (751) can train the AI agent using the scores, for example. The score of the question-answer behavior pair may indicate the suitability of the response behavior to the question. The score of the response behavior-reaction pair may indicate the suitability of the reaction to the response behavior.
[0130] For example, the relationship module (752) can train agents corresponding to relationships (e.g., a first agent (771) and / or a second agent (772)) based on the analysis results of the relationship analysis module (741). The relationship module (752) can fine-tune the AI agents. For example, the relationship module (752) can train separate AI agents based on conversation partners or relationships. For example, each AI agent may correspond to a different persona or relationship. The relationship module (752) can perform scoring on labeled input pairs (e.g., question-answer behavior pairs or response behavior-reaction pairs) in real time during the conversation. The relationship module (752) can train the AI agents using the scored input pairs at any point after the conversation ends.
[0131] According to one embodiment, the conversation assistance module (260) may be configured to provide recommended dialogue, recommended gestures, and / or recommended reactions using an AI agent. For example, the conversation assistance module (260) may generate recommended reactions by inputting the analysis results of the analysis module (240) into a trained AI agent. For example, when a user's query is identified, the conversation assistance module (260) may predict the other party's response behavior using a relationship-based AI agent. The relationship-based AI agent may include, for example, an AI agent configured according to the relationship between the user and the conversation partner (e.g., an AI agent trained to perform role-playing corresponding to the relationship). Based on the user's query, the conversation assistance module (260) may generate a prompt asking for the other party's response and predict the response behavior by inputting the generated prompt into the relationship-based AI agent. Alternatively, the conversation assistance module (260) may observe the response behavior using the analysis module (240). For example, the conversational assistance module (260) can generate a prompt for a response action to generate a line, gesture, and / or reaction that may evoke a specific emotion (e.g., trust, comfort, or emotional assimilation) in the conversation partner. The conversational assistance module (260) can generate a recommended reaction by inputting the generated prompt into a relationship-based AI agent.
[0132] In one embodiment, the conversation assistance module (260) may include a prompt module (761), a service control module (762), and / or a personalization module (764). The configurations of the conversation assistance module (260) illustrated are exemplary and the embodiments of the present disclosure are not limited thereto.
[0133] The prompt module (761) may be configured to generate prompts to be processed by the AI module (270). For example, the prompt module (761) may generate prompts based on the analysis results of the analysis module (240). In one example, the prompt module (761) may include a large language model (LLM) (not shown). The prompt module (761) may generate prompts for providing services. As described below, the prompt module (761) may generate prompts based on the services to be provided. The prompts may include phrases and queries that the AI agent designates to a specific persona, for example.
[0134] The service control module (762) may be configured to provide personalized services to the user of the electronic device (10). The service control module (762) may be configured to handle interactions between the user and the electronic device (10). For example, the service control module (762) may provide specified services based on user settings.
[0135] The service control module (762) may set weights based on the service to be provided. For example, the weights may include weights for at least one of emotion, agreement, focus, and / or confidence. In one example, for the first service, the weight of emotion may be set relatively high, and for the second service, the weight of confidence may be set relatively high. For example, for the first service provided in private conversation, a relatively high weight may be set for emotion. For example, for the second service to assist in public conversation (e.g., business meetings or interviews) or for the personality of an introverted user, a relatively high weight may be set for confidence. For the third service to assist in persuasion or information delivery (e.g., lectures), a relatively high weight may be set for focus and agreement. The weights may be set based on the user's settings (e.g., service designation or weight setting) or based on the conversation context analyzed by the conversation analysis module (742).
[0136] In one example, the prompt module (761) can generate a prompt based on a set weight. In an information delivery situation (e.g., a situation where a high weight is set for focus and consent), the prompt module (761) can generate a prompt asking for a way to increase the other party's focus and consent.
[0137] The personalization module (764) can generate recommended reactions (e.g., dialogue and / or gestures) by inputting the generated prompt to a relationship-based AI agent for the conversation partner. The personalization module (764) can provide the generated recommended reactions through an external electronic device connected via communication using a display (160), a speaker (184), or a communication circuit (190).
[0138] For example, the personalization module (764) may be configured to provide recommended conversations and / or gestures based on vocabulary, expressions, and gestures that the user actually uses frequently when providing conversation assistance features. The personalization module (764) may generate personalized recommended reactions using the prompt module (761). For example, the personalization module (764) may generate prompts using the prompt module (761) to generate recommended reactions that include the user's habits (e.g., frequently used vocabulary, expressions, and / or gestures). For example, the personalization module (764) may generate personalized recommended reactions by adjusting the output of a relationship-based AI agent (e.g., recommended reactions) based on the user's habits.
[0139] The AI module (270) may provide an interface between the AI agent and the conversational assistance module (260). The AI module (270) may include multiple AI agents (e.g., a first agent (771), a second agent (772), and / or a third agent (773)). For example, each of the multiple AI agents may correspond to a different persona or relationship. Each of the multiple AI agents may be configured to generate information using at least one AI model. For example, at least one AI model may include a large language model (LLM), a large multi-modal model (LMM), and / or a large vision model (LVM).
[0140] The AI agent can be personalized, for example, for a specific individual in a specific relationship. For example, the first agent (771) can be personalized for “John.” In this case, the first agent (771) can be dedicated to the relationship between “John” and the user. The first agent (771) can be set as a persona for the user or “John” depending on the prompt designation. The AI agent can be specific to a specific relationship, for example. For example, the second agent (772) can be specific to “friends.” The AI agent can be specific to a specific community or group, for example. For example, the third agent (773) can be specific to a group such as a company or church community.
[0141] With reference to FIGS. 1, 2, and 7 together, a learning method of an AI agent according to one embodiment of the present disclosure may be described. In the example of FIG. 1, a user (1) may have a private conversation with a counterpart (2). For example, the counterpart (2) may be “John,” a friend of the user (1). For example, an electronic device (10) may identify the relationship between the user (1) and the counterpart (2) using a relationship analysis module (741).
[0142] First, an utterance corresponding to a query (11) by the user (1) can be detected. For example, the content of the query (11) may be “What do you think about A?” Based on the end of the query (11), the electronic device (10) can identify turn-taking. In response to the query (11) by the user (1), the counterpart (2) can perform a response action. For example, the response action may include a response utterance (21) and a response gesture (22). The analysis module (240) can quantify emotions, agreement, focus, and / or confidence regarding the response action. Based on the end of the response utterance (21), the electronic device (10) can identify turn-taking. Based on the identification of turn-taking, the conversation analysis module (742) can generate question-response action pairs. The query (11) can be labeled as a query-type utterance, and the response action can be labeled as a response-type. The reward engine (762) can generate a score for a question-answer behavior pair using information quantified by the analysis module (240) (e.g., emotion, agreement, focus, and / or confidence).
[0143] In the example of FIG. 1, the conversation may be terminated according to the reaction (13) of the user (1). After the conversation is terminated, at any point, the electronic device (10) may train an AI agent using a labeled input pair. The electronic device (10) may use a prompt module (761) to generate a prompt that predicts the response behavior of the counterpart (2) corresponding to the query (11) of the input pair. The prompt may include the content of the query (11) and a question that predicts the response of the counterpart (2). For example, the electronic device (10) may input the prompt to a first agent (771) corresponding to the counterpart (2) and obtain the result of the first agent (771). The electronic device (10) may compare the score of the input pair with the result of the first agent (771) to perform fine-tuning or reinforcement learning on the first agent (771).
[0144] For example, if there is no AI agent corresponding to the counterpart (2), the electronic device (10) can select an AI agent corresponding to the relationship between the user (1) and the counterpart (2) (e.g., a second agent (772)) and train the selected AI agent using a labeled input pair.
[0145] With reference to FIGS. 1, 2, and 7 together, a method for providing a recommended reaction using an AI agent according to one embodiment of the present disclosure may be described. In the example of FIG. 1, a user (1) may have a public conversation with a counterpart (2). For example, the counterpart (2) may be in a business meeting with the user (1). For example, an electronic device (10) may identify the relationship between the user (1) and the counterpart (2) using a relationship analysis module (741).
[0146] First, an utterance corresponding to a query (11) by the user (1) can be detected. For example, the content of the query (11) may be “How about 110 USD per unit?” Based on the end of the query (11), the electronic device (10) can identify turn-taking. In response to the query (11) by the user (1), the counterparty (2) can perform a response action. For example, the response action may include a response utterance (21) and a response gesture (22). Based on the end of the response utterance (21), the electronic device (10) can identify turn-taking.
[0147] Based on the termination of the response utterance (21), the electronic device (10) can generate a recommended reaction using an AI agent. For example, the electronic device (10) can select an AI agent corresponding to a public relationship. The electronic device (10) can generate a prompt to be input to the selected AI agent. The prompt may include information about the persona of the counterpart (2) and information requesting a recommended reaction.
[0148] In one example, the prompt module (761) can generate prompts based on weights. For example, the weights can be set based on relationships. In the case of public conversation, a high weight may be set for confidence. In one example, the generated prompt may be: “Persona: business man, Query: You think I look confident when I speak with certain gestures or tone of voice?” The AI agent can generate a result (e.g., a recommended reaction) based on the prompt. The electronic device (10) can enable the AI agent to generate a response based on role-playing by setting a persona for the AI agent. The electronic device (10) can provide the generated result to the user (1). For example, the recommended reaction may be “Look the other person in the eye, nod slowly, and speak in a high tone.” The electronic device (10) may provide a recommended reaction, for example, through a virtual object (e.g., 8a, 8b), through a voice signal, or through a communication-connected external electronic device (not shown). The recommended reaction may be provided by a conversational assistance module (260). For example, a user may perform a reaction (13) by referring to the recommended reaction.
[0149] Hereinafter, conversational assistance methods of an electronic device (10) may be described with reference to FIGS. 8a through 8c. The operations described below in relation to FIGS. 8a through 8c may be referred to as operations performed in the processor (120) of the electronic device (10) of FIG. 2. The order of the operations described below in relation to FIGS. 8a through 8c is an example, and embodiments of the present disclosure are not limited thereto. For example, at least some of the operations may be performed differently from the order of FIGS. 8a through 8c, or may be performed substantially simultaneously with other operations of FIGS. 8a through 8c.
[0150] FIG. 8a is a flowchart of a conversational assistance triggering method of an electronic device according to one embodiment.
[0151] Referring to FIG. 2 and FIG. 8a, according to one embodiment, an electronic device (10) may determine the triggering of a conversational aid. For example, the electronic device (10) may perform the operation of FIG. 8a only when there is consent from the conversational partner (e.g., consent to data analysis and / or the use of sensor data). For example, the electronic device (10) may perform the operation of FIG. 8a after notifying the conversational partner that the conversational aid function is being executed. For example, the electronic device (10) may perform the operation of FIG. 8a based on user input to the electronic device (10).
[0152] According to one embodiment, in operation 805, the electronic device (10) can identify a conversation partner and determine a relationship. For example, the electronic device (10) can recognize a conversation partner using the voice recognition module (721) and / or face recognition module (726) of FIG. 7. For example, the electronic device (10) can identify a conversation partner based on user input (e.g., voice input, gesture input, etc.) indicating a conversation partner. Based on the identification of the conversation partner, the electronic device (10) can identify a relationship between the user and the conversation partner. For example, the electronic device (10) can identify a relationship using the relationship analysis module (741) of FIG. 7. For example, the electronic device (10) can identify a relationship based on a relationship with an identified conversation partner stored in memory (130). If no stored relationship exists, the electronic device (10) can update the relationship with the conversation partner using the user input or the relationship identified by the relationship analysis module (741).
[0153] According to one embodiment, in operation 810, the electronic device (10) can monitor the user and the counterpart. For example, the electronic device (10) can collect data from a camera (170), a first microphone (182), a second microphone (183), and / or a sensor circuit (140), and analyze the collected data. For example, the electronic device (10) can analyze the collected data using the recognition module (220) and analysis module (240) of FIG. 7. The electronic device (10) can analyze verbal expressions and physical expressions.
[0154] According to one embodiment, in operation 815, the electronic device (10) can determine whether turn-taking is detected. For example, the electronic device (10) can determine whether turn-taking has occurred if a silent period is identified in the conversation. Each of the utterances of the user or the counterpart may constitute a turn or sequence. For example, the electronic device (10) can determine whether the speaker has changed based on the recognition results of the voice recognition module (721) of FIG. 7 and / or the recognition results of the face movement recognition module (727). If a change of speaker following a silent period is identified, the electronic device (10) can determine that turn-taking has occurred.
[0155] If turn-taking is detected (e.g., operation 815-YES), the electronic device (10) can perform operations according to reference point A (e.g., operations of FIG. 8b). If turn-taking is not detected (e.g., operation 815-NO), the electronic device (10) can continue to monitor the user and the conversation partner.
[0156] FIG. 8b is a flowchart of a method for providing a recommended reaction of an electronic device according to one embodiment.
[0157] Referring to FIG. 2 and FIG. 8b, according to one embodiment, the electronic device (10) can provide a recommended reaction. For example, the electronic device (10) can provide a recommended reaction when a formatted conversation (e.g., a conversation including turn-taking) is identified.
[0158] According to one embodiment, in operation 820, the electronic device (10) can identify the speaker. For example, the electronic device (10) can identify the speaker before and after the turn-taking occurs based on the voice and facial movements (e.g., mouth shape).
[0159] According to one embodiment, in operation 825, the electronic device (10) can analyze the conversation. When a turn (e.g., utterance or sequence) ends, the electronic device (10) can determine the type of the utterance (e.g., question, response utterance, and / or reaction). For example, the electronic device (10) can determine the type of utterance using the conversation analysis module (742) of FIG. 7. The electronic device (10) can identify the relationship between the current sequence and the previous sequence based on the context of the conversation. For example, if the current sequence corresponds to a response utterance (e.g., response action), the electronic device (10) can identify the sequence among the previous sequences that corresponds to the question utterance and record these sequences in pairs. For example, if the current sequence corresponds to a reaction, the electronic device (10) can identify the sequence among the previous sequences that corresponds to the response utterance and record these sequences in pairs.
[0160] Once the type of sequence (e.g., type of utterance) and the relationship between the sequences are identified, the electronic device (10) can integrate and reinterpret non-verbal expressions (e.g., physical expressions and / or biometric data) with the verbal expressions of the sequences. For example, if a verbal expression carries the meaning of “affirmation” but the physical expression corresponding to the sequence carries “negation,” the value of “agreement” in the sequence may be reduced, or the meaning of the physical expression may be corrected to “affirmation.” For example, in the case of Indians, a gesture corresponding to negation may be similar to a nod. By analyzing the conversation by integrating verbal and physical expressions, interpretation errors due to cultural differences can be reduced. The electronic device (10) can store the sequences by integrating the utterance type, gestures, and / or responses (e.g., analysis results based on verbal and non-verbal expressions). The response may correspond to agreement, confidence, emotion, and / or concentration as described above in relation to the analysis module (240) of FIG. 7, for example.
[0161] According to one embodiment, in operation 830, the electronic device (10) may determine whether to provide conversational assistance. For example, the electronic device (10) may determine whether to provide conversational assistance based on conversation analysis. The electronic device (10) may determine whether to provide conversational assistance based on the response of the sequence. For example, if a change exceeding a threshold occurs among agreement, confidence, emotion, and / or concentration, the electronic device (10) may determine to provide conversational assistance. For example, the electronic device (10) may provide conversational assistance associated with the element where the change exceeding the threshold was detected.
[0162] In one example, the personalization module (764) of FIG. 7 may determine the type of service to be provided based on conversation analysis or identify the type of service selected by the user. In one example, the service to be provided may be set according to the relationship between the user and the conversation partner. Based on the type of service to be provided, the electronic device (10) may determine the provision of conversation assistance. The electronic device (10) may identify a threshold using weights for responses (e.g., agreement, confidence, emotion, and / or focus) set for the type of service to be provided. For example, if the type of service to be provided is for confidence, the electronic device (10) may set a relatively low threshold for confidence or monitor the sequence only for confidence.
[0163] If conversation assistance is not provided (e.g., operation 830-NO), the electronic device (10) can monitor the user and the other party according to operation 810. If it is decided to provide conversation assistance (e.g., operation 830-YES), the electronic device (10) can perform operation 835.
[0164] According to one embodiment, in operation 835, the electronic device (10) may generate a prompt. For example, the electronic device (10) may generate a prompt using the prompt module (761) of FIG. 7. The generated prompt may include a persona designation and a query that designates a persona of an AI agent (e.g., a relationship-based AI agent) corresponding to a relationship. For example, for a prompt intended to assist a user's psychological state, the persona may be designated as the user. For example, for a prompt intended to elicit a response from a counterpart, the persona may be designated as the counterpart's identifier or relationship. The designation of the persona may be performed based on a service to be provided (e.g., a service requested by the user or identified based on a relationship). The query may include text containing a query for an action to elicit a change in a specific response.
[0165] According to one embodiment, in operation 840, the electronic device (10) can generate a prompt-based recommendation reaction. For example, the electronic device (10) can generate a recommendation reaction based on the prompt-based output of an AI agent.
[0166] According to one embodiment, in operation 845, the electronic device (10) may provide a recommended reaction. For example, the electronic device (10) may provide the recommended reaction through an external electronic device connected via communication using a display (160), a speaker (184), or a communication circuit (190).
[0167] After providing a recommended reaction, the electronic device (10) can perform actions according to reference point B (e.g., actions of FIG. 8c).
[0168] FIG. 8c is a flowchart of a learning method of an electronic device according to one embodiment.
[0169] Referring to FIG. 2 and FIG. 8c, according to one embodiment, an electronic device (10) can train an AI agent. For example, when the end of a conversation is identified, the electronic device (10) can train an AI agent using training data obtained through the conversation.
[0170] According to one embodiment, in operation 850, the electronic device (10) can determine whether the conversation has ended. For example, the electronic device (10) can determine that the conversation has ended if no utterance is detected for more than a specified time. For example, the electronic device (10) can determine the end of the conversation based on explicit user input.
[0171] If the conversation has not ended (e.g., operation 850-NO), the electronic device (10) can monitor the user and the other party according to operation 810. If the conversation has ended (e.g., operation 850-YES), the electronic device (10) can perform operation 855.
[0172] According to one embodiment, in operation 855, the electronic device (10) can determine a relationship based on conversation. For example, even if the conversation partners are the same, the conversation between the user and the conversation partner may be a private conversation or a public conversation. If it is a public conversation, the content of the conversation may be used to train an AI agent corresponding to the public relationship. If the identified relationship is the same as the AI agent used for the recommendation reaction, the content of the conversation may be used to train the AI agent used for generating the recommendation reaction.
[0173] According to one embodiment, in operation 860, the electronic device (10) can score the response. For example, the electronic device (10) can perform scoring for the conversation (e.g., scoring for input pairs) using the reward engine (751) of 7. The electronic device (10) can determine the score of the response for each input pair based on the verbal and / or non-verbal expressions of each utterance (e.g., sequence). For example, the electronic device (10) can identify the score of the response based on emotion, confidence, agreement, and / or focus. In identifying the score of the response, the electronic device (10) may use weights set based on the relationship or conversational situation (e.g., weights set for emotion, confidence, agreement, and / or focus). For example, the electronic device (10) may apply weights to the elements of each response (e.g., emotion, confidence, agreement, and / or focus) and identify a score based on the average of the elements, the values for some elements, or the sum of the elements (e.g., weighted sum). The electronic device (10) may identify a score based on the amount of variation of the elements of each response between the input and the response in an input-response pair. The score identified for an input pair (e.g., input-response pair) may be used as a reward in the training of an AI agent.
[0174] According to one embodiment, in operation 865, the electronic device (10) can train an AI agent. For example, the electronic device (10) can train an AI agent corresponding to the relationship identified in operation 855. In this case, the electronic device (10) can perform fine-tuning and / or reinforcement learning on the selected AI agent using the score mapped to the input pair.
[0175] FIG. 9 illustrates an example of a recommended reaction of an electronic device according to one embodiment.
[0176] Referring to FIGS. 2 and FIGS. 9, according to one embodiment, the electronic device (10) may provide a recommended reaction. For example, the electronic device (10) may decide to provide conversation assistance (e.g., operation 830-YES in FIG. 8b) as a decrease in the conversation partner's concentration is detected. The electronic device (10) may provide a recommended reaction according to the method described above in relation to FIG. 8b (e.g., operation 845 in FIG. 8b).
[0177] For example, the electronic device (10) may display a recommendation UI (902) (e.g., a virtual graphic object). For example, the recommendation UI (902) may include sensor information (910), problem information (920), recommendation reaction information (930), and / or recommendation gesture information (940).
[0178] Sensor information (910) may indicate information used for monitoring the current conversation. For example, a microphone icon (911) may indicate that a microphone is used for monitoring the conversation, a camera icon (912) may indicate that a camera is used for monitoring the conversation, and a heart rate icon (913) may indicate that biometric information is used for monitoring the conversation. In the example of FIG. 9, an electronic device (10) in the form of an HMD is exemplified, but the electronic device (10) may display sensor information (910) through an external display connected via communication. For example, the electronic device (10) may inform the conversation partner of the type of information being monitored by providing sensor information (910) through an external display.
[0179] The problem information (920) may include feedback information regarding the problem situation. For example, the electronic device (10) may identify the problem situation in the step of determining whether to provide a conversational aid (e.g., operation 830 in FIG. 8b). The electronic device (10) may provide the identified problem information (920) along with a recommended reaction. In the example of FIG. 9, the problem information (920) may indicate a situation where the conversation partner's concentration has decreased.
[0180] The recommended reaction information (930) may include text information of a recommended reaction generated using an AI agent. In the example of FIG. 9, the recommended reaction information (930) may include a guide for resolving a problem situation.
[0181] The gesture information (940) may include visual information corresponding to the gesture suggested by the recommended reaction generated using an AI agent. For example, the electronic device (10) may provide an avatar animation corresponding to the gesture. The electronic device (10) may generate the avatar animation by, for example, generating a facial expression using Blendshape and generating a body gesture using a set of joints. Since the recommended reaction is generated using an AI agent dedicated to the relationship between the user (1) and the conversation partner, the gesture information (940) may be generated based on the user's (1) usual gestures.
[0182] In one example, the electronic device (10) may provide a recommended reaction (901) through an audio output. The recommended reaction (901) may include information corresponding to, for example, problem information (920), recommended reaction information (930), and / or gesture information (940).
[0183] In the example of FIG. 9, the conversation partner (not shown) may be assumed to be a person. The electronic device (10) may provide a recommended reaction corresponding to the relationship between the user (1) and the conversation partner based on an analysis of the conversation partner. For example, a finger snap may be used as a gesture to attract attention. If the relationship is not intimate, the conversation partner may feel uncomfortable due to the finger snap. Since the electronic device (10) generates the recommended reaction by considering the relationship between the user (1) and the conversation partner, it may not provide a finger snap as a recommended reaction for relationships with low intimacy. Since the electronic device (10) uses an AI agent trained using a score based on human feedback, inappropriate gestures, vocabulary, and / or tone may be excluded from the recommended reaction.
[0184] In one example, the electronic device (10) may provide information about the conversation after the conversation has ended. For example, a user (1) may use the electronic device (10) to review the conversation after the conversation has ended. For example, the user (1) may use the electronic device (10) to check the other party's reactions that the user was unaware of during the conversation. The user (1) may ask the AI agent to inform them of parts where the other party felt negative emotions during the conversation. In this case, the AI agent may provide information about parts of the other party's response that showed negative emotions based on the results of the conversation analysis. In one example, the electronic device (10) may provide a preemptive suggestion based on the conversation analysis. For example, the electronic device (10) may provide a preemptive suggestion such as, “Wouldn’t the other party feel better if you had talked about a topic they were more interested in?”
[0185] In the example of FIG. 9, the conversation partner may be a virtual person provided by the AI agent (e.g., a virtual persona that mimics a real person). For example, the persona may be referenced as information for the AI agent to give specific context and / or identity to the output based on the relationship between the user (1) and the conversation partner. For example, the AI agent's persona may be set based on the input of the user (1). For example, the user may set the AI agent's persona through input such as “You are my friend,” “You are my coworker,” or “From now on, you are me, and I am your friend.”
[0186] The electronic device (10) can provide human-like conversation based on a set persona. For example, the electronic device (10) can provide a response to a query from the user (1) that does not cause discomfort to the user (1). In this case, the electronic device (10) can provide an avatar corresponding to the persona (e.g., an animation generated based on Blendshape and Joint). The electronic device (10) can mimic human behavior (e.g., vocabulary, tone, and / or gestures). For example, the electronic device (10) can use hesitation, rhetoric of uncertainty, and / or tone to express a level of confidence. Additionally, the electronic device (10) can provide buffered expressions of affirmation or negation in consideration of the user's (1) reaction. For example, if the AI agent is set to a persona that uses direct speech, the electronic device (10) can provide a response using direct expressions. For example, if the AI agent is set to a persona that uses euphemisms, the electronic device (10) can provide a response using euphemisms. For example, if the AI agent's persona is set to a friend, the electronic device (10) can provide a response using private and informal vocabulary. For example, if the AI agent's persona is set to a coworker, the electronic device (10) can provide a response using public and formal vocabulary.
[0187] In the example of FIG. 9, a one-on-one conversation situation between a user (1) and a conversation partner is illustrated, but the electronic device (10) of the present disclosure may provide a recommended reaction for a one-to-many conversation situation. With reference to FIG. 10, the operation of the electronic device (10) in a one-to-many conversation situation may be described.
[0188] FIG. 10 illustrates a recommended reaction providing environment of an electronic device according to one embodiment.
[0189] Referring to FIGS. 2 and FIGS. 9, according to one embodiment, an electronic device (10) may provide a recommended reaction for a conversation between a user (1) and multiple counterparts. For example, the conversation counterparts may include a first counterpart (1011), a second counterpart (1012), a third counterpart (1013), and / or a fourth counterpart (1014). In the example of FIG. 9, the conversation counterparts are depicted as being in a video conference, but the examples described below may similarly apply to a situation where the user (1) is giving a presentation to multiple audience members.
[0190] In one example, a user (1) may give a presentation to multiple audience members. An electronic device (10) may use a relationship-based AI agent to provide auxiliary information (e.g., recommended reactions) corresponding to the presentation. The user (1) may provide the electronic device (10) with information regarding the purpose of the presentation and / or the type of audience, and request the creation of a script or the modification of a pre-written script. For example, the user's input may inform the electronic device (10) of the purpose of the presentation and the type of audience, such as, “The purpose is to convey the direction of the idea rather than technical details, and colleagues will be attending.” In this case, the electronic device (10) may use vocabulary and gestures corresponding to the type of audience (e.g., public relations) when creating and / or modifying the script. The script created by the electronic device (10) may include gestures to increase focus in public relations, for example, on important parts or parts to be emphasized.
[0191] According to one embodiment, the electronic device (10) may provide guide information (1001) based on the audience's response. For example, the electronic device (10) may aggregate the responses of a first counterpart (1011), a second counterpart (1012), a third counterpart (1013), and / or a fourth counterpart (1014). The electronic device (10) may recognize the sentiment (e.g., emotion, agreement, focus, and / or confidence) of each counterpart's response based on facial expression recognition and / or gesture recognition for the first counterpart (1011), the second counterpart (1012), the third counterpart (1013), and / or the fourth counterpart (1014), and may identify the audience's response based on the sum or average of the responses.
[0192] In identifying the audience's response, the audience's response can be identified based on the relationship between the user (1) and the audience (e.g., characteristics of the audience). For example, if the audience belongs to a group with a positive bias, an AI agent trained based on that relationship can interpret the audience's positive response as a lower level of agreement than usual. For example, if the audience belongs to a group with a negative bias, an AI agent trained based on that relationship can interpret the audience's positive response as a higher level of agreement than usual. Because an AI agent specialized in the relationship is used, the electronic device (10) can provide a guide (1001) that takes into account the organizational culture of the audience.
[0193] For example, in the example of FIG. 10, a user (1) can provide a lecture to a number of students via video conferencing. The electronic device (10) can collect the responses of the number of students and identify when the students' concentration has decreased. Based on identifying when the students' concentration has decreased, the electronic device (10) can provide a guide (1001), such as “Try telling a light joke to liven up the atmosphere.” For example, the electronic device (10) can proactively provide the guide (1001) based on the listener's response, even without user input.
[0194] FIG. 11 is a flowchart of a method for providing a recommended reaction of an electronic device according to one embodiment.
[0195] In the following, conversational assistance methods of an electronic device (10) may be described with reference to FIG. 2 and FIG. 11. An operation described below in relation to FIG. 11 may be referred to as an operation performed in the processor (120) of the electronic device (10) of FIG. 2. The order of the operations described below in relation to FIG. 11 is an example, and embodiments of the present disclosure are not limited thereto. For example, at least some of the operations may be performed differently from the order of FIG. 11 or may be performed substantially simultaneously with other operations of FIG. 11.
[0196] For example, the electronic device (10) may be a head mount device (HMD) (e.g., the electronic device (300) of FIGS. 3a and 3b, the electronic device (400) of FIGS. 4a and 4b, and / or the electronic device (500) of FIGS. 5a and 5b). The electronic device (10) of the present disclosure is not limited to an HMD, and as described above in relation to FIG. 6, the electronic device (10) may include any device configured to provide a recommended reaction.
[0197] According to one embodiment, in operation 1105, the electronic device (10) can acquire a first utterance of the user. For example, the electronic device (10) can acquire the first utterance of the user using at least one microphone (e.g., a first microphone (182) and / or a second microphone (183)). For example, the electronic device (10) can acquire the first utterance by detecting turn-taking (e.g., operation 815 of FIG. 8a) during monitoring of the conversation (e.g., operation 810 of FIG. 8a). For example, the electronic device (10) can identify that the type of the first utterance corresponds to a query through speaker identification (e.g., operation 820 of FIG. 8b) and conversation analysis (e.g., operation 825 of FIG. 8b). For example, the electronic device (10) can analyze the conversation using the analysis module (240) of FIG. 7.
[0198] According to one embodiment, in operation 1110, the electronic device (10) can obtain a response from the other party. For example, the electronic device (10) can obtain a response from the conversation partner using at least one microphone and / or at least one camera (e.g., camera (170)). For example, the response from the conversation partner may include a verbal response based on auditory information and / or a non-verbal response based on visual information. For example, the electronic device (10) can analyze the response using the analysis module (240) of FIG. 7. Through conversation analysis, the electronic device (10) can identify turn-taking between the user and the conversation partner and perform operation 1110 based on the identification of turn-taking.
[0199] According to one embodiment, the electronic device (10) may determine whether to provide conversation assistance based on the response. For example, the electronic device (10) may analyze the conversation according to operation 825 of FIG. 8b (e.g., using the analysis module (240) of FIG. 7) and determine whether to provide conversation assistance based on the analysis result (e.g., operation 830 of FIG. 8b). For example, the electronic device (10) may identify a score based on the sentiment of the conversation partner (e.g., confidence, agreement, focus, and / or emotion) using the response. For example, the electronic device (10) may identify the score using the analysis module (240) of FIG. 7. The electronic device (10) may identify multiple values corresponding to emotion, agreement, focus, and confidence from the response and identify the score based on the identified multiple values. In one example, the electronic device (10) may identify the score using weights. For example, the electronic device (10) may apply weights to multiple values corresponding to emotion, agreement, focus, and confidence, and identify a score using the multiple values to which the weights are applied. In one example, the weights may be values set based on the relationship between the user and the conversation partner. For example, the electronic device (10) may set the weights as described above in relation to the personalization module (764) of FIG. 7.
[0200] For example, the electronic device (10) may determine whether to generate a prompt based on an identified score. For example, the electronic device (10) may determine whether to generate a prompt by comparing the difference between a previously scored score and a score based on a response with a threshold value. For example, the electronic device (10) may refrain from generating a prompt if the difference is below the threshold value. For example, the electronic device (10) may generate a prompt if the difference is above the threshold value.
[0201] According to one embodiment, in operation 1115, the electronic device (10) can generate a prompt specific to the relationship based on the response. For example, the electronic device (10) can generate a prompt including the relationship between the user and the conversation partner based on the response. For example, the electronic device (10) can generate a prompt according to operation 835 of FIG. 8b. For example, the prompt may include a persona setting (e.g., Persona: business man) and a reaction query (e.g., Do you think I look confident when I speak with a certain gesture or tone?). The persona setting may include information for setting the role of the AI model according to the conversation partner or relationship. The reaction query may include a query for the user's reaction based on the persona setting. For example, the electronic device (10) can generate the prompt using the prompt module (761) of FIG. 7.
[0202] According to one embodiment, in operation 1120, the electronic device (10) can obtain a recommended reaction based on a prompt. For example, the electronic device (10) can obtain a recommended reaction by inputting a prompt into an artificial intelligence model specific to the relationship between the user and the conversation partner (e.g., the first agent (771), the second agent (772), or the third agent (773) of FIG. 270). For example, the recommended reaction may be a result generated by the artificial intelligence model based on the prompt.
[0203] For example, the electronic device (10) can identify a conversation partner (e.g., using the recognition module (220) of FIG. 7). The electronic device (10) can identify a conversation partner based on at least one image including the conversation partner and / or the voice of the conversation partner (e.g., voiceprint). If an artificial intelligence model trained in association with the identified conversation partner (e.g., an artificial intelligence model trained using the conversation content of the conversation partner) is identified, the electronic device (10) can select the corresponding artificial intelligence model and perform operation 1125. If an artificial intelligence model trained in association with the identified conversation partner is not identified, the electronic device (10) can select an artificial intelligence model corresponding to the relationship and perform operation 1125. The electronic device (10) can identify the relationship between the user and the conversation partner from the conversation content between the user and the conversation partner (e.g., using the relationship analysis module (741) of FIG. 7) and select an artificial intelligence model corresponding to the identified relationship.
[0204] For example, the electronic device (10) may store multiple artificial intelligence models (e.g., AI agents). The multiple artificial intelligence models may be trained in relation to a specific person or according to a specific relationship. The multiple artificial intelligence models may be classified according to a hierarchical structure. For example, the lowest level may include artificial intelligence models learned using conversations between a user and a specific conversation partner. For example, the higher level may include artificial intelligence models learned using conversations with people who have a specific relationship with the user (e.g., multiple people in that relationship). For example, the specific relationship may include categories that classify family, friends, company colleagues, and / or people in the lower level. For example, the top level may include domains containing multiple categories. For example, the top level may include private relationships and public relationships. The electronic device (10) may perform a hierarchical search when selecting an artificial intelligence model. For example, the electronic device (10) may search for an artificial intelligence model corresponding to an identified conversation partner in the lower level. If an AI model corresponding to a conversation partner exists in a lower layer, the electronic device (10) can generate a recommended reaction using the AI model. If an AI model corresponding to a conversation partner does not exist in a lower layer, the electronic device (10) can search for an AI model of a category corresponding to a relationship in an upper layer. If an AI model corresponding to a relationship exists, the electronic device (10) can generate a recommended reaction using the AI model of that upper layer. If an AI model corresponding to a relationship does not exist, the electronic device (10) can select an AI model in the top layer. In one example, the AI model in the upper layer can be trained using training data of AI models in the lower layer belonging to the same category.In one example, an upper-level artificial intelligence model can be generated through harmonic learning of lower-level artificial intelligence models belonging to the same category. In one example, a top-level artificial intelligence model can be trained using training data of upper-level and / or lower-level artificial intelligence models belonging to the same domain. In one example, a top-level artificial intelligence model can be generated through harmonic learning of upper-level and / or lower-level artificial intelligence models belonging to the same domain. Through the stratification of artificial intelligence models, the electronic device (10) can select an artificial intelligence model that corresponds to the relationship of the conversation partner.
[0205] According to one embodiment, in operation 1125, the electronic device (10) may provide a recommended reaction (e.g., operation 845 of FIG. 8b). For example, the electronic device (10) may provide a recommended reaction visually using a display (160) (e.g., the recommended UI (902) of FIG. 9). For example, the electronic device (10) may provide a recommended reaction using a speaker (184) (e.g., the recommended reaction (901) of FIG. 9 and / or the guide information (1001) of FIG. 10). In one example, the electronic device (10) may provide the recommended reaction visually and / or audibly through an external electronic device connected via communication using a communication circuit (190). For example, the electronic device (10) may provide the recommended reaction using the conversational assistance module (260) of FIG. 7.
[0206] FIG. 12 is a block diagram of an exemplary electronic device (1200) capable of performing the operations described in this document.
[0207] Referring to FIG. 12, the electronic device (1200) may be one of various forms of electronic devices, such as a notebook (1290), smartphones (1291) having various form factors (e.g., a bar-type smartphone (1291-1), a foldable-type smartphone (1291-2), or a sliderable (or rollable)-type smartphone (1291-3)), a tablet (1292), a cellular phone (not shown), and other similar computing devices (not shown). The components, their relationships, and their functions illustrated in FIG. 12 are illustrative only and are not intended to limit the implementations described or claimed herein. The electronic device (1200) may be referred to as a mobile device, a user device, a multifunction device, a portable device, or a server.
[0208] The electronic device (1200) may include components comprising at least one processor (1210) (hereinafter referred to as processor (1210)), at least one memory (1220) (hereinafter referred to as memory (1220)), at least one display (1240) (hereinafter referred to as display (1240)), at least one image sensor (1250) (hereinafter referred to as image sensor (1250)), at least one communication circuit (1260) (hereinafter referred to as communication circuit (1260)), and / or at least one sensor (1270) (hereinafter referred to as sensor (1270)). The components are merely exemplary. For example, the electronic device (1200) may include other components (e.g., power management integrated circuitry (PMIC), audio processing circuit, antenna, rechargeable battery, or input / output interface). For example, some components may be omitted from the electronic device (1200). For example, some components can be integrated into a single component.
[0209] The processor (1210) may be implemented as one or more IC (integrated circuit (or circuitry)) chips and may perform various data processing operations. The processor (1210) may include at least one electrical circuit and may process instructions (or programs, data, etc.) stored in memory (1220) individually or collectively in a distributed manner. The processor (1210) may include a processor assembly comprising one or more processing circuits. The processor (1210) may include any processing circuit that is operative to control the performance and operation of one or more components of the electronic device (1200) (e.g., memory (1220), display (1240), image sensor (1250), communication circuit (1260), and / or sensor (1270)). For example, a processor (1210) (e.g., an application processor (AP)) may be implemented as a system on chip (SoC) (e.g., a single chip or a chipset). For example, the processor (1210) may be implemented as a plurality of cores (or at least one core circuit), a plurality of chips, or a plurality of chipsets. For example, the processor (1210) may include one or more processing circuits. For example, the processor (1210) may include one or more processing circuits configured to perform the various functions of the present disclosure individually and / or collectively. As an example without limitation, at least a portion of the processor (1210) may be included in a first chip of the electronic device (1200), and at least another portion of the processor (1210) may be included in a second chip of the electronic device (1200) different from the first chip of the electronic device (1200).
[0210] For example, the processor (1210) may include a central processing unit (1211), a graphics processing unit (1212), a neural processing unit (1213), an image signal processor (1214), a display controller (1215), a memory controller (1216), a storage controller (1217), a communication processor (1218), and / or a sensor interface (1219). These components of the processor (1210) are merely exemplary. For example, the processor (1210) may include other components. For example, some components of the processor (1210) may be omitted from the processor (1210). For example, some components of the processor (1210) may be included as separate components of the electronic device (1200) outside of the processor (1210). For example, some components of the processor (1210) (e.g., memory controller (1216)) may be included in other components (e.g., at least part of memory (1220), an interface (e.g. available for connection to at least one component of the electronic device (100)), a display (1240) and / or an image sensor (1250)).
[0211] The processor (1210) may cause other components of the electronic device (1200) to perform various operations by executing instructions stored in memory (1220). The CPU (1211) (or central processing circuit) may be configured to control the components of the processor (1210) based on the execution of instructions stored in memory (1220) (e.g., volatile memory (1221) and / or non-volatile memory (1222)). The GPU (1212) (or graphics processing circuit) may be configured to execute parallel operations (e.g., rendering). The NPU (1213) (or neural processing circuit, or AI (artificial intelligence) chip) may be configured to execute operations for an artificial intelligence model (e.g., convolution computation). An ISP (1214) (or image signal processing circuit) may be configured to process a raw image acquired through an image sensor (1250) into a format suitable for a component within an electronic device (1200) or a component of a processor (1210). A display controller (1215) (or display control circuit, or DPU (display processing unit)) may be configured to process an image acquired from a CPU (1211), GPU (1212), ISP (1214), or memory (1220) (e.g., volatile memory (1221)) into a format suitable for a display (1240). A memory controller (1216) (or memory control circuit) may be configured to control reading data from volatile memory (1221) and writing data to volatile memory (1221). A storage controller (1217) (or storage control circuit) may be configured to control reading data from non-volatile memory (1222) and writing data to non-volatile memory (1222).The CP (1218) (communication processing circuit) may be configured to process data obtained from a component of the processor (1210) into a format suitable for transmitting to another electronic device via the communication circuit (1260), or to process data obtained from another electronic device via the communication circuit (1260) into a format suitable for processing by the component of the processor (1210). For example, the communication circuit (1260) may include one or more communication circuits. The sensor interface (1219) (or sensing data processing circuit, sensor hub) may be configured to process data regarding the state of the electronic device (1200) and / or the state around the electronic device (1200), obtained through the sensor (1270), into a format suitable for the component of the processor (1210).
[0212] Memory (1220) may include one or more storage media (or one or more storage devices). For example, memory (1220) may include a memory assembly comprising one or more storage media. For example, the one or more storage media may include a hard drive, a flash memory, a permanent memory such as ROM (read-only memory) (e.g., non-volatile memory (1222)), a semi-permanent memory such as RAM (random access memory) (e.g., volatile memory (1221)), any other suitable type of storage (or storage assembly), or any combination thereof. Memory (1220) may include a cache memory, which is one or more different types of memory used to temporarily store data for a function or feature of the electronic device (1200). As an example not limited to, the cache memory may be included within the processor (1210). The memory (1220) may be fixedly embedded within the electronic device (1200) or incorporated into one or more suitable types of components (e.g., a SIM (subscriber identity module) card and / or an SD (secure digital) card) that can be repeatedly inserted into and removed from the electronic device (1200).
[0213] For example, memory (1220) may store one or more software applications, such as operating system (or system) software applications, firmware software applications, driver software applications, plugin (e.g., add-in, add-on, and / or applet) software applications, and / or any other suitable software applications. For example, the one or more software applications may include instructions executable by the processor (1210). For example, memory (1220) may store instructions that can be called by an application programming interface (API). For example, memory (1220) may store instructions within a library.
Claims
1. In an electronic device (10), At least one camera (170); At least one microphone (182, 183); Memory (130); and It includes at least one camera, at least one microphone, and at least one processor (120) that is communically connected to the memory and includes at least one processing circuit, and When the above memory is executed individually or collectively by the above at least one processor, the electronic device: A first utterance of a user of the electronic device is obtained using at least one microphone, and Using at least one of the above-mentioned at least one camera or the above-mentioned at least one microphone, a response from a conversation partner to the above-mentioned first utterance is obtained, and Based on the above response, generate a prompt including the relationship between the user and the conversation partner, and Based on the above prompt, obtain a recommended reaction corresponding to the above response, and An electronic device that stores instructions for providing the above-mentioned recommended reaction to the user.
2. In Paragraph 1, When the above instructions are executed individually or in combination by the at least one processor, the electronic device: An electronic device that obtains the recommended reaction by inputting the above prompt into an artificial intelligence model specific to the above relationship.
3. In Paragraph 2, When the above instructions are executed individually or in combination by the at least one processor, the electronic device: Identifying the conversation partner in at least one image including the conversation partner or in the voice of the conversation partner, and If a first artificial intelligence model trained in association with the above-mentioned identified conversation partner is identified, the first artificial intelligence model is selected as an artificial intelligence model specific to the above relationship, and An electronic device that, when an artificial intelligence model trained in association with the identified conversation partner is not identified, identifies the relationship between the user and the conversation partner from the content of the conversation between the user and the conversation partner, and selects an artificial intelligence model corresponding to the identified relationship as an artificial intelligence model specific to the relationship.
4. In Paragraph 2, The above prompt is an electronic device comprising a persona setting for setting the role of an artificial intelligence model specific to the relationship according to the conversation partner or the relationship, and a user's reaction query based on the persona setting.
5. In Paragraph 1, The above response is an electronic device comprising a verbal response identified from auditory information and a non-verbal response identified from visual information.
6. In Paragraph 5, When the above instructions are executed individually or in combination by the at least one processor, the electronic device: Using the above response, identify a score based on the emotion of the conversation partner, and An electronic device that determines whether to generate the prompt based on the score of the conversation partner.
7. In Paragraph 6, When the above instructions are executed individually or in combination by the at least one processor, the electronic device, If the difference between the previously scored score and the score based on the above response exceeds a threshold, the above prompt is generated, and An electronic device that refrains from generating the prompt if the above difference is below the above threshold.
8. In Paragraph 6, When the above instructions are executed individually or in combination by the at least one processor, the electronic device, Identify multiple values corresponding to emotion, agreement, concentration, and confidence from the above response, and An electronic device that identifies the score based on the plurality of values above.
9. In Paragraph 8, When the above instructions are executed individually or in combination by the at least one processor, the electronic device, Applying weights corresponding to the relationship to the above multiple values, An electronic device that identifies the score using a plurality of values to which the above weights are applied.
10. In Paragraph 1, The above electronic device includes a head-mounted device (HMD), and The above user is an electronic device corresponding to the wearer of the above electronic device.
11. In a method for providing a recommended reaction for an electronic device (10), The operation of obtaining a first utterance of the user of the above electronic device; An operation to obtain a response from a conversation partner to the first utterance; An action of generating a prompt including the relationship between the user and the conversation partner based on the above response; An operation to obtain a recommended reaction corresponding to the response based on the above prompt; and A method comprising the action of providing the above-mentioned recommended reaction to the user.
12. In Paragraph 11, The action of obtaining the above-mentioned recommended reaction is, A method comprising the action of obtaining the recommended reaction by inputting the above prompt into an artificial intelligence model specific to the above relationship.
13. In Paragraph 12, An action of identifying the conversation partner in at least one image including the conversation partner or in the voice of the conversation partner; If a first artificial intelligence model trained in association with the identified conversation partner is identified, the operation of selecting the first artificial intelligence model as an artificial intelligence model specific to the relationship; and A method further comprising, when an artificial intelligence model trained in association with the identified conversation partner is not identified, identifying the relationship between the user and the conversation partner from the content of the conversation between the user and the conversation partner, and selecting an artificial intelligence model corresponding to the identified relationship as an artificial intelligence model specific to the relationship.
14. In Paragraph 12, A method comprising the above prompt, a persona setting for setting the role of an artificial intelligence model specific to the above relationship according to the conversation partner or the above relationship, and a reaction query of the user based on the persona setting.
15. In Paragraph 11, A method in which the above response includes a verbal response identified from auditory information and a non-verbal response identified from visual information.